Lightweight rice pest detection method based on GAFNet and electronic equipment

By optimizing feature extraction and target localization using the GAFNet network, the problem of pest detection difficulties caused by the complexity of the paddy field environment was solved, achieving efficient and accurate identification and classification of rice pests while reducing computational resource consumption.

CN120912872AActive Publication Date: 2025-11-07JILIN AGRICULTURAL UNIV

Patent Information

Application Number
CN202511431289.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-11-07
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

Traditional methods for detecting rice pests are difficult to use for accurate and efficient identification in complex and ever-changing rice paddy environments. In particular, under low light conditions, image quality deteriorates, backgrounds are complex, and pests are small and diverse, which can easily lead to missed or false detections.

Method used

A lightweight rice pest detection method based on GAFNet is adopted. By introducing a global attention fusion and spatial pyramid pooling module, a C3 efficient feature selection attention module, an enhanced Ghost detection head, and an enhanced loss function FECIoU, feature extraction and target localization are optimized, thereby improving detection accuracy and efficiency.

Benefits of technology

It significantly improves the accuracy and efficiency of rice pest detection in complex environments, reduces computational resource consumption, enhances the ability to accurately identify minute pests, and ensures the stability of real-time monitoring and early warning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912872A_ABST
    Figure CN120912872A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight rice pest detection method based on GAFNet and electronic equipment. A global attention fusion and spatial pyramid pooling module is designed to capture a global context and aggregate multi-scale features. On this basis, a C3-EFSA is provided, feature representation is optimized through depth separable convolution and a lightweight channel attention mechanism, and therefore the distinguishing ability under the complex background is improved. An enhanced Ghost detection head is constructed, and enhanced Ghost convolution (EGConv), an SE module and a SiLU activation function are integrated, so that redundancy is reduced, and a lightweight structure is further improved. And finally, an enhanced loss function FECIoU specially optimized for a complex pest detection scene is provided, the function is based on CIoU, a numerical stable term and a difficult sample weighting mechanism are introduced, and the positioning robustness of the shielded pests is further optimized. The lightweight target detection model can realize detection of different types of rice insect pests, and is suitable for actual deployment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to a lightweight rice pest detection method based on GAFNet and an electronic device. BACKGROUND

[0002] With the increasing demand for efficient and accurate pest monitoring in agricultural production, traditional manual inspection and old pest detection methods have been difficult to meet the challenges in modern agriculture, especially in rice pest control. There are many types of rice pests, and their living habits are complex, which often leads to farmers' inability to timely detect and handle pests, thereby affecting rice yield and quality. Especially the detection of small pests is difficult, and traditional methods are difficult to achieve accuracy and efficiency, and are prone to missed detection and false detection. These problems not only increase production costs, but also cause great economic losses.

[0003] In recent years, with the development of artificial intelligence technology, especially computer vision and deep learning technology, automated pest detection has gradually become an important tool in agricultural production. Deep learning models, especially object detection algorithms, have been widely used in crop disease and pest monitoring, with significant advantages. Through deep neural networks for pest identification and classification, the accuracy and efficiency of pest detection can be greatly improved, human errors can be reduced, and pest occurrence can be monitored in real time, thereby achieving accurate control. However, the environment in rice fields is complex and variable, including light, weather, background and pest size, etc. These factors pose challenges to the performance of existing deep learning models.

[0004] When applying deep learning technology in rice pest detection, there are still several challenges, mainly including the following aspects: (1) The environment in rice fields is complex and variable, and the instability of light conditions, especially in cloudy or early morning, evening and other low light conditions, leads to a decrease in image quality captured by the camera, thereby affecting recognition accuracy; (2) Rice pests are small in size, frequently change in posture, and the background environment is complex, so traditional deep learning models are easily disturbed by the background, leading to difficulty in recognition and easy misidentification or missed detection; (3) There are many types of pests in rice fields, and the appearance difference between similar pest species is small, which brings great classification difficulty to existing deep learning models, leading to a decrease in classification accuracy. SUMMARY

[0005] The technical solution of the present application to solve the above technical problems is to provide a lightweight rice pest detection method based on GAFNet, comprising the following steps:

[0006] S1, obtaining an image dataset containing multiple types of rice pests, and preprocessing and data enhancement of the image dataset to construct a training set, a validation set and a test set;

[0007] S2, a lightweight target detection model GAFNet is constructed, the GAFNet model takes YOLO11n network as a baseline, and integrates the following modules:

[0008] A global attention fusion and spatial pyramid pooling module is used to replace the SPPF module in the original network to capture global context information and aggregate multi-scale features.

[0009] A C3 efficient feature selection attention module optimizes feature representation through depthwise separable convolution (DWConv) and lightweight channel attention mechanism to improve feature discrimination ability in complex background.

[0010] An enhanced Ghost detection head EGDetect is used to replace the standard detection head in the original network, which integrates enhanced Ghost convolution (EGConv), Squeeze and Excitation (SE) module and SiLU activation function to reduce computational redundancy.

[0011] S3, the GAFNet model is trained using the training set, and an enhanced loss function FECIoU is used in the training process, which introduces a numerical stability term and a difficult sample weighting mechanism based on the CIoU loss function to optimize the positioning robustness of the occluded pests.

[0012] S4, the trained GAFNet model is used to detect pests in input rice images, and the class and location information of the pests are output.

[0013] Further, the execution process of the global attention fusion and spatial pyramid pooling module (GAM-SPP) includes:

[0014] 1x1 convolution is performed on the input feature map to compress the channel;

[0015] The compressed feature map is input into the maximum pooling layer with kernel size 5x5, 9x9 and 13x13 respectively to extract multi-scale features, and the extracted features are spliced with the original input features in the channel dimension;

[0016] 1x1 convolution is performed on the spliced features to fuse multi-scale information and compress the channel number;

[0017] The fused features are sequentially input into the channel attention (CA) submodule and the spatial attention (SA) submodule for weighted processing, the CA submodule adopts SE structure, and the SA submodule obtains spatial mapping map through maximum pooling and average pooling in channel direction, and generates spatial attention weight after convolution and Sigmoid activation.

[0018] Further, the C3 efficient feature selection attention module (C3-EFSA) adopts a three-way parallel branch structure:

[0019] The first branch sequentially includes a 1x1 convolution, a 3x3 depthwise separable convolution, batch normalization (BN), and a SiLU activation function;

[0020] The second branch sequentially includes a 1x1 convolution, a 5x5 depthwise separable convolution, batch normalization (BN), and a SiLU activation function;

[0021] The third branch sequentially includes a 3x3 grouped convolution, batch normalization (BN), and a SiLU activation function;

[0022] The output features of the three branches are spliced in the channel dimension;

[0023] The spliced features are subjected to 1x1 convolution, batch normalization (BN), and SiLU activation function for feature fusion;

[0024] The fused features are applied with an efficient channel attention (ECA) mechanism, which adaptively adjusts the channel attention degree through one-dimensional convolution;

[0025] When the input and output channel numbers are the same, the output of the module is connected with the input in residual connection.

[0026] Further, the execution process of the efficient channel attention (ECA) mechanism includes:

[0027] Global average pooling is performed on the input features to obtain a global response scalar for each channel;

[0028] The obtained scalar vector is subjected to cross-channel interaction using one-dimensional convolution, and the convolution kernel size is k, where k is determined by a channel number adaptive function;

[0029] The output of the one-dimensional convolution is passed through a Sigmoid activation function to generate channel attention weights;

[0030] The generated weights are multiplied by the original input features to complete the re-rating of the channel attention.

[0031] Further, the execution process of the enhanced Ghost convolution (EGConv) module includes:

[0032] A standard convolution is used to generate part of the output feature map;

[0033] The part of the output feature map is fed into a depthwise separable convolution to generate the remaining output channels;

[0034] The feature map generated by the standard convolution is spliced with the feature map generated by the depth separable convolution in the channel dimension to form a complete output feature;

[0035] The spliced feature is input into a Squeeze and Excitation (SE) module for channel attention weighting.

[0036] In the EGConv module, a SiLU activation function is used instead of a ReLU activation function.

[0037] Further, the enhanced Ghost detection head (EGDetect) is realized by replacing all 3x3 and 1x1 standard convolutions in the YOLO11n original detection head with the EGConv module.

[0038] Further, the calculation formula of the enhanced loss function FECIoU is as follows:

[0039] .

[0040] Further, the data enhancement operation includes one or more of random brightness adjustment, motion blur, random rectangular occlusion and salt and pepper noise addition, for simulating the complex environment in the rice field.

[0041] To solve the above technical problems, the present application further provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to realize the steps of the above method.

[0042] Compared with the prior art, the present application has the following advantages:

[0043] (1) The GAFNet network effectively improves the precision and efficiency of rice pest detection by introducing a global attention fusion module (GAM) and a spatial pyramid pooling (SPP). Traditional pest detection methods are difficult to deal with small size, high density or occluded pest targets, while the combination of GAM and SPP modules can enhance the model's perception ability of pest targets of different scales. The GAM module optimizes the fusion of features, so that the model can efficiently extract fine-grained target features when facing complex backgrounds; while the SPP module enhances the model's spatial perception ability of pest targets through multi-scale pooling strategy, avoiding information loss and misidentification. This optimization greatly improves the overall detection accuracy, especially in complex environments, ensuring fast and accurate identification of pest targets in rice fields.

[0044] (2) The GAFNet network introduces a C3-EFSA module, combining deep separable convolution and lightweight channel attention mechanism, enabling the model to more accurately identify tiny targets in complex backgrounds. Traditional pest detection methods are prone to missed or false detections when facing occlusion, complex backgrounds, or tiny targets. The C3-EFSA module, through the introduction of multi-scale receptive field modeling and ECA mechanism, can effectively improve the target recognition ability of the model. Especially in complex environments such as rice fields, the C3-EFSA module can improve the accuracy of pest targets through detailed feature extraction and adaptive weighting, ensuring accurate detection of tiny pests on rice leaves, thereby enhancing the practicality and robustness of the system.

[0045] (3) To reduce computational burden and improve processing speed, the GAFNet network introduces an enhanced Ghost convolution module (EGConv). This module combines the advantages of standard convolution and deep separable convolution, effectively reducing the consumption of computational resources while ensuring high-precision detection. By using lightweight GhostConv modules, EGConv not only optimizes the network structure but also significantly improves processing speed, making the model more advantageous in real-time monitoring and pest warning systems. In addition, the addition of channel attention mechanism further enhances the model's ability to identify key areas while maintaining efficient use of computational resources. This design ensures efficient operation of pest detection in resource-limited situations.

[0046] (4) The FECIoU loss function of the GAFNet network optimizes the limitations of traditional CIoU in handling small targets and occlusions. By introducing a small positive number ϵ to correct the aspect ratio difference term and using a difficult sample weighting mechanism, FECIoU can more accurately handle small targets and occluded pest targets in rice pest detection. Traditional loss functions are prone to numerical instability or large errors when dealing with irregular-shaped or heavily occluded pests. The optimization strategy of FECIoU enables the model to focus on difficult-to-detect samples during training, improving detection accuracy and model stability. This improvement enables GAFNet to effectively handle complex field environments and improve the positioning ability of tiny or occluded pests. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, without creative labor, other drawings can also be obtained from the structures shown in these drawings.

[0048] Figure 1Figure a is an enhanced Curculionidae image, figure b is an enhanced Delphacidae image, figure c is an enhanced Cicadellidae image, figure d is an enhanced Phlaeothripidae image, figure e is an enhanced Cecidomyiidae image, and figure f is an enhanced Crambidae image.

[0049] Figure 2 Figure is a schematic diagram of the GAM-SPP structure principle of the application.

[0050] Figure 3 Figure is a schematic diagram of the C3-EFSA module structure of the application.

[0051] Figure 4 Figure is a schematic diagram of the EGConv module structure of the application. DETAILED DESCRIPTION

[0052] The application provides a lightweight rice pest detection method based on GAFNet and an electronic device, aiming to improve the detection accuracy and efficiency of rice pests, reduce production costs, and ensure high-quality production of rice.

[0053] The lightweight rice pest detection method based on GAFNet proposed by the application will be described in specific embodiments as follows:

[0054] In the technical solution of this embodiment, a lightweight rice pest detection method based on GAFNet includes the following steps:

[0055] S1, obtain an image dataset containing multiple categories of rice pests, and pre-process and data enhance the image dataset to construct a training set, a validation set and a test set;

[0056] S2, construct a lightweight target detection model GAFNet, the GAFNet model takes YOLO11n network as baseline, and integrates the following modules:

[0057] A global attention fusion and spatial pyramid pooling module is used to replace the SPPF module in the original network to capture global context information and aggregate multi-scale features.

[0058] A C3 efficient feature selection attention module optimizes feature representation through depthwise separable convolution (DWConv) and lightweight channel attention mechanism to improve feature discrimination ability in complex background.

[0059] An enhanced Ghost detection head EGDetect is used to replace the standard detection head in the original network, which integrates enhanced Ghost convolution (EGConv), SqueezeandExcitation (SE) module and SiLU activation function to reduce computational redundancy;

[0060] S3, training the GAFNet model using the training set, and using an enhanced loss function FECIoU in the training process, the FECIoU loss function introduces a numerical stability term and a difficult sample weighting mechanism based on the CIoU loss function to optimize the positioning robustness of the occluded pests;

[0061] S4, using the trained GAFNet model to detect pests in the input rice image, and outputting the category and location information of the pests.

[0062] Further, the execution process of the global attention fusion and spatial pyramid pooling module (GAM-SPP) includes:

[0063] 1x1 convolution is performed on the input feature map to compress the channel;

[0064] The compressed feature map is input into the maximum pooling layer with kernel size of 5x5, 9x9 and 13x13 respectively to extract multi-scale features, and the extracted features are spliced with the original input features in the channel dimension;

[0065] 1x1 convolution is performed on the spliced feature to fuse multi-scale information and compress the channel number;

[0066] The fused features are sequentially input into the channel attention (CA) submodule and the spatial attention (SA) submodule for weighted processing, the CA submodule adopts the SE structure, and the SA submodule obtains the spatial mapping graph through the maximum pooling and average pooling in the channel direction, and generates the spatial attention weight after convolution and Sigmoid activation.

[0067] Specifically, the global attention fusion and spatial pyramid pooling module (GAM-SPP) significantly improves the precision and efficiency of rice pest detection by replacing the SPPF module in the YOLO11n model. This module addresses the problem of traditional models being unable to accurately extract key features when faced with small size, high density, and severe occlusion in rice pest detection. By introducing multi-scale pyramid pooling (SPP), GAM-SPP enhances the ability to perceive pest targets of different scales. The channel attention (CA) and spatial attention (SA) mechanisms optimize feature selectivity, helping the model focus on pest areas on rice leaves and filter irrelevant information. This design optimizes feature fusion while improving the accuracy and robustness of pest detection in complex backgrounds by adaptively adjusting the weights of channel and spatial features

[0068] Further, the C3 efficient feature selection attention module (C3-EFSA) adopts a three-parallel-branch structure:

[0069] The first branch includes 1x1 convolution, 3x3 depthwise separable convolution, batch normalization (BN), and SiLU activation function in sequence;

[0070] The second branch includes 1x1 convolution, 5x5 depthwise separable convolution, batch normalization (BN), and SiLU activation function in sequence;

[0071] The third branch includes 3x3 grouped convolution, batch normalization (BN), and SiLU activation function in sequence;

[0072] The output features of the three branches are concatenated in the channel dimension;

[0073] The concatenated features are subjected to 1x1 convolution, batch normalization (BN), and SiLU activation function for feature fusion;

[0074] The fused features are applied with the efficient channel attention (ECA) mechanism, which adaptively adjusts the channel attention degree through one-dimensional convolution;

[0075] When the input and output channel numbers are the same, the output of the module is connected with the input in a residual manner.

[0076] Specifically, the C3 high-efficiency feature selection attention module combines multi-scale receptive field modeling and the ECA mechanism, which can effectively reduce missed detection and false detection, and enhance the expression ability of pest features. The C3-EFSA adopts a three-way parallel branch design, extracts information of different scales through depth separable convolution and grouped convolution, then fuses features through 1x1 convolution, and introduces the ECA mechanism to strengthen the model's attention to significant areas. The ECA mechanism adaptively adjusts the channel attention degree through one-dimensional convolution, avoiding excessive computational overhead and improving the recognition ability of small targets. The reserved residual connection structure ensures the transmission of low-level feature information to enhance training stability and accelerate convergence. This design will significantly improve the rice pest detection accuracy in complex environments, especially in the case of occlusion and complex background.

[0077] Further, the execution process of the high-efficiency channel attention (ECA) mechanism includes:

[0078] Performing global average pooling on the input features to obtain a global response scalar for each channel;

[0079] Using one-dimensional convolution for cross-channel interaction on the obtained scalar vector, with a convolution kernel size of k, where k is determined by a channel number adaptive function;

[0080] Passing the output of the one-dimensional convolution through a Sigmoid activation function to generate channel attention weights;

[0081] Multiplying the generated weights with the original input features to complete the re-rating of channel attention.

[0082] Specifically, the enhanced Ghost convolution module (EGConv) improves the robustness and accuracy of rice pest detection by introducing a lightweight GhostConv module and a channel attention mechanism SE module. EGConv combines standard convolution and depth separable convolution to reduce computational overhead and enhance feature expression. The main branch extracts core features, and the cheap branch forges residual features, thereby optimizing the recognition ability of key areas. The SE module adjusts channel weights through the channel attention mechanism to enhance the model's attention to pest areas. The SiLU activation function is used instead of ReLU to improve the non-linear expression ability, especially in complex environments, which can effectively improve the small target detection accuracy while maintaining the lightweight of the network.

[0083] Further, the execution process of the enhanced Ghost convolution (EGConv) module includes:

[0084] Using one layer of standard convolution to generate partial output feature maps;

[0085] The partial output feature maps are fed into one layer of depth separable convolution to generate the remaining output channels;

[0086] The feature map generated by the standard convolution is spliced with the feature map generated by the depth separable convolution in the channel dimension to form a complete output feature;

[0087] The spliced feature is input into a Squeeze and Excitation (SE) module for channel attention weighting.

[0088] In the EGConv module, a SiLU activation function is used instead of a ReLU activation function.

[0089] Further, the enhanced Ghost detection head (EGDetect) is realized by replacing all 3x3 and 1x1 standard convolutions in the YOLO11n original detection head with EGConv modules.

[0090] Specifically, the enhanced Ghost detection head (EGDetect) replaces the 3x3 and 1x1 standard convolution modules in the YOLO11n original Detect module with a new EGDetect module to improve the model's ability in feature generation, attention fusion, and activation function optimization. EGDetect reduces the computational cost through the lightweight characteristics of the GhostConv module, while improving the recognition ability of key areas of rice pests. While optimizing performance, the network remains lightweight.

[0091] Further, the calculation formula of the enhanced loss function FECIoU is as follows:

[0092] .

[0093] Specifically, the enhanced loss function FECIoU improves the CIoU loss function in the original YOLO11n model by introducing an optimization mechanism that adapts to small targets, occlusions, and irregularly shaped pest detection. FECIoU addresses the numerical instability problem of traditional CIoU when dealing with extremely small targets or extreme proportions (such as elongated pests) in rice pest detection by introducing a small positive number ϵ to correct the aspect ratio difference term. In addition, FECIoU introduces a difficult sample weighting mechanism similar to Focal Loss to enhance the model's fitting ability for small targets, occluded or highly deviated samples, thereby improving the accuracy during training. The loss function adjusts the weight term α to smooth the gradient fluctuations and improve the training stability of the model in complex environments.

[0094] Traditional CIoU loss function As shown in equation (1):

[0095]

[0096] wherein, denotes the square of Euclidean distance between the center points of the predicted and ground truth boxes; denotes the square of the diagonal length of the minimum bounding box; , which is used to measure the aspect ratio difference; is the dynamic weight of the aspect ratio penalty term; is the width of the ground truth box; is the height of the ground truth box; is the width of the predicted box; is the height of the predicted box. However, in the aspect ratio consistency modeling part, the original CIoU uses function to measure the shape difference between the predicted and ground truth boxes. In rice pest detection, many targets are extremely small in size or have extreme proportions (e.g., long and thin pests), which can cause the height of the ground truth box to be close to zero. At this time, the value of function will become extremely large, causing numerical instability or gradient explosion problems. Such instability can seriously affect the training process, especially when dealing with these extreme targets, leading to poor convergence of the model.

[0097] To solve the problem, a small positive number is introduced in the calculation to improve the aspect ratio difference term of CIoU , as shown in equation (2):

[0098]

[0099] In fact, the weight term in equation (1) is prone to sharp fluctuations in the IoU minimum value. To this end, to smooth the gradient changes, define as equation (3),

[0100]

[0101] In addition, to further enhance the model's fitting ability for samples, a difficult sample weighting mechanism similar to FocalLoss is introduced in the loss function. This mechanism uses a simple IoU exponential weighting term to enhance the gradient influence of low-quality predictions, so that the model pays more attention to small targets, occluded or large prediction deviation samples during training. Here, is the focus parameter, which controls the strength of this adjustment. In the experiment, to not excessively disturb the training of normal samples, appropriately strengthen the contribution of difficult samples, set Finally, the FECIoU loss function is expressed as formula (4):

[0102]

[0103] In rice pest detection, FECIoU helps improve the accuracy of difficult-to-detect small pests, significantly improving the model's positioning ability, thereby improving detection accuracy and robustness in complex rice field environments.

[0104] Further, the data augmentation operation includes one or more of random brightness adjustment, motion blur, random rectangular occlusion, and salt and pepper noise addition, for simulating the complex environment of the rice field.

[0105] Through these improvements, the overall network will be optimized, enabling the method to effectively handle small, high-density, and occluded pest targets, enhancing the ability to recognize complex backgrounds and small targets. The optimized network structure not only reduces the computational burden but also improves the focus on pest areas. Through the improved loss function and weighting mechanism, the model exhibits stronger fitting ability when dealing with difficult samples, especially in the detection of small targets, occlusions, and irregularly shaped pests. This method provides strong support for rice pest detection in complex environments.

[0106] Embodiment 2: A lightweight rice pest detection method based on GAFNet, comprising the following steps:

[0107] Step 1, as Figure 1The dataset is shown for model training. The dataset contains 10 categories of pest samples, including Curculionidae, Delphacidae, Cicadellidae, Phlaeothripidae, Cecidomyiidae, Hesperiidae, Crambidae, Chloropidae, Ephydridae and Noctuidae. After de-duplication, cleaning and manual review by experts in the Plant Protection College, 4226 valid samples were finally retained. The dataset is divided into a training set of 3375, a validation set of 424 and a test set of 427 in a ratio of 8:1:1, and the distribution ratio of each pest in the three subsets is approximately the same to support effective training and fair evaluation of the model. To improve the generalization and robustness of the model, four data augmentation methods are implemented on the training set: random brightness, motion blur, random rectangular occlusion and salt and pepper noise. These methods simulate actual disturbances such as complex light changes in rice fields, plant swings and shooting shakes, leaf occlusions and imaging noise, etc. Specifically: random brightness simulates the brightness changes under weak light before sunrise, overexposure at noon and low illumination in the evening; motion blur reproduces the blur and trailing caused by device shaking or wind blowing the rice plants; random rectangular occlusion simulates partial or large-area occlusion of pests by rice leaves, ears, etc.; salt and pepper noise simulates imaging device or environmental noise. Through the superposition of the four enhancement operations, the number of training set samples is expanded to twice the original, reaching 6750, while the validation set and test set remain unchanged to improve the generalization ability and robustness of the model.

[0108] Step 2, the GAM-SPP module mainly consists of SPP branch, CA, SA three parts, as shown in Figure 2As shown in the GAM-SPP module, first, a 1x1 convolution is used to preliminarily compress the input feature map in the channel, and then multi-scale pyramid pooling is performed. Three two-dimensional maximum pooling operations with different receptive fields are introduced in turn, and the pooling kernel sizes of 5x5, 9x9 and 13x13 are used respectively, and stride=1 and corresponding padding are set to ensure that the spatial size of the feature map after pooling remains unchanged. This design can effectively model the multi-scale spatial context information, which helps to capture the significant areas of different size pest targets. Second, feature fusion and channel compression are performed. After the original feature map and the three-way pooled features are spliced in the channel dimension, a 1x1 convolution is used to compress the channel number back to the input size, avoiding channel dimension expansion, reducing parameter and calculation overhead, and at the same time completing the fusion of multi-scale semantic features. Third, the CA guided model adaptively allocates the importance of different channels. The CA uses the SE structure, which is specifically implemented by performing global two-dimensional adaptive average pooling on the fused features, followed by two 1x1 two-dimensional convolutions, and the two convolutions use ReLU and Sigmoid activation respectively. The weight map in the channel dimension is obtained, which is used to weight the channel response of the input feature. In addition, the SA can further model the significance of the spatial position. The SA is specifically implemented by obtaining two spatial mapping maps through maximum pooling and average pooling in the channel direction, splicing them, inputting them into a two-dimensional convolution layer with a 7x7 convolution kernel, and generating a spatial attention map through Sigmoid activation. The feature map is weighted to enable the model to focus more on the potential pest area in the rice leaf.

[0109] Step 3, the C3-EFSA module is designed to model light multi-scale receptive fields and ECA mechanisms as the core, aiming to enhance the model's expression ability for pest features under limited computing resources. As shown in Figure 3 As shown, the main module adopts a three-branch parallel design, each branch capturing different scale and spatial distribution of structural information from the input feature. Among them, two branches use 3x3 and 5x5 depth separable convolutions after a 1x1 two-dimensional convolution to extract local and contextual texture information, balancing accuracy and computational efficiency; the other branch uses 3x3 grouped convolution to reduce computational overhead while introducing cross-channel interaction, followed by batch normalization (BN) and SiLU activation function to enhance non-linear expression ability. The three features are spliced in the channel dimension and fused by a 1x1 two-dimensional convolution, and then BN and SiLU activation functions are also connected to enhance the non-linear expression ability.

[0110] After the fusion of the three features, the C3-EFSA module further introduces the ECA mechanism. Unlike traditional attention mechanisms such as SE and CBAM, ECA is based on a lightweight one-dimensional convolution, avoiding channel dimension compression and full connection operations. This attention module first performs global two-dimensional adaptive average pooling on the fused feature map to obtain a global response vector for each channel. Then, it performs local channel information interaction through one-dimensional convolution with a small convolution kernel (usually 3). Finally, it generates attention coefficients for each channel through the Sigmoid activation function and adjusts the original feature map by weighting. Therefore, the C3-EFSA module with ECA can significantly improve the module's attention to the significant areas of rice pests, suppress the invalid background areas, and enhance the model's discrimination performance for small targets.

[0111] In addition, C3-EFSA retains the residual connection structure. When the input and output channel numbers are equal, the module output is added element by element to the input, thus preserving the underlying feature information and promoting stable gradient propagation in deep networks. This design will further improve the convergence speed during training and the overall stability of the network.

[0112] Step 4, EGConv first uses a layer of standard convolution to perform preliminary feature extraction on the input, as shown in Figure 4 The output channel number of the main branch is a part of the final output channel. Then, this part of the feature is sent to a layer of depth separable convolution to generate the remaining output channels, achieving more rich expression. The outputs of the main branch and the cheap branch are then concatenated in the channel dimension to form the complete output.

[0113] To further enhance the model's attention to key target areas, a lightweight channel attention mechanism SE module is introduced in the EGConv module. SE module extracts the response statistics of each channel by performing global two-dimensional adaptive average pooling on the feature map, and generates channel weights through two layers of 1x1 two-dimensional convolution to re-scale the importance of each channel in the feature map. This mechanism can effectively enhance the significant features related to pest targets while suppressing background noise and irrelevant areas. The attention output will act on all channel features after concatenation, and the truncation operation will ensure that the final output channel number is consistent with the expected value, maintaining full compatibility with the original Detect module.

[0114] In addition, to improve the network's nonlinear expression ability, all ReLU activation functions used in EGConv are replaced with SiLU activation functions. The SiLU activation function is shown in equation (5):

[0115]

[0116] where, is each element of the input feature map, is a sigmoid activation function, and the SiLU activation function is superior to the traditional ReLU in terms of gradient continuity and expression ability, especially in target regression and small target detection tasks, which can bring more stable performance. The advantage of replacing the activation function is that it plays a positive role in optimizing the convergence and precision of the model without changing the size of the network structure.

[0117] EGDetect replaces all 3x3 and 1x1 standard convolution modules in YOLO11n original Detect with new EGConv modules to improve the enhanced features of the model in feature generation, attention fusion ability, and activation function optimization.

[0118] The calculation formula of the enhanced loss function FECIoU is as follows:

[0119] .

[0120] Compared with the existing Faster R-CNN, SSD, RT-DETR and YOLO series models, the GAFNet model exhibits overall advantages. In terms of performance indicators, the accuracy of GAFNet reaches 89.8%, the recall rate is 85.6%, and the mAP is 90.1%, which are all better than the corresponding 86.3%, 81.4%, 88.5% of YOLO11n, and also better than the corresponding 47.1%, 71.4%, 67.7% of Faster R-CNN, the corresponding 78.1%, 58.6%, 67.4% of SSD, the corresponding 86.2%, 80.9%, 83.7% of RT-DETR, the corresponding 85.3%, 77.6%, 85.2% of YOLOv5n, the corresponding 84.9%, 78.1%, 85.5% of YOLOv6n, the corresponding 83.4%, 75.1%, 82.7% of YOLOv7-tiny, the corresponding 85.6%, 78.5%, 85.0% of YOLOv8n, the corresponding 87.0%, 82.4%, 88.7% of YOLOv10n, and the corresponding 85.4%, 80.7%, 86.0% of YOLOv12n. In terms of model efficiency, the GAFNet model only requires 2.45M parameter amount, which is reduced by 5% compared with the 2.58M parameter amount of YOLO11n, and the model calculation amount is only 5.0 GFLOPs, which is reduced by 21% compared with the 6.3 GFLOPs of YOLO11n. The parameter amount and calculation amount are also better than the corresponding 137.10M, 370.2 GFLOPs of Faster R-CNN, the corresponding 23.75M, 60.9 GFLOPs of SSD, the corresponding 32.00M, 103.5 GFLOPs of RT-DETR, the corresponding 2.50M, 7.1 GFLOPs of YOLOv5n, the corresponding 4.23M, 11.8 GFLOPs of YOLOv6n, the corresponding 6.03M, 13.1 GFLOPs of YOLOv7-tiny, the corresponding 3.01M, 8.1 GFLOPs of YOLOv8n, the corresponding 2.70M, 8.2 GFLOPs of YOLOv10n, and the corresponding 2.56M, 6.3 GFLOPs of YOLOv12n. Compared with other models, the GAFNet model proposed in the application not only has obvious improvement in accuracy, recall rate and other aspects, but also is more lightweight.

[0121] Embodiment 3: An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the steps of embodiment 1 above when executing the program.

[0122] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any changes or replacements within the technical scope disclosed by the present application, which can be easily thought by those skilled in the art, should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A lightweight rice pest detection method based on GAFNet, characterized in that, The method comprises the following steps: S1, acquiring an image data set containing multiple categories of rice pests, and preprocessing and data enhancing the image data set to construct a training set, a verification set and a test set; S2, constructing a lightweight target detection model GAFNet, the GAFNet model taking a YOLO11n network as a baseline and integrating the following modules: a global attention fusion and spatial pyramid pooling module, used to replace an SPPF module in the original network to capture global context information and aggregate multi-scale features; a C3 efficient feature selection attention module, which optimizes feature representation through depth separable convolution and lightweight channel attention mechanism to improve feature discrimination ability in complex backgrounds; an enhanced Ghost detection head EGDetect, used to replace a standard detection head in the original network, the EGDetect integrating an enhanced Ghost convolution, a Squeeze and Excitation module and a SiLU activation function to reduce computational redundancy; S3, training the GAFNet model using the training set, and adopting an enhanced loss function FECIoU in the training process, the FECIoU loss function introducing a numerical stability term and a difficult sample weighting mechanism based on a CIoU loss function to optimize the positioning robustness to occluded pests; S4, using the trained GAFNet model to detect pests in an input rice image and output pest category and location information.

2. The method of claim 1, wherein, The execution process of the global attention fusion and spatial pyramid pooling module comprises: performing 1×1 convolution on the input feature map to compress the channels; inputting the compressed feature map into maximum pooling layers with kernel sizes of 5×5, 9×9 and 13×13 respectively to extract multi-scale features, and splicing the extracted features with the original input features in the channel dimension; performing 1×1 convolution on the spliced features to fuse multi-scale information and compress the channel number; inputting the fused features into a channel attention submodule and a spatial attention submodule in sequence for weighted processing, the channel attention submodule adopting an SE structure, and the spatial attention submodule obtaining a spatial mapping map through maximum pooling and average pooling in the channel direction, generating spatial attention weights after convolution and Sigmoid activation.

3. The method of claim 1, wherein, The C3 efficient feature selection attention module adopts a three-parallel-branch structure: the first branch comprises 1×1 convolution, 3×3 depth separable convolution, batch normalization and SiLU activation function in sequence; the second branch comprises 1×1 convolution, 5×5 depth separable convolution, batch normalization and SiLU activation function in sequence; the third branch comprises 3×3 grouped convolution, batch normalization and SiLU activation function in sequence; splicing the output features of the three branches in the channel dimension; performing 1×1 convolution, batch normalization and SiLU activation function on the spliced features for feature fusion; applying an efficient channel attention mechanism to the fused features, the efficient channel attention mechanism adaptively adjusting channel attention through one-dimensional convolution; when the input and output channel numbers are the same, performing residual connection on the output and the input of the module.

4. The method of claim 3, wherein, The execution process of the high-efficiency channel attention mechanism comprises: performing global average pooling on the input features to obtain a global response scalar of each channel; using one-dimensional convolution to perform cross-channel interaction on the obtained scalar vector, and the size of the convolution kernel is k, wherein k is determined by a channel number adaptive function; passing the output of the one-dimensional convolution through a Sigmoid activation function to generate channel attention weights; multiplying the generated weights with the original input features to complete the re-rating of the channel attention.

5. The method of claim 1, wherein, The execution process of the enhanced Ghost detection head EGDetect comprises: using one layer of standard convolution to generate partial output feature maps; feeding the partial output feature maps into one layer of depth separable convolution to generate remaining output channels; splicing the feature maps generated by the standard convolution and the feature maps generated by the depth separable convolution in the channel dimension to form complete output features; inputting the spliced features into a SqueezeandExcitation module to perform channel attention weighting; in the enhanced Ghost convolution, using a SiLU activation function instead of a ReLU activation function.

6. The method of claim 5, wherein, The enhanced Ghost detection head is realized by replacing all 3*3 and 1*1 standard convolutions in the YOLO11n original detection head with the enhanced Ghost convolution of claim 5.

7. The method of claim 1, wherein, The calculation formula of the enhanced loss function FECIoU is: 。 8. The method of claim 1, wherein, The data enhancement operation comprises one or more of random brightness adjustment, motion blur, random rectangular occlusion and salt and pepper noise addition, and is used to simulate the complex environment of a rice field.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the method of any one of claims 1 to 8 when executing the program. The processor implements the steps of the method of any one of claims 1 to 8 when executing the program.

Citation Information

Patent Citations

  • Digital printing defect detection method based on lightweight network

    CN120411118A

  • Systems and methods for detecting bad telematics device installations

    US12367688B1

  • Method for constructing pest detection model

    WO2021203505A1

Cited By

  • Visual anomaly detection method and system for rail transit operation and maintenance scene

    CN121074823A

  • Flame detection method for intelligent operation and maintenance fire-fighting inspection robot

    CN121582548A

  • A flame detection method for an intelligent operation and maintenance fire patrol robot

    CN121582548B