Agricultural pest larva detection method based on YOLOv8 improvement
By improving the YOLOv8 model and combining it with the LHF, QDDL and GLCM modules, the problems of low detection accuracy and poor effect in agricultural pest larvae detection are solved, and efficient and accurate pest larvae detection and classification are achieved.
Patent Information
- Application Number
- CN202510685099.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-26
AI Technical Summary
Existing methods for detecting agricultural pest larvae suffer from low accuracy and poor detection effects, especially in complex backgrounds and with changing lighting, making it difficult to effectively identify tiny targets.
A YOLOv8-LQG detection model is constructed, combining a learnable hue adaptive filter module (LHF), a quadruple downsampling detection layer (QDDL), and a contextual enhancement module (GLCM) of global semantic and local texture features to improve the model's detection accuracy and recognition ability for small targets.
It significantly enhances the target morphological characteristics of pest larvae, improves detection accuracy and recall rate, reduces missed detections and false detections, and achieves efficient and accurate detection of agricultural pest larvae.
Smart Images

Figure CN120707816A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection technology, and in particular to an improved agricultural pest larvae detection method based on YOLOv8. Background Art
[0002] Common agricultural pests, such as insect pests, have wide habitats, rapid migration, robust reproduction, and are challenging to control. Their larvae can devour over 80 varieties of grass crops, including rice, sugarcane, and corn, causing irreversible damage that can lead to yield reductions or even complete crop failure. Therefore, timely and accurate identification and detection of insect larvae is crucial for disaster prevention and control.
[0003] While traditional pest control methods can mitigate pest damage to a certain extent, they often come with high financial investment and potential environmental pollution. With the continuous development of technology and in-depth research, deep learning technology has opened up new possibilities for efficient and environmentally friendly pest monitoring and has demonstrated significant advantages in detecting crop pests. However, complex factors such as lighting variations, background interference, and target occlusion also make it difficult to improve small target detection.
[0004] The purpose of this invention is to provide an improved agricultural pest larvae detection method based on YOLOv8, which solves the problems of low detection accuracy and poor detection effect of existing detection methods, realizes the detection and classification of agricultural pest larvae, and the improved algorithm can realize efficient and accurate agricultural pest larvae detection, providing farmers with pest and disease monitoring and precise prevention and control measures. Summary of the Invention
[0005] This paper proposes an improved agricultural pest larvae detection method based on YOLOv8.
[0006] The present invention solves the above technical problems through the following technical solutions, which include the following steps:
[0007] S1: Constructed the APL8 dataset of 8 common agricultural pest larvae;
[0008] S2: Based on the YOLOv8 model, we built the YOLOv8-LQG detection model by combining the designed Learnable Hue Filter (LHF) module, the Quadruple Down-sampling Detection Layer (QDDL), and the Global-Local Context Module (GLCM) that integrates global semantic and local texture features.
[0009] S3: Train the YOLOv8-LQG model;
[0010] S4: Test the YOLOv8-LQG model;
[0011] Furthermore, in step S1, the APL8 dataset is constructed, and the processing flow is as follows:
[0012] We collected real agricultural pest images from natural scenes, covering eight common agricultural pest larvae: fall armyworm, corn borer, cutworm, beet armyworm, Spodoptera litura, rice leaf roller, white grub, and wireworm. The APL8 (Agricultural Pest Larvae 8) dataset was constructed. The APL8 dataset was divided into training, validation, and test sets, which were used for model training, validation, and testing, respectively. The larvae images were annotated with anchor boxes using the open-source annotation tool LabelImg. The larvae images were labeled as AWL for fall armyworm larvae, CBL for corn borer larvae, GBL for cutworm larvae, BAL for beet armyworm larvae, FAL for Spodoptera litura, RLFL for rice leaf roller larvae, GL for white grub larvae, and WL for wireworm larvae.
[0013] Furthermore, in step S2, a learnable hue filter (LHF) module is designed. Its processing flow is as follows:
[0014] S21: Convert the input RGB image to HSV space and extract the normalized hue (Hue, H). The mathematical expression is as follows:
[0015]
[0016]
[0017]
[0018]
[0019] V represents the brightest channel value in the color, Δ is the difference between the brightest and darkest channels, which is used to calculate saturation, and ϵ is a numerical stability term.
[0020] S22: Based on the extracted hue component, LHF is multiplied by two Sigmoidal layers to generate a bell-shaped curve mask, which replaces the traditional threshold segmentation and allows gradient backpropagation. The center hue c∈[0,1] and hue width w∈[0,0.5] are used as network parameters and are automatically optimized during training. The mathematical expression is as follows:
[0021]
[0022]
[0023] Here, σ is the sigmoid function, initially set to c = 0.5 (90 / 180) and width w = 0.167 (30 / 180) for the yellow-green background region. As more training data is collected, the mask parameters automatically adapt to the data-driven offset direction, thereby suppressing the response intensity of the background hue and enhancing the sensitivity of the target hue.
[0024] S23: Embed the LHF module into the preprocessing stage of YOLOv8, and its output feature X out Directly inputting the data into the backbone network enables the model to quickly suppress background interference in the early stages of training. At the same time, the color-sensitive area is continuously corrected through gradient backpropagation. When the hue distribution of pests deviates from the initial hypothesis, c and w automatically shift in the data-driven direction.
[0025] Furthermore, in step S2, a high-resolution detection head is added, and the processing flow is as follows:
[0026] Building on the original YOLOv8 architecture, this algorithm generates a high-resolution feature map of 160×160×128 through quadruple downsampling, reducing the receptive field of a single pixel to 4×4 pixels. This improves the model's ability to represent small objects, and adds a corresponding small object detection layer to the head. Furthermore, an upsampling module, a concat module, and a C2f module are added to the feature fusion stage. Furthermore, the feature details of these small objects are further transferred along the downsampling path to three other feature layers of different sizes, allowing for the concatenation of shallow and deep feature maps, thereby making better use of global contextual information during detection.
[0027] Furthermore, in step S2, a global-local context module (GLCM) is used to integrate global semantics and local texture features. The processing flow is as follows:
[0028] S21: A Global-Local Context Module (GLCM) that fuses global semantics and local textures is embedded in front of the newly added 160×160×128 detection head. Its core architecture consists of a global branch and a local branch in parallel. By extracting and collaboratively fusing multi-dimensional information of the input features, it enhances the expressive power of key features in the object detection task.
[0029] S22: The global branch first compresses the spatial dimension of the input feature map to 1×1 through global average pooling (GAP), and outputs channel features with a shape of B×C×1×1. Subsequently, two fully connected layers are used to implement channel compression (compression ratio is reduction=16) and dynamic weight adjustment for recovery. The first fully connected layer reduces the number of channels to c1 / 16, compresses the channel dimension, reduces the number of parameters, and extracts more compact global semantic information. After the ReLU activation function, the second fully connected layer is used to restore the original number of channels c1, and finally the Sigmoid activation function is used to generate global semantic weights for feature weighted fusion. Given an input feature map , global channel weight , its mathematical expression is as follows:
[0030]
[0031] S23: At the same time, the local branch performs fine-grained spatial feature extraction on the input feature map through a 5×5 depth-separable convolution, and cooperates with batch normalization and SiLU activation function to enhance local texture representation. Local feature extraction The mathematical expression is as follows:
[0032]
[0033] S24: In the feature fusion stage, a hybrid operation combining channel weighting and pixel-by-pixel addition is used. That is, the channel weights output by the global branch are channel-weighted with the original features to enhance the semantically significant areas. Then, the weights are added pixel-by-pixel with the local features extracted by the local branch to output the final feature Y. This fusion operation performs adaptive enhancement in both the channel and spatial dimensions, achieving complementary optimization of global and local features. Its mathematical expression is as follows:
[0034]
[0035] The GLCM module solves the common problems of insufficient global context perception and loss of local details in complex scenes due to the similarity between pest textures and plant textures by establishing the synergy between channel dependencies and spatial details, thereby improving the detection and recognition capabilities of small targets.
[0036] Furthermore, in step S2, a YOLOv8-LQG model is constructed, and its processing flow is as follows:
[0037] The LHF module is introduced in the image preprocessing stage, enabling the model to quickly suppress the interference of complex backgrounds at the beginning of training and significantly enhance the target morphological features of pest larvae. Subsequently, a higher-resolution detection layer (QDDL) of 160×160×128 obtained by fourfold downsampling is added to the original detection head, reducing the single-pixel receptive field to 4×4 pixels, directly improving the detection accuracy of tiny targets. At the same time, GLCM is embedded in front of this high-resolution detection layer. By establishing inter-channel dependencies and synergizing the capture of spatial detail information, the problems of insufficient global context perception and loss of local details caused by the easy confusion between pest textures and plant textures are solved.
[0038] Furthermore, in step S3, the YOLOv8-LQG model is trained, and the processing flow is as follows:
[0039] The training set images are input into the YOLOv8-LQG model, and the hyperparameters are set for training. The training process continues until the model converges and the optimal model weights are finally obtained.
[0040] Furthermore, in step S5, the YOLOv8-LQG model is tested, and the processing flow is as follows:
[0041] The optimal model weights are loaded and the divided test set images are input. After feature extraction and aggregation, the Detect layer obtains the bounding box coordinates, confidence level, and category of targets that may contain agricultural pest larvae. Non-maximum suppression is then used to remove redundant detection boxes to obtain the final detection results.
[0042] Compared with the prior art, the beneficial effects of the present invention are:
[0043] (1) The LHF module is introduced in the image preprocessing stage, which enables the model to quickly suppress the interference of complex background in the early stage of training and significantly enhance the target morphological characteristics of pest larvae.
[0044] (2) Add a higher resolution detection layer of 160×160×128 obtained by quadruple downsampling on the basis of the original detection head, reducing the single pixel receptive field to 4×4 pixels, directly improving the detection accuracy of small targets
[0045] (3) GLCM is embedded before the 160×160×128 detection layer. By establishing inter-channel dependencies and capturing spatial detail information, the problem of insufficient global context perception and loss of local details caused by the easy confusion between pest texture and plant texture is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of the present invention;
[0047] Figure 2This is a data example diagram of the present invention;
[0048] Figure 3 It is the LHF module structure used in the YOLOv8-LQG network of the present invention;
[0049] Figure 4 This is a comparison chart of the original image and the image after processing by the LHF module in the YOLOv8-LQG network of the present invention;
[0050] Figure 5 This is a three-dimensional grayscale comparison diagram of the original image and the image processed by the LHF module in the YOLOv8-LQG network of the present invention;
[0051] Figure 6 It is the QDDL module used in the YOLOv8-LQG network of the present invention;
[0052] Figure 7 It is the GLCM module used in the YOLOv8-LQG network of the present invention;
[0053] Figure 8 This is the overall structure diagram of the YOLOv8-LQG network of the present invention;
[0054] Figure 9 The detection effect diagram and heat map of the YOLOv8-LQG network of the present invention; DETAILED DESCRIPTION
[0055] The above and other technical features and advantages of the present invention are described in more detail below with reference to the accompanying drawings.
[0056] The present invention provides a technical solution: a method for detecting agricultural pest larvae based on an improved YOLOv8, referring to Figures 1 to 9 , the specific steps include:
[0057] (1) The APL8 dataset of 8 common agricultural pest larvae was constructed.
[0058] (2) Based on the YOLOv8 model, the YOLOv8-LQG detection model was constructed by combining the designed learnable hue filter (LHF) module, the quadruple down-sampling detection layer (QDDL), and the global-local context module (GLCM) that integrates global semantics and local texture features.
[0059] (3) Train the YOLOv8-LQG model.
[0060] (4) Test the YOLOv8-LQG model.
[0061] Constructing the APL8 dataset specifically includes:
[0062] The images of agricultural pest larvae were collected from natural scenes, covering 8 categories of common agricultural pest larvae, including fall armyworm, corn borer, cutworm, beet armyworm, leafworm, rice leaf roller, white grub and wireworm. Figure 2 As shown, the images were resized to 640 × 640 pixels, and a total of 2,929 qualified images were selected. The APL8 dataset was divided into training, validation, and test sets, which were used for model training, validation, and testing, respectively. Data augmentation was used only for the training set samples. The APL8 dataset contains a total of 10,460 images. The larvae images were annotated with anchor boxes using the open-source annotation tool LabelImg. The larvae of Spodoptera frugiperda were labeled as AWL, the larvae of Corn Borer were labeled as CBL, the larvae of Cutworm were labeled as GBL, the larvae of Spodoptera exigua were labeled as BAL, the larvae of Spodoptera litura were labeled as FAL, the larvae of Rice Leaf Roller were labeled as RLFL, the larvae of White Grub were labeled as GL, and the larvae of Wireworm were labeled as WL.
[0063] The designed Learnable Hue Filter (LHF) module includes the following steps:
[0064] In agricultural visual inspection tasks, the hue distribution of background areas such as crop canopies and soil often shows obvious clustering characteristics. By randomly sampling 300 samples from the APL8 dataset for analysis, it was found that the HSV hue of the target larvae and the background area had obvious hue distribution differences. This phenomenon shows that it is feasible to distinguish between targets and backgrounds by hue features, but there are two major key challenges: First, the traditional method based on fixed hue threshold is difficult to adapt to color distortion caused by factors such as complex field lighting changes, background heterogeneity (such as differences in crop maturity, variety diversity or different crop types), and different collection equipment; second, the color differences of leaves of different crop varieties and growth stages will change the background hue distribution, and the suppression interval needs to be dynamically adjusted. To this end, this paper proposes a learnable hue adaptive filtering module. Its core design ideas are as follows: Figure 3 As shown in the figure, end-to-end collaborative optimization of background interference suppression and target feature enhancement is achieved through differentiable HSV color space conversion and dynamic hue mask generation.
[0065] First, the input RGB image is converted to HSV space and the normalized hue is extracted. The purpose is to extract hue information through color space decomposition while maintaining the propagation of gradients to support subsequent hue filtering operations. The mathematical expression is as follows:
[0066]
[0067]
[0068]
[0069]
[0070] V represents the brightest channel value in the color, Δ is the difference between the brightest and darkest channels, which is used to calculate saturation, and ϵ is a numerical stability term.
[0071] Based on the extracted hue component, LHF is multiplied by two Sigmoidal layers to generate a bell-shaped curve mask, which replaces the traditional threshold segmentation and allows gradient backpropagation. The center hue c∈[0,1] and hue width w∈[0,0.5] are used as network parameters and are automatically optimized during training. The mathematical expression is as follows:
[0072]
[0073]
[0074] Here, σ is the sigmoid function, initially set to c = 0.5 (90 / 180) and width w = 0.167 (30 / 180) for the yellow-green background region. As more training data is collected, the mask parameters automatically adapt to the data-driven offset direction, thereby suppressing the response intensity of the background hue and enhancing the sensitivity of the target hue.
[0075] like Figure 4 As shown in the figure, when the target and background have similar hues, the LHF module proposed in this chapter adaptively adjusts the hue center and width during training, suppressing the background areas unrelated to the target and strengthening the feature response of the main target area, thereby achieving the purpose of separating the target area from the background. In the original image, the larvae and plant leaves exhibit similar hue distributions under natural light. However, after processing by the LHF module, the model dynamically suppresses the hue response intensity of the background area, significantly enhancing the features of the main larvae. While preserving the target details, it effectively reduces the interference of the complex background on the detection task, thus providing the subsequent network with an input image in which the target and background are more easily decoupled.
[0076] Figure 5A spatial domain comparative analysis of the grayscale distribution of the original image and the image processed by the LHF module was performed. The top row shows the image, and the bottom row shows the corresponding three-dimensional grayscale distribution map. The horizontal and vertical axes represent the width and height of the image, respectively, and the vertical axis represents the grayscale value. It can be found that the grayscale values of the background area of the original image do not show significant differences. However, the background area of the image processed by the LHF module suppresses the response intensity of the background hue area, resulting in a simultaneous decrease in the RGB channel values in this area. This is reflected in the grayscale map as a compression of the dynamic range of the background brightness value, while the target area maintains its original grayscale distribution characteristics due to the preservation of hue.
[0077] To add a high-resolution detection head, the following steps are included:
[0078] Based on the original structure of YOLOv8, a new quadruple downsampling detection layer (QDDL) is added, such as Figure 6 As shown in the figure, the detection layer generates a high-resolution feature map of 160×160×128 through a four-fold downsampling operation, and a corresponding small object detection layer is added to the detection head. At the same time, an upsampling module, a concat module, and a C2f module are added to the feature fusion stage. In addition, the feature details of these small objects can also be further transferred to the other three feature layers of different sizes along the downsampling path to splice the shallow and deep feature maps, enhancing the network's expressive power in the multi-scale feature fusion process.
[0079] The Global-Local Context Module (GLCM) that integrates global semantics and local texture features includes the following steps:
[0080] A Global-Local Context Module (GLCM) that fuses global semantics and local textures is embedded in front of the newly added 160×160×128 detection head. Its core architecture consists of a global branch and a local branch in parallel. By extracting and synergizing multi-dimensional information of the input features, it enhances the expression ability of key features in the target detection task. Its structure diagram is shown below. Figure 7 shown.
[0081] The global branch first compresses the spatial dimension of the input feature map to 1×1 through global average pooling (GAP), and outputs channel features with a shape of B×C×1×1. Subsequently, two fully connected layers are used to implement channel compression (compression ratio is reduction=16) and dynamic weight adjustment for recovery. The first fully connected layer reduces the number of channels to c1 / 16, compresses the channel dimension, reduces the number of parameters, and extracts more compact global semantic information. After the ReLU activation function, it is restored to the original number of channels c1 through the second fully connected layer, and finally the global semantic weight is generated by the Sigmoid activation function for feature weighted fusion. Given an input feature map , global channel weight , its mathematical expression is as follows:
[0082]
[0083] At the same time, the local branch performs fine-grained spatial feature extraction on the input feature map through a 5×5 depth-separable convolution, and strengthens the local texture representation with batch normalization and SiLU activation function. The mathematical expression is as follows:
[0084]
[0085] In the feature fusion stage, a hybrid operation combining channel weighting and pixel-by-pixel addition is used. That is, the channel weights output by the global branch are weighted with the original features to enhance the semantically significant areas. Then, they are added pixel by pixel with the local features extracted by the local branch to output the final feature Y. This fusion operation performs adaptive enhancement in both the channel dimension and the spatial dimension, achieving complementary optimization of global and local features. Its mathematical expression is as follows:
[0086]
[0087] The GLCM module solves the common problems of insufficient global context perception and loss of local details in complex scenes due to the similarity between pest textures and plant textures by establishing the synergy between channel dependencies and spatial details, thereby improving the detection and recognition capabilities of small targets.
[0088] To build the YOLOv8-LQG model, the following steps are included:
[0089] Introducing the LHF module in the image preprocessing stage enables the model to quickly suppress the interference of complex backgrounds at the beginning of training and significantly enhance the target morphological features of pest larvae; then, a higher resolution detection layer QDDL of 160×160×128 obtained by four times downsampling is added to the original detection head to reduce the single-pixel receptive field to 4×4 pixels, directly improving the detection accuracy of small targets; at the same time, GLCM is embedded in front of the high-resolution detection layer to solve the problems of insufficient global context perception and loss of local details caused by the easy confusion between pest texture and plant texture by establishing inter-channel dependencies and capturing spatial detail information. The improved module structure in the model is marked with a red frame in the figure. The specific structure is as follows Figure 8 shown.
[0090] Training the YOLOv8-LQG model involves the following steps:
[0091] The training set images were fed into the YOLOv8-LQG model, and the experimental hyperparameters were set as shown in Table 1. The model was then trained using the partitioned training set. The model was iterated through 200 full training cycles. After each training cycle, performance evaluation on the validation set was initiated, and a strategy for archiving optimal weights was implemented, retaining the model parameters with the highest validation set performance.
[0092] Table 1
[0093]
[0094] The detection accuracy of the improved YOLOv8-LQG model is evaluated using indicators such as precision, recall, and mean average precision (mAP). Precision directly reflects the recognition accuracy of the model. A higher precision means that the model can more effectively avoid false detections and reduce the misidentification of background or non-target objects. Recall reflects the model's ability to detect targets, that is, the completeness of the model's recognition of real pest larvae. A higher recall means that the model can more effectively reduce the missed detection rate and ensure that larvae targets are not missed. The comprehensive indicator mAP, by averaging performance at different IoU thresholds, examines the model's improvement in refined positioning and classification capabilities compared with the original YOLOv8 model. The comparison results are shown in Table 2 below:
[0095] Table 2
[0096]
[0097] As shown in Table 2 above, the average precision and recall of the improved YOLOv8-LQG are significantly improved, which improves the accuracy of the model in detecting and localizing agricultural pest larvae and reduces missed detections.
[0098] For testing the YOLOv8-LQG model, the following steps are included:
[0099] The optimal model weights after YOLOv8-LQG model training are loaded and the test set is input. After feature extraction and aggregation, the Detect layer obtains the bounding box coordinates, confidence level, and category of targets that may contain agricultural pest larvae. Non-maximum suppression is then used to remove redundant detection boxes to obtain the final detection results.
[0100] Depend on Figure 9 As shown in the figure, the YOLOv8-LQG model proposed in this chapter achieves relatively accurate detection results in complex scenes with variable target poses, environmental occlusion, and target overlap. Visual comparison of the results directly demonstrates that the proposed model is more accurate and stable than the original model. The focus areas in the heat map show higher confidence in the true target, and the number and location of focus areas are more consistent with the actual target distribution.
[0101] The above description is merely a preferred embodiment of the present invention and is intended to be illustrative rather than restrictive of the present invention. Those skilled in the art will appreciate that many changes, modifications, and even equivalents may be made to the present invention within the spirit and scope of the claims, all of which fall within the scope of protection of the present invention.
Claims
1. A method for detecting agricultural pest larvae based on an improved YOLOv8, characterized in that: The steps include: S1: Constructed the APL8 (Agricultural Pest Larvae 8) dataset of 8 common agricultural pest larvae. S2: Based on the YOLOv8 model, we built the YOLOv8-LQG detection model by combining the designed Learnable Hue Filter (LHF) module, the Quadruple Downsampling Detection Layer (QDDL), and the Global-Local Context Module (GLCM) that integrates global semantics and local texture features. S3: Train the YOLOv8-LQG model; S4: Test the YOLOv8-LQG model.
2. The improved agricultural pest larvae detection method based on YOLOv8 according to claim 1, characterized in that: In step S1, the APL8 dataset is constructed, and the processing flow is as follows: Real agricultural pest images in natural scenes were collected, covering 8 common agricultural pest larvae, including fall armyworm, corn borer, cutworm, beet armyworm, Spodoptera litura, rice leaf roller, white grub and wireworm, and the APL8 (Agricultural Pest Larvae 8) dataset was constructed. The APL8 dataset was divided into training, validation and test sets, which were used for model training, validation and test sets respectively. Only the training set samples were expanded by data augmentation method, and the open source labeling tool LabelImg was used to annotate the larvae images with anchor boxes. The fall armyworm larvae were labeled as AWL, corn borer larvae as CBL, cutworm larvae as GBL, beet armyworm larvae as BAL, Spodoptera litura larvae as FAL, rice leaf roller larvae as RLFL, white grub larvae as GL, and wireworm larvae as WL.
3. The improved agricultural pest larvae detection method based on YOLOv8 according to claim 1, characterized in that: In step S2, the Learnable Hue Filter (LHF) module is designed; its processing flow is as follows: S21: Convert the input RGB image to HSV space and extract the normalized hue (Hue, H). The mathematical expression is as follows: V represents the brightest channel value in the color, Δ is the difference between the brightest and darkest channels, which is used to calculate saturation, and ϵ is a numerical stability term; S22: Based on the extracted hue component, LHF is multiplied by two Sigmoidal layers to generate a bell-shaped curve mask, which replaces the traditional threshold segmentation and allows gradient backpropagation. The center hue c∈[0,1] and hue width w∈[0,0.5] are used as network parameters and are automatically optimized during training. The mathematical expression is as follows: Here, σ is the Sigmoid function, initially set to c = 0.5 (90 / 180) and width w = 0.167 (30 / 180) corresponding to the yellow-green background area. As the training data increases, the mask parameters automatically adapt to the data-driven offset direction, thereby suppressing the response intensity of the background hue and enhancing the sensitivity of the target hue. S23: Embed the LHF module into the preprocessing stage of YOLOv8, and its output feature X out Directly inputting the data into the backbone network enables the model to quickly suppress background interference in the early stages of training. At the same time, the color-sensitive area is continuously corrected through gradient backpropagation. When the hue distribution of pests deviates from the initial hypothesis, c and w automatically shift in the data-driven direction.
4. The method for detecting agricultural pest larvae based on improved YOLOv8 according to claim 1, characterized in that: In step S2, a quadruple downsampling detection layer (QDDL) is added, and its processing flow is as follows: Based on the original structure of YOLOv8, a high-resolution feature map of 160×160×128 is generated through a four-fold downsampling operation, which reduces the receptive field corresponding to a single pixel to 4×4 pixels, improving the model's ability to represent small targets, and adding a corresponding small target detection layer in the Head part; at the same time, an upsampling module, a Concat module and a C2f module are added to the feature fusion stage. In addition, the feature details of these small targets can also be further transmitted to the other three feature layers of different sizes along the downsampling path, so as to splice the shallow and deep feature maps, thereby making better use of global context information during detection.
5. The method for detecting agricultural pest larvae based on improved YOLOv8 according to claim 1, characterized in that: In step S2, the Global-Local Context Module (GLCM) that integrates global semantics and local texture features has the following processing flow: S21: A contextual enhancement module that fuses global semantics and local texture is embedded in front of the newly added 160×160×128 detection head. Its core architecture consists of a global branch and a local branch in parallel. By extracting and collaboratively fusing multi-dimensional information from input features, it enhances the expressive power of key features in object detection tasks. S22: The global branch first compresses the spatial dimension of the input feature map to 1×1 through global average pooling (GAP), and outputs channel features with a shape of B×C×1×1; Then, two fully connected layers are used to implement channel compression (compression ratio is reduction=16) and dynamic weight adjustment for restoration. The first fully connected layer reduces the number of channels to c1 / 16, compressing the channel dimension to reduce the number of parameters and extract more compact global semantic information. After the ReLU activation function, it is restored to the original channel number c1 through the second fully connected layer, and finally the Sigmoid activation function is used to generate the global semantic weight for feature weighted fusion; Given an input feature map , global channel weight , its mathematical expression is as follows: S23: At the same time, the local branch performs fine-grained spatial feature extraction on the input feature map through a 5×5 depth-separable convolution, and strengthens the local texture representation with batch normalization and SiLU activation function; local feature extraction The mathematical expression is as follows: S24: In the feature fusion stage, a hybrid operation combining channel weighting and pixel-by-pixel addition is used. That is, the channel weights output by the global branch are channel-weighted with the original features to enhance the semantically significant areas. Then, the weights are added pixel-by-pixel with the local features extracted by the local branch to output the final feature Y. This fusion operation performs adaptive enhancement in both the channel dimension and the spatial dimension to achieve complementary optimization of global and local features. Its mathematical expression is as follows: The GLCM module solves the common problems of insufficient global context perception and loss of local details in complex scenes due to the similarity between pest textures and plant textures by establishing the synergy between channel dependencies and spatial details, thereby improving the detection and recognition capabilities of small targets.
6. The improved agricultural pest larvae detection method based on YOLOv8 according to claim 1, characterized in that: In step S2, the YOLOv8-LQG model is constructed, and its processing flow is as follows: The LHF module is introduced in the image preprocessing stage, which enables the model to quickly suppress the interference of complex backgrounds in the early stages of training and significantly enhance the target morphological features of pest larvae. Subsequently, a higher-resolution detection layer (QDDL) of 160×160×128 obtained by fourfold downsampling is added to the original detection head, reducing the single-pixel receptive field to 4×4 pixels, directly improving the detection accuracy of small targets; at the same time, GLCM is embedded in front of this high-resolution detection layer. By establishing inter-channel dependencies and capturing spatial detail information in synergy, the problems of insufficient global context perception and loss of local details caused by the easy confusion between pest texture and plant texture are solved.
7. The improved agricultural pest larvae detection method based on YOLOv8 according to claim 1, characterized in that: In step S3, the YOLOv8-LQG model is trained, and the processing flow is as follows: The training set images are input into the YOLOv8-LQG model, and the hyperparameters are set for training. The training process continues until the model converges and the optimal model weights are finally obtained.
8. The improved agricultural pest larvae detection method based on YOLOv8 according to claim 1, characterized in that: In step S4, the YOLOv8-LQG model is tested, and the processing flow is as follows: The optimal model weights are loaded and the divided test set images are input. After feature extraction and aggregation, the Detect layer obtains the bounding box coordinates, confidence level, and category of targets that may contain agricultural pest larvae. Non-maximum suppression is then used to remove redundant detection boxes to obtain the final detection results.