YOLOv8-based pest and disease damage detection method
By improving the YOLOv8 object detection network model, combining the DenseNet and ResNet fusion architecture, local feature capture and global context modeling are introduced, which solves the problem of small target recognition difficulties and computing resource limitations in fruit pest detection, and achieves efficient and low-cost pest detection.
Patent Information
- Application Number
- CN202510338152.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has problems such as difficulty in identifying small targets, dense distribution of fruits, severe leaf occlusion, large ambient light changes and limited computing resources in fruit pest detection, which is difficult to meet the actual needs of eggplant pest detection.
The improved YOLOv8 object detection network model is adopted, combined with the DenseNet and ResNet fusion architecture, local feature capture, global context modeling and dynamic feature modulation are introduced, multi-branch network enhancement and DAE denoising optimization are adopted, model training is carried out through jump connection and feature fusion loss function, and loss function is iteratively solved by ADMM, and a quadratic matching strategy of high and low score detection boxes is introduced.
It improves the accuracy and efficiency of pest detection, adapts to different computing resources and environments, reduces false detection and missed detection, and is suitable for intelligent agriculture and pest monitoring.
Smart Images

Figure CN120279458A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the cross - fields of edge detection, computer vision, and intelligent agricultural monitoring, and particularly relates to a pest and disease detection method based on YOLOv8. Background Art
[0002] Fruit pest and disease detection is an important link in ensuring the healthy growth of crops and improving agricultural production efficiency. As an important cash crop, eggplant fruits are vulnerable to various pests and diseases during the growth process, such as anthracnose, brown spot, aphids, and whiteflies, which seriously affect fruit quality and yield. Traditional pest and disease detection methods mainly rely on manual inspections, which are greatly affected by factors such as experience, environment, and lighting, and have problems such as high labor intensity, low detection efficiency, and poor real - time performance, making it difficult to meet the needs of large - scale and refined agricultural management. In recent years, with the rapid development of computer vision, deep learning, and intelligent agriculture, the automatic recognition method of fruit pests and diseases based on object detection has gradually become a research hotspot. Compared with traditional manual detection, machine vision detection technology can achieve efficient and accurate fruit pest and disease recognition in complex agricultural environments, providing effective support for precision agriculture and intelligent agricultural machinery. However, fruit pest and disease detection in agricultural environments still faces many challenges, such as small disease symptom areas, severe leaf occlusion, similar fruit colors to the background, and large environmental lighting changes. These factors limit the detection accuracy and generalization ability of traditional object detection algorithms to a certain extent.
[0003] In the field of deep learning object detection, the YOLO (You Only Look Once) series, as a mainstream single-stage detection framework, features fast detection speed and high recognition accuracy, and has been widely applied to crop pest and disease detection tasks. In recent years, researchers have improved the performance of YOLO in aspects such as small object detection and complex background interference suppression by optimizing the network backbone structure, introducing attention mechanisms, and combining feature pyramid networks. For example, in crop disease detection tasks such as grapes, blueberries, and tomatoes, improved YOLO versions can enhance detection accuracy in environments with dense fruits, complex backgrounds, and leaf occlusion. However, most existing studies optimize for a single problem and have not yet formed a systematic improvement plan that takes into account multiple aspects such as small object detection, dense distribution, complex background interference, detection speed, and computational cost, making it difficult to meet the actual needs of eggplant pest and disease detection. In addition, the high computational cost of deep neural networks remains the main obstacle restricting their widespread application. Due to the large scale of deep learning model parameters, a large amount of labeled data is required during the training process, which is particularly restricted by the difficulty of data acquisition in the agricultural field. Transfer learning, as an effective solution, can transfer the knowledge of existing trained models to related tasks to reduce data requirements and training costs. Transfer learning has been applied in agricultural vision tasks such as fruit disease detection and pest and disease identification, and has improved the adaptability and detection efficiency of models by optimizing feature extraction networks, improving network pruning strategies, and introducing multi-source transfer methods. Summary of the Invention
[0004] In order to solve the problems existing in fruit pest and disease detection, such as difficult recognition of small objects, dense distribution of fruits, serious leaf occlusion, large environmental light changes, and limited computing resources, this application proposes a pest and disease detection method based on YOLOv8, which optimizes YOLOv8 and combines methods such as feature enhancement, attention mechanism, and adaptive training strategy to improve the accuracy and efficiency of eggplant fruit pest and disease detection, so as to achieve an efficient and low-cost intelligent monitoring and prevention solution.
[0005] The technical solution adopted in this application is: a pest and disease detection method based on YOLOv8, including the following steps:
[0006] S1: Construct a dataset and perform preprocessing;
[0007] S2: Construct a target detection network model that improves YOLOv8. The model includes a backbone network, a neck network, and multiple detection heads. The backbone network adopts a fusion architecture of DenseNet and ResNet, including variants of the C2f module, namely the C2f-OREPA module, the C2f-SPD module, and the C2f-DBB module. The neck network realizes dynamic feature sampling and splicing through Dysample and Concat, and the detection head adopts the design of a diversified branch block;
[0008] Moreover, a convolutional module that integrates local feature capture, global context modeling, and dynamic feature modulation, and an attention mechanism that integrates dynamic feature modulation and adaptive enhancement of features are adopted in the model. At the same time, a denoising network is combined based on DenseNet to construct the DAE-DenseNet network;
[0009] S3: Model training: The model is trained using a loss function that fuses the denoising result with the feature layer through skip connections;
[0010] S4: Model inference: The ADMM iteration is used to solve the loss function, and the solution result is fused with the features output by the DAE-DenseNet network to output the target tracking detection box.
[0011] Furthermore, the structure of the backbone network is as follows: The input data is processed by two Conv ROI modules and then undergoes feature enhancement through the C2f-OREPA module. Then, the spatial information is converted into depth information through the SPD module. After further feature extraction through the Conv ROI module, the feature is processed by the C2f-SPD module. After being processed by the SPD module, Conv ROI module, C2f-SPD module, SPD module, and Conv ROI module, the extracted features are input into the C2f-DBB module for further feature extraction and fusion. Finally, multi-scale feature extraction is performed through the SPPF module. The outputs of the two C2f-SPD modules and the SPPF module are all connected to the neck network for dynamic feature extraction and splicing.
[0012] Furthermore, the structure of the convolutional module that integrates local feature capture, global context modeling, and dynamic feature modulation is as follows: The input original image data is segmented into multiple small patches through the patch embedding module for embedding processing; the feature map after embedding processing undergoes a non-linear transformation through an activation function and then feature standardization through batch normalization; the input feature map is segmented into a series of 3×3 non-overlapping receptive field sliders; group convolution operations are performed on the segmented sliders to efficiently extract the features of each local region and expand the feature set; the feature map generated by group convolution matches the dimension of the input feature map to form a preliminary spatial feature representation; global average pooling is used to integrate the local features to obtain global context information, reducing the number of parameters and computational complexity; an attention weight map is generated through the Softmax function to weight the features of the receptive field sliders, emphasizing the feature information of key regions; the weighted attention map is element-wise multiplied with the spatial feature map to generate a weighted feature representation modulated by attention; the feature map input to the Conv ROI module is divided into three different paths and independently processed through Patch = 6, Patch = 7, and Patch = 8 in the CSMM module respectively; the feature maps output from the three paths are fused through an addition operation to integrate multi-scale feature information; the fused feature map undergoes global average pooling to further extract global context features; the first fully connected layer of the multi-layer perceptron is used to generate intermediate features; the second fully connected layer outputs the final aggregated features; through a channel expansion operation, the number of channels of the aggregated features is expanded; the expanded features are element-wise multiplied with the input feature map to achieve the fusion of the input and the channel expansion output; the visual context information of the feature map is encoded at multiple scales through sequential depth convolutional layers; a gating operation is performed on the multi-path features output from the CSMM module and the context focusing result to selectively aggregate context information; through a linear affine transformation, the modulator is applied to each feature element, and the optimized feature representation is element-wise added to the original input feature map to form a residual structure, retaining key information and enhancing training stability.
[0013] Furthermore, the attention mechanism that fuses feature dynamic modulation and adaptive enhancement realizes feature enhancement through a three-channel branch, including the X-channel, Y-channel of the original coordinate attention, and the global average pooling channel. The input feature map is subjected to average pooling in the height direction and width direction to obtain the feature maps of the X-channel and Y-channel respectively. Convolution operations are performed on the feature maps of the X-channel and Y-channel to extract fine-grained features in the height and width directions. The feature maps of the X-channel and Y-channel are activated through the sigmoid function to obtain weighted results, and are respectively multiplied element-wise with the original input feature map to strengthen the feature information. A third-channel branch is added, and global average pooling is used to perform overall downsampling on the input feature map, retaining the global background information and fusing it in parallel with the X and Y channels to achieve multi-path feature enhancement; visual context information is encoded at different levels through sequential depth convolutional layers, collecting spatial features of height and width from different scales to enhance the understanding of details and overall structure; gating operations are performed on the feature maps of the X-channel, Y-channel, and global pooling channel of the multi-path coordinate attention mechanism to filter and aggregate context information, preferentially retaining key details and suppressing redundant background information, and applying the gated and aggregated feature maps to each query element through linear affine transformation to achieve dynamic modulation and precise adjustment of features, enhancing the adaptability to complex visual content; through element-wise affine transformation, dynamic modulation and adaptive enhancement of features are realized.
[0014] Furthermore, the C2f-DBB module is formed by replacing the convolution operation in the original Bottleneck module of the C2f module in the traditional YOLOv8 backbone network with the DBB module; the DBB module includes multiple branches of different scales, and each branch uses a 1×1, 1×1-k×k composite convolution and a 1×1-AVG combined structure; among them, the 1×1-k×k composite convolution reduces the computational complexity by first performing a 1×1 convolution and then enhancing the local receptive field using a k×k convolution, thereby improving the feature expression ability while ensuring computational efficiency; the 1×1-AVG combined branch reduces the dimension using a 1×1 convolution and then uses average pooling to further enhance feature smoothness to reduce noise interference and improve the robustness of the network; and the structure reparameterization technology is introduced to simplify the multi-branch structure in the DBB module into an equivalent standard convolution operation during inference.
[0015] Furthermore, the C2f-OREPA module effectively reduces the computational and storage overhead during the training process through two key steps of block linearization and block compression, while maintaining the performance advantages in the inference stage. During the block linearization process, the multi-branch structure in the DBB module is merged through a linear transformation method, making it behave more like a single-branch convolutional network during the training process without losing the original expression ability; then, during the block compression process, the intermediate feature representation is optimized to reduce the computational amount and storage requirements, thereby achieving more efficient training.
[0016] Further, the loss function includes: denoising loss and object detection loss. The denoising loss includes image denoising loss, gradient branch loss, and fractional total variation loss. The object detection loss is implemented through the EloU loss function.
[0017] Further, the process of iteratively solving the loss function using the alternating direction optimization strategy ADMM is as follows. First, represent the input image data as a combination of a sparse noise matrix and a low-rank structure. Optimize the sparse noise matrix through the (2,1) norm, and iteratively solve the decomposition of the low-rank matrix and the sparse noise matrix through ADMM. Finally, generate a low-noise image structure and fuse it with the features output by the DAE-DenseNet network to separate the high-frequency noise and low-rank structure of the image.
[0018] Further, after the model outputs the target tracking detection box, the coherence of the target trajectory is maintained by introducing a quadratic matching strategy for high- and low-score detection boxes, including the following steps: Based on the detection results of the model, set a high-score threshold and a low-score threshold, divide the detection boxes into a high-score detection box set and a low-score detection box set, use a data association method to perform the first match on the high-score detection box set and the trajectory set, and update the successfully matched trajectories; perform a second match on the trajectory set that failed the first match and the low-score detection box set, update the successfully matched trajectories, and delete the unmatched low-score detection boxes; for the set of unmatched high-score detection boxes, if its confidence is greater than the tracking score threshold, create a new trajectory; for the set of lost trajectories, retain a set number of frames, perform a match when the trajectory reappears, and if it does not appear within the set number of frames, delete the trajectory;
[0019] Calculate the loU distance matrix between the high-score detection box set and the trajectory set; use the Hungarian algorithm to match the loU distance matrix. The successfully matched trajectories are updated through the Kalman filter and added to the current frame trajectory set; the unmatched trajectories are placed in the unmatched trajectory set, and the unmatched high-score detection boxes are placed in the unmatched detection box set.
[0020] Further, the second matching step includes calculating the loU distance matrix between the low-score detection box set and the unmatched trajectory set; using the Hungarian algorithm for matching. The successfully matched trajectories are updated through the Kalman filter and added to the current frame trajectory set; the unmatched trajectories are placed in the lost trajectory set, and the unmatched low-score detection boxes are directly deleted; for the set of unmatched high-score detection boxes, if its confidence is higher than the set tracking score threshold, create a new trajectory and merge it into the current frame trajectory set; the trajectories in the lost trajectory set are retained for a set number of frames, perform a match when the trajectory reappears within the set number of frames, otherwise delete the trajectory.
[0021] The beneficial effects of this application compared with the prior art are as follows: This application combines model pruning, lightweight design and dynamic optimization strategies to adapt to different computing resource environments and improve the real-time detection performance. By introducing local feature enhancement, global context modeling and dynamic feature modulation, the detection accuracy and stability of pest and disease targets are improved. Combining the attention mechanism, multi-branch network enhancement, and DAE denoising optimization, and adopting a feature fusion loss function based on skip connections to improve the detection accuracy and reduce false detection and missed detection problems. Through the secondary matching strategy of high and low score detection boxes, the stability and tracking accuracy of the detection boxes are improved, which is suitable for video stream detection tasks. The present invention can be widely applied to fields such as intelligent agriculture, pest and disease monitoring, precision plant protection, and crop health diagnosis, providing important technical support for the intelligent development of modern agriculture. Brief Description of the Drawings
[0022] The following further describes this application with reference to the drawings:
[0023] Figure 1 It is a schematic diagram of the model improvement framework of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application.
[0024] Figure 2 It is a schematic diagram of the edge device deployment of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application.
[0025] Figure 3 It is a flowchart of the convolution method of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application.
[0026] Figure 4 It is a schematic diagram of the backbone DenseNet and OREPA composite network of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application.
[0027] Figure 5 It is a schematic diagram of the process of the feature fusion DBB of the network output of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application.
[0028] Figure 6 It is a schematic diagram of the model operation result of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application.
[0029] Figure 7 It is a schematic diagram of the detection effect of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application.
[0030] Figure 8 It is a schematic diagram of the data training iteration of a pest and disease detection method based on YOLOv8 provided by an embodiment of this application. Detailed Embodiments
[0031] As shown Figures 1 to 8 in the figure, this application provides a pest and disease detection method based on YOLOv8, which uses an improved target detection network model of YOLOv8. Its backbone network adopts a fusion architecture of DenseNet and ResNet, which can improve the efficiency of gradient propagation and enhance the ability of feature reuse. Its neck network realizes dynamic feature sampling and splicing through Dysample and Concat. Its detection head adopts the design of a diversified branch block, and the diversified branch block adopts a lightweight Inception module to ensure that pest and disease targets of different scales can be effectively detected and adapted to various agricultural environments.
[0032] The overall framework of the model of this application is as Figure 1 shown in the figure. The overall process is that the input data is first processed by the Conv ROI module and then undergoes feature enhancement through the C2f-OREPA module. Then, it passes through the SPD module to convert spatial information into depth information. After further feature extraction through the Conv ROI module, it undergoes feature processing through the C2f-SPD module. After passing through the SPD module, Conv ROI module, C2f-SPD module, SPD module, and Conv ROI module, the extracted features are input into the C2f-DBB module for further feature extraction and fusion. Finally, multi-scale feature extraction is performed through the SPPF module. The outputs of the two C2f-SPD modules and the SPPF module are all output to the neck network for dynamic feature extraction and splicing, and then output in three paths to the detection head. The features of different paths are processed through multiple C2f modules and Conv ROI modules, and finally the target detection results are output in the detect module.
[0033] In the model of this application, a convolutional module that captures local features, models global context, and modulates dynamic features is adopted. The structure of this convolutional module is as follows: The input original image data is segmented into multiple small blocks through a patch embedding module for embedding processing. The feature map after embedding processing undergoes a non-linear transformation through an activation function (GELU), and then feature standardization is performed through batch normalization. The input feature map is divided into a series of non-overlapping receptive field windows (sliding blocks) of 3×3. Group convolution operations are performed on the segmented sliding blocks to efficiently extract the features of each local region and expand the feature set. The feature map generated by group convolution matches the dimension of the input feature map to form a preliminary spatial feature representation. Global average pooling is used to integrate local features to obtain global context information, reducing the number of parameters and computational complexity. An attention weight map is generated through the Softmax function to weight the features of the receptive field sliding blocks, emphasizing the feature information of key regions. The weighted attention map is multiplied element-wise with the spatial feature map to generate a weighted feature representation modulated by attention. The feature map input to the Conv ROI module is divided into three different paths and independently processed through Patch = 6, Patch = 7, and Patch = 8 in the CSMM module respectively. Each path uses convolutional kernels of different scales to perform multi-scale feature extraction on the feature map to capture context information at different spatial scales. The feature maps output by the three paths are fused through an addition operation (+) to integrate multi-scale feature information. The fused feature map undergoes global average pooling to further extract global context features. The first fully connected layer of a multi-layer perceptron (MLP structure) is used to generate intermediate features; the second fully connected layer outputs the final aggregated features. Through a channel expansion operation, the number of channels of the aggregated features is expanded. The expanded features are multiplied element-wise with the input feature map to achieve the fusion of the input and the output of channel expansion. The visual context information of the feature map is encoded at multiple scales through sequential depth convolutional layers. A gating operation is performed on the multi-path features output by the CSMM module and the context focusing result to selectively aggregate context information. Through a linear affine transformation, a modulator is applied to each feature element, and the optimized feature representation is added element-wise (+) to the original input feature map to form a residual structure, retaining key information and improving training stability.
[0034] Among them, local feature capture adopts an improved depthwise separable convolution to improve the real-time inference ability of edge devices by reducing computational complexity. Global context modeling adopts a fusion strategy of spatial attention mechanism and channel attention to enhance the distinguishability between the target and the background and improve detection accuracy. Dynamic feature modulation adopts variable receptive field convolution to adaptively adjust the feature extraction strategy and improve the adaptability of the model to different target scales.
[0035] In the model of this application, an attention mechanism of feature dynamic modulation and adaptive enhancement is adopted. Feature enhancement is achieved through three-channel branches, including the X-channel, Y-channel of the original coordinate attention, and the global average pooling channel. The input feature map is subjected to average pooling in the height direction and width direction to obtain the feature maps of the X-channel and Y-channel respectively. Convolution operations are performed on the feature maps of the X-channel and Y-channel to extract fine-grained features in the height and width directions. The feature maps of the X-channel and Y-channel are activated through the sigmoid function to obtain weighted results, and are respectively multiplied element-wise with the original input feature map to strengthen the feature information. A third-channel branch is added, and global average pooling is used to perform overall downsampling on the input feature map, retaining the global background information and fusing it in parallel with the X and Y channels to achieve multi-path feature enhancement. Visual context information is encoded at different levels through sequential depth convolutional layers, and spatial features of height and width are collected from different scales to enhance the understanding of details and overall structure; gating operations are performed on the feature maps of the X-channel, Y-channel, and global pooling channel of the multi-path coordinate attention mechanism to screen and aggregate context information, preferentially retaining key details and suppressing redundant background information, and applying the gated and aggregated feature maps to each query element through linear affine transformation to achieve dynamic modulation and precise adjustment of features, and improve the adaptability to complex visual content. Through element-wise affine transformation, dynamic modulation and adaptive enhancement of features are achieved, enhancing the flexibility and generalization ability of the network.
[0036] Among them, feature dynamic modulation optimizes the fusion of high-level and low-level features through a feature weight allocation strategy based on Transformers to ensure enhanced saliency of small targets. The adaptive enhancement attention mechanism combines channel attention (SE-block) and self-attention (Self-Attention) to improve the adaptability of the model to complex backgrounds and reduce false detections and missed detections.
[0037] The C2f-DBB module of this application is formed by replacing the convolution operation in the original Bottleneck module in the C2f module of the backbone network with the DBB module; the DBB module includes multiple branches of different scales, and each branch uses a 1×1, 1×1-k×k composite convolution and a 1×1-AVG combination structure. Among them, the 1×1-k×k composite convolution first performs a 1×1 convolution to reduce the computational complexity, and then uses a k×k convolution to enhance the local receptive field, thereby improving the feature expression ability while ensuring the computational efficiency. The 1×1-AVG combination branch uses a 1×1 convolution for dimensionality reduction and then uses average pooling (AVG Pooling) to further enhance the feature smoothness to reduce noise interference and improve the robustness of the network. The structural reparameterization technology is introduced to simplify the DBB structure into an equivalent standard convolution operation during inference to ensure the computational efficiency. As Figure 5 shown.
[0038] In terms of the mathematical expression form, the above calculation process can be expressed as follows:
[0039] For an input feature map I, if multiple different convolutional kernels F (1) and F (2) are used for calculation during training, then theoretically, this operation is equivalent to first calculating the two convolution results and then summing them up, that is:
[0040] I * F (1) + I * F (2) =I * (F (1) + F (2) );
[0041] Among them, the symbol * represents the convolution operation. Similarly, if a certain feature is scaled by p, this operation is equivalent to first scaling the convolutional kernel and then performing the convolution:
[0042] p(I * F) = I * (pF);
[0043] During the process of channel number conversion, the DBB module converts the input channel number from C1 to the output channel number C out , and the calculation formula is as follows:
[0044] C out = C · e;
[0045] Among them, the value of e is 0.5, indicating that the output channel number C out is reduced by half compared to the input channel number C, thereby further reducing the computational complexity. When the feature map is converted, the stride is set to 1, and when the number of groups g > 1, group convolution is introduced to enhance the diversity of feature expression, enabling different channel groups to capture information at different scales and improving the robustness of the model.
[0046] In the standard neural network architecture, the BN layer is mainly used to normalize the feature map to ensure the stability and fast convergence of the network. However, during the inference stage, the mean and variance of BN can be regarded as fixed parameters, so through mathematical equivalent transformation, the BN operation can be fused with the convolution weights, thereby reducing the additional computational overhead. Specifically, the fused convolution weights are calculated as follows:
[0047]
[0048] The offset fusion is calculated as follows:
[0049]
[0050] The final equivalent calculation can be expressed as:
[0051] BN(Conv(x)) = Wfused ·x + b fused ;
[0052] Where W and b represent the convolution kernel weights and bias parameters respectively, γ and β are learnable parameters in BN calculation, and mean and var represent the mean and variance statistically obtained during BN calculation.
[0053] In the backbone network structure of YOLO v8n, in order to achieve dimensionality reduction during feature extraction, a standard convolution module is adopted, with a kernel size of 3×3, a stride set to 2, and it is placed between two C2f modules to optimize the feature expression ability.
[0054] The internal convolution method of the DBB module in the C2f-DBB module mainly consists of two parts, namely the upper-layer calculation path and the lower-layer calculation path. Among them, the upper-layer calculation path is mainly responsible for extracting global information from the input feature map. The method is to use a 3×3 global average pooling (Global Average Pooling, GAP) operation with a stride set to 2 to ensure that the global information within the receptive field can be effectively captured. Next, 1×1 grouped convolution (GroupedConvolution) is used for information interaction, and then the features within the receptive field are normalized through the Softmax function to measure the importance of different features and generate a receptive field attention map (Receptive Field AttentionMap), which can be used to guide the subsequent feature screening and extraction. The main task of the lower-layer calculation path is to perform spatial feature extraction. Its calculation method is to use 3×3 grouped convolution with the same stride set to 2 to ensure that the finally generated receptive field spatial feature map has the same dimension as the receptive field attention feature map obtained from the upper-layer calculation for subsequent feature weighting operations. Since the grouped convolution technology is used in the calculation process, this module can effectively reduce the calculation cost, reduce the computational complexity, and improve the adaptability of the model on edge computing devices. After obtaining the receptive field attention feature map and the receptive field spatial feature map, we further fuse the two by element-wise multiplication to screen out the most discriminative features. Finally, through the dimension adjustment operation, the output feature map is adapted to the input requirements of the subsequent network structure, thus completing the entire receptive field attention convolution calculation.
[0055] In terms of mathematical expression, the above calculation process can be expressed as: First, let the input feature map be X, and the feature extraction in the upper-layer calculation path is through Performing global average pooling operation, where Represents the global average pooling operator with a kernel size of 3×3 and a stride of 3.
[0056] Next, group convolution is utilized to perform feature transformation on the global pooling result, and then obtain the normalized attention feature map through the Softmax operation F Softmax (·). Meanwhile, in the lower-level computational path, we apply a 3×3 group convolution with a stride of 2 to generate the receptive field spatial feature map.
[0057] Finally, we use element-wise multiplication to fuse the spatial feature and the attention feature, and perform further feature processing through the standard convolution F Conv (·), batch normalization (Batch Normalization, BN) F BN (·) and the non-linear activation function SiLU SiLU (·) to obtain the final output feature map F RFACov (X). Its mathematical expression is as follows:
[0058]
[0059] where, represents the group convolution operation with a kernel size of 3×3 and a stride of 2, is the group convolution operation with a kernel size of 1×1 and a stride of 1, represents the global average pooling with a kernel size of 3×3 and a stride of 3, F Conv (·) is the standard convolution operation, F BN (·) is the batch normalization, F SiLU (·) is the non-linear activation function SiLU, F ReLU (·) is the non-linear activation function ReLU, F Softmax (·) represents the Softmax normalization function, and represents the element-wise multiplication operation.
[0060] To ensure the gradient stability of the network, we introduce a skip connection in the DBB module for intermediate feature fusion to enhance the information flow transmission ability and mitigate the gradient vanishing problem. This connection method allows the intermediate layer features to be directly fused with the subsequent layers, improving the feature reuse rate and enhancing the depth expression ability of the model, enabling the DBB module to still maintain good convergence in large-scale neural networks.
[0061] The core structure of DenseNet consists of multiple Dense Blocks. Different from ResNet which mainly relies on residual connections, DenseNet adopts a fully connected feature reuse strategy. That is, within the same Dense Block, the output of each layer is passed not only to the next layer but also to all subsequent layers, making the feature representation of the network richer and the information flow more efficient. In a traditional convolutional neural network, for a network with L layers, the number of inter-layer connections is L. However, in the Dense Block structure, the number of inter-layer connections reaches L(L + 1) / 2. This dense connection method significantly enhances the efficiency of feature transfer and reduces the risk of gradient vanishing. For example, assume that the input image x0 is processed through a DenseBlock containing L layers, and each layer's non-linear transformation function H i consists of BN (Batch Normalization), ReLU (activation function), and 3×3 convolution (Conv). Then the output x i of the i-th layer is obtained by concatenating the feature maps of all previous layers and then inputting them into H i . The mathematical expression is as follows:
[0062] x i = H i ([x0, x1, …, x i-1 );
[0063] where [x0, x1, …, x i-1 represents the concatenation of the feature maps of all previous layers. This mechanism ensures that information can be fully utilized in the deep network, enabling both shallow features and deep features to directly affect the final output. Since the size of the feature maps within the Dense Block needs to be consistent, the network introduces a Transition Layer between different Dense Blocks to adjust the size and number of channels of the feature maps, ensuring the smooth transfer of features between different Dense Blocks. Through this structure, DenseNet significantly enhances the information flow, enabling the network to make full use of shallow information, thereby improving the generalization ability of the model. In addition, this dense connection method can provide more computational paths during the gradient backpropagation, effectively alleviating the problem of gradient vanishing and enabling the deep network to converge more stably during the training process. As Figure 4 shown.
[0064] During the optimization of deep neural networks, the Re-parameterization technique aims to enhance the model's representation ability through structural transformation without increasing the computational overhead during the inference phase. The core idea of this method is to introduce a more complex network topology during the training phase to improve the model's feature expression ability, and then transform the complex structure obtained during training into an equivalent simple model during the inference phase, thus achieving efficient computation during inference. This application proposes a method through two key steps: Block Linearization and Block Squeezing, which can effectively reduce the computational and storage overhead during training while maintaining the performance advantages during the inference phase. During the Block Linearization process, we merge the multi-branch structure through linear transformation methods, making it behave more like a single-branch convolutional network during training without losing its original expression ability. Then, during the Block Squeezing process, we optimize the intermediate feature representation to reduce the computational amount and storage requirements, thus achieving more efficient training. The advantage of this strategy is that it can avoid additional computational overhead during training while still maintaining the superior expression ability of the re-parameterized network, without affecting the performance during the inference phase.
[0065] During the process of model architecture optimization, Batch Normalization (BN) is one of the important factors affecting training efficiency. The main role of the BN layer is to normalize the feature map, making its mean approach 0 and variance approach 1, thereby stabilizing model training and accelerating convergence. However, the computational process of BN involves normalizing the feature maps of each independent branch. This process is not only highly non-linear but also requires calculating the mean and variance for each branch separately, resulting in high computational complexity and video memory occupancy. Therefore, in traditional re-parameterization methods, extensive use of BN layers often leads to additional computational burdens, increases training time, and exerts great pressure on GPU memory resources. Although BN plays an important role in optimizing the weight adjustment of different branches, directly removing BN may lead to a decline in the model's expression ability, thereby affecting the final detection accuracy.
[0066] The linear scaling layer can be perfectly integrated with the OREPA structure, making the calculation process more fluent and avoiding the gradient suppression problem brought by the traditional BN layer. To further verify the effectiveness of this method, we compared several different normalization strategies, including Layer Normalization (LN), Group Normalization (GN), and Instance Normalization (IN), and analyzed their respective calculation characteristics and applicable scenarios. LN often brings a high computational burden because it needs to perform calculations on all channel dimensions of a single sample; although GN can effectively reduce the computational burden by dividing the channels into multiple groups and calculating the mean and variance separately, it introduces an additional need for hyperparameter adjustment; IN only calculates the mean and variance within each channel of a single sample, so it cannot fully utilize the global information across channels and performs poorly in tasks such as defect detection. Therefore, compared with these normalization methods, the linear scaling layer can still maintain the optimization ability of BN while reducing the computational overhead, making it an ideal choice in the OREPA structure.
[0067] In the linearized block structure, we further optimized the calculation mode during training. By deleting all non-linear BN layers and replacing them with linear scaling layers, the computational complexity was reduced, enabling the calculations of different branches to be unified and integrated, and avoiding the additional storage requirements and computational bottlenecks caused by the BN layer. Specifically, the mathematical description of this optimization strategy is as follows: Assuming the input feature map is X, the traditional BN calculation method needs to perform the normalization operation separately for each branch where μ and σ represent the mean and standard deviation of the feature map respectively, and γ and β are learnable parameters. In our OREPA method, we directly use a linear scaling layer to replace it with X′ = αX + β, where α and β are learnable parameters and do not depend on the calculation of the mean and variance, thus greatly reducing the amount of calculation. In addition, since all calculations can be merged after linearization, we can integrate the calculations of multiple branches during the training phase, so that no additional conversion steps are required during inference, thereby further improving the computational efficiency.
[0068] To achieve high-quality denoising effects and retain the detailed features of the image, this application proposes a loss function, including: a denoising loss L denoise and an object detection loss L detect , specifically expressed as L total = L denoise + αL detect , where α is the target weight, and the denoising loss L denoiseIt includes the image denoising loss (LID), the gradient branch loss (LGB), and the fractional total variation loss (FFTV). These loss functions are optimized for different objectives, enabling the network to preserve edge information and overall texture details of the image as much as possible while denoising. Specifically, the image denoising loss (LID) is used to constrain the denoising ability of the network, making the denoised image as close as possible to the real clean image; the gradient branch loss (LGB) is used to enhance the gradient information of the image to ensure the integrity of the edge structure and texture; the fractional total variation loss (LFTV) is used to reduce noise and remove artifacts while preserving the local detailed texture of the image. The combined action of these three loss functions enables the denoising part of the network structure DGGNet (Denoising Generative Guided Network) to have stronger generalization ability and better visual effects in the denoising task.
[0069] In the image denoising task, our goal is to make the output of the model as close as possible to the real noise-free image. Therefore, we use the l1 loss function to measure the absolute error between the denoising result and the real image to ensure the convergence and stability of the model.
[0070] Suppose we have N sets of training samples, and each sample consists of the input noisy image u and the corresponding noise-free image y. That is, the training sample set is represented as:
[0071]
[0072] where, represents the denoised image generated by DGGNet. To measure the similarity between and the real image y, we use the l1 loss (absolute error loss) for constraint, and its mathematical expression is as follows:
[0073]
[0074] The optimization objective of this loss function is to minimize the pixel-level error between the denoised image output by the network and the real noise-free image, thereby improving the denoising effect and avoiding the possible blurring effect of the l2 loss. Compared with the l2 loss, the l1 loss has stronger robustness and can effectively reduce artifacts in the edge region.
[0075] Using the gradient branch loss is used to constrain the difference between the gradient result generated by the model and the real gradient e. The true value of the gradient information is calculated by the Sobel filter, which can effectively extract the edge features of the image. The mathematical expression of the gradient branch loss is as follows:
[0076]
[0077] Among them, represents the gradient result obtained by network prediction, and e represents the true gradient information calculated by the Sobel filter.
[0078] Total Variation (TV) regularization makes the denoised image smoother while retaining key texture details by constraining the gradient change of the image. The mathematical expression of the fractional-order total variation loss is as follows:
[0079]
[0080] Among them, and respectively represent the fractional-order gradients of the denoised image in the x-axis and y-axis directions.
[0081] To comprehensively optimize the performance of DGGNet in denoising, edge information preservation, and overall image quality, our final loss function is a weighted combination of the above three loss functions, that is:
[0082]
[0083] Among them, λ1, λ2, and λ3 are the corresponding weight hyperparameters used to control the contribution of different loss terms to the total loss. These weights can be tuned through experiments to ensure the optimal performance of the network in different application scenarios.
[0084] The running results using this model are as shown in Figure 6 , Figure 7 , Figure 8 .
[0085] The object detection loss L detect is implemented through the EloU loss function. The total EloU loss function L detect = L IoU + L dis + L asp , and EloU is specifically expressed as: Among them, ρ is the Euclidean distance (L2 norm), b is the center coordinate of the predicted bounding box, w is the width of the predicted bounding box, h is the height of the predicted bounding box, b gt is the center coordinate of the ground truth bounding box, w gt is the width of the ground truth bounding box, h gt is the height of the ground truth bounding box, w c is the width of the smallest enclosing box that surrounds the predicted bounding box and the ground truth bounding box, h cis the height of the smallest enclosing box that surrounds the predicted box and the ground truth box; The intermediate features of DGGNet and the denoising results are fused with the feature layers of YOLO through skip connections. The denoising network and the detection network are optimized through a joint training method, and the alternating direction optimization strategy ADMM is used to iteratively solve the loss function L total for iterative solution.
[0086] The alternating direction optimization strategy ADMM is used to iteratively solve the loss function L total The process of iterative solution is as follows. First, the input image data is represented as a combination of a sparse noise matrix and a low-rank structure. The sparse noise matrix Q is optimized through the (2,1) norm, where i is the row index of the image, j is the column index of the image, and Q ij is the sparse noise matrix. The decomposition of the low-rank matrix and the sparse noise matrix is iteratively solved through ADMM, and finally a low-noise image structure is generated and fused with the features output by the DAE-DenseNet network to separate the high-frequency noise and the low-rank structure of the image.
[0087] Mathematically, the above calculation process can be expressed as:
[0088] Decompose the input data matrix D subject to D = P + Q, where P is the low-rank matrix and Q is the sparse noise matrix. The alternating direction multiplier method is used to iteratively solve the objective function where λ is the trade-off parameter. Construct the Lagrangian function where β is the penalty parameter, F is the Frobenius norm, and update the low-rank matrix where β k is the penalty parameter of the current iteration step, Q k is the sparse noise matrix of the current iteration, M k is the Lagrange multiplier of the current iteration. The solution of P is obtained by performing singular value decomposition (SVD) on the matrix Update the sparse noise matrix where the solution of Q is achieved through a soft thresholding operation. The soft thresholding formula is: where (H) ij is the element of the H matrix, μ is the soft threshold, H is the current matrix, and update the Lagrange multiplier M: M k+1 = M k + β k (D - P k+1 - Q k+1 ).
[0089] This application also introduces dynamic capture processing of the secondary matching strategy for high and low score detection boxes to optimize the object tracking effect of the model. The specific implementation steps are as follows:
[0090] First, according to the output of the object detector, confidence screening is performed on the detection boxes. Two thresholds are set, and based on this, the detection results are divided into a high score detection box set D high and a low score detection box set D low . The former contains object detection results with higher confidence, while the latter covers those detection boxes with lower confidence but still may be valid objects. The system uses the Data Association method to perform the first match between the high score detection box set D high and the existing trajectory set. During this matching process, the intersection over union (IoU) between all high score detection boxes and the current trajectory set is calculated to construct an IoU distance matrix, and the Hungarian algorithm is used to solve the optimal matching relationship. For successfully matched trajectories, the state is updated using the Kalman Filter and added to the current frame trajectory set to ensure the continuity of the object state. At the same time, trajectories that fail to be successfully matched are retained in the unmatched trajectory set T remain , and unmatched high score detection boxes are stored in the unmatched detection box set D remain for subsequent processing.
[0091] After the initial matching is completed, the system enters the second matching stage, and a secondary matching is performed for the trajectory set T remain and the low score detection box set D low where no corresponding objects were found in the first round of matching. Similar to the first matching, this step also calculates the IoU distance matrix and uses the Hungarian algorithm for object allocation. For successfully matched trajectories, the state is updated again using the Kalman filter and incorporated into the current frame trajectory set to compensate for the temporary loss of objects caused by problems such as object occlusion and detector confidence fluctuations. For trajectories that fail to be successfully matched, they are stored in the lost trajectory set T lost , and a lost frame counter is set to attempt matching when the object reappears. At the same time, unmatched low score detection boxes are directly discarded due to their low confidence and not included in subsequent calculations to reduce the interference of misdetected objects on trajectory management.
[0092] To further enhance the stability of object tracking, this method also introduces a new trajectory creation mechanism to ensure that new objects can be promptly incorporated into the tracking system. After the second matching, if the set D of unmatched high score detection boxes remainIf the confidence of the detection box in it is higher than the set tracking score threshold, it is considered that the detection box may be a new target. Therefore, a new trajectory is created and merged into the trajectory set of the current frame to ensure that the new target can be effectively tracked.
[0093] In addition, to prevent trajectory breakage caused by the temporary loss of the target, this method sets a trajectory loss handling mechanism. Specifically, for the trajectories stored in the lost trajectory set, a 30-frame retention window is set. That is, within 30 frames, the trajectory is still stored in the system so that when the target reappears, matching recovery can be performed. If the target cannot be detected within 30 frames, the system will determine that the target has disappeared permanently, and thus delete the trajectory to avoid invalid trajectories affecting subsequent calculations.
[0094] This application also designs a hyperparameter optimization mechanism based on GridSearch for the visual inspection detection framework that improves YOLOv8 to further enhance its performance in complex object detection tasks. This method tunes around two key hyperparameters: the input image size and the batch size, explores the impact of different parameter combinations on the model detection accuracy, and selects the optimal configuration to improve the final detection effect.
[0095] The size of the input image directly affects the computational burden and detection accuracy of the model. A larger image size can usually provide richer detail information, which helps to improve the accuracy of object detection, but at the same time increases the computational overhead, resulting in a decrease in training and inference speed. To balance detection accuracy and computational efficiency, we considered three different image sizes during the hyperparameter optimization process: Size List = [240, 480, 640]. Among them, 240 represents a lower resolution and is suitable for scenarios with limited computing resources; 480 is used as a medium resolution to achieve a balance between accuracy and computational efficiency; 640 is used as a higher resolution and is suitable for detection tasks with high-precision requirements. The batch size determines the number of samples input to the model during each iteration and directly affects the training speed and convergence stability of the model. A larger batch size can improve training stability and reduce gradient fluctuations, but at the same time requires more video memory resources. A smaller batch size is suitable for environments with limited video memory, but may lead to unstable gradient updates during training. To comprehensively evaluate the impact of different batch sizes on the model performance, we selected five different batch sizes for testing during the search process:
[0096] Train using an annotation dataset in YOLO format, which is loaded through a.yaml configuration file. Ensure that the dataset contains diverse target categories and complex background environments to improve the model's generalization ability. Initialize different model versions of YOLOv8, including the lightweight version (YOLOv8n), the standard version (YOLOv8s), the medium-scale version (YOLOv8m), the large-scale version (YOLOv8l), and the strongest version (YOLOv8x). Set the hyperparameter search space, including SizeList and BatchList. Hyperparameter Optimization uses nested loops to iterate through different YOLOv8 versions, input image sizes, and batch sizes for training. After each training session, record the model's performance metrics on the validation set, including mean average precision (mAP50-95) and inference time. After all training is completed, select the parameter combination with the best performance based on the mAP(50-95) metric. Evaluate the final model using the test set to verify its generalization ability. Export the final optimized model (Model Export), export the model weights with the optimal configuration, and use them for object detection tasks in actual application scenarios.
[0097] In this study, since the existing detection datasets were not publicly available, before model development and training, data collection and preprocessing were first carried out to ensure the integrity and applicability of the experimental data. The data collection and processing process includes multiple steps. First, a Raspberry Pi (RPi) computing device was installed and connected to a camera to capture in real time during the process. In this way, complete process data can be obtained, ensuring the temporal continuity and hierarchical integrity of the data, thus providing high-quality training samples for subsequent detection.
[0098] In the data collection stage, to optimize the temporal information of the dataset, we adopted a strategy of sampling once per second, extracting single-frame images from each video to ensure that there is corresponding visual information for each process, thus constructing a more comprehensive fault detection sample set.
[0099] After data sampling and preliminary screening, we further used the Keras Image Generator of TensorFlow for data augmentation to expand the diversity of the dataset and enhance the model's generalization ability. The data augmentation methods include operations such as rotation, cropping, and flipping to simulate uncertainties in different angles, lighting changes, and printing environments.
[0100] In the data annotation stage, we use Label Studio for manual annotation to ensure the accuracy and consistency of the data. All annotation information is stored in XML format, and each image is linked to the corresponding XML file to clarify the bounding box information of the target, thus meeting the requirements of the deep learning model for the data format. After completing the data annotation, we divide the dataset into a training set, a validation set, and a test set according to a reasonable ratio. Among them, 70% of the data is used for training, 10% of the data is used for model validation, and the remaining 20% is used for testing to ensure that the performance of the model on different datasets can be comprehensively evaluated.
[0101] After completing the collation of the dataset, to further improve the detection ability of the model, this application uses a hyperparameter optimization mechanism to tune the improved YOLOv8 object detection network. During the optimization process, the grid search method is used to systematically explore the input image size and batch size to find the optimal hyperparameter combination. Specifically, in the selection of the input image size, we set three different options, namely 240, 480, and 640 pixels, to evaluate the impact of different image resolutions on the model performance. To ensure the stability of training and the effective use of computing resources, we set five different values for the batch size, namely 4, 8, 12, 16, and 32, so as to analyze the impact of different batch sizes on the training speed, convergence, and final detection performance.
[0102] The entire hyperparameter optimization process is carried out through multiple rounds of iteration. During each iteration, different YOLOv8 model versions (YOLOv8n, YOLOv8s, YOLOv8m, YOLOv8l, YOLOv8x) are trained with different combinations of image sizes and batch sizes, and the performance evaluation results after training are stored in a dictionary data structure (evalList) for subsequent analysis.
[0103] After all possible hyperparameter combinations have been trained, we select the hyperparameter combination with the best mAP value as the final optimization configuration based on the mean Average Precision (mAP 50-95) metric of the model on the validation set, and use this optimal parameter to evaluate the test set. Finally, the trained model weights are exported for subsequent detection applications.
[0104] This application also presents a pest and disease detection system based on YOLOv8, deploying the YOLOv8 model after training and evaluation to an embedded system environment. The entire deployment architecture consists of three core components, namely a vision edge device installed with Raspberry Pi, an edge server using NVIDIA Jetson Nano, and a local server using a personal computer (PC). The main task of Raspberry Pi is to be responsible for collecting video streams and transmitting them to the edge server for real-time object detection. This device uses Raspberry Pi 4 Model B and is connected to the vision module through a USB Type B interface, enabling Raspberry Pi to send control commands to the camera through Python scripts. When a fault is detected during the detection process, Raspberry Pi can pause the detection task to avoid detection failures.
[0105] As the edge server, Jetson Nano is responsible for receiving video frames from multiple cameras and performing fault detection using the optimized YOLOv8 model. This device uses the Jetson Nano Developer Kit - B01 development board, which has a built-in 128-core NVIDIA Maxwell GPU, a quad-core central processing unit (CPU), and 4GB of RAM, giving it powerful parallel computing capabilities and enabling it to efficiently execute deep learning inference tasks.
[0106] During the operation of the system, Jetson Nano continuously receives frame images from multiple cameras and inputs them into the improved YOLOv8 model for analysis to determine whether there are abnormalities in the current process and, if necessary, update the weights of the detection model to adapt to new fault modes. Finally, as the control center of the entire system, the local server uses a personal computer running the Linux Ubuntu operating system. The main functions of this server include training the YOLOv8 model, pushing new model weights to the edge server, and real-time monitoring of the working status. In terms of system communication, all devices are connected through sockets (Socket) and use specific IP addresses and ports for data transmission to ensure the stability and real-time performance of the system. As Figure 2 shown.
[0107] In this embodiment, for the edge device deployment method, the controller interacts with the edge server through the following data streams, d train for training the fault detection model; monitoring system data real-time monitoring of fault detection and system status; edge communication data For data exchange between the controller and the edge server. The edge server includes a model retrieval module and a fault detection module. The model retrieval module is used to retrieve the pre-trained model from the controller; the fault detection module uses Ultralytics YOLO and OpenCV technologies to perform fault detection on the collected data; the edge server uses the Jetson Nano device labeled E s The Raspberry Pi and the camera module are used to capture image data and video streams during the detection process; the image data and video streams are transmitted to the edge server through the data transmission channel. The data channel includes the first channel data The nth channel data through the data channel to interact with the device and give the processed results to the device. The system supports the Docker virtualization deployment environment, and the controller runs the Ubuntu operating system. The edge server can perform the dynamic optimization process of model retrieval and fault detection functions.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A pest and disease detection method based on YOLOv8, characterized in that: It includes the following steps: S1: Construct a dataset and perform preprocessing; S2: Construct an object detection network model that improves YOLOv8. The model includes a backbone network, a neck network, and multiple detection heads. The backbone network adopts a fusion architecture of DenseNet and ResNet, and contains variants of the C2f module, namely the C2f-OREPA module, the C2f-SPD module, and the C2f-DBB module. The neck network realizes dynamic feature sampling and splicing through Dysample and Concat, and the detection head adopts the design of a diversified branch block; And in the model, a convolutional module that fuses local feature capture, global context modeling, and dynamic feature modulation, and an attention mechanism that fuses feature dynamic modulation and adaptive enhancement are adopted. At the same time, a denoising network is combined on the basis of DenseNet to construct a DAE-DenseNet network; S3: Model training: Use a loss function that fuses the denoising result with the feature layer through skip connections to train the model; S4: Model inference: Use ADMM to iteratively solve the loss function, and fuse the solution result with the features output by the DAE-DenseNet network to output the target tracking detection box.
2. The pest and disease detection method based on YOLOv8 according to claim 1, characterized in that: The structure of the backbone network is as follows: The input data is processed by two Conv ROI modules and then undergoes feature enhancement through the C2f-OREPA module. Then, the SPD module converts the spatial information into depth information. After further feature extraction through the Conv ROI module, it undergoes feature processing through the C2f-SPD module. After being processed by the SPD module, Conv ROI module, C2f-SPD module, SPD module, and Conv ROI module, the extracted features are input into the C2f-DBB module for further feature extraction and fusion. Finally, multi-scale feature extraction is performed through the SPPF module. The outputs of the two C2f-SPD modules and the SPPF module are all connected to the neck network for dynamic feature extraction and splicing.
3. The pest and disease detection method based on YOLOv8 according to claim 1, characterized in that: The structure of the convolutional module that integrates local feature capture, global context modeling, and dynamic feature modulation is as follows: The input original image data is segmented into multiple small patches through the patch embedding module for embedding processing; the feature map after embedding processing undergoes non-linear transformation through the activation function and then feature standardization through batch normalization; the input feature map is segmented into a series of 3×3 non-overlapping receptive field sliders; group convolution operations are performed on the segmented sliders to efficiently extract the features of each local region and expand the feature set; the feature map generated by group convolution matches the dimension of the input feature map to form a preliminary spatial feature representation; global average pooling is used to integrate the local features to obtain global context information, reducing the number of parameters and computational complexity; an attention weight map is generated through the Softmax function to weight the features of the receptive field sliders, emphasizing the feature information of the key regions; the weighted attention map is multiplied element-wise with the spatial feature map to generate the attention-modulated weighted feature representation; the feature map input to the Conv ROI module is divided into three different paths and independently processed through Patch = 6, Patch = 7, and Patch = 8 in the CSMM module; the feature maps output from the three paths are fused through an addition operation to integrate multi-scale feature information; the fused feature map undergoes global average pooling to further extract global context features; The first fully connected layer of the multi-layer perceptron is used to generate intermediate features; The second fully connected layer outputs the final aggregated features; through the channel expansion operation, the number of channels of the aggregated features is expanded; the expanded features are multiplied element-wise with the input feature map to achieve the fusion of the input and the channel expansion output; The visual context information of the feature map is encoded at multiple scales through sequential depth convolutional layers; A gating operation is performed on the multi-path features output by the CSMM module and the context focus result to selectively aggregate context information; Through linear affine transformation, the modulator is applied to each feature element, and the optimized feature representation is added element-wise to the original input feature map to form a residual structure, retaining key information and enhancing training stability.
4. A pest and disease detection method based on YOLOv8 according to claim 1, characterized in that: The attention mechanism that fuses feature dynamic modulation and adaptive enhancement realizes feature enhancement through a three-channel branch, including the X channel, Y channel of the original coordinate attention, and the global average pooling channel. The input feature map is averaged in the height and width directions to obtain the feature maps of the X channel and Y channel respectively. Convolution operations are performed on the feature maps of the X channel and Y channel to extract fine-grained features in the height and width directions. The feature maps of the X channel and Y channel are activated through the sigmoid function to obtain weighted results, and are multiplied element-wise with the original input feature map respectively to strengthen the feature information. A third-channel branch is added, and global average pooling is used to perform overall downsampling on the input feature map, retaining the global background information and fusing it in parallel with the X and Y channels to achieve multi-path feature enhancement; visual context information is encoded at different levels through sequential depth convolution layers, and spatial features of height and width are collected from different scales to enhance the understanding of details and overall structures; gating operations are performed on the feature maps of the X channel, Y channel, and global pooling channel of the multi-path coordinate attention mechanism to filter and aggregate context information, giving priority to retaining key details and suppressing redundant background information. The gated and aggregated feature maps are applied to each query element through a linear affine transformation to achieve dynamic modulation and precise adjustment of features, improving the adaptability to complex visual content; through element-wise affine transformation, dynamic modulation and adaptive enhancement of features are realized.
5. A pest and disease detection method based on YOLOv8 according to claim 1, characterized in that: The C2f-DBB module is formed by replacing the convolution operation in the original Bottleneck module in the C2f module of the traditional YOLOv8 backbone network with the DBB module; the DBB module includes multiple branches of different scales, and each branch uses a 1×1, 1×1-k×k composite convolution and a 1×1-AVG combined structure; among them, the 1×1-k×k composite convolution reduces the computational complexity by first performing a 1×1 convolution, and then uses a k×k convolution to enhance the local receptive field, thereby improving the feature expression ability while ensuring computational efficiency; the 1×1-AVG combined branch reduces the dimension through a 1×1 convolution, and then uses average pooling to further enhance the feature smoothness to reduce noise interference and improve the robustness of the network; and the structure reparameterization technology is introduced to simplify the multi-branch structure in the DBB module into an equivalent standard convolution operation during inference.
6. The pest and disease detection method based on YOLOv8 according to claim 5, wherein: The C2f-OREPA module effectively reduces the computational and storage overhead during the training process through two key steps of block linearization and block compression, while maintaining the performance advantages in the inference stage. During the block linearization process, the multi-branch structure in the DBB module is merged through a linear transformation method, making it behave more like a single-branch convolutional network during the training process without losing the original expression ability; then, during the block compression process, the intermediate feature representation is optimized to reduce the computational amount and storage requirements, thereby achieving more efficient training.
7. The pest and disease detection method based on YOLOv8 according to claim 1, characterized in that: The loss function includes: denoising loss and object detection loss. The denoising loss includes image denoising loss, gradient branch loss, and fractional total variation loss. The object detection loss is implemented by the EloU loss function.
8. A pest and disease detection method based on YOLOv8 according to claim 7, characterized in that: The process of iteratively solving the loss function using the alternating direction optimization strategy ADMM is as follows. First, the input image data is represented as a combination of a sparse noise matrix and a low-rank structure. The sparse noise matrix is optimized by the (2,1) norm, and the decomposition of the low-rank matrix and the sparse noise matrix is solved iteratively by ADMM. Finally, a low-noise image structure is generated and fused with the features output by the DAE-DenseNet network to separate the high-frequency noise and low-rank structure of the image.
9. A pest and disease detection method based on YOLOv8 according to any one of claims 1-8, characterized in that: After the model outputs the target tracking detection box, a quadratic matching strategy for high- and low-score detection boxes is introduced to maintain the coherence of the target trajectory, including the following steps: Based on the detection results of the model, set a high-score threshold and a low-score threshold, divide the detection boxes into a high-score detection box set and a low-score detection box set, use a data association method to perform the first match on the high-score detection box set and the trajectory set, and update the successfully matched trajectories; perform a second match on the trajectory set that failed the first match and the low-score detection box set, update the successfully matched trajectories, and delete the unmatched low-score detection boxes; for the set of unmatched high-score detection boxes, if its confidence is greater than the tracking score threshold, create a new trajectory; for the set of lost trajectories, retain a set number of frames, perform a match when the trajectory reappears, and delete the trajectory if it does not appear within the set number of frames. Calculate the loU distance matrix between the high-score detection box set and the trajectory set; use the Hungarian algorithm to match the loU distance matrix. The successfully matched trajectories are updated by the Kalman filter and added to the current frame trajectory set; the unmatched trajectories are put into the unmatched trajectory set, and the unmatched high-score detection boxes are put into the unmatched detection box set.
10. A pest and disease detection method based on YOLOv8 according to claim 9, characterized in that: The second matching step includes calculating the loU distance matrix between the low-score detection box set and the unmatched trajectory set; using the Hungarian algorithm for matching. The successfully matched trajectories are updated by the Kalman filter and added to the current frame trajectory set; the unmatched trajectories are put into the lost trajectory set, and the unmatched low-score detection boxes are directly deleted; for the set of unmatched high-score detection boxes, if its confidence is higher than the set tracking score threshold, create a new trajectory and merge it into the current frame trajectory set; the trajectories in the lost trajectory set are retained for a set number of frames, perform a match when the trajectory reappears within the set number of frames, otherwise delete the trajectory.
Citation Information
Cited By
Rice disease and insect pest target detection algorithm based on Mamba and YOLOv8
CN120472149A
Building structure component comparison method and system oriented to project supervision
CN121121643A
Fruit disease detection method and system based on global memory and local comparison
CN121354093A