Pest and disease multi-target detection method and device, equipment and storage medium

By adopting the multi-objective detection method of pest and disease based on the YOLOv8 network in tomato leaf pest and disease detection, combined with the CPCA attention mechanism and improved Slimneck structure, the problems of low detection accuracy and slow model detection are solved, and more efficient and accurate pest and disease detection are achieved.

CN119942047APending Publication Date: 2025-05-06WUHAN POLYTECHNIC UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510022947.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has low detection accuracy in tomato leaf pest detection, model performance is easily disturbed, and it is difficult to cope with different environmental conditions.

Method used

Using the multi-object detection method of pest and disease based on YOLOv8 network, an initial object detection model is constructed by pre-processing the initial image dataset, and a CPCA attention mechanism module is added to the backbone network. The neck layer network is replaced with an improved Slimneck structure, including Focalmodulation module, upsampling, GSConv and VoVGSCSP.

Benefits of technology

The accuracy and speed of tomato leaf pest detection is improved, the model's adaptability to different environmental conditions is enhanced, and the problem of low detection accuracy and slow model detection speed is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942047A_ABST
    Figure CN119942047A_ABST
Patent Text Reader

Abstract

The invention discloses a pest and disease damage multi-target detection method and device, equipment and a storage medium. The method comprises the steps that a to-be-detected image and an initial data set are acquired; performing image preprocessing on the initial data set to obtain a preprocessed data set; the method comprises the following steps: constructing an initial target detection model based on a YOLOv8 network, replacing a backbone network of the initial target detection model with an internet structure, adding a CPCA attention mechanism behind the backbone network, and replacing a neck layer network of the initial target detection model with improved Slimcheck to obtain an improved target detection model; training the improved target detection model through the preprocessed data set to obtain an optimized target detection model; and detecting a to-be-detected image according to the optimized target detection model to obtain a detection result. According to the embodiment of the invention, the problems of low current target detection precision and slow model detection speed are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for multi-target detection of pests and diseases, equipment and storage medium. Background Art

[0002] When dealing with the impact of tomato pests and diseases, traditional pest control usually relies on the use of large amounts of pesticides, such as regular spraying of fungicides, removal of diseased leaves, and reasonable fertilization and irrigation to prevent tomato leaf pests and diseases. This not only increases the cost of planting, but also causes soil and water resource pollution, reduces soil fertility, and kills non-target beneficial insects. In addition, with the aging of the population and the shortage of agricultural labor, the detection and prevention of tomato leaf pests and diseases have been affected. If tomato pests and diseases cannot be effectively controlled, it will lead to a large-scale reduction in tomato production and affect farmers' income.

[0003] Pest and disease detection based on machine learning algorithms relies on manually extracted feature information, which makes it difficult to deal with pests and diseases in different environments. Although target detection algorithms based on deep learning technology can achieve high accuracy, they are currently limited by several key issues, such as complex network architecture, large number of parameters, and dependence on high-performance GPU hardware. Therefore, in real life, it is very necessary to improve the existing tomato leaf pest and disease detection methods. Summary of the invention

[0004] The main purpose of the present invention is to provide a method, device, equipment and storage medium for detecting tomato leaf diseases and insect pests, aiming to solve the technical problems of low detection accuracy and susceptibility of model performance to interference in the prior art.

[0005] To achieve the above object, the present invention provides a multi-target detection method for pests and diseases, the method comprising the following steps:

[0006] Obtain an initial image dataset of the target to be detected;

[0007] Performing image preprocessing on the initial image data set to obtain a preprocessed data set;

[0008] An initial target detection model is constructed based on the YOLOv8 network. The backbone network of the initial target detection model is a fasternet structure, and a CPCA attention mechanism module is added after the last C2f of the backbone network. The fasternet structure includes an embedding layer, a convolution layer, a merging layer, a pooling layer and a CPCA attention mechanism. The convolution layer includes a partial convolution and a one-dimensional convolution.

[0009] The neck layer network of the initial target detection model is an improved Slimneck structure, and an improved target detection model is obtained, wherein the improved Slimneck structure is originally constructed based on the Slimneck structure and the spatial pyramid pooling module of the Slimneck structure is replaced by a Focalmodulation module; including the Focalmodulation module, upsampling, GSConv and VoVGSCSP.

[0010] Using the preprocessed data set to train the initial target detection model to obtain an optimized target detection model;

[0011] The optimized target detection model is used to detect the target image to be detected to obtain a detection result.

[0012] Optionally, performing image preprocessing on the initial image data set to obtain a preprocessed data set includes:

[0013] Converting the original labels of the initial image data set into target labels to obtain a target data set;

[0014] The target data set is amplified by geometric transformation and color transformation to obtain a preprocessed enhanced mark data set, wherein the geometric transformation includes at least one of flipping, rotating, cropping, exposing, adding noise and blurring, and the color transformation includes at least one of color transformation, erasing, mosaic and filling.

[0015] Optionally, the applying the preprocessed data set to train the initial target detection model to obtain an optimized target detection model comprises:

[0016] Input the images in the preprocessed dataset into the improved FasterNet structure of the initial target detection model for feature extraction to obtain a reference feature map, wherein the FasterNet structure includes an embedding layer, a convolution layer, a merging layer, a pooling layer and a CPCA attention mechanism, and the convolution layer includes a partial convolution and a one-dimensional convolution;

[0017] Using CPCA attention mechanism, spatial information is aggregated from the target feature map through average pooling and maximum pooling operations to produce two independent spatial context descriptors;

[0018] Inputting two independent context descriptors into a shared multi-layer perceptron, combining the outputs of the shared MLP by element-wise summation to obtain a channel weight vector, wherein the shared MLP consists of a single hidden layer;

[0019] The channel weight vector is used to multiply each channel of the original feature map element by element, to enhance the characteristics of the target channel, and to suppress the output feature map of the non-target channel, so as to obtain feature maps of different scales;

[0020] Inputting the reference feature map into the neck layer network improved by Slimneck for fusion to obtain a target feature map;

[0021] Inputting the target feature map into the head network of the improved initial target detection model for detection, and outputting a marking box and a classification label;

[0022] Determine a target loss function based on the marked box and the classification label;

[0023] The parameters of the improved initial target detection model are optimized by the target loss function to obtain an optimized target detection model.

[0024] Optionally, the step of inputting the preprocessed image in the dataset into the FasterNet structure of the target detection model for feature extraction to obtain a reference feature map comprises:

[0025] Perform feature extraction on the images in the preprocessed data set to obtain the original feature map;

[0026] The original feature map is further convolved to extract features and obtain a number of sub-feature maps;

[0027] Connect the plurality of sub-feature maps according to the channel dimension to obtain an intermediate feature map;

[0028] The intermediate feature map is further input to the merging layer for feature merging, and then part of the convolution is repeatedly used for convolution to obtain a convolution feature map;

[0029] The convolutional feature map is input into the pooling layer and the fully connected layer of the fasternet for aggregation processing to obtain a reference feature map.

[0030] Using CPCA attention mechanism, spatial information is aggregated from the target feature map through average pooling and maximum pooling operations to produce two independent spatial context descriptors;

[0031] Inputting two independent context descriptors into a shared multi-layer perceptron, combining the outputs of the shared MLP by element-wise summation to obtain a channel weight vector, wherein the shared MLP consists of a single hidden layer;

[0032] The channel weight vector is used to multiply each channel of the original feature map element by element to enhance the features of the target channel and suppress the output feature maps of non-target channels to obtain feature maps of different scales.

[0033] Optionally, the step of inputting the reference feature map into the neck layer Slimneck network for fusion to obtain a target feature map comprises:

[0034] Finally, the reference feature map is input into the pyramid space fast pooling layer improved by focal modulation to extract feature information.

[0035] The feature maps of different scales are used as input, convolved through the GSConv module, fused by upsampling and concatenation, and enhanced by the GSConv module again;

[0036] The VoV-GSCSP module is used to perform cross-stage feature fusion and provide optimized feature maps for the head network.

[0037] Optionally, determining a target loss function according to the marked box and the classification label includes:

[0038] The position loss function is obtained by calculating based on the marked box and the real box;

[0039] Calculate according to the classification label and the preset label to obtain the classification loss function;

[0040] A target loss function is determined according to the position loss function and the classification loss function.

[0041] In addition, to achieve the above object, the present invention further provides a target detection device, the target detection device comprising:

[0042] An acquisition module is used to acquire the image to be detected and the initial data set;

[0043] A processing module, used for performing image preprocessing on the initial data set to obtain a preprocessed data set;

[0044] A construction module is used to construct an initial target detection model based on a YOLOv8 network. The backbone network of the initial target detection model is a fasternet structure, and a CPCA attention mechanism module is added after the last C2f of the backbone network. The fasternet structure includes an embedding layer, a convolution layer, a merging layer, a pooling layer, and a CPCA attention mechanism. The convolution layer includes a partial convolution and a one-dimensional convolution.

[0045] The neck layer network of the initial target detection model is an improved Slimneck structure, and an improved target detection model is obtained, wherein the improved Slimneck structure is constructed based on the original Slimneck structure and the spatial pyramid pooling module of the original Slimneck structure is replaced by a Focalmodulation module; including the Focalmodulation module, upsampling, GSConv and VoVGSCSP.

[0046] A training module, used to train the improved target detection model using the preprocessed data set to obtain an optimized target detection model;

[0047] The detection module is used to detect the image to be detected according to the optimized target detection model to obtain a detection result.

[0048] In addition, to achieve the above-mentioned purpose, the present invention also proposes a target detection device, which includes: a memory, a processor, and a target detection program stored in the memory and executable on the processor, and the target detection program is configured to implement the steps of the target detection method described above.

[0049] In addition, to achieve the above-mentioned purpose, the present invention further proposes a storage medium, on which a target detection program is stored, and when the target detection program is executed by a processor, the steps of the target detection method described above are implemented.

[0050] The present invention obtains an image to be detected and an initial data set; performs image preprocessing on the initial data set to obtain a preprocessed data set; constructs an initial target detection model based on a YOLOv8 network, replaces the backbone network of the initial target detection model with a fasternet structure, adds a CPCA attention mechanism to the backbone network of the initial target detection model, and replaces the neck layer network of the initial target detection model with an improved Slimneck to obtain an improved target detection model; trains the improved target detection model with the preprocessed data set to obtain an optimized target detection model; detects the image to be detected according to the optimized target detection model to obtain a detection result, and in the above manner, adds the fasternet structure, the CPCA attention mechanism and the improved Slimneck neck network structure to the initial target detection model to improve the model structure, and completes the optimization of the detection model by training the improved model, thereby solving the problems of low detection accuracy and slow model detection speed of tomato leaf diseases and pests, and improving the detection accuracy of tomato leaf diseases and pests. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1It is a schematic diagram of the structure of a target detection device in a hardware operating environment involved in an embodiment of the present invention;

[0052] Figure 2 It is a flowchart of the first embodiment of the target detection method of the present invention;

[0053] Figure 3 This is a structural diagram of an improved target detection model in the second embodiment of the target detection method of the present invention;

[0054] Figure 4 A schematic diagram of a flow chart of a second embodiment of a target detection method of the present invention;

[0055] Figure 5 This is a schematic diagram of the operation of FastNet in the second embodiment of the target detection method of the present invention;

[0056] Figure 6 This is a flowchart of CPCA channel attention implementation in the second embodiment of the target detection method of the present invention;

[0057] Figure 7 This is a flowchart of CPCA spatial attention implementation in the second embodiment of the target detection method of the present invention;

[0058] Figure 8 This is a schematic diagram of the specific structure of Focalmodulation in the second embodiment of the target detection method of the present invention;

[0059] Fig. 9 Schematic diagram of the improved Slimneck network in the second embodiment of the target detection method of the present invention;

[0060] Fig.10 It is a structural block diagram of the first embodiment of the target detection device of the present invention. DETAILED DESCRIPTION

[0061] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings.

[0062] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0063] Reference Figure 1 , Figure 1 The figure is a schematic diagram of the structure of a target detection device in the hardware operating environment involved in the embodiment of the present invention.

[0064] like Figure 1As shown, the target detection device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (Wireless-Fidelity, Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk storage. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0065] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the target detection device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0066] like Figure 1 As shown, the memory 1005 as a storage medium may include an operating system program, a network communication module, a user interface module, and a target detection program.

[0067] exist Figure 1 In the target detection device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the target detection device of the present invention can be set in the target detection device, and the target detection device calls the target detection program stored in the memory 1005 through the processor 1001, and executes the target detection method provided by the embodiment of the present invention.

[0068] The embodiment of the present invention introduces a target detection method, the specific process is as follows Figure 2 As shown in FIG. 1 , this figure is a flow chart of the first embodiment of the target detection method. In this embodiment, Figure 2 The target detection method mainly includes the following steps.

[0069] Step S10: Acquire the image to be detected and the initial data set.

[0070] It should be noted that the executor of this embodiment is a target detection device, and may also be other devices that can achieve the same or similar functions. This embodiment is not limited to this, and this embodiment is described by taking the target detection device as an example.

[0071] It can be understood that the image to be detected is a tomato leaf image that needs to be detected for pests and diseases. The object to be detected in the image to be detected is the pests and diseases of the tomato leaves. There are many kinds of pests and diseases on tomato leaves. This example selects ten different kinds of pests and diseases as the data set.

[0072] It is worth noting that a data set is a set consisting of data samples, and the initial data set is a set consisting of images to be detected. The initial data set may be a PlantVillage data set, and this embodiment does not impose any specific limitation on this.

[0073] Step S20: performing image preprocessing on the initial data set to obtain a preprocessed data set.

[0074] It should be noted that since deep learning requires learning through a large number of image samples, preprocessing the images in the initial data set can expand the data set, obtain richer data set images for model training, and improve the model's generalization ability.

[0075] Furthermore, in order to improve the accuracy of the detection model, the step S20 includes: converting the original labels of the initial data set into target labels to obtain a target data set; amplifying the target data set through geometric transformation and color transformation to obtain a preprocessed tomato leaf pest and disease data set, wherein the geometric transformation includes at least one of flipping, rotating, cropping, exposing, adding noise and blurring, and the color transformation includes at least one of color transformation, erasing, mosaic and filling.

[0076] It should be noted that the images in the initial data set all have original labels, which include: the horizontal coordinate of the upper left corner of the annotation box<bbox_left> , the vertical coordinate of the upper left corner of the annotation box<bbox_top> , Label box width<bbox_width> , Label box height<bbox_height> ,score <score>, target category<object_category> , cutoff rate <truncation>And the occlusion rate <occlusion>Etc., this embodiment does not impose any specific limitation on this.

[0077] It can be understood that the original labels of the initial data set are converted into label formats, and the original labels are converted into target labels. The target labels include: categories <c>, horizontal coordinate of the center of the label box <x>, the vertical coordinate of the center of the label box <y>, the annotation box is relatively wide <w>, the annotation box is relatively high <h>Etc., this embodiment does not impose any specific limitation on this.

[0078] In the specific implementation, the data augmentation includes geometric transformation and color transformation operations. The geometric transformation includes flipping, rotation, cropping, exposure, noise addition and blurring. The color transformation includes color transformation, erasing, mosaic and filling. Flipping includes horizontal flipping and vertical flipping. Noise addition is to add Gaussian noise. Mosaic is obtained by splicing the pictures of this data set, which can improve the robustness and generalization ability of the model.

[0079] Step S30: construct an initial target detection model based on the YOLOv8 network, replace the backbone network of the initial target detection model with a fasternet structure, replace the neck layer network of the initial target detection model with an improved Slimneck, and add a CPCA attention mechanism to the neck network to obtain an improved target detection model. The improved Slimneck structure is constructed based on the Slimneck structure and the spatial pyramid pooling module of the Slimneck structure is replaced by a Focalmodulation module;

[0080] It should be noted that the initial target detection model is structurally improved by replacing the backbone network with a fasternet structure, and the backbone network is used for feature extraction; the neck network of the initial target detection model is replaced with an improved Slimneck structure, and the neck network is used for feature fusion, wherein the improved Slimneck structure is constructed based on the Slimneck structure and the spatial pyramid pooling module of the Slimneck structure is replaced by a Focal modulation module.

[0081] like Figure 3 As shown, Figure 3 This is a structural diagram of an improved target detection model in the target detection method of this embodiment. The improved initial target detection model includes some convolution blocks, a C2f module, a CPCA attention mechanism, an upsampling layer upsample, an improved spatial pyramid pooling module layer Focalmodulation (replacing the original spatial pyramid pooling module with focalmodulation), a feature fusion layer Concat, a GSconv convolution, a VoVGSCSP module, a CBS module, a convolution layer, and a detection head Detect.

[0082] It is understandable that the CPCA attention mechanism module is added to the backbone network of the initial target detection model. The neck layer network is used for feature fusion. The CPCA attention mechanism module aims to enhance important features by dynamically adjusting the weights of each channel in the feature map and pruning channels in some cases to improve the efficiency of the model.

[0083] It can be understood that the neck network of the initial target detection model is replaced with the improved Slimneck, the spatial pyramid pooling module of the neck network is first replaced with a Focal modulation module, and the feature maps of different scales obtained are input into the Focalmodulation module for aggregation processing to obtain a reference feature map.

[0084] The embodiments of the present invention utilize the redundancy in feature maps and systematically apply regular convolution to some input channels for spatial feature extraction, while the remaining channels remain unchanged, thereby reducing redundant calculations and simultaneous memory accesses, so that partial convolution can more effectively extract spatial features.

[0085] Then replace the CBS of the neck network with GSconv, C2f with VoVGSCSP, and the neck network is used for feature fusion. Replacing the neck network with Slimneck can reduce the computational overhead while maintaining accuracy, which is suitable for real-time and resource-constrained environments. And through additional convolution, upsampling and downsampling operations, the features are enhanced to have stronger expression capabilities, thereby improving the positioning and classification performance of the model.

[0086] Step S40: training the improved target detection model using the preprocessed data set to obtain an optimized target detection model.

[0087] It should be noted that the improved target detection model is trained with the preprocessed data set, the loss function is calculated according to the training results, and the model parameters are optimized through the loss function to improve the model performance and obtain the optimized target detection model.

[0088] Step S50: Detect the image to be detected according to the optimized target detection model to obtain a detection result.

[0089] The image to be detected is detected according to the optimized target detection model, and the classification label and positioning box of the object to be detected in the detection image are output.

[0090] This embodiment obtains an image to be detected and an initial data set; performs image preprocessing on the initial data set to obtain a preprocessed data set; constructs an initial target detection model based on a YOLOv8 network, replaces the backbone network of the initial target detection model with a fasternet structure, adds a CPCA attention mechanism to the backbone network, replaces the neck layer network of the initial target detection model with an improved Slimneck, and obtains an improved target detection model; trains the improved target detection model with the preprocessed data set to obtain an optimized target detection model; and detects the image to be detected according to the optimized target detection model to obtain a detection result.

[0091] Through the above method, the fasternet structure, CPCA attention mechanism and improved Slimneck are added to the initial target detection model to improve the model structure, and the detection model optimization is completed by training the improved model, which solves the current problems of low accuracy of tomato leaf disease and pest detection and slow model detection speed, and improves the accuracy of tomato leaf disease and pest detection.

[0092] refer to Figure 4 , Figure 4 The flowchart of the second embodiment of the target detection method of the present invention is shown in FIG. Based on the first embodiment, the second embodiment is the steps in the target detection method, including:

[0093] Step S401: Input the image in the preprocessed data set into the FasterNet structure of the improved target detection model for feature extraction to obtain a reference feature map, wherein the FasterNet structure is as follows: Figure 5 As shown, it includes an embedding layer, a convolution layer, a merging layer, a pooling layer and a CPCA attention mechanism, and the convolution layer includes partial convolution and one-dimensional convolution;

[0094] It should be noted that feature extraction through the FasterNet structure can reduce redundant calculations and storage accesses, extract spatial features more efficiently, and will not affect the accuracy of various visual tasks.

[0095] Furthermore, the step S401 includes:

[0096] Performing feature extraction on the images in the preprocessed data set to obtain an original feature map;

[0097] Performing data conversion on the original feature map through the embedding layer embedding module in the improved fasternet structure to obtain a certain number of features;

[0098] The sub-feature map is input to the fasternet block for partial convolution and one-dimensional convolution. Each fasternet block has a partial convolution layer followed by two point convolution layers. The partial convolution only applies convolution to some input channels for spatial feature extraction, and does not extract features for other channels. Therefore, the FLOPs of partial convolution is only Where h is the height of the feature, w is the width of the feature, k is the size of the convolution kernel, and c is the p is the number of feature channels extracted by partial convolution. At this time, the FLOPs of partial convolution is only 1 / 16 of that of regular convolution, and the memory access amount is At this time, the memory access of some convolutions is only 1 / 4 of that of regular convolutions, which further optimizes the cost of the entire object detection network and reduces computational redundancy and memory access;

[0099] After partial convolution processing, a convolution feature map is obtained, and the convolution feature map is input into the merging module for aggregation processing; the fasternet block and the merging module are repeatedly input to process the features.

[0100] Among them, CPCA adopts the channel-spatial attention mechanism:

[0101] CPCA first gives an intermediate feature map as input, and the channel attention module first derives a 1D channel attention map;

[0102] Then it is element-wise multiplied with the input feature F and the channel attention value is broadcasted along the spatial dimension to obtain the refined features with channel attention. The spatial attention module processes F c To generate a 3D spatial attention map;

[0103] The final output features are obtained by multiplying Ms by Fc element-wise.

[0104] Among them, the generation of the channel attention map is the responsibility of the channel attention module, which achieves this goal by exploring the inter-channel relationship in the features. The entire attention process can be summarized as:

[0105]

[0106] Specifically, Figure 6 As shown, spatial information is aggregated from the feature map using average pooling and maximum pooling operations. This aggregation process produces two independent spatial context descriptors. These descriptors are then input into a shared multi-layer perceptron (MLP). The outputs of the shared MLP are combined by element-wise summation to obtain the channel attention map. At the same time, in order to reduce parameter overhead, the shared MLP consists of a single hidden layer, where the size of the hidden activation is set to, where R represents the reduction ratio. The calculation of channel attention can be summarized as:

[0107] CA(F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))

[0108] Where σ represents the sigmoid function.

[0109] Figure 7 As shown in Figure 1, the spatial attention map is generated by extracting the spatial mapping relationship. CPCA changes the practice of enforcing consistency in the spatial attention module of each channel and chooses to dynamically allocate attention weights in two dimensions: channel and space. Figure 1 shows that spatial attention uses deep convolution to capture the spatial relationship between features, ensuring that the relationship between channels is preserved while reducing computational complexity. It uses a multi-scale structure to enhance the ability of convolution operations to capture spatial relationships.

[0110] Channel mixing is performed by the tail of the spatial attention module using 1×1 convolution to generate a more refined attention map. The spatial attention is calculated as follows:

[0111]

[0112] Among them, DwConv represents deep convolution, Branch i , i∈{0,1,2,3} represents the i-th branch.

[0113] Step S402: Input the reference feature map into the improved neck layer network of the improved target detection model for encoding to obtain a target feature map, wherein the neck layer network includes a spatial pyramid pooling module replaced by a focal modulation module and a slimneck structure.

[0114] First, the reference feature map is input into the improved neck layer network of the improved target detection model for encoding, and the feature information is extracted using the spatial pyramid pooling module improved by Focalmodulation to obtain the target feature map. The specific structure of Focalmodulation is as follows Figure 8 As shown in the figure. First, the focus modulation mechanism is used to replace the self-attention mechanism to capture the long-distance dependency and contextual information in the image, reducing computational redundancy. And the focus modulation is used to process the high-level image features extracted by the backbone network, reducing the internal interaction of low-level features.

[0115] First, the input features are processed through a lightweight linear layer, and then the spatial context at different granularity levels is encoded into summary tokens through a hierarchical context module. Then, the gated aggregation module is used to selectively aggregate information based on the content of the query, and finally the modulator interacts with the query to generate output. This change simplifies the query process by aggregating knowledge at the same granularity level. Compared with the unimproved YOLOv8s, it can still fuse multi-scale features while maintaining efficient inference speed and low computational cost, and can better capture long-distance dependencies.

[0116] Then, the neck structure of yolov8 is replaced with an improved Slimneck, wherein the improved Slimneck structure is constructed based on the original Slimneck structure and the spatial pyramid pooling module of the original Slimneck structure is replaced with a Focalmodulation module; including the Focalmodulation module, upsampling, GSConv and VoVGSCSP, and the specific structure is as follows Fig. 9 shown.

[0117] The convolution structure in the original yolov8 neck is replaced by the lightweight convolution method GSConv. The GSonn can not only retain the detailed information of the target like the standard convolution, but also combine the advantages of depth-separable convolution to reduce the computational complexity, and achieve the preservation of channel connections with lower time complexity, avoiding the problem of losing details when processing large-scale models.

[0118] Specific:

[0119] First, the input features are subjected to standard convolution to obtain a feature map with C2 / 2 feature channels. Then, after being processed by a depthwise separable convolution operation, a second feature map with C2 / 2 feature channels is obtained. The two processed feature maps are then concat-joined to obtain a feature map with the expected number of feature channels of C2. Finally, the feature map is output after a Shuffle operation.

[0120] Secondly, a one-time aggregated VoV-GSCSP module is used to replace the C2f layer in the neck network. The feature map of GSConv and the feature map of standard convolution are concat-concatenated, and then standard convolution is performed to reduce the computational complexity and the number of model parameters, thereby improving the inference speed of the model.

[0121] Step S403: input the target feature map into the head network of the improved target detection model for detection, and output a marking box and a classification label.

[0122] It should be noted that the label frame is a rectangular frame marking the position of the object to be detected in the input image, and the classification label is a label of the category information of the object to be detected in the input image.

[0123] Step S404: determining a target loss function according to the marking box and the classification label.

[0124] It should be noted that the role of the loss function is to measure the distance between the neural network prediction information and the expected information (label). The closer the predicted information is to the expected information, the smaller the loss function value is. In this embodiment, the loss includes classification loss and position loss, and the target loss function is obtained according to the classification loss and position loss.

[0125] Further, the model evaluation accuracy index is determined according to the marked frame and the classification label, including: calculating the overlap (IoU) according to the marked frame and the real frame; calculating the true positive (TP), false positive (FP) and false negative (FN) according to the overlap, specifically, the true positive indicates the number of detection frames with a predicted result overlap greater than 0.5, the false positive indicates the number of detection frames with a predicted result overlap less than 0.5, and the false negative is the number of detection frames without any detection or with a low confidence level of the detection result; finally, according to the several indicators; on this basis, we can calculate the Precision, Recall, and mAP mean average precision. Determine the target loss function with the classification loss function.

[0126] It should be noted that the target average accuracy is determined based on the overlap, and the indicators are as follows:

[0127]

[0128] Among them, AP is the average precision, AP i Represents the average accuracy of the i-th category of objects. mAP (mean average precision) is a commonly used comprehensive indicator in target detection, which represents the average accuracy of the model under different overlap thresholds.

[0129] Step S405: Optimizing the parameters of the improved target detection model through the target loss function to obtain an optimized target detection model.

[0130] It should be noted that the model parameters are updated through the target loss function until the confidence of the model prediction exceeds the preset value, then the training and updating are stopped, and the current model is used as the optimized target detection model.

[0131] In this embodiment, the image in the preprocessed data set is input into the fasternet structure of the improved target detection model for feature extraction to obtain a reference feature map, wherein the fasternet structure includes an embedding layer, a convolution layer, a merging layer, a pooling layer and a CPCA attention mechanism, and the convolution layer includes a partial convolution and a one-dimensional convolution; the reference feature map is input into the improved neck layer network of the improved target detection model for encoding to obtain a target feature map, wherein the improved neck layer network includes a slimneck structure obtained by replacing a spatial pyramid pooling module with a focal modulation module; the target feature map is input into the head network of the improved target detection model for detection, and a marking box and a classification label are output; a target loss function is determined according to the marking box and the classification label; the parameters of the improved target detection model are optimized by the target loss function to obtain an optimized target detection model. In the above manner, the improved target detection model is trained by the preprocessed data set, the loss function is calculated according to the training result, and the model parameters are optimized by the loss function to improve the model performance to obtain an optimized target detection model.

[0132] Reference Fig.10 , which is a structural block diagram of the first embodiment of the target detection device of the present invention.

[0133] like Fig.10 As shown, the target detection device proposed in the embodiment of the present invention includes the following modules:

[0134] An acquisition module 10 is used to acquire an image to be detected and an initial data set;

[0135] A processing module 20, configured to perform image preprocessing on the initial data set to obtain a preprocessed data set;

[0136] A construction module 30 is a construction module for constructing an initial target detection model based on a YOLOv8 network, replacing the backbone network of the initial target detection model with a fasternet structure, adding a CPCA attention mechanism to the backbone network of the initial target detection model, and replacing the neck network of the initial target detection model with an improved Slimneck to obtain an improved target detection model, wherein the improved Slimneck structure is constructed based on the Slimneck structure and the spatial pyramid pooling module of the Slimneck structure is replaced with a Focalmodulation module;

[0137] A training module 40, configured to train the improved target detection model using the preprocessed data set to obtain an optimized target detection model;

[0138] The detection module 50 is used to detect the image to be detected according to the optimized target detection model to obtain a detection result.

[0139] In addition, the present invention also proposes a target detection device, which includes: a memory, a processor, and a target detection program stored in the memory and executable on the processor, wherein the target detection program is configured to implement the steps of the target detection method described above.

[0140] Since the target detection device adopts all the technical solutions of all the above embodiments, it has at least all the beneficial effects brought by the technical solutions of the above embodiments, which will not be described one by one here.

[0141] In addition, an embodiment of the present invention further provides a storage medium, on which a target detection program is stored. When the target detection program is executed by a processor, the steps of the target detection method described above are implemented.

[0142] Since the storage medium adopts all the technical solutions of all the above embodiments, it has at least all the beneficial effects brought by the technical solutions of the above embodiments, which will not be described one by one here.

[0143] It should be understood that the above is only an example and does not constitute any limitation on the technical solution of the present invention. In specific applications, technicians in this field can make settings as needed, and the present invention does not limit this.

[0144] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of the present invention. In practical applications, technicians in this field can select part or all of them according to actual needs to achieve the purpose of the present embodiment, and no limitation is made here.

[0145] In addition, for technical details not fully described in this embodiment, reference can be made to the target detection method provided in any embodiment of the present invention, and will not be repeated here.

[0146] In addition, it should be noted that, in this article, the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0147] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory (ROM) / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0149] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.< / h> < / w> < / y> < / x> < / c> < / occlusion> < / truncation> < / score>

Claims

1. A multi-target detection method for pests and diseases, characterized in that: The method comprises: Obtain an initial image dataset of the target to be detected; Performing image preprocessing on the initial image data set to obtain a preprocessed data set; An initial target detection model is constructed based on the YOLOv8 network. The backbone network of the initial target detection model is a fasternet structure, and a CPCA attention mechanism module is added after the last C2f of the backbone network. The fasternet structure includes an embedding layer, a convolution layer, a merging layer, a pooling layer and a CPCA attention mechanism. The convolution layer includes a partial convolution and a one-dimensional convolution. The neck layer network of the initial target detection model is an improved Slimneck structure, and an improved target detection model is obtained. The improved Slimneck structure is constructed based on the original Slimneck structure and the spatial pyramid pooling module of the original Slimneck structure is replaced by a Focalmodulation module, including a Focalmodulation module, upsampling, GSConv and VoVGSCSP; Using the preprocessed data set to train the initial target detection model to obtain an optimized target detection model; The optimized target detection model is used to detect the target image to be detected to obtain a detection result.

2. The multi-target detection method for pests and diseases according to claim 1, characterized in that: The performing image preprocessing on the initial image data set to obtain the preprocessed data set specifically includes: Converting the original labels of the initial image data set into target labels to obtain a target data set; The target data set is amplified by geometric transformation and color transformation to obtain a preprocessed enhanced mark data set, wherein the geometric transformation includes at least one of flipping, rotating, cropping, exposing, adding noise and blurring, and the color transformation includes at least one of color transformation, erasing, mosaic and filling.

3. The multi-target detection method for pests and diseases according to claim 1, characterized in that: The applying the preprocessed data set to train the initial target detection model to obtain an optimized target detection model specifically includes: The images in the preprocessed dataset are input into the improved FasterNet structure of the initial target detection model for feature extraction to obtain a reference feature map; Input the reference feature map to the improved Slimneck structure of the neck layer for fusion to obtain a target feature map, wherein the improved Slimneck structure is obtained by replacing the spatial pyramid pooling module of the original Slimneck structure with a Focalmodulation module, including a Focalmodulation module, upsampling, GSConv and VoVGSCSP; Inputting the target feature map into the head network of the improved initial target detection model for detection, and outputting a marking box and a classification label; Determine a target loss function based on the marked box and the classification label; The parameters of the improved initial target detection model are optimized by the target loss function to obtain an optimized target detection model.

4. The multi-target detection method for pests and diseases as claimed in claim 3, characterized in that: The step of inputting the image in the preprocessed data set into the FasterNet structure of the initial target detection model for feature extraction to obtain a reference feature map specifically includes: Perform feature extraction on the images in the preprocessed data set to obtain the original feature map; The original feature map is further convolved to extract features and obtain a number of sub-feature maps; Connect the plurality of sub-feature maps according to the channel dimension to obtain an intermediate feature map; The intermediate feature map is further input to the merging layer for feature merging, and then part of the convolution is repeatedly used for convolution to obtain a convolution feature map; The convolutional feature map is input into the CPCA attention mechanism, and spatial information is aggregated from the reference feature map through average pooling and maximum pooling operations to generate two independent spatial context descriptors; Inputting two independent spatial context descriptors into a shared multi-layer perceptron (MLP), combining the outputs of the shared MLP by element-wise summation to obtain a channel weight vector, wherein the shared MLP consists of a single hidden layer; The channel weight vector is used to multiply each channel of the original feature map element by element to enhance the features of the target channel and suppress the output feature maps of non-target channels to obtain feature maps of different scales.

5. The multi-target detection method for pests and diseases as claimed in claim 3, characterized in that: The step of inputting the reference feature map into the improved Slimneck network of the neck layer for fusion to obtain the target feature map specifically includes: Input the reference feature map into the pyramid space pooling layer improved by Focalmodulation to extract feature information; Inputting feature maps of different scales into the Focalmodulation module for aggregation processing to obtain a target feature map; Finally, convolution is performed through the GSConv module, feature maps are fused by upsampling and concatenation, and features are enhanced again through the GSConv module; The VoV-GSCSP module is used for cross-stage feature fusion to provide the head network with an optimized target feature map.

6. The multi-target detection method for pests and diseases as claimed in claim 3, characterized in that: Determining the target loss function according to the marking frame and the classification label specifically includes: The position loss function is obtained by calculating based on the marked box and the real box; Calculate according to the classification label and the preset label to obtain the classification loss function; A target loss function is determined according to the position loss function and the classification loss function.

7. A target detection device, characterized in that: The target detection device comprises: An acquisition module is used to acquire an initial image data set of the target to be detected; A processing module, used for performing image preprocessing on the initial image data set to obtain a preprocessed data set; A construction module is used to construct an initial target detection model based on a YOLOv8 network, wherein the backbone network of the initial target detection model is a fasternet structure, and a CPCA attention mechanism is added after the backbone network, and the neck layer network of the initial target detection model is an improved Slimneck, thereby obtaining an improved target detection model, wherein the improved Slimneck structure is constructed based on an original Slimneck structure and the spatial pyramid pooling module of the original Slimneck structure is replaced by a Focalmodulation module; including a Focalmodulation module, upsampling, GSConv, and VoVGSCSP; A training module, used to train the initial target detection model using the preprocessed data set to obtain an optimized target detection model; The detection module is used to detect the target image to be detected using the optimized target detection model to obtain a detection result.

8. A computing device, characterized in that The computing device comprises: a memory, a processor, and a target detection program stored in the memory and executable on the processor, wherein the target detection program is configured to implement the target detection method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium stores a target detection program, and when the target detection program is executed by the processor, the target detection method according to any one of claims 1 to 6 is implemented.