Weed seed detection model training method, application method, device and equipment

By adding new paths to the object detection model, deleting the sampling module, and replacing the backbone network module, the problem of insufficient detection accuracy of the weed seed detection model in small objects and excessive redundancy is solved, and efficient and lightweight weed seed detection is achieved.

CN120107566AActive Publication Date: 2025-06-06NORTHWEST A & F UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510578007.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing weed seed detection model is difficult to meet the real-time detection requirements of small targets due to insufficient detection accuracy and high model redundancy.

Method used

By adding new paths to the feature pyramid network of the object detection model, deleting the sampling module, and replacing the coarse-thin module in the backbone network with a reparameterized visual geometry group model, the model calculation amount is reduced and the detection accuracy is improved.

Benefits of technology

The detection accuracy of small targets such as weed seeds is improved. At the same time, the model calculation amount is reduced without significantly affecting the model performance, and the weed seed detection model is lightweight, which is suitable for real-time detection on edge devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107566A_ABST
    Figure CN120107566A_ABST
Patent Text Reader

Abstract

The invention provides a training method, an application method, a device and equipment of a weed seed detection model. The method comprises the steps of obtaining an image data set of weed seeds; in a feature pyramid network of a preset target detection model, up-sampling the first feature map to a second feature map, and deleting the first sampling module to obtain a modified feature pyramid network; adding a module for down-sampling the second feature map to the first feature map in the path aggregation network of the target detection model, and deleting the second sampling module to obtain a modified path aggregation network; replacing a coarse-fine module in the backbone network of the target detection model with a re-parameterized visual geometric group model, and abandoning the third sampling module to obtain a modified backbone network; and obtaining a weed seed detection model according to the modified feature pyramid network, the modified path aggregation network and the modified backbone network, and training the weed seed detection model by using a training set in the image data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and specifically relates to a training method, application method, device and equipment for a weed seed detection model. Background Art

[0002] At present, in recent years, agricultural technology has kept pace with the times, and the deep integration of scientific and technological innovation and agricultural industry has produced smart agriculture. With the help of advanced scientific and technological means, smart agriculture manages and controls agricultural production in an information-based, intelligent and automated manner to improve agricultural production efficiency, reduce resource consumption and protect the ecological environment. It has become an important development direction of modern agriculture.

[0003] The existing training methods of weed seed detection models rely on manual inspection or traditional computer vision technology. However, these methods often have insufficient detection accuracy for small targets such as weed seeds, resulting in a high rate of missed detection, and the model redundancy is too high to meet the needs of real-time detection. Summary of the invention

[0004] The present application aims to provide a training method, application method, device and equipment for a weed seed detection model, which can solve the problems of insufficient small target detection accuracy and high model redundancy.

[0005] In a first aspect, an embodiment of the present application discloses a method for training a weed seed detection model, the method comprising: Get a dataset of images of weed seeds; A first path is added to a feature pyramid network of a preset target detection model, and a first sampling module is deleted to obtain a modified feature pyramid network; the first path is used to upsample the first feature map to the size of the second feature map and then fuse it, and the size of the first feature map is smaller than the size of the second feature map; the first sampling module is used to upsample the third feature map to the size of the fourth feature map and fuse it, and the size of the third feature map is smaller than the size of the fourth feature map; A second path is added to the path aggregation network of the target detection model, and a second sampling module is deleted to obtain a modified path aggregation network; the second path is used to downsample the second feature map to the size of the first feature map and then fuse it; the second sampling module is used to downsample the fourth feature map to the size of the third feature map and fuse it; Replacing the coarse-fine module in the backbone network of the target detection model with a reparameterized visual geometry group model, and discarding the third sampling module to obtain a modified backbone network; A weed seed detection model is obtained according to the modified feature pyramid network, the modified path aggregation network and the modified backbone network, and the weed seed detection model is trained using a training set in the image data set.

[0006] In a second aspect, an embodiment of the present application discloses an application method of a weed seed detection model, the method comprising: Obtaining a seed image to be identified; The seed image to be identified is input into a trained weed seed detection model to obtain an output weed seed identification result, wherein the weed seed detection model is trained by the weed seed detection model training method described in the first aspect.

[0007] In a third aspect, an embodiment of the present application discloses a training device for a weed seed detection model, the device comprising: A first acquisition module is used to acquire an image dataset of weed seeds; A first network improvement module is used to add a first path in a feature pyramid network of a preset target detection model and delete a first sampling module to obtain a modified feature pyramid network; the first path is used to upsample the first feature map to the size of the second feature map and then fuse it, and the size of the first feature map is smaller than the size of the second feature map; the first sampling module is used to upsample the third feature map to the size of the fourth feature map and fuse it, and the size of the third feature map is smaller than the size of the fourth feature map; A second network improvement module is used to add a module for downsampling the second feature map to the first feature map in the path aggregation network of the target detection model, and delete the second sampling module to obtain a modified path aggregation network; the second sampling module is used to downsample the fourth feature map to the third feature map; A third network improvement module is used to replace the coarse-fine module in the backbone network of the target detection model with a reparameterized visual geometry group model and discard the third sampling module to obtain a modified backbone network; The model training module is used to obtain a weed seed detection model according to the modified feature pyramid network, the modified path aggregation network and the modified backbone network, and train the weed seed detection model using a training set in the image data set.

[0008] In a fourth aspect, an embodiment of the present application discloses an application device of a weed seed detection model, the device comprising: A second acquisition module is used to acquire a seed image to be identified; The model application module is used to input the seed image to be identified into a trained weed seed detection model to obtain an output weed seed identification result, wherein the weed seed detection model is trained by the weed seed detection model training method described in the first aspect.

[0009] In a fifth aspect, an embodiment of the present application discloses an electronic device, comprising a processor and a memory, wherein the memory stores a program or instruction running on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0010] In summary, in this embodiment, on the one hand, by adding a new path that is more suitable for small targets in the target detection model, the detection accuracy of small targets such as weed seeds is improved; on the other hand, some paths in the target detection model are discarded, and the amount of model calculation is reduced without significantly affecting the model performance, thereby achieving a lightweight weed seed detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flowchart of the steps of a training method for a weed seed detection model provided in an embodiment of the present application.

[0012] Figure 2 It is a flowchart of the specific steps of a training method for a weed seed detection model provided in an embodiment of the present application.

[0013] Figure 3 It is a network framework diagram of a feature pyramid network provided in an embodiment of the present application.

[0014] Figure 4 It is a network framework diagram of a path aggregation network provided in an embodiment of the present application.

[0015] Figure 5 It is a network framework diagram of a re-parameterized visual geometry group model provided in an embodiment of the present application.

[0016] Figure 6 It is a network framework diagram of a weed seed detection model provided in an embodiment of the present application.

[0017] Figure 7 It is a flowchart of the steps of an application method of a weed seed detection model provided in an embodiment of the present application.

[0018] Figure 8 It is a block diagram of a training device for a weed seed detection model provided in an embodiment of the present application.

[0019] Fig. 9 It is a block diagram of an application device of a weed seed detection model provided in an embodiment of the present application.

[0020] Fig.10 It is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0022] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0023] In the description of the present disclosure, unless otherwise specified, "multiple" means two or more than two, and other quantifiers are similar thereto; "at least one item", "one item or multiple items" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one item a can represent any number of a; for another example, one item or multiple items among a, b and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple; "and / or" is a kind of description of the association relationship of associated objects, indicating that there can be three kinds of relationships, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " indicates that the associated objects before and after are in an "or" relationship.

[0024] Although operations or steps are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be understood that it is required to perform these operations or steps in the specific order shown or in a serial order, or to perform all the operations or steps shown to obtain the desired results. In the embodiments of the present disclosure, these operations or steps can be performed in series; these operations or steps can also be performed in parallel; or some of these operations or steps can be performed.

[0025] First, the application scenarios of the present disclosure are described.

[0026] Take Amaranth as an example. The invasion of Amaranth into farmland can inhibit the growth of crops, resulting in serious crop yield reduction and quality degradation. For example, Amaranth caused a 65% decrease in cotton yield, a 79% decrease in soybean yield, and a 91% loss in corn yield in the North American agricultural region where it originated. Current quarantine weed seed detection methods mainly rely on manual inspection or traditional computer vision technology, but these methods have the following significant limitations: 1. Insufficient accuracy in small target detection: The P3-P5 detection head in the traditional YOLOv8n model (corresponding to a downsampling rate of 1 / 8 to 1 / 32) has insufficient feature resolution for tiny seeds (such as Amaranth seeds, which are only 1×1mm in size), resulting in a high missed detection rate; 2. Excessive model redundancy: The number of parameters in the standard convolutional layer is huge, and the inference speed is limited when deployed on edge devices, making it difficult to meet the needs of real-time detection.

[0027] In order to solve the above-mentioned problems, the present application provides a training method, application method, device and equipment for a weed seed detection model. The specific implementation methods of the present application are described in detail below in conjunction with the accompanying drawings.

[0028] Figure 1 This is a training method for a weed seed detection model provided in this embodiment, referring to Figure 1 , the method may include the following steps.

[0029] Step 101: Obtain an image dataset of weed seeds.

[0030] In some embodiments, step 101 may include: photographing a preset number of weed seed images by a camera, counting valid samples in the weed seed images and performing preprocessing to obtain a weed seed image dataset.

[0031] For example, for weed seeds Amaranth seeds, the preset number may be 1,000 images, and step 101 may include: using an industrial camera to separately photograph Amaranth seed groups under different lighting conditions, and collecting 1,000 images to construct a Amaranth seed dataset.

[0032] For another example, step 101 may include: simulating an actual application scenario, mixing seeds of other sizes into the images, and collecting a total of 1,000 images to construct a weed seed dataset.

[0033] In some embodiments, counting valid samples in a weed seed image and performing preprocessing to obtain a weed seed image dataset may include: adding annotation information to the weed seed image using an automatic annotation tool to count and obtain valid samples; and converting a label file in an original format corresponding to the image dataset into a label file in a standard format.

[0034] Optionally, the original format can be JavaScript Object Notation (JSON), and the standard format can be the common text (txt) format of YOLO (You Only Look Once). For example, use the X-Anylabeling annotation tool to add annotation information to the images of the image dataset, count the valid samples and generate the json format label file corresponding to the image dataset; convert the json format label file to the txt format label file commonly used by YOLO.

[0035] In some embodiments, step 101 may further include: dividing the image dataset into a training set and a validation set according to a preset ratio.

[0036] Optionally, the preset ratio may be 70% for the training set and 30% for the validation set. For example, if the image dataset consists of 1,000 images, the Amaranth seed dataset is divided into a training set and a validation set according to a ratio of 7:3. After the division, the training set includes 700 images, and the validation set includes 300 images.

[0037] Step 102: Add a first path to the feature pyramid network of the preset target detection model, and delete the first sampling module to obtain a modified feature pyramid network.

[0038] The first path is used to upsample the first feature map to the size of the second feature map and then fuse it, and the size of the first feature map is smaller than that of the second feature map. The first sampling module is used to upsample the third feature map to the size of the fourth feature map and fuse it, and the size of the third feature map is smaller than that of the fourth feature map.

[0039] In some embodiments, the preset target detection model can be a YOLOv8n model.

[0040] In the embodiment of the present application, Feature Pyramid Networks (FPN) is a network structure used for feature fusion. The main function of FPN is to fuse feature maps of different scales so that the model can capture target information of different scales. For example, in the YOLOv8n model, FPN will perform upsampling and fusion operations on multiple scale feature maps (P2, P3, P4, P5, etc.) output by the backbone network.

[0041] It should be noted that up sampling is also known as image enlargement and image interpolation; the main purpose of up sampling is to enlarge the original image so that it can be displayed on a display device with higher resolution; up sampling principle: image enlargement almost always adopts the interpolation method, that is, new elements are inserted between pixels based on the original image pixels using a suitable interpolation algorithm; interpolation algorithms also include traditional interpolation, edge image-based interpolation, and area-based image interpolation.

[0042] In some embodiments, the upsampling method in step 102 may include any one of the following: bilinear interpolation (bilinear), deconvolution (transposed convolution), and unpooling (un pooling).

[0043] In the embodiment of the present application, the size of the first feature map is smaller than the size of the second feature map. For example, the first feature map may be a P3 feature map, and the second feature map may be a P2 feature map.

[0044] It should be noted that P2 and P3 represent feature maps of different scales output by the backbone network. The feature map of P2 has the largest size, the least semantic information, but the richest spatial information; the size of the feature map of P3 is smaller than that of P2, and it has more semantic information. Upsampling will adjust the size of the feature map of P3 to the same size as that of P2. Figure 1 Similarly, fusing the two can make the model perform better in small target detection, because small targets are easier to detect in large-size feature maps (P2), and fusing the semantic information of P3 helps improve the accuracy of detection.

[0045] In the embodiment of the present application, the first sampling module is used to upsample the third feature map to a fourth feature map.

[0046] In the embodiment of the present application, the size of the third feature map is smaller than that of the fourth feature map. For example, the third feature map may be a P5 feature map, and the fourth feature map may be a P4 feature map.

[0047] It should be noted that the P5 feature map has the smallest size and the richest semantic information, but less spatial information; the P4 feature map is larger than the P5, and the semantic information and spatial information are at an intermediate level. In the original FPN structure, the P5 feature map is upsampled to the size of the P4 feature map and then fused. By discarding the first sampling module, the model no longer performs upsampling and fusion operations on the P5 and P4 feature maps, thus reducing the amount of calculation and parameters of the model.

[0048] Step 103: Add a second path to the path aggregation network of the target detection model, and delete the second sampling module to obtain a modified path aggregation network.

[0049] The second path is used to downsample the second feature map to the size of the first feature map and then fuse it. The second sampling module is used to downsample the fourth feature map to the size of the third feature map and fuse it.

[0050] In an embodiment of the present application, a path aggregation network (PAN) is a network structure used to enhance the multi-scale feature fusion capability of a target detection model. It is improved on the basis of FPN and can more effectively fuse feature information of different scales, thereby improving the detection performance of the target detection model on targets of different sizes.

[0051] It should be noted that down sampling is a process in deep learning and signal processing that reduces the spatial resolution and amount of data, reduces the computational complexity by reducing the size of feature maps, captures a wider range of contextual information, and prevents overfitting.

[0052] In some embodiments, the upsampling method in step 102 may include any one of the following: pooling, convolution with a step size greater than 1, etc.

[0053] It can be understood that P2 and P3 represent feature maps of different scales output by the backbone network. P2 has rich spatial information, and P3 feature map has relatively stronger semantic information. By downsampling the P2 feature map to the size of the P3 feature map and then performing feature fusion, the P3 feature map can retain the semantic information while incorporating more spatial information, thereby improving the model's ability to locate and detect small targets.

[0054] In the embodiment of the present application, the second sampling module is used to downsample the fourth feature map to the third feature map.

[0055] It should be noted that in the original PAN structure, the module of downsampling P4 to P5 is to transfer the information of P4 to P5 and further enhance the feature expression ability of P5. Since the semantic information of P4 and P5 is already rich enough, the operation of downsampling P4 to P5 does not contribute much to the improvement of model performance. By deleting the second sampling module, the amount of calculation can be reduced and the reasoning speed of the model can be improved without significantly affecting the performance of the model.

[0056] Step 104: Replace the coarse-fine module in the backbone network of the target detection model with the reparameterized visual geometry group model, and discard the third sampling module to obtain a modified backbone network.

[0057] In the embodiment of the present application, the backbone network is used to extract features from the input image to provide information for subsequent target detection tasks. The backbone network adopts a layered architecture to gradually downsample and extract features from the input image to generate feature maps of different scales, which will be passed to the subsequent FPN and PAN for further processing.

[0058] It should be noted that the backbone network includes a convolution (Conv) module, a coarse-to-fine (C2f) module, and a spatial pyramid pooling fast (SPPF) module. Among them, the C2f module is an improved version of the cross-stage partial network (CSP) module in the YOLOv8n model. The C2f module diverts the input features, extracts some features through multiple convolutional layers, and directly connects the other features across layers, and finally fuses the two features. Optionally, for the convenience of description, the C2f layer in this application can also be referred to as the "positioning classification layer."

[0059] In the embodiment of the present application, the re-parameterization visual geometry group model (Re-parameterization Visual Geometry Group, RepVGG) is a convolutional neural network architecture that performs well in the field of computer vision. Because a single-path structure is used during reasoning, the computational overhead caused by branching and addition operations in a multi-branch structure is avoided. Compared with some complex network structures, RepVGG has obvious advantages in memory usage and computing resource requirements, and is more suitable for running on resource-constrained devices; and, although the structure is simple during reasoning, through the multi-branch structure during training, RepVGG can learn rich feature information, and the performance will not be affected. Replacing the coarse-fine module in the backbone network of the target detection model with the re-parameterization visual geometry group model can achieve fast and accurate reasoning that is not limited by the hardware platform.

[0060] In the embodiment of the present application, the third sampling module may be a fifth downsampling module.

[0061] It should be noted that by discarding the third sampling module, the feature map that was originally obtained after 5 downsamplings now only undergoes 4 downsamplings, the number of downsampling times is reduced, the resolution of the feature map is improved, and more image details can be retained, which is conducive to the detection of small targets. At the same time, by discarding the fifth downsampling module, a layer of calculation in the network is reduced, the amount of calculation and parameters of the model are reduced, making the model lighter and running faster.

[0062] Step 105: obtain a weed seed detection model according to the modified feature pyramid network, the modified path aggregation network and the modified backbone network, and train the weed seed detection model using a training set in the image data set.

[0063] In the embodiment of the present application, the structures of the feature pyramid network, the path aggregation network and the backbone network are modified. The FPN is enhanced and optimized based on the original structure, some paths are deleted and new paths that are more suitable for small targets are added. The PAN is modified on the original structure, the information interaction direction is reconstructed, and the fusion of mid- and low-level features is focused. The C2f module in the backbone is replaced by the RepVGG model, and the number of sampling times is reduced. In this way, the YOLOv8n model is improved and a weed seed detection model is obtained.

[0064] In some embodiments, the weed seed detection model is trained by building a virtual environment for training the model on a Windows Server 2022 Datacenter server.

[0065] In summary, in this embodiment, on the one hand, by adding a new path that is more suitable for small targets in the target detection model, the detection accuracy of small targets such as weed seeds is improved; on the other hand, some paths in the target detection model are discarded, and the amount of model calculation is reduced without significantly affecting the model performance, thereby achieving a lightweight weed seed detection model.

[0066] Figure 2 is a flowchart of the specific steps of a training method for a weed seed detection model provided in an embodiment of the present application, with reference to Figure 2 , the training method of the above-mentioned weed seed detection model may include the following steps.

[0067] Step 201: Obtain an image dataset of weed seeds.

[0068] The method of this step has been described in the aforementioned step 101 and will not be repeated here.

[0069] Step 202: up-sampling is performed between the first preset layer and the second preset layer in the feature pyramid network through a first path to expand the size of the first feature map to the size of the second feature map.

[0070] In the embodiment of the present application, the first preset layer is the 15th layer, and the second preset layer is the 16th layer.

[0071] For example, see Figure 3 , the network framework diagram of the feature pyramid network is as follows Figure 3As shown. Between the 15th and 16th layers in the feature pyramid network, upsampling is performed through the first path to expand the P3 feature map to the size of the P2 feature map module. Figure 3 It includes 1 upsampling layer (Upsample), 1 feature fusion layer (Concat), 1 positioning classification layer (C2f layer), and the number of channels is 128. The P2 feature map refers to the feature map with a resolution of 160×160, that is, the feature map of 1 / 4 input image resolution (640×640). The first input source of Upsample sampling is the fused and spliced ​​P3 feature map in the FPN network, specifically the 15th layer in the feature pyramid network, which doubles the size of the original P3 feature map in the 15th layer. The second input source of Concat splicing is the feature map after the second downsampling in the backbone network, specifically the second layer in the feature pyramid network, which splices the P3 feature map that has been expanded by 2 times with the feature map after the second downsampling. C2f fuses the spliced ​​feature maps to enhance the expression ability of the P2 feature map.

[0072] Step 203: Delete the first upsampling layer, the feature fusion layer, and the positioning classification layer between the third preset layer and the fourth preset layer in the feature pyramid network to obtain a modified feature pyramid network.

[0073] In the embodiment of the present application, the third preset layer is the 10th layer, and the fourth preset layer is the 12th layer.

[0074] In the embodiment of the present application, the positioning classification layer is the C2f layer.

[0075] Step 204: Down-sampling is performed through the second path in the path aggregation network to reduce the size of the second feature map to the size of the first feature map.

[0076] For example, see Figure 4 , the network framework diagram of the path aggregation network is as follows Figure 4 As shown in Figure 2, the network structure of the module that downsamples the P2 feature map to the P3 feature map in the path aggregation network is as follows: Figure 4It includes 1 convolution layer (Conv), 1 feature fusion layer (Concat), 1 positioning classification layer (C2f layer), and the number of channels is 256. The first input source of Conv downsampling is the P2 feature map after C2f fusion enhancement in step 203, specifically the 13th layer in the path aggregation network, which is used to reduce the size of the P2 feature map of the 13th layer in the path aggregation network by half. The second input source of the feature fusion layer (Concat) splicing is the fused and spliced ​​P3 feature map in the FPN network, specifically the 15th layer in the feature pyramid network, which is used to splice the fused and spliced ​​P3 feature map in the FPN network with the P2 feature map with half the size. C2f fuses the spliced ​​feature maps to enhance the expression ability of the P3 feature map.

[0077] Step 205: Delete the second upsampling layer, the feature fusion layer, and the positioning classification layer between the third preset layer and the fourth preset layer in the path aggregation network to obtain a modified path aggregation network.

[0078] In the embodiment of the present application, the third preset layer is layer 10, and the fourth preset layer is layer 12. For example, the positioning classification layer may be a C2f module.

[0079] Step 206: Replace the coarse-fine module in the backbone network with a reparameterized visual geometry group model.

[0080] In some embodiments, the re-parameterized visual geometry group model includes: an initial module (Stem Block), three re-parameterized modules (Re-parameterized Block), two residual modules (Residual Block), an average pooling module (Avg Pool), a connection layer (Fully Connected, FC) and an output layer (Soft max); wherein the re-parameterized module includes: a group convolution (Group Conv), a composite activation function module (BatchNorm+SiLU), a shape adjustment module (Adjust Shape) and a fused convolution module (Concat+Conv).

[0081] It should be noted that the residual module is an important structure in deep learning. It solves the gradient vanishing and gradient exploding problems in deep neural network training by introducing skip connections or identity mapping. The core idea of ​​the residual module is to let the network learn the residual between the input and output, that is, the difference part, instead of directly learning the mapping relationship. The residual module usually consists of two or three convolutional layers, including a batch normalization layer (Batch Normalization, BN) and an activation function. In the residual module, the input is not only processed by these convolutional layers, but also passed directly to the output through skip connections.

[0082] For example, see Figure 5 , the network structure of RepVGG is as follows Figure 5 As shown in the figure, it is divided into different structures according to the two stages of training and inference. The structure during training consists of a main branch with a convolution kernel size of 3×3, a convolution branch with a convolution kernel size of 1×1, and an identity branch that only contains batch normalization operations. During inference, the 3×3 convolution operation of the main branch, the 1×1 convolution branch, and the identity branch are merged into a unified convolution operation to improve the inference speed. During the training phase, for the feature map with an input size of C×H×W, RepVGG first performs Group Conv, extracts spatial features through the 3×3 convolution of the main branch, generates a feature map of C′×H′×W′, and adjusts the number of channels through the 1×1 convolution branch, outputs C′′×H′′×W′′, and performs residual connection through the identity branch to directly pass the input feature map C×H×W. Normalization (Batch Normalization) is performed after the output of each branch. The Adjust Shape module combines the output feature maps of the three branches into a new feature map through the Concat operation, and finally performs a nonlinear transformation through the activation function (SiLU) to generate the final feature map. In the inference stage, the 3×3 convolution main branch, 1×1 convolution branch, and identity branch during training will be merged into a unified convolution operation, and multiple Batch Normalization layers are also fused to avoid the calculation of multiple branches, thereby improving the inference speed.

[0083] Step 207: Delete the convolutional layer located at the fifth preset layer and the coarse-fine module located at the sixth preset layer in the backbone network to obtain a modified backbone network.

[0084] In the embodiment of the present application, the fifth preset layer is the 5th layer, and the sixth preset layer is the 6th layer.

[0085] In some embodiments, step 207 may include: deleting the convolution layer located at the fifth preset layer and the C2f module located at the sixth preset layer in the backbone network, and deleting the corresponding Conv convolution layer.

[0086] It should be noted that since the convolutional layer at the fifth preset layer and the coarse-fine module at the sixth preset layer in the backbone network are deleted, the feature extraction and processing tasks originally undertaken by these two layers will no longer be performed, which will cause the feature map output by the backbone network to change. Since the semantic information of the feature map at a higher layer is richer, the amount of calculation and model complexity are reduced, the reasoning speed of the model is increased, and the real-time processing capability of the model is greatly improved. In addition, the simplified model structure is simpler, easier to understand and deploy, but it may reduce the expressiveness and generalization ability of the model.

[0087] Step 208: Obtain a weed seed detection model based on the modified feature pyramid network, the modified path aggregation network, and the modified backbone network, and use the training set in the image data set to train the weed seed detection model.

[0088] The method of this step has been described in the aforementioned step 105 and will not be repeated here.

[0089] For example, see Figure 6 , the network structure of the weed seed detection model is as follows Figure 6 See Figure 6 ,The weed seed detection model includes a modified feature pyramid network, a ,modified path aggregation network, a modified backbone network and a head network, ,wherein, a P2 feature layer enhancement module is added to the ,feature pyramid network, e.g. Figure 3 In the 15th and 16th layers, the P3 feature map is expanded to the size of the P2 feature map module through upsampling in the first path; the multi-scale feature interaction mode is optimized in the path aggregation network, such as Figure 4 In the path aggregation network, the P2 feature map is downsampled to the P3 feature map module; the redundant downsampling layers of the backbone network are pruned and the backbone network is replaced with the RepVGG reparameterized structure, for example Figure 5 In this way, the detection accuracy is improved while the model size is lightened. The detection frame rate on the edge device reaches 36.18fps. Compared with the traditional YOLOv8n model, the recall rate is increased by 16.99% and the model size is reduced by 68.7%.

[0090] In some embodiments, after step 208, the training method of the weed seed detection model may further include: Sub-step 209, using the validation set in the image data set, evaluating the average precision mean, detection speed, number of parameters, and model size of the weed seed detection model; Sub-step 210: When the average precision mean, detection speed, parameter quantity and model size all meet the preset conditions, the result of this iteration is used as the trained weed seed detection model.

[0091] In the embodiments of the present application, mean average precision (mAP) is an indicator for comprehensively evaluating the accuracy and recall of a model, and is closely related to precision and recall.

[0092] Optionally, precision represents the ratio of the number of samples correctly predicted as positive (True Positive, TP) to the number of samples predicted as positive TP+ (False Positive, FP). Recall represents the ratio of the number of samples correctly predicted as positive (TP) to the number of samples actually positive TP+ (False Negative, FN). Exemplarily, the calculation formula of precision is as follows:

[0093] The recall calculation formula is as follows:

[0094] In some embodiments, mean average precision (mAP) is one of the key indicators for measuring the performance of the object detection model. The precision of all categories is averaged when mAP is calculated. AP@0.5 represents the average precision for each category when the intersection over union (IoU) threshold is 0.5. mAP@0.5 is the average of the AP values ​​of all categories to reflect the ability of the model to maintain high precision at different recall rates. The higher the value of mAP@0.5, the better the model can maintain precision at high recall rates. The mAP calculation formula is as follows:

[0095] In the above formula, N represents the number of categories, and APi@0.5 represents the average precision value of the i-th category.

[0096] In the embodiment of the present application, the parameter quantity can be expressed as Giga Floating-point Operations Per Second (GFLOPs). GFLOPs can measure the ability of a central processing unit (CPU) or a graphics processing unit (GPU) to process graphics and deep learning models, and conversely can also be used to represent the CPU or GPU computing resources required for a neural network model.

[0097] In the embodiment of the present application, the unit of the model size is megabytes (MB).

[0098] In the embodiment of the present application, the detection frame rate (FPS) is used to measure the speed at which the model processes images (or video frames), that is, how many frames of images can be processed per second. It is a measure of inference efficiency and reflects the performance of the model on a given hardware. FPS is short for frame rate, and a high FPS means smoother movement and a better visual experience. The unit of FPS is Hertz (Hz). The FPS calculation formula is as follows:

[0099] In the above formula, the unit of average inference time is milliseconds.

[0100] In some embodiments, in order to understand the gain effect of each module on the target detection model, multiple groups of ablation experiments are set up. The ablation experiments are shown in Table 1 and are specifically shown as follows.

[0101] Table 1

[0102] First, the ablation experiment data shown in Table 1 shows that after introducing the feature pyramid network-path aggregation network (FAN-PAN) optimization, the precision and recall of the target detection model have been significantly improved, reaching 96.2% and 92.2% respectively. In other words, this optimization has greatly improved the detection accuracy of the weed seed detection model, especially when dealing with complex backgrounds and small targets.

[0103] Secondly, from the ablation experiment data shown in Table 1, it can be seen that the average precision is improved by 17.44%, which means that after optimization, compared with the traditional target detection model, the accuracy of the weed seed model in the embodiment of the present application is greatly improved.

[0104] Thirdly, backbone network optimization and RepVGG optimization successfully improved the inference speed by reducing the amount of computation and optimizing the backbone network, respectively, so that the FPS of the final weed seed model remained at a high level of 36.18, which was not significantly lower than the FPS of the initial model (38.33).

[0105] Finally, the size of the weed seed model was further reduced to 1.86MB, but the number of parameters did not increase significantly, only increasing from the initial 8.1GFLOPs to 9.3GFLOPs. In other words, while the weed seed model significantly reduced the model size, the demand for hardware computing power did not increase significantly. Therefore, it provides a more efficient deployment solution for memory-constrained edge devices, and the accuracy remains at a high level.

[0106] In summary, the weed seed model provided in the embodiment of the present application has a recall rate of 16.99% and an average precision of 19.10% compared to the initial model, and the model size is reduced by 68.69%, while the detection frame rate only decreases by 5.61%. This change shows that the optimization method used significantly improves the detection performance with only a small loss of detection speed, especially in small target detection. At the same time, the memory usage of the weed seed model is greatly reduced, making it more suitable for efficient real-time detection on memory-constrained edge devices.

[0107] Figure 7 This is an application method of a weed seed detection model provided in this embodiment, referring to Figure 7 , the method may include the following steps.

[0108] Step 301: Obtain a seed image to be identified.

[0109] Step 302: input the seed image to be identified into the trained weed seed detection model to obtain an output weed seed identification result.

[0110] In the embodiment of the present application, the weed seed detection model is trained by the training method of the weed seed detection model described in the above embodiment.

[0111] In summary, in the embodiments of the present application, the detection accuracy of small targets such as weed seeds is improved through the trained weed seed detection model, which is conducive to real-time recognition of seed images.

[0112] Figure 8 4 is a schematic diagram of the structure of a training device for a weed seed detection model provided in an embodiment of the present application. The training device 400 for a weed seed detection model may include the following modules.

[0113] A first acquisition module 401 is used to acquire an image dataset of weed seeds; The first network improvement module 402 is used to add a first path in the feature pyramid network of the preset target detection model and delete the first sampling module to obtain a modified feature pyramid network; the first path is used to upsample the first feature map to the size of the second feature map and then fuse it, and the size of the first feature map is smaller than the size of the second feature map; the first sampling module is used to upsample the third feature map to the size of the fourth feature map and fuse it, and the size of the third feature map is smaller than the size of the fourth feature map; The second network improvement module 403 is used to add a module for downsampling the second feature map to the first feature map in the path aggregation network of the target detection model, and delete the second sampling module to obtain a modified path aggregation network; the second sampling module is used to downsample the fourth feature map to the third feature map; The third network improvement module 404 is used to replace the coarse-fine module in the backbone network of the target detection model with the re-parameterized visual geometry group model and discard the third sampling module to obtain a modified backbone network; The model training module 405 is used to obtain a weed seed detection model according to the modified feature pyramid network, the modified path aggregation network and the modified backbone network, and train the weed seed detection model using a training set in the image data set.

[0114] Optionally, the first network improvement module 402 includes: An upsampling module, used to expand the size of the first feature map to the size of the second feature map by upsampling between the first preset layer and the second preset layer in the feature pyramid network; The first deletion module is used to delete the first upsampling layer, the feature fusion layer and the positioning classification layer between the third preset layer and the fourth preset layer in the feature pyramid network to obtain a modified feature pyramid network.

[0115] Optionally, the second network improvement module 403 includes: A downsampling module, configured to reduce the size of the second feature map to the size of the first feature map by downsampling in the path aggregation network; The second deletion module is used to delete the second upsampling layer, the feature fusion layer and the positioning classification layer between the third preset layer and the fourth preset layer in the path aggregation network to obtain a modified path aggregation network.

[0116] Optionally, the third network improvement module 404 includes: A network replacement module is used to replace the coarse-fine module in the backbone network with a reparameterized visual geometry group model; The third deletion module is used to delete the convolutional layer located at the fifth preset layer and the coarse-fine module located at the sixth preset layer in the backbone network to obtain a modified backbone network.

[0117] Optionally, the reparameterized visual geometry group model includes: an initial module, three reparameterized modules, two residual blocks, an average pooling module, a connection layer and an output layer; wherein the reparameterized module includes: a grouped convolution, a composite activation function module, a shape adjustment module and a fused convolution module.

[0118] Optionally, the training of the grass seed detection model further includes the following modules.

[0119] A model evaluation module is used to evaluate the mean average precision, detection speed, number of parameters, and model size of the weed seed detection model using a validation set in an image dataset; The model iteration module is used to use the result of this iteration as the trained weed seed detection model when the average precision mean, detection speed, parameter quantity and model size all meet the preset conditions.

[0120] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0121] Fig. 9 is a structural schematic diagram of an application device of a weed seed detection model provided in an embodiment of the present application. The application device 500 of the weed seed detection model may include the following modules.

[0122] The second acquisition module 501 is used to acquire a seed image to be identified.

[0123] The model application module 502 is used to input the seed image to be identified into the trained weed seed detection model to obtain the output weed seed identification result. The weed seed detection model is trained by the training method of the weed seed detection model.

[0124] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0125] Fig.10 600 is a block diagram of an electronic device 600 according to another embodiment of the present invention. For example, the electronic device 600 may be provided as a server. Figure 6The electronic device 600 includes a processing component 622, which further includes one or more processors, and a memory resource represented by a memory 632 for storing instructions executable by the processing component 622, such as an application. The application stored in the memory 632 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 622 is configured to execute instructions to perform a method for generating a text archive provided in an embodiment of the present application.

[0126] The electronic device 600 may also include a power supply component 626 configured to perform power management of the electronic device 600, a wired or wireless network interface 650 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 656. The electronic device 600 may operate based on an operating system stored in the memory 632, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, Free BSD TM or the like.

[0127] In an embodiment of the present application, the memory 632 can be used to store software programs and various data. The memory 632 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 632 may include a volatile memory or a non-volatile memory, or the memory 632 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 632 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0128] The embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the training method embodiment of the weed seed detection model is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0129] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0130] The embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the training method embodiment of the weed seed detection model as described above, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0131] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0132] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0133] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A training method for a weed seed detection model, characterized in that: The method comprises: Get a dataset of images of weed seeds; A first path is added to a feature pyramid network of a preset target detection model, and a first sampling module is deleted to obtain a modified feature pyramid network; the first path is used to upsample the first feature map to the size of the second feature map and then fuse it, and the size of the first feature map is smaller than the size of the second feature map; the first sampling module is used to upsample the third feature map to the size of the fourth feature map and fuse it, and the size of the third feature map is smaller than the size of the fourth feature map; A second path is added to the path aggregation network of the target detection model, and a second sampling module is deleted to obtain a modified path aggregation network; the second path is used to downsample the second feature map to the size of the first feature map and then fuse it; the second sampling module is used to downsample the fourth feature map to the size of the third feature map and fuse it; Replacing the coarse-fine module in the backbone network of the target detection model with a reparameterized visual geometry group model, and discarding the third sampling module to obtain a modified backbone network; A weed seed detection model is obtained according to the modified feature pyramid network, the modified path aggregation network and the modified backbone network, and the weed seed detection model is trained using a training set in the image data set.

2. The method according to claim 1, characterized in that: The step of adding a first path to a feature pyramid network of a preset target detection model and deleting a first sampling module to obtain a modified feature pyramid network includes: Upsampling is performed between a first preset layer and a second preset layer in the feature pyramid network through the first path to expand the size of the first feature map to the size of the second feature map; The first upsampling layer, the feature fusion layer, and the positioning classification layer between the third preset layer and the fourth preset layer in the feature pyramid network are deleted to obtain a modified feature pyramid network.

3. The method according to claim 1, characterized in that The step of adding a second path to the path aggregation network of the target detection model and deleting the second sampling module to obtain a modified path aggregation network includes: Downsampling the second path in the path aggregation network to reduce the size of the second feature map to the size of the first feature map; The second upsampling layer, the feature fusion layer, and the positioning classification layer between the third preset layer and the fourth preset layer in the path aggregation network are deleted to obtain a modified path aggregation network.

4. The method according to claim 1, characterized in that: The coarse-fine module in the backbone network of the target detection model is replaced with a reparameterized visual geometry group model, and the third sampling module is discarded to obtain a modified backbone network, including: Replacing the coarse-fine module in the backbone network with a reparameterized visual geometry group model; The convolutional layer located at the fifth preset layer and the coarse-fine module located at the sixth preset layer in the backbone network are deleted to obtain a modified backbone network.

5. The method according to claim 4, characterized in that The reparameterized visual geometry group model includes: an initial module, three reparameterized modules, two residual blocks, an average pooling module, a connection layer and an output layer; wherein the reparameterized module includes: a grouped convolution, a composite activation function module, a shape adjustment module and a fused convolution module.

6. The method according to claim 1, characterized in that The method further comprises: Using the validation set in the image dataset, evaluating the mean average precision, detection speed, number of parameters, and model size of the weed seed detection model; When the average precision mean, the detection speed, the parameter quantity and the model size all meet the preset conditions, the result of this iteration is used as the trained weed seed detection model.

7. An application method of a weed seed detection model, characterized in that: The method comprises: Obtaining a seed image to be identified; The seed image to be identified is input into a trained weed seed detection model to obtain an output weed seed identification result, wherein the weed seed detection model is trained by the weed seed detection model training method according to any one of claims 1 to 6.

8. A training device for a weed seed detection model, characterized in that: The device comprises: A first acquisition module is used to acquire an image dataset of weed seeds; A first network improvement module is used to add a first path in a feature pyramid network of a preset target detection model and delete a first sampling module to obtain a modified feature pyramid network; the first path is used to upsample the first feature map to the size of the second feature map and then fuse it, and the size of the first feature map is smaller than the size of the second feature map; the first sampling module is used to upsample the third feature map to the size of the fourth feature map and fuse it, and the size of the third feature map is smaller than the size of the fourth feature map; A second network improvement module is used to add a second path to the path aggregation network of the target detection model and delete the second sampling module to obtain a modified path aggregation network; the second path is used to downsample the second feature map to the size of the first feature map and then fuse it; the second sampling module is used to downsample the fourth feature map to the size of the third feature map and fuse it; A third network improvement module is used to replace the coarse-fine module in the backbone network of the target detection model with a reparameterized visual geometry group model and discard the third sampling module to obtain a modified backbone network; The model training module is used to obtain a weed seed detection model according to the modified feature pyramid network, the modified path aggregation network and the modified backbone network, and train the weed seed detection model using a training set in the image data set.

9. An application device of a weed seed detection model, characterized in that: The device comprises: A second acquisition module is used to acquire a seed image to be identified; A model application module is used to input the seed image to be identified into a trained weed seed detection model to obtain an output weed seed identification result, wherein the weed seed detection model is trained by the training method of the weed seed detection model according to any one of claims 1 to 6.

10. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and running on the processor, characterized in that when the processor executes the program, the method according to any one of claims 1 to 6 or claim 7 is implemented.

Citation Information

Patent Citations

  • Remote sensing image target detection method based on attention mechanism and multi-scale feature fusion

    CN118298263A

  • Model training method and device, target detection method and device and electronic equipment

    CN118521856A

  • Expressway outfield vehicle target detection method based on SSD algorithm

    CN118570742A

  • Weed classification detection method and system based on YOLOv8 improved algorithm

    CN118762286A

  • Yolov6-based directed target detection network, training method therefor, and directed target detection method

    WO2024119304A1