Method and system for detecting farmland weeds, and electronic device
By improving the backbone network structure and loss function of the YOLOv8 model, the problems of high computational complexity and large parameters in farmland weed recognition are solved, and more efficient weed recognition and precise removal are achieved.
Patent Information
- Application Number
- CN202410488772.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-04-22
AI Technical Summary
The prior art has problems in farmland weed recognition with low recognition accuracy, high computational complexity, large model parameters, and large model scale, which is difficult to meet the calculation and storage resource limitations of weeding robots.
The YOLOv8 model based on the RevColNet backbone network is used for improvement. By reconstructing the backbone network, introducing a fusion expandable residual attention module and a depth separation convolution module, the bounding box regression loss function is improved, the calculation complexity and parameter quantity are reduced, and feature extraction ability and recognition accuracy are improved.
While reducing the computational complexity and parameter quantity, the accuracy and efficiency of weed recognition are improved, achieving higher recognition accuracy and lower computing cost.
Smart Images

Figure CN118298406B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular, to a method and system for detecting farmland weeds, and an electronic device. Background Art
[0002] Weed control is one of the most critical tasks in agricultural production. Weeds have tenacious vitality and can have a huge impact on the yield and quality of crops by competing with crops for nutrients, water, light and other resources. According to statistics, the annual food loss caused by field weeds is about 13.2%, which is equivalent to the rations of 1 billion people for one year.
[0003] At present, the method of weed control is mainly completed by spraying herbicides over a large area. This indiscriminate spraying method will cause a large amount of pesticides to remain on the crops, which will not only affect the normal growth of the crops, but also cause certain damage to the field ecological environment.
[0004] Accurately identifying field weeds and achieving precise weeding play a huge role in improving the yield of crops and reducing the harm caused by pesticides to the ecology. Therefore, weeding robots that can accurately identify various weeds and implement weeding have gradually developed. It realizes intelligent weeding and plays an important role in increasing the yield of crops and reducing the impact of pesticides on the environment. Traditional methods for detecting field weeds mainly rely on manually designed features such as texture and shape, and use methods such as wavelet analysis, Bayesian discriminant models, and support vector machines to achieve the detection target. Since the manually designed features cannot well summarize various information of weeds, it is difficult to achieve a high recognition accuracy using these methods on more complex data sets.
[0005] In addition, the computing and storage resources of the core processing device of the weeding robot are limited. In view of the problems of high computational complexity, large number of model parameters, and large model scale of the current weeding robot, it is an urgent problem (very necessary) to reduce the number of model parameters and computational complexity while ensuring the recognition accuracy for weed recognition. Summary of the Invention
[0006] The present invention provides a method and system for detecting farmland weeds, and an electronic device, which are used to solve the defects in the prior art that when identifying farmland weeds, various information of weeds cannot be well described, it is difficult to achieve a high recognition accuracy, and it faces problems such as high computational complexity, large number of model parameters, and large model scale. The solution of the present application provides an improved model based on YOLOv8, which can identify weeds in farmland with higher accuracy, lower computational complexity, and higher weed recognition efficiency.
[0007] The present invention provides a method for detecting farmland weeds, including:
[0008] Collect the target images of weeds in the farmland;
[0009] Use YOLOv8 based on the RevColNet backbone network to build a weed detection model, and identify weeds based on the weed detection model;
[0010] Precisely remove the weeds.
[0011] According to the method for detecting farmland weeds provided by the present invention, using YOLOv8 based on the RevColNet backbone network to build a weed detection model, including:
[0012] Reconstruct the backbone network of YOLOv8 based on RevColNet to obtain a new backbone network RevCol;
[0013] Introduce a fused dilated residual attention module, and the fused dilated residual attention module improves the recognition ability for occluded targets;
[0014] Introduce the GSConv module and the VoV-GSCSPC module based on depthwise separable convolutions for model lightweight processing;
[0015] Improve the bounding box regression loss function of the YOLOv8 model based on the minimum point distance.
[0016] According to the method for detecting farmland weeds provided by the present invention, the backbone network RevCol includes multiple columns, each column represents an input, the starting position of each column contains low-level detail information, and as the image channels are compressed, high-level semantic information is extracted at the end of each column; a reversible connection design is adopted between columns to ensure lossless information transmission between columns, and at the same time, supervision is added at the very end of each column to constrain the feature extraction of each column.
[0017] According to the method for detecting farmland weeds provided by the present invention, the fused dilated residual attention module is used for:
[0018] Perform a standard 3×3 convolution operation on the data input to the weed detection model, and extract features through batch normalization and activation using an activation function;
[0019] After expanding the 3×3 convolution, obtain the semantic residual through the BN layer;
[0020] Connect all branch feature maps, use pointwise convolution to merge all feature maps, and generate the final residual corresponding to the data input to the weed detection model;
[0021] Fuse the final residual and the input data to construct the final feature representation.
[0022] According to the farmland weed detection method provided by the present invention, the dilatable residual attention module is fused and has a number of channels, and the number of dilated depth convolution channels with the lowest dilation rate is set to twice that of other channels.
[0023] According to the farmland weed detection method provided by the present invention, the GSConv module is used for:
[0024] Based on the number of input channels, obtain the first number of output channels through standard convolution;
[0025] Based on the first number of output channels, obtain the second number of output channels through depthwise separable convolution;
[0026] Connect and shuffle the first number of output channels and the second number of output channels to obtain the number of output channels.
[0027] According to the farmland weed detection method provided by the present invention, based on the minimum point distance, improve the bounding box regression loss function of the YOLOv8 model, including:
[0028] Determine the similarity between the predicted bounding box and the actual labeled bounding box during the bounding box regression process, and based on the similarity, calculate the key point distance between the predicted bounding box and the actual labeled bounding box to improve the accuracy of loss measurement.
[0029] According to the farmland weed detection method provided by the present invention, improve the bounding box regression loss function of the YOLOv8 model, including:
[0030] Introduce a scale factor ratio to control the size of the auxiliary bounding box to calculate the loss;
[0031] When the value of ratio is set to be greater than 1, generate an auxiliary bounding box with a larger scale relative to the actual bounding box to calculate the loss;
[0032] When the value of ratio is set to be less than 1, generate an auxiliary bounding box with a smaller scale to calculate the loss, so that the absolute value of the regression gradient is greater than the absolute value of the actual bounding box IoU gradient.
[0033] The present invention also provides a farmland weed detection system, including:
[0034] An image acquisition module that acquires a target image with weeds in the farmland;
[0035] A weed recognition module that uses YOLOv8 based on the RevColNet backbone network to construct a weed detection model and performs weed recognition based on the weed detection model;
[0036] A weed removal module for precisely removing weeds.
[0037] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the detection method of any one of the above-mentioned farmland weeds is implemented.
[0038] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the detection method of any one of the above-mentioned farmland weeds is implemented.
[0039] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the detection method of any one of the above-mentioned farmland weeds is implemented.
[0040] In the solution of this application, reconstructing the backbone network of YOLOv8 based on RevColNet can reduce the computational complexity and the number of parameters of the model while improving the model's ability to extract features. When using the improved weed detection model for weed recognition, weeds in farmland can be recognized with higher accuracy, and the computational complexity is low, and the efficiency of weed recognition is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 is a schematic flowchart of the detection method of farmland weeds provided by the embodiment of the present invention;
[0043] Figure 2 is one of the schematic structural diagrams of the weed detection model provided by the embodiment of the present invention;
[0044] Figure 3 is another schematic structural diagram of the weed detection model provided by the embodiment of the present invention;
[0045] Figure 4 is yet another schematic structural diagram of the weed detection model provided by the embodiment of the present invention;
[0046] Figure 5 is still another schematic structural diagram of the weed detection model provided by the embodiment of the present invention;
[0047] Figure 6 is yet still another schematic structural diagram of the weed detection model provided by the embodiment of the present invention;
[0048] Figure 7It is the sixth structural schematic diagram of the weed detection model provided by the embodiments of the present invention;
[0049] Figure 8 It is the seventh structural schematic diagram of the weed detection model provided by the embodiments of the present invention;
[0050] Figure 9 It is the eighth structural schematic diagram of the weed detection model provided by the embodiments of the present invention;
[0051] Figure 10 It is the structural schematic diagram of the detection system for farmland weeds provided by the embodiments of the present invention;
[0052] Figure 11 It is the entity structural schematic diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0053] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0054] With the rapid development of computer technology, convolutional neural networks have achieved good results in weed recognition. In recent years, the YOLO series of deep learning models have been more and more widely used in the field of target recognition and have achieved better performance than other models in multiple visual tasks. Therefore, some scholars have begun to apply the YOLO series of models to the field of agricultural recognition. Among them, Dong Hui et al. used an embedded SA module in the feature extraction part to optimize the feature extraction ability of the YOLOv4 model and improved the detection accuracy by optimizing the detection head. Guo Bozhang et al. proposed an improved YOLOv5 model integrating an attention mechanism and used stochastic gradient descent during model training to achieve accurate recognition in the case where the similarity between weeds and crops is relatively high. The latest detection model YOLOv8 achieved 92.1% in mAP50 and 62.3% in mAP50-95 on the weed25 dataset.
[0055] Through analysis, although the detection accuracy of various current models for recognition has achieved good results, the computing and storage resources of the core processing device of the weeding robot are limited. Therefore, it still faces problems such as high computational complexity, a large number of model parameters, and a large model size. To identify weeds, while ensuring the recognition accuracy, further research on solutions to reduce the number of model parameters and computational complexity is still needed. Therefore, this paper proposes an improved model based on the newly developed YOLOv8. By fixing the improved device to the vision module of the weeding robot for real-time scanning and detection, and returning the coordinates of the weeds through the positioning module, precise weeding is achieved.
[0056] Figure 1 It is a schematic flowchart of the method for detecting farmland weeds provided by the embodiment of the present invention.
[0057] As Figure 1 shown, this embodiment provides a method for detecting farmland weeds, including:
[0058] Step 101, collect a target image of the farmland with weeds;
[0059] Step 102, use YOLOv8 based on the RevColNet backbone network to construct a weed detection model, and perform weed recognition based on the weed detection model;
[0060] Step 103, precisely remove the weeds.
[0061] Among them, YOLOv8 is a commonly used object detection model, but there are still many problems with the YOLOv8 model. For example, the resolution of the small target annotation box is low, and the distribution is dense and easy to overlap; small target detection is easily affected by the image background and noise; it is not easy to calculate the classification and localization loss of small targets. Based on this, using the conventional YOLOv8 model cannot meet the work tasks of weed detection. Based on this, in the solution of this application, the conventional YOLOv8 model is improved based on RevColNet, while reducing the computational complexity and the number of parameters of the model, enhancing the ability of the model to extract features.
[0062] In practice, the backbones of the current YOLO series models are all top-down structures. During the feature extraction process, the information contained in the image will be lost to a certain extent, and the performance of the model will also be lost to a certain extent. In the solution of this application, the RevColNet (Reversible Column Networks) used is a reversible connection multi-column network, which is a multi-column structure.
[0063] In practical applications, in step 103, the weeds are precisely removed. Specifically, the position of the weeds can be accurately located in the form of coordinates, and the weeding robot is controlled to remove the weeds. That is to say, the weed recognition model provided in this embodiment can finally output the coordinates of the weeds in the target image of the farmland.
[0064] In an exemplary embodiment, a weed detection model is constructed by using YOLOv8 based on the RevColNet backbone network, including:
[0065] Reconstruct the backbone network of YOLOv8 based on RevColNet to obtain a new backbone network RevCol;
[0066] Introduce a fused dilated residual attention module, which improves the recognition ability of occluded targets;
[0067] Introduce the GSConv module and the VoV-GSCSPC module based on depthwise separable convolution for model lightweight processing;
[0068] Based on the minimum point distance, improve the bounding box regression loss function of the YOLOv8 model.
[0069] The solution of this embodiment has the following beneficial effects:
[0070] By redesigning the backbone of YOLOv8, the multi-scale fusion of feature information in different layers is strengthened, and by restricting the number of RevCol columns, the computational complexity and the number of model parameters of the model are significantly reduced.
[0071] Introducing a fused dilated residual attention module can help the model more effectively fuse features at different layers and improve the detection accuracy of the model.
[0072] By introducing the GSConv and VoVGSCSPC modules, while ensuring the detection accuracy and generalization ability of the model, the number of model parameters and the scale of the model are greatly reduced.
[0073] The bounding box regression loss function provided in this embodiment not only includes all relevant factors considered in the existing loss function, that is, overlapping or non-overlapping regions, the distance of the center point, and the deviation of width and height, but also simplifies the calculation process. And the bounding box regression loss function further improved on this basis helps the regression of samples by using the auxiliary bounding box to calculate the loss, and finally the proposed bounding box regression loss function effectively improves the detection accuracy of the model.
[0074] In an exemplary embodiment, the backbone network RevCol includes multiple columns, each column representing an input. The starting position of each column contains low-level detail information. As the image channels are compressed, high-level semantic information is extracted at the end of each column. A reversible connection design is adopted between columns to ensure lossless information transfer between columns. At the same time, supervision is added at the very end of each column to constrain the feature extraction of each column.
[0075] In practice, the low-level detail information can be represented by low-level information, which generally refers to some small detail information in an image, such as edges, corners, colors, pixels, gradients, etc. This information can be obtained through filters, SIFT, or HOG.
[0076] The high-level semantic information can be represented by feature. It is based on low-level information and can be used for the recognition and detection of objects or object shapes in an image, with richer semantic information. This high-level semantic information can be understood as the information obtained by integrating a series of information such as environmental information and texture information. Subsequently, this high-level semantic information can be used for judgment during classification or detection.
[0077] Figure 2 It is one of the structural schematic diagrams of the weed detection model provided by the embodiments of the present invention.
[0078] Figure 2 It exemplifies the macroscopic structure of RevColNet used in the solution of the present application. As Figure 2 shown, RevColNet adopts a multi-input design. The starting position of each column contains low-level information. As the image channels are compressed, the semantic information in the feature is extracted at the end of the column. A Reversible connection design between columns ensures lossless information transfer between columns. At the same time, supervision is added at the very end of each column to constrain the feature extraction of each column.
[0079] Figure 3 It is the second structural schematic diagram of the weed detection model provided by the embodiments of the present invention.
[0080] Figure 4 It is the third structural schematic diagram of the weed detection model provided by the embodiments of the present invention.
[0081] Figure 3 and Figure 4Illustrates the microstructure of RevColNet used in the solution of this application. In Figure a, each level module extracts features through downsampling and ConvNeXt. Figure b shows the Reversible connection design between columns.
[0082] In practical applications, the Reversible connection design between columns in Figure b conforms to the following calculation formulas (1) and (2):
[0083] Xt = Ft(Xt-1, Xt-m+1) + γXt-m (1)
[0084] Xt-m = γ -1 [X t - Ft(Xt-1, Xt-m+1)] (2)
[0085] Among them, formula (1) shows the interaction between each level in the second column, and the output X t is determined by three inputs. The output of the previous level is X t-1 , and the output of the next level in the previous column is X t-m+1 , and the two outputs are adjusted in shape through the F t operation to be consistent with the output X t-m of the previous column. The Ft() operation includes a fusion module and n convolutional modules, and finally the obtained features are added to γ times of X t-m for addition operation. Formula 2 shows the reversibility of the network, ensuring lossless information transmission.
[0086] Figure 5 is the fourth structural schematic diagram of the weed detection model provided by the embodiments of the present invention.
[0087] Figure 5 Illustrates the structure of a reconstructed backbone network RevCol.
[0088] As Figure 5 shown, in order to avoid the complexity and the number of parameters of the model increasing due to the over-bloated backbone network, in implementation, the number of columns of RevCol can be set to 2, and at the same time, the operations in the feature fusion module (Fusion Block) are reconstructed. For high-level semantic information, only one composite operation is performed to achieve downsampling, that is, convolution, batch normalization, and activation function. For low-level detailed information, convolution combined with upsampling is used to replace the original upsampling operation, and the C2f module of YOLOv8 is used to replace the ConvNeXt module in the level.
[0089] In practical applications, the above-mentioned feature fusion module is a module in the backbone network, which adjusts the feature channels of different input sizes to the same size of output.
[0090] In an exemplary embodiment, the fused dilated residual attention module is used for:
[0091] Performing a standard 3×3 convolution operation on the data input to the weed detection model, and extracting features by batch normalization and activation using an activation function;
[0092] After expanding the 3×3 convolution, obtaining a semantic residual through a BN layer;
[0093] Connecting all branch feature maps, using pointwise convolution to merge all feature maps, and generating a final residual corresponding to the data input to the weed detection model;
[0094] Fusing the final residual and the input data to construct a final feature representation.
[0095] Figure 6 It is the fifth structural schematic diagram of the weed detection model provided by the embodiments of the present invention.
[0096] In practical applications, due to the multi-scale characteristics of traditional YOLO series models, the recognition effect of occluded targets is not good, it is easy to make wrong classifications, and there are also certain deficiencies in small target detection. Therefore, this paper introduces a fused dilated residual (DWR) attention module, which is applied to the deep layer of the network. The multi-branch structure is used to adapt to the requirements of different receptive field sizes in one layer. Its structure is as Figure 6 shown. For the input feature map, first perform a standard convolution operation with a 3×3 kernel (convolution kernel), and then combine a batch normalization layer and a ReLU layer to extract features. Since each output channel contains several small spatial regions that need to be refined, the entire output is a collection of these regions. Then expand the depth 3×3 convolution, extract semantic information from these regions, and then obtain a semantic residual through a BN layer to further analyze the semantic information from the regional features. Then connect all branch feature maps and use pointwise convolution to merge all feature maps to generate a final residual corresponding to the input feature map. Finally, fuse the final residual and the input feature map to construct a stronger and more comprehensive feature representation.
[0097] In an exemplary embodiment, the fused dilated residual attention module is provided with a plurality of channels, and the number of dilated depth convolution channels with the lowest dilation rate is set to twice that of other channels.
[0098] In practical applications, regardless of the stage, the features extracted with small receptive fields are relatively important. Therefore, the number of dilated depth convolution channels with the lowest dilation rate is set to twice that of other channels.
[0099] Figure 7 It is the sixth structural schematic diagram of the weed detection model provided by the embodiments of the present invention.
[0100] In an exemplary embodiment, inspired by the DWR module, a C2fDWR module is further designed in this embodiment. Its structural diagram is as Figure 7 shown. In order to make up for the deficiency of the model in occluded target recognition, in this embodiment, the designed C2fDWR module is used to replace and reconstruct the C2f module of the backbone RevCol, so that the performance of the model is further improved.
[0101] In an exemplary embodiment, the GSConv module is used for:
[0102] Based on the number of input channels, obtain the first number of output channels through a standard convolution;
[0103] Based on the first number of output channels, obtain the second number of output channels through a depthwise separable convolution;
[0104] Connect and shuffle the first number of output channels and the second number of output channels to obtain the number of output channels.
[0105] Figure 8 It is the seventh structural schematic diagram of the weed detection model provided by the embodiments of the present invention.
[0106] In practical applications, in order to make the model better applicable to edge terminal devices, lightweight design is essential. The lightweight of the model can not only reduce the computational resource cost of the model, but also improve the detection speed of the model. In the solution of this embodiment, depthwise separable convolution is used to replace the traditional convolution module. Different from the traditional convolution method, depthwise separable convolution performs hierarchical processing on the feature layers of the input channels, which can effectively reduce the problem of large computational volume existing in multiple channels. However, there is a situation of information loss between channels in conventional depthwise separable convolution. For this reason, a GSConv lightweight convolution module based on depthwise separable convolution is introduced in this paper. Its main structure is as Figure 7 shown. The number of input channels is C1, and the number of output channels is C2. First, based on the input channel number C1, a standard convolution is performed to obtain the first number of output channels. Then, the same number of channels, that is, the second number of output channels, is obtained through a depthwise separable convolution. Finally, a Concat connection and shuffle operation are performed on the two results to obtain the final number of output channels. That is to say, both the first number of output channels and the second number of output channels are equivalent to one-half of the final number of output channels.
[0107] In the solution of this embodiment, GSConv can well preserve multi-channel information, while reducing computing resources and improving the expression ability of image features.
[0108] Figure 9 It is the eighth structural schematic diagram of the weed detection model provided by the embodiment of the present invention.
[0109] In an exemplary embodiment, Figure 9 The structure of a VoV-GSCSPC module is exemplified. In the solution of this embodiment, it is disclosed that a GSConv module can be introduced in the Neck layer, using GSConv to replace the standard convolution, reducing the number of parameters and the amount of computation of the neck module, and using the VoV cross-stage partial network module VoV-GSCSPC based on GSConv and the lightweight bottleneck layer GSbottleneck to replace the CSP module of the original model, further improving the performance of YOLOv8.
[0110] In an exemplary embodiment, based on the minimum point distance, the bounding box regression loss function of the YOLOv8 model is improved, including:
[0111] Determine the similarity between the predicted bounding box and the actual annotated bounding box during the bounding box regression process, and based on the similarity, calculate the key point distance between the predicted bounding box and the actual annotated bounding box to improve the accuracy of the loss metric.
[0112] In practical applications, the bounding box regression loss function used by the YOLOv8 model is the CIou loss. CIou takes the aspect ratio of the Bounding box into the loss function on the basis of DIou, further improving the regression accuracy. However, most BBR loss functions represented by CIou may have the same value under different prediction results, which reduces the convergence speed and accuracy of the bounding box regression. Based on this, in the solution of this embodiment, a new loss function MPDIou based on the minimum point distance is introduced as the bounding box regression loss function for improving the YOLOv8 model. By comparing the similarity between the predicted bounding box and the actual annotated bounding box during the bounding box regression process and directly calculating the key point distance between the predicted box and the true box, a more accurate loss metric method is provided. Specifically, the calculation formula of the loss function provided in this embodiment is as follows:
[0113]
[0114]
[0115]
[0116] LMPDloU = 1 - MPDIoU
[0117] Where A and B are the predicted bounding box and the ground truth bounding box respectively, and w and h represent the width and height of the input image respectively. and represent the coordinates of the upper left corner point of A and the coordinates of the lower right corner point of A respectively. and represent the coordinates of the upper left corner point of B and the coordinates of the lower right corner point of B respectively.
[0118] In an exemplary embodiment, the bounding box regression loss function of the improved YOLOv8 model includes:
[0119] Introduce a scale factor ratio to control the size of the auxiliary bounding box to calculate the loss;
[0120] When the value of ratio is set to be greater than 1, generate an auxiliary bounding box with a larger scale relative to the actual bounding box to calculate the loss;
[0121] When the value of ratio is set to be less than 1, generate an auxiliary bounding box with a smaller scale to calculate the loss, so that the absolute value of the regression gradient is greater than the absolute value of the actual bounding box IoU gradient.
[0122] In practical applications, the solution of this embodiment can also further add Inner IoU on the basis of the above loss function MPDIoU, and propose a new loss function InnerMPDIou as the model bounding box regression loss function. Specifically, different from traditional improvement methods, InnerIoU can effectively accelerate the bounding box regression by analyzing the use of auxiliary bounding boxes with different scales to calculate the loss during the regression process. Inner IoU introduces a scale factor ratio to control the size of the auxiliary bounding box to calculate the loss. Usually, the value range of the scale factor ratio is [0.5, 1.5]. When the value of ratio is set to be greater than 1, an auxiliary bounding box with a larger scale relative to the actual bounding box will be generated to calculate the loss, which can expand the effective range of the regression and has a certain gain for the regression of low IoU samples. When the value of ratio is set to be less than 1, an auxiliary bounding box with a smaller scale will be generated to calculate the loss, which can make the absolute value of the regression gradient greater than the absolute value of the actual bounding box IoU gradient, contribute to the regression of high IoU samples, and achieve the effect of accelerating convergence. The calculation formula of Inner MPDIoU is as follows:
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] union = (w gt * h gt ) * (ratio) 2 + (w * h) * (ratio) 2 - inter
[0133]
[0134] L InnerMPDIoU = L MPDIoU + IoU - IoU inner
[0135] On the weed25 dataset, compared with YOLOv8, the improved model proposed in this study reduces the computational complexity by 35.8%, the number of parameters by 35.36%, and the model size by 30.15%. At the same time, the mAP50 value and the mAP50 - 95 value are respectively increased to 93.8% and 63.4%, and the accuracy value is increased to 92.9%. The performance of the improved model exceeds that of the original YOLOv8 model in many aspects.
[0136] The following describes the farmland weed detection system provided by the present invention. The farmland weed detection system described below can be mutually corresponding and referred to with the farmland weed detection method described above.
[0137] Figure 10 It is a schematic structural diagram of the farmland weed detection system provided by an embodiment of the present invention.
[0138] As Figure 10 shown, the farmland weed detection system provided in this embodiment includes:
[0139] An image acquisition module 1001, which acquires a target image with weeds in the farmland;
[0140] A weed recognition module 1002, which constructs a weed detection model using YOLOv8 based on the RevColNet backbone network and performs weed recognition based on the weed detection model;
[0141] A weed removal module 1003, which is used to precisely remove weeds.
[0142] The specific implementation method of the farmland weed detection system provided in this embodiment can be implemented with reference to the above embodiment, and will not be elaborated here.
[0143] Figure 11 An example of a schematic physical structure diagram of an electronic device is shown as Figure 11 shown. The electronic device may include: a processor 1110, a communications interface 1120, a memory 1030, and a communication bus 1140. Among them, the processor 1110, the communications interface 1120, and the memory 1130 complete mutual communication through the communication bus 1140. The processor 1110 can call the logical instructions in the memory 1130 to execute the farmland weed detection method, and the method includes:
[0144] Collect a target image of the farmland with weeds;
[0145] Use YOLOv8 based on the RevColNet backbone network to build a weed detection model, and perform weed recognition based on the weed detection model;
[0146] Precisely remove the weeds.
[0147] The weed detection model is obtained through the following method:
[0148] Reconstruct the backbone network of YOLOv8 based on RevColNet to obtain a new backbone network RevCol;
[0149] Introduce a fusion dilated residual attention module, and the fusion dilated residual attention module improves the recognition ability of occluded targets;
[0150] Introduce the GSConv module and the VoV-GSCSPC module based on depthwise separable convolution for model lightweight processing;
[0151] Improve the bounding box regression loss function of the YOLOv8 model based on the minimum point distance.
[0152] In addition, when the logical instructions in the above-mentioned memory 1130 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0153] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the detection method of farmland weeds provided by the above-mentioned various methods. The method includes:
[0154] Collect a target image of the farmland with weeds;
[0155] Use YOLOv8 based on the RevColNet backbone network to build a weed detection model, and perform weed recognition based on the weed detection model;
[0156] Precisely remove the weeds.
[0157] The weed detection model is obtained in the following manner:
[0158] Reconstruct the backbone network of YOLOv8 based on RevColNet to obtain a new backbone network RevCol;
[0159] Introduce a fused dilated residual attention module, and the fused dilated residual attention module improves the recognition ability for occluded targets;
[0160] Introduce a GSConv module and a VoV-GSCSPC module based on depthwise separable convolutions for model lightweight processing;
[0161] Improve the bounding box regression loss function of the YOLOv8 model based on the minimum point distance.
[0162] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the detection method of farmland weeds provided by the above-mentioned various methods. The method includes:
[0163] Collect the target images of weeds in the farmland;
[0164] Use YOLOv8 based on the RevColNet backbone network to build a weed detection model, and identify weeds based on the weed detection model;
[0165] Precisely remove the weeds.
[0166] The weed detection model is obtained in the following way:
[0167] Reconstruct the backbone network of YOLOv8 based on RevColNet to obtain a new backbone network RevCol;
[0168] Introduce a fused dilated residual attention module, which improves the recognition ability of occluded targets;
[0169] Introduce the GSConv module and VoV-GSCSPC module based on depthwise separable convolution for model lightweight processing;
[0170] Improve the bounding box regression loss function of the YOLOv8 model based on the minimum point distance.
[0171] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0172] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for detecting farmland weeds, characterized in that, Including: Collecting target images of weeds in farmland; Using YOLOv8 based on the RevColNet backbone network to build a weed detection model and identifying weeds based on the weed detection model; Precisely removing the weeds; The using of YOLOv8 based on the RevColNet backbone network to build a weed detection model includes: Reconstructing the backbone network of YOLOv8 based on RevColNet to obtain a new backbone network RevCol; Introducing a fused dilated residual attention module, which improves the recognition ability for occluded targets; Introducing a GSConv module and a VoV-GSCSPC module based on depthwise separable convolution for model lightweight processing; Improving the bounding box regression loss function of the YOLOv8 model based on the minimum point distance.
2. The detection method of farmland weeds according to claim 1, characterized in that, The backbone network RevCol includes multiple columns, each column represents an input, the starting position of each column contains low-level detailed information, and as the image channels are compressed, high-level semantic information is extracted at the end of each column; a reversible connection design is adopted between columns to ensure lossless information transfer between columns, and at the same time, supervision is added at the very end of each column to constrain the feature extraction of each column.
3. The detection method of farmland weeds according to claim 1, wherein The fused dilated residual attention module is used for: Performing a standard 3×3 convolution operation on the data input to the weed detection model, and extracting features through batch normalization and activation using an activation function; After expanding the 3×3 convolution, obtaining a semantic residual through a BN layer; Connecting all branch connection feature maps, using pointwise convolution to merge all feature maps, and generating a final residual corresponding to the data input to the weed detection model; Fusing the final residual and the input data to construct a final feature representation.
4. The detection method of farmland weeds according to claim 3, characterized in that, The fused dilated residual attention module is provided with a number of channels, and the number of dilated depth convolution channels with the lowest dilation rate is set to twice that of other channels.
5. The detection method of farmland weeds according to claim 1, characterized in that, The GSConv module is used for: Based on the number of input channels, obtaining a first output channel number through a standard convolution; Based on the first output channel number, obtaining a second output channel number through a depthwise separable convolution; Connecting and shuffling the first output channel number and the second output channel number to obtain an output channel number.
6. The detection method of farmland weeds according to claim 1, characterized in that The improving of the bounding box regression loss function of the YOLOv8 model based on the minimum point distance includes: Determining the similarity between the predicted bounding box and the actual annotated bounding box during the bounding box regression process, and calculating the key point distance between the predicted bounding box and the actual annotated bounding box based on the similarity to improve the accuracy of loss measurement.
7. The detection method of farmland weeds according to claim 6, characterized in that, The improving of the bounding box regression loss function of the YOLOv8 model includes: Introducing a scale factor ratio to control the size of the auxiliary bounding box to calculate the loss; When the value of ratio is set to be greater than 1, generating an auxiliary bounding box with a larger scale relative to the actual bounding box to calculate the loss; When the value of ratio is set to be less than 1, a smaller-scale auxiliary bounding box is generated to calculate the loss, making the absolute value of the regression gradient greater than the absolute value of the actual bounding box IoU gradient.
8. A detection system for farmland weeds, applied to the detection method of farmland weeds according to any one of claims 1-7, characterized in that, Including: An image acquisition module that acquires a target image with weeds in the farmland; A weed recognition module that uses YOLOv8 based on the RevColNet backbone network to build a weed detection model and performs weed recognition based on the weed detection model; A weed removal module for precisely removing the weeds.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the detection method of farmland weeds as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Small target detection method and system for field corn weeds and application
CN117036950A