Crop seedling core real-time detection model and eggplant seedling core detection method

By improving the YOLOv8 network architecture, the lightweight backbone and neck network is built, the problem of accurate identification of eggplant seedlings in complex environments is solved, and efficient and fast eggplant seedlings detection is achieved, which is suitable for embedded equipment for agricultural machines.

CN120339794APending Publication Date: 2025-07-18GUANGZHOU INST OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510434753.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate identification of eggplant seedlings in complex environments, especially on embedded devices of agricultural machines, and the calculation volume is large, making it difficult to deploy on mobile devices.

Method used

A real-time detection model of crop seedling hearts based on the YOLOv8 network architecture is adopted. Through lightweight backbone network, lightweight neck network and head network, combined with MobileNetV3 block and C2f-Faster layer, a lightweight backbone network and neck network are built to improve feature extraction and fusion capabilities and reduce computing volume.

Benefits of technology

It realizes high-precision and fast eggplant seedling core detection on embedded devices, reduces computing resources, improves detection accuracy and speed, and is suitable for mobile devices and embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339794A_ABST
    Figure CN120339794A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of crop detection, and discloses a crop seedling core real-time detection model and an eggplant seedling core detection method, and the model comprises a lightweight backbone network, a lightweight neck network and a head network. The lightweight backbone network comprises a plurality of feature extraction modules constructed based on MobileNetV3 blocks, and the feature extraction modules are used for extracting features of input crop seedling images to obtain feature maps of multiple scales; the lightweight neck network comprises a plurality of fusion blocks formed by an up-sampling layer, a feature splicing layer and a C2f-Faster layer and a plurality of series blocks formed by a convolutional layer, a feature splicing layer and a C2f-Faster layer, and is used for bidirectionally fusing the feature maps extracted by the lightweight backbone network to obtain fusion features of multiple scales; the head network is composed of a convolutional layer and a two-dimensional convolutional layer and is used for decoding the fusion features of the multiple scales output by the lightweight neck network to obtain a detection result. The method can be deployed in an embedded device or a mobile device, and has high detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of crop detection, and particularly relates to a real-time detection model for crop seedling centers and a method for detecting eggplant seedling centers. Background Art

[0002] Real-time identification of eggplant seedling centers is a key link in intelligent and precise pesticide spraying for eggplant seedlings. The changes in eggplant seedlings from the early growth stage to the late growth stage are relatively large, and there are different seedling center characteristics in the early, middle, and late stages. The characteristics of different types of eggplant seedlings will also vary. Therefore, there are numerous characteristics of eggplant seedling centers, and they will change as the eggplant seedlings grow. In the natural environment, factors such as weeds beside the eggplant seedlings, the planting interval of eggplant seedlings, and natural light will all affect the target detection of eggplant seedling centers. Therefore, the complex growth environment and diverse appearance characteristics of eggplant seedlings make it difficult to achieve accurate identification of eggplant seedling centers.

[0003] Although deep learning technology has been applied in the agricultural field, in terms of object detection, the current mainstream detection models include: YOLO series, R-CNN series, SSD, etc. For example, Zhang Y et al. detected whether tomatoes were diseased based on an improved Faster-RCNN model, and the best average accuracy reached 98.54%. However, this model has a large amount of computation and high requirements for hardware performance, making it difficult to be deployed on mobile devices or embedded devices, and the identification of eggplant seedling centers is not accurate enough.

[0004] Currently, there is little research on the real-time detection of eggplant seedling centers in complex environments, and there is no detection model that is suitable for deployment on embedded devices of agricultural machines and has high detection accuracy. Summary of the Invention

[0005] The purpose of the present invention is to provide a real-time detection model for crop seedling centers, a method for detecting eggplant seedling centers, an electronic device, and a computer-readable storage medium, which are suitable for deployment on embedded devices of agricultural machines and have high detection accuracy.

[0006] The first aspect of the present invention discloses a real-time detection model for crop seedling centers based on the YOLOv8 network architecture, including:

[0007] A lightweight backbone network, a lightweight neck network, and a head network;

[0008] The lightweight backbone network includes multiple feature extraction modules constructed based on MobileNetV3 blocks, which are used to extract the features of the input crop seedling images and obtain feature maps of multiple scales;

[0009] The lightweight neck network includes multiple fusion blocks composed of an upsampling layer, a feature splicing layer, and a C2f-Faster layer, and multiple tandem blocks composed of a convolutional layer, a feature splicing layer, and a C2f-Faster layer, which are used to fuse the feature maps extracted by the lightweight backbone network bidirectionally to obtain fused features at multiple scales.

[0010] The head network consists of a convolutional layer and a two-dimensional convolutional layer, which is used to decode the fused features at multiple scales output by the lightweight neck network to obtain detection results.

[0011] In some embodiments, the head and tail of the C2f-Faster layer are 1×1 convolutional layers, and in the middle are a channel splitting layer, a Faster block layer, and a feature splicing layer in sequence. The feature splicing layer is used to splice the splitting information of the channel splitting layer and the Faster block layer.

[0012] In some embodiments, the Faster block layer has an inverted residual structure, including a partial convolution module and two consecutive 1×1 convolution modules.

[0013] In some embodiments, when training the real-time crop seedling center detection model, Grad-CAM heatmaps are used to visually analyze the feature information output by the real-time crop seedling center detection model.

[0014] In some embodiments, an SE attention module is further deployed after the depth convolution in the feature extraction module. The SE attention module is used to compress the feature map of W×H×C containing global information into a feature vector of 1×1×C.

[0015] In some embodiments, the number of output channels of the first fully connected layer in the SE attention module is 1 / 4 of the number of channels in the dilation layer in the SE attention module.

[0016] In some embodiments, h-swish is used as the activation function in the deep network of the MobileNetV3 block, and ReLU is used as the activation function in the shallow network of the MobileNetV3 block.

[0017] The second aspect of the present invention discloses a method for detecting the eggplant seedling center, including:

[0018] Collecting eggplant seedling images;

[0019] Inputting the eggplant seedling images into any of the above-mentioned trained real-time crop seedling center detection models to obtain eggplant seedling center recognition results.

[0020] A third aspect of the present invention discloses an electronic device, including a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the eggplant seedling center detection method disclosed in the second aspect.

[0021] A fourth aspect of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the eggplant seedling center detection method disclosed in the second aspect.

[0022] The beneficial effects of the present invention are as follows: on the basis of YOLOv8, a lightweight backbone network is formed by introducing a feature extraction module Mblock based on MobileNetV3, and a simple and efficient lightweight neck network using C2f-Faster layers is designed, realizing the miniaturization of the model, improving the inference speed, being able to better obtain diverse appearance features of the eggplant seedling center, meeting the requirements of real-time performance and low power consumption on the premise of ensuring the model accuracy, being capable of being deployed in embedded devices or mobile devices, and having high detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings herein show specific examples of the technical solutions of the present invention and form a part of the specification together with the specific embodiments, for explaining the technical solutions, principles and effects of the present invention.

[0024] Unless otherwise specified or defined, in different drawings, the same reference numerals represent the same or similar technical features, and for the same or similar technical features, different reference numerals may also be used for representation.

[0025] Figure 1 is the network architecture diagram of the real-time detection model of crop seedling centers in the embodiments of the present invention;

[0026] Figure 2 is the network structure diagram of Mblock equipped with the SE module in the embodiments of the present invention;

[0027] Figure 3 is the network structure diagram of C2f-faster in the embodiments of the present invention;

[0028] Figure 4 is the network structure diagram of a partial convolution module in the embodiments of the present invention;

[0029] Figure 5 is the schematic diagram of the seedling centers of eggplant seedlings at different growth stages in the embodiments of the present invention;

[0030] Figure 6 is the schematic diagram of the seedling centers of eggplant seedlings in different weathers in the embodiments of the present invention;

[0031] Figure 7 It is a visualization comparison chart of Grad-CAM heatmaps with different output scales of YOLOv8 and EggYOLOPlant;

[0032] Figure 8 It is a comparison chart of the ablation detection effects of Mblock and C2f-faster;

[0033] Figure 9 It is a comparison chart of the detection effects of different detection models. Detailed implementation manners

[0034] Unless otherwise specified or defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs. In the case of combining the technical solutions of the present invention with real scenarios, all technical and scientific terms used herein may also have meanings corresponding to the purpose of implementing the technical solutions of the present invention. The "first, second..." used herein are only for distinguishing names and do not represent specific quantities or orders. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0035] It should be noted that when an element is considered to be "fixed to" another element, it can be directly fixed to the other element or there can be an intermediate element; when an element is considered to be "connected to" another element, it can be directly connected to the other element or there can be an intermediate element at the same time; when an element is considered to be "mounted on" another element, it can be directly mounted on the other element or there can be an intermediate element at the same time. When an element is considered to be "provided in" another element, it can be directly provided in the other element or there can be an intermediate element at the same time.

[0036] Unless otherwise specified or defined, the "said" and "the" used herein refer to the technical features or technical contents mentioned or described before the corresponding positions, and the technical features or technical contents can be the same as or similar to the technical features or technical contents they mention. In addition, the terms "including" and "having" used herein and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0037] Complex model structures often exhibit excellent detection performance in complex scenarios, but they also limit the model's speed, such as two-stage object detection algorithms like Faster RCNN. On the other hand, one-stage object detection algorithms, such as YOLO and SSD, offer faster detection speeds but may sacrifice some accuracy.

[0038] Considering the performance of the embedded devices of agricultural machines and the diversity of the environment they face, the fast detection characteristics of one-stage algorithms are more applicable. Therefore, the present invention proposes a real-time detection model for crop seedling centers, namely a deep learning model called EggYOLOPlant, which aims to achieve a good balance between detection accuracy and inference speed as much as possible.

[0039] In this embodiment, EggYOLOPlant is deployed on the embedded device of an agricultural machine to detect the eggplant seedling center in real time, also known as the real-time detection model for the eggplant seedling center, to solve the object detection problem caused by the diverse appearance features of the eggplant seedling center in a complex growth environment. The real-time detection model for the eggplant seedling center is based on the improved YOLOv8 architecture, and a lightweight backbone network based on Mblock and a slim-Neck network (lightweight neck network) are constructed to replace the original backbone network (Backbone) and neck network (Neck) respectively, so as to improve the feature extraction and fusion capabilities and enhance the detection speed. Through the analysis of the interpretable Gradient-weighted Class Activation Mapping (Grad-CAM) technology, it is found that EggYOLOPlant can well filter the background attention area, and the attention area for the eggplant seedling center is relatively concentrated, making it achieve accurate, fast, and computationally resource-saving effects in the eggplant seedling center object detection.

[0040] The experimental results show that compared with the original YOLOv8, EggYOLOPlant has improved the precision (P) and average precision (AP) by 4 percentage points and 3.8 percentage points respectively, while the number of parameters has decreased by about 30%, and the real-time detection speed has increased by about 11.8%. Compared with Faster-RCNN, YOLOv5s, and YOLO11s, the AP of this model has increased by 3.2 percentage points, 1.8 percentage points, and 1.5 percentage points respectively, and the detection speed has increased by about 24 times, 1.59 times, and 1.61 times respectively. Therefore, the EggYOLOPlant model is suitable for scenarios with limited memory and computing power but requiring high accuracy, and can perform eggplant seedling center object detection on mobile devices or embedded devices. It should be noted that EggYOLOPlant is not limited to detecting the eggplant seedling center and can also perform object detection on similar plants.

[0041] Specifically, YOLOv8 is a single-stage target detection model. With its advanced network structure and training strategy techniques, it has good performance in speed and accuracy and is currently the mainstream target detection model. YOLOv8 inherits the original YOLO target detection network, and its network architecture consists of three parts: Backbone, Neck, and Head. Backbone of YOLOv8 is responsible for feature extraction, and is composed of blocks (Blocks) composed of multiple convolutional layers (Conv) and C2f (CrossStage Partial-faster) network modules stacked in series; the Neck part is mainly used to achieve feature fusion, and adopts the PAN-FPN structure that combines the feature pyramid network (FPN) with the path aggregation network (PANet); the Head part is responsible for decoding the features and finally generates the detection results through the decision-making process.

[0042] The network architecture of the real-time detection model of crop seedlings in this embodiment is as follows: Figure 1 As shown, it includes a lightweight backbone network, a lightweight neck network and a head network. The lightweight backbone network is used to extract the features of the input eggplant seedling image to obtain feature maps of multiple scales; the lightweight neck network is used to bidirectionally fuse the feature maps extracted by the lightweight backbone network to obtain fused features of multiple scales; the head network is used to decode the fused features of multiple scales output by the lightweight neck network to obtain the detection results.

[0043] Based on the original YOLOv8 network architecture, a lightweight Backbone with MobileNetV3 block (abbreviated as Mblock) was redesigned to replace the original Backbone and a slim-Neck (lightweight Neck) to replace the original Neck. The improved Backbone is built with lightweight and efficient Mblocks. Mblock, as a feature extraction module, can better capture the diverse appearance feature information of eggplant seedlings in complex environments.

[0044] Based on multiple factors such as convolution method, feature fusion structure and spatial pyramid pooling structure, slim-Neck is constructed to replace the original Neck. In slim-Neck, there are multiple fusion blocks consisting of upsampling layer, feature concatenation layer and C2f-Faster layer ( Figure 1 The left side of the lightweight neck network in the figure) and multiple series blocks consisting of convolutional layers, feature concatenation layers, and C2f-Faster layers ( Figure 1 By using the C2f-faster module instead of the original C2f module, redundant calculations and memory accesses can be reduced to achieve faster inference speed.

[0045] The Head part remains unchanged and consists of a convolutional layer and a two-dimensional convolutional layer.

[0046] The following elaborates on the efficiency improvement of the eggplant seedling heart real-time detection model by comparing the standard convolution in the original YOLOv8's Backbone with the depthwise convolution used in the Mblock.

[0047] Specifically, MobileNet is a network structure that uses depthwise separable convolutions. Depthwise separable convolutions are divided into depthwise convolution and pointwise convolution. Its feature extraction process is different from that of standard convolution, which directly extracts features through a convolution kernel. Depthwise convolution splits the convolution kernel into single-channel forms and performs convolution operations on each channel without changing the depth of the input feature map, obtaining an output feature map with the same number of channels as the input feature map. For example, for an input feature map F of D F ×D F ×M and passing it through a convolution kernel K of size D K ×D K ×M×N, the output feature map G of D F ×D F ×N is obtained, where D F is the height and width of the square input and output feature maps, D K is the spatial dimension of the square convolution kernel, and M and N are the number of channels of the input and output feature maps respectively.

[0048] Assuming a stride of 1, the working method of the standard convolution output feature map can be expressed as:

[0049]

[0050] The amount of computation required to use the standard convolution is:

[0051] D K ·D K ·M·N·D F ·D F .

[0052] And one filter in each channel number in the depthwise convolution can be expressed as:

[0053]

[0054] Among them, is a depthwise convolution kernel of size D K ×D K ×M, and the mth filter among them is used in the mth channel of the feature map F to extract the feature map in the mth channel

[0055] Then, the features output from the depthwise convolution are combined through 1×1 convolution (pointwise convolution) to generate new features, resulting in an output of D F ×D F ×N feature map G.

[0056] The computational cost of using depthwise separable convolution is:

[0057] D K ·D K ·M·D F ·D F +M·N·D F ·D F .

[0058] By comparing the computational costs of the two convolutions, it can be obtained that:

[0059]

[0060] Therefore, when using a 3×3 convolution kernel, the computational cost of depthwise separable convolution is about one-ninth to one-eighth less than that of standard convolution.

[0061] MobileNetV3 uses the Neural Architecture Search (NAS) technology in its network structure. It searches for the network to optimize each block under the premise of limited computational cost and number of parameters, and then uses the NetAdapt algorithm to optimize the number of convolution kernels in each layer to obtain a small network model. At the same time, MobileNetV3 also introduces a new activation function h-swish and a lightweight attention module Squeeze-and-Excite (SE).

[0062] Among them, the h-swish non-linear activation function is the hard version of swish. The swish non-linear function has the characteristics of having no upper bound but a lower bound, being smooth, and non-monotonic. When swish is used as the activation function instead of ReLU, it can significantly improve the accuracy of the network model. The formula of swish is as follows:

[0063] swish(x) = x·σ(x).

[0064] However, in the mobile environment, the cost of calculating the sigmoid function is relatively large, so its piecewise linear hard simulation is used Therefore, the formula of the hard version of swish, that is, h-swish, is as follows:

[0065]

[0066] Although optimized, the introduction of h-swish still incurs some latency costs. However, in deep networks, the cost of applying non-linearity decreases, and most of the benefits of swish are achieved by using them in the deep layers. Therefore, in the real-time eggplant seedling heart detection model of this embodiment, h-swish is only applied as the activation function in the deep network of MobileNetV3, and ReLU is still used as the activation function in the shallow network to reduce the computational cost.

[0067] Among them, the SE attention module is divided into global information embedding (Squeeze) and adaptive recalibration (Excitation). Global information embedding uses channel global average pooling to compress the feature map U of W×H×C containing global information into a feature vector Z of 1×1×C. The channel features of C feature maps are compressed into a single value, so that the generated channel-level statistical data Z contains global information. The definition formula is as follows:

[0068]

[0069] Then, an adaptive recalibration operation is required to obtain the relationship between channels. This operation uses two fully connected layers, defined as follows:

[0070] s = hard-σ(W2ReLU(W1z))

[0071] Among them, The first fully connected layer (FC) plays a role in dimensionality reduction, and the dimensionality reduction coefficient is the hyperparameter r. Then, it is activated by ReLU, and then the original dimension is restored in the second fully connected layer (FC), and activated by Hard-sigmoid.

[0072] Finally, the activation values of each learned channel are multiplied by the original features on the feature map U:

[0073]

[0074] In the real-time eggplant seedling heart detection model of this embodiment, the SE attention module is deployed behind the depth convolution in some Mblocks, effectively enhancing the feature extraction ability of Mblocks. At the same time, the size of the SE attention module is related to the size of the convolution bottleneck. Therefore, the output channel number of the first fully connected layer in the SE attention module is fixed at 1 / 4 of the channel number of the dilation layer in the SE attention module, which can improve the accuracy of the model while moderately increasing the number of parameters, and there is no obvious latency cost. The structure of the Mblock equipped with the SE module is as Figure 2 shown.

[0075] For the neck network, in order to reduce the computational redundancy and memory access of the neck network, considering the convolution method and the feature fusion structure, without destroying the original gradient path, this embodiment proposes a C2f-faster module to replace the original C2f module, forming a lightweight neck network. The structure of C2f-faster is as Figure 3 shown. Its head and tail are the same as those of C2f, both being a 1×1 convolutional layer, which are used to amplify and downscale the channels respectively. The middle part is successively a channel splitting layer, a Faster block layer, and a feature concatenation layer. First, the "Channel split" operation is used to split the channels. After "Channel split" in C2f, it is followed by a Bottleneck block, while after "Channel split" in the C2f-faster module, it is followed by a FasterNet block (composed of n Faster blocks). The feature concatenation layer concatenates the split information of the channel splitting layer and the FasterNet block layer, that is, uses the idea of gradient splitting to transfer the feature information of each block, and concatenates it with the previous split information. Finally, it passes through a 1×1 convolution to adjust the output channels.

[0076] Compared with the Bottleneck block, using the FasterNet block can obtain stronger feature extraction ability and lower latency. Different from the structure composed of two 3×3 standard convolutions in the Bottleneck block, as Figure 4 shown, each Faster block is composed of a partial convolution module (Partial Convolution: PConv) and two consecutive 1×1 convolution modules, forming an inverted residual structure. Among them, the 1×1 convolution can effectively utilize the information of all channels. PConv is a simple and fast convolution method. It only uses standard convolution for spatial feature extraction on some input channels, while the other input channels remain unchanged, thus reducing redundant calculations and memory access.

[0077] Assuming that the number of input channels is equal to the number of output channels, that is, M = N, the number of channels for which PConv needs to perform convolution operations is N p , and the computational cost of PConv can be obtained as which is only 1 / r of that of the standard convolution: 1 / r = (N p / N) 2 .

[0078] When training the real-time detection model for the eggplant seedling center, the eggplant seedling center image dataset comes from the eggplant planting experimental field of XX Agricultural Engineering College, with the variety being South China Long Eggplant, and it was collected in September 2023. Figure 5It is a close-up view of the seedling center of eggplant seedlings at different growth stages. Figure 5 Among them, (a) is the initial stage of seedling growth. At this stage, the seedlings are short and small, and the purple color of the seedling center is relatively light, making it easy to be mixed with weeds. Figure 5 Among them, (b) is the middle growth stage. The leaves are large and have a certain height, and can stand among the weeds. Figure 5 Among them, (c) is the late growth stage. Its center point is larger, the leaves are more and larger, and the plant is also taller.

[0079] To ensure the diversity of the dataset, we collected the eggplant seedling center dataset in the range of 0.5 meters to 1.3 meters in height, and a total of 948 image data were collected. Figure 6 Among them, (a) is an example of data collected on a sunny day. Figure 6 Among them, (b) is an example of data collected on a cloudy day. Figure 6 Among them, (c) is an example of data collected under the setting sun. 760 images were selected as the original training set, 96 as the original validation set, and 92 as the original test set. To improve the generalization ability of the model, we performed data augmentation operations such as brightness adjustment, image flipping, adding noise, Gaussian blur, and hybrid enhancement on some images in the original training set and validation set, and only used data augmentation operations such as rotation and translation on some images in the test set. After augmentation, the training set increased to 2832, the validation set increased to 282, and the test set increased to 276.

[0080] The LabelImg annotation tool was used to manually annotate the center position of each seedling with an appropriate rectangle. During the annotation process, the unified label was the "xin" category.

[0081] A series of platform parameters were used for model training and testing. The detailed information of these parameters is shown in Table 1. The PyTorch deep learning framework was used as the main platform, and Python 3.8 was used as the programming language. In addition, libraries such as CUDA and OpenCV were used to provide image processing and acceleration functions.

[0082] The input image size during training was set to 640×640 pixels, the training period of the model was set at 200 epochs, and the batch size for each epoch was 16. At the same time, the Stochastic Gradient Descent (SGD) optimization algorithm was used to train the model, with the momentum factor set to 0.937. The initial learning rate was set to 0.01, and a weight decay coefficient of 0.0005 was applied to control the complexity of the model and prevent overfitting.

[0083] By setting the above parameters, the model can effectively optimize the weights and parameters during training, thereby improving the performance and accuracy of the model in the object detection task. The selection of these platform parameters has undergone multiple experiments and adjustments to ensure the best performance of the model in the given task and dataset.

[0084] Table 1 Hardware and Software Configuration

[0085]

[0086] This embodiment uses a series of metrics to evaluate the performance of the eggplant seedling heart recognition model, including P (Precision) and AP (Average Precision). The calculation formulas for these metrics are as follows:

[0087]

[0088] In the object detection of eggplant seedling hearts, T P represents the number of predicted bounding boxes that are positive samples and overlap with the ground truth bounding boxes, and F P represents the number of predicted bounding boxes that are positive samples but do not overlap with the ground truth bounding boxes, and F N represents the number of ground truth bounding boxes that are not predicted. The precision metric P is used to evaluate the prediction correctness of the model for eggplant seedling hearts, indicating the proportion of correctly predicted positive samples among the predicted positive samples. The recall rate R represents the proportion of correctly predicted positive samples among the actual positive samples, which is used to measure the detection ability of the model to ensure that the model can predict all the eggplant seedling hearts. AP is the area under the Precision-Recall curve, which is used to evaluate the detection performance of the model for eggplant seedling hearts. In addition to the accuracy of the model, the real-time detection ability of the model is also very important for pesticide spraying. Therefore, this embodiment also uses the inference time (InferenceTime) and the number of parameters (Number of Parameters) of the model to evaluate other performance metrics of the eggplant seedling heart object detection model.

[0089] By using the above metrics, the accuracy, recall rate, and detection ability of the model in the eggplant seedling heart object recognition task can be comprehensively evaluated. At the same time, factors such as the real-time performance and the number of parameters of the model are also considered, and the performance of the model is comprehensively evaluated from multiple perspectives. The selection of these metrics is based on extensive research and practical experience and is widely used in the evaluation of object detection tasks.

[0090] In the experiments evaluating the Backbone and slim-Neck in EggYOLOPlant, we adopted Grad-CAM (Gradient-weighted Class Activation Mapping) heatmap visualization to analyze the performance of the model. Grad-CAM is a gradient-based class activation map generation method that assigns importance values to each neuron according to the gradient information flowing into the last convolutional layer of the CNN for attention decision-making in specific regions. Specifically, by calculating the gradient of the target class and performing element-wise multiplication with the corresponding feature map, the weights of the activation map are obtained. Then, the weights are spatially weighted averaged and non-linearly processed through ReLU (Rectified Linear Unit) to finally obtain the class activation map.

[0091] This Grad-CAM method can help researchers visualize the importance of each neuron in the convolutional neural network, especially for the network prediction results. By observing the heatmap, one can clearly understand the attention of the model in specific regions, and then analyze the performance of the network in different categories and positions. Through Grad-CAM heatmap visualization, we can gain a deeper understanding of how the Backbone and slim-Neck structures of the EggYOLOPlant model work in the task of eggplant seedling heart recognition, providing valuable references for further optimization and improvement. The calculation method of Grad-CAM is as follows:

[0092]

[0093] where ReLU represents the rectified linear unit, which changes the negative values of the activation map output to 0 to suppress the uninteresting weight parts. A k represents the data of the k-th channel in the feature layer A, represents the feature map A of the target class c k 's important weights, which are calculated from the gradient information and are shown as follows:

[0094]

[0095] where y c represents the score corresponding to the class c before the model performs the softmax operation, Denote the data at coordinates (u, v) on the k-th channel of the feature layer A, and Z represents the area of the feature layer (u×v). Finally, the upsampled weighted feature map after ReLU activation is upsampled to the same size as the input image for visualization. By obtaining the heatmap through Grad-CAM and overlaying it on the original image, the attention area of the model on the eggplant seedling heart can be intuitively shown. The darker the color, the higher the attention in that area. Using the Grad-CAM heatmap visualization technique can analyze the feature information output by the EggYOLOPlant model. The experimental results show that the EggYOLOPlant model can well filter the background attention area and fuse the effective context feature information, thus reducing the required computing resources and inference time.

[0096] Similar to YOLOv8, the EggYOLOPlant model also has three levels of feature outputs in the Backbone layer, namely shallow feature output, middle feature output, and deep feature output. The feature attention areas of the two models on the eggplant seedling heart are concentrated in the shallow and middle layers of the model output features. The visualization comparison effect of the Grad-CAM heatmaps of different output scales of YOLOv8 and EggYOLOPlant is as Figure 7 shown, Figure 7 in which Layers[-2], Layers[-3], and Layers[-4] represent the deep, middle, and shallow feature outputs of the model respectively. From Figure 7 it can be seen that compared with the YOLOv8 model, the EggYOLOPlant model can better filter background information and also focus the attention area on the eggplant seedling heart, reducing both computing resources and improving the accuracy of the model in identifying the eggplant seedling heart.

[0097] The test set is the data that has not been used for model training or validation. To objectively evaluate the performance of the eggplant seedling heart real-time detection model, the test set is used as the standard for finally evaluating the model performance to ensure the accuracy and generalization ability of the model. Therefore, we conducted a series of experiments on the test set to prove the superiority of the improved module. When testing, the input image is set to 640×640 pixels, and the IoU is set to 0.7.

[0098] The following shows the impact of the size of r on the model performance through a simple ablation experiment, and selects the optimal size of r through this method.

[0099] When selecting the backbone network, the computing speed and trainability of the model are comprehensively considered. Especially when applied to agriculture, a backbone network with fewer parameters and less computational complexity is more suitable. Therefore, we selected Mblock as the Backbone of the model. In the ablation experiment, Mblock and the C2f-faster module were mainly tested.

[0100] The experimental results in Table 2 show that, compared with the original YOLOv8, using Mblock as the Backbone of YOLOv8, the number of model parameters has decreased to 2.5M, the AP has increased by 3 percentage points, the single-frame inference time has decreased by 6.9%, and the detection accuracy gain for the eggplant seedling heart is the largest. After introducing the C2f-faster module (set p = 8) into the Neck, compared with the original YOLOv8, the number of model parameters has decreased to 1.9M, the AP has increased by 3.8 percentage points, and the single-frame inference time has decreased by 13.8%. The introduction of the C2f-faster module further reduces the inference time of the model. At the same time, because Mblock extracts enough features, the C2f-faster module makes good use of these features and further improves the detection accuracy of the model.

[0101] Ablation experiment results of each module in Table 2

[0102]

[0103] On the basis of the improvement, a brief ablation experiment on the convolution ratio 1 / r of PConv was carried out, which proved that when r = 8, the model would have a higher AP value and faster inference speed. The results are summarized in Table 3. Empirically, it is found that when the value of r is set too small, PConv will degenerate into a conventional Conv. The experimental results show that using the C2f-faster module can improve the computational efficiency of the model, thus accelerating the inference speed of the model, and when r = 8, the model has a higher AP value and faster inference speed.

[0104] Simple ablation experiments for some r values in PConv in Table 3

[0105]

[0106] Referring to Table 4, compared with the original YOLOv8, the number of parameters of EggYOLOPlant is 1.9M, and the weight volume is only 3.91MB, which is about 73% of YOLOv8. There is also a good improvement in the detection accuracy of the model. The precision rate for the eggplant seedling heart reaches 94.7%, an increase of 4 percentage points, the recall rate reaches 91.9, an increase of 3.8 percentage points, the AP value reaches 95.9, an increase of 3.8 percentage points, and the speed of detecting pictures per second reaches 294fps, with a speed increase of about 11.8%.

[0107] Comparison of the overall performance of the model before and after improvement in Table 4

[0108]

[0109] In Figure 8 shows the effect comparison of the ablation detection experiment.Figure 1 and Figure 3 there are 2 and 1 eggplant seedling heart targets respectively, and Figure 3 the eggplant seedlings are relatively small and mixed with the surrounding weeds, Figure 1 in which a part of the heart of an eggplant seedling is blocked by the bamboo pole fixing the eggplant seedling. In Figure 2 there are 6 eggplant seedling heart targets, and at the same time there are many weeds in the background. In Figure 1 all the eggplant seedling hearts are recognized by all three models, but the confidence of YOLOv8 in recognizing the partially blocked eggplant seedling heart is relatively low. In Figure 2 YOLOv8+Mblock misjudges the middle weed as an eggplant seedling heart, while the other two models can complete the detection target well. In Figure 3 when recognizing small eggplant seedlings, YOLOv8 wrongly judges the adjacent weeds as eggplant seedling hearts, while the other two models can complete the detection target well. It can be seen from the results that YOLOv8 is prone to misjudgment in recognizing small targets. After introducing Mblock, the feature extraction ability of the model is enhanced, and its misjudgment in small targets is reduced. However, misjudgment occurs in the case of multiple targets. After introducing Mblock and C2f-faster into YOLOv8, the feature extraction and feature fusion abilities of the model are enhanced, the detection ability of the model for small targets and multiple eggplant seedling hearts on the same screen is improved, and the misjudgment of the model is reduced. Generally speaking, in the case of the existence of small targets and multiple eggplant seedling hearts, EggYOLOPlant using Mblock and C2f-faster will have better robustness and generalization ability.

[0110] We also compared EggYOLOPlant with other mainstream object detection models, and the results are shown in Table 5. EggYOLOPlant has the best comprehensive performance in terms of detection performance. Its number of parameters is 1.9M, which is about 6.7%, 20.9%, and 20.2% of Faster R-CNN, YOLOv5s, and YOLO11s respectively, and the weight is only 3.91MB. In the test of the test set, the AP of the EggYOLOPlant model is 95.5%, the number of frames detected per second is 294fps, and the average inference time per image is 2.5ms. Compared with Faster R-CNN, YOLOv5s, and YOLO11s, the AP of the EggYOLOPlant model has increased by 3.2, 1.8, and 1.5 percentage points respectively, and the detection speed has increased by about 24 times, 1.59 times, and 1.61 times in terms of fps. Among these four detection models, the EggYOLOPlant model has the highest detection speed, and the average precision AP also ranks first. However, the detection precision P and recall rate R are respectively 0.8 percentage points behind YOLOv5s and 3.9 percentage points behind Faster-RCNN, ranking second.

[0111] Table 5 Comprehensive comparison results of EggYOLOPlant and different models

[0112]

[0113] Figure 9 shows the comparison of the detection effects of EggYOLOPlant and different detection models on the eggplant seedling hearts. In Figure 1 , Figure 2 small object detection appears. Because Figure 1 a single eggplant seedling is too small and the characteristics of the seedling heart are not obvious. Only EggYOLOPlant barely recognized it, while the other three models ignored it. In Figure 2 , only YOLO11s failed to detect the seedling heart of the small eggplant seedling next to the large eggplant seedling. In Figure 3 there are four eggplant seedlings and numerous background weeds. Faster-RCNN and YOLO11s both missed the detection of the seedling heart of one eggplant seedling, while YOLOv5s missed the detection of two. In Figure 4 , there are two large eggplant seedlings under the sunset light. Faster-RCNN missed the detection of the seedling heart of one eggplant seedling, and YOLOv5s missed the detection of the seedling heart of one eggplant seedling and misidentified the seedling heart of one seedling as two seedlings of different sizes. In Figure 5Among them, the core of one eggplant seedling was partially blocked, and both Faster-RCNN and YOLO11s missed detecting the core of this seedling. Generally speaking, both Faster-RCNN and YOLOv5s missed detecting the cores of 4 eggplant seedlings, YOLO11s missed detecting the cores of 3 seedlings, while EggYOLOPlant successfully completed the target detection task of the eggplant seedling cores in the above 5 scenarios. Experiments show that compared with the other three object detection models, EggYOLOPlant has fewer cases of missed and misjudged detections when detecting the cores of eggplant seedlings. This model has the best performance and is more adaptable to the challenges of the complex growth environment and diverse appearance features of eggplant seedlings.

[0114] Based on the experimental results, the advantages of the EggYOLOPlant model of this embodiment in terms of detection accuracy and inference speed are verified, enabling it to meet the requirements of precise pesticide spraying operations. This model has a low computational load and model size, making it more suitable for working in scenarios with limited memory and computing power but with requirements for accuracy. Therefore, this model can be applied to mobile devices and embedded devices, providing an effective object detection method for eggplant production in agriculture and valuable research experience for other similar plant object detections.

[0115] In summary, the EggYOLOPlant model introduces the feature extraction module Mblock based on MobileNetV3 on the basis of YOLOv8 and designs a simple and efficient slim-Neck, achieving the miniaturization of the model and improving the inference speed. It can better capture the diverse appearance features of the eggplant seedling cores, optimize feature fusion and computational efficiency, reduce the model size and memory access volume, and meet the requirements of real-time performance and low power consumption on the premise of ensuring the model accuracy.

[0116] Based on the above real-time crop core detection model, an embodiment of the present invention discloses a method for detecting the cores of eggplant seedlings. The real-time crop core detection model is pre-deployed in the embedded device of an agricultural machine. The steps include:

[0117] Collect images of eggplant seedlings;

[0118] Input the images of eggplant seedlings into any of the above-mentioned trained real-time crop core detection models to identify the cores of eggplant seedlings.

[0119] An embodiment of the present invention discloses an electronic device, including a memory storing executable program code and a processor coupled to the memory;

[0120] Among them, the processor calls the executable program code stored in the memory to execute the method for detecting the cores of eggplant seedlings described in the above embodiments.

[0121] An embodiment of the present invention also discloses a computer-readable storage medium storing a computer program, wherein the computer program causes a computer to execute the eggplant seedling heart detection method described in each of the above embodiments.

[0122] The purpose of the above embodiments is to exemplarily reproduce and deduce the technical solution of the present invention, and to completely describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to understand the disclosed content of the present invention more thoroughly and comprehensively, and it does not limit the protection scope of the present invention.

[0123] The above embodiments are not exhaustive listings based on the present invention. In addition, there may be multiple other embodiments not listed. Any substitution and improvement made on the basis of not violating the inventive concept of the present invention fall within the protection scope of the present invention.

Claims

1. A real-time detection model for crop seedling centers based on the YOLOv8 network architecture, characterized in that, Including: A lightweight backbone network, a lightweight neck network, and a head network; The lightweight backbone network includes multiple feature extraction modules constructed based on MobileNetV3 blocks, which are used to extract the features of the input crop seedling image and obtain feature maps of multiple scales; The lightweight neck network includes multiple fusion blocks composed of an upsampling layer, a feature splicing layer, and a C2f-Faster layer, and multiple concatenation blocks composed of a convolutional layer, a feature splicing layer, and a C2f-Faster layer, which are used to bidirectionally fuse the feature maps extracted by the lightweight backbone network and obtain fused features of multiple scales; The head network is composed of a convolutional layer and a two-dimensional convolutional layer, which are used to decode the fused features of multiple scales output by the lightweight neck network to obtain detection results.

2. The real-time detection model for crop seedling centers based on the YOLOv8 network architecture as described in claim 1, characterized in that, The head and tail of the C2f-Faster layer are 1×1 convolutional layers, and the middle part is successively a channel splitting layer, a Faster block layer, and a feature splicing layer. The feature splicing layer is used to splice the splitting information of the channel splitting layer and the Faster block layer.

3. The real-time detection model for crop seedling centers based on the YOLOv8 network architecture according to claim 2, wherein, The Faster block layer has an inverted residual structure and includes a partial convolution module and two consecutive 1×1 convolution modules.

4. The real-time detection model for crop seedling centers based on the YOLOv8 network architecture according to claim 1, characterized in that, When training the real-time crop seedling center detection model, the Grad-CAM heatmap is used to visually analyze the feature information output by the real-time crop seedling center detection model.

5. The real-time detection model for crop seedling centers based on the YOLOv8 network architecture according to claim 1, characterized in that, An SE attention module is also deployed after the depth convolution in the feature extraction module. The SE attention module is used to compress the feature map of W×H×C containing global information into a feature vector of 1×1×C.

6. The real-time detection model for crop seedling centers based on the YOLOv8 network architecture according to claim 5, characterized in that, The output channel number of the first fully connected layer in the SE attention module is 1 / 4 of the channel number of the dilation layer in the SE attention module.

7. The real-time detection model for crop seedling centers based on the YOLOv8 network architecture according to claim 1, characterized in that, The h-swish is used as the activation function in the deep network of the MobileNetV3 block, and the ReLU is used as the activation function in the shallow network of the MobileNetV3 block.

8. A method for detecting the heart of eggplant seedlings, characterized in that, Including: Collecting eggplant seedling images; Inputting the eggplant seedling image into the trained real-time crop seedling center detection model according to any one of claims 1-7 to obtain an eggplant seedling center recognition result.

9. An electronic device, characterized in that, Including a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the eggplant seedling center detection method according to claim 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program enables the computer to execute the eggplant seedling center detection method according to claim 8.

Citation Information

Patent Citations

  • C2f-Faster-based YOLOv8n power transmission line fault detection method

    CN117636125A

  • Fire detection method fusing YOLOv8 and RT-DETR

    CN117974973A

  • Safety helmet and reflective vest detection method based on unmanned aerial vehicle image

    CN118230197A

  • Target detection method for young grape fruit bunches

    CN118710880A

  • Improved YOLOv8n-based grape cluster young fruit lightweight detection method

    CN119339367A