A Fast Detection Method for Pear Fruits in Natural Environments Based on a Lightweight Transformer Architecture

By adopting the detection method of the lightweight Transformer architecture in the rapid detection of pear fruits in natural environments, the problems of difficulty in identification and slow detection in the prior art are solved, and efficient and accurate pear fruit detection is achieved, meeting the requirements of picking robots for real-time detection.

CN119648998BActive Publication Date: 2025-06-27HEBEI AGRICULTURAL UNIV.
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411720677.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-06-27
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The prior art has problems in the rapid detection of pear fruits in natural environments, especially under the influence of light changes, branch and leaf occlusion, fruit overlap and distance environment changes.

Method used

The detection method based on the lightweight Transformer architecture is adopted, including the use of a lightweight FasterNet backbone network, a hybrid encoder and a transformer decoder with auxiliary prediction heads, combined with the HiLo attention mechanism and GSConv convolution, and replace the loss function to Shape_NWD to improve detection performance.

Benefits of technology

It improves detection speed and accuracy, reduces the complexity of the model, and can accurately identify pear fruits in complex environments, meeting the requirements of picking robots for real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648998B_ABST
    Figure CN119648998B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for rapid detection of pear fruits in a natural environment based on a lightweight Transformer architecture, comprising the following steps: A. Collecting pear fruit images in different environments; B. Making a pear fruit data set; C. Constructing a pear fruit detection model; D. Inputting the image to be recognized into the pear fruit detection model for rapid detection of pear fruits. The present invention can improve the deficiencies of the prior art, reduce the complexity of the detection model, and improve the detection speed and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target image detection, and in particular, relates to a fast detection method for pear fruits in a natural environment based on a lightweight Transformer architecture. Background Art

[0002] The mechanization level of the pear industry is low. The picking link requires a large amount of labor, accounting for 35%-45% of the total labor input, and the picking cost even accounts for 50%-70% of all links. The overall process of intelligentization and mechanization in the picking link can generally be divided into two parts. One is to realize the rapid recognition of pear fruits, and the other is to accurately and non-destructively grasp the recognized fruits. However, due to the influence of light changes, branch and leaf occlusion, fruit overlap, and distance environment changes, the model in the actual picking environment is prone to problems such as difficult recognition and slow detection speed. Therefore, how to quickly and accurately identify pear fruits has become the primary problem in automated picking.

[0003] The YOLO model shows certain advantages in both accuracy and speed, but usually requires non-maximum suppression (NMS) for post-processing, and introduces hyperparameters that make the accuracy and speed of the model unstable. In addition, a reasonable NMS threshold needs to be selected, which hinders the development of real-time target detection.

[0004] RT-DETR is slightly inferior to YOLO in small target detection, but is superior to current real-time detectors of similar scale in terms of speed and accuracy. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a fast detection method for pear fruits in a natural environment based on a lightweight Transformer architecture, which can solve the deficiencies of the existing technology, reduce the complexity of the detection model, and improve the detection speed and accuracy.

[0006] To solve the above technical problems, the technical solutions adopted by the present invention are as follows.

[0007] A fast detection method for pear fruits in a natural environment based on a lightweight Transformer architecture includes the following steps:

[0008] A. Collect pear fruit images in different environments;

[0009] B. Make a pear fruit dataset;

[0010] C. Construct a pear fruit detection model;

[0011] D. Input the image to be recognized into the pear fruit detection model for fast detection of pear fruits.

[0012] Preferably, in step A, pear fruit images are collected at the mature stage and the picking stage respectively. The image collection time is from 7:00 to 12:00 and from 13:00 to 19:00 every day, and the image shooting distances are 100 - 500 mm and 700 - 1350 mm.

[0013] Preferably, in step B, first, labelImg is used to annotate several pear fruit images in different scenarios. After training the annotated images with the YOLOv7 model, the optimal weight file is obtained. Then, the remaining pear fruit images are semi-automatically annotated, and finally, labelImg is used to manually adjust the annotation results.

[0014] Preferably, the pear fruit detection model consists of a lightweight FasterNet backbone network, a hybrid encoder, and a transformer decoder with an auxiliary prediction head.

[0015] Preferably, the lightweight FasterNet backbone network includes four hierarchical stages. Before each stage of the FasterNet Block, an embedding layer or a merging layer is connected, which is used for spatial downsampling and channel number expansion respectively. The stride of the embedding layer is 4, and the size of the convolutional kernel is Conv4×4. The stride of the merging layer is 2, and the size of the convolutional kernel is Conv2×2. Each directional residual block consists of a front PConv layer and two rear PWConv layers, which are used to make full use of the information from all channels. A shortcut link composed of a batch normalization and a rectified linear unit activation function is placed between the two PWConv layers to strengthen the input features of the image, thereby improving the performance and training speed of the model.

[0016] Preferably, the hybrid encoder consists of an intra-scale feature interaction module and a cross-scale feature fusion module, which is used to convert multi-scale features into an image feature sequence.

[0017] Preferably, in the intra-scale feature interaction module, the HiLo attention mechanism is used to divide the multi-head self-attention layer into two components. One component is used to encode the high-frequency attention branch Hi-Fi with a local self-attention mechanism and a high-resolution feature map, and the other component is used to encode the low-frequency attention branch Lo-Fi by downsampling the features. The output of each HiLo attention mechanism is the concatenation of the high-frequency attention branch and the low-frequency attention branch, that is, HiLo = [Hi-Fi(X); Lo-Fi(X)]. The scaled dot product attention in the improved spatial reduction strategy is used to adapt the two high- and low-frequency attention branches. One path encodes the high-frequency interaction through the scaled dot product attention and the relatively high-resolution feature map, and the other path encodes the low-frequency interaction through the dot product attention and the downsampled features, which is used to improve the information extraction efficiency.

[0018] Preferably, the cross-scale feature fusion module includes a lightweight convolutional GSConv. The lightweight convolutional GSConv takes the downsampling of a normal convolution as input, then uses a depthwise convolution (DWConv) to concatenate the output results of the two convolutions, and finally performs a data shuffle operation to concatenate the corresponding channel numbers of the previous two convolutions, so as to overcome the problem of partial loss of semantic information caused by spatial compression and channel expansion of the feature map.

[0019] Preferably, a loss function and a small object detection mechanism are jointly used for perceptual query, and a certain number of image features are selected from the encoder output sequence as the initial object query of the decoder. The loss function after combining the loss function and the small object detection mechanism is

[0020] Shape_NWD = (1 - IoU)(1 - 0.5NWD) + IoU(1 - 0.5ShapeIoU), where

[0021] Shape - IoU = 1 - IoU + distance shape + 0.5×Ω shape ,

[0022]

[0023]

[0024]

[0025] IoU is the actual intersection over union, and ww and hh represent the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT box.

[0026] Preferably, in step D, the rapid detection of pear fruits includes the following steps

[0027] Input the fruit image into the lightweight FasterNet backbone network, and input the features of the last three stages of the backbone network into the hybrid encoder. The hybrid encoder converts multi-scale features into an image feature sequence through an intra-scale feature interaction module and a cross-scale feature fusion module. The Shape_NWD small object detection mechanism is used for joint perceptual query to select a certain number of image features from the encoder output sequence as the initial object query of the decoder. The decoder with an auxiliary prediction head iteratively optimizes the object query to generate the bounding box and confidence score.

[0028] The beneficial effects brought by the above technical solution are as follows: In the present invention, on the basis of RT-DETR, the backbone of the original model is replaced with the lightweight FasterNet network. Secondly, the HiLo efficient attention mechanism with high and low frequency information extraction is used to improve the in-scale feature interaction module, and a GSConv convolution and a cross-stage partial network module VoVGSCSP are introduced in the cross-scale feature fusion module. Finally, the loss function in the original model is replaced with Shape_NWD, improving the detection performance of the model. Description of the Drawings

[0029] Figure 1 It is a flowchart of the steps of the present invention.

[0030] Figure 2 It is a block diagram of the network model structure of the present invention.

[0031] Figure 3 It is a schematic diagram of the Shape-IoU parameter of the present invention.

[0032] Figure 4 It is the detection result of the target of Xinli No. 7 fruits using the detection method of the present invention in different environments. Detailed Embodiment

[0033] Referring to Figure 1 , the image data collection location of this embodiment is the modern pear orchard demonstration base of Hebei Wokang Agricultural Technology Co., Ltd. The variety of pears collected is Xinli No. 7, and the collection time is from July to August 2024. Two modes of close-range shooting (100 - 500mm) and long-range shooting (700 - 1350mm) are designed for data collection. To make the time distribution of the collected data set closer to the actual picking time of fruit farmers, the shooting time of this data set is selected in two time periods: 7:00 - 12:00 and 13:00 - 19:00. A total of 7468 Xinli No. 7 images in different environments and scenarios are collected, including 4001 close-range images and 3467 long-range images.

[0034] To improve the annotation efficiency, YOLO semi-automatic annotation is adopted. First, 2000 Xinli No. 7 fruit images in different scenarios are annotated using labelImg. After training the annotated images with the YOLOv7 model, the optimal weight file is obtained. Then, the remaining Xinli No. 7 fruit images are semi-automatically annotated, and finally, labelImg is used to manually adjust mislabeled and missed-labeled situations in the annotation results.

[0035] After manual and semi-automatic annotation, a total of 7468 XML files of Xinli No. 7 are obtained and randomly divided according to the ratio of 8:1:1 to form a training set, a validation set, and a test set. Among them, the training set contains 5974 images, and the validation set and the test set each contain 747 images.

[0036] Refer to Figure 2 , the pear fruit detection model consists of a lightweight FasterNet backbone network, a hybrid encoder, and a transformer decoder with an auxiliary prediction head.

[0037] We improved a fast pear fruit detection model in a natural environment based on RT-DETR. First, we replaced the backbone of the original model with a lightweight FasterNet network. Second, we improved the scale-internal feature interaction module using the efficient attention mechanism HiLo with high and low frequency information extraction, and introduced a GSConv convolution and a cross-stage partial network module VoVGSCSP in the cross-scale feature fusion module. Finally, we replaced the loss function in the original model with Shape-NWD, improving the detection performance of the model.

[0038] The FasterNet backbone network has four hierarchical stages. Before each stage of FasterNet Block, an embedding layer (stride 4, Conv4×4) or a merging layer (stride 2, Conv2×2) is connected, which are used for spatial downsampling and channel number expansion respectively. To make full and efficient use of information from all channels, each directional residual block consists of a front-end PConv layer and two PWConv layers at the back end. A shortcut link composed of a batch normalization (BatchNormalization, BN) and a rectified linear unit (Rectified Linear Unit, ReLU) activation function is placed between the two PWConv layers to strengthen the input features of the new No. 7 pear fruit image, thereby improving the performance and training speed of the model.

[0039] The HiLo attention mechanism divides the MSA layer into two components: one is a high-frequency attention branch Hi-Fi for encoding local self-attention mechanism and high-resolution feature maps, and the other is a low-frequency attention branch Lo-Fi for global encoding of downsampled features. The output of each HiLo attention mechanism is the concatenation of the high-frequency attention branch and the low-frequency attention branch, that is, HiLo is HiLo = [Hi-Fi(X); Lo-Fi(X)].

[0040] And we re-improved the scaled dot-product attention in the spatial reduction strategy to adapt to the two high and low frequency attention branches. One path encodes high-frequency interactions through scaled dot-product attention and relatively high-resolution feature maps, while the other path encodes low-frequency interactions through dot-product attention and downsampled features, thus effectively improving the information extraction efficiency and being more effective than the standard MSA.

[0041] To further enable the pear fruit picking robot to meet the requirements of real-time detection, we introduce lightweight convolutional GSConv in the cross-scale feature fusion module (CCFM), which further reduces the complexity of the model while maintaining the accuracy and solves the problem of the prediction calculation speed in the convolutional neural network. First, GSConv takes the downsampling of a normal convolution as input, then uses depthwise convolution (DWConv) to concatenate the output results of the two convolutions, and finally performs a data assignment shuffle operation to concatenate the corresponding number of channels of the previous two convolutions. Therefore, when the spatial information of the input image is gradually transferred to the channels, GSConv convolution avoids the phenomenon of partial loss of semantic information caused by the compression of the feature map in space (width and height) and the expansion of channels each time.

[0042] In the RT-DETR algorithm, the prediction box regression loss function uses GIoU. When the heights and widths of two prediction boxes are the same and they are on the same horizontal plane, GIoU degenerates into the IOU loss function, resulting in problems such as slow convergence and inaccurate regression. To address the above problems, we use the Shape_NWD loss function to replace the GIoU loss function used in RT-DETR. This method can calculate the loss by focusing on the shape and scale of the bounding box itself, thus making the bounding box regression more accurate. Figure 3 Schematic diagram of Shape-IoU parameters.

[0043] Select the allocation ratio strategy of 50% suitable for the detection target of this model to maximize the detection evaluation index of this model. The Shape_NWD loss function is

[0044] Shape_NWD = (1 - IoU)(1 - 0.5NWD) + IoU(1 - 0.5ShapeIoU),

[0045] where Shape-IoU = 1 - IoU + distance shape + 0.5×Ω shape ,

[0046]

[0047]

[0048]

[0049] IoU is the actual intersection over union, and ww and hh represent the weight coefficients in the horizontal and vertical directions respectively, and their values are related to the shape of the GT box.

[0050] Shape_NWD better handles the influence of shape and scale factors on the regression results in bounding box regression compared with traditional loss functions. Especially in small object detection tasks, this strategy can not only improve the object detection accuracy but also ensure the model detection speed.

[0051] The specific detection process is as follows: Input the fruit image into the lightweight FasterNet backbone network, and input the features of the last three stages of the backbone network into the hybrid encoder; The hybrid encoder converts multi-scale features into an image feature sequence through the in-scale feature interaction module and the cross-scale feature fusion module; The Shape_NWD small object detection mechanism is used to jointly sense the query, and a certain number of image features are selected from the encoder output sequence as the initial object query for the decoder; The decoder with an auxiliary prediction head iteratively optimizes the object query to generate the bounding box and confidence score.

[0052] A comparative experiment was conducted between RT-DETR and YOLOv5m, YOLOv7, YOLOv8m, and YOLOv10m models. The results show that RT-DETR can achieve a good balance in terms of model lightweight and recognition accuracy compared with other models, and can realize the rapid detection of new pear No. 7 fruits.

[0053] The precision, recall, and mean average precision of the model of the present invention reach 93.7%, 91.9%, and 98% respectively, and the number of parameters, computational amount, and model size are reduced by 48.47%, 56.2%, and 48.31% respectively compared with the original model. The research shows that the model of the present invention provides technical support for the detection of new pear No. 7 fruits and the deployment and application of the model in complex environments. The results of the ablation experiment are shown in the following table.

[0054]

[0055]

[0056] Figure 4 The effects of the original model and the improved model on the target detection of new pear No. 7 fruits under different scene conditions are shown. Under the front light and near view, both the RT-DETR model and the improved model can accurately identify the new pear No. 7 fruit target in the image, but various degrees of missed detections occur under conditions such as branch and leaf occlusion, backlight and near view, fruit overlap, front light and far view, and backlight and far view, as Figure 4 shown by the thick solid line box, while the improved model can detect the missed fruits; and the improved model has a high confidence under different environments, such as Figure 4For the fruits of Xinli No. 7 blocked by the fruits shown in column b, the fruit confidence detected by the improved model is 17% higher than that of RT-DETR. The detection results show that in the complex environment of foliage occlusion, fruit overlap, and near and far views with reverse and direct light, the improved model shows better detection effect and higher confidence, meeting the requirements of the Xinli No. 7 fruit picking robot for recognition accuracy.

[0057] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0058] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the description in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for rapid detection of pear fruit in natural environment based on lightweight Transformer architecture, characterized in that The following steps are involved: A. Collect pear fruit images in different environments; B. Create a pear fruit dataset; C. Build a pear fruit detection model, which consists of a lightweight FasterNet backbone network, a hybrid encoder, and a transformer decoder with an auxiliary prediction head; The lightweight FasterNet backbone network includes four hierarchical stages. The FasterNetBlock of each stage is connected to an embedding layer or a merging layer, which are used for spatial downsampling and channel number expansion respectively. The step size of the embedding layer is 4, and the size of the convolution kernel is Conv4×4. The step size of the merging layer is 2, and the size of the convolution kernel is Conv2×2. Each directional residual block consists of a front-end PConv layer and two back-end PWConv layers, which are used to make full use of information from all channels. A shortcut link consisting of batch normalization and rectified linear unit activation function is placed between the two PWConv layers to strengthen the input features of the image, thereby improving the performance and training speed of the model. The hybrid encoder consists of an intra-scale feature interaction module and a cross-scale feature fusion module, which is used to convert multi-scale features into image feature sequences; In the intra-scale feature interaction module, the HiLo attention mechanism is used to divide the multi-head self-attention layer into two components, one component is used to encode the high-frequency attention branch Hi-Fi with the local self-attention mechanism and the high-resolution feature map, and the other component is used to encode the low-frequency attention branch Lo-Fi through the down-sampled features; the output of each HiLo attention mechanism is the concatenation of the high-frequency attention branch and the low-frequency attention branch, that is, HiLo = [Hi-Fi (X); Lo-Fi (X)]; the scaled dot product attention in the improved space reduction strategy is used to adapt the two high- and low-frequency attention branches, one path encodes the high-frequency interaction through the scaled dot product attention and the relatively high-resolution feature map, and the other path encodes the low-frequency interaction through the dot product attention and the down-sampled features to improve the efficiency of information extraction; The cross-scale feature fusion module includes a lightweight convolution GSConv, which inputs a downsampled ordinary convolution, then uses DWConv deep convolution to splice the output results of the two convolutions, and finally performs a data distribution shuffle operation to connect the corresponding channels of the previous two convolutions in series; it is used to overcome the problem of partial loss of semantic information caused by spatial compression and channel expansion of feature maps; D. Input the image to be identified into the pear fruit detection model for rapid pear fruit detection.

2. The method for rapid detection of pear fruit in natural environment based on lightweight Transformer architecture according to claim 1 is characterized in that: In step A, pear fruit images are collected during the ripening period and the picking period, respectively. The image collection time is 7:00-12:00 and 13:00-19:00 every day, and the image shooting distance is 100-500 mm and 700-1350 mm.

3. The method for rapid detection of pear fruit in natural environment based on lightweight Transformer architecture according to claim 1 is characterized in that: In step B, labelImg is first used to annotate several pear fruit images in different scenes. The annotated images are trained using the YOLOv7 model to obtain the optimal weight file, and then the remaining pear fruit images are semi-automatically annotated. Finally, labelImg is used to manually adjust the annotation results.

4. The method for rapid detection of pear fruit in natural environment based on lightweight Transformer architecture according to claim 3 is characterized in that: The loss function and the small target detection mechanism are combined to form a perceptual query, and a certain number of image features are selected from the encoder output sequence as the initial object query of the decoder; the loss function after the loss function and the small target detection mechanism are combined is: Shape_NWD=(1-IoU)(1-0.5NWD)+IoU(1-0.5ShapeIoU), in, Shape-IoU=1-IoU+distance shape +0.5×Ω shape , IoU is the actual intersection-over-union ratio, ww and hh represent the weight coefficients in the horizontal and vertical directions respectively, and their values ​​are related to the shape of the GT box.

5. The method for rapid detection of pear fruit in natural environment based on lightweight Transformer architecture according to claim 4 is characterized in that: In step D, the rapid detection of pear fruit comprises the following steps: The fruit image is input into the lightweight FasterNet backbone network, and the features of the last three stages of the backbone network are input into the hybrid encoder; the hybrid encoder converts multi-scale features into image feature sequences through the intra-scale feature interaction module and the cross-scale feature fusion module; the Shape_NWD small target detection mechanism is used to jointly perceive the query, and a certain number of image features are selected from the encoder output sequence as the initial object query of the decoder; The decoder with an auxiliary prediction head iteratively refines the object query, producing bounding boxes and confidence scores.

Citation Information

Patent Citations

  • Rapid pest detection method based on improved YOLO V4

    CN114220035A

  • Lightweight target detection method and system based on YOLOv8 improvement

    CN116740538A