Real-time Tomato Pose Detection Method Based on DCT-YOLOv5 Model

By introducing the DCT-YOLOv5 model into the YOLOv5 model, combining the convolution layer, Batch Normalization layer, LeakyRelu layer, upsampling layer, transblock layer, CA attention mechanism layer and dynamic convolution layer, the accuracy problem when detecting occluded targets or tiny targets in the prior art is solved, and high-precision and fast real-time detection under variable lighting conditions are achieved.

CN114782360BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210409195.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-06-13
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

The prior art cannot accurately detect the attitude of the target when detecting an obstructed target or a tiny target, and the detection accuracy is affected under variable lighting conditions, making it difficult to meet the needs of fast and real-time applications.

Method used

A real-time tomato pose detection method based on the DCT-YOLOv5 model is proposed. This model combines convolutional layer, Batch Normalization layer, LeakyRelu layer, upsampling layer, transblock layer, CA attention mechanism layer and dynamic convolution layer to achieve accurate detection of tomato pose through these technical means.

Benefits of technology

It realizes rapid real-time detection while ensuring high accuracy, and can effectively handle the detection of occluded targets and tiny targets, adapt to variable lighting conditions, and meet fast and real-time application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782360B_ABST
    Figure CN114782360B_ABST
Patent Text Reader

Abstract

The present invention relates to a real-time tomato pose detection method based on the DCT-YOLOv5 model. The method includes the following steps: Step 1: Design the DCT-YOLOv5 backbone network and loss function; Step 2: Collect image data of tomatoes with different angles, different sizes, and different growth conditions by manual shooting; Step 3: Make a tomato dataset and train it; Step 4: Deploy the DCT-YOLOv5 compressed model to the AGX Xavier embedded system and use TensorRT to accelerate model inference; Step 5: Use a realsense camera to perform real-time tomato detection on the AGX Xavier. The present invention is used to perform real-time tomato detection on the NVIDIA Jetson AGX Xavier embedded development board, ensuring the real-time performance of detection and the high efficiency of model operation while guaranteeing the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to an image processing technology based on deep learning, and particularly relates to a real-time tomato pose detection method based on the DCT-YOLOv5 model. Background Technique

[0002] Real-time object detection technology has been a research hotspot in the field of computer vision in recent years. This technology includes the design of lightweight object detection networks, the production of object datasets, and the research on model deployment carriers. At present, real-time object detection technology based on image sequences can enable a computer to observe and detect objects in image sequences, and this technology is representative in future intelligent driving and robot intelligent sorting. Among them, one of the most potential applications lies in the field of real-time and fast intelligent sorting, such as the intelligent robot picking system in orchards.

[0003] In recent years, the application of agricultural robots in fruit and vegetable picking has been rich and complex. Irie et al. designed a harvesting robot that first uses a 3D sensor to detect whether asparagus can be harvested, and then uses a robotic arm and an end effector to grasp and harvest the asparagus. Bachche et al. proposed a grasping and cutting robot controlled by a servo motor to pick sweet peppers in a horticultural greenhouse. Liu et al. trained deep networks such as YOLOv3, ResNet50, and ResNet152 to detect fruits, thus demonstrating the effectiveness of deep neural networks in fruit recognition.

[0004] In the intelligent robot fruit and vegetable picking system in orchards, the accuracy of detection is the first factor to be considered. In the early object detection tasks based on convolutional neural networks, Ross Girshick et al. proposed an object detection method that pre-extracts a series of candidate regions and extracts features on the candidate regions. This method laid the foundation for the R-CNN series of methods and gave rise to more perfect object detection models such as Fast R-CNN, Faster R-CNN, and Mask R-CNN. The R-CNN series, including the state-of-the-art Faster R-CNN model, has the highest image recognition accuracy in object detection and recognition. However, convolutional network models have a large number of layers and nodes, and the parameters used reach millions or even billions. The computational intensity and storage intensity of the network will bring huge computational and memory consumption, and cannot meet the requirements of fast and real-time applications; it is difficult to be applied to mobile devices with small computational power and small storage space.

[0005] The second key point of the intelligent robot vegetable and fruit picking system is real-time performance. The previous object detection models could not meet the requirements of real-time performance. To address the drawbacks of the previous models, such as excessive model parameters and slow detection speed, Joseph Redmon et al. proposed the YOLO network, from which YOLOv2, YOLOv3, YOLOv5 and other networks were derived. This series of networks directly treats the tomato detection task as a regression problem, combining the two stages of candidate region selection and detection into one. The YOLO series combines recognition and localization into one, with a simple structure and fast detection speed.

[0006] Although the YOLO series of models has greatly improved the detection speed and ensured a certain model accuracy, it cannot accurately detect the pose of occluded or tiny objects. At the same time, the variable lighting in the tomato growth environment also affects the detection accuracy. Adding a dynamic convolution structure, an attention mechanism, and a transblock structure to the original model can effectively solve the above problems. Summary of the Invention

[0007] The present invention overcomes the drawbacks of the prior art and proposes a DCT-YOLOv5 tomato pose detection model that is easy to implement and highly applicable. This network can achieve fast real-time detection while ensuring high accuracy.

[0008] The present invention takes an image sequence as input. First, the DCT-YOLOv5 model is used to perform object detection and recognition on each frame of the image. The basic units of this model are composed of a convolutional layer, a Batch Normalization layer (BN layer), a LeakyRelu layer, an upsampling layer, a transblock layer, a CA attention mechanism layer, and a dynamic convolution layer. The network model structure diagram is shown in the appendix Figure 1 . The network structure of DCT-YOLOv5 can be divided into four parts: the input end, the Backbone, the Neck, and the head. Among them, the input end includes techniques such as Mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling; the Backbone includes structures such as the Focus structure and CSP; the Neck includes CA attention, FPN, and PAN structures; the head includes techniques such as the transblock structure and GIOU_Loss. The DCT-YOLOv5 model is deployed on the Jetson AGX Xavier development board, and TensorRT is used to accelerate the inference. The Jetson AGX Xavier development board uses a realsense depth camera to collect tomato RGB image data. The data is input into the DCT-YOLOv5 object detection model in the form of an image sequence. The model performs object detection and recognition on each frame of the image, and outputs the detection and recognition results of the tomatoes in the image, including the center position of the tomatoes and the depth of the tomato center from the camera, which is convenient for the manipulator to perform subsequent grasping.

[0009] The technical solution adopted by the present invention is as follows: A real-time tomato pose detection method based on the DCT-YOLOv5 model, which is characterized by the following steps:

[0010] Step 1: Design the DCT-YOLOv5 backbone network and loss function;

[0011] Step 2: Collect image data of tomatoes in various growth forms by manual shooting;

[0012] Step 3: Make a tomato dataset and perform training;

[0013] Step 4: Deploy the DCT-YOLOv5 model to an embedded system and use TensorRT to accelerate model inference;

[0014] Step 5: Use a realsense camera to perform real-time tomato position detection and depth detection on Jetson AGX Xavier.

[0015] The specific steps of the said Step 1 are as follows:

[0016] 1.1): Design of the DCT-YOLOv5 backbone network;

[0017] 1.1.1): Draw on the shortcut design in the ResNet network to deepen the depth of the DCT-YOLOv5 main network, and achieve downsampling of the convolutional layer by setting the stride parameter in the convolutional layer. Except for the last three convolutional layers used for prediction, Batch Normalization (BN) operations are added after the rest of the convolutional layers, and the LeakyRelu activation function is connected after the BN layer. Use dynamic convolution to eliminate the influence of variable illumination. The CSP module is adopted in the network to first divide the feature map of the basic layer into two parts, and then merge them through a cross-stage hierarchical structure, ensuring the accuracy while reducing the computational amount. Draw on the model structures of the FPN and PAN networks, and perform concat fusion on the three feature maps output by the network through upsampling operations to achieve the purpose of multi-scale prediction. Use the CA attention mechanism to fuse vertical and horizontal attention to distinguish the interfering branches in the image. Add the tranblock module to capture the global attention of the image and accurately identify the growth pose of tomatoes;

[0018] 1.1.2): Use the K-meas clustering method and genetic algorithm to cluster the real boxes in the dataset to obtain nine anchor boxes, and every three anchor boxes correspond to a feature map of a scale. The purpose of this method is to accelerate the regression of the prediction box;

[0019] 1.1.3): The prediction formula in the forward inference of the network is as follows:

[0020] b x = σ(t x ) + c x (1)

[0021] b y = σ(t y ) + c y (2)

[0022]

[0023]

[0024] b x , b y are the relative center coordinate values of the prediction box on the feature map of the corresponding size. b w , b h are the width and height of the prediction box. c x , c y is the upper left corner coordinate of the grid cell of the output feature map, p w , p h are the width and height of the anchor box. t x , t y are the predicted coordinate offset values, t w , t h are the predicted scale scaling factors;

[0025] 1.1.4): The implementation formula of dynamic convolution is as follows:

[0026]

[0027]

[0028]

[0029] β k (x) are the weights of the k convolutional kernels calculated by the network, and the weight values are between 0 and 1 and the sum is 1. represents each convolutional kernel, represents the bias of each convolution. represents the final convolutional kernel, represents the final bias. g represents the BN layer and activation function operations, and y represents the feature map output after dynamic convolution;

[0030] 1.1.5): The implementation formula of the CA attention mechanism is as follows:

[0031]

[0032]

[0033]

[0034] x c (i, j) is the feature value at the (i, j) position in the feature map, H and W are the length and width of the feature map, and z c is the information embedding at each position in the calculated feature map. This step enables the module to capture features with precise position information in two directions. T 1 , T 2 are two linear connection layers, which can learn important channels in the feature map. RELU is the activation function, and σ is the sigmoid activation function. X is the original feature map, is the processed feature map. The weighted feature map is more sensitive to horizontal and vertical information. It is beneficial for the model to recognize the branches and the growth postures of tomatoes.

[0035] 1.1.6): The implementation formula of the transblock structure is as follows:

[0036] Q = W Q (W(x)), K = W K (W(x)), V = W V (W(x)) (11)

[0037] y = W(x) + MLP(Dropout(MultiHead(Q, K, V)) + W(x)) (12)

[0038] W(x) is the input feature map passing through a convolutional layer and then passing through W Q , W K , W V three different fully connected layers to obtain the query vector Q, the key vector K, and the value vector V. y is the output of a Transformer Encoder structure. Any number of Transformer Encoders can be stacked in the transblock. Concatenating the output of the final Transformer Encoder structure with the input feature map can obtain the final output feature map.

[0039] 1.2): Design the DCT - YOLOv5 loss function;

[0040] 1.2.1): Design the object confidence loss function;

[0041] 1.2.2): Design the object class loss function;

[0042] 1.2.3): Design the object localization loss function;

[0043] 1.2.4): Obtain the final loss function through the weight coefficient;

[0044] The specific steps of step 3 are as follows:

[0045] 3.1): Preprocess the collected tomato image samples and establish a tomato detection target database;

[0046] 3.2): Manually annotate the detection objects in the image using the labelImg software to generate an xml file. The xml file contains the corresponding coordinate value information of the real bounding boxes of tomatoes manually annotated by labelImg, as well as the label information corresponding to each box;

[0047] 3.3): Input the annotated image data into the model for training;

[0048] The specific steps of step 4 are as follows:

[0049] 4.1): Vertically integrate the backbone network structure of DCT - YOLOv5, and fuse the convolutional layer, BN layer, and Relu layer into one layer.

[0050] 4.2): Horizontally integrate the backbone network structure of DCT - YOLOv5, and fuse the tensors with the same input dimension and the layers performing the same operations together.

[0051] 4.3): Directly send the input of the concat layer in the backbone to the subsequent operations to reduce the transmission throughput.

[0052] 4.4): Quantize the model parameters of DCT - YOLOv5, changing the format from float32 to float16 to accelerate the inference speed of the model.

[0053] In summary, the advantages of the present invention are that the original DC - TYOLOv5 model already has a high - precision detection effect. On this basis, TensorRT is used for inference acceleration, enabling it to be successfully deployed on an embedded development board with low configuration; and during the TensorRT acceleration process, the fused layers have the same performance as the layers before fusion, and will not have too much impact on the model performance; this model is accelerated through TensorRT on Jetson AGX Xavier to obtain the final detection model. This model realizes the function of real - time detection on a low - configuration embedded development board. Description of the Drawings

[0054] Figure 1 is the structural diagram of the DCT - YOLOv5 model in the present invention;

[0055] Figure 2 It is a flowchart for accelerating model inference using TensorRT in the present invention;

[0056] Figure 3 It is a flowchart for real-time relasense detection in the present invention. Specific embodiments

[0057] The present invention will be further described below with reference to the accompanying drawings.

[0058] The specific process of the real-time tomato pose detection method based on the DCT-YOLOv5 model of the present invention is as follows:

[0059] 1.1): Design of the DCT-YOLOv5 backbone network, as Figure 1 shown;

[0060] 1.1.1): Theoretically, the deeper the network, the better the detection effect and the higher the accuracy. However, experimental results show that excessive increase in the number of network layers will cause the network to fall into overfitting, slow down network convergence, reduce detection accuracy, and make it more difficult to deploy on embedded devices due to the increased computational cost of the model. To solve this problem, the DCT-YOLOv5 backbone network draws on the skip connection structure of the deep residual network. To reduce the impact on gradient calculation caused by the pooling layer, downsampling operations in the network are all implemented through convolutional layers, and the stride of the convolutional layers is set to 2. A large number of experiments show that there will be a problem of inconsistent data distribution between each layer in the neural network, which will make it difficult for the network to converge and train. To solve this problem, the DCT-YOLOv5 network performs Batch Normalization operations on the outputs of the remaining convolutional layers except for the last three convolutional layers used for prediction, as a method to solve gradient disappearance and gradient explosion, accelerate network convergence, and avoid overfitting. After each BN layer, the network introduces the LeakyRelu function as the activation function, and the role of this layer is to introduce non-linearity into the network. The convolutional layer, BN layer, and LeakyRelu layer together constitute the smallest component of the network. To accurately detect tomatoes of different sizes, DCT-YOLOv5 draws on the Feature Pyramid Network FPN (feature pyramid network) and PAN, fuses features through upsampling operations, and makes predictions at three scales for the detection targets based on the images collected by the realsense camera on the Jetson AGX Xavier. In the present invention, the output feature map with a size of 20*15 has the largest receptive field and is specifically used to detect large targets, and the output feature map with a size of 80*60 has the smallest receptive field and is specifically used to detect small targets;

[0061] 1.1.2): Use the K-Means++ algorithm and the genetic algorithm to perform anchor box clustering on the tomato dataset according to the three detection scales described in step 1-1, generating 9 types of anchor boxes of different sizes, with 3 types of anchor boxes assigned to each detection scale. The role of the anchor boxes is to more quickly and accurately regress the detection boxes;

[0062] 1.1.3): In the forward inference of the network, through the formula:

[0063] b x = σ(t x ) + c x (1)

[0064] b y = σ(t y ) + c y (2)

[0065]

[0066]

[0067] Perform the prediction of the target detection box, and finally obtain the relative center coordinate values b x , b y , as well as the width and height b w , b h , c x , c y is the upper left coordinate of the grid cell of the output feature map, p w , p h is the width and height of the anchor box. t x , t y is the coordinate offset value predicted by the network, t w , t h is the scale scaling multiple predicted by the network.

[0068] 1.1.4): In the dynamic convolution, through the formula:

[0069]

[0070]

[0071]

[0072] Perform the adaptive transformation of the convolution kernel. β k (x) is the weight of the k convolution kernels calculated by the network, and the weight size is between 0 and 1, and the sum is 1. represents each convolution kernel, represents the bias of each convolution. represents the final convolution kernel, Represents the final bias. g represents the BN layer and activation function operations, and y represents the feature map output after dynamic convolution;

[0073] 1.1.5): In the CA attention mechanism, through the formula:

[0074]

[0075]

[0076]

[0077] Realize the extraction of horizontal and vertical attentions of the network. x c (i, j) is the feature value at the (i, j) position in the feature map, H and W are the length and width of the feature map, and z c is the information embedding at each position in the calculated feature map. This step enables the module to capture features with precise position information in both directions. T 1 , T 2 are two linear connection layers, which can learn important channels in the feature map, RELU is the activation function, and σ is the sigmoid activation function. X is the original feature map, is the processed feature map. The weighted feature map is more sensitive to horizontal and vertical information. It is beneficial for the model to identify the branches and the growth postures of tomatoes.

[0078] 1.1.6): In the transblock structure, through the formula:

[0079] Q = W Q (W(x)), K = W K (W(x)), V = W V (W(x)) (11)

[0080] y = W(x) + MLP(Dropout(MultiHead(Q, K, V)) + W(x)) (12)

[0081] Realize the capture of global features of the image and improve the accuracy of tomato posture recognition. W(x) is the input feature map passing through a convolutional layer, and then passing through W Q , W K , W V Three different fully connected layers to obtain the query vector Q, the key vector K, and the value vector V. y is the output of a Transformer Encoder structure. Any number of TransformerEncoders can be stacked in the transblock. Concatenate the output of the final Transformer Encoder structure with the input feature map to obtain the final output feature map.

[0082] 1.2) Design of the DCT - YOLOv5 loss function;

[0083] 1.2.1) For the object confidence, which is the probability of an object existing in the object detection box, the binary cross - entropy loss function is adopted. The designed object confidence loss function is as follows:

[0084]

[0085] Where The network output c i Is obtained through the Sigmoid function

[0086] 1.2.2) For the object class loss function, the binary cross - entropy is also used. The designed object class loss function is as follows:

[0087]

[0088] Where, The network output c i Is obtained through the Sigmoid function Represents the Sigmoid probability that the j - th class object exists in the object detection box i:

[0089] 1.2.3) The object localization loss function adopts the MSE loss function, as follows:

[0090]

[0091] Where:

[0092]

[0093]

[0094]

[0095] Where Represents the coordinate offset of the predicted box (DCT - YOLOv5 predicts the coordinate offset value), Represents the coordinate offset of the ground - truth box, (b x ,b y ,b w ,b h ) are the parameters of the predicted box, (c x ,c y ,p w ,p h ) are the parameters of the anchor box, (g x ,g y ,gw , g h ) are the parameters of the ground truth box;

[0096] 1.2.4): Add up all the above loss functions with weights to obtain the total loss function:

[0097] L(O, o, C, c, l, g) = λ conf L conf (o, c) + λ cla L cla (O, C) + λ loc L loc (l, g) (16)

[0098] 2): Collect image data of tomatoes by taking manual photos. When collecting, tomatoes with different lighting, different sizes, different angles, and different distances need to be photographed.

[0099] 2.1): Perform data augmentation on the collected target images. Expand the dataset through image flipping, stretching, rotation, and cropping to establish a tomato detection dataset.

[0100] 2.2): Use the labelImg software to annotate the tomatoes in the images to generate xml files. The xml files contain the coordinate information of the ground truth boxes manually annotated using labelImg, as well as the labels corresponding to each box.

[0101] 3): Input the annotated dataset into the model for normal training. Set the initial learning rate to 0.01 and the Batch_size value to 16.

[0102] 4): Deploy the model to the Jetson AGX Xavier embedded development board and perform forward inference acceleration through TensorRT, as Figure 2 shown;

[0103] 4.1): Vertically integrate the backbone network structure of the DCT - YOLOv5 model, and fuse the convolutional layer, BN layer, and Relu layer into one layer.

[0104] 4.2): Horizontally integrate the backbone network structure of the DCT - YOLOv5, and fuse the tensors with the same input dimensions and the layers performing the same operations together.

[0105] 4.3): Directly send the input of the concat layer in the backbone to the subsequent operations to reduce the transmission throughput.

[0106] 4.4): Quantize the model parameters of the DCT - YOLOv5, changing the format from float32 to float16 to accelerate the inference speed of the model.

[0107] 5): The AGX Xavier is externally connected to the realsense camera module. The realsense camera is used to collect RGB images. OpenCV is used to process the video stream, and the accelerated model is used to detect the real-time position and depth of tomatoes, as Figure 3 shown.

[0108] The content described in the embodiments of this specification is only an example of the implementation form of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art according to the inventive concept.

Claims

1. A real-time tomato pose detection method based on the DCT-YOLOv5 model, characterized in that it includes the following steps: Step 1: Design the DCT-YOLOv5 backbone network and loss function; Step 2: Collect image data of tomatoes with different angles, sizes, and growth conditions through manual shooting; Step 3: Make a tomato dataset and train it; Step 4: Deploy the DCT-YOLOv5 compressed model to the AGX Xavier embedded system and use TensorRT to accelerate model inference; Step 5: Use a realsense camera to perform real-time tomato detection on the AGX Xavier embedded development board; The specific steps of the said Step 1 are as follows: 1.1): Design of the DCT-YOLOv5 backbone network; 1.1.1): Draw on the shortcut design in the ResNet network to deepen the depth of the DCT-YOLOv5 main network, and realize downsampling of the convolutional layer by setting the stride parameter in the convolutional layer; except for the last three convolutional layers used for prediction, Batch Normalization (BN) operations are added after the rest of the convolutional layers, and the LeakyRelu activation function is connected after the BN layer; dynamic convolution is used to eliminate the influence of variable illumination; in the network, the CSP module first divides the feature map of the basic layer into two parts, and then merges them through a cross-stage hierarchical structure, ensuring accuracy while reducing the amount of calculation; draw on the model structures of the FPN and PAN networks, and perform concat fusion on the three feature maps output by the network through upsampling operations to achieve the purpose of multi-scale prediction; use the CA attention mechanism to fuse vertical and horizontal attention to distinguish the interfering branches in the image; add a tranblock module to capture the global attention of the image and accurately identify the growth pose of tomatoes; 1.1.2): Use the K-meas clustering method and genetic algorithm to cluster the real boxes to obtain nine anchor boxes, and every three anchor boxes correspond to a feature map of a scale; the purpose of this method is to accelerate the regression of the prediction box; 1.1.3): In the forward inference of the network, the prediction formula is as follows: b x =σ(t x )+c x (1) b y =σ(t y )+c y (2) b x ,b y is the relative center coordinate value of the prediction box on the feature map of the corresponding size; b w ,b h are the width and height of the prediction box; c x ,c y is the upper left coordinate of the grid cell of the output feature map, p w ,p h are the width and height of the anchor box; t x ,t y is the predicted coordinate offset value, t w ,t h is the predicted scale scaling factor; 1.1.4): The implementation formula of the dynamic convolution is as follows: β k (x) is the weight of the k convolution kernels calculated by the network, and the weight size is between 0 and 1, and the sum is 1; represents each convolution kernel, represents the bias of each convolution; represents the final convolution kernel, represents the final bias; g represents the BN layer and the activation function operation, and y represents the feature map output after the dynamic convolution; 1.1.5): The implementation formula of the CA attention mechanism is as follows: x c (i,j) is the feature value at the (i,j) position in the feature map, and H and W are the length and width of the feature map, z c is the information embedding of each position calculated in the feature map; this step enables the module to capture features with precise position information in two directions; T 1 ,T 2 are two linear connection layers, which can learn the important channels in the feature map, RELU is the activation function, and σ is the sigmoid activation function; X is the original feature map, is the processed feature map; the weighted feature map is more sensitive to horizontal and vertical information; it is beneficial to the model's recognition of the branches and the growth postures of tomatoes; 1.1.6): The implementation formula of the transblock structure is as follows: Q = W Q (W(x)), K = W K (W(x)), V = W V (W(x)) (11) y = W(x) + MLP(Dropout(MultiHead(Q, K, V)) + W(x)) (12) W(x) is the input feature map passing through a convolutional layer and then through W Q , W K , W V three different fully connected layers to obtain the query vector Q, the key vector K, and the value vector V; y is the output of a Transformer Encoder structure, and any number of Transformer Encoders can be stacked in the transblock; concatenating the output of the final Transformer Encoder structure with the input feature map can obtain the final output feature map; 1.2): Design the DCT - YOLOv5 loss function; 1.2.1): Design the object confidence loss function as follows: where the network output c i is obtained through the Sigmoid function 1.2.2): The design target category loss function is as follows: Among them, Network output c i Obtained through the Sigmoid function Represents the Sigmoid probability that the j-th class of target exists in the target detection box i; 1.2.3): The design target localization loss function is as follows: Wherein: Among them represents the coordinate offset of the predicted bounding box, represents the coordinate offset of the ground truth bounding box, (b x , b y , b w , b h ) are the parameters of the predicted bounding box, (c x , c y , p w , p h ) are the parameters of the anchor box, (g x , g y , g w , g h ) are the parameters of the ground truth bounding box; 1.2.4): Obtain the final loss function through the weight coefficient: L(O,o,C,c,l,g) = λ conf L conf (o,c) + λ cla L cla (O,C) + λ loc L loc (l,g) (16).

2. The real-time tomato pose detection method based on the DCT-YOLOv5 model according to claim 1, characterized in that: The specific steps of step 3 are as follows: 3.1): Preprocess the collected tomato image samples and establish a tomato detection target database; 3.2): Manually annotate the detection objects in the image with the labelImg software to generate an xml file, which contains the corresponding coordinate value information of the real bounding boxes of tomatoes manually annotated by labelImg and the label information corresponding to each box; 3.3): Input the annotated image data into the model for training.

3. The real-time tomato pose detection method based on the DCT-YOLOv5 model according to claim 1, characterized in that: The specific steps of step 4 are as follows: 4.1): Vertically integrate the backbone network structure of DCT-YOLOv5, and fuse the convolutional layer, BN layer, and Relu layer into one layer; 4.2): Horizontally integrate the backbone network structure of DCT-YOLOv5, and fuse the tensors with the same input dimension and the layers performing the same operations together; 4.3): Directly send the input of the concat layer in the backbone to the subsequent operations to reduce the transmission throughput; 4.4): Quantize the model parameters of DCT-YOLOv5, change the format from float32 to float16 to speed up the inference speed of the model.

Citation Information

Patent Citations

  • Real-time medicine box detection method based on YOLOv3 pruning network and embedded development board

    CN112597919A

  • Night tomato recognition system and method based on improved yolk

    CN113326808A