Lightweight Yolov8-based 6-Dof attitude estimation method
Through the lightweight Yolov8 algorithm based on deep learning, combined with key point detection and feature fusion technology, the challenge of quickly and accurately estimating the 6-Dof pose from a single RGB image is solved, and a fast, lightweight and high-precision pose estimation is achieved.
Patent Information
- Application Number
- CN202510053694.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has challenges in estimating the 6-Dof pose of an object from a single RGB image, especially in the absence of object textures and complex backgrounds, where traditional methods match slowly and poorly robustly.
Using the lightweight Yolov8 algorithm based on deep learning, the network structure is improved through key point detection and feature fusion to achieve fast, lightweight and accurate 6-Dof pose estimation.
It realizes that while reducing the amount of parameters and calculations, it maintains high key point extraction capabilities and 6-Dof pose estimation accuracy, and is suitable for platforms such as cloud services, edge devices and mobile devices.
Smart Images

Figure SMS_1 
Figure SMS_4 
Figure SMS_5
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of 6-Dof posture estimation in computer vision and relates to a 6-Dof posture estimation method based on a deep learning method. Background Art
[0002] Target pose estimation has important applications in many real-world applications such as robotics, virtual reality, augmented reality, etc. It can be implemented using RGB-D, point cloud, RGB image, and some combined data. Obviously, in practical applications, RGB images are the easiest to obtain and relatively more accurate, but due to the lack of depth information, the lack of object texture, and cluttered scenes, it is a challenging task to quickly and accurately obtain the 6-Dof pose information of an object from a single RGB image. Traditional methods are based on feature point matching algorithms between two-dimensional images and three-dimensional object models, such as SURF, SIFT, etc. Such methods have slow matching speeds and require rich texture features. They are not robust to different image qualities and background complexity.
[0003] With the development of deep learning, the method of estimating the 6-Dof pose information (3D rotation and 3D translation) of an object through a single RGB image has developed rapidly. There are currently four mainstream methods. One is the direct method, which directly predicts the 6-Dof pose of an object through an image, such as PoseCNN and SSD6d. The speed of this method is considerable, but the accuracy is usually not ideal. One is to predict the key points of the object in the image and then use various PnP algorithms to solve the pose information, such as Yolo-6d and BB8. This method is simple to implement, easy to transplant, relatively fast in detection speed, and relatively few parameters. One is a voting-based method, such as PVNet, which is essentially an indirect detection of the key points of the object. The key points are obtained by voting for each pixel, and then the pose is solved by the PnP algorithm. This method performs well on severely occluded images, but the computation is large. One is a dense method, which establishes a dense 2D-3D correspondence by predicting the 3D coordinates of the object coordinate system, such as CDPN and GDR-Net. This method usually has a large computational workload. The main purpose of this invention is to design a deep learning-based target 6-Dof pose estimation algorithm based on key point detection using a single RGB image as input, and the algorithm must meet the requirements of being fast, lightweight and accurate. Summary of the invention
[0004] To solve the above problems, the object of the present invention is to provide a lightweight object 6-Dof posture estimation algorithm based on deep learning method.
[0005] The object of the present invention is to provide a method for estimating the 6-Dof posture of an object based on a lightweight Yolov8 algorithm, characterized in that it includes:
[0006] Step 1: Define the key points of the identified object;
[0007] Step 2: Implement a 6-Dof pose estimation network using Yolov8;
[0008] Step 3: The 6-Dof pose estimation network using Yolov8 is lightweight and improved to obtain a network structure named Yolov8-lite-6D-v1;
[0009] Step 4: Improve the feature fusion of the lightweight network Yolov8-lite-6D-v1 to increase the accuracy and obtain the final lightweight network Yolov8-lite-6D-v2;
[0010] Step 5: Quantitatively evaluate the algorithm results.
[0011] In the object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm of the present invention, the step 1 is specifically:
[0012] Step 1.1: Use the FPS sampling algorithm and the model corner points to jointly select the three-dimensional coordinates of key points that are evenly distributed on the surface of the target model. The principle of the FPS algorithm is as follows:
[0013] First, the input point cloud (model) has N points. A point P0 is selected from the point cloud as the starting point to obtain a sampling point set S = {P0}. The initial point selection needs to select the farthest point from the center of gravity of the point cloud; calculate the distance from all points to P0 to form an N-dimensional array L, select the point corresponding to the maximum value as P1, and update the sampling point set to S = {P0, P1}; calculate the distance from all points to P1, for each point Pi, if its distance from P1 is less than L[i], then update L[i] = d(Pi, P1), and the array L always stores the closest distance from each point to the sampling point set S; select the point corresponding to the maximum value in L as P2, and update the sampling point set S = {P0, P1, P2}; repeat steps 2 to 4 until the number of sampling points reaches the upper limit.
[0014] Step 1.2: The three-dimensional corner points of the model are selected as: [Xmin, Ymin, Zmin], [Xmin, Ymin, Zmax], [Xmin, Ymax, Zmin], [Xmin, Ymax, Zmax], [Xmax, Ymin, Zmin], [Xmax, Ymin, Zmax], [Xmax, Ymax, Zmin], [Xmax, Ymax, Zmax]. Among them, Xmin, Ymin, and Zmin are the minimum values of X, Y, and Z in the model coordinate system respectively, and Xmax, Ymax, and Zmax are the maximum values of X, Y, and Z in the model coordinate system respectively.
[0015] Step 1.3: Combine all the selected corner points with the position information of the target object in each image to project the target model and map it to the pixel coordinate system. The mapping method is:
[0016]
[0017] Among them, u and v are the pixel coordinates obtained by transformation, X W , Y W , Z W It is the three-dimensional coordinate in the world coordinate system. is the camera internal parameter, is the camera external parameter, that is, the target object pose information.
[0018] In the object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm of the present invention, the step 2 is specifically:
[0019] Step 2.1: Use Yolov8's 6-Dof pose estimation network Yolov8-6D. The backbone network of Yolov8 is Darknet-53, in which the C2f module replaces the C3 module in the previous Yolo series. The detection head of the head network is modified so that it can detect the two-dimensional coordinate values of the center point and eight corner points of the target object and n FPS sampling points. For the coordinates of the center point and the corner point, since Yolov8 is an anchor-free structure, it is not necessary to obtain the coordinate values of the corner point by predicting the offset. For training of objects of different sizes, use a confidence function that adapts to training of different sizes:
[0020]
[0021] Where α is the control parameter, D T (x) is the Euclidean distance between the predicted coordinates and the true coordinates, d th is the confidence threshold.
[0022] Step 2.2: The loss function consists of multiple parts:
[0023] L=λ ciou L ciou +λ cls L cls +λ dfl L dfl +λ kobj L kobj +λ pose L pose (3)
[0024] Where L ciou It is the complete intersection-over-union loss proposed in "Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation", L cls is the category cross entropy loss, L dfl It is the probability density loss proposed in "Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection", L kobj is the key point loss, which measures whether the key point exists, L pose is the key point position loss, is the sum of the Euclidean distance differences between the predicted key points and the true values of the key points. λ is the weight of each loss. pose Set to the maximum, in practice it should be greater than or equal to 10.
[0025] Step 2.3: To be more suitable for the 6-Dof pose estimation task, the evaluation function used in the network training process is set to:
[0026]
[0027] Among them, m is the number of targets, R and T are the actual rotation and translation, The predicted and solved rotation and translation of the current key point, x is the three-dimensional coordinate of the point, and M represents the target object.
[0028] Step 2.4: For all the predicted key points, use the EPnP algorithm in the opencv library to calculate the pose information of the target object.
[0029] In the object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm of the present invention, the step 3 is specifically:
[0030] Step 3.1: Replace all convolutional modules and C2f modules in the backbone detection part with the shuffle Block lightweight module. This lightweight module mainly feeds the input into the depth-separable convolutional layer DWConvBlock and the 1x1 convolutional layer, and rearranges the channel information of the grouped output, thereby reducing the amount of calculation and the number of parameters of the model.
[0031] Step 3.2: Replace all C2f modules in the Neck part with depthwise separable convolutional layers. The replacement mainly requires determining the number of input and output channels and size of each C2f module, and then constructing a depthwise separable convolutional layer module. The depthwise separable convolutional module is mainly divided into two parts, one is the depthwise convolution part, and the other is the pointwise convolution part. Construct a 1x1 convolutional layer to adjust the number of channels. Finally, combine the depthwise convolution and pointwise convolution and call them in sequence.
[0032] Step 3.3: A StemBlock layer is connected after the front-end input layer of the network. Its main function is to improve the efficiency of initial feature extraction, reduce data dimension and calculation amount, and adapt to the structural requirements of different networks. The main structure is the combination of convolutional layers and pooling operations.
[0033] Step 3.4: Get the lightweight network Yolov8-1ite-6D-v1.
[0034] In the object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm of the present invention, the step 4 is specifically:
[0035] Step 4.1: Directly mix the three low-level feature maps with the three high-level feature maps respectively, and then calculate the weights according to CGA. Use the complementary weights as the weights of the two feature maps, and remix the original low-level feature maps and the high-level feature maps after weighting to obtain a fused feature map with content-guided attention mechanism.
[0036] Step 4.2: Get the final network Yolov8-lite-6D-v2 after feature fusion.
[0037] In the object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm of the present invention, the step 5 is specifically:
[0038] Step 5.1: Use the mAP indicator to evaluate the key point prediction ability of the model. The calculation of the mAP indicator is usually based on OKS, and the calculation method is:
[0039]
[0040] Among them, i is the serial number of the key point, d is the Euclidean distance between the predicted value of the key point and the true value of the key point, S is the scale factor, w and h are the width and height of the detection box. σ is the normalization factor of the key point, which is calculated by the standard deviation between the manual annotation and the true value of the real key points in all samples. The larger σ is, the more difficult it is to annotate the key point. δ(v i >0) is the indicator function, if v i >0 is 1, otherwise it is 0, which usually indicates the visibility of the key point. When OKS is greater than the threshold, the prediction is considered correct. Sort the predicted key points according to the confidence of the prediction results. Starting from high confidence, calculate the accuracy and recall of each prediction result in turn to obtain a series of accuracy-recall data points. AP is calculated by integrating the data points. Average the AP of all different key points to get mAP. The mAP when the threshold is 0.5 is mAP50. The average mAP calculated from the threshold starting from 0.5 and increasing in steps of 0.05 to 0.95 is mAP50-95.
[0041] Step 5.2: Use the reprojection error to evaluate the performance of 6-Dof pose estimation. Use the predicted 6-Dof pose information to reproject each point in the known object model into the new image. The specific projection method is consistent with the projection calculation method in step 1.3. Calculate the overlapping pixels of the target object in the original image and the newly obtained projected image. If the number of non-overlapping pixels is less than 5, the pose estimation result is considered correct.
[0042] Step 5.3: Use the ADD indicator to evaluate the performance of 6-Dof pose estimation. The three-dimensional point cloud in the object model is transformed into the corresponding rotation and translation using the real 6-Dof pose information and the predicted 6-Dof pose information, and the average distance between each identical point pair in the two is calculated. When the average distance does not exceed 10% of the maximum diameter of the model, the pose estimation result is considered correct. This indicator is called ADD-0.1d. For symmetrical objects, the ADD-S indicator is used, in which the average distance is calculated based on the distance between the closest point pairs.
[0043] The object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm of the present invention has at least the following beneficial effects:
[0044] (1) A 6-Dof pose estimation method for target objects using Yolov8 is proposed. Compared with most other detection methods based on key points, it has better key point extraction capabilities. The anchor-free strategy and multi-scale feature maps can better detect key points of different sizes and can effectively detect key points of different distances and sizes.
[0045] (2) A lightweight Yolov8 network is proposed for 6-Dof pose estimation, which greatly reduces the number of parameters and computation while only slightly reducing the accuracy. It is easy to deploy on various platforms, including cloud services, edge devices, and mobile devices.
[0046] (3) A lightweight Yolov8 network using feature fusion is proposed and used for 6-Dof pose estimation. It improves the accuracy while only slightly increasing the number of parameters and computation without losing lightweight, achieving better 6-Dof pose estimation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is the overall flow chart of 6-Dof pose estimation based on lightweight Yolov8
[0048] Figure 2 It is the basic structure of Yolov8-6D
[0049] Figure 3 Is the structure of StemBlock and Shuffle_Block
[0050] Figure 4 It is the lightweight Yolov8 network Yolov8-lite-6D-v1 structure diagram
[0051] Figure 5 The principle of CGA
[0052] Figure 6 It is based on the principle of CGA feature fusion mechanism
[0053] Figure 7 It is the feature fusion structure of Yolov8 based on CGA
[0054] Figure 8 It is a lightweight Yolov8 network Yolov8-lite-6D-v2 structure using CGA feature fusion
[0055] Fig. 9 is the reprojection error and ADD index of Yolov8-6D
[0056] Fig.10 This is the performance comparison of Yolov8-6D and ADD-0.1d of different methods of the same type and conditions Fig.11 Is the lightweight performance of the model
[0057] Fig.12 is the ADD performance index of the lightweight method
[0058] Fig.13 It is an indicator of lightweight and key point detection capability after feature fusion
[0059] Fig.14 is the 6-Dof pose estimation performance after feature fusion ADD-0.1d
[0060] Fig.15 It is an Ape-type RGB image.
[0061] Fig.16 Comparison between Yolov8-lite-6D-v2 predicted Ape objects and true values
[0062] Fig.17 It is a Cat class RGB image
[0063] Fig.18 Comparison between Yolov8-lite-6D-v2 prediction of Cat objects and the true value
[0064] Fig.19 This is a schematic diagram of the Ape class using PVNet posture estimation
[0065] Fig. 20 This is a diagram of Cat class using PVNet posture estimation DETAILED DESCRIPTION
[0066] The object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm of the present invention comprises:
[0067] Step 1: Define the key points of the identified object;
[0068] Step 1.1: Use the FPS sampling algorithm and the model corner points to jointly select the three-dimensional coordinates of key points that are evenly distributed on the surface of the target model. The principle of the FPS algorithm is as follows:
[0069] First, the input point cloud (model) has N points. A point P0 is selected from the point cloud as the starting point to obtain the sampling point set S = {P0}. The initial point selection needs to select the farthest point from the center of gravity of the point cloud; calculate the distance from all points to P0 to form an N-dimensional array L, select the point corresponding to the maximum value as P1, and update the sampling point set to S = {P0, P1}; calculate the distance from all points to P1, for each point Pi, if its distance from P1 is less than L[i], then update L[i] = d(Pi, P1), and the array L always stores the closest distance from each point to the sampling point set S. Select the point corresponding to the maximum value in L as P2, and update the sampling point set S = {P0, P1, P2}; repeat steps 2 to 4 until the number of sampling points reaches the upper limit.
[0070] Step 1.2: The three-dimensional corner points of the model are selected as: [Xmin, Ymin, Zmin], [Xmin, Ymin, Zmax], [Xmin, Ymax, Zmin], [Xmin, Ymax, Zmax], [Xmax, Ymin, Zmin], [Xmax, Ymin, Zmax], [Xmax, Ymax, Zmin], [Xmax, Ymax, Zmax]. Among them, Xmin, Ymin, and Zmin are the minimum values of X, Y, and Z in the model coordinate system respectively, and Xmax, Ymax, and Zmax are the maximum values of X, Y, and Z in the model coordinate system respectively.
[0071] Step 1.3: Combine all the selected corner points with the position information of the target object in each image to project the target model and map it to the pixel coordinate system. The mapping method is:
[0072]
[0073] Among them, u and v are the pixel coordinates obtained by transformation, X W , Y W , Z W It is the three-dimensional coordinate in the world coordinate system. is the camera internal parameter, is the camera external parameter, that is, the target object pose information.
[0074] Step 2: Implement a 6-Dof pose estimation network using Yolov8;
[0075] Step 2.1: Use Yolov8's 6-Dof pose estimation network Yolov8-6D. The backbone network of Yolov8 is Darknet-53, in which the C2f module replaces the C3 module in the previous Yolo series. The detection head of the head network is modified so that it can detect the two-dimensional coordinate values of the center point and eight corner points of the target object and n FPS sampling points. For the coordinates of the center point and the corner point, since Yolov8 is an anchor-free structure, it is not necessary to obtain the coordinate values of the corner point by predicting the offset. For training of objects of different sizes, use a confidence function that adapts to training of different sizes:
[0076]
[0077] Where α is the control parameter, D T (x) is the Euclidean distance between the predicted coordinates and the true coordinates, d th is the confidence threshold.
[0078] Step 2.2: The loss function consists of multiple parts:
[0079] L=λ ciou Lciou +λ cls L cls +λ dfl L dfl +λ kobj L kobj +λ pose L pose (3)
[0080] Where L ciou It is the complete intersection-over-union loss proposed in "Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation", L cls is the category cross entropy loss, L dfl It is the probability density loss proposed in "Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection", L kobj is the key point loss, which measures whether the key point exists, L pose is the key point position loss, is the sum of the Euclidean distance differences between the predicted key points and the true values of the key points. λ is the weight of each loss. pose Set to the maximum, in practice it should be greater than or equal to 10.
[0081] Step 2.3: To be more suitable for the 6-Dof pose estimation task, the evaluation function used in the network training process is set to:
[0082]
[0083] Among them, m is the number of targets, R and T are the actual rotation and translation, The predicted and solved rotation and translation of the current key point, x is the three-dimensional coordinate of the point, and M represents the target object.
[0084] Step 2.4: For all the predicted key points, use the EPnP algorithm in the opencv library to calculate the pose information of the target object.
[0085] Step 3: The 6-Dof pose estimation network using Yolov8 is lightweight and improved to obtain a network structure named Yolov8-1ite-6D-v1;
[0086] Step 3.1: Replace all convolutional modules and C2f modules in the backbone detection part with the shuffle_Block lightweight module. This lightweight module mainly feeds the input into the depth-separable convolutional layer DWConvBlock and the 1x1 convolutional layer, and rearranges the channel information of the grouped output, thereby reducing the amount of calculation and the number of parameters of the model.
[0087] Step 3.2: Replace all C2f modules in the Neck part with depthwise separable convolutional layers. The replacement mainly requires determining the number of input and output channels and size of each C2f module, and then constructing a depthwise separable convolutional layer module. The depthwise separable convolutional module is mainly divided into two parts, one is the depthwise convolution part, and the other is the pointwise convolution part. Construct a 1x1 convolutional layer to adjust the number of channels. Finally, combine the depthwise convolution and pointwise convolution and call them in sequence.
[0088] Step 3.3: A StemBlock layer is connected after the front-end input layer of the network. Its main function is to improve the efficiency of initial feature extraction, reduce data dimension and calculation amount, and adapt to the structural requirements of different networks. The main structure is the combination of convolutional layers and pooling operations.
[0089] Step 3.4: Get the lightweight network Yolov8-lite-6D-v1.
[0090] Step 4: Improve the feature fusion of the lightweight network Yolov8-lite-6D-v1 to increase the accuracy and obtain the final lightweight network Yolov8-lite-6D-v2;
[0091] Step 4.1: Directly mix the three low-level feature maps with the three high-level feature maps respectively, and then calculate the weights according to CGA. Use the complementary weights as the weights of the two feature maps, and remix the original low-level feature maps and the high-level feature maps after weighting to obtain a fused feature map with content-guided attention mechanism.
[0092] Step 4.2: Get the final network Yolov8-lite-6D-v2 after feature fusion.
[0093] Step 5: Quantitatively evaluate the algorithm results;
[0094] Step 5.1: Use the mAP indicator to evaluate the key point prediction ability of the model. The calculation of the mAP indicator is usually based on OKS, and the calculation method is:
[0095]
[0096] Among them, i is the serial number of the key point, d is the Euclidean distance between the predicted value of the key point and the true value of the key point, S is the scale factor, w and h are the width and height of the detection box. σ is the normalization factor of the key point, which is calculated by the standard deviation between the manual annotation and the true value of the real key points in all samples. The larger σ is, the more difficult it is to annotate the key point. δ(v i >0) is the indicator function, if v i >0 is 1, otherwise it is 0, which usually indicates the visibility of the key point. When OKS is greater than the threshold, the prediction is considered correct. Sort the predicted key points according to the confidence of the prediction results. Starting from high confidence, calculate the accuracy and recall of each prediction result in turn to obtain a series of accuracy-recall data points. AP is calculated by integrating the data points. Average the AP of all different key points to get mAP. The mAP when the threshold is 0.5 is mAP50. The average mAP calculated from the threshold starting from 0.5 and increasing in steps of 0.05 to 0.95 is mAP50-95.
[0097] Step 5.2: Use the reprojection error to evaluate the performance of 6-Dof pose estimation. Use the predicted 6-Dof pose information to reproject each point in the known object model into the new image. The specific projection method is consistent with the projection calculation method in step 1.3. Calculate the overlapping pixels of the target object in the original image and the newly obtained projected image. If the number of non-overlapping pixels is less than 5, the pose estimation result is considered correct.
[0098] Step 5.3: Use the ADD indicator to evaluate the performance of 6-Dof pose estimation. The three-dimensional point cloud in the object model is transformed into the corresponding rotation and translation using the real 6-Dof pose information and the predicted 6-Dof pose information, and the average distance between each identical point pair in the two is calculated. When the average distance does not exceed 10% of the maximum diameter of the model, the pose estimation result is considered correct. This indicator is called ADD-0.1d. For symmetrical objects, the ADD-S indicator is used, in which the average distance is calculated based on the distance between the closest point pairs.
[0099] The present invention is further described below with reference to examples.
[0100] Embodiment 1:
[0101] In order to prove the effectiveness of various methods of 6-Dof pose estimation proposed in this patent, the mAP indicator on the COCO key point detection dataset is used to verify the key point detection performance of the network. The performance of 6-Dof pose estimation is evaluated on the LINEMOD dataset using the 2D reprojection error proj2D-5px and ADD-0.1d indicators. The experiment should follow the following rules: use the same dataset control group. The experiment compares different types of the same algorithms, that is, different key point-based, non-end-to-end 6-Dof pose estimation algorithms with single RGB images as input to show the effectiveness of the method. According to Fig. 9 and Fig.10 The results show that in the case of a two-stage, non-end-to-end method that uses a single RGB image input and does not use pose refinement, the method using Yolov8's 6-Dof pose estimation performs well among methods of exactly the same type, and is comparable to the SOTA method of the same type, CDPN.
[0102] Based on Yolov-6D, it is lightweighted and improved. The network after lightweight improvement using StemBlock, Shuffle_Block and DWConvBlock is named Yolov8-lite-6D-v1. Fig.11 The COCO dataset verifies that the improved network has good key point detection capabilities while being lightweight. The lightweight model reduces the number of parameters by about 90% and the amount of computation by more than 80%. In terms of key point detection capabilities, the coarse detection capability mAP50 has basically remained unchanged, and the fine detection capability indicator mAP50-95 has decreased by 6%. Compared with the degree of lightweighting, the degree of accuracy reduction is very small, and a good lightweighting has been achieved.
[0103] Fig.12 The 6-Dof pose estimation metrics of Yolov8-6D and the lightweight network Yolov8-lite-6D-v1 are compared. In addition, the performance of these models is compared with PoseCNN and PVNet, as well as the equally lightweight method Yolov7-6D. The reason for lightweight development and optimization is that in addition to meeting the requirements of some edge devices, the current highest-precision methods, such as RNNPose, are rendering methods that cannot run independently and require other methods for initial rough pose estimation. The more accurate the initial estimate, the faster and more accurate the rendering methods. Existing methods such as PoseCNN and PVNet are often used for such initial estimates. Compared with these methods, Yolov8-6D achieves the highest accuracy and speed among them. Yolov8-lite-6D-v1 has a 5% performance reduction, but the speed is the fastest of all methods. So lightweight processing makes good sense.
[0104] On the basis of the lightweight network Yolov8-lite-6D-v1, in order to improve some of the accuracy, a feature fusion mechanism based on CGA is added, and the network after feature fusion is named Yolov8-lite-6D-v2. Fig.13 This shows that the model has improved its key point detection capability after adding this feature fusion mechanism. Fig.14 This method also improves the performance of 6-Dof pose estimation. It can be shown that this feature fusion mechanism and method can improve the network's key point detection ability and 6-Dof pose estimation ability. Most importantly, this fusion method only increases the amount of calculation by a very small amount, about 0.3Gflops, and reduces the estimation speed by about 10FPS, while still maintaining the network's good lightweight characteristics.
[0105] Embodiment 2:
[0106] The evaluation of this patent is based on the LINEMOD dataset, the most classic and authoritative public dataset of 6-Dof. The training of the model, improved model and comparison model are all under the same parameter conditions, and the factors that can affect the training effect, such as the training hyperparameters, the number of rounds, and data preprocessing, are strictly controlled. The input of the algorithm is a single RGB image, and the output is 6-Dof posture information.
[0107] The trained Yolov8-lite-6D-v2 network is used to predict the RGB image. First, eight corner points plus 16 key points of the farthest point sampling are obtained, and the predicted pose is obtained through the EPnP algorithm. Fig.15 , Fig.17 is one of the RGB images in the test set, where Fig.15 For Ape objects, Fig.17 It is a Cat object. Fig.16 , Fig.18 The accuracy of the object's 6-Dof posture estimation is intuitively displayed through point cloud visualization, that is, the ADD indicator is intuitively displayed. The green point cloud is the true value of the object's posture point cloud, and the red point cloud is the predicted value of the object's posture point cloud. The more the red point cloud and the green point cloud overlap, the more accurate the predicted value of the posture estimation is.
[0108] In the comparative example, the data used is kept unchanged, and the method Yolov8-lite-6D-v2 proposed in the patent in the embodiment is replaced by the same type of method PVNet for experiment.
[0109] Fig.19 and Fig. 20 and Fig.16 and Fig.18In comparison, the overlap of point clouds is worse, the dispersion between green point clouds and red point clouds is more serious, and the pose estimation effect is worse. The experimental results show that the method of using Yolov8-lite-6D-v2 method to perform 6-Dof pose estimation on objects has better accuracy than other methods of the same type, and the obtained point cloud overlap is higher.
[0110] This patent achieves high precision and real-time performance of 6-Dof posture estimation, and has great practical application significance.
Claims
1. A 6-Dof pose estimation method for an object based on a lightweight Yolov8 algorithm, characterized in that: include: Step 1: Define the key points of the identified object; Step 2: Implement a 6-Dof pose estimation network using Yolov8; Step 3: The 6-Dof pose estimation network using Yolov8 is lightweight and improved to obtain a network structure named Yolov8-lite-6D-v1; Step 4: Improve the feature fusion of the lightweight network Yolov8-lite-6D-v1 to increase the accuracy and obtain the final lightweight network Yolov8-lite-6D-v2; Step 5: Quantitatively evaluate the algorithm results.
2. The object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm as claimed in claim 1, characterized in that: The step 1 is specifically as follows: Step 1.1: Use the FPS sampling algorithm and the model corner points to jointly select the three-dimensional coordinates of key points that are evenly distributed on the surface of the target model. The principle of the FPS algorithm is as follows: First, the input point cloud (model) has N points, and a point P0 is selected from the point cloud as the starting point to obtain a sampling point set S = {P0}. The initial point selection needs to select the farthest point from the center of gravity of the point cloud; Calculate the distances from all points to P0 to form an N-dimensional array L, select the point corresponding to the maximum value as P1, and update the sampling point set to S = {P0, P1}; Calculate the distance from all points to P1. For each point Pi, if its distance from P1 is less than L[i], update L[i] = d(Pi, P1). The array L always stores the shortest distance from each point to the sampling point set S. Select the point corresponding to the maximum value in L as P2, and update the sampling point set S = {P0, P1, P2}; Repeat the steps until the number of sampling points reaches the upper limit. Step 1.2: The three-dimensional corner points of the model are selected as: [Xmin, Ymin, Zmin], [Xmin, Ymin, Zmax], [Xmin, Ymax, Zmin], [Xmin, Ymax, Zmax], [Xmax, Ymin, Zmin], [Xmax, Ymin, Zmax], [Xmax, Ymax, Zmin], [Xmax, Ymax, Zmax]. Among them, Xmin, Ymin, and Zmin are the minimum values of X, Y, and Z in the model coordinate system respectively, and Xmax, Ymax, and Zmax are the maximum values of X, Y, and Z in the model coordinate system respectively. Step 1.3: Combine all the selected corner points with the pose information of the target object in each image to project the target model to the pixel coordinate system. The mapping method is: Among them, u and v are the pixel coordinates obtained by transformation, X W , Y W , Z W It is the three-dimensional coordinate in the world coordinate system. is the camera internal parameter, is the camera external parameter, that is, the target object pose information.
3. The object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm as claimed in claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1: Use Yolov8's 6-dof pose estimation network Yolov8-6D. The backbone network of Yolov8 is Darknet-53, in which the C2f module replaces the C3 module in the previous Yolo series, and the detection head of the head network is modified so that it can detect the two-dimensional coordinate values of the center point and eight corner points of the target object and n FPS sampling points. For the coordinates of the center point and the corner point, since Yolov8 is an anchor-free structure, it is not necessary to obtain the coordinate values of the corner point by predicting the offset. For training of objects of different sizes, use a confidence function that adapts to training of different sizes: Where α is the control parameter, D T (x) is the Euclidean distance between the predicted coordinates and the true coordinates, d th is the confidence threshold. Step 2.2: The loss function consists of multiple parts: L=λ ciou L ciou +λ cls L cls +λ dfl L dfl +λ kobj L kobj +λ pose L pose (3) Where L ciou It is the complete intersection-over-union loss proposed in "Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation", L cls is the category cross entropy loss, L dfl It is the probability density loss proposed in "Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection", L kobj is the key point loss, which measures whether the key point exists, L pose is the key point position loss, is the sum of the Euclidean distance differences between the predicted key points and the true values of the key points. λ is the weight of each loss. pose Set to the maximum, in practice it should be greater than or equal to 10. Step 2.3: To be more suitable for the 6-dof pose estimation task, the evaluation function used in the network training process is set to: Among them, m is the number of targets, R and T are the actual rotation and translation, The predicted and solved rotation and translation of the current key point, x is the three-dimensional coordinate of the point, and M represents the target object. Step 2.4: For all the predicted key points, use the EPnP algorithm in the opencv library to calculate the pose information of the target object.
4. The object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm as claimed in claim 1, characterized in that: The step 3 is specifically as follows: Step 3.1: Replace all convolutional modules and C2f modules in the backbone detection part with the shuffle_Block lightweight module. This lightweight module mainly feeds the input into the depth-separable convolutional layer DWConvBlock and the 1x1 convolutional layer, and rearranges the channel information of the grouped output, thereby reducing the amount of calculation and the number of parameters of the model. Step 3.2: Replace all C2f modules in the Neck part with depthwise separable convolutional layers. The replacement mainly requires determining the number of input and output channels and size of each C2f module, and then constructing a depthwise separable convolutional layer module. The depthwise separable convolutional module is mainly divided into two parts, one is the depthwise convolution part, and the other is the pointwise convolution part. Construct a 1x1 convolutional layer to adjust the number of channels. Finally, combine the depthwise convolution and pointwise convolution and call them in sequence. Step 3.3: A StemBlock layer is connected after the front-end input layer of the network. Its main function is to improve the efficiency of initial feature extraction, reduce data dimension and calculation amount, and adapt to the structural requirements of different networks. The main structure is the combination of convolutional layers and pooling operations. Step 3.4: Get the lightweight network Yolov8-lite-6D-v1.
5. The object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm as claimed in claim 1, characterized in that: The step 4 is specifically as follows: Step 4.1: Directly mix the three low-level feature maps with the three high-level feature maps respectively, and then calculate the weights according to CGA. Use the complementary weights as the weights of the two feature maps, and remix the original low-level feature maps and the high-level feature maps after weighting to obtain a fused feature map with content-guided attention mechanism. Step 4.2: Get the final network Yolov8-lite-6D-v2 after feature fusion.
6. The object 6-Dof posture estimation method based on the lightweight Yolov8 algorithm as claimed in claim 1, characterized in that: The step 5 is specifically as follows: Step 5.1: Use the mAP indicator to evaluate the key point prediction ability of the model. The calculation of the mAP indicator is usually based on OKS, and the calculation method is: Among them, i is the serial number of the key point, d is the Euclidean distance between the predicted value of the key point and the true value of the key point, S is the scale factor, w and h are the width and height of the detection box. σ is the normalization factor of the key point, which is calculated by the standard deviation between the manual annotation and the true value of the real key points in all samples. The larger σ is, the more difficult it is to annotate the key point. δ(v i >0) is the indicator function, if v i >0 is 1, otherwise it is 0, which usually indicates the visibility of the key point. When OKS is greater than the threshold, the prediction is considered correct. Sort the predicted key points according to the confidence of the prediction results. Starting from high confidence, calculate the accuracy and recall of each prediction result in turn to obtain a series of accuracy-recall data points. AP is calculated by integrating the data points. Average the AP of all different key points to get mAP. The mAP when the threshold is 0.5 is mAP50. The average mAP calculated from the threshold starting from 0.5 and increasing in steps of 0.05 to 0.95 is mAP50-95. Step 5.2: Use the reprojection error to evaluate the performance of 6-Dof pose estimation. Use the predicted 6-dof pose information to reproject each point in the known object model into the new image. The specific projection method is the same as the projection calculation method in step 1.
3. Calculate the overlapping pixels of the target object in the original image and the newly obtained projected image. If the number of non-overlapping pixels is less than 5, the pose estimation result is considered correct. Step 5.3: Use the ADD indicator to evaluate the performance of 6-Dof pose estimation. The three-dimensional point cloud in the object model is transformed into the corresponding rotation and translation using the real 6-dof pose information and the predicted 6-dof pose information, and the average distance between each identical point pair in the two is calculated. When the average distance does not exceed 10% of the maximum diameter of the model, the pose estimation result is considered correct. This indicator is called ADD-0.1d. For symmetrical objects, the ADD-S indicator is used, in which the average distance is calculated based on the distance between the closest point pairs.