Transparent object grabbing method and system based on 6D pose estimation

By using the TOG-Net object detection module and particle swarm optimization algorithm in the robotic arm grasping system, the problem of difficulty in identifying and positioning of transparent objects is solved, and high-precision and stable transparent object grasping is achieved.

CN120023837AActive Publication Date: 2025-05-23HUNAN INSTITUTE OF ENGINEERING

Patent Information

Application Number
CN202510517680.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

When facing transparent objects, robotic arm grasping technology based on 6D pose estimation has difficulties in accurately identifying and positioning, especially when transparent objects lack obvious color or texture features.

Method used

A transparent object grasping method based on 6D pose estimation is adopted, and the object detection is performed through the TOG-Net object detection module, the target features are extracted and similarly matched with the reference image library, and the preliminary 6D pose is generated and refined and adjusted. Combining particle swarm optimization algorithm and polynomial interpolation function, the trajectory of the robotic arm is planned for precise grasping.

Benefits of technology

The recognition and positioning accuracy of transparent objects is improved, ensuring that the robotic arm can track smoothly and continuously, avoiding sudden changes, violent acceleration or jitter in joint movement, thereby achieving stable grasping of transparent objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120023837A_ABST
    Figure CN120023837A_ABST
Patent Text Reader

Abstract

The invention discloses a transparent object capturing method and system based on 6D pose estimation, and the method comprises the steps: 1, inputting an original input picture containing a transparent object into a target detection module, and obtaining a 2D bounding box and a confidence score of a target object; 2, extracting target features from the preprocessed original input picture, and performing similarity matching on the target features and features in a reference image library to obtain a reference image with the highest similarity; 3, generating a preliminary 6D pose of the target feature, and then performing refined adjustment on the preliminary 6D pose to obtain a final 6D pose of the target feature; 4, dividing a path between the mechanical arm and the target features into three sections, converting the three sections into a polynomial interpolation function, and solving the polynomial interpolation function by using a particle swarm optimization algorithm to obtain an optimized trajectory; and 5, the mechanical arm grabs the target features according to the optimized track. According to the method, the accuracy of estimating the 6D pose of the transparent object is improved, and the grabbing precision of the transparent object is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target grasping, and in particular to a transparent object grasping method and system based on 6D pose estimation. Background Art

[0002] In recent years, with the improvement of industrial automation, computer vision, as an important research field of artificial intelligence, has been widely used in various industries. Among them, vision-based robotic grasping has gradually become a current research hotspot.

[0003] When faced with some special objects, such as transparent objects, machine vision may not be able to accurately identify and locate the target object. Sajjan et al. proposed a two-stage method for depth recovery of transparent objects, which first estimates surface normals, occlusion boundaries and segmentation from RGB images, and then calculates the refined depth through global optimization. However, the optimization is very time-consuming and heavily dependent on previous network predictions. Transparent objects lack obvious color or texture features and cannot be accurately identified.

[0004] In addition, vision-based robotic grasping is also a current research hotspot, but the grasping of transparent objects is not mature. ChenWang et al. proposed a network for estimating the 6D pose of objects from RGB-D images. The pose of objects is directly predicted based on RGB-D image information. Its distinctive features include dense prediction, confidence setting and new pose iterative optimization algorithm. However, there is a certain limitation that its generalization ability is limited. It can only accurately predict the trained targets that have been seen, but it is difficult to guarantee the accuracy of the prediction results for new targets that have never been seen.

[0005] In summary, robotic grasping based on 6D pose estimation methods has made significant progress. However, 6D pose estimation and grasping still face great challenges when facing transparent and texture-less target objects. Summary of the invention

[0006] The present invention provides a transparent object grasping method and system based on 6D pose estimation to solve the technical problems mentioned in the background technology.

[0007] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0008] The present invention provides a transparent object grasping method based on 6D pose estimation, comprising: S1. Input the original input image containing transparent objects into the TOG-Net target detection module to obtain the category, 2D bounding box and confidence score of the target object in the original input image; S2, extracting target features from the preprocessed original input image according to the category, 2D bounding box and confidence score of the target object, and performing similarity matching with the features in the reference image library to obtain the reference image with the highest similarity; S3, based on the reference features in the reference image with the highest similarity, estimate the rotation and translation information of the plane where the target feature is located, generate a preliminary 6D pose of the target feature, and then refine and adjust the preliminary 6D pose of the target feature to obtain the final 6D pose of the target feature; S4, dividing the path between the robot arm and the target feature into three segments, converting the three segments into polynomial interpolation functions, and solving the polynomial interpolation functions using a particle swarm optimization algorithm to obtain an optimized trajectory; S5. The robot arm grasps the target features according to the optimized trajectory.

[0009] On the other hand, the present invention also provides a transparent object grasping system based on 6D pose estimation, including a robotic arm, a visual perception device, an electric gripper and a built-in PC unit, the visual perception device and the electric gripper are both installed at the end of the robotic arm, the built-in PC unit is installed below the robotic arm, and the robotic arm, the visual perception device and the electric gripper are all electrically connected to the built-in PC unit; the robotic arm grasps the transparent object according to the transparent object grasping method.

[0010] Beneficial effects of the present invention: 1. The present invention discloses a transparent object grasping method based on 6D pose estimation, which uses a hybrid encoder. The hybrid encoder includes an intra-scale feature interaction module AIFI and a cross-scale fusion module CCFM connected in sequence. The intra-scale feature interaction module AIFI achieves a more efficient feature representation capability by recombine different parts of the input feature map, and the cross-scale fusion module CCFM achieves a more efficient feature representation capability by learning local and global contexts.

[0011] 2. In the transparent object grasping method based on 6D pose estimation disclosed in the present invention, a method combining a particle swarm optimization algorithm and a polynomial interpolation function is used to provide a smooth and continuous trajectory for the robotic arm, avoiding sudden changes, violent acceleration or jitter in joint movement. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a logic block diagram of the transparent object grasping method in the present invention; Figure 2 It is a structural block diagram of the backbone network in the TOG-Net target detection module of the present invention; Figure 3 It is a structural block diagram of the ConvNeXt network in the backbone network of the present invention; Figure 4It is a structural block diagram of the cross-scale fusion module CCFM in the present invention; Figure 5 A physical diagram of a transparent object grasping system in an embodiment of the present invention; Figure 6 exemplified diagram for comparing different methods for 6D pose estimation in an embodiment of the present invention with the present method; Figure 7 It is a graph of the unoptimized position, velocity and acceleration of each joint in the trajectory planning experiment in an embodiment of the present invention; Figure 8 It is a curve diagram of the optimized position, velocity and acceleration of each joint in the trajectory planning experiment in an embodiment of the present invention; Fig. 9 It is a motion trajectory curve diagram of the end of the robot arm in the trajectory planning experiment in the embodiment of the present invention; Fig.10 It is a convergence curve diagram of the robot arm joint 1 in the trajectory planning experiment in the embodiment of the present invention; Fig.11 It is a convergence curve diagram of the robot arm joint 2 in the trajectory planning experiment in the embodiment of the present invention; Fig.12 It is a convergence curve diagram of the robot arm joint 3 in the trajectory planning experiment in the embodiment of the present invention; Fig.13 It is a convergence curve diagram of the robot arm joint 4 in the trajectory planning experiment in the embodiment of the present invention; Fig.14 These are example diagrams of target grasping in different lighting environments in the embodiments of the present invention, where (a) is an example diagram of unobstructed grasping in normal lighting; (b) is an example diagram of unobstructed grasping in dark light; (c) is an example diagram of obstructed grasping in normal lighting; and (d) is an example diagram of obstructed grasping in dark light. DETAILED DESCRIPTION

[0013] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Preferred embodiments of the present invention are provided in the drawings. However, the present invention can be implemented in many other different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present invention more thorough and comprehensive.

[0014] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0015] Reference Figure 1, the embodiment of the present application provides a transparent object grasping method based on 6D pose estimation, comprising: S1. Input the original input image containing transparent objects into the TOG-Net target detection module to improve the diversity of features, thereby obtaining the category, 2D bounding box and confidence score of the target object in the original input image; S2, extracting target features from the preprocessed original input image according to the category, 2D bounding box and confidence score of the target object, and performing similarity matching with the features in the reference image library to obtain the reference image with the highest similarity; S3, based on the reference features in the reference image with the highest similarity, estimate the rotation and translation information of the plane where the target feature is located, generate a preliminary 6D pose of the target feature, and then refine and adjust the preliminary 6D pose of the target feature to obtain the final 6D pose of the target feature; S4, dividing the path between the robot arm and the target feature into three segments, converting the three segments into polynomial interpolation functions, and solving the polynomial interpolation functions using a particle swarm optimization algorithm to obtain an optimized trajectory; S5. The robot arm grasps the target features according to the optimized trajectory.

[0016] In some embodiments, the TOG-Net target detection module includes a backbone network, a network layer, and a prediction layer connected in sequence; Reference Figure 2 The backbone network includes a convolutional sampling layer and four stage units connected in series. The four stage units are the first stage unit to the fourth stage unit. The first stage unit includes a sequentially connected LN normalization layer and a ConvNeXt network (a computer vision model). The second stage unit to the fourth stage unit are sequentially connected downsampling layers and ConvNeXt networks. The specific structure of the ConvNeXt network is referenced Figure 3 As shown in the figure, it specifically includes a sequentially connected depth-separable convolutional layer, two convolutional sampling layers, and a layer scaling layer. The ConvNeXt network significantly improves the local feature extraction capability by improving the traditional convolutional layer CNN. The depth-separable convolutional layer in the ConvNeXt network can more effectively identify the contours and surface characteristics of transparent objects, and capture the reflection of glass or the refraction pattern of transparent materials through high-order feature layers.

[0017] The network layer is a hybrid encoder, which can achieve cross-channel and cross-scale fusion on the basis of retaining more information, reducing unnecessary redundant calculations; the hybrid encoder includes an intra-scale feature interaction module AIFI (intra-scale interaction) and a cross-scale fusion module CCFM (cross-scale fusion) connected in sequence; the input of the intra-scale feature interaction module AIFI is the output of the fourth stage unit, that is, the feature map The output of the intra-scale feature interaction module AIFI is the feature map The input of the cross-scale fusion module CCFM is the output of the second to fourth stage units, which are feature maps , feature map , feature map , expressed by the formula, as follows: ; in, Represents the features after fusion by the cross-scale fusion module CCFM; Reference Figure 4 As shown in the figure, the cross-scale fusion module CCFM inserts a fusion block composed of multiple convolutional layers into the fusion path, and fuses two adjacent scale features into a new feature; specifically, the cross-scale fusion module CCFM contains two 1×1 convolutional layers, two 1×1 convolutional layers to adjust the number of channels, and uses N RepBlocks composed of RepConv for feature fusion, and fuses the outputs of the two paths by element-wise addition.

[0018] The cross-scale fusion module CCFM improves the network's detection accuracy of transparent objects by adaptively fusing cross-scale and cross-channel feature information. Especially in the detection of transparent objects, the cross-scale fusion module CCFM uses multi-scale feature pyramid fusion to help the model better capture the details of transparent objects at different scales. At the same time, it introduces a channel attention mechanism to enhance the expression of important channel features, effectively suppress background interference, and improve the recognition of transparent objects.

[0019] In some embodiments, the S1 specifically includes the following steps: S11. First, the original input image of the transparent object Input into the convolution sampling layer in the backbone network for multi-level feature extraction to obtain the initial feature map , expressed by the formula, as follows: ; in, It is a 4×4 convolution operation; represents the real three-dimensional space, are the height and width of the image respectively; Indicates the number of channels of the convolution sampling layer; S12, the initial feature map The input is sent to four stage units to perform multi-stage multi-scale feature extraction to obtain multiple feature maps, namely, feature maps , feature map , feature map ;in, At that time, The multi-scale feature operation of the stage unit is expressed by the formula as follows: ; in, Indicates Feature map obtained by multi-scale feature operation of stage unit; It is The number of channels of the stage unit; Represents the set of operations of the LN normalization layer and the ConvNeXt network, or the set of operations of the downsampling layer and the ConvNeXt network; S13, feature map Flatten to a one-dimensional vector, generating a query matrix , key matrix Sum Matrix The input data is input into the intra-scale feature interaction module AIFI, the input data is processed by the intra-scale feature interaction module AIFI, and then the Reshape operation is used to restore the feature map. The same shape, get the feature map , expressed by the formula, as follows: ; in, Represents a flattening operation; represents the shape recovery operation, i.e. The inverse operation of S14, feature map , feature map And feature map All are input into the cross-scale fusion module CCFM to obtain multi-scale features; S15. Input the multi-scale features into the prediction layer to obtain the category, 2D bounding box and confidence score of the target object in the input image.

[0020] In some embodiments, S2 specifically includes the following steps: S21, preprocessing the original input image to obtain a preprocessed input image; S22, according to the category, 2D bounding box and confidence score output by the TOG-Net target detection module, crop the corresponding target area from the preprocessed input image to obtain a query image, and then perform multi-feature extraction in the corresponding target area (i.e., the query image) to obtain the extracted target features; S23, matching the extracted target features with the features in the reference image library, calculating the similarity scores between the target features and the reference features using a similarity metric, and selecting the first N reference images with the highest similarity scores as candidate viewpoints; The similarity score calculation process is as follows: First, the initial similarity score is calculated using the following formula; ; in, is the initial similarity score; represents the initial similarity score calculation operation; For two features, the initial similarity scores need to be calculated; Then the final similarity score is calculated based on the initial similarity score. The calculation formula of the final similarity score is as follows: ; in, represents the final similarity score, represents the weight coefficient; Represents the refraction path consistency score based on ray tracing; represents the query image; Indicates Reference images.

[0021] S24. Introduce refraction path consistency as an additional selection criterion, calculate the refraction path differences between the preprocessed input image and N reference images, and select the viewpoint with the smallest difference. The reference image corresponding to the viewpoint with the smallest difference is the reference image with the highest similarity.

[0022] In some embodiments, S3 specifically includes the following steps: S31, based on the reference features in the reference image with the highest similarity, estimate the rotation and translation information of the plane where the target feature is located, combine the rotation and translation information, and generate a preliminary 6D pose of the target feature; S32, training the posture refiner to obtain a trained posture refiner; Specifically, to train the pose refiner, 323 voxel points are sampled in the unit cube of the target coordinate system. Then, these voxel points are transformed to the input camera coordinate system by the input pose. The loss is the difference between the distance of the sample point transformed by the ground truth similarity transformation and the distance of the sample point transformed by the predicted similarity transformation; ; in are the sampling point coordinates in the input camera coordinate system, represents the set of real numbers, and are the predicted scale and the true scale, and are the predicted rotation and the true rotation respectively, and are the predicted 2D translation and the actual 2D translation, respectively. is the error term based on light refraction, Weight coefficient. It can be calculated by simulating the refraction path of light on the surface of a transparent object. Specifically, the difference in the refraction path between the predicted pose and the actual pose can be calculated by a ray tracing algorithm and used as an error term. This ensures that the final 6D pose estimation not only considers geometric transformations, but also combines the optical properties of transparent objects, thereby achieving more accurate and robust pose prediction.

[0023] S33. Use the trained pose refiner to adjust the preliminary 6D pose to obtain the output of the pose refiner, that is, the final 6D pose of the target feature.

[0024] In some embodiments, the output of the posture refiner in S33 is expressed by a formula, which is as follows: ; in, represents the output of the pose refiner, i.e., the final 6D pose of the target feature; Represents feature extraction operation; represents the preliminary 6D pose; Represents the 2D bounding box of the target object in the original input image.

[0025] Specifically, the adjustment process of the pose refiner for the preliminary 6D pose is as follows: First, the bounding box detected from the original input image The features are extracted and expressed by the formula as follows: ; in, Represents feature extraction operation; Then, based on the features extracted in the previous step, the initial 6D pose is optimized to obtain the pose correction , expressed by the formula, as follows: ; in, Indicates optimization operation; Finally, according to the attitude correction The final 6D pose of the target feature is calculated and expressed by the formula as follows: .

[0026] In some embodiments, the S4 specifically includes the following steps: S41, taking the initial position of the robot arm as the starting point and the final 6D pose of the target feature as the middle and end points; adding two path points between the starting point and the end point, so as to divide the path between the robot arm and the target feature into three sections, namely the first section to the third section; S42, using a cubic polynomial interpolation function in the first and third segments, and using a quintic polynomial interpolation function in the second segment, thereby completing the construction of the polynomial interpolation function; S43. A particle swarm optimization algorithm is used to solve the constructed polynomial interpolation function to obtain an optimized trajectory.

[0027] In some embodiments, the cubic polynomial interpolation of the first and third segments and the quintic polynomial interpolation of the second segment in S42 are specifically as follows: ; ; ; in, , and They are respectively The first segment of cubic polynomial, the second segment of quintic polynomial and the third segment of cubic polynomial trajectory of each joint; , , , Respectively The 3rd, 2nd, 1st, and 0th coefficients of the interpolation function of the first trajectory of each joint; , , , , , Respectively The 5th, 4th, 3rd, 2nd, 1st, and 0th coefficients of the interpolation function of the second trajectory of each joint; , , , Respectively The 3rd, 2nd, 1st, and 0th coefficients of the interpolation function of the third trajectory of each joint; , , Respectively The time of the first cubic polynomial, the second quintic polynomial and the third cubic polynomial trajectory corresponding to each joint.

[0028] In some embodiments, the S43 specifically includes the following steps: S431, initialize the particle swarm, generate the position and velocity of each particle, divide the sub-population, and then set the maximum number of iterations D, initialize the individual optimal value and the global optimal value; S432, update inertia weight , Group Learning Factor and self-learning factor , construct particle trajectory equation; inertia weight The calculation formula is as follows: ; in, , Inertia weight The maximum and minimum values ​​of ; k represents the number of iterations; Group learning factor and self-learning factor The calculation formula is as follows: ; ; S433, by deriving the particle trajectory equation, the velocity and acceleration equations of each particle are obtained, and then the velocity and acceleration of each particle are calculated by the velocity and acceleration equations of each particle; S434, judging whether the speed and acceleration of each particle satisfy the constraint conditions; if not, setting the particle fitness value to infinity until the conditions are satisfied; if all conditions are satisfied, entering S435; S435, based on self-learning factor Calculate the fitness and update the individual optimal value based on the fitness; then based on the inertia weight , individual optimal value and group learning factor Update the global optimal value; S436. Determine whether the current number of iterations reaches the maximum number of iterations D. If so, output the global optimal value, which is the optimized trajectory. If not, return to S432.

[0029] The following is a comparison and verification of the grasping effect of a transparent object grasping method based on 6D pose estimation provided by the present invention by combining a variety of different grasping methods: First, the hardware used in the present invention is described in detail, including a 6D pose grasping system of a robotic arm. The 6D pose grasping system of a robotic arm mainly includes a deep learning platform and a transparent object grasping system including a robotic arm. The GPU of the deep learning platform is NVIDIA Geforce RTX2080ti, and the built-in software environment is Ubuntu18.04 LTS system. The transparent object grasping system is a DobotCR series CR3 intelligent robot of Yuejiang Technology. The main components include: Realsense D435i depth camera, EPG50-060 series electric parallel gripper, NUC onboard microcomputer, etc. The transparent object grasping system is as follows: Figure 5 shown.

[0030] Due to the light transmittance of transparent objects, the Trans10K dataset only has plane-level label information such as 2D bounding boxes or masks, so it is only suitable for target detection and instance segmentation; the ClearGrasp dataset contains deep label information of transparent objects, which is mainly used for depth estimation, but still lacks specific 6D pose labels, so it is still not suitable for 6D pose estimation of transparent objects. Existing generalizable pose estimators either require high-quality object models or require additional depth maps or object masks during testing, which greatly limits the scope of application. Therefore, the present invention adjusts the strategy for making datasets, and the complex reflection and refraction phenomena of transparent objects do not need to be considered. Only some pose images of the target object are required, and the dataset GLASS8K is made for transparent objects. In addition, the images captured by the dataset GLASS8K have not been cropped, and only a single transparent object target is retained in scenes with other interferences, which is closer to the real scene than other datasets.

[0031] In order to test the effectiveness of the present invention and the dataset produced under different material objects, a transparent object was placed among a variety of objects including three interference objects and without occlusion, and a 6D pose estimation comparison experiment was conducted on the self-made dataset GLASS8K using the PVNet algorithm, Pi2Pose algorithm, 6-PACK algorithm and the method provided by the present invention. Figure 6 As shown: Figure 6 It can be seen that under normal lighting, the pose prediction effects of the 6-PACK algorithm and the Pi2Pose algorithm are poor, and large-scale pose offsets occur. Obviously, the depth problem of transparent objects affects the pose refinement of the ICP algorithm. Secondly, the PVNet algorithm predicts the target pose better in the 12th and 43rd frames, but there are large offsets in the 80th and 56th frames. The last line is the result of the method provided by the present invention. It can be seen that the method provided by the present invention is better and more accurate than the first three methods at different viewing angles.

[0032] The following quantitative analysis experiments are conducted by combining a variety of different data sets and a variety of different crawling methods: The quantitative analysis is divided into ablation experiments of target detection algorithms and comparative experiments with different algorithms on the GLASS8K dataset, Trans10K dataset, and ClearGrasp dataset.

[0033] TOG-Net target detection module ablation experiment: In order to verify the effectiveness of adding the ConvNeXt network and the cross-scale fusion module CCFM, the method provided by the present invention is compared with four algorithms using the YOLOV11 algorithm as a prototype: 1) Prototype YOLOV11 backbone network (C3K2); 2) In the YOLOV11 framework, FasterNet is used to replace the backbone network; 3) Add a lightweight cross-scale feature fusion module CCFM under the YOLOV11 framework; 4) In the TOG-Net target detection module provided by the present invention, the backbone network is replaced by the ConvNeXt network, and a lightweight cross-scale feature fusion module CCFM is added.

[0034] In order to quantitatively analyze the target detection performance, several common evaluation indicators are used, such as precision, recall, detection accuracy, etc. The results are shown in Table 1: Table 1: Ablation experiment of TOG-Net target detection module

[0035] As shown in Table 1, the performance of the prototype YOLOV11 backbone network and the YOLOV11 framework using the ConvNeXt network to replace the backbone network are first compared. With little change in the accuracy, recall and average precision indicators, the number of parameters of the model has dropped significantly from 5.10M to 4.00M, indicating that the computational complexity has been significantly reduced. Secondly, the effects of adding a lightweight cross-scale fusion module CCFM and a TOG-Net target detection module to the YOLOV11 framework are compared. The results show that, under the premise of only increasing 0.02M parameters, the accuracy of the TOG-Net target detection module is improved by about 1.2% compared with YOLOV11; among them, the detection accuracy of transparent objects is significantly improved. Therefore, under the condition that the GPU computing resources of the grasping platform are limited, the TOG-Net target detection module proposed in the present invention is more suitable for application in robotic arm grasping tasks.

[0036] Pose estimation experiment: In order to quantitatively analyze the method provided by the present invention, comparative analysis experiments were carried out on the PVNet algorithm, Pi2Pose algorithm, 6-PACK algorithm and the method provided by the present invention on the GLASS8K dataset, Trans10K dataset and ClearGrasp dataset, as shown in Table 2.

[0037] Table 2: Quantitative evaluation of 6D pose on the GLASS8K dataset

[0038] As can be seen from Table 2, the average prediction accuracy of the object of the method provided by the present invention reaches 95.6%, which is 6.4%, 6.9% and 2.8% higher than that of PVNet, Pi2Pose and 6-PACK algorithms respectively. It can be seen that the detection accuracy of target 3 is lower than that of 6-PACK algorithm, mainly because the surface texture of target 3 is more, while the transparency of other target objects is higher.

[0039] Next, some objects were selected on the Trans10K dataset and the method provided by the present invention was compared with the PVNet algorithm, Pi2Pose algorithm and 6-PACK algorithm respectively. The comparison results are shown in Table 3.

[0040] Table 3: Quantitative evaluation of 6D pose on the Trans10K dataset

[0041] As shown in Table 3, the method provided by the present invention has little improvement compared with other algorithms. The reason is that the research objects of the Trans10K dataset are mostly other items and do not contain many transparent objects. Therefore, they do not have the characteristics of transparent objects that will produce erroneous depth.

[0042] Table 4: Quantitative evaluation of 6D pose on the ClearGrasp dataset

[0043] The method provided by the present invention is compared with the PVNet algorithm, the Pi2Pose algorithm and the 6-PACK algorithm in the ClearGrasp data set, and the results are shown in Table 4. When the target object is a transparent object, the method provided by the present invention is significantly improved by 13.5% compared with PVNet and 6.2% compared with 6-PACK. For the pose estimation of transparent objects, the test results show that the accuracy of the method provided by the present invention is significantly improved.

[0044] Robotic arm trajectory planning experiment: In order to verify the robotic arm grasping method based on piecewise polynomial trajectory planning combined with particle swarm optimization and joint space planning, this embodiment is verified by simulation experiments using MATLAB Robotics Toolbox and Optimtool toolbox.

[0045] The improved power function operator is added to the inertia weight, and each particle is allowed to expand the search space based on the total number of iterations during the search, thereby increasing the population diversity. The operating parameters of the basic particle swarm algorithm are set as follows: Group learning factor The value is 1.5, the self-learning factor The value is 1.7, the population size S pop The value is 20, the population dimension is 10, and the experiment is run 20 times, with 500 iterations each time. The optimal trajectory planning is obtained as follows Figures 7 to 13 As shown: Reference Figure 7 and Figure 8 As shown in the figure, during the movement of the robot arm, each joint remains stable, the angular velocity and angular acceleration change smoothly without mutation, and the constraints of each joint of the robot arm are met. After optimization by the improved particle swarm algorithm, the position of joint 1 converges after 34 iterations, and the shortest time required to pass through three polynomial trajectories is 0.6385s, 0.7162s, and 0.446s, respectively, and the total time is less than the basic particle swarm algorithm. The position of joint 2 converges after 30 iterations, while the positions of joints 3 and 4 converge after 15 and 6 iterations, respectively, and the number of iterations required for convergence is significantly lower than the basic PSO algorithm. The results show that this method can effectively meet the grasping requirements of the robot arm in complex scenes.

[0046] Robotic arm grasping experiment: Considering the impact of transparent objects under different lighting conditions, this experiment mainly discusses the robotic arm grasping in complex scenes under the influence of lighting changes and obstacles. The grasped objects mainly use glass bottle reagents and other transparent objects in the GLASS8K dataset. In order to verify the performance of the method provided by the present invention in practical applications, this experiment conducted two groups of grasping: the first group: the robotic arm grasped in normal lighting and dark lighting environments without obstacles; the second group: the robotic arm grasped in normal lighting and dark lighting environments with obstacles. First, the depth camera detects the object and enables the robotic arm to grasp the object, and then performs the next grasp.

[0047] Table 5: Robotic arm grasping success rate in different environments (%)

[0048] As can be seen from Table 5, in the case of no obstacles; under normal light, the average grasping success rate of the robotic arm is 88.7%, while under low light, the average grasping success rate is only 79%, which is 9.7% lower than that under normal indoor light. Specifically, the clarity of the images captured by the camera in the low-light environment is not high, which is not conducive to feature extraction, so the grasping rate is higher under normal light. Secondly, in the environment with obstacles, the average grasping rate of the robotic arm under normal light is 85.7%, which is only 3% lower than that without obstacles, proving the effectiveness of the 3-5-3 particle swarm optimization algorithm proposed in this paper.

[0049] In addition, as can be seen from the results in the figure: under normal light conditions, the grasping success rate for transparent glass is relatively high. Specifically, the recognition of the pose will not be affected by the reflection of light on the glass, but the detection by the depth camera will be affected by light, resulting in a low grasping rate in the low-light environment. In the low-light environment, situations of grasping failure will occur, such as Fig.14 shown. Coupled with the occlusion of obstacles, the entire target object has no obvious texture and feature loss, resulting in a final grasping failure. The failure of grasping is also related to the force of the gripper. There are cases where the glass bottle is damaged during grasping. On this basis, in the experiment, foam is added on both sides of the gripper to effectively control the force and reduce the problems brought. The grasping success rate will decrease in the presence of obstacles, but only by 3%. Therefore, the method proposed in this paper can achieve successful grasping with or without obstacles.

[0050] Refer to Figure 5 , on the other hand, the present invention also provides a transparent object grasping system based on 6D pose estimation, including a robotic arm, a visual perception device, an electric gripper (EPG50-060 series electric parallel gripper), and a built-in PC unit (NUC airborne microcomputer). The visual perception device and the electric gripper are both installed at the end of the robotic arm, and the built-in PC unit is installed below the robotic arm. Moreover, the robotic arm, the visual perception device, and the electric gripper are all electrically connected to the built-in PC unit; the robotic arm grasps the transparent object according to the transparent object grasping method.

[0051] The visual perception device includes a network camera, a Realsence camera (Realsense D435i depth camera), a Kinect peripheral device, and a light field camera. The electric gripper is wrapped with cotton, and the robotic arm is a six-axis robotic arm.

[0052] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered within the protection scope of the present invention. In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the ability of ordinary technicians in the field to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A transparent object grasping method based on 6D pose estimation, characterized in that: include: S1. Input the original input image containing transparent objects into the TOG-Net target detection module to obtain the 2D bounding box and confidence score of the target object in the original input image. S2, extracting target features from the preprocessed original input image according to the 2D bounding box and confidence score of the target object, and performing similarity matching with the features in the reference image library to obtain the reference image with the highest similarity; S3, based on the reference features in the reference image with the highest similarity, estimate the rotation and translation information of the plane where the target feature is located, generate a preliminary 6D pose of the target feature, and then refine and adjust the preliminary 6D pose of the target feature to obtain the final 6D pose of the target feature; S4, dividing the path between the robot arm and the target feature into three segments, converting the three segments into polynomial interpolation functions, and solving the polynomial interpolation functions using a particle swarm optimization algorithm to obtain an optimized trajectory; S5. The robot arm grasps the target features according to the optimized trajectory.

2. The transparent object grasping method based on 6D pose estimation according to claim 1, characterized in that: The TOG-Net target detection module includes a backbone network, a network layer and a prediction layer connected in sequence; The backbone network includes a convolutional sampling layer and four stage units connected in series. The four stage units are the first stage unit to the fourth stage unit, wherein the first stage unit includes a sequentially connected LN normalization layer and a ConvNeXt network, and the second stage unit to the fourth stage unit are sequentially connected downsampling layers and ConvNeXt networks; The network layer is a hybrid encoder, which includes an intra-scale feature interaction module AIFI and a cross-scale fusion module CCFM connected in sequence; the input of the intra-scale feature interaction module AIFI is the output of the fourth stage unit, that is, the feature map The output of the intra-scale feature interaction module AIFI is the feature map The input of the cross-scale fusion module CCFM is the output of the second to fourth stage units, which are feature maps , feature map , feature map .

3. The transparent object grasping method based on 6D pose estimation according to claim 2 is characterized in that: The S1 specifically includes the following steps: S11. First, the original input image of the transparent object Input into the convolution sampling layer in the backbone network for multi-level feature extraction to obtain the initial feature map , expressed by the formula, as follows: ; in, It is a 4×4 convolution operation; represents the real three-dimensional space, are the height and width of the image respectively; Indicates the number of channels of the convolution sampling layer; S12, the initial feature map The input is sent to four stage units to perform multi-stage multi-scale feature extraction to obtain multiple feature maps, namely, feature maps , feature map , feature map ;in, At that time, The multi-scale feature operation of the stage unit is expressed by the formula as follows: ; in, Indicates Feature map obtained by multi-scale feature operation of stage unit; It is The number of channels of the stage unit; Represents the set of operations for downsampling layers and ConvNeXt networks; S13, feature map Flatten to a one-dimensional vector, generating a query matrix , key matrix Sum Matrix The input data is input into the intra-scale feature interaction module AIFI, the input data is processed by the intra-scale feature interaction module AIFI, and then the Reshape operation is used to restore the feature map. The same shape, get the feature map , expressed by the formula, as follows: ; in, Represents a flattening operation; represents the shape recovery operation, i.e. The inverse operation of S14, feature map , feature map And feature map All are input into the cross-scale fusion module CCFM to obtain multi-scale features; S15. Input the multi-scale features into the prediction layer to obtain the category, 2D bounding box and confidence score of the target object in the input image.

4. The transparent object grasping method based on 6D pose estimation according to claim 3 is characterized in that: The S2 specifically includes the following steps: S21, preprocessing the original input image to obtain a preprocessed input image; S22, according to the category, 2D bounding box and confidence score output by the TOG-Net target detection module, crop the corresponding target area from the preprocessed input image, and then perform multi-feature extraction in the corresponding target area to obtain the extracted target features; S23, matching the extracted target features with the features in the reference image library, calculating the similarity scores between the target features and the reference features using a similarity metric, and selecting the first N reference images with the highest similarity scores as candidate viewpoints; S24. Introduce refraction path consistency as an additional selection criterion, calculate the refraction path differences between the preprocessed input image and N reference images, and select the viewpoint with the smallest difference. The reference image corresponding to the viewpoint with the smallest difference is the reference image with the highest similarity.

5. The transparent object grasping method based on 6D pose estimation according to claim 4 is characterized in that: The S3 specifically includes the following steps: S31, based on the reference features in the reference image with the highest similarity, estimate the rotation and translation information of the plane where the target feature is located, combine the rotation and translation information, and generate a preliminary 6D pose of the target feature; S32, training the posture refiner to obtain a trained posture refiner; S33. Use the trained pose refiner to adjust the preliminary 6D pose to obtain the output of the pose refiner, that is, the final 6D pose of the target feature.

6. The transparent object grasping method based on 6D pose estimation according to claim 5, characterized in that: The output of the attitude refiner in S33 is expressed by a formula, which is as follows: ; in, represents the output of the pose refiner, i.e., the final 6D pose of the target feature; Represents feature extraction operation; represents the preliminary 6D pose; Represents the 2D bounding box of the target object in the original input image.

7. The transparent object grasping method based on 6D pose estimation according to claim 6, characterized in that: The S4 specifically includes the following steps: S41, taking the initial position of the robot arm as the starting point and the final 6D pose of the target feature as the middle and end points; adding two path points between the starting point and the end point, so as to divide the path between the robot arm and the target feature into three sections, namely the first section to the third section; S42, using a cubic polynomial interpolation function in the first and third segments, and using a quintic polynomial interpolation function in the second segment, thereby completing the construction of the polynomial interpolation function; S43. A particle swarm optimization algorithm is used to solve the constructed polynomial interpolation function to obtain an optimized trajectory.

8. The transparent object grasping method based on 6D pose estimation according to claim 7, characterized in that: The cubic polynomial interpolation of the first and third segments and the quintic polynomial interpolation of the second segment in S42 are specifically as follows: ; ; ; in, , and They are respectively The first segment of cubic polynomial, the second segment of quintic polynomial and the third segment of cubic polynomial trajectory of each joint; , , , Respectively The 3rd, 2nd, 1st, and 0th coefficients of the interpolation function of the first trajectory of each joint; , , , , , Respectively The 5th, 4th, 3rd, 2nd, 1st, and 0th coefficients of the interpolation function of the second trajectory of each joint; , , , Respectively The 3rd, 2nd, 1st, and 0th coefficients of the interpolation function of the third trajectory of each joint; , , Respectively The time of the first cubic polynomial, the second quintic polynomial and the third cubic polynomial trajectory corresponding to each joint.

9. The transparent object grasping method based on 6D pose estimation according to claim 8, characterized in that: The S43 specifically includes the following steps: S431, initialize the particle swarm, generate the position and velocity of each particle, divide the sub-population, and then set the maximum number of iterations D, initialize the individual optimal value and the global optimal value; S432, update inertia weight , Group Learning Factor and self-learning factor , construct particle trajectory equation; inertia weight The calculation formula is as follows: ; in, , Inertia weight The maximum and minimum values ​​of Indicates the number of iterations; Group learning factor and self-learning factor The calculation formula is as follows: ; ; S433, by deriving the particle trajectory equation, the velocity and acceleration equations of each particle are obtained, and then the velocity and acceleration of each particle are calculated by the velocity and acceleration equations of each particle; S434, judging whether the speed and acceleration of each particle satisfy the constraint conditions; if not, setting the particle fitness value to infinity until the conditions are satisfied; if all conditions are satisfied, entering S435; S435, based on self-learning factor Calculate the fitness and update the individual optimal value based on the fitness; then based on the inertia weight , individual optimal value and group learning factor Update the global optimal value; S436. Determine whether the current number of iterations reaches the maximum number of iterations D. If so, output the global optimal value, which is the optimized trajectory. If not, return to S432.

10. A transparent object grasping system based on 6D pose estimation, characterized in that: The invention comprises a robotic arm, a visual perception device, an electric gripper and a built-in PC unit, wherein the visual perception device and the electric gripper are both installed at the end of the robotic arm, the built-in PC unit is installed below the robotic arm, and the robotic arm, the visual perception device and the electric gripper are all electrically connected to the built-in PC unit; the robotic arm grasps the transparent object according to the transparent object grasping method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for calculating 6D attitude parameters of transparent object

    CN113313810A

  • Method for estimating 6D attitude of transparent object grabbed by mechanical arm

    CN114119753A

  • Transparent object tracking model and construction method and application thereof

    CN117115208A

  • Transparent object grabbing method based on depth completion

    CN117975165A

  • Handheld transparent object pose estimation method and robot grabbing control method

    CN118700130A

Cited By

  • Dizziness lamp control method and system based on human body depth estimation and target detection

    CN120786770A