Method and device for quickly estimating 6D attitude of target object
By constructing the cube bag of the target object and fitting the penalty function, and combining it with the gradient descent algorithm, the 6D pose of the target object can be quickly estimated directly from RGB and depth images. This solves the problems of computational complexity and real-time performance in existing technologies and achieves efficient 6D pose estimation.
Patent Information
- Application Number
- CN202511104436.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing methods for 3D model pose estimation of target objects face challenges in terms of computational complexity, generalization, real-time performance, and computational cost, especially in meeting real-time requirements on edge devices.
By acquiring the RGB and depth images of the target object, a binary mask and a surface point cloud matrix are constructed. The gradient vectors of the center point and Euler angles of the target object are calculated using the cubic bag and the fitting penalty function. The 6D pose is then iteratively solved using the gradient descent algorithm, which simplifies the calculation process and improves accuracy and real-time performance.
It can quickly estimate 6D pose without using neural networks and detailed 3D models, reducing computation and latency, and improving real-time performance and accuracy on edge devices.
Smart Images

Figure CN120976317A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic digital data processing in the field of machine vision and perception technology, and particularly to a method and apparatus for fast 6D pose estimation of a target object. Background Technology
[0002] Existing pose estimation methods for 3D models of target objects mainly fall into two categories: 1) feature matching-based methods; 2) template matching-based methods.
[0003] Feature-matching-based methods extract local features (such as point features, line features, or edge features) from images and match them with the 3D model of the target object. Finally, the 6D pose parameters are solved using the Perspective-n-Point (PnP) algorithm. The specific process includes: Feature extraction: Detecting key points or geometric structures from the input image to generate high-dimensional feature representations. Feature matching: Associating the extracted 2D features with corresponding features in the 3D model point cloud or CAD model to establish a 2D-3D correspondence. Pose solving: Based on the matching correspondence, the PnP algorithm is used to minimize reprojection errors and optimize rotation and translation parameters. This method has the advantage of being robust to partial occlusion and background noise, but it relies on accurate feature extraction and matching algorithms, resulting in low computational efficiency. Especially in complex scenes, it is prone to pose deviations due to matching errors.
[0004] Template matching-based methods first generate synthetic rendered images (templates) of the target 3D model from different viewpoints. Then, a deep neural network is used to calculate the similarity between the input image and the template, and the pose of the template with the highest similarity is taken as the estimation result. The core steps include: template generation: offline rendering of multi-view images of the 3D model, covering the complete viewpoint space; similarity calculation: using a CNN or Transformer network to extract depth features between the template and the input image, and quantifying the similarity score; pose regression: directly outputting the 6D pose by retrieving the pose parameters of the template with the highest score, or combining iterative optimization to improve accuracy. This method is highly effective for objects with rich textures, but requires pre-storing a large amount of template data, resulting in high storage and computational costs, and its generalization ability is limited, making it difficult to adapt to unseen or irregularly shaped objects.
[0005] As can be seen from the above, current technologies generally face the following challenges: Both methods are complex and highly dependent: they require detailed 3D models of the target object or multi-view image libraries, increasing the deployment threshold; template matching based on deep learning requires training a dedicated network, further exacerbating the complexity.
[0006] Generalization and real-time issues: Feature matching methods perform poorly on low-texture objects, and the computational complexity of the matching process increases quadratically with the number of feature points, making it difficult to meet real-time requirements. Template matching methods consume a lot of resources for rendering and similarity calculation, resulting in significant latency on edge devices, and are sensitive to changes in lighting and viewpoint, exhibiting weak generalization ability.
[0007] High computational load: Deep learning models have a large number of parameters, requiring high computing power during inference, and have low energy efficiency when deployed at the edge, affecting the feasibility of practical applications. Summary of the Invention
[0008] To address the above problems, this invention provides a fast 6D pose estimation method for a target object, comprising the following steps: Obtain the RGB image and depth image of the target object, and input the RGB image into the segmentation model to obtain the binary mask of the target object; Obtain the camera intrinsic parameter matrix, and construct the surface point cloud matrix of the target object based on the depth image, binary mask, and camera intrinsic parameter matrix; Construct the cube bag of the target object, and construct a fitting penalty function based on the surface point cloud matrix and the cube bag; Based on the surface point cloud matrix, cube bag, and fitting penalty function, the gradient vectors of the center point and Euler angles of the target object are calculated. The gradient vectors of the center point and Euler angles are iteratively solved using the gradient descent algorithm to obtain the optimal solutions for the center point and Euler angles. The optimal solutions for the center point and Euler angles are then used as the 6D pose of the target object.
[0009] Optionally, a surface point cloud matrix is formed by the coordinates of each pixel of the target object in the 3D camera coordinate system, where the X-axis of the 3D camera coordinate system is the vertical direction, the Y-axis is the horizontal direction, and the Z-axis is the optical axis direction of the RGBD camera.
[0010] Optionally, a fitting penalty function is constructed based on the surface point cloud matrix and the cubic bag, specifically including: S11: The cube bag of the target object is the smallest cube that encloses the target object. The height H, width W, and depth D of the cube bag are obtained through a measuring tool. S12: Extract the i-th pixel P from the surface point cloud matrix. i coordinates [X i ,Y i Z i ]; S13: Based on the height H, width W, depth D, and pixel P of the cube bag i coordinates [X i ,Y i Z i ], calculate to obtain pixel point P i The sum of the distances M to the six faces of the cube bagi ; S14: Repeat steps S12-S14 until the sum of the distances of N pixels is obtained. Loss is the fitting penalty function, where N represents the total number of pixels in the surface point cloud matrix.
[0011] Optionally, the independent variable for the center point of the target object includes the X-axis coordinate of the center point. c Y-axis coordinate of the center point c and the center point Z-axis coordinate z c The independent variables of the Euler angles of the target object include the first Euler angle ω0, the second Euler angle ω1, and the third Euler angle ω2.
[0012] Optionally, the calculation process for the gradient vector of the center point of the target object specifically includes: The height H, width W, and depth D of the cube bag are obtained through measurement tools. The coordinates of each pixel in the surface point cloud matrix are obtained. The unit vector n1 of the target object in the 3D camera coordinate system is obtained, which is oriented towards the X-axis, the unit vector n2 towards the Y-axis, and the unit vector n3 towards the Z-axis. Based on the height H, width W, depth D of the cube, the coordinates of each pixel, and the unit vector n1, the fitting penalty function Loss is calculated for the center point's X-axis coordinate x. c partial derivatives ; Based on the height H, width W, depth D of the cube, the coordinates of each pixel, and the unit vector n2, the fitting penalty function Loss is calculated for the center point's Y-axis coordinate y. c partial derivatives ; Based on the height H, width W, depth D of the cube, the coordinates of each pixel, and the unit vector n3, the fitting penalty function Loss is calculated for the center point's Z-axis coordinate z. c partial derivatives ; partial derivatives Partial derivatives and partial derivatives The combined row vectors ( , , The gradient vector is used as the center point.
[0013] Optionally, the calculation process of the gradient vector of the Euler angles of the target object specifically includes: The height H, width W, and depth D of the cube bag are obtained through measurement tools. The coordinates of each pixel in the surface point cloud matrix are obtained. The unit vector n1 of the target object in the 3D camera coordinate system is obtained, which is oriented towards the X-axis, the unit vector n2 towards the Y-axis, and the unit vector n3 towards the Z-axis. Based on the height H, width W, depth D of the cube bag, the coordinates of each pixel, and the unit vector n1, the partial derivative of the fitting penalty function Loss with respect to the first Euler angle ω0 is calculated. ; Based on the height H, width W, depth D of the cube bag, the coordinates of each pixel, and the unit vector n2, the partial derivative of the fitting penalty function Loss with respect to the second Euler angle ω1 is calculated. ; Based on the height H, width W, depth D of the cube bag, the coordinates of each pixel, and the unit vector n3, the partial derivative of the fitting penalty function Loss with respect to the third Euler angle ω2 is calculated. ; partial derivatives Partial derivatives and partial derivatives The combined row vectors ( , , ) is used as the gradient vector of Euler angles.
[0014] Optionally, the gradient vectors of the center point and Euler angles are iteratively solved using the gradient descent algorithm to obtain the optimal solutions for the center point and Euler angles. These optimal solutions are then used as the 6D pose of the target object, specifically including: Set the step size, step size decay factor, and iteration stopping parameters of the gradient descent algorithm, and calculate the optimal solution for the center point and Euler angles that minimizes the gradient vectors of the center point and Euler angles using the set gradient descent algorithm. The optimal solution for the center point includes: the optimal center point X-axis coordinate, the optimal center point Y-axis coordinate, and the optimal center point Z-axis coordinate; The optimal solutions for Euler angles include: the optimal first Euler angle, the optimal second Euler angle, and the optimal third Euler angle; The optimal center point X-axis coordinates, optimal center point Y-axis coordinates, optimal center point Z-axis coordinates, optimal first Euler angle, optimal second Euler angle, and optimal third Euler angle are used as the 6D pose of the target object.
[0015] The present invention also provides a fast 6D pose estimation device for a target object, used to implement the fast 6D pose estimation method for the target object, the device comprising: The binary mask acquisition module obtains the binary mask of the target object by inputting the RGB image and depth image of the target object into the segmentation model. The surface point cloud matrix acquisition module constructs the surface point cloud matrix of the target object based on the depth image, binary mask, and camera intrinsic parameter matrix by using the camera intrinsic parameter matrix to obtain the camera intrinsic parameter matrix. The fitting penalty function construction module constructs a fitting penalty function based on the surface point cloud matrix and the cubic bag used to construct the target object. The gradient vector acquisition module is used to calculate the gradient vectors of the center point and Euler angles of the target object based on the surface point cloud matrix, the cube bag, and the fitting penalty function. The 6D pose acquisition module is used to iteratively solve the gradient vectors of the center point and Euler angles using the gradient descent algorithm to obtain the optimal solution for the center point and Euler angles, and then use the optimal solution for the center point and Euler angles as the 6D pose of the target object.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for fast 6D pose estimation of the target object.
[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for fast 6D pose estimation of the target object.
[0018] The present invention has the following beneficial effects: 1. By taking the center point and Euler angles of the cube hull of the target object as the 6D pose of the target object, the 6D pose of the target object can be obtained simply by constructing the cube hull of the target object. There is no need to use a neural network or obtain the 3D model of the target object, which greatly simplifies the calculation process of the 6D pose of the target object. 2. Construct a fitting penalty function based on the surface point cloud matrix and the cube bag of the target object. The fitting penalty function can calculate the sum of the distances from each pixel in the surface point cloud matrix to the six faces of the cube bag. The fitting penalty function can fully reflect the relative position and containment relationship between the cube bag and the target object in the 3D camera coordinate system, thereby improving the calculation accuracy of the center point and Euler angles of the target object. 3. The gradient vectors of the center point and Euler angles are iteratively solved by the gradient descent algorithm to obtain the optimal solution of the center point and Euler angles. The optimal solution of the center point and Euler angles is used as the 6D pose of the target object. The 6D pose can be calculated iteratively by only the gradient descent algorithm with a small amount of computation, without the need to process a large amount of data. This greatly improves the real-time performance of obtaining the 6D pose and significantly reduces the latency of displaying the 6D pose on edge devices. Attached Figure Description
[0019] Figure 1 This is a flowchart of a method according to an embodiment of the present invention; Figure 2 A schematic diagram of a cubic bag for a yogurt beverage container; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0021] Reference Figure 1 This invention provides a fast 6D pose estimation method for a target object, comprising the following steps: Obtain the RGB image and depth image of the target object, and input the RGB image into the segmentation model to obtain the binary mask of the target object; In some embodiments, for RGB images captured by an RGBD camera, a trained instance segmentation model (such as the YOLOv8-seg model) is used to obtain the binary mask of the target object. The YOLOv8-seg model extends pixel-level segmentation capability on the basis of object detection, and supports simultaneous output of target bounding boxes, class labels and masks. It is suitable for fine-grained recognition tasks in complex scenes. The network consists of a Backbone module, a Neck module and a Head module. The Backbone module is built based on the CSPDarknet53 feature extraction network and introduces the C2f module to replace the C3 module of YOLOv5 to enhance gradient flow efficiency. The Neck module adopts PANet (Path Aggregation Network) or FPN (Feature Pyramid Network) to fuse multi-scale features. The Head module adopts a dual-branch design of detection head plus segmentation head. The detection head is used to detect the bounding box and class label of the target object. The segmentation head is used for segmentation operation based on mask coefficients (Box+Cls) + segmentation head (MaskCoefficients) dual-branch design. The core segmentation mechanism module is the Proto layer. The Proto layer uses an 80×80 feature map to generate a 160×160 32-channel prototype mask through convolutional upsampling. Each detector head outputs 32-dimensional coefficients, which are then multiplied by the Proto layer to generate the final binary mask.
[0022] Alternatively, a segmentation model based on cue words (SAM) can be used to estimate the binary mask of the target object. SAM adopts a Transformer architecture, supports zero-shot transfer and adaptive segmentation capabilities, and can generate pixel-level masks based on cues such as points, boxes, and text. SAM consists of three core modules working together: an image encoder, a cue encoder, and a mask decoder. The image encoder is used to extract image features; the cue encoder is used to encode interactive cues (points / boxes / text) into vectors, with points / boxes using positional encoding, text using CLIP encoding, and mask cues processed by convolution; the mask decoder is used to fuse image and cue features to generate a segmentation mask. The mask decoder is a lightweight Transformer decoder that outputs the mask through a cross-attention mechanism.
[0023] Obtain the camera intrinsic parameter matrix, and construct the surface point cloud matrix of the target object based on the depth image, binary mask, and camera intrinsic parameter matrix; In some embodiments, a surface point cloud matrix is formed by the coordinates of each pixel of the target object in the three-dimensional camera coordinate system, wherein the X-axis of the three-dimensional camera coordinate system is the vertical direction, the Y-axis is the horizontal direction, and the Z-axis is the optical axis direction of the RGBD camera.
[0024] In some embodiments, given a binary mask of the target object, the row coordinate vector U and column coordinate vector V of the target object's pixels can be obtained. Using Python, U and V can be obtained by calling the `numpy.where()` function. The length of the row coordinate vector U is N, and the depth of the vector [U[i], V[i]] of the i-th pixel is the Z-axis coordinate Z[i] of the pixel, where i represents the index of the row coordinate vector U (i∈[0,N-1]), U[i] represents the i-th real number in the row coordinate vector U, and V[i] represents the i-th real number in the column coordinate vector. Assuming the depth image acquired by the RGBD camera is called Depth, according to the above definition, then Z[i] = Depth[U[i], V[i]]. Based on the mapping relationship between the camera coordinate system and the image coordinate system, the X-axis and Y-axis coordinates of the i-th pixel in the surface point cloud matrix of the target object in the 3D camera coordinate system are represented as follows:
[0025]
[0026] in, This represents element-wise vector dot product. , , and The camera intrinsic parameters in the camera intrinsic parameter matrix K are represented as follows: The surface point cloud matrix S=[X,Y,Z] of the target object with dimensions of N rows and 3 columns can be constructed from the coordinates of all pixels. X represents the X-axis coordinate column vector of the pixel, Y represents the Y-axis coordinate column vector of the pixel, and Z represents the Z-axis coordinate column vector of the pixel.
[0027] Construct the cube bag of the target object, and construct a fitting penalty function based on the surface point cloud matrix and the cube bag; In some embodiments, a fitting penalty function is constructed based on the surface point cloud matrix and the cubic bag, specifically including: S11: The cube bag of the target object is the smallest cube that encloses the target object. The height H, width W, and depth D of the cube bag are obtained through a measuring tool. S12: Extract the i-th pixel P from the surface point cloud matrix. i coordinates [X i ,Y i Z i ]; S13: Based on the height H, width W, depth D, and pixel P of the cube bag i coordinates [X i ,Y i Z i ], calculate to obtain pixel point P i The sum of the distances M to the six faces of the cube bag i ; In some embodiments, pixel P i The sum of the distances M to the six faces of the cube bag i The expression is as follows:
[0028] Let P be the i-th pixel [X[i], Y[i], Z[i]] in the surface point cloud matrix S, and L be the pixel [X[i], Y[i], Z[i]]. j [i] can understand vectors In the normal vector n jThe projection onto the target object is shown, with O' as the center point of the cube bag and j=1,2,3. The unit vectors of the target object in the 3D camera coordinate system are n1 towards the X-axis, n2 towards the Y-axis, and n3 towards the Z-axis. If point P is inside the cube bag, the sum of its distances to the six faces reaches the minimum value H+W+D; otherwise, if it is outside the cube bag, the sum of its distances to the six faces will exceed H+W+D. Therefore, it is easy to predict that when the loss reaches the minimum value H+W+D, the cube bag will completely enclose the surface of the target object, meaning there cannot be any points protruding from the cube bag. Furthermore, if all six faces of the cube bag contain points from the point cloud S, then the center point of the cube bag coincides with the center point of the target object, and the orientation and pose are geometrically coincident with the target object (assuming the cube bag of the target object is unique; irregular objects are usually unique, only symmetrical objects such as spheres and cylinders have non-unique cube bags).
[0029] S14: Repeat steps S12-S14 until the sum of the distances of N pixels is obtained. Loss is the fitting penalty function, where N represents the total number of pixels in the surface point cloud matrix.
[0030] In some embodiments, the physical meaning of the fitting penalty function Loss is to calculate the sum of the distances from each pixel P in the surface point cloud matrix S to the six faces of the cube (including the two parallel faces to the left and right, the two parallel faces to the top and bottom, and the two parallel faces to the front and back, for a total of six planes).
[0031] Based on the surface point cloud matrix, cube bag, and fitting penalty function, the gradient vectors of the center point and Euler angles of the target object are calculated. In some embodiments, the independent variable of the center point of the target object includes the X-axis coordinate of the center point. c Y-axis coordinate of the center point c and the center point Z-axis coordinate z c The independent variables of the Euler angles of the target object include the first Euler angle ω0, the second Euler angle ω1, and the third Euler angle ω2.
[0032] In some embodiments, the 3D rotation matrix can be obtained based on Euler angles ω=[ω0,ω1,ω2]. Figure 2 This diagram shows a cubic bag of yogurt beverage containers, where O' = [x c ,y c ,z c [] represents the center point of the cube (yogurt drink), [H,W,D] represent the height, width, and depth of the yogurt drink, respectively, and the unit vectors n1, n2, and n3 represent the orientation of the yogurt drink along the X, Y, and Z axes in the 3D camera coordinate system, and are mutually orthogonal. These correspond to the three row vectors of the rotation matrix, i.e.:
[0033]
[0034]
[0035] In some embodiments, the fitting penalty function Loss is differentiable; the key point is to calculate the fitting penalty function Loss for the six independent variables [x]. c ,y c ,z c The partial derivatives of [ω0, ω1, ω2].
[0036] The calculation process of the gradient vector at the center point of the target object specifically includes: The height H, width W, and depth D of the cube bag are obtained through measurement tools. The coordinates of each pixel in the surface point cloud matrix are obtained. The unit vector n1 of the target object in the 3D camera coordinate system is obtained, which is oriented towards the X-axis, the unit vector n2 towards the Y-axis, and the unit vector n3 towards the Z-axis. Based on the height H, width W, depth D of the cube, the coordinates of each pixel, and the unit vector n1, the fitting penalty function Loss is calculated for the center point's X-axis coordinate x. c partial derivatives ; Based on the height H, width W, depth D of the cube, the coordinates of each pixel, and the unit vector n2, the fitting penalty function Loss is calculated for the center point's Y-axis coordinate y. c partial derivatives ; Based on the height H, width W, depth D of the cube, the coordinates of each pixel, and the unit vector n3, the fitting penalty function Loss is calculated for the center point's Z-axis coordinate z. c partial derivatives ; In some embodiments, let d1=W, d2=H, d3=D, , , For the gradient vector at the center point, according to the chain rule, we have the following formula:
[0037]
[0038] Among them, L j [i]= <S[i,:]-O’,n j >(j=1,2,3),<·> denotes the vector inner product operator, and S[i,:] denotes the i-th row vector in the surface point cloud matrix; partial derivatives Partial derivatives and partial derivatives The combined row vectors ( , , The gradient vector is used as the center point.
[0039] In some embodiments, the calculation process of the gradient vector of the Euler angles of the target object specifically includes: The height H, width W, and depth D of the cube bag are obtained through measurement tools. The coordinates of each pixel in the surface point cloud matrix are obtained. The unit vector n1 of the target object in the 3D camera coordinate system is obtained, which is oriented towards the X-axis, the unit vector n2 towards the Y-axis, and the unit vector n3 towards the Z-axis. Based on the height H, width W, depth D of the cube bag, the coordinates of each pixel, and the unit vector n1, the partial derivative of the fitting penalty function Loss with respect to the first Euler angle ω0 is calculated. ; Based on the height H, width W, depth D of the cube bag, the coordinates of each pixel, and the unit vector n2, the partial derivative of the fitting penalty function Loss with respect to the second Euler angle ω1 is calculated. ; Based on the height H, width W, depth D of the cube bag, the coordinates of each pixel, and the unit vector n3, the partial derivative of the fitting penalty function Loss with respect to the third Euler angle ω2 is calculated. ; In some embodiments, let d1=W, d2=H, d3=D. For the gradient vector of Euler angles, according to the chain rule, we have:
[0040] in:
[0041] Among them, L j [i]= <S[i,:]-O’,n j >(j=1,2,3),<·> denotes the vector inner product operator, and S[i,:] denotes the i-th row vector in the surface point cloud matrix;
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050] partial derivatives Partial derivatives and partial derivatives The combined row vectors ( , , ) is used as the gradient vector of Euler angles.
[0051] The gradient vectors of the center point and Euler angles are iteratively solved using the gradient descent algorithm to obtain the optimal solutions for the center point and Euler angles. The optimal solutions for the center point and Euler angles are then used as the 6D pose of the target object.
[0052] In some embodiments, the gradient vectors of the center point and Euler angles are iteratively solved using a gradient descent algorithm to obtain the optimal solutions for the center point and Euler angles. These optimal solutions are then used as the 6D pose of the target object. Specifically, this includes: Set the step size, step size decay factor, and iteration stopping parameters of the gradient descent algorithm, and calculate the optimal solution for the center point and Euler angles that minimizes the gradient vectors of the center point and Euler angles using the set gradient descent algorithm. In some embodiments, the Backtracking Line Search gradient descent algorithm is used to obtain the optimal solution. Backtracking Line Search is an adaptive step size selection strategy used in gradient descent to optimize convergence speed and stability. It avoids divergence or slow convergence caused by a fixed step size by dynamically adjusting the step size in each iteration. The core of Backtracking Line Search is to introduce the Armijo condition to ensure that the objective function value decreases sufficiently in each iteration. Using a step size of 0.01, a step size decay factor of 0.8, and the Armijo condition as the iteration stopping parameter, the optimal solution for the convergence center point and Euler angles is obtained within 20 iterations. The calculation process for the optimal solution of the center point and Euler angles is as follows:
[0053] Where η is the learning rate, k is the number of iterations, and c (k) Let c represent the center point of the k-th iteration. (k+1) This represents the center point of the (k+1)th iteration. Let θ be the gradient vector at the center point. (k) Let θ be the Euler angle for the k-th iteration. (k+1) For the (k+1)th iteration, Let be the gradient vector of the Euler angles. Update the parameters iteratively using gradient descent until convergence or the maximum number of iterations is reached, outputting the optimal solution for the center point and Euler angles. As a 6D pose.
[0054] The optimal solution for the center point includes: the optimal center point X-axis coordinate, the optimal center point Y-axis coordinate, and the optimal center point Z-axis coordinate; The optimal solutions for Euler angles include: the optimal first Euler angle, the optimal second Euler angle, and the optimal third Euler angle; The optimal center point X-axis coordinates, optimal center point Y-axis coordinates, optimal center point Z-axis coordinates, optimal first Euler angle, optimal second Euler angle, and optimal third Euler angle are used as the 6D pose of the target object.
[0055] The present invention also provides a fast 6D pose estimation device for a target object, used to implement the fast 6D pose estimation method for the target object, the device comprising: The binary mask acquisition module obtains the binary mask of the target object by inputting the RGB image and depth image of the target object into the segmentation model. The surface point cloud matrix acquisition module constructs the surface point cloud matrix of the target object based on the depth image, binary mask, and camera intrinsic parameter matrix by using the camera intrinsic parameter matrix to obtain the camera intrinsic parameter matrix. The fitting penalty function construction module constructs a fitting penalty function based on the surface point cloud matrix and the cubic bag used to construct the target object. The gradient vector acquisition module is used to calculate the gradient vectors of the center point and Euler angles of the target object based on the surface point cloud matrix, the cube bag, and the fitting penalty function. The 6D pose acquisition module is used to iteratively solve the gradient vectors of the center point and Euler angles using the gradient descent algorithm to obtain the optimal solution for the center point and Euler angles, and then use the optimal solution for the center point and Euler angles as the 6D pose of the target object.
[0056] This application provides an electronic device, including a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements a fast 6D pose estimation method for a target object according to any of the above schemes.
[0057] Specifically, the processor may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor may also include onboard memory for caching purposes. The processor may be a single processing unit or multiple processing units for performing different actions of the method flow according to embodiments of this application.
[0058] Memory can be any medium capable of containing, storing, transmitting, propagating, or transmitting instructions. For example, memory can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, instruments, or propagation media. Specific examples of memory include: magnetic storage devices such as magnetic tape or hard disk drives (HDDs); optical storage devices such as optical discs (CD-ROMs); and also random access memory (RAM) or flash memory; and / or wired / wireless communication links.
[0059] This application also provides a computer-readable medium storing a computer program that, when executed by a processor, implements a method for rapid 6D pose estimation of a target object according to any of the above-described schemes. This computer-readable medium may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into that device / apparatus / system. The aforementioned computer-readable medium carries one or more programs, which, when executed, implement the methods as described in the embodiments of this application.
[0060] According to embodiments of this application, a computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wired, optical fiber, radio frequency signals, etc., or any suitable combination thereof.
[0061] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application. Therefore, the scope of this application should not be limited to the above embodiments, but should be defined not only by the appended claims, but also by their equivalents. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for fast 6D pose estimation of a target object, characterized in that, The method comprises the steps of: obtaining an RGB image and a depth image of a target object, inputting the RGB image into a segmentation model to obtain a binary mask of the target object; obtaining a camera intrinsic matrix, and constructing a curved surface point cloud matrix of the target object according to the depth image, the binary mask and the camera intrinsic matrix; constructing a cubic package of the target object, and constructing a fitting penalty function according to the curved surface point cloud matrix and the cubic package; calculating a gradient vector of a center point and Euler angles of the target object according to the curved surface point cloud matrix, the cubic package and the fitting penalty function; obtaining an optimal solution of the center point and the Euler angles by iterative solving of the gradient vector of the center point and the Euler angles through a gradient descent algorithm, and taking the optimal solution of the center point and the Euler angles as the 6D pose of the target object.
2. The method of Claim 1, wherein, The curved surface point cloud matrix is composed of coordinates of each pixel point of the target object in a three-dimensional camera coordinate system, wherein an X-axis of the three-dimensional camera coordinate system is a vertical direction, a Y-axis is a horizontal direction, and a Z-axis is an optical axis direction of an RGBD camera.
3. The method of Claim 1, wherein, The fitting penalty function is constructed according to the curved surface point cloud matrix and the cubic package, and specifically comprises: S11: the cubic package of the target object is a smallest cube wrapping the target object, and the height H, the width W and the depth D of the cubic package are obtained through a measuring tool; S12: Extract the coordinate [X, Y, Z] of the i-th pixel point P in the curved surface point cloud matrix i . i i i S13: Based on the height H, width W, depth D, and pixel P of the cube bag i coordinates [X i ,Y i Z i ], calculate to obtain pixel point P i The sum of the distances M to the six faces of the cube bag i ; S14: repeat steps S12-S14 until the sum of distances of N pixel points is obtained, and the sum of distances of N pixel points is obtained As the fitting penalty function Loss, N represents the total number of pixel points in the surface point cloud matrix.
4. The method of Claim 1, wherein, The independent variables of the center point of the target object include a center point X-axis coordinate x c , a center point Y-axis coordinate y c , and a center point Z-axis coordinate z c The independent variables of the Euler angles of the target object include a first Euler angle ω0, a second Euler angle ω1, and a third Euler angle ω2.
5. The method of Claim 4, wherein, The calculation process of the gradient vector of the center point of the target object specifically comprises: the height H, the width W and the depth D of the cubic package are obtained through a measuring tool, the coordinates of each pixel point in the curved surface point cloud matrix are obtained, and a unit vector n1 towards the X-axis, a unit vector n2 towards the Y-axis and a unit vector n3 towards the Z-axis in the three-dimensional camera coordinate system of the target object are obtained; According to the high H, the wide W, the deep D of the cubic package, the coordinates of each pixel point and the unit vector n1, the partial derivative of the fitting penalty function Loss to the x-axis coordinate x of the center point is calculated c ; According to the high H, the wide W, the deep D of the cubic package, the coordinates of each pixel point and the unit vector n2, the partial derivative of the fitting penalty function Loss with respect to the Y-axis coordinate y of the center point is calculated c ; According to the high H, the wide W, the deep D of the cubic package, the coordinates of each pixel point and the unit vector n3, the partial derivative of the fitting penalty function Loss with respect to the central point Z-axis coordinate z is calculated c ; partial derivatives , partial derivatives , and partial derivatives are combined into a row vector ( , , ) as a gradient vector of the center point.
6. The method of Claim 4, wherein, The calculation process of the gradient vector of the Euler angles of the target object specifically comprises: the height H, the width W and the depth D of the cubic package are obtained through a measuring tool, the coordinates of each pixel point in the curved surface point cloud matrix are obtained, and a unit vector n1 towards the X-axis, a unit vector n2 towards the Y-axis and a unit vector n3 towards the Z-axis in the three-dimensional camera coordinate system of the target object are obtained; According to the high H, the wide W, the deep D of the cubic package, the coordinates of each pixel point and the unit vector n1, the partial derivative of the fitting penalty function Loss with respect to the first Euler angle ω0 is calculated ; According to the high H, the wide W, the deep D of the cubic package, the coordinates of each pixel point and the unit vector n2, a partial derivative of the fitting penalty function Loss with respect to the second Euler angle ω1 is calculated and obtained ; According to the high H, the wide W, the deep D of the cubic package, the coordinates of each pixel point and the unit vector n3, the partial derivative of the fitting penalty function Loss with respect to the third Euler angle ω2 is calculated ; partial derivatives , partial derivatives , and partial derivatives are combined into a row vector ( , , ) as the gradient vector of the Euler angles.
7. The method of claim 4, wherein, The optimal solution of the center point and the Euler angles is obtained by iterative solving of the gradient vector of the center point and the Euler angles through the gradient descent algorithm, and the optimal solution of the center point and the Euler angles is taken as the 6D pose of the target object, and specifically comprises: a step size, a step size attenuation factor and an iteration stop parameter of the gradient descent algorithm are set, and the optimal solution of the center point and the Euler angles that minimizes the gradient vector of the center point and the Euler angles is obtained through the set gradient descent algorithm; the optimal solution of the center point comprises: an optimal center point X-axis coordinate, an optimal center point Y-axis coordinate and an optimal center point Z-axis coordinate; the optimal solution of the Euler angles comprises: an optimal first Euler angle, an optimal second Euler angle and an optimal third Euler angle; the optimal center point X-axis coordinate, the optimal center point Y-axis coordinate, the optimal center point Z-axis coordinate, the optimal first Euler angle, the optimal second Euler angle and the optimal third Euler angle are taken as the 6D pose of the target object.
8. A device for fast 6D pose estimation of a target object, configured to implement the method for fast 6D pose estimation of a target object according to any one of claims 1 to 7, characterized in that, The device comprises: a binary mask obtaining module, which is used for obtaining an RGB image and a depth image of a target object, and inputting the RGB image into a segmentation model to obtain a binary mask of the target object; a curved surface point cloud matrix obtaining module, which is used for obtaining a camera intrinsic matrix, and constructing a curved surface point cloud matrix of the target object according to the depth image, the binary mask and the camera intrinsic matrix; The fitting penalty function construction module constructs a fitting penalty function according to the curved surface point cloud matrix and the cuboid package through a cuboid package used for constructing the target object; The gradient vector acquisition module is configured to calculate and obtain the gradient vector of the center point and the Euler angle of the target object according to the curved surface point cloud matrix, the cuboid package and the fitting penalty function; The 6D pose acquisition module is configured to obtain the optimal solution of the center point and the Euler angle by iteratively solving the gradient vector of the center point and the Euler angle through a gradient descent algorithm, and take the optimal solution of the center point and the Euler angle as the 6D pose of the target object.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the 6D pose fast estimation method of the target object according to any one of claims 1 to 7 when executing the program.
10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the 6D pose fast estimation method of the target object according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Monocular 6D attitude estimation method and device based on deep convolutional neural network
CN112767486A
6D pose estimation method of target object
CN117011380A
Class-level 6D attitude estimation method and system, medium and program product
CN120182365A
Robot grabbing posture generation method and related device
CN120307282A
Method for estimating poses of multiple targets, robot control method and system, and product
WO2025073306A2
Cited By
Robot arm 6D pose estimation method and system based on fuzzy visual angle prior
CN121391997A
A robot arm 6D pose estimation method and system based on a fuzzy perspective prior
CN121391997B
Humanoid robot water pouring method, device and equipment based on visual servo and medium
CN121625158A