6D object pose estimation method and system based on multi-scale deformable attention, terminal and storage medium
Through the 6D object position estimation method of multi-scale deformable attention, the accuracy of object position estimation in occlusion and multi-scale scenarios is solved, and the stable and accurate estimation of object posture in complex environments is achieved.
Patent Information
- Application Number
- CN202510311364.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
AI Technical Summary
In the process of object recognition and pose estimation, the prior art has limitations when facing occlusion and multi-scale scenarios, resulting in inaccurate estimation results of object position estimation.
Using a 6D object position estimation method based on multi-scale deformable attention, the target RGB image is obtained for feature extraction, semantic annotation and translation estimation, combined with the multi-scale feature pyramid and attention mechanism, the fusion of semantic context information and spatial detail features is achieved, and 3D translation and rotation regression is performed to obtain accurate 6D pose estimation results.
In complex scenes, the 6D posture of an object can be sturdily estimated, and even accurately estimated even when the object is blocked, improving the stability and accuracy of the object posture estimation.
Smart Images

Figure CN120182380A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a 6D object pose estimation method, system, terminal and computer-readable storage medium based on multi-scale deformable attention. Background Art
[0002] In the rapid development process of robotics, endowing robots with the ability to accurately perceive and understand the three-dimensional environment has become one of the key challenges for achieving advanced artificial intelligence. Especially in the field of object recognition and pose estimation, the ability to accurately locate and understand objects in three-dimensional space is crucial for the effective interaction between robots and the real world.
[0003] However, in the process of object recognition and pose estimation in the prior art, there are limitations when facing occlusion and multi-scale scenarios, resulting in inaccurate object pose estimation results.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide a 6D object pose estimation method, system, terminal and computer-readable storage medium based on multi-scale deformable attention, aiming to solve the problem that in the process of object recognition and pose estimation in the prior art, there are limitations when facing occlusion and multi-scale scenarios, resulting in inaccurate object pose estimation results.
[0006] To achieve the above purpose, the present invention provides a 6D object pose estimation method based on multi-scale deformable attention, and the 6D object pose estimation method based on multi-scale deformable attention includes the following steps: Obtain a target RGB image, and perform feature extraction processing on the target RGB image to obtain a plurality of feature images; Perform semantic annotation processing and translation estimation processing on the plurality of feature images to obtain a plurality of masks of semantic labels and a plurality of vectors of implicit center point positions, and obtain a 3D translation vector according to the plurality of masks of semantic labels and the plurality of vectors of implicit center point positions; Perform rotation regression processing on the plurality of feature images to obtain a 3D rotation estimation value, and obtain a 6D pose estimation result according to the 3D rotation estimation value and the 3D translation vector.
[0007] Optionally, in the 6D object pose estimation method based on multi-scale deformable attention, wherein the step of obtaining a target RGB image and performing feature extraction processing on the target RGB image to obtain a plurality of feature images specifically includes: Obtain a target RGB image and input the target RGB image into a 6D object pose estimation model; Perform feature encoding processing and spatial detail feature extraction processing on the target RGB image through a multi-scale feature pyramid in the 6D object pose estimation model to obtain semantic context information and spatial detail features; Perform feature fusion processing on the semantic context information and the spatial detail features to obtain multiple feature images.
[0008] Optionally, in the 6D object pose estimation method based on multi-scale deformable attention, the semantic annotation processing includes dimensionality reduction processing and attention weight adjustment processing; the translation estimation processing includes upsampling processing and feature enhancement processing; Performing semantic annotation processing and translation estimation processing on the multiple feature images to obtain masks of multiple semantic labels and vectors of multiple implicit center point positions, and obtaining a 3D translation vector according to the masks of the multiple semantic labels and the vectors of the multiple implicit center point positions, specifically including: Perform dimensionality reduction processing and attention weight adjustment processing on the multiple feature images to obtain masks of multiple semantic labels; Perform upsampling processing and feature enhancement processing on the multiple feature images to obtain vectors of multiple implicit center point positions; Perform loss calculation on the masks of the multiple semantic labels and the vectors of the multiple implicit center point positions to obtain a 3D translation vector.
[0009] Optionally, in the 6D object pose estimation method based on multi-scale deformable attention, the performing dimensionality reduction processing and attention weight adjustment processing on the multiple feature images to obtain masks of multiple semantic labels specifically includes: Determine multiple projection layers in the 6D object pose estimation model, and select corresponding target projection layers from the multiple projection layers according to the scale of each feature image; Input each feature image into the corresponding target projection layer, and perform dimensionality reduction processing on each corresponding feature image through the target projection layer to obtain multiple dimensionality-reduced feature maps; Perform attention weight adjustment processing and feature conversion processing on the multiple dimensionality-reduced feature maps to obtain masks of multiple semantic labels.
[0010] Optionally, in the 6D object pose estimation method based on multi-scale deformable attention, the performing attention weight adjustment processing and feature conversion processing on the multiple dimensionality-reduced feature maps to obtain masks of multiple semantic labels specifically includes: Input multiple of the dimensionality-reduced feature maps into the multi-scale deformable attention module in the 6D object pose estimation model, and adjust the attention weights of the multiple dimensionality-reduced feature maps through the multi-scale deformable attention module to obtain multiple adjusted feature maps; Input multiple of the adjusted feature maps into the segmentation head module in the 6D object pose estimation model, and perform feature transformation processing on the multiple adjusted feature maps through the segmentation head module to obtain masks of multiple semantic labels.
[0011] Optionally, in the 6D object pose estimation method based on multi-scale deformable attention, the upsampling processing and feature enhancement processing of the multiple feature images to obtain vectors of multiple implicit center point positions specifically include: Input multiple of the feature images into the upsampling module in the 6D object pose estimation model, and perform upsampling processing on the multiple feature images through the upsampling module to obtain multiple high-dimensional feature images; Input multiple of the high-dimensional feature images into the first convolutional block attention module in the 6D object pose estimation model, and perform feature enhancement processing on the multiple high-dimensional feature images through the first convolutional block attention module to obtain vectors of multiple implicit center point positions.
[0012] Optionally, in the 6D object pose estimation method based on multi-scale deformable attention, the rotation regression processing of the multiple feature images to obtain 3D rotation estimation values, and obtaining the 6D pose estimation result according to the 3D rotation estimation values and the 3D translation vectors specifically includes: Input multiple of the feature images into the pooling layer in the 6D object pose estimation model, and perform pooling processing on the multiple feature images through the pooling layer to obtain multiple feature descriptors; Input multiple of the feature descriptors into the second convolutional block attention module in the 6D object pose estimation model, and perform feature enhancement processing on the feature descriptors through the second convolutional block attention module to obtain multiple enhanced feature descriptors; Input multiple of the enhanced feature descriptors into the fully connected layer in the 6D object pose estimation model, and perform abstract feature extraction processing and dimensional vector output processing on the multiple enhanced feature descriptors through the fully connected layer to obtain 3D rotation estimation values; Calculate the loss of the 3D rotation estimation values to obtain target quaternions, and obtain the 6D pose estimation result according to the target quaternions and the 3D translation vectors.
[0013] In addition, to achieve the above object, the present invention further provides a 6D object pose estimation system based on multi-scale deformable attention, wherein the 6D object pose estimation system based on multi-scale deformable attention includes: A feature extraction processing module, configured to obtain a target RGB image and perform feature extraction processing on the target RGB image to obtain a plurality of feature images; A 3D translation vector acquisition module, configured to perform semantic annotation processing and translation estimation processing on a plurality of the feature images to obtain masks of a plurality of semantic labels and vectors of a plurality of implicit center point positions, and obtain a 3D translation vector according to the masks of the plurality of semantic labels and the vectors of the plurality of implicit center point positions; A 6D pose estimation result generation module, configured to perform rotation regression processing on a plurality of the feature images to obtain a 3D rotation estimation value, and obtain a 6D pose estimation result according to the 3D rotation estimation value and the 3D translation vector.
[0014] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and a 6D object pose estimation program based on multi-scale deformable attention stored on the memory and executable on the processor, and when the 6D object pose estimation program based on multi-scale deformable attention is executed by the processor, the steps of the 6D object pose estimation method based on multi-scale deformable attention as described above are implemented.
[0015] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a 6D object pose estimation program based on multi-scale deformable attention, and when the 6D object pose estimation program based on multi-scale deformable attention is executed by a processor, the steps of the 6D object pose estimation method based on multi-scale deformable attention as described above are implemented.
[0016] In the present invention, a target RGB image is obtained, and feature extraction processing is performed on the target RGB image to obtain a plurality of feature images; semantic annotation processing and translation estimation processing are performed on the plurality of feature images to obtain masks of a plurality of semantic labels and vectors of a plurality of implicit center point positions, and a 3D translation vector is obtained according to the masks of the plurality of semantic labels and the vectors of the plurality of implicit center point positions; rotation regression processing is performed on the plurality of feature images to obtain a 3D rotation estimation value, and a 6D pose estimation result is obtained according to the 3D rotation estimation value and the 3D translation vector. By performing semantic annotation processing, translation estimation processing, and rotation regression processing on the feature images obtained after feature processing of the target RGB image, the present invention can effectively solve the problems of difficult pose estimation in the case of object occlusion and multi-scale scenes. Even if the target object is occluded by other objects, pixels can stably project the center, enabling accurate estimation of the 6D pose of an object in a complex scene. At the same time, the stability of object pose estimation is also ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a preferred embodiment of the 6D object pose estimation method based on multi-scale deformable attention of the present invention; Figure 2 is a schematic diagram of the MSDAPose network framework of a preferred embodiment of the 6D object pose estimation method based on multi-scale deformable attention of the present invention; Figure 3 is a schematic diagram of the overall 6D pose estimation framework of MSDAPose of a preferred embodiment of the 6D object pose estimation method based on multi-scale deformable attention of the present invention; Figure 4 is a schematic diagram of the process of projecting the three-dimensional center point of an object onto a two-dimensional image plane by a pinhole camera in a preferred embodiment of the 6D object pose estimation method based on multi-scale deformable attention of the present invention; Figure 5 is a schematic diagram of the process of projecting the instance center in a preferred embodiment of the 6D object pose estimation method based on multi-scale deformable attention of the present invention; Figure 6 is a structural diagram of a preferred embodiment of the 6D object pose estimation system based on multi-scale deformable attention of the present invention; Figure 7 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer and more explicit, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0019] In the rapid development of robotics, endowing robots with the ability to accurately perceive and understand the three-dimensional environment has become one of the key challenges in achieving advanced artificial intelligence. Especially in the field of object recognition and pose estimation, the ability to precisely locate and understand objects in three-dimensional space is crucial for robots to interact effectively with the real world. This ability plays a central role in many application scenarios, ranging from industrial automation, service robots to smart home assistants, and even extending to complex scientific research environments. For example, in robotic manipulation tasks, accurately grasping the three-dimensional position (i.e., coordinates) and orientation (usually represented by Euler angles or axis angles) of an object enables precise grasping, placement, or manipulation of the target object. In human-robot interaction scenarios, such as learning skills from human demonstrations, understanding the three-dimensional pose of an object helps the robot interpret action intentions, imitate behavior patterns, or predict the consequences of actions. Therefore, the real-time perception of the three-dimensional pose of objects is regarded as a core element in achieving advanced robotic intelligence.
[0020] However, this task faces multiple challenges. The diversity of object shapes, rich three-dimensional shapes, sizes, and material properties in the real world directly affect the way they are presented in two-dimensional images (such as the visual information captured by cameras). Environmental factors such as changes in lighting conditions, cluttered backgrounds within the scene, and mutual occlusion between objects can all cause significant changes in the appearance of objects in the image, increasing the difficulty of recognition and pose estimation. For example, strong light or shadows may obscure the detailed features of an object; background clutter may confuse the object boundaries and interfere with contour recognition; while occlusion phenomena may result in partial or complete invisibility of the object, rendering traditional recognition methods based on complete visual features ineffective.
[0021] In the prior art, 6D object pose estimation mainly relies on two strategies: template matching-based methods and feature point matching-based methods. Among them, the template matching method usually relies on an accurate representation of the object model. By comparing the model with candidate regions in the image, the best match is found to determine the object pose. This method is relatively effective when dealing with objects with rich textures and easy to distinguish. However, when faced with objects without textures, similar textures, or large lighting variations, its recognition performance will significantly decline. In addition, the template matching method has a weak ability to handle occlusion situations because the presence of occluders may cause the matching degree between the template and the image to drop sharply. On the contrary, the feature point matching-based method focuses on extracting local features from the image, such as (SIFT, Scale Invariant Feature Transform, an algorithm widely used in the fields of image processing and computer vision), (SURF, Speeded Up Robust Features, an algorithm used for feature detection and description in the field of computer vision), etc., and matches these features with the feature points pre-recorded on the three-dimensional model of the object, thereby establishing the correspondence between the pixels in the image and the corresponding points on the model. Based on these two-dimensional to three-dimensional correspondences, mathematical methods (such as the EPnP algorithm: an accurate positioning algorithm widely used in the fields of computer vision and robotics, the ICP algorithm: a classic algorithm for point cloud registration, widely used in fields such as 3D scanning, robot navigation, SLAM (Simultaneous Localization and Mapping), etc.) can be used to solve the 6D pose of the object. The advantage of this type of method is its certain robustness to occlusion because even if the object is partially occluded, it is still possible to find a sufficient number of visible feature points for matching. However, this strategy requires the object surface to have rich textures in order to extract and match sufficient feature points. For objects without textures or with sparse textures, the success rate of feature point extraction and matching will be greatly reduced, thus limiting the application scope of this method. To address the recognition problem of textureless objects, recent research has started to turn to machine learning techniques, especially deep learning, to learn descriptors that describe the surface features of objects, and these descriptors can still distinguish objects in the absence of obvious textures. Some other research attempts to directly regress from image pixels to the three-dimensional coordinates of the object to establish a 2D-3D correspondence, and then estimate the 6D pose of the object. Although this method avoids the dependence on textures, it will encounter difficulties when dealing with symmetric objects because the multi-solution nature of the three-dimensional coordinate mapping caused by symmetry makes it difficult for the direct regression method to uniquely determine the exact pose of the object.
[0022] To solve the above problems, the present invention proposes a novel neural network architecture - MSDAPose (A 6D Object Pose Estimation Framework Based on Multi-Scale, that is, a 6D object pose estimation framework based on multi-scale deformable attention). The present invention utilizes deep learning to overcome their respective deficiencies. In recent years, 6D pose estimation technology has received extensive attention, and relevant datasets and solutions have emerged continuously. Against this background, MSDAPose demonstrates excellent generality and adaptability. The present invention will first review the related work and elaborate in detail on the advantages and limitations of existing methods. Subsequently, the network architecture design and working principle of MSDAPose of the present invention will be introduced. Finally, a series of experiments will be conducted to evaluate the performance of MSDAPose.
[0023] The 6D object pose estimation method based on multi-scale deformable attention according to a preferred embodiment of the present invention, as Figure 1 shown, the 6D object pose estimation method based on multi-scale deformable attention includes the following steps: Step S10: Obtain a target RGB image, and perform feature extraction processing on the target RGB image to obtain a plurality of feature images.
[0024] In the prior art, the methods for 6D object pose estimation can be mainly divided into two categories: template-based methods and feature-based methods. These two types of methods each have their own advantages and limitations and exhibit different application values in different scenarios.
[0025] Among them, the template-based methods mainly achieve pose estimation by constructing and matching templates. With the development of deep learning, some studies have begun to combine deep learning techniques with template matching. By using an autoencoder to learn the orientation representation of an object, the performance of template matching has been effectively improved. However, the template-based methods often perform poorly when dealing with occlusion scenarios. In complex scenarios, traditional template matching methods are difficult to handle the situation where an object is partially occluded. To solve this problem, some researchers have begun to explore the possibility of combining template methods with other techniques.
[0026] The feature-based methods have made significant progress in recent years due to the application of deep learning techniques. Early work attempted to directly regress the 6D pose of an object through a deep neural network, while SSD-6D extended the single-stage object detector SSD to the 6D pose estimation task. However, the direct regression methods often have difficulty obtaining high-precision pose estimation results.
[0027] To improve the estimation accuracy, many researchers have turned to methods based on key-point prediction, such as locating key points through a pixel-level voting mechanism or recovering the pose by predicting the eight corner points of the object bounding box. These methods usually need to be combined with geometric optimization algorithms such as PnP to solve the final pose. In addition, directly integrating PnP optimization into the deep learning framework enables end-to-end training, and this idea of combining geometric optimization has significantly improved the accuracy of pose estimation.
[0028] Recent research trends have begun to explore the complementarity of different representations. For example, DPOD and CDPN use dense pixel-level 2D-3D correspondences to improve the robustness of pose estimation. Pix2Pose establishes dense correspondences by predicting pixel-level 3D coordinates. These methods attempt to address the limitations of traditional key-point methods in dealing with occlusion and illumination changes. In practical application scenarios, the real-time performance of the algorithm is crucial for tasks such as robotic grasping.
[0029] As Figure 2 shown, the present invention proposes a network framework for 6D object pose estimation based on multi-scale features: MSDAPose, which can perform three tasks: semantic annotation, 3D translation estimation, and 3D rotation regression.
[0030] Given an input image, the task of 6D object pose estimation is to calculate the rigid transformation from the object coordinate system to the camera coordinate system. Assuming that the three-dimensional model of the object is known and the object coordinate system has been defined in the three-dimensional space of the model, this rigid transformation consists of an SE transformation, including a three-dimensional rotation matrix and a 3D translation vector . Among them, represents the rotation angles of the object coordinate system around the axis, axis, and axis, and represents the coordinate values of the origin of the coordinate system in the camera coordinate system . During the imaging process, determines the position and scale of the object in the image, while affects the apparent form of the object in the image according to the three-dimensional shape and surface texture of the object. Given that these two parameters have different visual characteristics, the present invention proposes a framework for joint learning of detection and pose estimation based on multi-scale deformable attention (i.e., MSDAPose).
[0031] Specifically, obtain the target RGB image and input the target RGB image into a 6D object pose estimation model; perform feature encoding processing and spatial detail feature extraction processing on the target RGB image through the multi-scale feature pyramid in the 6D object pose estimation model to obtain semantic context information and spatial detail features; perform feature fusion processing on the semantic context information and the spatial detail features to obtain multiple feature images.
[0032] As Figure 3 shown, MSDAPose in the present invention is an end-to-end 6D object pose estimation framework that realizes joint optimization of object detection, instance segmentation, and pose estimation through a multi-scale feature sharing and task collaboration mechanism. This framework constructs a multi-scale feature pyramid with ResNet as the backbone, where deep features are processed through encoding to obtain semantic context information, and shallow features retain spatial detail features, providing multi-granularity representations for downstream tasks through cross-level feature fusion. The core of solving downstream tasks includes three functionally coupled branches: instance segmentation, translation estimation, and rotation regression. During the training phase, a dynamic weight strategy is adopted to adaptively balance multi-task learning, and a stable warm-up phase is used to improve convergence stability. During inference, pose calculation is achieved through cascaded annotation (i.e., semantic annotation), 3D positioning, and the quaternion after regression.
[0033] As Figure 3 shown, for semantic annotation, the present invention designs a semantic annotation branch ( Figure 3 which is divided into three branches, and the semantic annotation branch belongs to the branch where multiple projection layers and the MSDA module are located) to achieve efficient multi-scale semantic segmentation. This branch can process feature maps of different resolutions and adopts the multi-scale deformable attention mechanism (MSDA, Multi-Scale Deformable Attention) in DETR to enhance the feature representation ability. As Figure 3 shown: This branch first performs dimensionality reduction processing on the input feature map (i.e., the multiple feature images obtained after the target RGB image is processed by ResNet) through a series of projection layers (each scale uses a 2D convolutional layer for dimensionality reduction to ensure that the channels of all scales are consistent). These projection layers correspond to feature maps of different scales (allowing the model to extract features at different scales to better capture detailed information in the image). Subsequently, a multi-scale deformable attention module (i.e., the Figure 3 MSDA module in
[0034] Next, the segmentation head consists of a series of convolutional layers for converting the feature map (through Softmax2D) into a class probability map. After enhancing the feature representation through convolutional operations and activation functions, a convolutional layer is used to reduce the number of channels to the number of classes plus one (including the background class), and bilinear interpolation upsampling operation is used to restore the output to the original image size. Finally, the Softmax function is applied to generate the probability distribution of each pixel.
[0035] In addition, the present invention also designs an auxiliary function to generate the bounding boxes (BBX, Bounding Boxes) of each instance according to the segmentation labels, achieving an efficient and accurate semantic annotation task. At the same time, BBX also provides a segmentation basis for subsequent translation estimation and rotation regression, enhancing the model's understanding ability of complex scenes.
[0036] Step S20: Perform semantic annotation processing and translation estimation processing on multiple said feature images to obtain masks of multiple semantic labels and vectors of multiple implicit center point positions, and obtain 3D translation vectors according to the masks of multiple said semantic labels and the vectors of multiple said implicit center point positions. The semantic annotation processing includes dimensionality reduction processing and attention weight adjustment processing; the translation estimation processing includes upsampling processing and feature enhancement processing.
[0037] Specifically, determine multiple projection layers in the 6D object pose estimation model, and select corresponding target projection layers from multiple said projection layers according to the scale of each said feature image; input each said feature image into the corresponding target projection layer, and perform dimensionality reduction processing on each corresponding feature image through the target projection layer to obtain multiple dimensionality-reduced feature maps; input multiple said dimensionality-reduced feature maps into the multi-scale deformable attention module in the 6D object pose estimation model, and perform attention weight adjustment processing on multiple said dimensionality-reduced feature maps through the multi-scale deformable attention module to obtain multiple adjusted feature maps; input multiple said adjusted feature maps into the segmentation head module in the 6D object pose estimation model, and perform feature conversion processing on multiple said adjusted feature maps through the segmentation head module to obtain masks of multiple semantic labels. Input multiple said feature images into the upsampling module in the 6D object pose estimation model, and perform upsampling processing on multiple said feature images through the upsampling module to obtain multiple high-dimensional feature images; input multiple said high-dimensional feature images into the first convolutional block attention module in the 6D object pose estimation model, and perform feature enhancement processing on multiple said high-dimensional feature images through the first convolutional block attention module to obtain vectors of multiple implicit center point positions; calculate the loss of multiple said masks of semantic labels and multiple said vectors of implicit center point positions to obtain 3D translation vectors.
[0038] AsFigure 4 As shown in the figure, for the translational estimation process, the present invention adopts a method for estimating the translation of an object in a three-dimensional space (i.e., the displacement of the object relative to the camera coordinate system). This method aims to achieve accurate 3D translational estimation by locating the center of the object in the image and predicting the distance between this center point and the camera.
[0039] As Figure 4 shown, the 3D translation is the coordinate of the origin of the object in the camera coordinate system. An intuitive method for estimating is to directly regress the image features to .
[0040] In the prior art, the displacement of an object on the and axes is calculated using the pinhole camera model, and the formula is as follows: and ; specifically, considering the center point of an object in its two-dimensional projection, , is the abscissa of the center point, is the ordinate of the center point, is the matrix transpose. If the network can accurately identify this point and estimate the corresponding depth information , this involves using the camera focal length and and the optical center coordinates to convert the spatial coordinates from 2D to 3D. If the origin of the object is the centroid of the object, that is, the 2D center of the object is obtained, then and can be restored.
[0041] However, this method is not universal because the object may appear at any position in the image. To overcome the detection failure caused by occlusion that may occur when directly regressing the center position of the object, the present invention combines the strategy of PoseCNN and regresses each pixel to the direction vector pointing to the center of the object instead of the direct distance vector. As Figure 5 shown, this means that for any given pixel , it will regress to a unit-length direction vector , and this vector represents the direction from the pixel to the center of the object. Among them, and represent the vectors of each pixel point pointing to the center point.
[0042] As Figure 3As shown in the figure, in terms of the network architecture, this branch increases the channel dimension for different object categories to adapt to more complex output requirements and also restores to the original image size. In particular, before mapping the high-dimensional features to a unified 128-dimensional space, a Convolutional Block Attention Module (CBAM) is introduced to enhance the feature representation ability. The CBAM module helps the network focus on important feature regions, thereby improving the accuracy of the center point localization to ensure sufficient expressive ability to accurately regress the task, and finally obtaining the three variables that each class of instance objects needs to regress. , where the regression process can be expressed as: .
[0043] Step S30: Perform rotation regression processing on the multiple feature images to obtain a 3D rotation estimate value, and obtain a 6D pose estimate result based on the 3D rotation estimate value and the 3D translation vector.
[0044] Specifically, input the multiple feature images into the pooling layer in the 6D object pose estimation model, perform pooling processing on the multiple feature images through the pooling layer to obtain multiple feature descriptors; input the multiple feature descriptors into the second convolutional block attention module in the 6D object pose estimation model, and perform feature enhancement processing on the feature descriptors through the second convolutional block attention module to obtain multiple enhanced feature descriptors.
[0045] As Figure 3 shown, for the rotation regression processing, in the 3D rotation regression framework proposed by the present invention, first, multi-scale feature maps and a Region of Interest (RoI) pooling layer are used to crop and pool the visual features generated in the first stage of the network. Specifically, for each RoI, the present invention uses convolutional kernels of different scales to process the corresponding feature maps and converts them into fixed-size feature representations through the RoI pooling layer. These pooled feature maps are then superimposed together to form a comprehensive feature descriptor, which not only contains the local detail information of the object but also retains its global structural features. On this basis, the CBAM module is introduced to further enhance the model's learning ability for key features.
[0046] Input the multiple enhanced feature descriptors into the fully connected layer in the 6D object pose estimation model, perform abstract feature extraction processing and dimension vector output processing on the multiple enhanced feature descriptors through the fully connected layer to obtain a 3D rotation estimate value; calculate the loss of the 3D rotation estimate value to obtain a target quaternion, and obtain a 6D pose estimate result based on the target quaternion and the 3D translation vector.
[0047] As Figure 3 shown, in order to regress accurate 3D rotation parameters from these comprehensive features, the present invention designs a sub-network containing three fully connected (FC) layers. The first two FC layers have 4096 neurons and are used to extract high-level abstract features, while the last FC layer outputs a vector with a dimension of , where represents the number of object categories. For each category, the final FC layer can output a 3D rotation estimate represented by quaternion, thus realizing end-to-end 3D pose estimation.
[0048] To train quaternion regression, the present invention introduces PoseLoss and ShapeMatch-Loss in PoseCNN, where one is specifically designed to handle symmetric objects. The first loss function PoseLoss (i.e., PLoss) is used to measure the mean squared distance between the model points under the estimated pose and the corresponding points under the true pose. Its definition is as follows: ; where represents the set of 3D model points, is the number of points, is the randomly sampled point in the 3D model. and represent the rotation matrices calculated from the estimated quaternion and the true quaternion respectively. However, PLoss has limitations in dealing with symmetric objects. For symmetric objects, there may be multiple correct 3D rotation directions, and PLoss will unnecessarily penalize the network for predicting other possible correct directions. To solve this problem, in this section, ShapeMatch-Loss (i.e., SLoss) is introduced simultaneously. This is a loss function that does not require specifying symmetry, and its definition is as follows: ; where are the three-dimensional points sampled in the model corresponding to the predicted pose, are the three-dimensional space points selected in the corresponding model under the true pose. Formally, SLoss is similar to ICP and measures the offset between each point in the estimated model direction and the nearest point on the true model. When the two 3D models match, SLoss reaches the minimum. By considering all possible point correspondences, the symmetric properties of the object are automatically processed, thus providing a more robust training objective.
[0049] In the present invention, the MSDAPose is set up. By separating the regression of translation and rotation, it effectively solves the problems of severe occlusion or too small instances. Even if an object is occluded by other objects, pixels can stably deliver the center. Experimental results show that accurate estimation of the 6D pose of an object can be achieved in complex scenes only using visual data, which paves the way for using cameras with resolutions and fields of view far exceeding those of current depth camera systems.
[0050] In summary, the present invention provides a neural network architecture MSDAPose for solving the 6D object pose estimation problem. The core feature of this architecture is the fusion of multi-scale feature hierarchies, which realizes bottom-up pixel-level object annotation and top-down pose regression. This design enables MSDAPose to achieve robust multi-instance object detection and 6D pose estimation in complex scenes.
[0051] Experimental verification: The present invention has conducted a large number of experiments on the PROPS dataset and the self-made dataset, verifying the high robustness of MSDAPose to symmetric objects in occlusion scenarios and its ability to achieve accurate pose estimation only using color image input. It has achieved the current optimal performance on the challenging PROPS dataset, and the experimental results are detailed in Tables 1 and 2. Among them, MSDAPose is superior to the baseline PoseCNN on multiple test instances. Thanks to the segmentation of the multi-scale attention mechanism, the model can accurately identify symmetric small objects or occluded objects such as "tuna_fish_can" and "large_marker".
[0052] Although object-centered datasets can provide accurate annotations for object pose and segmentation tasks, since most of these annotations rely on manual completion, the scale of such datasets is often greatly limited. Taking the famous LineMOD dataset as an example, it provides approximately 1000 manually annotated images for each of the 15 objects. Although this dataset is extremely important for evaluating model-based pose estimation algorithms, its data volume is much smaller than the standard datasets required for training modern deep neural networks, with a gap of several orders of magnitude. To address this challenge, the present invention uses synthetic images for data augmentation, thus greatly expanding the data volume, and uses algorithms to achieve high-precision automatic annotation of masks, which not only greatly reduces the dataset cost but also meets the actual engineering implementation requirements.
[0053] In addition, the PROPS dataset was created specifically to support the development of deep learning models in the field of robot perception. It is constructed based on the data collected by the ProgressLabeller annotation tool. This dataset mainly focuses on desktop scenarios that simulate the operating environment of domestic service robots and uses the model objects in YCB-Vedio. The PROPS dataset can be used for semantic segmentation, object detection, and 6D pose estimation. The dataset for the 6D pose estimation task contains a training set and a validation set. Among them, each class contains 500 annotated 640x480 RGB images, and aligned depth images and segmentation masks are provided for more accurate research.
[0054] Table 1: Results of Quantitative Evaluation and Comparison Based on the PROPS Dataset
[0055] Table 2: Results of Quantitative Evaluation and Comparison Based on the Synthetic Dataset
[0056] In addition, the present invention also sets up an ablation experiment: to verify the roles of the multi-scale feature fusion and the CBAM module proposed in the present invention, the present invention trains networks without using multi-scale features as input and without adding the CBAM module respectively. Due to the limitation of computing resources, the present invention only conducts experiments on the dog model in the self-made dataset. The data results of the ablation experiment are shown in Table 3 below.
[0057] Table 3: Ablation Experiment to Verify the Influence Degree of Multi-scale Feature Fusion and CBAM Module on the Rotation or Translation Branch
[0058] As shown in Table 3, multi-scale feature fusion is a key element of the present invention. Since after being processed by the MSDA module and only relying on single-scale feature maps for training, it is difficult for the model to capture the details of rotation. In addition, in the enhancement experiments using the CBAM module only for translation or rotation, the results without enhancing the rotation branch are very similar. This fully shows that the CBAM module has a greater contribution effect on 3D pose regression and also plays a certain role in 3D translation estimation.
[0059] Experiments show that under similar training data conditions, as a method for directly estimating 6D poses, the performance of MSDAPose can be comparable to that of traditional 2D detection PnP methods. More importantly, the method of the present invention has a significant computational efficiency advantage: its processing time hardly increases significantly with the increase in the number of objects in the scene. This characteristic makes MSDAPose applicable to practical application scenarios, such as robot grasping and autonomous driving, which require processing multiple targets simultaneously and have strict real-time requirements. It can be seen that MSDAPose not only demonstrates excellent performance and versatility but also has good practical value.
[0060] In summary, this paper proposes an end-to-end framework - MSDAPose, which uses a multi-scale deformable attention mechanism to achieve instance center localization and camera distance prediction. This mechanism enables the model to adaptively focus on different regions of the object at different scales, enhancing its ability to handle occlusions and multi-scale objects. In addition, the bottom-up method can capture the fine-grained details of the object, while the top-down method provides global pose information, thus improving the stability of pose estimation in complex scenes. At the same time, the experimental results on the PROPS benchmark dataset prove the superiority of MSDAPose. Using only RGB input, it achieves the current optimal accuracy of 99.08% in the ADD(-S) metric. Ablation experiments further verify the importance of multi-scale feature fusion and convolutional block attention modules in the framework. In short, MSDAPose provides a reliable and efficient solution for 6D object pose estimation in complex scenes and shows great application potential in the field of robotics and other related fields.
[0061] Furthermore, as Figure 6 shown, based on the above 6D object pose estimation method based on multi-scale deformable attention, the present invention also correspondingly provides a 6D object pose estimation system based on multi-scale deformable attention. Among them, the 6D object pose estimation system based on multi-scale deformable attention includes: A feature extraction processing module 51, configured to obtain a target RGB image and perform feature extraction processing on the target RGB image to obtain a plurality of feature images; A 3D translation vector acquisition module 52, configured to perform semantic annotation processing and translation estimation processing on a plurality of the feature images to obtain a plurality of masks of semantic labels and a plurality of vectors of implicit center point positions, and obtain a 3D translation vector according to the plurality of masks of semantic labels and the plurality of vectors of implicit center point positions; A 6D pose estimation result generation module 53, configured to perform rotation regression processing on a plurality of the feature images to obtain a 3D rotation estimation value, and obtain a 6D pose estimation result according to the 3D rotation estimation value and the 3D translation vector.
[0062] Further, as Figure 7 shown, based on the above 6D object pose estimation method and system based on multi-scale deformable attention, the present invention also correspondingly provides a terminal, which includes a processor 10, a memory 20, and a display 30. Figure 7 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0063] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a 6D object pose estimation program 40 based on multi-scale deformable attention is stored on the memory 20, and this 6D object pose estimation program 40 based on multi-scale deformable attention can be executed by the processor 10, so as to implement the 6D object pose estimation method based on multi-scale deformable attention in the present application.
[0064] The processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing the 6D object pose estimation method based on multi-scale deformable attention, etc.
[0065] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface. The terminals communicate with each other through a system bus.
[0066] In one embodiment, when the processor 10 executes the 6D object pose estimation program 40 based on multi-scale deformable attention in the memory 20, the steps of the 6D object pose estimation method based on multi-scale deformable attention as described above are implemented.
[0067] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a 6D object pose estimation program based on multi-scale deformable attention, and when the 6D object pose estimation program based on multi-scale deformable attention is executed by a processor, the steps of the 6D object pose estimation method based on multi-scale deformable attention as described above are implemented.
[0068] In summary, the present invention provides a 6D object pose estimation method, system and terminal based on multi-scale deformable attention. The method includes: obtaining a target RGB image, and performing feature extraction processing on the target RGB image to obtain a plurality of feature images; performing semantic annotation processing and translation estimation processing on the plurality of feature images to obtain masks of a plurality of semantic labels and vectors of a plurality of implicit center point positions, and obtaining a 3D translation vector according to the masks of the plurality of semantic labels and the vectors of the plurality of implicit center point positions; performing rotation regression processing on the plurality of feature images to obtain a 3D rotation estimation value, and obtaining a 6D pose estimation result according to the 3D rotation estimation value and the 3D translation vector. By performing semantic annotation processing, translation estimation processing and rotation regression processing on the feature images after the target RGB image feature processing, the present invention can effectively solve the problems of object occlusion and difficult pose estimation in multi-scale scenarios. Even if the target object is occluded by other objects, pixels can stably project the center, and accurate estimation of the 6D pose of the object in a complex scene can be achieved. At the same time, the stability of object pose estimation is also ensured.
[0069] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.
[0070] Certainly, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer, and when the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disc, etc.
[0071] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A 6D object pose estimation method based on multi-scale deformable attention, characterized in that: The 6D object pose estimation method based on multi-scale deformable attention includes: Acquire a target RGB image, and perform feature extraction processing on the target RGB image to obtain multiple feature images; Performing semantic labeling and translation estimation processing on the plurality of feature images to obtain masks of a plurality of semantic labels and vectors of a plurality of implicit center point positions, and obtaining a 3D translation vector according to the masks of the plurality of semantic labels and the vectors of the plurality of implicit center point positions; A rotation regression process is performed on the plurality of feature images to obtain a 3D rotation estimation value, and a 6D pose estimation result is obtained according to the 3D rotation estimation value and the 3D translation vector.
2. The 6D object pose estimation method based on multi-scale deformable attention according to claim 1, characterized in that: The step of acquiring a target RGB image and performing feature extraction processing on the target RGB image to obtain a plurality of feature images specifically includes: Obtain a target RGB image, and input the target RGB image into a 6D object pose estimation model; Performing feature encoding processing and spatial detail feature extraction processing on the target RGB image through a multi-scale feature pyramid in the 6D object posture estimation model to obtain semantic context information and spatial detail features; The semantic context information and the spatial detail features are subjected to feature fusion processing to obtain a plurality of feature images.
3. The 6D object pose estimation method based on multi-scale deformable attention according to claim 2, characterized in that: The semantic annotation processing includes dimensionality reduction processing and attention weight adjustment processing; the translation estimation processing includes upsampling processing and feature enhancement processing; The semantic annotation processing and translation estimation processing are performed on the plurality of feature images to obtain masks of a plurality of semantic labels and vectors of a plurality of implicit center point positions, and a 3D translation vector is obtained according to the masks of the plurality of semantic labels and the vectors of the plurality of implicit center point positions, specifically including: Performing dimensionality reduction processing and attention weight adjustment processing on the plurality of feature images to obtain masks of a plurality of semantic labels; Performing upsampling and feature enhancement processing on the plurality of feature images to obtain vectors of a plurality of implicit center point positions; Loss calculation is performed on the masks of the multiple semantic labels and the vectors of the multiple implicit center point positions to obtain a 3D translation vector.
4. The 6D object pose estimation method based on multi-scale deformable attention according to claim 3, characterized in that: The step of performing dimensionality reduction processing and attention weight adjustment processing on the plurality of feature images to obtain masks of a plurality of semantic labels specifically includes: Determine a plurality of projection layers in the 6D object pose estimation model, and select a corresponding target projection layer from the plurality of projection layers according to the scale of each of the feature images; Input each of the feature images into a corresponding target projection layer, and perform dimensionality reduction processing on each corresponding feature image through the target projection layer to obtain a plurality of reduced-dimensional feature maps; Attention weight adjustment processing and feature conversion processing are performed on the multiple dimensionality reduction feature maps to obtain masks of multiple semantic labels.
5. The 6D object pose estimation method based on multi-scale deformable attention according to claim 4, characterized in that: The step of performing attention weight adjustment processing and feature conversion processing on the plurality of dimension reduction feature maps to obtain masks of a plurality of semantic labels specifically includes: Inputting the plurality of reduced-dimensionality feature maps into a multi-scale deformable attention module in the 6D object posture estimation model, and performing attention weight adjustment processing on the plurality of reduced-dimensionality feature maps through the multi-scale deformable attention module to obtain a plurality of adjusted feature maps; The plurality of adjusted feature maps are input into a segmentation head module in the 6D object posture estimation model, and feature conversion processing is performed on the plurality of adjusted feature maps by the segmentation head module to obtain masks of a plurality of semantic labels.
6. The 6D object pose estimation method based on multi-scale deformable attention according to claim 4, characterized in that: The upsampling and feature enhancement processing is performed on the plurality of feature images to obtain a plurality of vectors of implicit center point positions, specifically including: Inputting the plurality of feature images into an upsampling module in the 6D object posture estimation model, and performing upsampling processing on the plurality of feature images through the upsampling module to obtain a plurality of high-dimensional feature images; The multiple high-dimensional feature images are input into the first convolutional block attention module in the 6D object posture estimation model, and the multiple high-dimensional feature images are subjected to feature enhancement processing by the first convolutional block attention module to obtain vectors of multiple implicit center point positions.
7. The 6D object pose estimation method based on multi-scale deformable attention according to claim 2, characterized in that: The performing rotation regression processing on the plurality of feature images to obtain a 3D rotation estimation value, and obtaining a 6D pose estimation result according to the 3D rotation estimation value and the 3D translation vector, specifically includes: Inputting the plurality of feature images into a pooling layer in the 6D object posture estimation model, and performing pooling processing on the plurality of feature images through the pooling layer to obtain a plurality of feature descriptors; Inputting the plurality of feature descriptors into a second convolutional block attention module in the 6D object posture estimation model, and performing feature enhancement processing on the feature descriptors through the second convolutional block attention module to obtain a plurality of enhanced feature descriptors; Inputting the plurality of enhanced feature descriptors into a fully connected layer in the 6D object posture estimation model, performing abstract feature extraction processing and dimension vector output processing on the plurality of enhanced feature descriptors through the fully connected layer to obtain a 3D rotation estimation value; A loss calculation is performed on the 3D rotation estimation value to obtain a target quaternion, and a 6D pose estimation result is obtained according to the target quaternion and the 3D translation vector.
8. A 6D object pose estimation system based on multi-scale deformable attention, characterized in that: The 6D object pose estimation system based on multi-scale deformable attention includes: A feature extraction processing module is used to obtain a target RGB image and perform feature extraction processing on the target RGB image to obtain multiple feature images; A 3D translation vector acquisition module, used to perform semantic annotation processing and translation estimation processing on the plurality of feature images, obtain masks of the plurality of semantic labels and vectors of the plurality of implicit center point positions, and obtain a 3D translation vector according to the masks of the plurality of semantic labels and the vectors of the plurality of implicit center point positions; The 6D pose estimation result generation module is used to perform rotation regression processing on the multiple feature images to obtain a 3D rotation estimation value, and obtain a 6D pose estimation result based on the 3D rotation estimation value and the 3D translation vector.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a 6D object pose estimation program based on multi-scale deformable attention stored in the memory and executable on the processor. When the 6D object pose estimation program based on multi-scale deformable attention is executed by the processor, the steps of the 6D object pose estimation method based on multi-scale deformable attention are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a 6D object pose estimation program based on multi-scale deformable attention. When the 6D object pose estimation program based on multi-scale deformable attention is executed by a processor, the steps of the 6D object pose estimation method based on multi-scale deformable attention are implemented as described in any one of claims 1 to 7.
Citation Information
Cited By
6D pose estimation method and device fusing attention mechanism, equipment and medium
CN121095347A