Class-level object 6D pose estimation method and system based on orientation perception and guidance
By introducing object orientation description in text prompts and combining it with a cross-projection feature interaction mechanism, the problems of poor cross-scene adaptability and low accuracy in existing methods are solved, and a more efficient pose estimation effect is achieved.
Patent Information
- Application Number
- CN202510911100.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-10
AI Technical Summary
Existing category-level object 6D pose estimation methods based on text cues lack consideration of fine-grained semantic attributes related to the pose estimation task and with strong robustness, making it difficult to achieve more effective and accurate pose estimation, and have poor cross-scene adaptability.
By introducing object orientation description in text prompts, contrastive learning is used to estimate the object orientation category, and its text description is integrated into the text prompt to guide the feature extraction and pose estimation process. The cross-projection feature interaction mechanism is combined to provide geometric learning supervision, thereby enhancing the model's cross-scene adaptability and pose estimation accuracy.
It significantly improves the model's cross-scene adaptability and pose estimation accuracy, provides semantic features that are robust to instance and environment changes, enhances the guidance effect of text prompts, and improves the accuracy of pose estimation.
Smart Images

Figure CN120765746A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot visual perception technology, and in particular to a method and system for estimating the 6D pose of a category-level object based on orientation perception and guidance. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Category-level 6D pose estimation is a vision-based perception task that aims to predict the pose of a target object in the camera coordinate system from an input image, including its 3D rotation, 3D translation, and 3D size. This technology can generalize to unseen objects of the same category, enabling cross-instance pose estimation. It is currently widely used in vision-guided robotic systems such as grasping, intelligent assembly, and human-robot interaction.
[0004] Currently, most deep learning-based methods for category-level 6D pose estimation rely primarily on single-image modality information, failing to fully leverage the complementary advantages of multimodal data. Current research demonstrates that textual modality information can effectively enhance visual task performance, prompting some research to utilize textual cues to improve category-level 6D pose estimation. Existing text-based 6D pose estimation methods for category-level objects can be broadly categorized into two technical paradigms: methods based on object appearance descriptions and methods based on object category descriptions. Essentially, both methods encode object attribute information as textual cues and extract cross-modal semantic features through multimodal models to improve pose estimation accuracy.
[0005] However, the current technical paradigm based on text prompts still has the following problems: (1) Methods based on object appearance description usually involve describing the object's color and texture attributes. These attributes are essentially irrelevant to the goal of pose estimation and are easily affected by differences in object instances and environmental changes. This not only leads to a decrease in feature quality, but also may cause false correlations between features and pose labels, thereby weakening the model's cross-scene adaptability and overall performance.
[0006] (2) The method based on object category description uses the object category label as a text prompt, but it can only provide global semantic information and has insufficient ability to represent pose changes, making it difficult to effectively guide the pose estimation task.
[0007] In other words, existing category-level object 6D pose estimation methods based on text prompts lack consideration of fine-grained semantic attributes related to the pose estimation task and with strong robustness, making it difficult to achieve more effective and accurate pose estimation. Summary of the Invention
[0008] To address the shortcomings of the above-mentioned prior art, the present invention provides a category-level 6D pose estimation method and system based on orientation perception and guidance. This method introduces modeling of the task-relevant attribute of object orientation in text prompts. By classifying orientation into multiple discrete categories, contrastive learning is used to estimate the object's orientation category and integrate its textual description into the text prompt to guide feature extraction and pose estimation. By introducing object orientation descriptions, the present invention provides task-relevant and robust fine-grained semantic features, effectively enhancing the guiding effect of text prompts and significantly improving the model's cross-scenario adaptability and pose estimation accuracy.
[0009] In a first aspect, the present invention provides a category-level object 6D pose estimation method based on orientation perception and guidance.
[0010] A method for estimating 6D pose of category-level objects based on orientation perception and guidance, comprising: Obtain the original scene image, identify the target object in the image, and extract the RGB image and point cloud of the target object; The RGB image of the target object and the corresponding orientation description text set of the object are input into the orientation perception network. The feature similarity between the RGB image and each text in the orientation description text set is calculated through comparative learning to generate the orientation description text prompt of the target object. The RGB image, point cloud and orientation description text prompt of the target object are input into the orientation-guided pose estimation network, and the orientation description text prompt features and point cloud features are extracted. The orientation description text prompt features are used to guide the extraction of image features. The point cloud and image features are fused and cross-projection feature interaction is performed to obtain the Nocturnal Orientation Sequence (NOCS) coordinates of the target object. Pose regression is performed based on the fused features, image features, point cloud features, orientation description text prompt features and NOCS coordinates to obtain the 6D pose and 3D size of the target object.
[0011] In a second aspect, the present invention provides a category-level object 6D pose estimation system based on orientation perception and guidance.
[0012] A category-level object 6D pose estimation system based on orientation perception and guidance, comprising: Image acquisition module, used to acquire original scene images, identify target objects in the images, and extract RGB images and point clouds of target objects; The orientation perception module is used to input the RGB image of the target object and the corresponding orientation description text set into the orientation perception network. Through comparative learning, the feature similarity between the RGB image and each text in the orientation description text set is calculated to generate a text prompt describing the orientation of the target object. The orientation-guided pose estimation module is used to input the target object's RGB image, point cloud, and orientation description text prompt into the orientation-guided pose estimation network, extract the orientation description text prompt features and point cloud features, and use the orientation description text prompt features to guide the extraction of image features; fuse the point cloud and image features and perform cross-projection feature interaction to obtain the target object's Nocturnal Orientation Coordinates (NOCS) coordinates; perform pose regression based on the fused features, image features, point cloud features, orientation description text prompt features, and NOCS coordinates to obtain the target object's 6D pose and 3D size.
[0013] In a third aspect, the present invention also provides an electronic device, comprising: a memory for storing executable instructions; and a processor for implementing the above-mentioned category-level object 6D pose estimation method based on orientation perception and guidance when executing the executable instructions stored in the memory.
[0014] In a fourth aspect, the present invention also provides a computer-readable storage medium storing executable instructions for causing a processor to execute the executable instructions to implement the above-mentioned category-level object 6D pose estimation method based on orientation perception and guidance.
[0015] In a fifth aspect, the present invention also provides a computer program product, which includes executable instructions, and the executable instructions are stored in a computer-readable storage medium; wherein, when the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the above-mentioned category-level object 6D pose estimation method based on orientation perception and guidance is implemented.
[0016] One or more of the above technical solutions have the following beneficial effects: 1. The present invention provides a method and system for category-level 6D pose estimation of objects based on orientation perception and guidance. This approach addresses the problems of poor cross-scene adaptability and limited guidance provided by text-based prompts, which are caused by existing methods based on text prompts and rely on task-irrelevant appearance attributes or coarse-grained category attributes. In the proposed method, orientation is divided into multiple discrete categories. Contrastive learning is used to estimate the orientation category of the object and its text description is integrated into the text prompt to guide feature extraction and pose estimation. The orientation description reflects the rotational relationship between the object and the camera coordinate system. The fine-grained orientation description and coarse-grained category information are integrated into the text prompt to construct a multi-granular text prompt. This design can provide task-related fine-grained semantic attributes to guide object pose estimation. At the same time, because the orientation description is not affected by changes in object instances and the environment, it can provide stable and robust semantic features in complex scenes. That is, it provides semantic features that are robust to instance and environmental changes to guide feature extraction and pose regression, thus avoiding the adverse effects of task-irrelevant appearance descriptions on cross-scene adaptability. In summary, the above features not only enhance the guiding effect of text prompts, but also improve the cross-scene adaptability of the model, thereby significantly improving the accuracy of pose estimation.
[0017] 2. Considering that textual prompts cannot directly provide explicit geometric supervision, the model's robustness to objects with significant shape variations is limited. Therefore, this paper proposes to enhance the model's geometric learning supervision capabilities by utilizing a cross-projection feature interaction mechanism to calculate the object's coordinate values in the Normalized Object Coordinate Space (NOCS). This can effectively improve the robustness to objects with significant shape variations and has high engineering practical value.
[0018] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0020] Figure 1 Flowchart of the method for class-level object 6D pose estimation based on orientation perception and guidance proposed in an embodiment of the present invention; Figure 2 Schematic diagram of a method for marking an object orientation according to an embodiment of the present invention; Figure 3 This is a diagram showing the network structure of orientation perception and guided pose estimation in an embodiment of the present invention; Figure 4This is a structural diagram of a cross-projection feature interaction module in an embodiment of the present invention; Figure 5 2 is a structural diagram of a pose estimation module in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] It should be noted that the following detailed descriptions are exemplary only and are intended to describe specific embodiments and provide further explanation of the present invention, and are not intended to limit the exemplary embodiments according to the present invention. Unless otherwise indicated, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the art to which the present invention belongs. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0022] Example 1 Considering that existing methods based on text prompts rely on task-irrelevant appearance attributes or coarse-grained category attributes, lack modeling of task-related attributes and explicit geometric supervision, the guiding role of text prompts is limited and the model finds it difficult to fully learn features related to pose, and cannot effectively adapt to object instances with significant shape changes, which in turn leads to poor cross-scene adaptability and limits the model's pose estimation performance. To this end, this embodiment proposes a category-level object 6D pose estimation method based on orientation perception and guidance, which guides pose estimation by integrating fine-grained orientation semantic attributes that are task-related and highly robust into text prompts. At the same time, combined with an effective geometric learning supervision mechanism, cross-scene adaptability and improved pose estimation accuracy are achieved.
[0023] This embodiment proposes a category-level object 6D pose estimation method based on orientation perception and guidance, such as Figure 1 As shown, the specific steps include: Step S1: Acquire the original scene image, identify the target object in the image, and extract the RGB image and point cloud of the target object; Step S2: Input the RGB image of the target object and the corresponding orientation description text set of the object into the orientation perception network, calculate the feature similarity between the RGB image and each text in the orientation description text set through comparative learning, and generate a text prompt describing the orientation of the target object; Step S3: Input the RGB image, point cloud, and orientation description text prompt of the target object into the orientation-guided pose estimation network, extract the orientation description text prompt features and point cloud features, and use the orientation description text prompt features to guide the extraction of image features; fuse the point cloud and image features and perform cross-projection feature interaction to obtain the NOCS coordinates of the target object; perform pose regression based on the fused features, image features, point cloud features, orientation description text prompt features, and NOCS coordinates to obtain the 6D pose and 3D size of the target object.
[0024] The above method proposed in this embodiment is described in more detail through the following content.
[0025] First, before performing 6D pose estimation of category-level objects, it is necessary to construct an orientation description text set for each type of object to facilitate subsequent applications.
[0026] Specifically, for different categories of objects, the object orientation categories are labeled and a set of orientation description texts is constructed. The orientation categories are labeled by calculating the angle between the object coordinate system and the pixel coordinate system in the 2D image, and an orientation description text set is constructed for each category of objects. The specific steps include: First, for each category of objects, obtain the 2D image of the object and mark the pixel coordinate system. Figure 2 As shown, the upper left corner of the image is taken as the origin and the pixel coordinate system is marked along the image direction.
[0027] Secondly, the angle between the coordinate systems is calculated. Based on the camera intrinsic parameters and the object's true position, the angle between the object coordinate system and the pixel coordinate system in the 2D image is calculated. , which can be expressed as: ; in, is the camera internal reference; is the true pose of the object; The orientation axis of the object is usually selected from the three coordinate axes of the object coordinate system, indicating the main semantic direction of the object, such as the handle direction of a cup or the opening direction of a bottle; Represents the pixel coordinate system axis; is the normalization factor.
[0028] Then, mark the orientation category. Based on the angle calculated above, , marking the discretized object orientation category. In this embodiment, for continuous rotationally symmetrical objects, the orientation categories of such objects are divided into two categories, "Upward" and "Downward", according to the angle; for non-continuous rotationally symmetrical objects, the orientation categories of such objects are divided into sixteen categories according to the angle, and each orientation category corresponds to The range is as Figure 2 shown.
[0029] Finally, the orientation description text set is constructed. Specifically, the orientation description text prompt template is used, such as "A photo of a [CLS], it is [ORI]." Among them, [CLS] represents the object category, and [ORI] represents the orientation description of the object. Combined with the object category determined in the previous step and the discretized several object orientation categories, the orientation description text prompts of all orientation categories are constructed to form the orientation description text set of the object of this category. Each object category corresponds to an orientation description text set containing the text description of all its orientation categories. For continuous rotational symmetry objects, the orientation description text set of each object category contains two text prompts, and for non-continuous rotational symmetry objects, the orientation description text set of each object category contains sixteen text prompts.
[0030] In step S1, the original scene image is obtained, the target object pixels are identified from the original scene image, and the corresponding region is cropped to extract the RGB image and point cloud of the target object. Specifically, the following steps are included: Step S11, a target segmentation algorithm such as Mask R-CNN series algorithm is used to process the original scene image to estimate the mask and class label of the target object.
[0031] Step S12, based on the target object mask, the background information is filtered out and the target object region is cropped, which is scaled to a predefined size, for example, 192x192 pixels, to generate the RGB image of the target object.
[0032] Step S13, combined with the mask of the target object and the camera intrinsic parameter, the depth image of the target object is converted into a point cloud.
[0033] In step S2, the target object orientation perception is performed, that is, the RGB image of the target object and the corresponding orientation description text set of the object are input into the orientation perception network, the feature similarity between the RGB image and the text in the orientation description text set is calculated through contrastive learning, and the text with the highest feature similarity with the image is selected as the orientation description text prompt of the object in the orientation description text set.
[0034] As Figure 3 shown, wherein ( a ) shows the generation of the orientation description text prompt of the object using the orientation perception network, which includes a CLIP model with frozen parameters and two adapter layers with updated parameters, wherein the CLIP model includes an image encoder and a text encoder. The orientation perception network is used for feature extraction, feature fine-tuning, and similarity calculation of images and texts in the orientation description text set, which includes the following steps: Step S21: Feature extraction of image and orientation description text set. Specifically, the RGB image of the target object and its corresponding orientation description text set are used as input, and the image encoder and text encoder in the CLIP model with frozen parameters are used to extract the features of the input image respectively. And the features of each orientation description text in the corresponding orientation description text set .
[0035] Step S22: Feature fine-tuning. Specifically, two adapter layers are used to fine-tune the image features and the features of each orientation description text in the orientation description text set, respectively, to enhance the perception of orientation semantics and generate fine-tuned image features. and orientation description text set features . Each adapter layer consists of two fully connected layers with residual connections. The mathematical expression of feature fine-tuning is: ; ; in, is the fusion weight, represents a fully connected layer.
[0036] Step S23: Similarity calculation: Specifically, by calculating the cosine similarity between the fine-tuned image features and each orientation description text feature in the orientation description text set, the text with the highest similarity is selected as the orientation description text prompt of the target object.
[0037] In step S3, the orientation-guided object pose estimation is performed. The RGB image, point cloud, and orientation description text of the target object are fed into the orientation-guided pose estimation network to calculate the 6D pose and 3D dimensions of the target object. Furthermore, the object's Nonce Coordinate System (NOCS) coordinates are calculated for geometric feature learning.
[0038] like Figure 3 As shown, where ( b ) demonstrates that the orientation-guided pose estimation network uses orientation text descriptions to guide feature extraction and pose regression to enhance the model's cross-scene adaptability. The orientation-guided pose estimation network is used to obtain the NOCS coordinates, 6D pose, and 3D size of the target object. Specifically, the following steps are included: Step S31: Extract point cloud and text prompt features. Specifically, the orientation description text prompt output in step S2 is input into the text encoder in the orientation perception network to extract text features. ; Input the object point cloud into PointNet++ and a fully connected layer respectively to extract point cloud features and location characteristics .
[0039] Step S32: Extract image features. Specifically, the RGB image is downsampled to 1 / 4 of the original image size, and the downsampled image and text features are input into U-Net. The image and text features are fused at different resolutions through cross-attention. The image features are aligned with the orientation semantic features, and the text features are used to guide the model to focus on orientation-sensitive image areas, such as the handle area of the cup. Finally, PSP-Net is used to restore the fused features to the original image size, and the image features are obtained through sampling operations. .
[0040] Step S33: Multimodal feature fusion. Specifically, the point cloud features extracted in step S31 are fused and location characteristics , and the image extracted in step S32 , generate multimodal fusion features This step specifically includes: Step S331: 、 and Concatenate in the channel dimension and fuse through a fully connected layer; Step S332: Pass the fused features through an average pooling layer to obtain global features; Step S333: Concatenate the global features and fusion features in the channel dimension and input them into a fully connected layer to obtain multimodal fusion features. .
[0041] Step S34: NOCS coordinate reconstruction. Specifically, the multimodal fusion features output from step S33 are Input to Figure 4 In the cross-projection feature interaction module shown in , the NOCS coordinates of the object are reconstructed to provide geometric information for subsequent pose estimation. Figure 4 The cross-projection feature interaction module shown reconstructs the NOCS coordinates of the object to assist the model in learning the geometric features of the object, specifically including the following steps: Step S341: Multimodal fusion features output from step S33 As input, self-attention is used to aggregate global context information to obtain self-attention weighted fusion features, which can be expressed as: ; ; in, 、 and is the linear transformation matrix; is the dimension of the query vector in the self-attention mechanism.
[0042] Step S342: The self-attention weighted fusion feature outputted from step S341 Mapped to three subspaces through three fully connected layers 、 and In , each subspace corresponds to the high-dimensional features of the orthogonal projection plane of the NOCS coordinates. When regressing the coordinates of NOCS on the orthogonal projection plane, there are inherent consistency constraints between different projection planes, such as Plane and Flat The coordinate components are consistent. Therefore, this embodiment uses the cross-attention mechanism for plane feature interaction to ensure information aggregation across subspaces and learn the inherent consistency constraints of the coordinate components between projection planes, which can be expressed as: ; in, Indicates concatenating two features along the length of the sequence; 、 and is the linear transformation matrix; Represents the dimension of the query vector in the crisscross attention mechanism.
[0043] Step S343: The features output from step S342 、 and Input to three fully connected layers to reconstruct the projection of NOCS coordinates on three orthogonal planes 、 and , and then the final NOCS coordinates are obtained by taking the mean of the corresponding coordinate components between the projection planes , which can be expressed as: ; in, and Respectively On the projection plane and Coordinate set, and so on for other symbols.
[0044] Furthermore, step S35, posture and size regression. Specifically, the orientation description text prompt feature outputted in step S31 is used. , point cloud features , location features , the image features output in step S32 , the fusion feature output in step S33 And the NOCS coordinates generated in step S34 are used as input, using Figure 5The pose estimation module shown calculates the 6D pose and 3D size of the object. This process specifically includes the following steps: Step S351: Input the NOCS coordinates into a fully connected layer to generate NOCS encoding features and compare it with 、 、 as well as Splicing is done in the channel dimension; the spliced features are passed through a fully connected layer to generate fused features, which are then input into the global average pooling layer to obtain global features.
[0045] Step S352: Concatenate the global features with the fused features generated in step S351 in the channel dimension and input them into a fully connected layer to obtain compressed features.
[0046] Step S353: Input the compressed features generated in step S352 into two independent fully connected layers to regress the translation matrix and size respectively; since the orientation description reflects the rotation relationship between the object coordinate system and the camera coordinate system, the orientation description text is used to guide the regression of the rotation matrix, that is, the orientation description text prompt feature The 6D pose and 3D size of the target object are obtained through this step.
[0047] As an implementation method, the aforementioned model is trained in two phases: the orientation-aware network and the orientation-guided pose estimation network are trained independently. In the first phase, the orientation-aware network is trained. During training, the adapter layer parameters are iteratively updated by maximizing the feature similarity between the input image and the corresponding orientation description text, while minimizing the feature similarity between the input image and other unmatched orientation description texts. In the second phase, the orientation-guided pose estimation network is trained. The network parameters are iteratively updated by calculating the difference between the output NOCS coordinates, 6D pose, and 3D size and the true values.
[0048] Among them, the contrast loss function of the first stage training is defined as: ; in, is the batch size for training; For the The characteristics of the image; Features describing the textual hint for orientation matching with the image; Describe features of textual cues for orientations that do not match the image; is the cosine similarity; is the temperature coefficient.
[0049] The pose and size loss functions for the second stage training are defined as: ; in, 、 and Represent the real rotation matrix, translation matrix and size respectively; 、 and Represent the predicted rotation matrix, translation matrix and size respectively; represents the L1 norm. The NOCS coordinate loss for the second stage training is defined as: ; in, and are the true and predicted NOCS coordinates, respectively.
[0050] The consistency loss of the second stage training is defined as: ; in, Indicates NOCS coordinates in Plane projection A collection of coordinate components, and so on for other symbols.
[0051] Then, the multi-task learning loss function for the second stage of training is constructed as: 。
[0052] To further verify the superiority of the method proposed in this embodiment, the following experiments were performed. Tables 1 and 2 below show the comparison results of this embodiment with the most advanced methods on the REAL275 and CAMERA25 datasets, respectively. Both datasets contain six types of objects: bottle, bowl, camera, can, laptop, and mug. The REAL275 dataset contains 2,750 test images, and the CAMERA25 dataset contains 25,000 test images. The methods listed in the table all use 3D intersection over union (IoU) and mean average precision (mAP) of the combined rotation and translation errors as evaluation indicators, where the threshold of 3D IoU is 75%, the threshold of rotation error is 5° and 10°, and the threshold of translation error is 2cm and 5cm.
[0053] Table 1: Performance comparison with state-of-the-art methods on the REAL275 dataset
[0054] Table 2: Performance comparison with state-of-the-art methods on the CAMERA25 dataset
[0055] As can be seen from the above, the method of this embodiment performs better than the most advanced methods on the dataset. By integrating the object orientation description into the text prompt, it provides task-related fine-grained robust semantics for pose estimation, significantly improving the model's cross-scene adaptability and overall performance.
[0056] Example 2 This embodiment provides a category-level object 6D pose estimation system based on orientation perception and guidance, including: Image acquisition module, used to acquire original scene images, identify target objects in the images, and extract RGB images and point clouds of target objects; The orientation perception module is used to input the RGB image of the target object and the corresponding orientation description text set into the orientation perception network. Through comparative learning, the feature similarity between the RGB image and each text in the orientation description text set is calculated to generate a text prompt describing the orientation of the target object. The orientation-guided pose estimation module is used to input the target object's RGB image, point cloud, and orientation description text prompt into the orientation-guided pose estimation network, extract the orientation description text prompt features and point cloud features, and use the orientation description text prompt features to guide the extraction of image features; fuse the point cloud and image features and perform cross-projection feature interaction to obtain the target object's Nocturnal Orientation Coordinates (NOCS) coordinates; perform pose regression based on the fused features, image features, point cloud features, orientation description text prompt features, and NOCS coordinates to obtain the target object's 6D pose and 3D size.
[0057] Example 3 This embodiment provides an electronic device, including: a memory for storing executable instructions; and a processor for implementing the above method provided in this embodiment when executing the executable instructions stored in the memory.
[0058] Example 4 This embodiment further provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by a processor, the processor will be caused to execute the above method provided in this embodiment.
[0059] Example 5 This embodiment provides a computer program product including executable instructions, which are computer instructions stored in a computer-readable storage medium. When a processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the electronic device performs the method provided in this embodiment.
[0060] The steps involved in the above embodiments 2 to 5 correspond to those in embodiment 1. For detailed implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media that includes one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to perform any method of the present invention.
[0061] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0062] The above description is only a preferred embodiment of the present invention. Although the specific implementation of the present invention is described in conjunction with the accompanying drawings, it does not limit the scope of protection of the present invention. Those skilled in the art should understand that on the basis of the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.
Claims
1. A method for class-level object 6D pose estimation based on orientation perception and guidance, characterized in that: include: Obtain the original scene image, identify the target object in the image, and extract the RGB image and point cloud of the target object; The RGB image of the target object and the corresponding orientation description text set of the object are input into the orientation perception network. The feature similarity between the RGB image and each text in the orientation description text set is calculated through comparative learning to generate the orientation description text prompt of the target object. The RGB image, point cloud, and orientation description text prompt of the target object are input into the orientation-guided pose estimation network, the orientation description text prompt features and point cloud features are extracted, and the orientation description text prompt features are used to guide the extraction of image features; The point cloud and image features are fused and cross-projection feature interaction is performed to obtain the NOCS coordinates of the target object; pose regression is performed based on the fused features, image features, point cloud features, orientation description text prompt features and NOCS coordinates to obtain the 6D pose and 3D size of the target object.
2. The method for class-level object 6D pose estimation based on orientation perception and guidance according to claim 1, wherein: Construction of a text set describing the orientation of different types of objects, including: For each category of objects, obtain the object's 2D image and annotate the pixel coordinate system; Calculate the angle between the object coordinate system and the pixel coordinate system in the 2D image based on the camera intrinsic parameters and the object's true pose; Based on the calculated angle, the discretized object orientation categories are labeled; Based on the orientation description text prompt template, the object category and several discretized object orientation categories are combined to construct the orientation description text prompts of all orientation categories, forming the orientation description text set of objects in this category.
3. The method for class-level object 6D pose estimation based on orientation perception and guidance according to claim 1, wherein: The orientation-aware network includes a CLIP model with frozen parameters and two adapter layers with updated parameters, wherein the CLIP model includes an image encoder and a text encoder; The RGB image of the target object and the corresponding orientation description text set of the object are input into the orientation perception network to generate the orientation description text prompt of the target object, including: Taking the RGB image of the target object and its corresponding orientation description text set as input, the image encoder and text encoder are used to extract the features of the input image and the text features of each orientation description text in the orientation description text set respectively; Two adapter layers are used to fine-tune the two types of extracted features respectively, and the similarity between the fine-tuned image features and the features of each orientation description text in the orientation description text set is calculated. The orientation description text with the highest similarity is selected as the orientation description text prompt of the target object.
4. The method for class-level object 6D pose estimation based on orientation perception and guidance according to claim 1, wherein: The RGB image, point cloud, and orientation description text of the target object are input into the orientation-guided pose estimation network to obtain the NOCS coordinates of the target object, including: Extract point cloud features and position features; Extracting text features that are oriented towards describing textual cues; The downsampled RGB image and text features are input into U-Net, which is fused at different resolutions through cross-attention. The text features are used to guide U-Net to focus on sensitive image areas and extract image features. The extracted point cloud features, position features and image features are fused to obtain multimodal fusion features; The multimodal fusion features are input into the cross-projection feature interaction module, and the global context information is aggregated using self-attention to obtain the self-attention weighted fusion features; The self-attention weighted fusion features are mapped into three subspaces through three fully connected layers, each of which corresponds to the high-dimensional features of the orthogonal projection plane of the NOCS coordinates; Utilize the cross-attention mechanism for plane feature interaction and learn the consistency constraints of coordinate components between projection planes; The features after interactive learning are passed through three fully connected layers to reconstruct the projection of the NOCS coordinates on three orthogonal planes, and then the final NOCS coordinates are obtained by taking the mean of the corresponding coordinate components between the projection planes.
5. The method for class-level object 6D pose estimation based on orientation perception and guidance according to claim 4, wherein: Based on the fusion features, image features, point cloud features, orientation description text prompt features and NOCS coordinates, pose regression is performed to obtain the 6D pose and 3D size of the target object, including: The NOCS coding features of NOCS coordinates are extracted using a fully connected layer, and the NOCS coding features are concatenated with image features, point cloud features, position features, and multimodal fusion features in the channel dimension. The fusion features are then passed through a fully connected layer to generate fusion features, and the fusion features are passed through a global average pooling layer to obtain global features. The global features and fusion features are spliced in the channel dimension, and compressed features are obtained after the fully connected layer; The compressed features are input into two independent fully connected layers to regress the translation matrix and size respectively. The orientation description text prompt features and the compressed features are then spliced in the channel dimension and input into a fully connected layer to regress the rotation matrix to obtain the 6D pose and 3D size of the target object.
6. The method for class-level object 6D pose estimation based on orientation perception and guidance according to claim 1, wherein: Independently train the orientation perception network and the orientation guidance pose estimation network, where: The first stage trains the orientation-aware network. During training, a contrastive loss function is constructed based on maximizing the feature similarity between the input image and the corresponding orientation description text, while minimizing the feature similarity between the input image and other non-matching orientation description texts. This is used to iteratively train and update the parameters of the adapter layer. The second stage trains the orientation-guided pose estimation network. During the training process, based on the difference between the output NOCS coordinates, 6D pose and 3D size and the true value, the pose and size loss, NOCS coordinate loss and consistency loss are constructed to form a multi-task learning total loss function, thereby iteratively training and updating the network parameters.
7. A category-level object 6D pose estimation system based on orientation perception and guidance, characterized in that: include: Image acquisition module, used to acquire original scene images, identify target objects in the images, and extract RGB images and point clouds of target objects; The orientation perception module is used to input the RGB image of the target object and the corresponding orientation description text set into the orientation perception network. Through comparative learning, the feature similarity between the RGB image and each text in the orientation description text set is calculated to generate a text prompt describing the orientation of the target object. The orientation-guided pose estimation module is used to input the RGB image, point cloud, and orientation description text prompt of the target object into the orientation-guided pose estimation network, extract the orientation description text prompt features and point cloud features, and use the orientation description text prompt features to guide the extraction of image features; The point cloud and image features are fused and cross-projection feature interaction is performed to obtain the NOCS coordinates of the target object; pose regression is performed based on the fused features, image features, point cloud features, orientation description text prompt features and NOCS coordinates to obtain the 6D pose and 3D size of the target object.
8. An electronic device, characterized in that: include: a memory for storing executable instructions; The processor is configured to implement the method for category-level object 6D pose estimation based on orientation perception and guidance as described in any one of claims 1 to 6 when executing the executable instructions stored in the memory.
9. A computer-readable storage medium, characterized in that Executable instructions are stored, which are used to cause the processor to execute the executable instructions to implement the category-level object 6D pose estimation method based on orientation perception and guidance as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product includes executable instructions stored in a computer-readable storage medium; When the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the category-level object 6D pose estimation method based on orientation perception and guidance described in any one of claims 1 to 6 is implemented.