6d pose estimation method and system for unseen object, and device and medium

By fusing 3D models and RGB-D images, a rotational invariant descriptor is generated and a matching relationship between objects and images is established, the problem of estimating poses without seeing objects is solved, and a high-precision and low-computation overhead pose estimation is achieved.

WO2025131018A1PCT designated stage expired Publication Date: 2025-06-26XYZ ROBOTICS CHINA INC

Patent Information

Application Number
PCT/CN2024/140811
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

It is difficult to effectively estimate the 6D pose of an unseen object, especially when processing multiple objects or generalizing to objects that are not presented in the training data, resource-intensive training and high computational overhead become bottlenecks.

Method used

By fusing the 3D model of an unseen object with the RGB-D image, a general color-aware rotation-invariant 3D descriptor is generated, and a matching relationship between the object and the image is established, and the pose parameters are optimized to estimate the 6D pose of the object.

Benefits of technology

High-precision 6D pose estimation of unseen objects is achieved, reducing calculation overhead, enhancing the generalization ability of new objects, and capturing the natural symmetry of objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140811_26062025_PF_FP_ABST
    Figure CN2024140811_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are an RGB-D image-based 6D pose estimation method and system for an unseen object, and a device and a medium. The method comprises: performing fusion processing on a 3D model and an RGB-D image of an unseen object to obtain fusion information; on the basis of the fusion information, generating a universal color-aware rotation-invariant 3D descriptor; and on the basis of the universal color-aware rotation-invariant 3D descriptor, establishing a matching relationship between the 3D model and the RGB-D image of the unseen object, and estimating the 6D pose of the unseen object. In the method in the present application, the natural symmetry of an object can be inherently captured and modeled without relying on explicit symmetry annotations, thereby reducing the pose ambiguity; and a descriptor that combines texture and geometric information to describe the appearance and shape features of an object in a complementary and universal manner makes the matching between a 3D model and an RGB-D image of an unseen object more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

6D pose estimation method, system, device and medium for unseen objects Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, system, device, and medium for estimating the 6D pose of an unseen object based on RGB-D images. Background Art

[0002] Estimating the 6D pose of objects is a critical task in computer vision applications such as robotic manipulation, augmented reality, and autonomous driving. Frameworks based on template matching, keypoint detection, surface mapping, and reconstruction have achieved high-accuracy recognition of single objects. However, these methods are not designed to handle multiple objects or generalize to objects not present in the training data.

[0003] Current object pose estimation learning methods require resource-intensive training for each individual object instance or category, hindering their scalability for real-world applications when faced with previously unseen objects. Conventional category-level pose estimation methods struggle to handle novel categories when dealing with multiple objects in a dataset, especially when faced with novel instances. These methods are ill-suited for unseen object pose estimation in real-world applications, where 3D models are only available at inference time, as they are designed to overfit to the specific distribution of an object, category, or dataset. Some one-shot learning methods attempt to use template matching to align object models or capture structure from motion (SfM) of unseen objects. However, these methods often require object-specific preprocessing steps. Classic approaches to pose estimation for unseen objects employ hand-crafted features and correspondences between CAD models and observed RGB-D images. However, these methods introduce many highly ambiguous assumptions that require iterative ranking and refinement, resulting in computational overhead.

[0004] While general pose estimation has been extensively studied recently by proposing learned models that evaluate and refine pose hypotheses, their typical designs involve time-consuming rendering and comparison strategies. This approach limits their usefulness in real-world applications. An alternative approach formulates the general 6D pose estimation problem as a point cloud registration task, which benefits from point cloud representation backbones but ignores critical texture information in RGB images. As a result, descriptors learned solely from point clouds often have less discriminative power.

[0005] Public content

[0006] In response to the defects in the prior art, the purpose of this application is to provide a 6D pose estimation method, system, device and medium for unseen objects based on RGB-D images.

[0007] In a first aspect, embodiments of the present application provide a method for estimating the 6D pose of an unseen object based on an RGB-D image, comprising:

[0008] Fuse the 3D model of the unseen object with the RGB-D image to obtain fusion information;

[0009] generating a universal color-aware rotation-invariant 3D descriptor based on the fused information;

[0010] A matching relationship between a 3D model of an unseen object and an RGB-D image is established based on a universal color-aware rotation-invariant 3D descriptor, and the 6D pose of the unseen object is estimated.

[0011] Optionally, fusing the 3D model of the unseen object with the RGB-D image to obtain fusion information includes:

[0012] The 3D model of the unseen object is represented as a point cloud P, the partial point cloud obtained by the depth channel of the RGB-D image is recorded as Q, and the RGB of the cropped local object is recorded as K;

[0013] Extract 3D local features from point clouds P and Q, and extract 2D local features from the cropped RGB image K;

[0014] Fusion of 3D local features and 2D local features in the latent space where the superscript d represents the dimension of the latent space.

[0015] Optionally, generating a universal color-perceptual rotation-invariant 3D descriptor according to the fusion information includes:

[0016] Extracting latent space The general descriptor Φ corresponding to the point in the point cloud P P , and the universal descriptor Φ corresponding to the points of the point cloud Q Q ;

[0017] Calculate Φ P and Φ Q The similarity between them and construct the corresponding set in Represents point p i and q j Match descriptor of , where: p i represents the i-th point in the point cloud P, 1≤i≤n, n represents the total number of point clouds in the 3D model, q j Represents the j-th point in the depth point cloud Q, 1≤j≤m, and m represents the number of point clouds sampled in the depth image.

[0018] Optionally, establishing a matching relationship between a 3D model of an unseen object and an RGB-D image based on a universal color-aware rotation-invariant 3D descriptor, and estimating a 6D pose of the unseen object, includes:

[0019] According to the input point cloud P and point cloud Q, and the cropped RGB image K, the mapping function Ψ maps the point cloud P and point cloud Q to the universal descriptor Φ P and Φ Q , and fuse the cross-modal information from (Q, K) and (P, K) respectively;

[0020] By matching universal descriptors, we establish correspondences between unseen objects and RGB-D images, and estimate the 6D pose of the unseen objects by optimizing pose parameters.

[0021] Optionally, extracting 3D local features from the point cloud P and the point cloud Q includes:

[0022] The Transformer-based architecture RoITr builds the backbone of the encoder and samples the input point cloud P and point cloud Q into super points P′ and Q′ through the farthest point sampling of the encoder, where and n′ and m′ represent the number of superpoints in P′ and Q′, respectively;

[0023] For each superpoint p i ′ and q j ′, extract local geometric features from neighboring points within a radius of r, so that a local rotation-invariant 3D descriptor can be obtained by projecting the local geometric features into the latent space.

[0024] The Transformer in this embodiment is a deep learning model that uses an attention mechanism, which can assign different weights according to the importance of each part of the input data.

[0025] Optionally, extracting 2D local features from the cropped RGB image K includes:

[0026] The improved encoder of FPN is used as the convolutional neural network (CNN). The cropped RGB image K is sampled into a local feature map of a preset size, and the local texture information is projected into a d-dimensional latent space consistent with the 3D geometric features. Where H represents the length of the image and W represents the width of the image.

[0027] Flatten the local feature map of the image, where the 2D superpixel is represented as K′ and the 2D superpixel feature is represented as Φ K′ .

[0028] Optionally, 3D local features and 2D local features are fused in the latent space Including:

[0029] The extracted 3D local features and the context of the 2D local features are fused in the latent space through a latent fusion attention module; wherein the latent fusion attention module includes: a 3D-2D fusion block and a 2D-3D fusion block;

[0030] The 3D-2D fusion block is specifically configured as follows:

[0031] Aggregate Φ using a global transformer P′ and Φ Q′ , so that the RGB features are sequentially fused with the global perception depth and 3D features, and the cross-modal 2D features of each superpoint are obtained through the potential fusion transformer, where Φ P′ is Φ P The corresponding super-point feature, Φ Q′ is Φ Q Corresponding super-point features;

[0032] The 2D-3D fusion block is specifically configured as follows:

[0033] We use RGB features to enhance Φ separately through latent fusion transformer P′ and Φ Q′ The super-point features are jointly injected into the enhanced super-point features through the global transformer to generate 3D super-point features containing 2D perception and

[0034] Optionally, establish correspondences between unseen objects and RGB-D images by matching universal descriptors, and estimate the 6D pose of the unseen objects by optimizing pose parameters, including:

[0035] Establish 3D-3D matching between point clouds P and Q through the extracted features, and use covariance analysis to measure the similarity of normalized features;

[0036] By determining the top κ most relevant putative 3D-3D matches for the set of corresponding points C, we create η pose hypotheses from them;

[0037] For each 1≤ν≤η hypothesis, we write Randomly select s corresponding relations from the set C, where s < < κ;

[0038] The 6D pose hypothesis of the object is solved by the obtained correspondence set C, and all pose hypotheses are ranked according to the average score between the 3D and RGB verification process, where the 3D score is calculated based on the Euclidean point-to-point distance between the converted 3D model and the extracted depth map, and the RGB score is output through a deep extension network. The final prediction is the pose with the highest score

[0039] In a second aspect, an embodiment of the present application provides a 6D pose estimation device for an unseen object based on an RGB-D image, comprising: a processor and a memory, wherein the memory stores executable program instructions, and when the processor calls the program instructions in the memory, the processor is configured to:

[0040] Perform the steps of the method for estimating 6D pose of an unseen object based on RGB-D images as described in any one of the first aspects.

[0041] In a third aspect, an embodiment of the present application provides a computer-readable storage medium configured to store a program, which, when executed, implements the steps of the method for 6D pose estimation of an unseen object based on RGB-D images as described in any one of the first aspects.

[0042] In a fourth aspect, an embodiment of the present application provides a program product, comprising a computer program, wherein the computer program is stored in a readable storage medium, and at least one processor of a robot can read the computer program from the readable storage medium, and the at least one processor executes the computer program so that the robot implements the steps of the 6D pose estimation method for an unseen object based on RGB-D images in the first aspect.

[0043] In a fifth aspect, an embodiment of the present application provides a system for estimating the 6D pose of an unseen object based on an RGB-D image, comprising:

[0044] The fusion unit is used to fuse the 3D model of the unseen object and the RGB-D image to obtain fusion information;

[0045] The descriptor generation unit is configured to generate a universal color perception rotation invariant 3D descriptor according to the fusion information;

[0046] The matching and pose estimation unit is used to establish a matching relationship between the 3D model of the unseen object and the RGB-D image based on the universal color-aware rotation-invariant 3D descriptor, and to estimate the 6D pose of the unseen object.

[0047] Compared with the prior art, this application has the following beneficial effects:

[0048] This application uses a fusion-description-matching strategy to estimate the 6D pose of an unseen object from a single RGB-D image, extracting rotation-invariant descriptors that can be shared across various objects, thereby promoting generalization to unseen objects. The method in this application can inherently capture and model the natural symmetry of objects without relying on explicit symmetry annotations, effectively reducing pose ambiguity. In addition, by combining texture and geometric information to describe the descriptors of the appearance and shape characteristics of the object in a complementary and universal way, and by introducing a loss function, the association between texture and geometric features is strengthened, thereby making the matching between the CAD model and the RGB-D image of the unseen object more precise and accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on the provided drawings without inventive work. By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes and advantages of the present disclosure will become more apparent:

[0050] FIG1 is a schematic diagram illustrating the principle of a method for estimating the 6D pose of an unseen object based on RGB-D images according to an embodiment of the present application;

[0051] FIG2 is a flow chart of a method for estimating 6D pose of an unseen object based on RGB-D images provided in an embodiment of the present application;

[0052] FIG3 is a schematic diagram of the principle architecture of estimating the 6D pose of an unseen object based on RGB-D images according to an embodiment of the present application;

[0053] Figure 4 is a diagram showing the quantitative results of the average recall rate (AR) on the BOP-5 core benchmark dataset for the unseen object pose estimation task;

[0054] FIG5 is a schematic diagram showing qualitative results of 6D pose estimation using the method in the embodiment of the present application compared with Megapose and ZeroPose;

[0055] FIG6 is a schematic diagram of the t-SNE visualization results of the object descriptors provided in an embodiment of the present application;

[0056] FIG7 is a schematic diagram of t-SNE visualization comparison results of regular tetrahedrons with different colored faces provided in an embodiment of the present application;

[0057] FIG8 is a schematic diagram showing the results of quantitatively evaluating the 6D pose (ADD-0.1d) of the visible object pose estimation task on the LM-O dataset provided by an embodiment of the present application;

[0058] FIG9 is a schematic diagram of ablation study results of a key design on the BOP-5 dataset provided by an embodiment of the present application;

[0059] Figure 10(a) is a schematic diagram of AR with different numbers of hypotheses;

[0060] Figure 10(b) is a schematic diagram of the AR of the final pose and the HR of the pose hypotheses with different numbers of corresponding samples;

[0061] FIG11 is a schematic diagram of the structure of a device for estimating 6D pose of an unseen object based on RGB-D images according to an embodiment of the present application;

[0062] FIG12 is a schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0063] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0064] It should be noted that when a component is referred to as being "fixed to" another component, it may be directly on the other component or there may be an intermediate component. When a component is referred to as being "connected to" another component, it may be directly connected to the other component or there may be an intermediate component.

[0065] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit this disclosure. The term "and / or" as used in this disclosure includes any and all combinations of one or more of the relevant listed items.

[0066] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of the present disclosure and in the accompanying drawings are configured to distinguish similar objects and are not necessarily configured to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present disclosure described herein, for example, can be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0067] The following specific embodiments are used to describe in detail the technical solutions of the present disclosure and how the technical solutions of the present application solve the above technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0068] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features therein may be combined with each other.

[0069] The task of unseen object pose estimation aims to estimate the 6D pose between a CAD model not seen during training and its partial observations from RGB and / or depth images.

[0070] This application proposes a method for 6D pose estimation of unseen objects based on RGB-D images, which can be abbreviated as MatchU. MatchU is a general method that fuses 2D texture and 3D geometric cues and is configured to predict the 6D pose of unseen objects. First, when estimating the 6D pose, it relies on learning geometric 3D descriptors, which are rotation-invariant by design. By encoding pose-independent geometry, the learned descriptors are naturally generalized to unseen objects and capture symmetries. To resolve ambiguous associations using only 3D geometry, the method in this application incorporates additional RGB information into the descriptors. This is achieved through a novel attention-based mechanism that incorporates cross-modal information, and a matching loss that utilizes a latent space learned from RGB data to guide the descriptor learning process. Extensive experiments have demonstrated that MatchU far outperforms all existing methods in both accuracy and speed, without even requiring expensive retraining or rendering.

[0071] In this embodiment, the pose of an unseen object is estimated by matching descriptors between learned RGB-D data and its CAD model. In this embodiment, by extracting the fused RGB-D input features, the unseen CAD model can be registered to the object in the scene.

[0072] Figure 1 is a schematic diagram of the principle of a method for estimating the 6D pose of an unseen object based on an RGB-D image provided by an embodiment of the present application. As shown in Figure 1, first, in the first step, the unseen 3D model (CAD model) and the RGB-D image (RGB image and corresponding depth image) are fused to obtain fused information; in the second step, a universal color-perceived rotation-invariant 3D descriptor is generated based on the fused information; in the third step, a matching relationship is established between the 3D model of the unseen object and the RGB-D image based on the universal color-perceived rotation-invariant 3D descriptor to estimate the 6D pose of the unseen object.

[0073] For example, as shown in Figure 1, this embodiment aims to extract rotation-invariant descriptors that can be shared across various objects, promoting generalization to unseen objects. The extraction of rotation-invariant descriptors is crucial because they can inherently capture and model the natural symmetry of objects without relying on explicit symmetry annotations. However, rotation invariance still has some ambiguity, that is, a point can match multiple geometrically similar points. To address the ambiguity problem caused by rotation invariance, a novel 2D-3D fusion module, called the latent fusion attention module, is introduced in this embodiment. This module effectively combines texture and geometric information. This leads to the extraction of descriptors that describe the appearance and shape features of objects in a complementary and universal manner. In addition, this embodiment also proposes a novel bridging coarse-level matching loss, which utilizes RGB information to enhance the learning of geometric descriptors. This loss function strengthens the association between texture and geometric features, thereby making the matching between CAD models and RGB-D images of unseen objects more precise and accurate.

[0074] FIG2 is a flow chart of a method for estimating 6D pose of an unseen object based on RGB-D images provided in an embodiment of the present application. As shown in FIG2 , the method in this embodiment may include:

[0075] Step S201: Fusing the 3D model of the unseen object with the RGB-D image to obtain fusion information.

[0076] In this embodiment, the 3D model of the unseen object is represented as a point cloud P, the partial point cloud obtained from the depth channel of the RGB-D image is recorded as Q, and the RGB of the cropped local object is recorded as K; 3D local features are extracted from the point cloud P and point cloud Q, and 2D local features are extracted from the cropped RGB image K; the 3D local features and 2D local features are fused in the latent space where the superscript d represents the dimension of the latent space.

[0077] In an optional embodiment, the backbone of the encoder can be constructed based on the transformer architecture RoITr, and the input point cloud P and point cloud Q are sampled into super points P′ and Q′ through the farthest point sampling of the encoder, where and n′ and m′ represent the number of superpoints in P′ and Q′ respectively. Then, for each superpoint p i ′ and q j ′, extract local geometric features from neighboring points within a radius of r, so that a local rotation-invariant 3D descriptor can be obtained by projecting the local geometric features into the latent space.

[0078] In this embodiment, a superpoint refers to a method for downsampling an original dense point cloud into a sparse point cloud capable of representing its geometric features through certain specific sampling methods. Points in this sparse point cloud are called superpoints, and each superpoint is represented as a coordinate in space and has a feature descriptor. A superpixel refers to a method for downsampling an original image into a pixel map with a lower resolution through certain specific sampling methods. Each pixel at this resolution is called a superpixel, and each superpixel has a feature descriptor.

[0079] In another optional embodiment, the improved encoder of FPN can be used as the convolutional neural network CNN, and the cropped RGB image K is sampled to a preset size (for example Where H represents the length of the image and W represents the width of the image), and the local texture information is projected into a d-dimensional latent space consistent with the 3D geometric features; Then, the local feature map of the image is flattened, where the 2D superpixel is represented as K′ and the 2D superpixel feature is represented as Φ K′ .

[0080] For example, the context of the extracted 3D local features and 2D local features can be fused in the latent space through a latent fusion attention module, wherein the latent fusion attention module includes: a 3D-2D fusion block and a 2D-3D fusion block. The 3D-2D fusion block is specifically configured as follows: using a global transformer to aggregate Φ P′ and Φ Q′ , so that the RGB features are sequentially fused with the global perception depth and 3D features, and the cross-modal 2D features of each superpoint are obtained through the potential fusion transformer, where Φ P′ is Φ P The corresponding super-point feature, Φ Q′ is Φ QThe corresponding super-point features; 2D-3D fusion block, specifically configured as follows: using RGB features through the potential fusion transformer to enhance Φ P′ and Φ Q′ The super-point features are jointly injected into the enhanced super-point features through the global transformer to generate 3D super-point features containing 2D perception and

[0081] Step S202: Generate a universal color-perceptual rotation-invariant 3D descriptor based on the fused information.

[0082] In this embodiment, first, the latent space is extracted The general descriptor Φ corresponding to the point in the point cloud P P , and the universal descriptor Φ corresponding to the points of the point cloud Q Q Then, calculate Φ P and Φ Q The similarity between them and construct the corresponding set in Represents point p i and q j Match descriptor of , where: p i represents the i-th point in the point cloud P, 1≤i≤n, n represents the total number of point clouds in the 3D model, q j Represents the j-th point in the depth point cloud Q, 1≤j≤m, and m represents the number of point clouds sampled in the depth image.

[0083] Step S203: Establish a matching relationship between the 3D model of the unseen object and the RGB-D image based on the universal color-aware rotation-invariant 3D descriptor, and estimate the 6D pose of the unseen object.

[0084] In this embodiment, according to the input point cloud P and point cloud Q, and the cropped RGB image K, the mapping function Ψ maps the point cloud P and point cloud Q to the universal descriptor Φ P and Φ Q , and fuse cross-modal information from (Q, K) and (P, K) respectively; establish correspondences between unseen objects and RGB-D images by matching universal descriptors, and estimate the 6D pose of unseen objects by optimizing pose parameters.

[0085] For example, a 3D-3D match is established between point clouds P and Q using the extracted features, and the similarity of the normalized features is measured using covariance analysis. η pose hypotheses are created from the top κ most relevant hypothetical 3D-3D matches of the corresponding point set C. For each hypothesis with 1≤ν≤η, denoted as Randomly select s correspondences from the set C, where s < < κ. In addition, the 6D pose can be optimized using RANSAC (a least squares optimization algorithm is modeled through the resulting correspondence set C to solve the 6D pose of the object), and all pose hypotheses are ranked according to the average score between the 3D and RGB verification process, where the 3D score is calculated based on the Euclidean point-to-point distance between the converted 3D model and the extracted depth map, and the RGB score is output through a deep probing network. The final prediction is the pose with the highest score.

[0086] In this embodiment, the 3D model of the unseen object is represented as a point cloud n represents the total number of point clouds of the model, and the part of the point cloud obtained from the depth channel is recorded as m represents the total number of point clouds output by the depth channel, and the RGB of the cropped local object is recorded as K. are collected in a set of predicted corresponding points C, which is used to estimate the 6D pose of the new object by optimizing the mutual 3D correspondence targets. The 6D pose of the new object in the special Euclidean group SE(3) representing rigid transformations in 3D space.

[0087] In order to solve the object pose, we first extract the latent space The universal descriptor Φ of the points in point cloud P and point cloud Q P and Φ Q To calculate the corresponding relationship. By calculating Φ P and Φ Q The similarity between them, construct the corresponding set in Represents point p i and q j The matching descriptor of . Formula (1) is formulated as a least squares problem and can be robustly solved using outlier-aware consensus algorithms such as RANSAC.

[0088] For example, according to the input point cloud P and point cloud Q, and the cropped RGB image K, the mapping function Ψ maps the point cloud P and point cloud Q to a universal descriptor: and The cross-modal information from (Q, K) and (P, K) is fused separately. By matching universal descriptors, a correspondence is established between unseen objects and their partial observations, and the pose of the object is finally estimated.

[0089] Figure 3 is a schematic diagram of the principle architecture of estimating the 6D pose of an unseen object based on RGB-D images provided by an embodiment of the present application. As shown in Figure 3, 3D and 2D local features are first extracted from point clouds P and Q and the cropped RGB image K, and then the 3D and 2D local features are fused in the latent space. The fusion attention module bridges the coarse-level matching loss. The updated 3D descriptor is denoted as and They are then passed to the decoder, which extracts them into higher resolution descriptors Φ P and Φ Q , and matching losses at a fine level Finally, the pose parameters can be optimized by matching the descriptors in the latent space To estimate the 6D pose of unseen objects and continuously optimize the pose parameters

[0090] In this embodiment, the latest Transformer-based architecture RoITr is used as the encoder backbone to extract rotation-invariant 3D local features from the CAD point cloud P and the partially observed point cloud Q in the depth image. The inherent rotation invariance of the descriptor provides robust feature extraction for geometric cues and guarantees generality for unseen objects. Given point cloud P and point cloud Q, the encoder downsamples the input point cloud into superpoints by farthest point sampling (FPS). They represent a well-distributed coarse representation of the spatial structure from the underlying dense point cloud and are defined as and Where n′ and m′ represent the number of superpoints in P′ and Q′ respectively. i ′ and q j ′, first extract local geometric features from neighboring points within a radius of r; then, project the local geometric features into the latent space through a series of attention blocks, and obtain the inherent rotation-invariant local 3D geometric descriptor represented as and where d is the dimension of the latent space.

[0091] In this embodiment, a convolutional neural network (CNN) is configured for local visual feature extraction. Following LoFTR, an improved FPN encoder is used as the CNN backbone. The 2D encoder crops and samples the input image of size H×W into a size of The local feature map of the image is then flattened into Among them, the 2D superpixel is represented as K′, and the 2D superpixel feature is represented as ΦK′ .

[0092] In this embodiment, after extracting 3D and 2D local features, the encoded 3D and 2D contexts are fused in the latent space through the latent fusion attention module. To maintain the versatility of the network and avoid overfitting to specific object features, the latent fusion transformer and the global transformer can be used for fusion in these two fusion blocks.

[0093] In an optional implementation, spatial awareness is incorporated by using position encoding in the attention mechanism and implicitly aligning different modalities. For 2D features, the spatial information of the 2D feature map is encoded into the feature space according to DETR. For 3D features, pose-independent point pair features (PPFs) can be used as the subsequent position representation instead of encoding the original position of the points, which ensures geometric rotation invariance and versatility in pose estimation of unseen objects.

[0094] The Latent Fusion Transformer aims to fuse 2D superpixel features and 3D superpoint features in the latent space. The Latent Fusion Transformer consists of a series of self-attention and cross-attention layers. In practice, g self-attention and cross-attention layers can be stacked for each Latent Fusion Transformer. The Global Transformer aims to aggregate the global context of 3D and 3D features.

[0095] 3D-2D fusion block and 2D-3D fusion block. For the 3D-2D fusion block, the global transformer is first used to aggregate Φ P′ and Φ Q′ Then, the RGB features are sequentially fused with the global perception depth and CAD features to obtain the final cross-modal 2D features of each superpoint through the latent fusion transformer. For the 2D-3D fusion block, the CADΦ is ​​first enhanced separately using the RGB features through the latent fusion transformer. P′ and depth Φ Q′ Superpoint features. The global transformer then jointly injects this information to provide 2D-aware 3D superpoint features for CAD and depth. and

[0096] Exemplarily, several loss functions can be used to guide the learning of fusion descriptors. For example, using the latent features learned from the RGB image as a bridge between the CAD latent space and the depth point cloud (BCM Loss, i.e., bridging coarse-level matching loss) can significantly promote the unification of the two different 3D-based latent spaces and help generate more robust and reliable correspondences between superpoints. In addition, a fine-level matching loss is introduced in the embodiment of the present application to guide the refinement of superpoint matching to point correspondences.

[0097] In this embodiment, to ensure the validity of RGB-based 2D information in the latent space, it is crucial to provide supervisory signals from both 2D and 3D modalities by establishing cross-modal matching between 2D and 3D features. The alignment between superpoints P′ and Q′ can be obtained using the groundtruth transformation matrix and nearest neighbor search. The true 3D-2D correspondence between superpoints is inherent in the RGB-D pair.

[0098] For example, circle loss is used to maximize the similarity of superpoint positive pairs and minimize the similarity of superpoint negative pairs. i ′∈P′ and q j ′∈Q′, where p i ′ and q j The overlap calculation formula between ′ is as follows:

[0099] in: Indicates the corresponding relationship, is assigned from P′ to p i ′’s point group i is grouped by point-to-node strategy, is assigned from Q′ to q j ′’s point group j is grouped by point-to-node strategy. A pair of superpoints p i ′ and q j ′, if and only if When p i ′ and q j ′ is considered as a positive pair, where τ r Is the overlap threshold. When a set of positive superpoints is sampled from Q′ and a set of negative superpoints is sampled from P′. The coarse-level superpoint Circle Loss loss of P′ can be calculated using the overlap weight, which is expressed as

[0100] The same loss for Q′ is similarly defined, and the overall loss between superpoints P′ and Q′ is defined as follows:

[0101] Similarly, when Circle Loss is configured as a 3D-2D coarse-level matching, the 3D positive and negative samples are first projected onto the 2D plane, and then the positive and negative pairs between the 2D superpoints and the 3D superpoints are obtained. The loss function between K′ and Q′ is defined as The loss function between K′ and P′ is defined as Then the overall bridging coarse-level matching loss between P′, Q′ and K′ is defined as follows:

[0102] Among them, λ b are the weights of 3D-3D and 3D-2D matching losses, respectively.

[0103] Exemplarily, in order to improve the accuracy of 3D-3D correspondence, in this embodiment, a fine-level matching loss is applied to the point cloud P and the point cloud Q, where denser points P' and Q' are generated from the coarse-level super point cloud P and the point cloud Q. Given the super-point correspondence, a set of fine point features is assigned to each super point through a point-to-node strategy, and the similarity matrix between the corresponding groups can be calculated. The fine-level matching is formulated as an optimal transfer problem, which can be solved by the Sinkhorn algorithm. Negative log-likelihood is applied to the similarity matrix to obtain the fine matching loss between P and Q.

[0104] In summary, the overall loss function of training is defined as follows:

[0105] in: is the bridging coarse-level matching loss, is the fine-level matching loss, λ c is the weight to balance the coarse and fine level training.

[0106] In this example, 3D-3D matching is established between the CAD point cloud P and the observation Q using the extracted features. Covariance analysis is used to measure the similarity of the normalized features. By determining the top k most relevant hypothetical 3D-3D matches for the corresponding point set C, n pose hypotheses are created from them. For each hypothesis 1≤ν≤η First, we randomly select correspondences with s < < κ from C and solve Equation (1). We use RANSAC to optimize the 6D object pose. This process speeds up the prediction process and provides a control parameter η to determine the efficiency of MatchU. All hypotheses are then ranked according to the average score between the 3D and RGB validation processes. For 3D, the score is calculated based on the Euclidean point-to-point distance between the converted CAD model and the lifted depth map. The final prediction is the pose with the highest score.

[0107] Figure 4 shows quantitative results of average recall (AR) on the BOP-5 core benchmark dataset for the unseen object pose estimation task. MH indicates whether multiple hypotheses are used. Instance location indicates whether the object detector or segmenter is trained on the test object. Refine indicates whether depth and / or RGB images are used to refine the results.

[0108] For example, the method in this embodiment is first evaluated on the five BOP core datasets and compared with other unseen object pose estimation methods. The learned descriptors are then visualized to highlight their ability to simultaneously capture symmetry and encode texture without the need for external annotations. Since recent RGB-D fusion methods are designed only for seen objects, the method in this embodiment is compared with methods on the LM-O dataset.

[0109] For example, the model in this embodiment was trained using the Google-Scanned-Objects (GSO) dataset provided by MegaPose. 850 GSO objects with approximately 800,000 rendered images were used for training, while the remaining 94 objects and approximately 200,000 images were used for validation. To evaluate the method for unseen object pose estimation in this embodiment, five core BOP datasets were used as test sets: LM-O, T-LESS, TUD-L, IC-BIN, and YCB-V. The LM-O dataset was used for the ablation study in this embodiment. All CAD models and images in the test set were guaranteed to be unseen during training.

[0110] This embodiment uses the average recall (AR) from the standard benchmark BOP as its primary evaluation metric. It calculates the average recall across three pose errors by varying the threshold within a certain range. This embodiment also uses the average distance metric (ADD) as an auxiliary metric to provide a fair comparison with the baseline in Figure 8. ADD is configured to calculate the average point distance between the point clouds of the object's CAD model with the ground truth and estimated poses.

[0111] Figure 5 illustrates the qualitative results of 6D pose estimation using the method in this embodiment compared to Megapose and ZeroPose. As shown in Figure 5, the first row of images shows an egg carton that is severely occluded in the LM-O dataset. Experiments demonstrate that the method in this embodiment is robust to occlusion, while other methods incorrectly flip the pose. The second row of images shows a highly blurred object. Compared to other methods that invert the pose, the method in this embodiment achieves more accurate predictions.

[0112] FIG6 is a schematic diagram of the t-SNE visualization corresponding results of the object descriptor provided in an embodiment of the present application. As shown in FIG6 , the method in this embodiment is capable of capturing continuous and discrete symmetries without external annotations.

[0113] Figure 7 is a schematic diagram of the t-SNE visualization comparison results of regular tetrahedrons with different color faces provided in an embodiment of the present application. As shown in Figure 7, the method in this embodiment can extract different descriptors (right), while RoITr cannot (middle).

[0114] As shown in Figure 4, our method achieves state-of-the-art performance on the unseen object pose estimation task compared to all baselines. For fair comparison, our method is run against the baselines in five specific settings. In group (a), our method is compared to Zero-Pose, OSOP, and MegaPose, using only one hypothesis at inference and without any refinement. In group (a), our method performs significantly better and faster on average. Note that in group (a), a generic object detector is used to detect unseen objects, demonstrating our method's robustness to noisy detection initialization, a common occurrence in real-world applications. In group (b), our results are further improved by introducing more pose hypotheses (default: 20). Our method is compared to OSOP with the same number of hypotheses and demonstrates superior performance. In group (c), by adding ICP refinement, our model achieves the best results among all methods in both overall and per-dataset evaluations. Note that in panel (c), our method requires less time to perform an inference compared to other rendering-based refinement baselines. In panel (d), our method is compared with OVE6D and GCPose, visible object detectors. Our method consistently outperforms the baselines on the T-LESS and TUD-L datasets. In panel (e), our results are improved by using a trained detector combined with ICP refinement.

[0115] For example, as shown in Figure 5, the descriptor learned in this embodiment can capture both continuous and discrete symmetries, thanks to the rotation invariance designed by the method in this embodiment. Compared with GC-Pose, which relies on symmetry label supervision, the descriptor of this embodiment can recognize symmetry even without any external symmetry annotations. In addition, this embodiment visualizes "tetX" in the SYMSOL dataset. Due to geometric similarity, the descriptors extracted by RoITr show the same distribution on all 4 faces, introducing matching ambiguity and potentially leading to incorrect pose estimation. In contrast, the method in this embodiment extracts unique descriptors on textured faces. This shows that our method not only describes the geometric properties of the object, but also captures texture information, thereby further eliminating ambiguity in pose estimation.

[0116] Figure 8 is a schematic diagram of the results of a quantitative evaluation of 6D pose (ADD-0.1d) for the visible object pose estimation task on the LM-O dataset provided by an embodiment of the present application. Figure 9 is a schematic diagram of the results of an ablation study of a key design on the BOP-5 dataset provided by an embodiment of the present application.

[0117] To demonstrate the effectiveness of the RGB-D fusion mechanism proposed in this embodiment, MatchU was compared with recent RGB-D fusion methods on the LM-O dataset. As shown in Figure 3, although the method in this embodiment is specifically designed for unseen object pose estimation, it still outperforms PVN3D and FFB6D, and achieves comparable results to the current state-of-the-art method DFTr by increasing the number of hypotheses, which demonstrates the effectiveness of the RGB-D fusion mechanism in this embodiment. In addition, the method in this embodiment, trained without test objects, even outperforms most baselines that have seen them. This result further confirms the universality of the method in this embodiment.

[0118] Figure 10(a) shows AR with different numbers of hypotheses, and Figure 10(b) shows AR of the final pose and HR of pose hypotheses with different numbers of corresponding samples. As shown in the first row of Figure 9, performance drops significantly, demonstrating the effectiveness of BCM Loss in guiding descriptor learning. Secondly, by masking the input RGB image, performance drops sharply (row 2), demonstrating the effectiveness of RGB-D fusion and the key role RGB information plays in this embodiment's pipeline. This embodiment uses the original RoITr model and initializes point features with RGB values ​​to further demonstrate the superiority of RGB-D fusion. RoITr with RGB outperforms RoITr without RGB in this embodiment, but still significantly falls behind the full pipeline of this embodiment. The impact of the number of hypotheses. As shown in Figure 10(a), this embodiment's method benefits from increasing the number of hypotheses. However, performance saturates when the number exceeds 50, indicating that 50 hypotheses can cover the true pose in most cases. To balance computational cost and performance, this embodiment uses 20 as the default value.

[0119] In this embodiment, in order to investigate the quality of the pose hypothesis of this embodiment, the hit recall rate (HR) is defined as the ratio of the test set whose ground truth pose is included in the hypothesis proposed by this embodiment. Specifically, the first 128 correspondences are used as a sample pool, and a hypothesis is calculated by randomly selecting 3 correspondences from them. This process is repeated to generate multiple hypotheses. The comparison of HR and average recall rate (AR) is reported by changing the number of corresponding samples. As shown in Figure 10(b), the AR number is lower because it only considers the 1 pose hypothesis with the highest score. When considering all hypotheses in the evaluation, the method of this embodiment achieves a HR of more than 80%, which reflects the potential of the descriptor in this embodiment in generating correct poses.

[0120] In this embodiment, by estimating the pose of an unseen object from a single RGB-D image, rotation-invariant descriptors are first extracted from the 3D point cloud of the CAD model and the depth map. Then, multimodal fusion of texture and geometry is achieved through the potential fusion attention module. A bridging coarse-level matching loss is introduced to utilize the latent features of the RGB image to connect partial observations and descriptions of the complete object geometry. MatchU inherently captures object symmetry without the need for explicit annotations. MatchU significantly surpasses all existing unseen object pose estimation methods on standard benchmarks. Of course, it relies on external object positioning and may be affected by its erroneous results. In the future, incorporating such modules into the pipeline to build end-to-end training may further improve the results of this embodiment.

[0121] Figure 11 is a structural diagram of a device for estimating the 6D pose of an unseen object based on RGB-D images provided in an embodiment of the present application. The device 1100 for estimating the 6D pose of an unseen object based on RGB-D images in this embodiment may include: a processor 1101 and a memory 1102.

[0122] Memory 1102 is configured to store programs. Memory 1102 may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc. Memory may also include non-volatile memory, such as flash memory. Memory 1102 is configured to store computer programs (such as applications and functional modules that implement the above-mentioned methods), computer instructions, etc. These computer programs and computer instructions may be partitioned and stored in one or more memories 1102. Furthermore, these computer programs, computer instructions, data, etc. may be called by processor 1101.

[0123] The aforementioned computer programs, computer instructions, etc. may be partitioned and stored in one or more memories 1102 . Furthermore, the aforementioned computer programs, computer instructions, data, etc. may be called by the processor 1101 .

[0124] The processor 1101 is configured to execute the computer program stored in the memory 1102 to implement the various steps in the method involved in the above embodiment.

[0125] For details, please refer to the relevant description in the previous method embodiment.

[0126] The processor 1101 and the memory 1102 may be independent structures or integrated structures. When the processor 1101 and the memory 1102 are independent structures, the memory 1102 and the processor 1101 may be coupled via a bus 1103 .

[0127] The 6D pose estimation device 1100 for unseen objects based on RGB-D images in this embodiment can execute the technical solution in the method shown in Figure 2. Its specific implementation process and technical principles can be found in the relevant description of the method shown in Figure 2, which will not be repeated here.

[0128] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "platforms."

[0129] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When at least one processor of a user device executes the computer-executable instructions, the user device executes the various possible methods described above.

[0130] Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of computer programs from one location to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. Alternatively, the ASIC can be located in a user device. Of course, the processor and storage medium can also exist as discrete components in a communication device.

[0131] The present application also provides a program product, which includes a computer program stored in a readable storage medium. At least one processor of a server can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the server implements any method of the above-mentioned embodiments of the present disclosure.

[0132] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0133] FIG12 is a schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present disclosure. Referring to FIG12 , a program product 1200 configured to implement the above-described method according to an embodiment of the present disclosure is described. The program product 1200 may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0134] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0135] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program configured to be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0136] Program code configured to perform the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0137] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present disclosure. Various modifications to these embodiments will be obvious to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features disclosed herein.

[0138] The above describes the specific embodiments of the present disclosure. It should be understood that the present disclosure is not limited to the above specific embodiments, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the essence of the present disclosure.

Claims

1. A 6D pose estimation method for an unseen object based on an RGB-D image, comprising: The 3D model of the unseen object and the RGB-D image are fused to obtain fusion information; generating a universal color-aware rotation-invariant 3D descriptor based on the fused information; A matching relationship between a 3D model of an unseen object and an RGB-D image is established based on a universal color-aware rotation-invariant 3D descriptor, and a 6D pose of the unseen object is estimated.

2. The method for estimating 6D pose of an unseen object based on RGB-D images according to claim 1, wherein: The fusion process of the 3D model of the unseen object and the RGB-D image to obtain fusion information includes: The 3D model of the unseen object is represented as a point cloud P, the partial point cloud obtained by the depth channel of the RGB-D image is recorded as Q, and the RGB of the cropped local object is recorded as K; Extract 3D local features from point clouds P and Q, and extract 2D local features from the cropped RGB image K; Fusion of 3D local features and 2D local features in latent space where the superscript d represents the dimension of the latent space.

3. The 6D pose estimation method for unseen objects based on RGB-D images according to claim 2, wherein: Generating a universal color-perceived rotation-invariant 3D descriptor according to the fusion information includes: Extracting latent space The general descriptor Φ corresponding to the points in the point cloud P P , and the universal descriptor Φ corresponding to the points of the point cloud Q Q ; Calculate Φ P and Φ Q The similarity between them and construct the corresponding set in Respectively represent point p i and q j The matching descriptor of i represents the i-th point in the point cloud P, 1≤i≤n, n represents the total number of point clouds in the 3D model, q j Represents the jth point in the depth point cloud Q, 1≤j≤m, and m represents the number of point clouds sampled in the depth image.

4. The method for estimating 6D pose of an unseen object based on RGB-D images according to claim 3, wherein: A matching relationship between a 3D model of an unseen object and an RGB-D image is established based on a universal color-aware rotation-invariant 3D descriptor, and a 6D pose of the unseen object is estimated, including: According to the input point cloud P and point cloud Q, and the cropped RGB image K, the mapping function Ψ maps the point cloud P and point cloud Q to the general descriptor Φ P and Φ Q , and fuse the cross-modal information from (Q, K) and (P, K) respectively; Correspondences are established between unseen objects and RGB-D images by matching universal descriptors, and the 6D pose of unseen objects is estimated by optimizing pose parameters.

5. The method for estimating 6D pose of an unseen object based on RGB-D images according to claim 2, wherein: The extracting of 3D local features from point cloud P and point cloud Q includes: The transformer-based architecture RoITr builds the encoder backbone and samples the input point cloud P and point cloud Q into super points P′ and Q′ through the encoder’s farthest point sampling, where And Q′={q j ′∈ n′ and m′ represent the number of superpoints in P′ and Q′, respectively; For each superpoint p i ′ and q j ′, extract local geometric features from neighboring points within the radius r, so that a local rotation-invariant 3D descriptor can be obtained by projecting the local geometric features into the latent space.

6. The method for estimating 6D pose of an unseen object based on RGB-D images according to claim 5, wherein: The step of extracting 2D local features from the cropped RGB image K comprises: The improved encoder of FPN is used as the convolutional neural network CNN, the cropped RGB image K is sampled into a local feature map of a preset size, and the local texture information is projected into a d-dimensional latent space consistent with the 3D geometric features; where H represents the length of the image and W represents the width of the image; Flatten the local feature map of the image, where the 2D superpixel is represented as K′ and the 2D superpixel feature is represented as Φ K′ .

7. The method for estimating 6D pose of an unseen object based on RGB-D images according to claim 5, wherein: Fusion of 3D local features and 2D local features in latent space Including: The context of the extracted 3D local features and 2D local features are fused in the latent space through a latent fusion attention module; wherein the latent fusion attention module includes: a 3D-2D fusion block and a 2D-3D fusion block; The 3D-2D fusion block is specifically configured as follows: Aggregate Φ using global transformers P′ and Φ Q′ , so that the RGB features are fused with the global perception depth and 3D features in turn, and the cross-modal 2D features of each superpoint are obtained through the potential fusion transformer, where Φ P′ Φ P The corresponding super-point feature, Φ Q′ Φ Q The corresponding super-point features; The 2D-3D fusion block is specifically configured as follows: The latent fusion transformer uses RGB features to enhance Φ separately. P′ and Φ Q′ The super-point features are then injected into the enhanced super-point features through the global transformer to generate 3D super-point features containing 2D perception. and 8. The method for estimating 6D pose of an unseen object based on RGB-D images according to claim 4, wherein: We establish correspondences between unseen objects and RGB-D images by matching universal descriptors, and estimate the 6D pose of unseen objects by optimizing pose parameters, including: Establish 3D-3D matching between point clouds P and Q through the extracted features, and use covariance analysis to measure the similarity of normalized features; By determining the top κ most relevant putative 3D-3D matches of the corresponding point set C, creating η pose hypotheses from them; For each assumption 1≤ν≤η, we write Randomly select s corresponding relations from the set C, where s<<κ; The 6D pose hypothesis of the object is solved through the resulting correspondence set C, and all pose hypotheses are ranked according to the average score between the 3D and RGB verification process, where the 3D score is calculated based on the Euclidean point-to-point distance between the converted 3D model and the extracted depth map, and the RGB score is output through a deep extension network; the final prediction is the pose with the highest score 9. A system for estimating 6D pose of an unseen object based on RGB-D images, comprising: The fusion unit is used to fuse the 3D model of the unseen object and the RGB-D image to obtain fusion information; The descriptor generation unit is used to generate a universal color perception rotation invariant 3D descriptor according to the fusion information; The matching and posture estimation unit is used to establish a matching relationship between the 3D model of the unseen object and the RGB-D image according to the universal color-aware rotation-invariant 3D descriptor, and to estimate the 6D posture of the unseen object.

10. A 6D pose estimation device for an unseen object based on an RGB-D image, comprising: A processor and a memory, wherein the memory stores executable program instructions, and when the processor calls the program instructions in the memory, the processor is configured as follows: Execute the steps of the method for estimating 6D pose of an unseen object based on RGB-D images as described in any one of claims 1 to 8.

11. A computer-readable storage medium, wherein: The method is configured to store a program, which, when executed, implements the steps of the method for estimating the 6D pose of an unseen object based on an RGB-D image as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for estimating unknown object grabbing positions and posture on basis of mixed information input network models

    CN108280856A

  • A man-machine cooperation oriented real-time posture detection method for hand-held objects

    CN109255813A

  • Object six-degree-of-freedom pose estimation method based on color and depth information fusion

    CN111179324A

  • Class level 6D attitude estimation method based on monocular RGB-D image

    CN114863573A

  • Graph convolution-based double-flow shape prior adaptive category-level object attitude estimation method

    CN116958958A

Cited By

  • Multi-modal drainage wire clamp pose estimation method, system and device based on cross-modal comparative learning and medium

    CN120997300A

  • 6D pose estimation method and device fusing attention mechanism, equipment and medium

    CN121095347A

  • Dense point level prediction method and device for large-scale point cloud, medium and equipment

    CN121259367A

  • Intelligent monitoring method and system for forest wild animals based on multi-source data fusion

    CN121660505A