Object 6D attitude estimation method and system, electronic equipment and storage medium

Through the method of step-by-step pixel-to-pixel correspondence learning, the correspondence between RGB observation images and templates is optimized, and the problem of insufficient pose estimation accuracy in the prior art is solved, and a higher precision 6D pose prediction is achieved.

CN119941847AActive Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202411904461.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-06
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The existing 6D pose estimation method based on RGB images is difficult to achieve competitive performance, and is affected by the corresponding relationship noise and outliers, resulting in insufficient pose prediction.

Method used

A method of step-by-step pixel-to-pixel correspondence learning is proposed. By obtaining the template set and target observation image of the target object model, feature matching, affine transformation matrix regression and regression offset block learning is performed, the correspondence between the RGB observation image and the template is gradually optimized, and noise and outliers are filtered to obtain an accurate 6D pose.

Benefits of technology

Through step-by-step pixel-to-pixel corresponding learning, the accuracy of pose estimation is significantly improved, and the object's 6D pose can be predicted more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941847A_ABST
    Figure CN119941847A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an object 6D attitude estimation method and system, electronic equipment and a storage medium, and belongs to the technical field of computer vision. According to the method, in the first stage, an optimal matching template of a target observation image in a 3D template set is analyzed, and meanwhile, a rough corresponding relation between the optimal matching template and the target observation image is determined; in the second stage, an affine transformation matrix between the target observation image and the optimal matching template is determined according to the rough corresponding relation, and a smooth corresponding relation between the target observation image and the optimal matching template is determined based on the affine transformation matrix, so that corresponding relation noise and outliers are filtered out; in the third stage, the coordinate offset is determined according to a regression offset block between the target observation image and the feature map of the optimal matching template, the smooth corresponding relation is updated according to the coordinate offset, and the accurate corresponding relation is obtained; and determining the 6D attitude of the object according to the accurate corresponding relation. According to the method, the precision of attitude estimation is improved through gradual pixel-to-pixel corresponding learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a method, system, electronic device and storage medium for estimating 6D posture of an object. Background Art

[0002] The pose of an object is usually represented by six degrees of freedom (6DoF) parameters, including 3D rotation and 3D translation, to define the transformation from standard object space to camera space. Object pose estimation has attracted much attention in practical applications such as robotic manipulation and augmented reality, and has therefore been widely explored in research.

[0003] Early research on object pose estimation focused on using the same object CAD model for pose estimation during training and testing, but lacked flexibility for objects not seen in training. Subsequent research solved the pose estimation problem for unseen objects in known categories by defining a class-normalized object coordinate space, but difficulties still exist for completely new categories. With the advancement of basic models, more recent research has increasingly focused on handling completely new objects to achieve zero-shot 6D pose estimation, which poses a major challenge to the generalization ability of the model.

[0004] In the task of zero-shot pose estimation for new objects, related techniques have achieved remarkable performance using RGB-D image methods through template matching combined with pose updating or point cloud registration for pose calculation. The success of these methods is largely due to the geometric support provided by depth maps, which not only provide key features for matching, but also enhance the ability to locate objects in three-dimensional space through geometric priors. However, the high cost of depth sensors often limits their popularity in practical applications, making methods that rely solely on RGB images a more attractive option. Despite this, RGB-based methods remain understudied and generally fail to achieve competitive performance. For example, the 6D pose estimation methods of GigaPose and FoundPose, which establish the correspondence between the observed scene and the rendered template through simple feature matching, are often affected by correspondence noise and outliers, resulting in inaccurate pose predictions. Summary of the invention

[0005] The main purpose of the embodiments of the present application is to propose a method, system, electronic device and storage medium for estimating the 6D posture of an object, aiming to improve the accuracy of 6D posture prediction of an object in an image.

[0006] To achieve the above object, an embodiment of the present application provides a method for estimating a 6D pose of an object, comprising the following steps:

[0007] Obtain a template set of a target object model and a target observation image;

[0008] Perform feature matching between the target observation image and each template in the template set to determine the best matching template, and determine a rough corresponding relationship between the best matching template and the target observation image;

[0009] Determine an affine transformation matrix between the target observation image and the best matching template according to the rough correspondence relationship, and apply the affine transformation matrix to the best matching template for pixel alignment to determine a smooth correspondence relationship between the target observation image and the best matching template;

[0010] Determine a coordinate offset according to a regression offset block between the target observation image and the feature map of the best matching template, and update the smooth correspondence according to the coordinate offset to obtain an accurate correspondence between the target observation image and the best matching template;

[0011] The 6D posture of the object in the target observation image is determined according to the precise corresponding relationship.

[0012] In some embodiments, the template set is obtained by the following steps:

[0013] Get the target object model;

[0014] Rendering the target object model respectively using a plurality of desired viewing angles to obtain a template set;

[0015] The target observation image is obtained by the following steps:

[0016] Get scene image;

[0017] A zero-sample segmentation technique is used to identify a target object in the scene image, and a detection frame of the target object is cut out from the scene image to obtain a target observation image.

[0018] In some embodiments, the step of performing feature matching between the target observation image and each template in the template set to determine the best matching template comprises the following steps:

[0019] Performing feature extraction on the target observation image and each template in the template set respectively to obtain a first feature of the target observation image and a second feature of each template;

[0020] The matching scores between each template and the target observation image are determined by the similarity between the first feature and the second feature, and the best matching template is determined according to the matching scores of each template.

[0021] In some embodiments, determining the affine transformation matrix between the target observation image and the best matching template according to the rough correspondence relationship comprises the following steps:

[0022] Encoding the rough corresponding relationship to obtain a corresponding relationship graph;

[0023] Regressing the affine transformation relationship between the target observation image and the best matching template according to the corresponding relationship diagram to obtain an affine transformation matrix;

[0024] Among them, the affine transformation matrix is ​​expressed as follows

[0025]

[0026] Among them, α represents the rotation angle in the plane, and s represents and The relative ratio between u ,t v ) represents the horizontal and vertical two-dimensional translation relative to the center of mass of the object in the image.

[0027] In some embodiments, applying the affine transformation matrix to the best matching template for pixel alignment to determine a smooth correspondence between the target observation image and the best matching template comprises the following steps:

[0028] Determine, according to the affine transformation matrix and the pixel position of the target observation image, the corresponding position of the pixel position on the best matching template;

[0029] According to the corresponding position of each pixel position of the target observation image on the best matching template, a smooth corresponding relationship between the target observation image and the best matching template is obtained.

[0030] In some embodiments, the determining of the coordinate offset according to the regression offset block between the target observation image and the feature map of the best matching template comprises the following steps:

[0031] Inputting the target observation image and the best matching template into a dense prediction transformer respectively, to obtain a plurality of first hierarchical feature maps of the target observation image and a plurality of second hierarchical feature maps of the best matching template;

[0032] Determine a regression offset block of the corresponding layer according to the first hierarchical feature map and the second hierarchical feature map of the corresponding layer;

[0033] Scaling and adjusting the two-dimensional position of the first hierarchical feature map in the regression offset block, and aligning and adjusting the second hierarchical feature map with the adjusted first hierarchical feature map to obtain a third hierarchical feature map;

[0034] Performing a correlation search on the first hierarchical feature map and the second hierarchical feature map in the regression offset block to obtain a correlation feature map;

[0035] The coordinate offset of each pixel is determined according to the first hierarchical feature map, the third hierarchical feature map and the correlation feature map of different layers.

[0036] In some embodiments, determining the 6D posture of the object in the target observation image according to the precise correspondence includes the following steps:

[0037] Get the confidence of the pixel coordinate offset;

[0038] When the confidence of the coordinate offset of the pixel is greater than the expected threshold, the corresponding position of the pixel in the target observation image in the best matching template is determined according to the precise corresponding relationship, and a two-dimensional and three-dimensional associated point pair set is determined according to the corresponding position of each pixel;

[0039] The 6D posture of the object in the target observation image is determined according to the set of associated point pairs.

[0040] To achieve the above object, another aspect of the embodiment of the present application provides a 6D pose estimation system for an object, comprising:

[0041] The first module is used to obtain a template set of a target object model and a target observation image;

[0042] The second module is used to perform feature matching between the target observation image and each template in the template set to determine the best matching template, and determine a rough corresponding relationship between the best matching template and the target observation image;

[0043] A third module is used to determine an affine transformation matrix between the target observation image and the best matching template according to the rough corresponding relationship, and apply the affine transformation matrix to the best matching template for pixel alignment to determine a smooth corresponding relationship between the target observation image and the best matching template;

[0044] A fourth module is used to determine a coordinate offset according to a regression offset block between the target observation image and the feature map of the best matching template, and to update the smooth correspondence according to the coordinate offset to obtain an accurate correspondence between the target observation image and the best matching template;

[0045] The fifth module is used to determine the 6D posture of the object in the target observation image according to the precise corresponding relationship.

[0046] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application proposes an electronic device, which includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory, and the program implements the method described in the above embodiment when executed by the processor.

[0047] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium used for computer-readable storage, and the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in the above embodiment.

[0048] The object 6D posture estimation method, system, electronic device and storage medium proposed in the present application analyze the best matching template with the target observation image in the 3D template set in the first stage, and determine the rough correspondence between the best matching template and the target observation image; in the second stage, the affine transformation matrix between the target observation image and the best matching template is determined according to the rough correspondence, and the affine transformation matrix is ​​applied to the best matching template for pixel alignment to determine the smooth correspondence between the target observation image and the best matching template, thereby filtering out the correspondence noise and outliers; in the third stage, the coordinate offset is determined according to the regression offset block between the feature map of the target observation image and the best matching template, and the smooth correspondence is updated according to the coordinate offset to obtain an accurate precise correspondence; finally, the 6D posture of the object in the target observation image is determined according to the precise correspondence. The present application improves the accuracy of posture estimation through step-by-step pixel-to-pixel correspondence learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flow chart of a method for estimating a 6D posture of an object provided in an embodiment of the present application;

[0050] Figure 2 It is a schematic diagram of the PicoPose learning framework provided in an embodiment of the present application;

[0051] Figure 3 It is a schematic diagram of the process of three stages of pixel correspondence learning provided by an embodiment of the present application;

[0052] Figure 4 It is a feature correspondence diagram between the target observation image and the template under the affine transformation provided in the embodiment of the present application;

[0053] Figure 5 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] It should be noted that, although the functional modules are divided in the system and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0057] First, the terms involved in the embodiments of the present application are explained as follows:

[0058] DINOv2, an advanced self-supervised model for computer vision released by Meta. DINOv2 is a new method for training high-performance computer vision models that uses self-supervised learning to learn from any collection of images and produce powerful visual representations. These representations perform well on downstream tasks such as classification, segmentation, image retrieval, and depth estimation without fine-tuning.

[0059] ViT-L is a version of the Vision Transformer (ViT) model, which stands for Large model. It extracts global features of the image by dividing the image into a series of small patches and inputting these patches into the Transformer model for processing. This processing method makes ViT-L perform well in computer vision tasks such as image classification, object detection, and semantic segmentation.

[0060] Dense Prediction Transformer (DPT), an architecture that uses visual transformers instead of convolutional networks as the backbone for dense prediction tasks. Dense prediction tasks refer to predicting each pixel or image patch in an image and assigning corresponding labels to them. Such tasks usually require the model to capture detailed information in the image and accurately predict the category, depth, edge and other attributes of each pixel or image patch.

[0061] The framework of the method of the embodiment of the present application is generally described. The embodiment of the present application proposes a new step-by-step pixel-to-pixel correspondence learning framework, which is called PicoPose, to achieve accurate new object pose estimation based on RGB images. PicoPose gradually optimizes the correspondence between the RGB observation image and the template in three stages, greatly improving the accuracy of the object pose calculated through these correspondences. Specifically, Figure 2 As shown in the figure, given an RGB image (Scene image) containing a complex scene and an object CAD model (CAD model), the object is first rendered from multiple perspectives of the CAD model to obtain multiple templates (Templates). These templates are then combined with the Zero-Shot Segmentation technique to detect the target object in the RGB scene and obtain the target observation image. Next, PicoPose adopts a three-stage learning process of progressive pixel-to-pixel correspondence learning, that is, the best matching template of the detected object is determined through template mapping (Template Matching), and fine-grained pixel-to-pixel correspondence is learned. Since each pixel in the template corresponds to a three-dimensional surface point on the CAD model, a pairing is established between the two-dimensional position in the observed image and the corresponding three-dimensional point in the template, and then the 6D pose (Object pose) of the image object is calculated through algorithms such as PnP / RANSAC. The specific description of the three stages of progressive pixel-to-pixel correspondence learning is as follows:.

[0062] The first stage is feature matching to obtain a rough correspondence (Feature Matching for Coarse Correspondence). In this stage, PicoPose can use the visual transformer backbone network to extract the features of the RGB target observation image and the rendered template, identify the best matching template, and obtain a rough correspondence between the template and the observation image.

[0063] The second stage is global transformation estimation to optimize smooth correspondence (Trans formation Estimation for Smooth Correspondence). In this stage, PicoPose represents the rough correspondence as a correspondence graph and regresses the two-dimensional affine transformation (including in-plane rotation, scaling, and two-dimensional translation) to smooth the rough correspondence to obtain a smooth correspondence to filter outliers in the first stage.

[0064] The third stage is local refinement for fine correspondence. In this stage, PicoPose applies affine transformation to the feature map of the best matching template and learns the corresponding offsets in the local area through multiple offset regression modules, gradually achieving fine pixel-to-pixel correspondence.

[0065] Based on the above learning framework, the embodiments of the present application provide a method, system, electronic device and storage medium for estimating the 6D posture of an object, aiming to improve the accuracy of 6D posture prediction of an object in an image.

[0066] The object 6D posture estimation method, system, electronic device and storage medium provided in the embodiments of the present application are specifically explained through the following embodiments. First, the object 6D posture estimation method in the embodiments of the present application is described.

[0067] The object 6D posture estimation method provided in the embodiment of the present application relates to the field of computer vision technology. The object 6D posture estimation method provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the object 6D posture estimation method, etc., but is not limited to the above forms.

[0068] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0069] Figure 1 is an optional flow chart of the object 6D posture estimation method provided in the embodiment of the present application, Figure 1 The method may include but is not limited to steps S101 to S105.

[0070] Step S101, obtaining a template set of a target object model and a target observation image;

[0071] Step S102, performing feature matching between the target observation image and each template in the template set to determine the best matching template, and determining a rough correspondence between the best matching template and the target observation image;

[0072] Step S103, determining an affine transformation matrix between the target observation image and the best matching template according to the rough correspondence, and applying the affine transformation matrix to the best matching template for pixel alignment to determine a smooth correspondence between the target observation image and the best matching template;

[0073] Step S104, determining a coordinate offset according to a regression offset block between the target observation image and the feature map of the best matching template, and updating a smooth correspondence according to the coordinate offset to obtain an accurate correspondence between the target observation image and the best matching template;

[0074] Step S105 , determining the 6D posture of the object in the target observation image according to the precise correspondence relationship.

[0075] In step S101 of some embodiments, the target object model is a three-dimensional model of the target object, specifically a three-dimensional CAD model, wherein the object posture of the target object model is represented by six degrees of freedom parameters, including three-dimensional rotation and three-dimensional translation. The target observation image refers to an RGB image of the target object.

[0076] In some embodiments, in step S101, the template set is obtained by the following steps:

[0077] Step S201, obtaining a target object model;

[0078] Step S202, using multiple desired viewing angles to render the target object model respectively to obtain a template set;

[0079] In some embodiments, in step S101, the target observation image is obtained by the following steps:

[0080] Step S301, acquiring a scene image;

[0081] Step S301 , using zero-sample segmentation technology to identify a target object in a scene image, and cutting out a detection frame of the target object from the scene image to obtain a target observation image.

[0082] For example, for a given clustered scene RGB image and a CAD model of a target object, the object template is first rendered from each viewpoint of the CAD model to obtain the templates of the target object under each viewpoint, and each template forms a template set. Then, the zero-sample segmentation technique is used to identify and crop the area containing the target object from the scene image (i.e., the RGB image) to obtain a target observation image. Furthermore, the size of the target observation image and the target object template is adjusted to a fixed size H×W, and the target observation image is represented as The template collection is represented as Where N is the number of templates. In the subsequent process, PicoPose can use the target observation image and the template set as input to search for the best matching template (expressed as ), and gradually learn in three stages and The pixel-to-pixel correspondence between them. Each foreground pixel on the CAD model corresponds to a 3D surface point on the CAD model, enabling Create 2D point pairs on The corresponding 3D point pairs are established, and then the 6D object posture can be calculated by using the above point pairs.

[0083] In step S102 of some embodiments, in the first stage of PicoPose, a visual transformer backbone network can be used to extract features of the target observation image and each rendering template, and the best matching template can be identified by feature matching, thereby obtaining a rough correspondence between the template and the target observation image with respect to pixels.

[0084] In some embodiments, in step S102, the step of performing feature matching between the target observation image and each template in the template set to determine the best matching template may include, but is not limited to, steps S401 to S402:

[0085] Step S401, performing feature extraction on the target observation image and each template in the template set respectively to obtain a first feature of the target observation image and a second feature of each template;

[0086] Step S402: determining the matching scores between each template and the target observation image according to the similarity between the first feature and the second feature, and determining the best matching template according to the matching scores of each template.

[0087] For example, please refer to Figure 3 In the first stage (Stage 1), for the resized target object image (RGB observation) and object templates from various viewpoints of the CAD model The ViT-L backbone network pre-trained by DINOv2 is used to extract the features of the target observation image and each object template respectively, and the corresponding feature representation is obtained as follows: and Where D is the feature dimension and M is the number of patches. It is understandable that other network models for image feature extraction can be used for feature extraction, and the embodiments of the present application do not impose specific restrictions. Then, feature matching is performed based on the extracted features to identify the best matching template, and at the same time, a rough correspondence between the target observation image and the best matching template is determined. Feature matching can be achieved by calculating the similarity between all features of the target observation image and the object template, and establishing a rough correspondence through feature similarity. Specifically, by matching each template Observation image with target The similarity between the features is scored to retrieve the best matching template For each template By comparing the foreground pixels of the observed image (determined by the above zero-sample segmentation step) with The maximum feature cosine similarity of the pixels in the template is averaged to obtain the template matching score c. i , from the template collection In the template matching, select the template with the highest score As the best matching template, a rough pixel-to-pixel correspondence of obvious features between the best matching template and the target observation image is established. In addition, and The feature similarity between Each pixel in The most similar pixels in the image are used to provide a rough correspondence.

[0088] In step S103 of some embodiments, in the first stage of PicoPose, feature matching is used to obtain and There is a rough and sparse correspondence between , however, these correspondences are Figure 3 As shown, it usually shows a messy distribution, containing noise and many outliers. Based on this, in the second stage of PicoPose, the two-dimensional affine transformation (including in-plane rotation, scaling and two-dimensional translation) relationship between the target observation image and the best matching template is regressed according to the rough correspondence relationship to obtain the affine transformation matrix, and the affine transformation matrix is ​​applied to the best matching template for pixel alignment to determine the smooth correspondence between the target observation image and the best matching template, thereby smoothing the rough correspondence and filtering outliers.

[0089] In some embodiments, in step S103, the step of determining the affine transformation matrix between the target observation image and the best matching template according to the rough correspondence relationship may include, but is not limited to, steps S501 to S502:

[0090] Step S501, encoding the rough correspondence relationship to obtain a correspondence relationship graph;

[0091] Step S502, regressing the affine transformation relationship between the target observation image and the best matching template according to the correspondence relationship graph to obtain an affine transformation matrix.

[0092] For example, please refer to Figure 3 The second stage (Stage 2) part takes into account and The rough correspondence between them can effectively capture the changes in image rotation angle, scaling, and translation, so that the affine transformation relationship can be learned (the affine transformation relationship uses the affine transformation matrix ) basic mode, reducing the difficulty of related learning. Therefore, this embodiment does not directly connect the related features and Come back Instead, the rough correspondence obtained in the first stage is used for regression. Specifically, the rough correspondence in the first stage is encoded into a correspondence map to meet the input format of the subsequent network. Then, the correspondence map is input into the constructed multi-layer perception mechanism (MLP) based network to perform in-plane rotation (In-plane rotation) angle α, scale (Scale) s and translation (Translation) (t u ,t v ) is regressed to obtain the affine transformation matrix 2D affine transformation matrix It can be parameterized using four degrees of freedom (DoF) as follows:

[0093]

[0094] Among them, α represents the rotation angle in the plane, and s represents and The relative ratio between u ,t v ) represents the 2D translation of the object’s center of mass in the two images. By applying To change the template Helps with Pixel alignment is performed to achieve smooth correspondence. Various affine transformations (including in-plane rotation, scaling, and 2D translation) visualize the correspondence between the features of the RGB observation points (marked by stars) and the template features. Figure 4 In addition, by and The viewpoint rotation combined with the CNN can give a 6D object pose.

[0095] In some embodiments, in step S103, the step of applying the affine transformation matrix to the best matching template for pixel alignment and determining a smooth correspondence between the target observation image and the best matching template may include but is not limited to steps S601 to S602:

[0096] Step S601, determining the corresponding position of the pixel position on the best matching template according to the affine transformation matrix and the pixel position of the target observation image;

[0097] Step S602: Obtain a smooth correspondence between the target observation image and the best matching template according to the corresponding position of each pixel position of the target observation image on the best matching template.

[0098] For example, please refer to Figure 3 The second stage, for For any pixel position (u,v) on get The corresponding position (u′, v′) on is as follows:

[0099]

[0100] The deviation between (u,v) and (u′,v′) can be interpreted as the common “optical flow”. All pixels in The corresponding position in is expressed as To represent the smooth correspondence. As an index to collect features, to achieve Affine transformations applied to the feature map, including rotation, scaling, and translation, are generated to produce Aligned transformed feature maps.

[0101] In step S104 of some embodiments, in the third stage of PicoPose, the affine transformation matrix is ​​applied to multiple hierarchical feature maps of the best matching template to obtain an offset regression block, and the corresponding offsets in the local area are learned through multiple offset regression modules. To smooth the correspondence Updated to Implement fine-grained correspondence adjustment in local areas and gradually achieve refined pixel-to-pixel correspondence.

[0102] In some embodiments, in step S104, the step of determining the coordinate offset according to the regression offset block between the target observation image and the feature map of the best matching template may include but is not limited to steps S701 to S705:

[0103] Step S701, inputting the target observation image and the best matching template into a dense prediction transformer respectively, to obtain a plurality of first hierarchical feature maps of the target observation image and a plurality of second hierarchical feature maps of the best matching template;

[0104] Step S702, determining a regression offset block of the corresponding layer according to the first hierarchical feature map and the second hierarchical feature map of the corresponding layer;

[0105] Step S703, scaling and adjusting the two-dimensional position of the first hierarchical feature map in the regression offset block, and aligning and adjusting the second hierarchical feature map with the adjusted first hierarchical feature map to obtain a third hierarchical feature map;

[0106] Step S704, performing a correlation search on the first hierarchical feature map and the second hierarchical feature map in the regression offset block to obtain a correlation feature map;

[0107] Step S705, determining the coordinate offset of each pixel according to the first hierarchical feature map, the third hierarchical feature map and the correlation feature map of different layers.

[0108] For example, please refer to Figure 3 In the third stage (Stage 3), the dense prediction transformer (DPT) is first applied to the best matching template and target observation image Generate separately The L first layer feature maps and The L second layer feature maps Among them, H l ×W l Indicates the size of the space, D l It is the first th Each first-layer feature map and the corresponding second-layer feature map form an offset regression block (Offset Regression Block), and L offset regression blocks are used to iteratively update

[0109] For the first th offset regression blocks, currently The size is adjusted to H l ×W l× 2 and scaled by dividing each 2D position of the first layer feature map by the image size to ensure spatial consistency. The resized and scaled version is denoted as Use it as an index to retrieve feature maps from the second layer Collect features in order to obtain the converted third-layer feature map With Then, the correlation lookup module is used to analyze the correlation between the first layer feature and the second layer feature to obtain the correlation feature map. to make the correlation explicit and to make it easier to learn the offset. and Connected together to form the input of two stacked convolution sequences, which are used to regress the coordinate offsets of each feature map output by the mapping. It is also possible to output a certainty map of each predicted coordinate offset. Then Interpolate to size H×W×2, scale by multiplying the 2D coordinates by the image size, and add to Update. Deterministic Graph represents the confidence of the regression offset and is upsampled to produce This embodiment uses L offset regression blocks, which can be updated to obtain a more refined Its deterministic mapping is

[0110] In step S105 of some embodiments, since each pixel in the template corresponds to a three-dimensional surface point on the CAD model, after an accurate correspondence is established between the two-dimensional position in the target observation image and the corresponding three-dimensional point in the template, the 6D posture of the object can be calculated by algorithms such as PnP / RANSAC.

[0111] In some embodiments, step S105 may include but is not limited to steps S801 to S803:

[0112] Step S801, obtaining the confidence of the pixel coordinate offset;

[0113] Step S802, when the confidence of the pixel coordinate offset is greater than the expected threshold, the corresponding position of the pixel in the target observation image in the best matching template is determined according to the precise corresponding relationship, and a two-dimensional and three-dimensional associated point pair set is determined according to the corresponding position of each pixel;

[0114] Step S803: determining the 6D posture of the object in the target observation image according to the set of associated point pairs.

[0115] In this embodiment, for For each foreground pixel in , if its corresponding certainty (i.e., confidence) exceeds 0.5, the precise Find the location in The corresponding pixel in is associated with the surface point of the 3D object. Therefore, all pixel-to-pixel correspondences generate relevant 2D-3D associated point pairs, and the final object pose is calculated based on the associated point pair set through the PnP / RANSAC algorithm.

[0116] According to some embodiments of the present application, the beneficial effects of the object 6D posture estimation method of the embodiments of the present application are described below in combination with experimental results.

[0117] (1) Experimental setup

[0118] The model is trained on the ShapeNet-Objects and Google-Scanned-Objects synthetic datasets provided in the MegaPose paper, using a total of 2,000,000 training images. The metric evaluation is performed on seven core datasets of the BOP benchmark, including LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HB, and YCB-V. The standard BOP evaluation protocol is used to report the average recall (AR) relative to three error functions, namely visible surface difference (VSD), maximum symmetry-aware surface distance (MSSD), and maximum symmetry-aware projection distance (MSPD).

[0119] (2) Comparison with existing methods

[0120] The PicoPose learning framework method proposed in this embodiment is evaluated against existing methods on seven core datasets of the BOP benchmark. The robustness of PicoPose is enhanced by using the first five templates from the second and third stages to learn fine correspondences and selecting the poses that best match these correspondences. The quantitative results are shown in Table 1, where PicoPose significantly outperforms other methods, highlighting its superior zero-sample capability for novel object pose estimation through progressive correspondence learning. For example, a single PicoPose model using the first five templates achieves 19.6% and 9.2% higher AR than GigaPose and FoundPose single models, respectively. In Table 1, the results under the iterative refinement proposed by MegaPose are also reported, where PicoPose consistently outperforms other methods.

[0121] Table 1 Pose estimation results of different methods on seven core datasets of the BOP benchmark

[0122]

[0123] (3) Ablation experiment

[0124] Effectiveness of Progressive Correspondence Learning. PicoPose crucially lies in its design of progressive pixel-to-pixel correspondence learning. To evaluate this, we first examine the impact of improving the quality of correspondences by testing the results obtained from the respective correspondences of the three stages using PnP / RANSAC. As shown in Table 2, pose accuracy improves with finer correspondences, supporting our core claim. Results from the second stage improve accuracy by an average of 5.7% on the LM-O, T-LESS, and YCB-V datasets by smoothing the correspondences from the first stage, while the third stage further improves accuracy on all three datasets by locally refining the correspondences from the second stage, with an AR improvement of 14.7%.

[0125] Table 2 Results of the three stages on three different datasets

[0126]

[0127] Effect of the number of templates. Following the settings in the GigaPose paper, N = 162 templates are used for each object in the evaluation experiments. In Table 3, additional quantitative results of GigaPose and PicoPose proposed in this application with different numbers of templates are shown. The results show that the effect improves with the increase in the number of templates, because both methods rely on template matching to select the best matching template for the target object. However, the improvement rate slows down as the number of templates increases. It is worth noting that PicoPose is more effective than GigaPose when fewer templates are used, which further highlights the advantage of PicoPose.

[0128] Table 3 Quantitative comparison of different template numbers on three different datasets

[0129]

[0130] Effectiveness of stacked offset regression blocks in the third stage. In the third stage, L offset regression blocks are used. Specifically, L = 3 offset regression blocks are used in the experiments, and the feature space size in DPT is set to 16×16, 32×32, and 64×64 to refine the correspondence within the local region. Table 4 reports the results of different offset regression blocks. As shown in Table 4, the performance gradually improves as more blocks are used, indicating that finer correspondences can be achieved through gradual local refinement.

[0131] Table 4 Quantitative results of different offset regression blocks (denoted as “OR Block”) in the third stage

[0132]

[0133] The embodiment of the present application also provides a system for estimating a 6D posture of an object, comprising:

[0134] The first module is used to obtain a template set of a target object model and a target observation image;

[0135] The second module is used to perform feature matching between the target observation image and each template in the template set to determine the best matching template, and to determine a rough correspondence between the best matching template and the target observation image;

[0136] The third module is used to determine the affine transformation matrix between the target observation image and the best matching template according to the rough correspondence relationship, and apply the affine transformation matrix to the best matching template for pixel alignment to determine the smooth correspondence relationship between the target observation image and the best matching template;

[0137] The fourth module is used to determine the coordinate offset according to the regression offset block between the target observation image and the feature map of the best matching template, and update the smooth correspondence according to the coordinate offset to obtain the accurate correspondence between the target observation image and the best matching template;

[0138] The fifth module is used to determine the 6D posture of the object in the target observation image based on the precise correspondence.

[0139] It can be understood that the contents of the above-mentioned object 6D posture estimation method embodiment are all applicable to the present system embodiment, the functions specifically implemented by the present system embodiment are the same as those in the above-mentioned object 6D posture estimation method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned object 6D posture estimation method embodiment.

[0140] The embodiment of the present application also provides an electronic device, the electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory, and the program is executed by the processor to realize the above-mentioned object 6D posture estimation method. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0141] See also Figure 5 , Figure 5 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:

[0142] The processor 501 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0143] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 502, and the processor 501 calls and executes the object 6D posture estimation method of the embodiment of this application;

[0144] Input / output interface 503, used to implement information input and output;

[0145] Communication interface 504, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);

[0146] A bus 505 that transmits information between the various components of the device (e.g., the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);

[0147] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via the bus 505 .

[0148] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium used for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned object 6D posture estimation method.

[0149] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0151] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0152] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0153] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0154] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0155] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0156] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.

[0157] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0160] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A method for estimating 6D pose of an object, characterized in that: The following steps are involved: Obtain a template set of a target object model and a target observation image; Perform feature matching between the target observation image and each template in the template set to determine the best matching template, and determine a rough corresponding relationship between the best matching template and the target observation image; Determine an affine transformation matrix between the target observation image and the best matching template according to the rough correspondence relationship, and apply the affine transformation matrix to the best matching template for pixel alignment to determine a smooth correspondence relationship between the target observation image and the best matching template; Determine a coordinate offset according to a regression offset block between the target observation image and the feature map of the best matching template, and update the smooth correspondence according to the coordinate offset to obtain an accurate correspondence between the target observation image and the best matching template; The 6D posture of the object in the target observation image is determined according to the precise corresponding relationship.

2. The object 6D posture estimation method according to claim 1, characterized in that: The template set is obtained by the following steps: Get the target object model; Rendering the target object model respectively using a plurality of desired viewing angles to obtain a template set; The target observation image is obtained by the following steps: Get scene image; A zero-sample segmentation technique is used to identify a target object in the scene image, and a detection frame of the target object is cut out from the scene image to obtain a target observation image.

3. The object 6D posture estimation method according to claim 1, characterized in that: The step of performing feature matching between the target observation image and each template in the template set to determine the best matching template comprises the following steps: Performing feature extraction on the target observation image and each template in the template set respectively to obtain a first feature of the target observation image and a second feature of each template; The matching scores between each template and the target observation image are determined by the similarity between the first feature and the second feature, and the best matching template is determined according to the matching scores of each template.

4. The object 6D posture estimation method according to claim 1, characterized in that: Determining the affine transformation matrix between the target observation image and the best matching template according to the rough corresponding relationship comprises the following steps: Encoding the rough corresponding relationship to obtain a corresponding relationship graph; Regressing the affine transformation relationship between the target observation image and the best matching template according to the corresponding relationship diagram to obtain an affine transformation matrix; Among them, the affine transformation matrix is ​​expressed as follows Among them, α represents the rotation angle in the plane, and s represents and The relative ratio between u ,t v ) represents the horizontal and vertical two-dimensional translation relative to the center of mass of the object in the image.

5. The object 6D posture estimation method according to claim 4, characterized in that: The applying the affine transformation matrix to the best matching template to perform pixel alignment and determining a smooth correspondence between the target observation image and the best matching template comprises the following steps: Determine, according to the affine transformation matrix and the pixel position of the target observation image, the corresponding position of the pixel position on the best matching template; According to the corresponding position of each pixel position of the target observation image on the best matching template, a smooth corresponding relationship between the target observation image and the best matching template is obtained.

6. The object 6D posture estimation method according to claim 1, characterized in that: The step of determining the coordinate offset according to the regression offset block between the target observation image and the feature map of the best matching template comprises the following steps: Inputting the target observation image and the best matching template into a dense prediction transformer respectively to obtain a plurality of first hierarchical feature maps of the target observation image and a plurality of second hierarchical feature maps of the best matching template; Determine a regression offset block of the corresponding layer according to the first hierarchical feature map and the second hierarchical feature map of the corresponding layer; Scaling and adjusting the two-dimensional position of the first hierarchical feature map in the regression offset block, and aligning and adjusting the second hierarchical feature map with the adjusted first hierarchical feature map to obtain a third hierarchical feature map; Performing a correlation search on the first hierarchical feature map and the second hierarchical feature map in the regression offset block to obtain a correlation feature map; The coordinate offset of each pixel is determined according to the first hierarchical feature map, the third hierarchical feature map and the correlation feature map of different layers.

7. The object 6D posture estimation method according to claim 6, characterized in that: Determining the 6D posture of the object in the target observation image according to the precise corresponding relationship comprises the following steps: Get the confidence of the pixel coordinate offset; When the confidence of the coordinate offset of the pixel is greater than the expected threshold, the corresponding position of the pixel in the target observation image in the best matching template is determined according to the precise corresponding relationship, and a two-dimensional and three-dimensional associated point pair set is determined according to the corresponding position of each pixel; The 6D posture of the object in the target observation image is determined according to the set of associated point pairs.

8. A 6D pose estimation system for an object, characterized in that: include: The first module is used to obtain a template set of a target object model and a target observation image; The second module is used to perform feature matching between the target observation image and each template in the template set to determine the best matching template, and determine a rough corresponding relationship between the best matching template and the target observation image; A third module is used to determine an affine transformation matrix between the target observation image and the best matching template according to the rough corresponding relationship, and apply the affine transformation matrix to the best matching template for pixel alignment to determine a smooth corresponding relationship between the target observation image and the best matching template; A fourth module is used to determine a coordinate offset according to a regression offset block between the target observation image and the feature map of the best matching template, and to update the smooth correspondence according to the coordinate offset to obtain an accurate correspondence between the target observation image and the best matching template; The fifth module is used to determine the 6D posture of the object in the target observation image according to the precise corresponding relationship.

9. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the method described in any one of claims 1 to 7 are realized.

10. A storage medium, the storage medium being a computer-readable storage medium, used for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Object 6D attitude estimation method and device based on 2D detection and computer equipment

    CN112614184A

  • Object six-degree-of-freedom attitude estimation method, system and device and medium

    CN114821125A

  • Method executed by electronic equipment, electronic equipment and computer readable storage medium

    CN118864588A

  • System and method for model-free, one-shot object pose estimation via coordinate regression

    US20240404104A1

Cited By

  • Object attitude estimation method and system based on three-dimensional feature mapping and visual angle aggregation

    CN121354087A

  • Object pose estimation method and system based on three-dimensional feature mapping and perspective aggregation

    CN121354087B