Relative attitude estimation method and system based on structure perception corresponding relation learning

Through the corresponding relationship learning method based on structure perception, image features are extracted and processed, key points are obtained and improved to 3D space, 3D correspondence relationships are established and the optimal rotation matrix is ​​solved, and the limitations of the relative pose estimation method in the prior art are solved when dealing with objects with significant viewing angle changes, and relative pose estimation with high precision and low computing overhead is achieved.

CN119991818AInactive Publication Date: 2025-05-13DEEP SPACE EXPLORATION LABORATORY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510473206.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing relative pose estimation methods have limitations when dealing with objects with significant viewing angle changes, including the difficulty of establishing reliable correspondences in the 2D correspondence method, assuming that the verification method cannot effectively model the continuous pose space, and the calculation overhead based on the 3D correspondence method is large and unreliable.

Method used

Using a corresponding relationship learning method based on structure perception, feature maps are extracted through a pre-trained backbone network, and processed through a two-step attention mechanism, and fine feature extraction is performed in combination with object mask. The structured perceptual key point extraction module is used to obtain the coordinates and features of the key point space, and the key points are upgraded to 3D space through the structured correspondence estimation module to establish a 3D correspondence relationship, and finally, the weighted singular value decomposition method is used to solve the optimal rotation matrix.

Benefits of technology

It effectively avoids the problems of unreliability in matching and large calculation overhead in traditional methods, and improves the processing ability of large changes in object shape and the accuracy of corresponding relationship estimation under large perspective changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991818A_ABST
    Figure CN119991818A_ABST
Patent Text Reader

Abstract

The invention discloses a relative attitude estimation method and system based on structure perception corresponding relation learning, and belongs to the field of computer vision. The method comprises the steps that a query image and a reference image are collected, feature maps are extracted from the query image and the reference image through a pre-trained backbone network, and updated feature maps are obtained through two-step attention mechanism processing; multiplying with an object mask element by element to obtain fine feature maps of the query image and the reference image; obtaining key point space coordinates and corresponding key point features from the fine feature map by using a structured sensing key point extraction module; lifting the key points to a 3D space in a query coordinate system by using a structured correspondence estimation module, regressing corresponding 3D coordinates of the key points in a reference coordinate system, and establishing a 3D correspondence for relative attitude estimation; and based on the established 3D corresponding relation, solving an optimal rotation matrix by adopting a weighted singular value decomposition method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a relative posture estimation method and system based on structure-aware correspondence relationship learning. Background Art

[0002] Object pose estimation aims to estimate the 3D translation and rotation information of an object from a single image, and is widely used in augmented reality, robotics, autonomous driving and other fields. Early methods focused on instance-level pose estimation, which can only be trained and predicted for specific objects but cannot be generalized to other object categories. In order to improve versatility, relative pose estimation methods have emerged in recent years, which only require a reference image to estimate the pose of a new image.

[0003] However, existing relative pose estimation methods are still limited by the following issues: 2D correspondence-based methods rely on key point matching, but due to large perspective changes, it is difficult to establish reliable correspondences in small overlapping areas. Hypothesis-verification-based methods rely on a discrete sampling process and cannot effectively model continuous pose spaces, resulting in low estimation accuracy. 3D correspondence-based methods require 2D features to be upgraded to 3D voxel features for matching, but due to the lack of additional information in invisible areas, unreliable 3D matching results in high computational overhead. Therefore, existing methods still have significant limitations when dealing with objects with significant perspective changes. Summary of the invention

[0004] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a relative posture estimation method and system based on structure-aware correspondence learning, which solves the problems in the prior art.

[0005] The purpose of the present invention can be achieved by the following technical solutions: The relative pose estimation method based on structure-aware correspondence learning includes the following steps: Collect the query image and the reference image, use the pre-trained backbone network to extract feature maps from the query image and the reference image respectively, and obtain the updated feature maps through the two-step attention mechanism; then multiply them element-by-element with the object mask to obtain the refined feature maps of the query image and the reference image; Use the structured perception key point extraction module to obtain the key point spatial coordinates and corresponding key point features from the fine feature map; Based on the key point spatial coordinates and corresponding key point features, the structured correspondence estimation module is used to lift the key points to the 3D space in the query coordinate system, and regress their corresponding 3D coordinates in the reference coordinate system to establish a 3D correspondence relationship for relative pose estimation; Based on the established 3D correspondence, the weighted singular value decomposition method is used to solve the optimal rotation matrix.

[0006] Furthermore, the two-step attention mechanism includes: multi-head self-attention and multi-head cross-attention.

[0007] Furthermore, the calculation formula of the fine feature map of the query image and the reference image is: in, and are the fine feature maps of the query image and the reference image, respectively. Indicates the multiplication of corresponding elements of the matrix; To query the image corresponding to the object mask, is the object mask corresponding to the reference image, represents a convolutional neural network, is the updated feature map of the query image, is the updated feature map of the reference image.

[0008] Furthermore, the steps for obtaining the key point spatial coordinates of the query image and the corresponding key point features are as follows: 1) Initialize a set of query vectors ; Then use the attention mechanism to update the query vector , adapting it to the image content to generate image-specific keypoint detectors : 2) Calculate key point detector Similarity between image features, generating key point heatmap ; 3) Heat map of key points Row weighted average to obtain the spatial coordinates of the key points of the query image And the corresponding features .

[0009] Furthermore, using the structured correspondence estimation module, the steps of establishing the 3D correspondence for relative pose estimation are: 1) Use self-attention mechanism with rotational position encoding to extract keypoint features from query images To optimize: in, is the optimized key point feature, represents the rotational position encoding, Represents the ROPE position fusion operation; Represents the self-attention mechanism; 2) Apply the cross-attention mechanism to aggregate structural information from the reference image to query the key point features of the image As a query, take the key point features of the reference image As key and value: in, is the updated query image key point feature set; 3) Obtain the updated query image key point feature set Then, the query image Pseudo depth value of key points Concatenate with the corresponding 2D coordinates to obtain the 3D coordinates of each key point in the query coordinate system; in, To query the 2D coordinates of the key points of the image, To query the 3D coordinates of each key point in the coordinate system; Represents the key point feature set from the updated query image No. Key point features; Multilayer perceptron network for predicting pseudo depth; 4) Use MLP to estimate the 3D coordinates of the corresponding key points in the reference coordinate system: in, Indicates the reference coordinate system The estimated 3D coordinates of key points, represents the confidence of the key point, is the position embedding of the 3D coordinates, A multi-layer perceptron network for predicting 3D coordinates in a reference coordinate system.

[0010] Furthermore, the optimal rotation matrix The calculation formula is: in, and are the query and reference coordinate systems. The 3D coordinates of the key points, represents the confidence of the key point, is the three-dimensional rotation matrix, is the number of key points in the query image.

[0011] A relative pose estimation system based on structure-aware correspondence learning, including: Feature extraction unit: collects query images and reference images, extracts feature maps from the query images and reference images respectively using the pre-trained backbone network, and obtains updated feature maps through a two-step attention mechanism. Then, it multiplies the elements of the query images and reference images by the object mask to obtain refined feature maps. Key point extraction unit: Use the structured perception key point extraction module to obtain the key point spatial coordinates and corresponding key point features from the fine feature map; Correspondence estimation unit: Based on the key point spatial coordinates and the corresponding key point features, the structured correspondence estimation module is used to lift the key points to the 3D space in the query coordinate system, and regress their corresponding 3D coordinates in the reference coordinate system to establish a 3D correspondence relationship for relative pose estimation; And, the posture estimation unit: based on the established 3D correspondence, a weighted singular value decomposition method is used to solve the optimal rotation matrix.

[0012] A computer storage medium stores a readable program, which can execute the above-mentioned relative posture estimation method based on structure-aware correspondence relationship learning when the program is running.

[0013] An electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned relative posture estimation method based on structure-aware correspondence relationship learning.

[0014] A computer program product includes computer instructions, wherein the computer instructions instruct a computing device to perform operations corresponding to the above-mentioned relative posture estimation method based on structure-aware correspondence relationship learning.

[0015] Beneficial effects of the present invention: 1. Compared with the existing methods based on 2D or 3D feature matching, the present invention avoids the problems of unreliable matching or high computational overhead in traditional methods by directly regressing the 3D correspondence.

[0016] 2. The structure-aware key point extraction module proposed in the present invention can effectively cope with the challenge of large changes in object shape, while the structure-based correspondence estimation module solves the problem that the correspondence relationship is difficult to accurately estimate under large viewing angle changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 is a flow chart of the relative posture estimation method of the present invention; Figure 2 It is a framework diagram of the structured key point extraction module of the present invention; Figure 3 It is a framework diagram of the structured correspondence estimation module of the present invention. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] Example 1 like Figure 1 As shown, the relative posture estimation method based on structure-aware correspondence learning includes the following steps: S1, collect query images and reference images, and use the pre-trained backbone network to extract feature maps from the query images and reference images respectively and ; and processed by a two-step attention mechanism to obtain an updated feature map and ; The feature map will be updated later and Multiply element-wise with the object mask to obtain the fine feature map of the query image and the reference image , ; Get fine feature map , The process is: 1) First, use the pre-trained backbone network to extract image features from the query image and the reference image to obtain the feature map ; In this embodiment, the backbone network is CroCoNet (Cross-view Completion Network), which includes: a feature encoding module for extracting high-dimensional features from the input image; a cross-view attention module for modeling interactive information between features of different perspectives; and a feature decoding module included in self-supervised training, which is used to reconstruct a complete object representation from the fused features, thereby achieving cross-perspective feature completion and structure-aware representation.

[0021] 2) Then, the feature map The feature map of the query image is processed through a two-step attention mechanism, including multi-head self-attention (MHSA) and multi-head criss-cross attention (MHCA) layers. For example, update the features: in, is the query image feature map of the previous layer, is the reference image feature map of the previous layer, is the query image feature map updated by the self-attention mechanism, Updated query image feature map for the current layer.

[0022] Feature map of the reference image The same operation is performed. Different from previous work, the parameters in MHSA and MHCA are ) and the reference image ( ) to ensure the consistency of learned features. After such interaction, the updated feature map is obtained. and .

[0023] 3) In order to suppress the influence of background features, a lightweight mask predictor is applied to generate object masks, thereby refining the feature map; specifically, the mask and Respectively by and get: in, Convolutional neural network representing a simple structure, predicting object masks.

[0024] The mask prediction loss is calculated using the binary cross entropy (BCE) loss and the true mask. : in, To query the real mask of the object corresponding to the image, is the real mask of the object corresponding to the reference image; Finally, by multiplying the object mask element by element, we can obtain the fine feature maps of the query image and the reference image respectively. and , effectively retaining object-related features and reducing background interference: in, Indicates the multiplication of corresponding elements of the matrix; To query the image corresponding to the object mask, is the object mask corresponding to the reference image, is the updated feature map of the query image, is the updated feature map of the reference image.

[0025] S2, using the structured perception key point extraction module from the fine feature map , In the example above, we obtain the spatial coordinates of the key points of the query image and the reference image. , And the corresponding key point features , ; In order to solve the problem of how to effectively represent the structure of object components, this embodiment proposes the following Figure 2 The structured keypoint extraction module shown in the figure aims to adaptively select keypoints with structural significance. This method reduces the computational burden by focusing on keypoints with stable structures and effectively aggregates structure-aware features, thereby improving the accuracy of pose estimation and enhancing the model's generalization ability for unknown object categories.

[0026] like Figure 2 As shown in the structured key point extraction module, the following takes the query image as an example to introduce the key point spatial coordinates of the query image. Corresponding key point features Steps to obtain: 1) Initialize a set of learnable query vectors ,in is the number of key points, is the feature dimension; these query vectors are then updated using an attention mechanism to adapt them to the image content, generating image-specific keypoint detectors : 2) Calculate keypoint detector Similarity between image features, generating key point heatmap ; 3) Key point heat map Row weighted average to obtain the spatial coordinates of the key points of the query image and the corresponding features ; in, represents the spatial coordinates of the key points of the query image, Indicates the corresponding features; However, unconstrained key point extraction often results in key points being concentrated in local areas, which cannot effectively capture the comprehensive structural information of object parts. To solve this problem, this embodiment introduces image reconstruction loss, which reconstructs the object foreground using only the features and coordinates of the key points, so that the key points cover the semantically rich areas of the object: in, Is a lightweight decoder, represents the reconstructed query image; image reconstruction loss By pixel-level similarity The loss and the perceptual loss based on the VGG network are composed of: in, Represents the VGG network The feature map of the layer, is the weighting factor, is the query image; The loss ensures pixel accuracy, and the perceptual loss ensures semantic consistency.

[0027] Through image reconstruction, the key point distribution can be optimized end-to-end to ensure that the key points cover the semantically rich areas on the surface of the object, thereby enhancing the structural representation. By applying the above process to the query image and the reference image, key points with structural significance can be obtained to effectively represent the structure of the object.

[0028] For the reference image, the key point extraction process is similar to that of the query image. Through the same structured-aware key point extraction module, the key point detector is first generated using the learnable query vector, and then the similarity between the query vector and the image features is calculated to generate a heat map. Subsequently, the spatial coordinates of the key points in the reference image and their corresponding features are obtained by weighted averaging of the heat map. At the same time, the image reconstruction loss is also introduced to optimize the key point distribution to ensure that the key points cover the semantically salient areas in the reference image, thereby improving the integrity and robustness of the structural expression.

[0029] S3, spatial coordinates of key points based on query image and reference image , And the corresponding key point features , , the key points are lifted to the 3D space in the query coordinate system using the structured correspondence estimation module, and their corresponding 3D coordinates in the reference coordinate system are regressed to establish the 3D correspondence relationship for relative pose estimation; Given the 2D keypoint coordinates of the query image and the reference image , And its corresponding features , , extract structure-aware features, use the structured correspondence estimation module to lift 2D key points to 3D space in the query coordinate system, and regress their corresponding 3D coordinates in the reference coordinate system, thereby establishing a set of 3D correspondences for relative pose estimation.

[0030] like Figure 3 As shown, using the structured correspondence estimation module, the steps of establishing the 3D correspondence for relative pose estimation are: 1) For key point features extracted from the query image , using the self-attention mechanism with rotational position encoding (ROPE) for feature optimization, so that the key point features Able to perceive the structural information in the image; the rotation position encoding is recorded as , Represents the ROPE position fusion operation.

[0031] 2) Apply the cross-attention mechanism to aggregate structural information from the reference image to query the key point features of the image As a query, take the key point features of the reference image As key and value: in, is the updated query image key point feature set; These attention mechanisms help capture the relationship between inside and outside the image, making keypoint features more robust and facilitating the estimation of 3D correspondences.

[0032] 3) Obtain the updated query image key point feature set Finally, the 2D keypoints are lifted to 3D space in the query coordinate system by regressing the pseudo depth value of each keypoint: in, Indicates the query image The pseudo depth value of key points, Represents the key point feature set from the updated query image No. Key point features; Multilayer perceptron network for predicting pseudo depth; The query image Pseudo depth value of key points Concatenate with the corresponding 2D coordinates to obtain the 3D coordinates of each key point in the query coordinate system: in, To query the 2D coordinates of the key points of the image, To query the 3D coordinates of each key point in the coordinate system; 4) Use another MLP to estimate the 3D coordinates of the corresponding key points in the reference coordinate system. The input of this MLP is the key point feature set from the updated query image. No. Key point features And the 3D coordinates of each keypoint in the query coordinate system: in, Indicates the reference coordinate system The estimated 3D coordinates of key points, represents the confidence of the key point, is the position embedding of the 3D coordinates, A multi-layer perceptron network is used to predict the 3D coordinates in the reference coordinate system. Based on the 3D coordinates of the same key points in the query and reference coordinate systems, a 3D-3D correspondence can be naturally established to help determine the relative pose.

[0033] In order to ensure the accuracy of these 3D correspondences, this embodiment proposes a loss function to supervise the predicted 3D coordinates. Specifically, the ground truth 3D coordinates in the reference coordinate system are calculated by the ground truth rotation matrix: in, are the ground truth 3D coordinates in the reference coordinate system, To query the 3D coordinates of each key point in the coordinate system; is the ground truth rotation matrix.

[0034] In order to The estimated 3D coordinates of the key points The real 3D coordinates in the reference coordinate system Alignment, defining 3D keypoint loss as follows: in, is the number of key points, represents the confidence of the key point, is a hyperparameter that controls the impact of confidence; is the 3D key point error, 3D key point error Measures the difference between the estimate and the ground truth coordinates: in, This symmetric loss penalizes the deviation between the two coordinate systems and ensures that the 3D coordinate and Finally, a reliable 3D correspondence is obtained.

[0035] S4, based on the 3D correspondence established in S3, uses the weighted singular value decomposition (wSVD) method to solve the optimal rotation matrix ; After obtaining the 3D correspondence, the weighted singular value decomposition (wSVD) method is used to solve the relative rotation ; in, and are the query and reference coordinate systems. The 3D coordinates of the key points, represents the confidence of the key point, is the three-dimensional rotation matrix, To query the number of key points in the image; the optimization goal is to find the optimal rotation matrix , so that and The error between these two sets of 3D coordinates is minimal. The specific steps are as follows: 1) First calculate the covariance matrix : 2) Then the covariance matrix Perform SVD decomposition: in, and is an orthogonal matrix, representing two orthogonal bases; 3) Optimal rotation matrix Calculated by the following formula: In order to estimate the optimal rotation matrix The ground truth value of the relative rotation matrix Alignment, the present invention adopts loss: in, and They are and 6D representation.

[0036] In addition, the query image and the reference image are used symmetrically during training, and the key points in the reference image are also used to obtain a set of 3D correspondences, which effectively provides more training data and improves efficiency and robustness. At inference time, the model extracts key points from the query and reference images, but only regresses the 3D coordinates of the query image for relative pose estimation.

[0037] This embodiment proposes a relative pose estimation method based on structure-aware correspondence learning, which is particularly suitable for robot grasping tasks. Compared with the traditional relative pose estimation method, the method of this embodiment avoids the common matching error and high computational overhead problems by extracting key points that characterize the structure of the object and directly regressing its 3D correspondence. The method of this embodiment can effectively perform pose estimation without an object CAD model, only through a small number of reference images. Through the key point extraction module of structure perception, it is possible to solve the challenges brought by the variable shape and large texture changes of the object, and still maintain high accuracy in scenes with large viewing angle changes or partial occlusion. Unlike traditional methods, the present invention does not rely on a large amount of training data or a specific object model, and has a stronger generalization ability. This means that even in the face of objects with variable shapes and complex structures, the system can still perform pose estimation stably and accurately. Especially in robot grasping tasks, the device only needs to use a small number of reference images to complete the estimation of the object pose in real time, without retraining or pre-building the object model, which greatly improves the efficiency and robustness of the grasping task.

[0038] Based on similar inventive concepts, an embodiment of the present invention further provides a computer storage medium storing a readable program, which can execute the above-mentioned relative posture estimation method based on structure-aware correspondence relationship learning when the program is running.

[0039] Based on similar inventive concepts, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned relative posture estimation method based on structure-aware correspondence relationship learning.

[0040] Based on similar inventive concepts, an embodiment of the present invention further provides a computer program product, including computer instructions, which instruct a computing device to execute operations corresponding to the above-mentioned relative posture estimation method based on structure-aware correspondence relationship learning.

[0041] Example 2 Based on the relative posture estimation method based on structure-aware correspondence learning proposed in Example 1, in this embodiment, a relative posture estimation system based on structure-aware correspondence learning is proposed, which specifically includes: Feature extraction unit: collects query images and reference images, extracts feature maps from the query images and reference images respectively using the pre-trained backbone network, and obtains updated feature maps through a two-step attention mechanism. Then, it multiplies the elements of the query images and reference images by the object mask to obtain refined feature maps. Key point extraction unit: Use the structured perception key point extraction module to obtain the key point spatial coordinates and corresponding key point features from the fine feature map; Correspondence estimation unit: Based on the key point spatial coordinates and the corresponding key point features, the structured correspondence estimation module is used to lift the key points to the 3D space in the query coordinate system, and regress their corresponding 3D coordinates in the reference coordinate system to establish a 3D correspondence relationship for relative pose estimation; And, the posture estimation unit: based on the established 3D correspondence, a weighted singular value decomposition method is used to solve the optimal rotation matrix.

[0042] Example 3 In this embodiment, the relative posture estimation method based on structure-aware correspondence learning of the present invention is experimentally verified to reflect the significant improvement in accuracy of the method; the experimental content includes: 1) Experimental Dataset This experiment selected three public and widely used relative pose estimation datasets: CO3D, Objaverse, and LineMOD. The CO3D dataset contains 18,619 video sequences spanning 51 categories of objects. In the experiment, 41 categories were selected as training sets, and the remaining 10 categories were selected as test sets to verify the generalization ability of the model; the Objaverse dataset consists of synthetic images rendered from 3D models. The experiment selected 128 categories for testing and the rest for training; the LineMOD dataset consists of real images of 13 household objects. The test set contains 5 categories of objects and is completely isolated from the training set to ensure test independence.

[0043] 2) Implementation parameter settings The model is trained using the Adam optimizer, and the initial learning rate is set to 2×10⁻ 4, and decays to the original 0.1 after every 200 epochs. The entire training process is carried out for 400 epochs, and the batch size is set to 80. All experiments are run on 4 NVIDIA RTX3090 graphics cards, and the total training time is about 36 hours. During training, the image area is cropped with the real annotated target bounding box to enhance the model's attention to the target area.

[0044] 3) Comparison method In order to comprehensively evaluate the superiority of the proposed method, seven representative existing mainstream relative pose estimation methods including SuperGlue, LoFTR, ZSP, RelPose, RelPose++, 3DAHV, and DVMNet were selected as comparison benchmarks. These methods cover different paradigms such as local feature matching, hypothesis testing, and deep learning modeling, and are highly representative.

[0045] In this embodiment, three indicators are used to evaluate the performance of different methods in the task of 3D object pose estimation, including: mean angle error (mAE, the lower the better), which indicates the average deviation between the predicted pose and the true pose; and accuracy Acc@30° and Acc@15° (both the higher the better), which respectively indicate the proportion of samples with angle errors within 30 degrees and 15 degrees, to measure the performance of the methods under different accuracy requirements. The experimental results are shown in Table 1 below: Table 1 Comparison results of this method with the existing state-of-the-art methods on three relative pose estimation datasets As can be seen from Table 1, the present invention has achieved the best results in all evaluation indicators on the three authoritative relative pose estimation datasets of CO3D, Objaverse and LineMOD, showing a comprehensive performance advantage over the existing technology. On the CO3D dataset, the present invention not only achieved the lowest average pose error (mAE 14.2°), but also achieved 93.6% and 80.2% accuracy in Acc@30° and Acc@15°, respectively, far exceeding the existing methods. In the Objaverse and LineMOD datasets, the present invention continues to maintain the best performance in all indicators, especially in LineMOD, Acc@30° is increased to 76.2%, Acc@15° is increased to 41.8%, showing strong generalization ability and robustness in complex scenarios.

[0046] In summary, the comprehensive optimal performance of the present invention on various typical data sets fully demonstrates the advanced nature and practical value of its technical solution in attitude estimation tasks, and has significant technological advancement and industrial application potential.

[0047] The method of the present invention may be implemented in hardware, firmware, or as software or computer code that may be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein may be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, processor, or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0048] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements all fall within the scope of the present invention to be protected.

Claims

1. A relative pose estimation method based on structure-aware correspondence learning, characterized in that: The following steps are involved: Collect query images and reference images, use the pre-trained backbone network to extract feature maps from the query images and reference images respectively, and obtain updated feature maps through a two-step attention mechanism; Then, it is element-wise multiplied with the object mask to obtain the fine feature maps of the query image and the reference image; Use the structured perception key point extraction module to obtain the key point spatial coordinates and corresponding key point features from the fine feature map; Based on the key point spatial coordinates and corresponding key point features, the structured correspondence estimation module is used to lift the key points to the 3D space in the query coordinate system, and regress their corresponding 3D coordinates in the reference coordinate system to establish a 3D correspondence relationship for relative pose estimation; Based on the established 3D correspondence, the weighted singular value decomposition method is used to solve the optimal rotation matrix.

2. The relative posture estimation method based on structure-aware correspondence learning according to claim 1, characterized in that: The two-step attention mechanism includes: multi-head self-attention and multi-head cross-attention.

3. The relative posture estimation method based on structure-aware correspondence learning according to claim 1, characterized in that: The calculation formula for the fine feature map of the query image and the reference image is: in, and are the fine feature maps of the query image and the reference image, respectively. Indicates the multiplication of corresponding elements of the matrix; To query the image corresponding to the object mask, is the object mask corresponding to the reference image, represents a convolutional neural network, is the updated feature map of the query image, is the updated feature map of the reference image.

4. The relative posture estimation method based on structure-aware correspondence learning according to claim 1, characterized in that: Steps to obtain the key point spatial coordinates and corresponding key point features of the query image: 1) Initialize a set of query vectors ; Then use the attention mechanism to update the query vector , adapting it to the image content to generate image-specific keypoint detectors : 2) Calculate keypoint detector Similarity between image features, generating key point heatmap ; 3) Heat map of key points Row weighted average to obtain the spatial coordinates of the key points of the query image And the corresponding features .

5. The relative posture estimation method based on structure-aware correspondence learning according to claim 4, characterized in that: Using the structured correspondence estimation module, the steps to establish 3D correspondences for relative pose estimation are: 1) Use self-attention mechanism with rotational position encoding to extract keypoint features from query images To optimize: in, is the optimized key point feature, represents the rotational position encoding, Represents the ROPE position fusion operation; Represents the self-attention mechanism; 2) Apply the cross-attention mechanism to aggregate structural information from the reference image to query the key point features of the image As a query, take the key point features of the reference image As key and value: in, is the updated query image key point feature set; 3) Obtain the updated query image key point feature set Then, the query image Pseudo depth value of key points Concatenate with the corresponding 2D coordinates to obtain the 3D coordinates of each key point in the query coordinate system; in, To query the 2D coordinates of the key points of the image, To query the 3D coordinates of each key point in the coordinate system; Represents the key point feature set from the updated query image No. Key point features; Multilayer perceptron network for predicting pseudo depth; 4) Use MLP to estimate the 3D coordinates of the corresponding key points in the reference coordinate system: in, Indicates the reference coordinate system The estimated 3D coordinates of key points, represents the confidence of the key point, is the position embedding of the 3D coordinates, A multi-layer perceptron network for predicting 3D coordinates in a reference coordinate system.

6. The relative posture estimation method based on structure-aware correspondence learning according to claim 1, characterized in that: The optimal rotation matrix The calculation formula is: in, and are the query and reference coordinate systems. The 3D coordinates of the key points, represents the confidence of the key point, is the three-dimensional rotation matrix, is the number of key points in the query image.

7. A relative pose estimation system based on structure-aware correspondence learning, characterized in that: include: Feature extraction unit: collects query images and reference images, extracts feature maps from the query images and reference images respectively using the pre-trained backbone network, and obtains updated feature maps through a two-step attention mechanism. Then, it multiplies the elements of the query images and reference images by the object mask to obtain refined feature maps. Key point extraction unit: Use the structured perception key point extraction module to obtain the key point spatial coordinates and corresponding key point features from the fine feature map; Correspondence estimation unit: Based on the key point spatial coordinates and the corresponding key point features, the structured correspondence estimation module is used to lift the key points to the 3D space in the query coordinate system, and regress their corresponding 3D coordinates in the reference coordinate system to establish a 3D correspondence relationship for relative pose estimation; And, the posture estimation unit: based on the established 3D correspondence, a weighted singular value decomposition method is used to solve the optimal rotation matrix.

8. A computer storage medium storing a readable program, characterized in that: When the program is running, it can execute the relative posture estimation method based on structure-aware correspondence relationship learning as described in any one of claims 1 to 6.

9. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the relative posture estimation method based on structure-aware correspondence relationship learning as described in any one of claims 1-6.

10. A computer program product comprising computer instructions, characterized in that: The computer instructions instruct the computing device to perform operations corresponding to the relative posture estimation method based on structure-aware correspondence learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Key point detection method and device, neural network, electronic equipment and storage medium

    CN115063598A

  • Deep learning algorithm for 6D pose estimation based on attention mechanism

    CN116580085A

  • Method and device for directly regressing object 6D pose based on image joint attention

    CN119205911A

  • Self-supervised 3D keypoint learning for ego-motion estimation

    US20210237764A1

  • Systems and methods for generic visual odometry using learned features via neural camera models

    US20220084231A1