Feature point matching method based on three-dimensional point position coding

By adopting a three-dimensional point position encoding method in image feature matching, combining depth estimation and camera position information, the problem of inaccurate three-dimensional position information in the prior art is solved, and a higher quality feature matching result is achieved.

CN120032144APending Publication Date: 2025-05-23XIDIAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510177538.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the prior art uses three-dimensional spatial information to match image features, it lacks accurate depth information, resulting in unreliable three-dimensional position information and affects the quality of the matching results.

Method used

The feature point matching method based on three-dimensional point position encoding is adopted. By estimating the relative depth of the input image and fusing it with the image features, the features of the depth estimation task are extracted using a convolutional network, direct depth estimation and depth distribution probability estimation are performed, and three-dimensional point position encoding is performed in combination with the camera's relative position information to generate image features containing three-dimensional spatial information.

Benefits of technology

Improve the accuracy and reliability of image feature matching, especially in large-view angle changes or weak texture scenarios, which can provide higher quality matching results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032144A_ABST
    Figure CN120032144A_ABST
Patent Text Reader

Abstract

The invention discloses a feature point matching method based on three-dimensional point position coding. The method comprises the following steps: 1, estimating relative depth of an input image, and encoding to obtain depth guide features; 2, fusing the depth guide features with image features of an input image, and extracting features of a depth estimation task of the image through a convolutional network; 3, respectively carrying out direct depth estimation and depth distribution probability estimation on the features of the depth estimation task, and weighting an estimated result to obtain an estimated real depth; 4, performing three-dimensional point position coding by using the estimated real depth to obtain a three-dimensional position feature; and 5, performing element-by-element addition on the three-dimensional position features and the image features, and sending the image features containing three-dimensional space information to a matching module to obtain final two-dimensional matching point coordinates. The method has the advantages that scene depth information is efficiently utilized, reliable three-dimensional space information is introduced into image features, and high-quality matching point extraction is carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of feature matching between images, and in particular relates to a feature point matching method based on three-dimensional point position coding. Background Art

[0002] Feature matching between images is an important task in many computer vision tasks, including camera calibration, structure from motion (SfM), visual localization, simultaneous localization and mapping (SLAM), stereo matching, etc. In general, the local feature matching problem can be solved in three stages: feature point detection, descriptor extraction, and feature matching. Most existing methods follow the above process in sequence to handle feature matching. The feature point detector narrows the focus to a set of interest points on the image. The descriptor extraction stage generates corresponding feature descriptors for each interest point. Finally, the correspondence between image interest points is found through feature matching algorithms. With the development of deep learning, many image feature matching methods based on deep learning have been generated, among which the detector-free method does not go through the feature point detection stage and regards each pixel as a potential interest point to achieve pixel-level matching.

[0003] Chenhao Li et al. proposed a new image feature matching method in their paper "PA-LoFTR: Local Feature Matching with 3D Position-Aware Transformer". This method introduces 3D position information, adds depth information generated by the depth estimation module to enhance feature representation, associates 3D spatial information with depth features, and learns image features containing 3D spatial information. This information helps determine accurate image matching and provides high-quality matching results in challenging scenes. However, the camera-ray-based 3D position encoding used in this method cannot provide reliable 3D position information due to the lack of accurate depth information.

[0004] The prior art solution introduces three-dimensional spatial information on the basis of extracting local visual features, which helps the matching method to obtain higher quality matching results in scenes with large viewing angle changes or weak textures. However, the three-dimensional information encoding method used is to encode the ray direction from the optical center of the camera to the pixel on the image plane. Specifically, given a depth range Rd = [Dmin, Dmax], the camera ray position encoding first divides a fixed maximum depth value into Nd intervals, so the 3D position information of the pixel is represented as Nd points along the camera ray direction. The ray direction only provides rough positioning information of 2D image features, and does not contain depth information. In the prior art solution, the depth information obtained by the depth estimation network has not been fully utilized in combination with the relative position information of the camera, making the role of the depth information relatively limited. Summary of the invention

[0005] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a feature point matching method based on three-dimensional point position coding, which has the characteristics of efficiently utilizing scene depth information, introducing reliable three-dimensional spatial information into image features, and performing high-quality matching point extraction.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is:

[0007] A feature point matching method based on three-dimensional point position coding, characterized in that it includes the following steps:

[0008] Step 1: Estimate the relative depth of the input image and encode it to obtain depth-guided features;

[0009] Step 2: Fusing the depth-guided features with the image features of the input image, and extracting features of the depth estimation task of the image through a convolutional network;

[0010] Step 3: perform direct depth estimation and depth distribution probability estimation on the features of the depth estimation task respectively, and weight the estimated results to obtain the estimated true depth;

[0011] Step 4: Encode the three-dimensional point position using the estimated true depth to obtain a three-dimensional position feature;

[0012] Step 5: Add the three-dimensional position features and image features element by element to generate image features containing three-dimensional spatial information, and send the image features containing three-dimensional spatial information to the matching module to obtain the final two-dimensional matching point coordinates.

[0013] In the step 1, an RGB image of the scene is obtained as an input image, and a relative depth value of the scene is generated by a relative depth estimation module; the relative depth value of the scene is converted into a weighted feature vector corresponding to adjacent integers after linear mapping to obtain a relative depth guidance feature.

[0014] The step 1 is specifically as follows:

[0015] First, the RGB image of the shape [3, H, W] is processed by the relative depth estimation module DepthAnytingv2 model to generate the scene relative depth value of the shape [1, H, W], whose value is a floating point number in the range [0, 1]. Then, the above [0, 1] floating point number is mapped to a floating point number in the range [0, 255].

[0016] The learnable embedding layer is responsible for mapping integers in the range [0, 255] into feature vectors with a dimension of 256. Then, the linear interpolation method is used to convert the linearly mapped relative depth values ​​in the range [0, 255] into the weighted feature vectors corresponding to adjacent integers to obtain the relative depth guidance feature. The relative depth guidance feature is used as the scene geometry prior information to be combined with the image features to enhance the depth estimation capability of the depth estimation module.

[0017] The step 2 is specifically as follows:

[0018] The RGB image outputs image features through the image feature extraction module, the relative depth guide feature is multiplied element by element with the output image feature of the same dimension, the combined feature is obtained, and the combined feature is linearly transformed through a learnable linear layer to obtain a fusion feature;

[0019] The fused features are passed through a CNN convolution module for feature extraction to generate features for depth estimation tasks.

[0020] In the step 3, the features used for the depth estimation task are respectively passed through two branches, namely a direct depth estimation module and a depth distribution probability module, for depth estimation: the depth estimation module is able to retain the advantages of the two depth estimation methods, dynamically adjust the weights of the two methods according to the characteristics of the scene, and obtain the final depth estimation by weighting the two depth estimation results for three-dimensional coordinate calculation and subsequent three-dimensional point position encoding.

[0021] The step 3 is specifically as follows:

[0022] The pixel position in each RGB image can obtain the direct depth estimation value and the indirect depth distribution probability respectively. The depth range of [0,15] meters is evenly divided into 64 intervals. The above probability distribution estimates the corresponding probability for each depth interval. The midpoint value of each interval is used as the depth value of the interval, which is multiplied and added with the probability value corresponding to each interval. Finally, a depth expectation value is obtained as the result of the indirect depth estimation branch.

[0023] Set up a learnable weight α and transform the depth results D obtained by the above two branches r , D p According to the formula D = α * D r +(1-α)*D p , calculate the final estimated depth;

[0024] Use the true depth value of the dataset to calculate the loss of the estimated results of the two depth estimation branches to supervise the training of module parameters;

[0025] For the direct depth estimation branch, the L1 norm mean of the estimated value and the true depth value of each pixel is calculated; for the depth interval distribution probability estimation branch, the formula is used:

[0026] loss df =∑α(1-p i ) γ logp i

[0027] Calculate the sum of the depth estimation losses at each pixel corresponding to the position, where p i Represents the probability of each pixel position being estimated in the correct depth interval. α and γ are the set hyperparameters, and the final depth estimation loss is the sum of the two.

[0028] In step 4, the homogeneous form of the two-dimensional coordinates of each pixel in the image is multiplied by the corresponding depth value to obtain a column vector: the two-dimensional pixel coordinates are back-projected into the three-dimensional camera coordinate system, and the three-dimensional space coordinates corresponding to each pixel of the two images are converted to the same coordinate system using the relative camera pose information corresponding to the two images;

[0029] For a three-dimensional space coordinate, the maximum and minimum values ​​in each coordinate direction are used for normalization. After sine and cosine position encoding, the three-dimensional position feature is obtained, and then the linear encoding layer is used to realize the alignment of the three-dimensional position feature with the image feature dimension.

[0030] The step 4 is specifically as follows:

[0031] Multiply the homogeneous form of the pixel's two-dimensional coordinates [x, y, 1] by the corresponding depth value d to obtain a column vector:

[0032] p h =[x*d,y*d,d] T

[0033] Using the formula p 3d =K -1 p h , that is, let the inverse of the camera intrinsic parameter matrix be multiplied by vector p h , back-project the two-dimensional pixel coordinates into the three-dimensional camera coordinate system, and use the relative camera pose information corresponding to the two images, namely the rotation matrix R and the translation matrix t, according to the formula:

[0034] p 3d '=Rp 3d +t=[x,y,z] T ;

[0035] For a three-dimensional space coordinate p 3d =[x,y,z] T , use the maximum and minimum values ​​in each coordinate direction to perform normalization:

[0036]

[0037] Among them, u and v represent the two-dimensional coordinates of the pixel. For the x coordinate of a certain three-dimensional point, sine and cosine encoding is used to convert it into a 128-dimensional vector:

[0038] Similarly, the three-dimensional points are converted into position features of 128*3 dimensions, and a learnable linear layer is set up to map the above position feature dimensions to 256 to achieve alignment with the image feature dimensions.

[0039] The step 5 is specifically as follows:

[0040] The obtained three-dimensional position features are added element by element to the image features output by the image feature extraction module as the input of the Transformer feature encoding module to generate image features containing three-dimensional spatial information;

[0041] Image features containing three-dimensional spatial information, the shapes are [L0, C], [L1, C], respectively, image features F3d0, F3d1, where L0 = H0*W0, L1 = H1*W1, i.e., the number of feature vectors corresponding to each image, and C represents the number of feature channels;

[0042] The matching module calculates the two images to be matched, Figure 0 and Figure 1 , Figure 0 and Figure 1 The cosine similarity of image features F3d0 and F3d1 is used to obtain a similarity matrix with a shape of [L0, L1]. A threshold t is set and the mutual nearest neighbor algorithm is used to screen out all matching pairs in the similarity matrix that meet the following two conditions, and the feature index is converted into a two-dimensional coordinate as a coarse matching result. Based on the coarse matching result, the feature vector in the corresponding position window is extracted from the feature map. Similarly, the similarity score of the feature vector is calculated to obtain the fine matching result in the window. After the two matching stages of coarse matching and fine matching, the final two-dimensional matching point coordinates are obtained.

[0043] A feature point matching system based on three-dimensional point position coding, comprising a relative depth estimation module, an image feature extraction module, a true depth estimation module, a three-dimensional point position coding module, a Transformer feature coding module and a feature matching module;

[0044] The relative depth estimation module is responsible for performing monocular relative depth estimation on the two RGB image scenes to obtain relative depth values;

[0045] The image feature extraction module is used to extract features from the relative depth value output by the relative depth estimation module to obtain image features;

[0046] The relative depth value and image features are used as inputs to the real depth estimation module, which is responsible for combining the relative depth value with the image features and outputting the depth features.

[0047] The 3D point position encoding module is used to convert the depth features into 3D point position features;

[0048] The Transformer feature encoding module converts 3D point position features and depth features into image features containing 3D spatial information;

[0049] The feature matching module is used to match image features containing three-dimensional spatial information.

[0050] The application of feature point matching methods based on 3D point position encoding is used for 3D reconstruction, image stitching and panorama generation, and visual SLAM.

[0051] In 3D reconstruction: multiple images (usually taken from different perspectives) are used to restore the 3D structure of a scene or object. When the camera pose is known, it can be used to reduce the amount of calculation and improve accuracy.

[0052] Image stitching and panorama generation: stitching multiple images into a complete image (such as a panorama). Although the camera pose is known, feature point matching is still crucial because it is necessary to identify corresponding points between images to align the images. Known relative poses can help quickly locate and align images, reduce the amount of calculation, and enhance the accuracy of stitching.

[0053] In visual SLAM: In the visual perception system of a mobile robot or drone, the camera captures environmental information, dynamically builds a map and locates itself. When the relative position is known, feature point matching can be used to help correct the camera's trajectory or optimize the map. Feature point matching can enhance the stability of the system, especially in long-term operation or large-scale environments.

[0054] Beneficial effects of the present invention:

[0055] Compared with the camera ray position encoding method under the prior art, the three-dimensional point position encoding method of the present invention can provide more accurate three-dimensional spatial information. At the same time, the present application proposes to introduce relative depth guidance information to the depth estimation module. Since this information can provide global geometric structure prior, the relative depth information is encoded as a feature and combined with the image feature, which can realize multimodal feature fusion and enhance the generalization ability of the real depth estimation module, thereby reducing the training difficulty of the depth estimation module.

[0056] The hybrid depth estimation module of the present invention performs direct depth estimation and classification depth estimation at the same time, and sets a learnable weighting coefficient to fuse the depth estimation results of the two. The former directly obtains the depth value of each pixel through regression estimation, which is suitable for processing areas with smooth depth changes in the scene, but is more sensitive to noise and local details; the latter discretizes the depth value into multiple categories, estimates the probability on each category and multiplies it with the depth value represented by each category and performs weighted calculation to perform depth estimation. It is more robust to noise and local details, but due to discretization, the resolution of the depth estimation is low. By combining these two methods, the advantages of both are retained, and the model dynamically adjusts the weights of the two methods according to the characteristics of the scene, thereby helping the model to generate more reliable depth estimates in complex scenes, generate reliable depth features and three-dimensional space information, and ultimately improve the performance of the feature point matching model, and produce more reliable matching results than the existing technical solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is the overall block diagram of the model of the present invention.

[0058] Figure 2 It is a schematic diagram of the process method of the present invention. DETAILED DESCRIPTION

[0059] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0060] like Figure 1 As shown, the feature matching model based on three-dimensional point position coding includes a relative depth estimation module, an image feature extraction module, a true depth estimation module, a three-dimensional point position coding module, a Transformer feature coding module and a feature matching module;

[0061] The relative depth estimation module and image feature extraction module provide scene relative depth prior and basic image information; the true depth estimation module and 3D point position encoding are responsible for depth and 3D information processing; the Transformer feature encoding module is responsible for the fusion of image features, 3D position features and depth features; and the feature matching module is responsible for the final matching task.

[0062] Relative depth estimation module: takes grayscale images I0 and I1 of size [1, H, W] (number of channels, height, width) as input, is responsible for performing monocular relative depth estimation on the two RGB image scenes respectively, obtains relative depth values ​​of size [1, H, W] between 0 and 1, and converts them into corresponding relative depth features Frd0 and Frd1 as guidance information through the encoding layer.

[0063] Image feature extraction module (CNN): takes grayscale images I0 and I1 of size [1, H, W] as input, is responsible for extracting local features of the image through multi-layer convolution blocks, and outputs feature maps of size [128, H / 2, W / 2] and [256, H / 8, W / 8] as image features for each image.

[0064] Real depth estimation module: takes the relative depth features output by the relative depth estimation module and the image features output by the image feature extraction module as input, is responsible for combining the relative depth guidance information with the image features, and obtains the real depth map D0 and D1 of size [1, H / 8, W / 8] as the output of this module through the hybrid depth estimation algorithm, and outputs the intermediate feature vector as the depth feature.

[0065] 3D point position encoding module: takes the real depth estimation output by the real depth estimation module as input, and is responsible for projecting the pixels in the two images to the 3D coordinates in the same camera coordinate system with the help of the camera's intrinsic parameter matrix K and the relative pose between the two cameras, i.e., the rotation matrix R and the translation matrix t, and uses the sine-cosine encoding algorithm and the linear encoding layer to obtain the 3D point position features as the output of this module.

[0066] Transformer feature encoding module: takes the image features output by the image feature extraction module and combined with the 3D point position features, and the depth features output by the true depth estimation module as input;

[0067] Responsible for using the attention algorithm to calculate Figure 0, Figure 1 The mutual attention between the image feature Fi and the deep feature Fd, Figure 0 and Figure 1 The self-attention of its own image feature Fi and the image feature Fi0 of Figure 0 and Figure 1 The mutual attention of the image feature Fi1 eventually generates image features F3d0 and F3d1 containing three-dimensional spatial information as the output of this module.

[0068] Feature matching module: The output F3d0 and F3d1 of the Transformer feature encoding module with a size of [L=H*W,C] are used as the module input, which is responsible for densely matching the two sets of feature vectors.

[0069] Specifically, this module calculates the L0 image features in Figure 0 and Figure 1 The feature cosine similarity of the L1 image features in the image is obtained to obtain a similarity matrix with a shape of [L0, L1]. A threshold t is set and the mutual nearest neighbor algorithm is used to filter out all matching pairs in the similarity matrix that meet the following two conditions: Figure 1 are the two images to be matched;

[0070] The current value is the maximum value of the row and column;

[0071] The current value is greater than the threshold t;

[0072] This module first obtains the corresponding integer pixel matching point coordinates based on the selected matching pairs. Specifically, if the [i, j] coordinates in the similarity matrix meet the matching pair conditions, it means that the i-th point in Figure 0 is Figure 1 The j-th point in is a pair of matching points. According to the shape and scale of the feature map, the two-dimensional coordinates of the whole pixel matching point can be obtained as P0: [s*i / W, s*(i%W)], P1: [s*j / W, s*(j%W)], where s represents the scale of the image feature map relative to the original image, and W represents the width of the feature map. The above is the coarse matching stage. Based on the coarse matching results, the module extracts the image features at the corresponding position of the image 1 / 2 feature map for a layer of refined matching. A feature vector in the image 1 / 8 feature map corresponds to a feature vector in a window of size [5, 5] in its 1 / 2 feature map. If the i-th feature vector in Figure 0 is Figure 1 The jth feature vector in is the matching vector, then the feature vectors in the window corresponding to i and j are taken from the 1 / 2 feature map corresponding to the two images, recorded as Fw0 and Fw1, with a shape of [5,5,128]. Select the vector located at the center of the window in Fw0, and calculate the similarity score with all the feature vectors in Fw1 as the weight, multiply the similarity scores of all feature vectors in Fw1 with their two-dimensional coordinate offset relative to the center of the window, calculate the final two-dimensional offset value [Δx, Δy], add the offset value to P1, and obtain the final two-dimensional coordinates of the matching point [s*j / W+Δx,s*(j%W)+Δy], realizing the conversion of the integer pixel→integer pixel matching result into the integer pixel→sub-pixel matching result.

[0073] Using 900 scenes from the indoor dataset Scannet, with the same parameters, different 3D position encoding methods were used: camera ray-based and 3D point-based position encoding, and the same 10 iterations of training were performed. The attitude errors of the two models were evaluated by indicators.

[0074] The experiment calculates the AUC of the pose error at the thresholds (5, 10, 20), and the pose error is defined as the maximum value of the angle error in rotation and translation. The AUC indicator of the pose error is expressed as a percentage. The pose estimation error in the ScanNet indoor scene is shown in the following table:

[0075]

[0076] Experiments show that the quality of matching points obtained using a three-dimensional position encoding algorithm based on three-dimensional point coordinates is much better than the three-dimensional position encoding method used in the prior art.

[0077] like Figure 2 As shown, the feature point matching method based on three-dimensional point position coding includes the following steps:

[0078] The specific algorithm is as follows:

[0079] (1) The relative depth estimation module first processes the [3, H, W] RGB image through the DepthAnytingv2 model to generate a [1, H, W] scene relative depth value, which is a floating point number in the range [0, 1]. The [0, 1] floating point number is then mapped to a floating point number in the range [0, 255] so that it can correspond to the number of embedded codes set by the learnable embedding layer, which is 256.

[0080] (2) Establish a learnable embedding layer, which is responsible for mapping integers in the range [0, 255] to feature vectors with a dimension of 256. Then, using the linear interpolation method, the relative depth values ​​in the range [0, 255] after linear mapping are converted into the weighted feature vectors corresponding to adjacent integers to obtain the relative depth guidance feature, which is used as the scene geometry prior information to be combined with the image features to enhance the depth estimation capability of the depth estimation module;

[0081] (3) The relative depth-guided feature corresponding to each pixel is multiplied element-by-element with the image feature of the same dimension output by the image feature extraction module, and then the combined features are linearly transformed through a learnable linear layer to further enhance the feature expression capability, and the output of the linear layer is used as the fusion feature;

[0082] (4) The fused features are further extracted through the established CNN convolution module, so that the generated features are suitable for depth estimation tasks;

[0083] (5) The features obtained in step (4) are respectively passed through two branches for depth estimation: a direct depth estimation module and a depth distribution probability module, so that the depth estimation module can retain the advantages of the two depth estimation methods and dynamically adjust the weights of the two methods according to the characteristics of the scene, thereby making the depth estimation generated by the model in complex scenes more reliable;

[0084] Therefore, each pixel position can obtain a direct depth estimation value and an indirect depth distribution probability. This algorithm divides the depth range of [0,15] meters into 64 intervals on average. The above probability distribution estimates the corresponding probability for each depth interval. The midpoint value of each interval is used as the depth value of the interval, which is multiplied and added with the probability value corresponding to each interval. Finally, an expected depth value is obtained as the result of the indirect depth estimation branch.

[0085] Set up a learnable weight α and transform the depth results D obtained by the above two branches r , Dp Using the formula D = α * D r +(1 - α) * D p , the final estimated depth is calculated. The loss is calculated using the true depth values of the dataset for the estimation results obtained from the two depth estimation branches to supervise the training of the module parameters. For the direct depth estimation branch, the mean of the L1 norm between each pixel's estimated value and the true depth value is calculated; for the depth interval distribution probability estimation branch, the formula is used:

[0086] loss df = ∑α(1 - p i ) γ logp i

[0087] The sum of the depth estimation losses corresponding to each pixel's position is calculated, where p i represents the probability estimated for each pixel position in the correct depth interval, and α and γ are set hyperparameters. The final depth estimation loss is the sum of the two;

[0088] (6) Multiply the homogeneous form [x, y, 1] of the pixel's two - dimensional coordinates by the corresponding depth value d to obtain a column vector:

[0089] p h = [x * d, y * d, d] T

[0090] Using the formula p 3d = K -1 p h , that is, multiply the inverse of the camera's internal parameter matrix by the vector p h , to back - project the two - dimensional pixel coordinates into the three - dimensional camera coordinate system. Using the relative pose information of the cameras corresponding to the two images, that is, the rotation matrix R and the translation matrix t, according to the formula:

[0091] p 3d ’ = Rp 3d + t = [x, y, z] T

[0092] Transform the three - dimensional space coordinates corresponding to each pixel of the two images into the same coordinate system;

[0093] (7) For a spatial coordinate p 3d = [x, y, z] T , perform a normalization operation using the set maximum and minimum values in each coordinate direction:

[0094]

[0095] where u and v represent the two - dimensional pixel coordinates. For the x - coordinate of a certain three - dimensional point, transform it into a 128 - dimensional vector using sine - cosine encoding:

[0096] x→[sin(x / 10000^(0 / 128)),cos(x / 10000^(0 / 128)),sin(x / 10000^(2 / 128)),co s(x / 10000^(2 / 128)),…,sin(x / 10000^(128 / 128)),cos(x / 10000^(128 / 128))].

[0097] Similarly, the 3D points can be converted into position features of 128*3 dimensions. A learnable linear layer is set up to map the above position feature dimensions to 256 to achieve alignment with the image feature dimensions;

[0098] (8) The obtained three-dimensional position features and image features are added element by element as the input of the Transformer feature encoding module to generate image features F3d0 and F3d1 containing three-dimensional spatial information and the shapes are [L0, C] and [L1, C] respectively, where L0 = H0*W0 and L1 = H1*W1, i.e., the number of feature vectors corresponding to each image, and C represents the number of feature channels;

[0099] (9) Send the result obtained in the previous step to the matching module to calculate the graph 0 and Figure 1 The cosine similarity of image features F3d0 and F3d1 is used to obtain a similarity matrix of shape [L0, L1]. A threshold t is set and the mutual nearest neighbor algorithm is used to screen out all matching pairs that meet the following two conditions in the similarity matrix, and the feature index is converted into a two-dimensional coordinate as a coarse matching result. Based on the coarse matching result, the feature vector in the 5×5 window at the corresponding position is extracted in the 1 / 2 feature map. Similarly, the similarity score of the feature vector is calculated to obtain the fine matching result in the window. After the two matching stages of coarse matching and fine matching, the final two-dimensional matching point coordinates are obtained.

[0100] The present invention optimizes the three-dimensional spatial information encoding method and depth estimation module of the prior art solution, and introduces a hybrid depth estimation module guided by relative depth information and a three-dimensional point position encoding module into the feature point matching model based on deep learning, so as to enhance the model's feature extraction capability for spatial information. The features extracted by the convolutional neural network (CNN) have a limited receptive field, and the extracted features are local. Experiments show that the Transformer has a global receptive field by utilizing the attention mechanism. In the image feature matching model, the Transformer gives the model the ability to associate two features at any position in the image, thereby improving the matching quality of the model. Position encoding makes the features output by the Transformer related to the position, thereby improving the matching performance of the model when facing weak texture areas where the features are not significant.

[0101] The present invention encodes the actual three-dimensional spatial position of pixels, and adds position features based on three-dimensional point position encoding to the image features before inputting the image features into the Transformer for feature encoding operation under the attention mechanism, so that the feature point matching neural network can generate more robust image features in scenarios such as camera rotation and viewing angle changes.

[0102] The hybrid depth estimation module guided by relative depth information proposed in the present invention is the key to realizing three-dimensional point coordinate encoding. The present invention does not need to provide an actual depth map at input, but predicts the depth information of the scene through a depth estimation module based on deep learning, and introduces the scene relative depth information generated by the depthanythingv2 monocular deep network model into the image features to help the hybrid depth estimation module generate a true scene depth value. The hybrid depth estimation module obtains the image features extracted by the convolutional neural network and containing relative depth guidance information as input, and performs direct depth estimation and distribution probability estimation of different depth intervals respectively. According to the obtained probability in each depth interval, the corresponding depth expectation value can be obtained. By setting up a learnable weight parameter, the two depth estimation results can be weighted as the final depth estimation.

[0103] The present invention uses the Transformer encoder to perform global information interaction under the attention algorithm on the three-dimensional spatial features corresponding to the two images, and uses the encoding results for dense matching. Figure 1 The feature similarity of each pixel feature is calculated pair by pair, and the matching item with a high confidence score is selected from the above dense matching pairs as the final matching result. The overall block diagram of the model is shown in the figure below:

[0104] The present invention improves the depth estimation module design to improve the depth estimation performance. By using the estimated depth value and camera parameter information and the proposed three-dimensional point position encoding method to introduce more reliable three-dimensional point position features into the image features, the feature point matching model can still obtain high-quality feature point matching results by extracting three-dimensional space information with strong robustness when facing matching scenes with camera rotation, weak texture or large angle changes.

[0105] The present invention obtains the position of image pixels in three-dimensional space through an improved depth estimation module, and generates position features containing reliable three-dimensional spatial information through back projection, coordinate system conversion, normalization, sine-cosine encoding and linear layer mapping, so that the Transformer-based feature encoder can generate image features with strong robustness in scenarios such as camera rotation and viewing angle change, thereby improving the quality of feature point matching results.

Claims

1. A feature point matching method based on three-dimensional point position coding, characterized in that: The steps include: Step 1: Estimate the relative depth of the input image and encode it to obtain depth-guided features; Step 2: Fusing the depth-guided features with the image features of the input image, and extracting features of the depth estimation task of the image through a convolutional network; Step 3: perform direct depth estimation and depth distribution probability estimation on the features of the depth estimation task respectively, and weight the estimated results to obtain the estimated true depth; Step 4: Encode the three-dimensional point position using the estimated true depth to obtain a three-dimensional position feature; Step 5: Add the three-dimensional position features and image features element by element to generate image features containing three-dimensional spatial information, and send the image features containing three-dimensional spatial information to the matching module to obtain the final two-dimensional matching point coordinates.

2. A feature point matching method based on three-dimensional point position coding according to claim 1, characterized in that: In the step 1, an RGB image of the scene is obtained as an input image, and a relative depth value of the scene is generated by a relative depth estimation module; the relative depth value of the scene is converted into a weighted feature vector corresponding to adjacent integers after linear mapping to obtain a relative depth guidance feature.

3. The feature point matching method based on three-dimensional point position coding according to claim 2, characterized in that: The step 1 is specifically as follows: First, the RGB image of the shape [3, H, W] is processed by the relative depth estimation module DepthAnytingv2 model to generate the scene relative depth value of the shape [1, H, W], whose value is a floating point number in the range [0, 1]. Then, the above [0, 1] floating point number is mapped to a floating point number in the range [0, 255]. The learnable embedding layer is responsible for mapping integers in the range [0, 255] into feature vectors with a dimension of 256. Then, the linear interpolation method is used to convert the linearly mapped relative depth values ​​in the range [0, 255] into the weighted feature vectors corresponding to adjacent integers to obtain the relative depth guidance feature. The relative depth guidance feature is used as the scene geometry prior information to be combined with the image features to enhance the depth estimation capability of the depth estimation module.

4. The feature point matching method based on three-dimensional point position coding according to claim 1, characterized in that: The step 2 is specifically as follows: The RGB image outputs image features through the image feature extraction module, the relative depth guide feature is multiplied element by element with the output image feature of the same dimension, and then a linear transformation is performed through a learnable linear layer to obtain a fusion feature; The fused features are passed through a CNN convolution module for feature extraction to generate features for depth estimation tasks.

5. The feature point matching method based on three-dimensional point position coding according to claim 1, characterized in that: In the step 3, the features used for the depth estimation task are respectively passed through two branches, namely a direct depth estimation module and a depth distribution probability module, for depth estimation: the depth estimation module is able to retain the advantages of the two depth estimation methods, dynamically adjust the weights of the two methods according to the characteristics of the scene, and obtain the final depth estimation by weighting the two depth estimation results for three-dimensional coordinate calculation and subsequent three-dimensional point position encoding.

6. The feature point matching method based on three-dimensional point position coding according to claim 5, characterized in that: The step 3 is specifically as follows: The pixel position in each RGB image can obtain the direct depth estimation value and the indirect depth distribution probability respectively. The depth range of [0,15] meters is evenly divided into 64 intervals. The above probability distribution estimates the corresponding probability for each depth interval. The midpoint value of each interval is used as the depth value of the interval, which is multiplied and added with the probability value corresponding to each interval. Finally, a depth expectation value is obtained as the result of the indirect depth estimation branch. Set up a learnable weight α and transform the depth results D obtained by the above two branches r , D p According to the formula D = α * D r +(1-α)*D p , calculate the final estimated depth; Use the true depth value of the dataset to calculate the loss of the estimated results of the two depth estimation branches to supervise the training of module parameters; For the direct depth estimation branch, the L1 norm mean of the estimated value and the true depth value of each pixel is calculated; for the depth interval distribution probability estimation branch, the formula is used: loss df =∑α(1-p i ) γ logp i Calculate the sum of the depth estimation losses at each pixel corresponding to the position, where p i Represents the probability of each pixel position being estimated in the correct depth interval. α and γ are the set hyperparameters, and the final depth estimation loss is the sum of the two.

7. The feature point matching method based on three-dimensional point position coding according to claim 1, characterized in that: In step 4, the homogeneous form of the two-dimensional coordinates of each pixel in the image is multiplied by the corresponding depth value to obtain a column vector: the two-dimensional pixel coordinates are back-projected into the three-dimensional camera coordinate system, and the three-dimensional space coordinates corresponding to each pixel of the two images are converted to the same coordinate system using the relative camera pose information corresponding to the two images; For a three-dimensional space coordinate, the maximum and minimum values ​​in each coordinate direction are used for normalization, and the three-dimensional position feature is obtained after sine and cosine position encoding. Then, the linear encoding layer is used to align the three-dimensional position feature with the image feature dimension. The step 4 is specifically as follows: Multiply the homogeneous form of the pixel's two-dimensional coordinates [x, y, 1] by the corresponding depth value d to obtain a column vector: p h =[x*d,y*d,d] T Using the formula p 3d =K -1 p h , that is, let the inverse of the camera intrinsic parameter matrix be multiplied by vector p h , back-project the two-dimensional pixel coordinates into the three-dimensional camera coordinate system, and use the relative camera pose information corresponding to the two images, namely the rotation matrix R and the translation matrix t, according to the formula: p 3d '=Rp 3d +t=[x,y,z] T ; For a three-dimensional space coordinate p 3d =[x,y,z] T , use the maximum and minimum values ​​in each coordinate direction to perform normalization: Among them, u and v represent the two-dimensional coordinates of the pixel. For the x coordinate of a certain three-dimensional point, sine and cosine encoding is used to convert it into a 128-dimensional vector: Similarly, the three-dimensional points are converted into position features of 128*3 dimensions, and a learnable linear layer is set up to map the above position feature dimensions to 256 to achieve alignment with the image feature dimensions.

8. The feature point matching method based on three-dimensional point position coding according to claim 1, characterized in that: The step 5 is specifically as follows: The obtained three-dimensional position features are added element by element to the image features output by the image feature extraction module as the input of the Transformer feature encoding module to generate image features containing three-dimensional spatial information; Image features containing three-dimensional spatial information, the shapes are [L0, C], [L1, C], respectively, image features F3d0, F3d1, where L0 = H0*W0, L1 = H1*W1, i.e., the number of feature vectors corresponding to each image, and C represents the number of feature channels; The matching module calculates the cosine similarity of the two images to be matched, Figure 0 and Figure 1, and the image features F3d0 and F3d1 of Figure 0 and Figure 1, and obtains a similarity matrix with a shape of [L0, L1]. It sets a threshold t and uses the mutual nearest neighbor algorithm to screen out all matching pairs that meet the following two conditions in the similarity matrix, and converts the feature index into a two-dimensional coordinate as a coarse matching result. Based on the coarse matching result, the feature vector in the corresponding position window is extracted from the feature map. Similarly, the fine matching result within the window is obtained by calculating the similarity score of the feature vector. After performing the two matching stages of coarse matching and fine matching, the final two-dimensional matching point coordinates are obtained.

9. A feature point matching system based on three-dimensional point position coding, characterized in that: It includes relative depth estimation module, image feature extraction module, true depth estimation module, 3D point position encoding module, Transformer feature encoding module and feature matching module; The relative depth estimation module is responsible for performing monocular relative depth estimation on the two RGB image scenes to obtain relative depth values; The image feature extraction module is used to extract features from the relative depth value output by the relative depth estimation module to obtain image features; The relative depth value and image features are used as inputs to the real depth estimation module, which is responsible for combining the relative depth value with the image features and outputting the depth features. The 3D point position encoding module is used to convert the depth features into 3D point position features; The Transformer feature encoding module converts 3D point position features and depth features into image features containing 3D spatial information; The feature matching module is used to match image features containing three-dimensional spatial information.

10. An application of a feature point matching method based on three-dimensional point position coding, characterized in that: Feature point matching methods based on 3D point position encoding are used for 3D reconstruction, image stitching and panorama generation, as well as visual SLAM.

Citation Information

Cited By

  • Depth estimation method and device and computer equipment

    CN121391953A

  • Depth estimation method, apparatus, and computer device

    CN121391953B