Image feature matching model, estimation method and system based on space geometric constraint
By introducing spatial geometric constraints in the image feature matching model, the problem of insufficient matching accuracy and robustness in the prior art in strong light and low texture scenarios is solved, and a higher precision camera pose estimation is achieved.
Patent Information
- Application Number
- CN202510160395.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art lacks accuracy and robustness of image feature matching in strong light transformation and low-texture image scenes, especially in camera pose estimation tasks.
A image feature matching model based on spatial geometric constraints is proposed, including feature extraction module, differentiable soft sampling module and spatial pose regression module. This model uses spatial geometric relationships to improve the robustness and matching accuracy of image features by introducing three-dimensional spatial affine transformation information.
It significantly improves the accuracy and robustness of image feature matching, especially in severe lighting changes or low texture environments, improving the accuracy of camera pose estimation, showing higher F1 scores, internal point proportions and average polar line errors.
Smart Images

Figure CN120088514A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and visual navigation, and particularly relates to an image feature matching model, an estimation method and a system based on spatial geometric constraints. Background Art
[0002] Image feature extraction and matching are algorithms for finding corresponding relationships of similar structures from image pairs, and are widely used in advanced vision tasks such as Simultaneous Localization and Mapping (SLAM), Structure from Motion (SFM), and three-dimensional reconstruction. As the basis for calculating three-dimensional space information from two-dimensional image features, the accuracy of the results of image feature extraction and matching algorithms will directly affect subsequent advanced vision tasks. Although the problem of image feature extraction and matching has attracted extensive research by researchers, most of the existing research represents features by calculating image gray information, resulting in a serious reduction in the matching accuracy of the algorithm on strongly illuminated transformed and low-texture image pairs, and the robustness has not yet reached the ideal performance.
[0003] Currently, the most advanced methods in the research direction of image feature matching in extreme scenarios use various deep learning techniques, combined with a variety of innovative methods to improve the accuracy and robustness of image feature matching. Chinese Patent with Application No. 202410875933.9 discloses an image stitching processing system based on feature matching, which uses a VGG encoder and two decoders to perform feature extraction and description respectively, and completes feature modeling through the Transformer method, and combines the nearest neighbor points to complete feature matching and outlier removal. However, this patent does not optimize for strong illumination transformation and low-texture image scenarios, and its robustness needs to be improved. Chinese Patent with Application No. 202411084909.X discloses an aviation vision navigation positioning method based on semantic topological feature matching. By constructing a high-level semantic topological relationship feature vector, the distance ratio, angle difference, and digital semantic label between the reference object in the image and the image center are encoded as a feature vector to accurately describe the semantic topological structure between visual reference objects, and realize the matching and positioning of the route image and the reference database. This patent adds prior information to the field of aviation vision navigation positioning, but it is difficult to obtain semantic information of images in extreme environments and it is difficult to be widely applied in other fields. Chinese Patent 202410967253.X discloses a monocular SLAM image feature matching method based on the improvement of the ORB algorithm. Through a multi-level threshold screening and feature optimization process, including using FAST corner detection, Harris response value, Gaussian pyramid, and rotation-invariant BRIEF descriptor, combined with the RANSAC algorithm, the optimal matching result is output. This method effectively improves the efficiency and accuracy of feature matching, and provides an efficient and reliable image matching solution for the monocular SLAM system. This patent is optimized for the edge deployment of SLAM and has completed the landing inspection of actual projects. However, the final accuracy is not ideal compared with the deep learning method, and it is suitable for deployment in scenarios with low accuracy requirements. Chinese Patent with Application No. 202411123523.5 discloses a method and device for UAV positioning based on image feature matching. By changing the style of the original sub-map of the target map, a variety of generalized sub-maps are generated, and combined with feature extraction technology, the initial features of each sub-map are extracted, the common features are screened out, and a robust feature group is formed and stored in the database. In practical applications, the UAV realizes accurate matching and positioning by comparing the real-time image with the robust feature group in the database. This method can provide higher matching accuracy in complex environments or under adverse conditions, thus enhancing the reliability and robustness of UAV positioning. However, the feature extraction process of this patent still relies on gray information, and it is easy to produce navigation drift when the UAV's perspective changes greatly.
[0004] Traditional image feature extraction and matching are based on the local gray-scale features of images for feature extraction and matching, and combined with the image gray-scale pyramid model to increase the receptive field of the local feature extractor and thus improve the robustness of the algorithm. In recent years, researchers have applied deep learning methods to image feature extraction and matching tasks, using the more powerful non-linear fitting capabilities of neural network structures such as CNN and Transformer to complete a more dense feature matching function, as Figure 3 shown. However, it should be noted that the deep learning-based matching method also directly encodes image features based on image gray scale, and there are many incorrect matches in the camera pose estimation task with a large spatial perspective transformation, seriously affecting the accuracy of camera pose estimation. Summary of the Invention
[0005] The present invention precisely aims at the problems existing in the prior art, and provides an image feature matching model, estimation method and system based on spatial geometric constraints, which at least includes a feature extraction module, a differentiable soft sampling module and a spatial pose regression module; the feature extraction module is composed of a gray-scale feature pyramid encoder, a global coarse feature decoder and a sub-pixel fine matcher; wherein the gray-scale feature pyramid encoder uses a U-Net structure based on a convolutional neural network as the backbone network; the global coarse feature decoder at least includes a global attention module; the sub-pixel fine matcher at least includes a local attention module; the feature extraction module uses the U-Net backbone network as the feature encoder to output feature maps at different scales, combines the global attention module as the feature decoder to obtain a preliminary matching result, and adjusts the matching result through the local attention module to achieve a high-precision sub-pixel matching relationship; the differentiable soft sampling module introduces epipolar geometry constraints during the training of the model, end-to-end supervises the process of feature point extraction and matching, and ensures the differentiability of the sampling process from two-dimensional matching to three-dimensional pose estimation; the spatial pose regression module is used to regress the fundamental matrix for camera pose estimation, and the fundamental matrix describes the epipolar geometry relationship between corresponding points in two images, provides an algebraic expression form for the matching point constraint, and is the basis for camera pose estimation and three-dimensional reconstruction. The method of the present invention can directly output the camera pose and shows higher accuracy in terms of indicators such as F1 score, inlier ratio and average epipolar line error.
[0006] To achieve the above object, the technical solution adopted by the present invention is: an image feature matching model based on spatial geometric constraints, which at least includes a feature extraction module, a differentiable soft sampling module and a spatial pose regression module;
[0007] The feature extraction module: is composed of a gray-scale feature pyramid encoder, a global coarse feature decoder and a sub-pixel fine matcher;
[0008] The grayscale feature pyramid encoder uses a U-Net structure based on a convolutional neural network as the backbone network;
[0009] The global coarse feature decoder includes at least a global attention module;
[0010] The sub-pixel fine matcher includes at least a local attention module;
[0011] The feature extraction module uses the U-Net backbone network as a feature encoder to output feature maps at different scales, combines the global attention module as a feature decoder to obtain a preliminary matching result, and adjusts the matching result through the local attention module to achieve a high-precision sub-pixel matching relationship;
[0012] The differentiable soft sampling module: introduces epipolar geometry constraints during the training of the model, end-to-end supervises the process of feature point extraction and matching, and ensures the differentiability of the sampling process from two-dimensional matching to three-dimensional pose estimation;
[0013] The spatial pose regression module: regresses to obtain the fundamental matrix for camera pose estimation. The fundamental matrix is used to describe the epipolar geometry relationship between corresponding points in two images, maps points in the first image to epipolar lines in the second image, thereby providing an algebraic expression form for matching point constraints, and is the basis for camera pose estimation and three-dimensional reconstruction.
[0014] As an improvement of the present invention, the grayscale feature pyramid encoder uses a U-Net structure based on a convolutional neural network as the backbone network to extract image grayscale features, specifically:
[0015] F r &F p ∈R H′×W′×C =UNet(I),I∈R H×W×3
[0016] Wherein, I represents the input image, with a width of W, a height of H, and 3 channels; UNet represents a U-Net structure based on a convolutional neural network for feature extraction; F r ∈H / 8×W / 8×64 represents the extracted rough feature map; F p ∈H / 2×W / 2×64 represents the extracted fine feature map, with a width of W’, a height of H’, and C channels.
[0017] As an improvement of the present invention, the method for building the global coarse feature decoder is specifically:
[0018] Calculate the attention message M:
[0019]
[0020] Among them, M represents the feature fusion result containing the self-attention layer and the cross-attention layer, φ and are activation functions with non-negative value ranges; q, k, v correspond to the query, key, and value vectors of the attention mechanism Q, K, V, R i and R j represent rotational position encoding, and i and j represent two feature vectors for calculating attention;
[0021] The feed-forward network MLP processes the attention message M to update the features, and the process is as follows:
[0022]
[0023] Among them, represents the rough feature map of the i-th layer, and the symbol || represents the concatenation operation;
[0024] The global matching cost matrix S is calculated by measuring the inner product of the rough feature maps, representing the confidence of feature matching, and the sharpness of the matrix is controlled by combining the temperature coefficient t:
[0025]
[0026] where t is the temperature coefficient and S is the global matching cost matrix.
[0027] As another improvement of the present invention, in the sub-pixel fine matcher, the Dual-Softmax layer is used to extract the rough position relationship of feature points, and the final confidence matrix P is:
[0028] P(i,j) = softmax(S(i,·)) j ·softmax(S(·,j)) i
[0029] At each rough feature point position (i r , j r ), a grid of size w×w is divided, and the self-attention and cross-attention features are calculated for F p A and F p B respectively to generate a heat map representing the sub-pixel matching probability, and to generate a heat map representing the sub-pixel matching probability. By calculating the expectation matrix of the probability distribution, the exact matching point pairs (i p , j p ) based on the gray-scale change are obtained.
[0030] As yet another improvement of the present invention, in the differentiable soft sampling module, the probability size of the confidence of feature point pairs is calculated through the softmax function:
[0031]
[0032] Among them, t is the temperature coefficient, c j is the sampling point, G j is the Gumbel perturbation function.
[0033] As another improvement of the present invention, in the spatial pose regression module, according to the epipolar geometry constraint, there is:
[0034]
[0035] where x a and x b represent the already matched feature points, R and t represent the rotation matrix and the translation matrix respectively, F represents the fundamental matrix, and when directly expanded and written in matrix form, it is:
[0036] AF flatten = 0
[0037]
[0038] where the superscript represents 7 groups of already matched feature points, A is the coefficient matrix, and F flatten is the expanded fundamental matrix.
[0039] To achieve the above object, the technical solution adopted by the present invention is also: an image feature matching estimation method based on spatial geometric constraints, establishing an image feature matching model based on spatial geometric constraints and performing model training, and using the trained model to complete the estimation of image feature matching.
[0040] As a further improvement of the present invention, the model training of the image feature matching based on spatial geometric constraints specifically includes the following steps:
[0041] S100: Construct a feature matching image sequence data set, using the time stamp as the image serial number, and including the internal parameter matrix of the camera;
[0042] S110: Work under the supervision of the homography matrix, using the U-Net backbone network as the feature encoder to output feature maps at different scales, combining the global attention module as the feature decoder to obtain a preliminary matching result, and fine-tuning the matching result through the local attention module to achieve a high-precision sub-pixel matching relationship;
[0043] S120: Build a differentiable soft sampling module to ensure the differentiability of the entire model training process while endowing the sampler with random exploration to simulate distribution generation;
[0044] S130: Build a spatial pose regression module;
[0045] S140: Homography transformation matrix supervised feature point matching training, using Euclidean distance for sub-pixel level refinement adjustment;
[0046] S150: Fundamental matrix supervised pose estimation training. At the sub-pixel level, the average symmetric epipolar line error is used to quantify the overall deviation of the ideal epipolar line geometry relationship between two perspectives; at the relative pose level error, the pose estimation loss is defined.
[0047] To achieve the above object, the technical solution adopted by the present invention is also: an image feature matching system based on spatial geometric constraints, including a computer program, and when the computer program is executed by a processor, it implements the steps of any one of the above-mentioned estimation methods.
[0048] Compared with the prior art, the technical advantages and effects of the present invention are:
[0049] (1) The present invention proposes an image feature matching model and estimation method based on spatial geometric constraints. By introducing three-dimensional space affine transformation information and using spatial geometric relationships, the robustness and matching accuracy of image features are improved. This method embeds three-dimensional geometric information into the image feature extraction process, making the extracted features have higher description ability and stability, especially significantly improving the matching quality in environments with drastic illumination changes or low texture.
[0050] (2) The present invention proposes a differentiable pose regression module, which can be seamlessly integrated into any advanced feature extraction network. By detecting high regression state feature point pairs, it directly estimates the fundamental matrix of the camera and performs three-dimensional space pose calculation. This module is designed as an end-to-end architecture, greatly improving the robustness and adaptability of the pose estimation task, and enabling high-precision pose inference in dynamic scenes and complex environments.
[0051] (3) The present invention designs multi-level loss functions for coarse-grained matching, sub-pixel level matching, epipolar constraint matching, and pose level. These loss functions comprehensively supervise the model from the perspectives of multi-scale and multi-task, ensuring that the entire process from feature extraction to pose estimation has an efficient optimization objective, and further enhancing the model performance and application applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a structural diagram of the image feature matching model based on spatial geometric constraints of the present invention;
[0053] Figure 2 It is a flowchart of the model training of the image feature matching model based on spatial geometric constraints of the present invention;
[0054] Figure 3 It is a structural diagram of the traditional model paradigm for learning-based image feature matching;
[0055] Figure 4 This is the flowchart of the image feature matching estimation method based on spatial geometric constraints of the present invention;
[0056] Figure 5 This is the simulation sampling comparison diagram of the differentiable soft sampling module of the present invention;
[0057] Figure 6 This is the pose regression comparison diagram of the image feature matching model combined with spatial geometric constraints in the test comparative example of the present invention;
[0058] Figure 7 This is the feature matching effect comparison diagram of the image feature matching model combined with spatial geometric constraints in the test comparative example of the present invention. Detailed implementation manners
[0059] The present invention will be further clarified below in conjunction with the accompanying drawings and detailed implementation manners. It should be understood that the following detailed implementation manners are only used to illustrate the present invention and not to limit the scope of the present invention.
[0060] Embodiment 1
[0061] In the task of estimating the spatial pose of a camera, it is advanced to perform pixel-level feature matching based on the attention mechanism of image grayscale. The present invention follows the use of self-attention and cross-attention mechanisms to expand the receptive field of feature descriptors to the entire image, combines a coarse-to-fine dense matching strategy and a cross-correlation Dual-Softmax layer to extract and match more robust and dense feature points. The results directly obtained by associating image grayscale and the homography transformation matrix have a certain number of incorrect matches, and these incorrect results will generate noise interference for subsequent pose estimation tasks, accumulate regression errors in a long image sequence, and perform poorly on challenging large-viewpoint change datasets. Under the ideal optical assumption, the true values of feature point matches show geometric consistency in the corresponding relationships, while incorrect matches lack this consistency. Therefore, the present invention discloses an image feature matching model based on spatial geometric constraints.
[0062] As Figure 1 shown, an image feature matching model based on spatial geometric constraints includes a feature extraction module, a differentiable soft sampling module, and a spatial pose regression module. For the image pair to be matched, the feature extraction module is used to detect feature points to solve the repeatability problem of feature points and maintain system stability; the differentiable soft sampling module is used to ensure the differentiability of the sampling process from two-dimensional matching to three-dimensional pose estimation, thereby ensuring the end-to-end backpropagation process; finally, based on the spatial pose regression module, the fundamental matrix for camera pose estimation is regressed.
[0063] The image feature matching model based on spatial geometric constraints proposed by the present invention aims to achieve high-precision feature matching and camera pose estimation. Methodologically, the model expands the feature receptive field through self-attention and cross-attention mechanisms, combines a coarse-to-fine feature matching strategy and a Dual-Softmax layer, extracts dense and robust feature points step by step, and uses epipolar geometric constraints and homography matrices to supervise the matching process, significantly reducing the cumulative error caused by false matches. At the same time, the feature extraction module constructs a gray feature pyramid encoder with a U-Net backbone, and combines a global coarse feature decoder and a sub-pixel fine matcher to optimize the matching relationship. On this basis, a differentiable soft sampling module is introduced, and the discrete probability distribution is simulated through the Gumbel distribution and the Softmax function, enhancing the random exploration of the training process and avoiding overfitting. The model supports end-to-end regression to calculate the six-degree-of-freedom spatial pose of the camera, and the entire process from feature extraction to pose estimation is differentiable, ensuring optimization efficiency and accuracy. In complex scenes and extreme environments, the model achieves sub-pixel matching accuracy and robustness, providing an efficient solution for dynamic scene monitoring and 3D reconstruction.
[0064] Embodiment 2
[0065] Based on the image feature matching estimation method of spatial geometric constraints, on the basis of establishing the image feature matching model based on spatial geometric constraints in Embodiment 1, a model is established and the model is trained, as Figure 2 shown, the model training method specifically includes the following steps:
[0066] Step S100: Construct a feature matching image sequence dataset, use the timestamp as the image serial number, and include the internal parameter matrix of the camera.
[0067] Step S110: It works under the supervision of the homography matrix, uses the U-Net backbone network as the feature encoder to output feature maps at different scales. Combine the global attention module as the feature decoder to obtain the preliminary matching result, and fine-tune the matching result through the local attention module to achieve a high-precision sub-pixel matching relationship. In particular, the feature extraction module consists of the following three modules: a gray feature pyramid encoder, a global coarse feature decoder, and a sub-pixel fine matcher. This modular design can effectively improve the accuracy and robustness of feature matching, thus providing high-quality matching results for subsequent tasks.
[0068] S111: Adopt the U-Net structure based on the convolutional neural network as the backbone network to extract the image gray feature, which helps the model understand the overall gray feature of the image.
[0069] F r &F p ∈R H′×W′×C =UNet(I),I∈RH×W×3
[0070] Among them, I represents the input image, with width W, height H, and 3 channels; UNet represents the U-Net structure based on convolutional neural network for feature extraction; F r ∈H / 8×W / 8×64 represents the extracted rough feature map; F p ∈H / 2×W / 2×64 represents the extracted fine feature map, with width W', height H', and C channels.
[0071] S112: The method for building the global rough feature decoder includes: calculating the attention message M:
[0072]
[0073] Among them, M represents the feature fusion result including the self-attention layer and the cross-attention layer, φ and are activation functions with non-negative value ranges; q, k, v correspond to the query, key, and value vectors of the attention mechanism Q, K, V, R i and R j represent the rotational position encoding, and i and j represent two feature vectors for calculating attention;
[0074] The multi-layer perceptron (MLP) processes the attention message M to update the features, and the process is as follows:
[0075]
[0076] Among them, represents the rough feature map of the i-th layer, and the symbol || represents the concatenation operation.
[0077] Calculate the global matching cost matrix S by measuring the inner product of the rough feature maps, which characterizes the confidence of feature matching, and combine the temperature coefficient t to control the sharpness of the matrix:
[0078]
[0079] where t is the temperature coefficient and S is the global matching cost matrix.
[0080] S113: To ensure the accuracy of feature matching, build a sub-pixel fine matcher, use the Dual-Softmax layer to complete the extraction of the rough position relationship of feature points, and the final confidence matrix P is:
[0081] P(i,j) = softmax(S(i,·)) j ·softmax(S(·,j)) i
[0082] At each rough feature point position (ir , j r ) Divide a grid of size w×w around it. Within this local window, for F p A and F p B Calculate self-attention and cross-attention features, generate a heatmap representing the sub-pixel matching probability, and generate a heatmap representing the sub-pixel matching probability. Finally, by calculating the expected matrix of the probability distribution, obtain the exact matching point pairs (i p , j p ) based on the gray-scale change.
[0083] Step S120: Build a differentiable soft sampling module, introduce an explicit constraint from 2D to 3D during the training of the model, i.e., the epipolar geometry constraint, and end-to-end supervise the process of feature point extraction and matching. While ensuring the differentiability of the entire model training process, endow the sampler with a certain degree of random exploration to simulate distribution generation, so as to avoid model overfitting.
[0084] Given the confidence sequence of the matching feature point pairs Extract k minimum sample sequences as (c 1 , c 2 , …, c k ). Since the extreme values of the exponential family distribution all follow the Gumbel distribution, its probability density function is:
[0085]
[0086] where z = (x - μ) / β, so add a random Gumbel distribution to the sample sequence to simulate the discrete probability distribution:
[0087]
[0088] Finally, calculate the probability of the confidence of the feature point pairs through the softmax function:
[0089]
[0090] where t is the temperature coefficient. As the model is trained, the temperature coefficient t decreases from large to small, so that the calculation result tends to the One-Hot distribution more, solving the problem of the difference between forward propagation and backward propagation. When the temperature coefficient is large, the output is closer to the uniform distribution; when the temperature coefficient is small, the output is closer to the One-Hot distribution, as Figure 5As shown, this patent defines a multinomial distribution, plots its true probability density function, and simulates the effectiveness of different sampling methods. The abscissa represents a multinomial distribution with ten categories, and shows its true density function and the density functions obtained after adding different types of noise. The experimental results show that when sampling with Gumbel noise, the obtained samples are closest to the true distribution, and its accuracy is better than that of sampling with normal distribution noise. The result of sampling with uniform distribution noise is the worst. It can be seen that the Gumbel distribution can be used in the reparameterization method to ensure that the generated sample points can accurately match the target distribution while maintaining the differentiability of the computational graph, thus realizing a more effective sampling process.
[0091] Step S130: Build a spatial pose regression module: According to the epipolar geometric constraint, we have:
[0092]
[0093] where x a and x b represent the matched feature points, R and t represent the rotation matrix and the translation matrix respectively, F represents the fundamental matrix, and expanding and writing it in matrix form gives:
[0094] AF flatten = 0
[0095]
[0096] where the superscript represents 7 groups of matched feature points, A is the coefficient matrix, and F flatten is the expanded fundamental matrix.
[0097] Step S140: Homography transformation matrix supervised feature point matching training: The output of the global coarse feature decoder is the confidence matrix P. Therefore, calculate the negative log-likelihood loss based on the confidence, and define the true coarse matching relationship M as the nearest neighbor relationship obtained by comparing two sets of grids at 1 / 8 resolution. In this process, the distance between any two grids is evaluated by the reprojection distance of their respective center points.
[0098]
[0099] In this embodiment, the Euclidean distance is used for sub-pixel level refinement. For each query point i, the uncertainty of the point is quantified by calculating the sum of the heat map variances. The goal of the present invention is to preferentially refine the positions with lower uncertainty, thus introducing a final weighted loss function (Weighted Loss Function) to optimize the model performance:
[0100]
[0101] Among them, L c and L f are the loss functions on the rough matching and fine matching scales of the features respectively, P is the confidence matrix output by the double softmax matching layer, and M f are the true value sets of the features in the rough matching and fine matching modes respectively, σ 2 (i) is the variance of the heat map corresponding to feature point i, j' and j' gt are the homography deviation value and the true value of feature point j respectively;
[0102] Step S150: Fundamental matrix supervised pose estimation training: At the sub-pixel level, the average symmetric epipolar error is used to quantify the overall deviation of the ideal epipolar geometry relationship between two views:
[0103]
[0104] L is the inlier set selected by the true model θ, Φ i is the relevant geometric feature or feature under the specific match i, ∈ epi represents the corresponding residual error, which is used to quantify the deviation of the ideal or expected relationship based on these geometric features.
[0105] At the relative pose level error, the pose estimation loss is defined as follows:
[0106]
[0107] Among them, ∈ R and ∈ t represent the rotation error and the translation error respectively.
[0108] After completing the model training, use the trained model to complete the estimation of image feature matching, as Figure 4 shown, which specifically includes the following steps:
[0109] Step S200: Construct a feature matching image sequence dataset; specifically, it can be shot by a drone, a ground mobile platform or other sensing devices in an extreme environment to collect representative image data, so as to obtain the image sequence to be matched. During the collection process, each image is attached with a timestamp as the sequence number, and at the same time, the internal parameter matrix of the camera (such as focal length, principal point coordinates, distortion coefficient, etc.) is recorded to ensure that the data has high consistency and precise geometric constraint relationships. Each image sequence includes T frames, and the height of each frame of image is H and the width is W; in this embodiment, T = 150, H = 960, W = 640, ensuring that the image sequence can cover enough view changes of the target scene, and then extract image features.
[0110] Step S210: Construct a feature extraction module, which consists of the following three modules: a grayscale feature pyramid encoder, a global coarse feature decoder, and a sub-pixel fine matcher.
[0111] S211: Use the U-Net structure based on a convolutional neural network as the backbone network to extract the grayscale features of the image.
[0112] S212: The method for building the global coarse feature decoder includes calculating the cost matrix of global matching by measuring the inner product of the rough feature maps.
[0113] S213: Build a sub-pixel fine matcher and use the Dual-Softmax layer to complete the extraction of the rough position relationship of the feature points.
[0114] Step S220: Build a differentiable soft sampling module to simulate the discrete probability distribution.
[0115] Step S230: Calculate the six-degree-of-freedom three-dimensional space pose of the camera through the epipolar geometry constraint regression to complete the camera pose estimation task.
[0116] Embodiment 3
[0117] This embodiment also proposes a system for an image feature matching model based on spatial geometric constraints. By adopting the estimation method proposed in Embodiment 2, it supports local deployment and hardware acceleration implementation. The system optimizes the feature matching process by introducing geometric constraints, and cooperates with the HUAWEI Ascend series of AI processors and multi-modal sensors integration, including binocular cameras, depth cameras, lidars, IMUs, and high-precision positioning modules, to ensure the feature matching accuracy and real-time performance in complex environments. The built-in efficient data synchronization mechanism and optimized algorithm architecture of the system can significantly improve the operation efficiency of the matching model, and at the same time support rapid deployment on a variety of hardware platforms, providing a reliable technical solution for feature extraction and pose estimation in extreme scenarios.
[0118] The system for the image feature matching model based on spatial geometric constraints provided by the present invention can quickly and accurately achieve the extraction, matching, and spatial position estimation of image features, and at the same time can effectively reduce the matching error in complex environments, improve the robustness and adaptability of the model; the system supports multiple data source inputs, can process multi-view and multi-scale image information, and maintains high computational performance in extreme scenarios, providing important support for accurate reconstruction and pose inference.
[0119] Test comparative example
[0120] To evaluate the quality of the estimated fundamental matrix and camera pose, this test case used three evaluation metrics in ten challenging scenarios by extracting and matching feature points of a public dataset: F1 score, inlier ratio, and mean epipolar error (unit: pixel). The F1 score provides a balance between precision and recall, indicating the accuracy of feature correspondences. The inlier ratio reflects the proportion of correctly matched key points, which is crucial for robust pose estimation. The mean epipolar error measures the average distance between the projected points and the corresponding epipolar lines, providing an evaluation of the geometric consistency of the estimated fundamental matrix. The above metrics provide a comprehensive assessment of the precision and robustness of the model in estimating camera pose and fundamental matrix. Thus, the test result metrics are as Figure 6 shown, and the test effect is Figure 7 shown.
[0121] The calculation methods of the metrics are as follows:
[0122]
[0123] where R represents different metric improvement values, n is the number of test scenarios, and r our represents the test result of the method of this patent, and r LoFTR+OpenCV-RANSAC represents the test result of the comparative method.
[0124] Among them, the average F1 score of LoFTR is:
[0125]
[0126] The average F1 score of the method of this patent:
[0127]
[0128] Percentage increase in F1 score:
[0129]
[0130] The average inlier of LoFTR:
[0131]
[0132] The average inlier of the method of this patent:
[0133]
[0134] Percentage increase in the number of inliers:
[0135]
[0136] The average mean epipolar error of LoFTR:
[0137]
[0138] Average value of the mean epipolar error of the method of this patent:
[0139]
[0140] Percentage reduction in the mean epipolar error:
[0141]
[0142] Therefore, compared with the LoFTR+OpenCV-RANSAC method, the method of this patent has an average increase of 12.4% in the F1 score, indicating a better balance between accuracy and recall in camera pose estimation. Further, the method of this patent increases the number of inliers by 23.9%, showing a higher proportion of correctly matched key points, which is crucial for robust pose estimation. In addition, the method of this patent significantly reduces the mean epipolar error of all landmarks, with an average reduction of 63.4%, highlighting its ability to maintain geometric consistency when estimating the fundamental matrix. These improvements show that the method of this patent not only provides more accurate key point correspondences but also ensures reliable camera pose estimation, thus representing a significant advancement compared to the baseline method.
[0143] The ultimate core idea of this patent is to abandon the black-box design model paradigm and provide a deep learning method designed for specific downstream tasks. While advocating modular design of deep learning methods, this patent enhances the interpretability of each module, thereby supporting the robustness of the entire model. This patent hopes that in future research, more such interpretable module designs for specific downstream tasks can be seen, promoting the applicability of the algorithm in real-world environments.
[0144] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. Image feature matching model based on spatial geometric constraints, characterized by: At least includes a feature extraction module, a differentiable soft sampling module and a spatial pose regression module; The feature extraction module is composed of a grayscale feature pyramid encoder, a global coarse feature decoder and a sub-pixel fine matcher; The grayscale feature pyramid encoder uses a U-Net structure based on a convolutional neural network as a backbone network; The global coarse feature decoder comprises at least a global attention module; The sub-pixel fine matcher includes at least a local attention module; The feature extraction module uses the U-Net backbone network as a feature encoder to output feature maps at different scales, combines the global attention module as a feature decoder, obtains preliminary matching results, and adjusts the matching results through the local attention module to achieve high-precision sub-pixel matching relationships; The differentiable soft sampling module: introduces epipolar geometry constraints in the process of training the model, supervises the feature point extraction and matching process end-to-end, and ensures the differentiability of the sampling process from two-dimensional matching to three-dimensional pose estimation; The spatial pose regression module regresses the basic matrix of the camera pose estimation, which describes the extreme geometric relationship between corresponding points in two images and provides an algebraic expression for matching point constraints.
2. The image feature matching model based on spatial geometric constraints as claimed in claim 1, characterized in that: The grayscale feature pyramid encoder uses the U-Net structure based on convolutional neural network as the backbone network to extract image grayscale features, specifically: F r &F p ∈R H′×W′×C =UNet(I),I∈R H×W×3 Where I represents the input image, with a width of W, a height of H, and 3 channels; UNet represents the U-Net structure based on the convolutional neural network, which is used for feature extraction; F r ∈H / 8×W / 8×64 represents the extracted coarse feature map; F p ∈H / 2×W / 2×64 represents the extracted fine feature map, whose width is W', height is H', and number of channels is C.
3. The image feature matching model based on spatial geometric constraints as claimed in claim 1, characterized in that: The specific method of building a global coarse feature decoder is: Calculate the attention message M: Among them, M represents the feature fusion result including the self-attention layer and the cross-attention layer, φ and is an activation function with non-negative range; q, k, v correspond to the query, key and value vectors of the attention mechanism Q, K, V, R i and R j represents the rotation position encoding, i and j represent the two feature vectors for calculating attention; The feedforward network MLP processes the attention message M to update the features as follows: in, represents the coarse feature map of the i-th layer, and the symbol || represents the concatenation operation; The cost matrix S of global matching is calculated by measuring the inner product of the coarse feature map to characterize the confidence of feature matching, and the sharpness of the matrix is controlled by combining the temperature coefficient t: Where t is the temperature coefficient and S is the cost matrix of global matching.
4. The image feature matching system based on spatial geometric constraints as claimed in claim 1, characterized in that: In the sub-pixel fine matcher, the Dual-Softmax layer is used to complete the rough position relationship extraction of feature points, and the final confidence matrix P is: P(i,j)=softmax(S(i,·)) j ·softmax(S(·,j)) i At each coarse feature point position (i r ,j r ) is divided into w×w grids around and Calculate the self-attention and cross-attention features to generate a heat map representing the sub-pixel matching probability. By calculating the expected matrix of the probability distribution, the exact matching point pairs based on the grayscale change (i p ,j p ).
5. The image feature matching system based on spatial geometric constraints as claimed in claim 1, characterized in that: In the differentiable soft sampling module, the probability of the confidence of the feature point pair is calculated by the softmax function: Where t is the temperature coefficient, c j is the sampling point, G j is the gumbel perturbation function.
6. The image feature matching system based on spatial geometric constraints as claimed in claim 1, characterized in that: In the spatial pose regression module, according to the epipolar geometry constraints: where x a and x b Describe the matched feature points, R and t represent the rotation matrix and translation matrix respectively, F represents the basic matrix, which can be directly expanded and written in matrix form as follows: OF flatten =0 The superscripts represent 7 sets of matched feature points, A is the coefficient matrix, and F flatten is the unfolded basic matrix.
7. An image feature matching estimation method based on spatial geometric constraints, using the model as claimed in claim 1, characterized in that: Under the supervision of a multi-level loss function, an image feature matching model based on spatial geometric constraints is established and model training is performed, and the image feature matching estimation is completed using the trained model; the multi-level loss function at least includes coarse-grained matching, sub-pixel matching, epipolar constraint matching, and pose-level multi-level loss functions, wherein, The loss function for coarse-grained matching is: The loss function for sub-pixel matching is: Among them, P is the confidence matrix output by the double softmax matching layer, and M f are the feature truth value sets in the coarse matching and fine matching modes, σ 2 (i) is the variance of the heat map corresponding to feature point i, j' and j' gt The feature point j is the homography deviation value and the true value respectively; The loss function for epipolar constraint matching is: Where L is the set of interior points selected by the true model θ, Φ i is the relevant geometric feature or features under a specific match i, ∈ epi represents the corresponding residual error; The loss function at the pose level is: Among them, ∈ R and ∈ t represents the rotation error and translation error.
8. The image feature matching estimation method based on spatial geometric constraints as claimed in claim 7, characterized in that: The model training of image feature matching based on spatial geometric constraints specifically includes the following steps: S100: construct a feature matching image sequence data set, using the timestamp as the image sequence number, including the camera's internal parameter matrix; S110: Works under the supervision of the homography matrix, uses the U-Net backbone network as a feature encoder to output feature maps at different scales, combines the global attention module as a feature decoder, obtains preliminary matching results, and fine-tunes the matching results through the local attention module to achieve high-precision sub-pixel matching relationships; S120: Build a differentiable soft sampling module to ensure that the entire model training process is differentiable while giving the sampler random exploration to simulate distribution generation; S130: Build a spatial pose regression module; S140: homography transformation matrix supervises feature point matching training, using Euclidean distance for sub-pixel level refinement; S150: Fundamental matrix supervised pose estimation training. At the sub-pixel level, the average symmetric epipolar error is used to quantify the overall deviation from the ideal epipolar geometry between two views; at the relative pose level error, the pose estimation loss is defined.
9. An image feature matching system based on spatial geometric constraints, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps of the estimation method as described in any one of claims 8 to 9 are implemented.
Citation Information
Patent Citations
Unmanned aerial vehicle positioning method and device based on image feature matching
CN118644556A
Improved monocular SLAM image feature matching method based on ORB algorithm
CN118747811A
An image stitching processing system based on feature matching
CN118864237B
Aviation visual navigation positioning method based on semantic topological feature matching
CN118864906A
Cited By
LoFTR stereo matching method based on epipolar constraint
CN120525934A
Multi-modal image matching method and system based on learning features and epipolar geometric constraints
CN120726352A
A multimodal image matching method and system based on learned features and epipolar geometric constraints
CN120726352B
Visual pose estimation method based on data driving
CN121353412A
Data-driven visual pose estimation method
CN121353412B