Visual relocation method based on saliency scene coordinate regression
Through the regression network of significance scene coordinates, the significant feature points in the image are extracted and matched, and the problem that image areas cannot be used for positioning and feature matching errors in the prior art is solved, thereby achieving more efficient and accurate visual relocation.
Patent Information
- Application Number
- CN202510210624.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the existing visual repositioning methods, the image area cannot be used for positioning, and the extraction and matching features from the image area are prone to mismatch, resulting in reduced positioning accuracy and efficiency.
The visual relocation method based on significance scene coordinate regression is adopted to obtain 2D image key points through the significance prediction network, and 3D map points are obtained in combination with the hierarchical scene coordinate regression network, and the camera's six-degree of freedom pose is solved through the PnP-RANSAC pose algorithm.
This method can effectively reduce outliers, improve positioning accuracy and efficiency, reduce calculation amount, and exhibit smaller positioning errors on indoor and outdoor data sets.
Smart Images

Figure CN120147419A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual positioning. Specifically, it relates to a visual relocalization method based on saliency scene coordinate regression. Background Art
[0002] The task of visual relocalization refers to calculating the 6-degree-of-freedom pose of a given query image in a known environment. There are many structural systems according to the construction of the scene model, the matching of environmental information, and the solution of the camera pose. Currently, the known environmental expression methods can be divided into: the scheme based on 3D structural features, the scheme based on 2D image retrieval, and the scheme based on deep learning. From the perspective of the environmental information matching technology, traditional visual relocalization methods are usually based on 3D structural features. That is to say, they rely on the two-dimensional to three-dimensional mapping relationship between the two-dimensional image position and the three-dimensional scene coordinates. In such schemes, the mapping relationship is usually obtained by matching local features, and many matching and filtering technologies have been proposed on this basis, which can achieve efficient and robust city-scale positioning, such as the Active Search (AS) algorithm.
[0003] Some researchers have also begun to apply image retrieval technology to visual relocalization. They use a 2D database to represent environmental information. The pose recovery of the query image can be directly approximated by the most similar retrieved database image. Since a compact image global descriptor is used for matching, the image retrieval method can be extended to a very large environment. Among them, the descriptor of the image can be structure-based or learned using a deep neural network. The retrieval method can be combined with the structure-based method or relative pose estimation to predict a more accurate pose. Usually, the retrieval step helps to limit the search space and form a faster and more accurate positioning.
[0004] For the visual relocalization algorithm based on the mapping relationship between the two-dimensional image and the three-dimensional scene, in actual application deployment, such methods must maintain a three-dimensional model online, and the requirements for memory and computing power of this model increase with the increase of the positioning range. The descriptor matching step for obtaining the corresponding relationship is also an expensive and time-consuming procedure. Not only that, the obtained corresponding relationship is not robust, and as the model grows, the number of outlier noise points also increases, which requires an increase in the running time of RANSAC, thereby increasing the running time of the positioning task.
[0005] Many image regions, such as the sky, occlusions, and repeated indistinguishable image regions, cannot be used for positioning. In addition to increasing unnecessary computational work, extracting and matching features from these regions will generate many false matches, which in turn reduces the accuracy and efficiency of positioning. Summary of the Invention
[0006] Aiming at the problem that the image regions in existing visual relocalization methods cannot be used for localization and that many false matches are easily generated when extracting and matching features from image regions, the present invention provides a visual relocalization method based on saliency scene coordinate regression.
[0007] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:
[0008] A visual relocalization method based on saliency scene coordinate regression, comprising the steps of:
[0009] S1. Obtain the image to be queried;
[0010] S2. Input the image to be queried into the saliency prediction network to obtain the prediction result of 2D image key points;
[0011] S3. Input the image to be queried into the hierarchical scene coordinate regression network to obtain the prediction result of 3D map points;
[0012] S4. Match the prediction result of 2D image key points with the prediction result of 3D map points to obtain the 2D-3D matching relationship of the feature points with higher saliency values within the image scene;
[0013] S5. Solve the pose of the camera with six degrees of freedom based on the PnP-RANSAC pose algorithm.
[0014] Further, the detailed steps of predicting 2D image key points are as follows:
[0015] Extract the saliency dataset obtained from geometric semantic information to train the saliency model;
[0016] Construct a two-dimensional image and saliency descriptor, and select the most valuable feature points for localization from the image regions;
[0017] Adopt the non-maximum suppression method to select the key points with confidence scores higher than a specific threshold, and respectively select their corresponding three-dimensional coordinates;
[0018] Perform projection analysis on the result after non-maximum suppression of the saliency value and the original image.
[0019] Further, the detailed steps of constructing a two-dimensional image and saliency descriptor include:
[0020] Make a saliency dataset and extract the geometric feature G of the input image i ;
[0021] Use the semantic segmentation network RefineNet to obtain the semantic mask S of the image I i : S i = RefineNet(I i )
[0022] Filter geometric information using semantic masks: for do:
[0023] Gaussian blur:
[0024] Train the saliency model DINet on the saliency dataset;
[0025] Calculate the saliency map: The saliency map
[0026] Saliency map normalization:
[0027] Calculate the saliency descriptor: The saliency descriptor
[0028] Furthermore, the hierarchical scene coordinate regression network includes a RegionLabels classification layer, a SubregionLabels classification layer, and a SceneCoordinates basic regression layer;
[0029] RegionLabels performs a relatively coarse classification of the 3D map scene, dividing it into four categories and assigning a class label to each 3D point;
[0030] SubregionLabels performs a detailed classification of the 3D map scene, with about a dozen categories, and also assigns a label to each 3D point;
[0031] The SceneCoordinates basic regression layer is used to output the 3D coordinates of the scene.
[0032] Furthermore, the label information generated by the coarse-grained layer can be used in the prediction of the fine-grained layer. By receiving the input of the class label through an adjustment parameter generator, two parameters are then generated: namely, γ and β. Then, the adjustment layer applies these two parameters to the output given by the previous layer of convolution. The application method formula is:
[0033] f(x, l) = γ(l) ⊙ x + β(l)
[0034] where ⊙ represents the Hadamard product, that is, multiplying the corresponding positions of the matrices.
[0035] Furthermore, different loss functions are used for the classification label output and the regression coordinate output;
[0036] The loss function for the classification output is defined in the form of cross-entropy, and its expression is as follows:
[0037]
[0038] For the regression task, it is performed by minimizing the Euclidean distance between the predicted scene coordinates and the true scene coordinates y, and the formula is as follows:
[0039]
[0040] The camera projection equation is added to the scene coordinate network to form a geometric constraint, and the equation of its loss function is as follows:
[0041]
[0042] where x i is the 2D key point selected in the previous section, π represents the camera projection equation, k is the camera internal parameter, and the corresponding R and t are the true values of the camera pose and translation.
[0043] Furthermore, the scene coordinate regression network with geometric constraints, combined with the extraction of significant key points, plays a role in outlier rejection. The formula of the final loss function is as follows:
[0044]
[0045] where ω 1 , ω 2 , ω 3 , ω 4 are the weights of the loss function.
[0046] During training, we set the weights ω 1 , ω 2 of the scene classification loss to 1, while the weights ω 3 , ω 4 of the regression loss and geometric constraint are set to 10.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] This method uses image regions with high discriminability in the localization task to reduce outliers and perform localization accurately and effectively. Our method can learn to select good features, avoid occlusions and other unreliable regions, such as the features of the sky, streets, trees, shrubs, pedestrians, and cars, for the relocalization task. In addition, the algorithm of the present invention focuses more on the accuracy of the scene coordinate network itself. After adding geometric constraints and saliency screening, the matching point pairs obtain more accurate localization results with fewer numbers. It has been verified on indoor and outdoor datasets that our algorithm can obtain smaller localization errors while having less computational complexity compared to other algorithms. The main innovations and contributions are reflected in the following aspects:
[0049] The work of the present invention adds a detection process for significant feature points to the positioning task based on the scene coordinate regression network. The importance of feature points in the task is judged by the level of significance value and sent to the scene coordinate regression network for learning. Experiments show that our method can reduce the number of outliers in the positioning task and enable the network to learn from some valuable regions in the image.
[0050] For the scene coordinate regression network itself, the condition of geometric constraint is added. Experiments show that the algorithm of the present invention does not require a pose optimization module, thus saving computing resources. The number of 2D-3D matching points screened by the significance prediction network is smaller but of better quality. Brief Description of the Drawings
[0051] Figure 1 It is the overall flowchart of a visual relocalization method based on significant scene coordinate regression in an embodiment of the present invention;
[0052] Figure 2 It is the schematic diagram of the visualization result of 2D image key points in an embodiment of the present invention;
[0053] Figure 3 It is the schematic diagram of the scene coordinate regression network in an embodiment of the present invention;
[0054] Figure 4 It is the schematic diagram of the conditional parameter generation network in an embodiment of the present invention. Detailed Embodiment
[0055] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with embodiments and drawings. The content mentioned in the embodiments does not limit the present invention.
[0056] As Figure 1 shown, this embodiment provides a visual relocalization method based on significant scene coordinate regression, including the steps:
[0057] S1. Obtain the image to be queried;
[0058] S2. Input the image to be queried into the significance prediction network to obtain the prediction result of 2D image key points;
[0059] S3. Input the image to be queried into the hierarchical scene coordinate regression network to obtain the prediction result of 3D map points;
[0060] S4. Match the prediction result of 2D image key points with the prediction result of 3D map points to obtain the 2D-3D matching relationship of the feature points with higher significance value in the image scene;
[0061] S5. Solve the pose of the camera with six degrees of freedom based on the PnP-RANSAC pose algorithm.
[0062] Different from traditional structure - based or hierarchical localization strategies, which require real - time maintenance of a 3D map and related feature - point descriptors and also need to perform matching between descriptors, the algorithm proposed in the present invention can directly obtain such a relationship. A prediction network of salient points is trained through a saliency dataset and corresponding to the coordinate points in the 3D scene. This direct 2D - 3D matching relationship can effectively extract static - related information in the scene and effectively avoid the phenomenon of false matching existing in the traditional descriptor - matching process. In addition, the results of related ablation experiments of this method show that the proposed algorithm naturally acts as an outlier filter without the need for more training parameters like the differentiable RANSAC - based PnP pose - optimization module.
[0063] Detailed steps for 2D image key - point prediction:
[0064] Extract the geometric semantic information to obtain a saliency dataset for training a saliency model;
[0065] Construct a two - dimensional image and saliency descriptors, and select the most valuable feature points for localization from the image region;
[0066] Adopt non - maximum suppression method to select key points with confidence scores higher than a specific threshold and respectively select their corresponding three - dimensional coordinates;
[0067] Perform projection analysis on the result after non - maximum suppression of the saliency value and the original image.
[0068] The saliency dataset obtained by extracting geometric semantic information is used to train a saliency model. In addition, after obtaining the saliency map, since the points required in feature - point detection are sparse feature points, the non - maximum suppression method needs to be adopted to select key points with confidence scores higher than a specific threshold and respectively select their corresponding three - dimensional coordinates. In fact, screening key points according to the magnitude of the key - point saliency value is approximately an outlier - filtering process during the localization process. To verify our view, it will be verified in the subsequent experimental part. Note that only the salient points with saliency values greater than a certain threshold in the saliency map will enter the network for 3D scene coordinate regression to learn. Therefore, compared with the dense scene coordinate regression network, the number of matching pairs and the computational amount obtained by our algorithm are significantly reduced, but what is reduced is the quantity, and the quality of selection is significantly improved.
[0069] Detailed steps for constructing a two - dimensional image and saliency descriptors include:
[0070] Make a saliency dataset and extract the geometric feature G of the input image i ;
[0071] Use the semantic segmentation network RefineNet to obtain image I i Semantic mask: S i =RefineNet(I i );
[0072] Filtering geometric information using semantic masks: for do:
[0073] Gaussian Blur:
[0074] Train the saliency model DINet on the saliency dataset;
[0075] Computing Saliency Maps: Saliency Maps
[0076] Saliency map normalization:
[0077] Compute saliency descriptors: saliency descriptors
[0078] like Figure 2 As shown in the figure, it is the visualization result of the feature points selected by the saliency model. We splice the saliency map predicted by the saliency model with the original image and give different weights to the high and low saliency values, as shown in the left column of the figure. It can be seen that some fixed feature points suitable for positioning tasks such as buildings and street lights in the figure belong to the area of high saliency values, which are the 2D key points we selected. Next, we made a more detailed visualization, that is, the result after the saliency value is suppressed by non-maximum value and the original image are projected and analyzed. As shown in the right column of the figure, it can be seen that these points all cleverly fall on the edge focus area of the building, avoiding low-value areas such as the sky, lawn and pedestrians in the picture.
[0079] The hierarchical scene coordinate regression network includes the RegionLabels classification layer, the SubregionLabels classification layer, and the SceneCoordinates basic regression layer;
[0080] RegionLabels performs a rough classification of the 3D map scene into four categories and assigns a category label to each 3D point;
[0081] SubregionLabels classifies the 3D map scenes in detail, with about a dozen categories, and also assigns a label to each 3D point;
[0082] The SceneCoordinates base regression layer is used to output scene 3D coordinates.
[0083] The scene coordinate network is no longer classified and represents the coordinates of each 3D point. The corresponding deep learning network has two classification layers ( Figure 3 within the red bounding box in) and a basic regression layer ( Figure 3 within the green bounding box in), which corresponds to the above three-layer structure from coarse to fine. The classification branches of the upper two layers take the class labels as the final outputs, and the basic regression layer takes the scene 3D coordinates as the final outputs. This network design from coarse to fine, with the support of perfect data, can enable the network to have a smaller receptive field in the finer output layer, that is, the network can obtain more global information while also being able to distinguish local information in the deeper feature maps, preventing overfitting. To enable the label information generated by the coarse-grained layer to be used in the prediction of the fine-grained layer, HSCNet specifically designs an adjustment layer (as shown on the Figure 3 right side).
[0084] The label information generated by the coarse-grained layer can be used in the prediction of the fine-grained layer. By receiving the input of the class label through an adjustment parameter generator, two parameters are then generated: namely, γ and β. Then, the adjustment layer applies these two parameters to the output given by the previous layer of convolution. The formula for the application method is:
[0085] f(x, l) = γ(l) ⊙ x + β(l)
[0086] where ⊙ represents the Hadamard product, that is, multiplying the corresponding positions of the matrices.
[0087] The classification label output and the regression coordinate output use different loss functions;
[0088] The loss function of the classification output is defined in the form of cross-entropy, and its expression is as follows:
[0089]
[0090] For the regression task, it is performed by minimizing the Euclidean distance between the predicted scene coordinates and the real scene coordinates y. The formula is as follows:
[0091]
[0092] A differentiable RANSAC model is used to add a pose optimization process. This process will obtain a pose hypothesis pool based on the 2D-3D matching relationship established by the previous scene coordinates, and then select the best pose that conforms to the model. However, our model can actually use the prior pose and 3D model. It is very reasonable and has development space to add the camera projection equation to the scene coordinate network to form geometric constraints.
[0093] The camera projection equation is added to the scene coordinate network to form geometric constraints, and the equation of its loss function is as follows:
[0094]
[0095] Among them, x i is the 2D key point selected in the previous section, π represents the camera projection equation, k is the camera internal parameter, and the corresponding R and t are the true values of the camera pose and translation.
[0096] As Figure 4 shown, the scene coordinate regression network with added geometric constraints, combined with the extraction of significant key points, plays a role in outlier rejection. The formal formula of the final loss function is as follows:
[0097]
[0098] Among them, ω 1 , ω 2 , ω 3 , ω 4 are the weights of the loss function.
[0099] Since the regression task still dominates the whole task, followed by the geometric constraints achieved by using the reprojection error, and finally the loss function of the classification term, the corresponding weight sizes are also sorted according to the importance degree. During training, we set the weights ω 1 , ω 2 of the scene classification loss to 1, while the weights ω 3 , ω 4 of the regression loss and geometric constraints are set to 10.
[0100] The 2D feature points in the query image and the matching 3D space points are obtained. This completes the key issue of 2D-3D matching in the positioning algorithm, and our algorithm does not require descriptor matching nor does it need to invest in a massive matching relationship to establish a pose hypothesis pool. The subsequent task is to directly use the PnP algorithm in the RANSAC scheme to estimate the six-degree-of-freedom pose. We directly use the algorithm interface in OpenCV to implement the whole process.
[0101] Compared with the prior art, the present invention has the following beneficial effects:
[0102] This method uses image regions with high discriminability in the positioning task to reduce outliers and perform positioning accurately and effectively. Our method can learn to select good features, avoid occlusions and other unreliable regions, such as feature points of the sky, streets, trees, shrubs, pedestrians, and cars, for the relocalization task. In addition, the algorithm of the present invention focuses more on the accuracy of the scene coordinate network itself. After adding geometric constraints and saliency screening, the matching point pairs with fewer numbers yield more accurate positioning results. Verified on indoor and outdoor datasets, our algorithm can obtain smaller positioning errors while having less computational complexity compared to other algorithms. The main innovations and contributions are as follows:
[0103] The work of the present invention adds a saliency feature point detection process to the positioning task based on the scene coordinate regression network. The importance of feature points in the task is judged by the level of saliency values and sent to the scene coordinate regression network for learning. Experiments show that our method can reduce the number of outliers in the positioning task and enable the network to learn from some valuable regions in the image.
[0104] For the scene coordinate regression network itself, geometric constraints are added. Experiments show that the algorithm of the present invention does not require a pose optimization module, thus saving computational resources. The number of 2D-3D matching points screened by the saliency prediction network is smaller but of better quality.
[0105] The algorithm proposed in the present invention is tested on the Cambridge landmarks dataset and the 7Scenes dataset to verify the effectiveness of the algorithm. The datasets used contain scenes such as different landmark buildings in the city, and also include variations in weather, light, etc. In addition, the indoor datasets also include different application scenarios such as offices, kitchens, and living rooms, which can effectively test the algorithm. The operating system is Ubuntu16.04, the graphics card is NVIDIA GeForce RTX2070, and the CPU is AMD Ryzen 5-2600(8GB).
[0106] 1. Positioning results of the Cambridge landmarks dataset:
[0107] The proposed method is compared with other advanced algorithms. The Active Search (AS) algorithm is a traditional geometric-based method with relatively high accuracy and the most representative at present. PoseNet and VLocNet are methods based on deep learning pose regressors with relatively high accuracy at present. These two methods directly perform pose regression through prior poses. The main reason for choosing these two methods is to verify our view that the scheme for learning scene coordinates is better than the method of directly regressing the pose estimator. DSAC++ and DSAC* are algorithms based on deep learning scene coordinate regression. They directly obtain the 2D-3D matching relationship and use it for pose solving. However, such algorithms do not have a saliency mechanism. They believe that all coordinate points obtained through the network are valuable for the positioning task and bring all the matching point pairs obtained in the scene coordinate regression network into a pose optimization module based on differentiable RANSAC for calculation. HSCNet is the selected baseline algorithm, which is a scheme that focuses more on the scene coordinate regression network itself but does not add saliency information. The positioning results of these algorithms on this dataset are shown in Table 1. The median positioning errors (including translation error and rotation error) of different algorithms in different scenarios on the Cambridge landmark dataset are shown in the table. Generally speaking, the proposed algorithm ranks among the leading levels in all scenarios.
[0108] Table 2 Positioning Results on Cambridge landmark Dataset
[0109]
[0110] As can be seen from the table, the median error of the proposed algorithm is at a relatively low level in the four scenarios of this dataset compared with other algorithms. In this section, HSCNet is used as the baseline. Compared with the baseline, after adding the prediction of salient 2D key points, the median error of the algorithm is further reduced, which first verifies the effectiveness of the algorithm proposed in this chapter. From the perspective of method categories, AS is a traditional structure-based scheme. Its main contribution lies in proposing an efficient algorithm for establishing 2D-3D matching relationships. Therefore, its overall performance is very good. From the data, the gaps between different scenarios of this type of algorithm in the dataset are not large, and the overall performance is relatively good. However, due to the need to establish matching relationships from a large 3D model, its computational time and memory are not the best choices. For example, the running time of AS for one frame is about 400 ms, while the running time of the algorithm proposed in the present invention is about 20 ms.
[0111] The present invention focuses on the scene coordinate network itself and the selection of 2D key points, indicating that improvements to this single component can already enhance the positioning performance and exceed the level of some previous advanced algorithms in the field. Compared with traditional algorithms such as AS, the algorithm of the present invention does not require online maintenance of the 3D model of the scene, as well as the associated key points and their descriptors. Therefore, the work of balancing search speed and accuracy in traditional methods does not occur in our algorithm. The algorithm proposed by the present invention directly predicts 2D salient key points through two neural networks and sends them into the scene coordinate network to calculate the matching 3D coordinates, without processes such as descriptor matching calculations.
[0112] 2. Localization results of the 7Scenes dataset:
[0113] According to the logic in the first point, the present invention conducts a localization test on the indoor dataset, the 7Scenes dataset. The present invention still selects different types of methods as control objects for the experiment. First of all, overall, due to the relatively small scene scale of the indoor dataset, its translational error term is one order of magnitude lower than that of the outdoor dataset, while the rotational error has little difference.
[0114] The present invention still conducts a comparative analysis according to different types of methods. First is the traditional method AS. It can be seen that the algorithm based on the traditional method has actually achieved a very small median error effect, and the gap with other algorithms is not large. Analyzing the reasons, mainly because compared with the outdoor scene, the scale of the key points in the indoor dataset has become smaller, offsetting the disadvantages of the traditional method. Therefore, some researchers will use the results obtained by the traditional method as a reference for their own algorithms. Achieving the accuracy and stability of the traditional method means that the effect of their own algorithm is worthy of recognition. From the perspective of learning-based methods, it is the same as our previous conclusion: Methods such as PoseNet that learn the entire localization algorithm process have higher median localization errors in different scenarios than those that learn part of it, such as the scene coordinate regression network in this chapter. From the perspective of stability, its average localization accuracy is also much lower. Although the localization accuracy of the method based on partial learning varies greatly in different scenarios, overall it is still better than methods such as PoseNet.
[0115] Table 2 Localization results in the 7Scenes dataset
[0116]
[0117] On this indoor dataset, the minimum median positioning error can be basically obtained in each scenario, and its average positioning accuracy is also significantly better than other algorithms. When the algorithm of the present invention is tested on the indoor dataset, although there are differences in the extraction of significant features because we have always emphasized the extraction of significant regions for driverless driving, some of the point, line, and plane features are common. The algorithm in this chapter still obtains a lower median error and a higher average positioning accuracy with fewer 2D-3D matching relationships. For the baseline algorithm selected in this chapter, our work adds the screening of significant regions and obtains a lower median error in the seven scenarios of this indoor dataset, but the improvement amplitude varies in each scenario. For example, the improvement amplitude on Heads is significantly smaller than that in other scenarios. Our algorithm allows for slight differences between different scenarios. After all, there will be errors in the significant model itself during the training and prediction phases. One possible future research direction for us is to improve the extraction process of the significant model and use data-driven significant prediction instead of using a self-made significant dataset to train the model.
[0118] 3. Analysis of noise point removal
[0119] NG-RANSAC proposes an extension to the classic RANSAC algorithm by training a network to guide the hypothesis sampling of RANSAC. Through this network, outlier points are assigned low weights while inlier points are assigned high weights. This type of method realizes the algorithm architecture for relocalization by combining the architecture of DSAC* for scene regression and NG-RANSAC for pose hypothesis sampling. The difference between our algorithm and this type of algorithm is that we focus more on scene coordinate regression itself. At the same time, we use the significant model to screen the points in the positioning task. This process is actually the same as the purpose of NG-RANSAC. Table 4 shows the median positioning errors of our algorithm and NG-RANSAC on the Cambridge landmark data. We hope to first compare with such algorithms to prove that the outlier removal strategy of the present invention can reach an advanced level. By using reliable significant key points, we have obtained quite good results. Specifically, we achieved the best positioning in the Old Hospital scenario. In this scenario, the work in this chapter successfully avoided selecting correspondences that fall in occluded areas, cars, trees, bushes, repetitive flat areas, and reflective windows. By learning to avoid these areas in these scenarios, a large number of outlier sources can be avoided, thereby improving the positioning accuracy.
[0120] Table 3 Comparison of positioning errors of the NG-RANSAC algorithm on the Cambridge landmark dataset
[0121]
[0122] The method proposed by the present invention performs positioning by selecting a set of highly discriminative 2D-3D matching points. The selected salient key points are located in the highly recognizable parts of the image. This step helps to select a minimum number of corresponding points while ensuring a low level of outliers. To show the importance of the selected key points, we compare the pose positioning errors by selecting two sets of corresponding points, each with 200 pairs of matching points. The first set is selected from the points with a predicted saliency value higher than 0.7, while the other set is composed of the points with a saliency value of 0.4 after non-maximum suppression. The present invention directly runs the PnP-RANSAC implementation of OpenCV for 100 iterations with a reprojection error of 3px. The results are shown in Table 4.
[0123] Table 4 Positioning errors of different saliency thresholds on the Cambridge landmark dataset
[0124]
[0125] It can be seen that the key points selected in Setting 1 of the method of the present invention basically belong to reliable regions. In most cases, trees, sidewalks with similar appearances, streets, reflective windows, and the sky are avoided, and this process was also visually analyzed in the 2D key point prediction in the previous section. This helps the network learn the scenes corresponding to the reliable regions and reliably predict the corresponding three-dimensional coordinates. Therefore, the corresponding positioning error of the algorithm proposed in this chapter will also be reduced. On the contrary, if more salient points fall into the above regions, it will make it difficult for the network to reliably predict the three-dimensional scene coordinates. This in turn reduces the positioning accuracy. Compared with DSAC++ and DSAC* that select thousands of key points, these results further illustrate that the key factor in the positioning process is not the number of matching points, but the quality of the matching points.
[0126] The above provides a detailed introduction to a visual relocalization method based on salient scene coordinate regression provided by the present application. The description of the specific embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A visual relocalization method based on salient scene coordinate regression, characterized in that: Includes steps: S1, obtain the image to be queried; S2, input the query image into the saliency prediction network to obtain the prediction result of its 2D image key points; S3, input the query image into the hierarchical scene coordinate regression network to obtain its 3D map point prediction result; S4, matching the 2D image key point prediction results with the 3D map point prediction results to obtain a 2D-3D matching relationship of feature points with higher significance values in the image scene; S5. Solve the six-degree-of-freedom pose of the camera based on the PnP-RANSAC pose algorithm.
2. The visual relocalization method based on salient scene coordinate regression according to claim 1, characterized in that: Detailed steps for 2D image key point prediction: Extract the saliency dataset obtained by geometric semantic information to train the saliency model; Construct a two-dimensional image and saliency descriptor, and select the most valuable feature points for positioning from the image area; Using the non-maximum suppression method, we select key points with confidence scores above a certain threshold and select their corresponding three-dimensional coordinates respectively; The saliency value is projected and analyzed with the result after non-maximum suppression and the original image.
3. The visual relocalization method based on salient scene coordinate regression according to claim 2, characterized in that: The detailed steps of constructing a two-dimensional image and saliency descriptor include: Create a saliency dataset to extract geometric features of the input image; Use the semantic segmentation network RefineNet to obtain the semantic mask of the image; Filtering geometric information using semantic masks; Gaussian blur processing; Train a saliency model on a saliency dataset; Calculate the saliency map and normalize the saliency map; Compute saliency descriptors.
4. The visual relocalization method based on salient scene coordinate regression according to claim 3, characterized in that: The hierarchical scene coordinate regression network includes the Region Labels classification layer, the Subregion Labels classification layer, and the Scene Coordinates basic regression layer; RegionLabels performs a rough classification of the 3D map scene into four categories and assigns a category label to each 3D point; SubregionLabels classifies the 3D map scene in detail and also assigns a label to each 3D point; The SceneCoordinates base regression layer is used to output scene 3D coordinates.
5. The visual relocalization method based on salient scene coordinate regression according to claim 4, characterized in that: The label information generated by the coarse-grained layer can be used in the prediction of the fine-grained layer. An adjustment parameter generator receives the input of the category label and then generates two parameters: γ and β. These two parameters are then applied to the output given by the previous convolution layer. The formula for its action is: f(x, l) = γ(l)⊙x+β(l) Where ⊙ represents the Hadamard product, that is, the multiplication of corresponding positions of the matrices.
6. The visual relocalization method based on salient scene coordinate regression according to claim 5, characterized in that: Different loss functions are used for classification label output and regression coordinate output; The loss function for classification output is defined in the form of cross entropy, which is expressed as follows: represents the jth element in the true label vector of sample i. Represents the predicted probability distribution of sample i The j-th element in the vector. For regression tasks, by minimizing the predicted scene coordinates It is performed by the Euclidean distance between the real scene coordinate y and the real scene coordinate y, and the formula is as follows: The camera projection equation is added to the scene coordinate network to form a geometric constraint. The equation of the loss function is as follows: Among them, x i is the 2D key point selected in the previous section, π represents the projection equation of the camera, k is the camera intrinsic parameter, and the corresponding R and t are the true values of the camera's position and translation.
7. The visual relocalization method based on salient scene coordinate regression according to claim 6, characterized in that: The scene coordinate regression network with geometric constraints is combined with the extraction of significant key points to remove outliers. The final loss function is formally formulated as follows: Among them, ω1, ω2, ω3, ω4 are the weights of the loss function.
Citation Information
Cited By
Visual relocation and pose estimation method based on space-time dual compression and related device
CN121837372A