Indoor repositioning method based on virtual view synthesis
By generating virtual views and utilizing the RenderNet network to optimize features, the problem of performance degradation in image retrieval and feature matching in visual relocalization is solved, achieving a more efficient indoor relocalization effect.
Patent Information
- Application Number
- CN202510877662.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-18
AI Technical Summary
Existing visual relocation methods suffer from reduced image retrieval and feature matching performance in large-scale indoor environments due to changes in viewpoint and appearance, and the rendering process of high-quality 3D models is time-consuming and labor-intensive.
By generating virtual views, global and local features are generated from existing database images using the RenderNet network, optimizing the image retrieval and feature matching process, including view enhancement and pose optimization stages, and directly learning to generate features for image retrieval and matching.
It improves the effectiveness of image retrieval and feature matching, reduces the impact of artifacts, and enhances the overall performance of visual relocalization, especially under conditions of large changes in viewing angle.
Smart Images

Figure CN120976309A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric communication, more particularly, to an indoor repositioning method based on virtual view synthesis. BACKGROUND
[0002] To meet the growing demand of various spatial intelligence related applications, such as navigation, augmented reality (AR) and SLAM, visual repositioning is an indispensable technology
[39] . Essentially, it is a registration problem to localize a query image in a visual database. By establishing 2D-3D correspondences between the query image and database images, the 6DoF camera pose corresponding to the query image can be estimated.
[0003] Specifically, the goal of visual repositioning is to predict the 6DoF camera pose in a pre-built 3D map from an RGB image. Current localization pipelines can be divided into three categories:
[0004] 1) End-to-end pose regression methods
[14]
[13] [5]
[12] . Directly infer the absolute camera pose from a single RGB image. It is later found that these methods are closely related to image retrieval.
[0005] 2) Coordinate regression methods
[28] [3][4] . Regress the dense 3D scene coordinates of the query image and obtain the final camera pose through dense 2D-3D correspondences. Most of these methods are specific to a particular scene and need to be trained for new scenes.
[0006] 3) Image matching based methods. These methods first retrieve similar images from the database through image retrieval
[10] [2] , then establish sparse
[24] or dense
[23] correspondences between the query image and the retrieved database images, forming 2D-3D correspondences. Finally, the pose is solved through PnP algorithm.
[0007] Most of the current state-of-the-art methods for large-scale environment localization are based on image matching. Many studies focus on local feature learning [7][8]
[16]
[15]
[17] , sparse feature matching
[35]
[36]
[25] and dense feature matching
[23]
[22] to improve the performance of image matching based localization methods. However, if there is a small overlap between the database image and the query image, the image matching based method will still suffer from severe performance degradation, which is common in large-scale environments.
[0008] The current state-of-the-art scene-agnostic localization method
[26]
[24] is composed of an image retrieval module and an image matching module, where the image retrieval module searches the candidate database images as localization references, and the image matching module establishes the 2D-3D correspondence. Then the camera pose is calculated by the subsequent RANSAC+PnP algorithm. However, large view angle and appearance changes have a great impact on the performance of these two modules, and this problem is very common in large-scale indoor scenes, because the view angle and appearance will change rapidly when the camera moves.
[0009] To solve this problem, an intuitive idea is to synthesize new views to enrich the query database, because the database images containing a larger range of overlaps will make image retrieval and image matching easier, thus improving the overall performance of relocalization. View synthesis aims to generate photorealistic images of arbitrary new views from 3D models, sparse views
[18] or point clouds
[29]
[20] View synthesis has been proven to be useful for many pose estimation related fields. As Torii et al.
[33] show that the image retrieval performance can be enhanced by enriching the database with new views; Sibbing et al.
[29] render views from point clouds for visual localization; while Zhang et al.
[37] improve pose estimation in the localization problem by view synthesis.
[0010] However, image synthesis requires high-quality 3D models, which are usually captured by high-cost 3D scanners, and this process is time-consuming and labor-intensive. For example, Inloc
[34] relocalization dataset is captured by using an expensive laser radar scanner [1] , which takes more than 30 minutes for a single scan, so this process is very time-consuming.
[0011] References:
[0012] [1] Faro 3d scanner. https: / / www.faro.com / .
[0013] [2] Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. Netvlad: Cnn architecture for weakly supervised place recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5297-5307, 2016.
[0014] [3] Eric Brachmann, Alexander Krull, Sebastian Nowozin, Jamie Shotton, Frank Michel, Stefan Gumhold, and Carsten Rother. Dsac - differentiable ransac for camera localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6684-6692, 2017.
[0015] [4] Eric Brachmann and Carsten Rother. Expert sample consensus applied to camera re-localization. In Proceedings of the IEEE International Conference on Computer Vision, pages 7525-7534, 2019.
[0016] [5] Samarth Brahmbhatt, Jinwei Gu, Kihwan Kim, James Hays, and Jan Kautz. Geometry-aware learning of maps for camera localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2616-2625, 2018.
[0017] [6] Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nister. ScanNet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5828-5839, 2017.
[0018] [7] Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superpoint: Self-supervised interest point detection and description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 224-236, 2018.
[0019] [8] Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, and Torsten Sattler. D2-net: A trainable cnn for joint detection and description of local features. arXiv preprint arXiv:1905.03561, 2019.
[0020] [9] Huanhuan Fan, Yuhao Zhou, Ang Li, Shuang Gao, Jijunnan Li, and Yandong Guo. Visual localization using semantic segmentation and depth prediction. arXiv preprint arXiv:2005.11922, 2020.
[0021]
[10] Albert Gordo, Jon Almazán, Jerome Revaud, and Diane Larlus. Deep image retrieval: Learning global representations for image search. In European conference on computer vision, pages 241-257. Springer, 2016.
[0022]
[11] Martin Humenberger, Yohann Cabon, Nicolas Guerin, Julien Morat, Revaud, Philippe Rerole, Noé Pion, Cesar de Souza, Vincent Leroy, and Gabriela Csurka. Robust image retrieval-based visual localization using kapture. arXiv preprint arXiv:2007.13867, 2020.
[0023]
[12] Alex Kendall and Roberto Cipolla. Modelling uncertainty in deep learning for camera relocalization. In 2016 IEEE international conference on Robotics and Automation (ICRA), pages 4762-4769. IEEE, 2016.
[0024]
[13] Alex Kendall and Roberto Cipolla. Geometric loss functions for camera pose regression with deep learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5974-5983, 2017.
[0025]
[14] Alex Kendall, Matthew Grimes, and Roberto Cipolla. PoseNet: A convolutional network for real-time 6-dof camera relocalization. In Proceedings of the IEEE international conference on computer vision, pages 2938-2946, 2015.
[0026]
[15] Zixin Luo, Tianwei Shen, Lei Zhou, Jiahui Zhang, Yao Yao, Shiwei Li, Tian Fang, and Long Quan. Contextdesc: Local descriptor augmentation with cross-modality context. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 2527-2536, 2019.
[0027]
[16] Zixin Luo, Tianwei Shen, Lei Zhou, Siyu Zhu, Runze Zhang, Yao Yao, Tian Fang, and Long Quan. Geodesc: Learning local descriptors by integrating geometry constraints. In Proceedings of the European conference on computer vision (ECCV), pages 168-183, 2018.
[0028]
[17] Zixin Luo, Lei Zhou, Xuyang Bai, Hongkai Chen, Jiahui Zhang, Yao Yao, Shiwei Li, Tian Fang, and Long Quan. Aslfeat: Learning local features of accurate shape and localization. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 6589-6598, 2020.
[0029]
[18] Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. arXiv preprint arXiv:2008.02268, 2020.
[0030]
[19] Norio Nakata, Nobuhiro Takeda, and Norihiro Tokitoh. Synthesis and properties of the first stable germabenzene. Journal of the American Chemical Society, 124(24):6914-6920, 2002.
[0031]
[20] Francesco Pittaluga, Sanjeev J Koppal, Sing Bing Kang, and Sudipta N Sinha. Revealing scenes by inverting structure from motion reconstructions. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 145-154, 2019.
[0032]
[21] Jerome Revaud, Jon Almazán, Rafael S Rezende, and Cesar Roberto de Souza. Learning with average precision: Training image retrieval with a listwise loss. In Proceedings of the IEEE International Conference on Computer Vision, pages 5107-5116, 2019.
[0033]
[22] Ignacio Rocco, Relja Arandjelovic, and Josef Sivic. Efficient neighbourhood consensus networks via submanifold sparse convolutions. In European Conference on Computer Vision, pages 605-621. Springer, 2020.
[0034]
[23] Ignacio Rocco, Mircea Cimpoi, Relja Arandjelovic, Akihiko Torii, Tomas Pajdla, and Josef Sivic. Neighbourhood consensus networks. arXiv preprint arXiv:1810.10510, 2018.
[0035]
[24] Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. From coarse to fine: Robust hierarchical localization at large scale. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 12716-12725, 2019.
[0036]
[25] Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 4938-4947, 2020.
[0037]
[26] Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, et al. Benchmarking 6dof outdoor visual localization in changing conditions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8601-8610, 2018.
[0038]
[27] Torsten Sattler, Qunjie Zhou, Marc Pollefeys, and Laura Leal-Taixe. Understanding the limitations of cnn-based absolute camera pose regression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3302-3312, 2019.
[0039]
[28] Greg Schohn and David Cohn. Less is more: Active learning with support vector machines. In ICML, volume 2, page 6. Citeseer, 2000.
[0040]
[29] Dominik Sibbing, Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Sift realistic rendering. In 2013 International Conference on 3D Vision - 3DV 2013, pages 56-63. IEEE, 2013.
[0041]
[30] Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Loftr: Detector-free local feature matching with transformers. arXiv preprint arXiv:2104.00680, 2021.
[0042]
[31] Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomas Pajdla, and Akihiko Torii. Inloc: Indoor visual localization with dense matching and view synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7199-7209, 2018.
[0043]
[32] Hajime Taira, Ignacio Rocco, Jiri Sedlar, Masatoshi Okutomi, Josef Sivic, Tomas Pajdla, Torsten Sattler, and Akihiko Torii. Is this the right place? geometric-semantic pose verification for indoor visual localization. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 4373-4383, 2019.
[0044]
[33] Akihiko Torii, Relja Arandjelovic, Josef Sivic, Masatoshi Okutomi, and Tomas Pajdla. 24 / 7 place recognition by view synthesis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1808-1817, 2015.
[0045]
[34] Erik Wijmans and Yasutaka Furukawa. Exploiting 2d floorplan for building-scale panorama rgbd alignment. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 308-316, 2017.
[0046]
[35] Kwang Moo Yi, Eduard Trulls, Yuki Ono, Vincent Lepetit, Mathieu Salzmann, and Pascal Fua. Learning to find good correspondences. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2666-2674, 2018.
[0047]
[36] Jiahui Zhang, Dawei Sun, Zixin Luo, Anbang Yao, Lei Zhou, Tianwei Shen, Yurong Chen, Long Quan, and Hongen Liao. Learning two view correspondences and geometry using order-aware network. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 5845-5854, 2019.
[0048]
[37] Zichao Zhang, Torsten Sattler, and Davide Scaramuzza. Reference pose generation for visual localization via learned features and view synthesis. arXiv preprint arXiv:2005.05179, 2020.
[0049]
[38] Qunjie Zhou, Torsten Sattler, and Laura Leal-Taixe. Patch2pix: Epipolar guided pixel-level correspondences. arXiv preprint arXiv:2012.01909, 2020.
[0050]
[39] Cai Xudong, Wang Yongcai, Bai Xuewei, Li Deying. A review of visual relocalization methods based on prior maps. Journal of Software, 2024, 35(2):975–1009. http: / / www.jos.org.cn / 1000-9825 / 6946.htm Summary of the Invention
[0051] The technical problem to be solved by the present invention is to provide an indoor relocation method based on virtual view synthesis to address the shortcomings of the existing technology, which greatly improves the image retrieval and feature matching effect in the visual relocation process.
[0052] The present invention discloses an indoor relocation method based on virtual view synthesis, the method comprising:
[0053] Obtain the query image and the actual database image, and generate a virtual view;
[0054] The query image and real database images are processed by a retrieval network and a feature extraction network to generate global features G. q G r and local features L q L r These are used for image retrieval and image matching, respectively; the virtual view is processed using RenderNet to generate global features G. v and local features L v ;
[0055] Calculate the global feature G q With global features G v The distance between them is used to find the k most similar real database images to the query image, and a 2D-3D correspondence is established between the query image and the found real database images. The initial camera pose is estimated based on the 2D-3D correspondence.
[0056] The camera pose is optimized by using the virtual view under the initial camera pose as a relocalization reference.
[0057] Preferably, the RenderNet includes:
[0058] Global features G used for image retrieval v RenderNet G Branches, and
[0059] A set of local features L used to generate key points in the virtual view v RenderNet for sparse matching L Branch.
[0060] Preferably, the RenderNetG Branch adopts widely used generalized mean pooling layer to extract global feature G v , the global feature G v is an m-dimensional global feature vector.
[0061] Preferably, the distance between the global feature G q and the global feature G v is calculated by the following formula:
[0062] argmin φ ||RenderNet G (I p , F p )-RetrievalNet(I) || 2 ;
[0063] In the formula, Φ is the parameter of RenderNet G , I p and F p are the projected color and projected feature of the virtual view respectively, and I is the real image collected under the virtual view perspective.
[0064] Preferably, the key points of the database image near the virtual viewpoint are projected to the virtual view according to the database information, and then the enhanced local descriptor is generated by using RenderNet L , the enhanced local descriptor being an n-dimensional feature map, and the local feature of the key point is interpolated from the feature map.
[0065] Preferably, the distance between the local feature L r and the local feature L v is calculated by the following formula:
[0066] argmin ψ ||RenderNet L (I p , F p )-DescNet(I)·M || 2 ;
[0067] In the formula, Ψ is the parameter of RenderNet L ; M is a mask, I p and F p are the projected color and projected feature of the virtual view respectively, and I is the real image collected under the virtual view perspective.
[0068] Preferably, the optimization method of the camera pose is as follows:
[0069] A new virtual view is rendered according to the initial camera pose, and then the enhanced local descriptor is generated by using RenderNet LThe local features are generated again for sparse matching, and camera pose estimation is performed at this time.
[0070] Advantages
[0071] The present application has the advantages of:
[0072] 1. Unlike the method of rendering new views from pre-constructed high-quality 3D models, the present application proposes RenderNet to generate virtual new views from existing database images. Unlike the usual view synthesis process aiming at rendering visually realistic views
[18]
[19] , the neural network proposed by the present application directly learns to obtain global features and local features of virtual viewpoints for image retrieval modules and image matching modules. The features required to be generated are selected, instead of generating images, in order to reduce the influence of inevitable artifacts in the synthesized images.
[0073] 2. With the generated virtual views, the traditional localization process is optimized by two operations, namely, view enhancement operation and pose optimization operation. In the view enhancement stage, a set of possible viewing angles in the scene are first manually pre-sampled as virtual views, so that the query image can retrieve those views using the global features rendered by it, and establish the corresponding relationship through local features. These generated virtual views can greatly improve the process of image retrieval and corresponding relationship establishment. In the pose optimization stage, a new virtual view is generated using the rough camera pose calculated in the previous stage. Then 2D-3D correspondence can be matched and the accurate pose can be obtained. BRIEF DESCRIPTION OF DRAWINGS
[0074] Figure 1 The overall architecture diagram of the visual relocalization method based on prior map;
[0075] Figure 2 The overall flowchart of the visual relocalization method of the present application;
[0076] Figure 3 The training flowchart of the RenderNet of the present application;
[0077] Figure 4 Comparison of matching results of the HLoc method and the RenderNet proposed by the present application;
[0078] Figure 5 The virtual 3D earth model displayed in the real environment by the user through AR glasses or mobile phone;
[0079] Figure 6 The example diagram of SLAM map construction and localization;
[0080] Figure 7 The schematic diagram of omnidirectional stereo surround viewing and operating the running state of the satellite in the real scene. Detailed Implementation
[0081] The present invention will be further described below with reference to embodiments, but this does not constitute any limitation on the present invention. Any limited modifications made by any person within the scope of the claims of the present invention are still within the scope of the claims of the present invention.
[0082] This invention provides an indoor relocalization method based on virtual view synthesis. Following traditional methods, it first retrieves similar images (i.e., real database images) from a database by measuring the distance of global features of the images. Then, it estimates the camera pose through 2D-3D correspondences based on local feature matching. Furthermore, it generates a virtual view to improve the image retrieval and matching processes. Utilizing the features of the generated virtual view, this invention designs view enhancement and pose optimization stages to improve localization performance, such as… Figure 2 As shown in the figure, the content illustrates the overall process of the proposed visual relocalization method. The top left corner shows the generation of a virtual view by projecting colors and local features from nearby real database images; the top right corner shows the use of the virtual view together with the real database images for view augmentation. Specifically, for both the query image and the real database image, RetrievalNet and DescNet are used to generate global features G. q G r and local features L q L r These are used for image retrieval and image matching, respectively. RenderNet is used for virtual views. G and RenderNet L Generate the corresponding global feature G v and local features L v The bottom section illustrates the pose optimization process to improve relocalization performance. First, RANSAC+PnP is used to estimate the camera pose of the query image. Then, the pose is further optimized by using a virtual view under the initial camera pose as a relocalization reference.
[0083] The indoor relocation method of this invention will be described in detail below. First, a relocation method using virtual viewpoints for view enhancement and pose optimization will be introduced. Then, the implementation details of RenderNet will be introduced, which generates global and local features of the virtual view, which can greatly improve the image retrieval and feature matching effects in the visual relocation process.
[0084] Regarding view augmentation. In this stage, rendered virtual views will be generated at a set of poses selected by a heuristic algorithm to enrich the real database. These virtual views have more overlap with the query image than the existing real database images. Specifically, we compute the overlap area between each pair of existing viewpoints, and then compute a set of new viewpoints according to the overlap area. Then we generate their global and local features at these new viewpoints. As shown in Figure 2 , for query and real database images, we generate global feature vectors using the retrieval network RetrievalNet (we use AP-GeM
[21] in this paper). In addition, we detect and describe keypoints using the feature extraction network DesNet (we use SuperPoint [7] in this paper). For virtual views, we generate their global feature vectors using RenderNet G , and generate their local features using RenderNet L . In the retrieval stage, the query image finds the top-k database images that are most similar to it by measuring the distance of global features. Next, we establish 2D-3D correspondences between the query image and the retrieved images by the sparse feature matching method. Finally, we estimate the rough camera pose from the established correspondences by RANSAC+PnP algorithm.
[0085] Regarding pose optimization. To optimize the estimated camera pose, we render a new virtual view at the rough camera pose estimated by the above procedure, and then perform sparse matching and PnP pose estimation again. This time, we only generate local features using RenderNet L . As will be shown later, pose optimization can greatly improve the accuracy of the pose, especially in the strict threshold region.
[0086] Regarding the virtual viewpoint generation network RenderNet. Specifically, we will introduce how to render any new virtual view using RenderNet in the following. Instead of rendering real database images, we propose to directly generate the features needed for visual relocalization. This approach not only makes the relocalization procedure more compact, but also avoids artifacts when generating images. Specifically, RenderNet contains two branch networks: 1) the RenderNet G branch, which is used to generate global features for image retrieval; 2) the RenderNet L branch, which is used to generate a set of local features of keypoints in this view for sparse matching. Instead of re-detecting keypoints in the virtual view, the local feature generation branch network proposed by the present invention reuses the projected keypoints in the existing view and only generates enhanced descriptors that can better describe this view.
[0087] Input of RenderNet. The input of RenderNet contains dense projected color and local feature information from real database images, which are tensors of HxWx3 and H / 8xW / 8xn, respectively, where n is the channel dimension. If there is no overlap between viewpoints, it is filled with zeros. Since the local feature descriptor is a representation method of image patches, the encoded image patch information can help solve the artifacts in the projected color image. At the same time, the coarse 3D model of the scene database can be used to provide a rough depth during projection to remove occluded points. The pre-trained DescNet is used as the local feature network in the present invention, and the local feature used here is SuperPoint.
[0088] Output of RenderNet. The output of RenderNet G Branch network aims to extract global feature representation for visual relocalization from virtual views. The present invention uses a widely used generalized mean pooling (GeM) layer to extract global features. It outputs an m-dimensional global feature vector for image retrieval. In the retrieval stage, the distance is calculated by measuring the similarity between the real feature vector and the virtual feature vector. In order to train this network while ensuring that the distance comparison makes sense, we train RenderNet G to get its output features, and at the same time get the features of the real database image under this view generated by RetrievalNet, minimize the distance between the two features. Just like the following formula:
[0089] argmin φ ||RenderNet G (I p , F p )-RetrievalNet(I)| 2 (1)
[0090] In the formula, Φ is the parameter of RenderNet G , I p and F p are projected color and projected feature, respectively, and I is the real image collected at the current viewpoint.
[0091] For the local feature of this new viewpoint, the present invention proposes an enhanced local descriptor based on projected key points. Specifically, the present invention first projects the key points of the database image near the virtual viewpoint to the virtual view according to the database information, and then uses RenderNet LGenerate enhanced local descriptors. Although the original local descriptors are learned to be invariant to view angles, we find that it is better to optimize their feature representations to adapt to the current novel view, especially when the viewpoint changes greatly. It outputs an n-dimensional feature map, and then interpolates the local features of the key points from this feature map. Similarly, the method trains RenderNet L to minimize the distance between its output and the local features of the real database images generated by DescNet.
[0092] argmin ψ ||RenderNet L (I p , F p )-DescNet(I)·M|| 2 (2)
[0093] In the formula, Ψ is the parameter of RenderNet L . The role of the mask M is to only consider the loss of the projection position.
[0094] It should be noted that RenderNet L and DescNet respectively output virtual view local features and real database image local features under the same view, and the distance between the two features is calculated, and the smaller the distance between the two, the stronger and more accurate the virtual view local feature ability of RenderNet L output. Therefore, the goal of this network is to minimize the local feature distance between RenderNet L and DescNet. The local features of the image refer to the local region descriptors extracted from the image, which have uniqueness and robustness, such as corner points, edge points, or texture significant regions, etc. These features are usually represented as high-dimensional vectors (such as SIFT, ORB, HOG, etc. Features), and each dimension in the vector corresponds to a certain attribute of the feature (such as gradient direction, intensity, color, etc.). The distance between local features is essentially the geometric distance between two high-dimensional feature vectors in vector space, and the smaller the distance, the more similar the features; the larger the distance, the greater the difference.
[0095] Network architecture of RenderNet. The present application uses AP-GeM and SuperPoint as the image retrieval and local feature extraction network of the real database image. For the proposed RenderNet, a modified ResNet18 is used as the backbone network. As Figure 3As shown, compared to the traditional ResNet18, the RenderNet proposed in this invention has two main features. First, it adds a concatenated network. This concatenated network takes an input of H×W×3, the color of the projected points, and outputs features of size 1 / 8. These 1 / 8 local features are then concatenated for further processing. Second, it adds two branch networks, namely RenderNet. G and RenderNet L branch, RenderNet G and RenderNet L The branches output global and local feature maps respectively. For RenderNet G The branch, after generating features of size 1 / 32, uses a GeM pooling layer and a fully connected layer as an AP-GeM network to output an m-dimensional global feature. RenderNet L The branch generates a local feature map of size 1 / 8, using the same convolutional layers as SuperPoint.
[0096] Regarding the training process of RenderNet. For example... Figure 3 As shown, Figure 3 The system is divided into two layers. The upper layer consists of DescNet and RetrieveNet. These two networks take a real image I as input and then obtain the global and local features of the real image from the same viewpoint, i.e., the global and local features in the image. These two networks are pre-trained and are called teacher models, mainly used to train the two student models, RenderNet. G and RenderNet L For teaching purposes. Here, "teaching" refers to RenderNet. G and RenderNet L The training objective for both networks is to make their outputs as close as possible to the outputs of DescNet and RetrieveNet. That is... Figure 3 The lower layer shown here contains some of the lower-level content, and the final output is RenderNet. G and RenderNet L The outputs of the two networks, Global and Local features, are used to calculate the L2 Loss along with the outputs of DescNet and RetrieveNet. Reducing the L2 Loss is the entire training process. Essentially, the outputs of DescNet and RetrieveNet are the target values, and RenderNet... Gand RenderNet L The output values are training values.
[0097] The indoor relocalization method of the present application can be applied in the fields of virtual reality / augmented reality, unmanned autonomous driving, etc. The following takes an augmented reality product as an example to introduce the practical application of the method in the product landing. Other products such as unmanned autonomous driving, etc., the positioning and role of the algorithm in them are basically the same.
[0098] As Figure 4 shown, the current most advanced method HLoc
[24] and the RenderNet method proposed in this paper are compared. The top image pair shows the retrieval result with the real database image, and 11 feature points are retrieved. The bottom image pair shows the retrieval result with the virtual synthesized view, and 338 feature points are retrieved. Both of the two image pairs show the hypothetical matching items from the SuperGlue matcher
[25] . From the matching result of the top image pair, it can be seen that the most similar image in the database may still have insufficient overlap with the query image, making it difficult to retrieve and match these images, thus leading to performance degradation. From the matching result of the bottom image pair, it can be seen that the virtual view not explicitly generated can have greater overlap with the query image, thus can help image retrieval and image matching, and obtain better relocalization performance.
[0099] Augmented reality is a technology that superimposes virtual information onto the real world in real time. First, the AR device needs to perceive and understand the actual environment in which the user is located. This is usually done through various sensors on the device, such as cameras, depth sensors, etc. The camera can obtain the image of the environment, and the depth sensor can obtain the distance between the object and the device. In some advanced devices, there may also be a LiDAR to obtain more accurate environmental data. Next, the AR system needs to understand the characteristics of the environment by processing the collected data. For example, it can use image recognition technology to identify certain objects in the environment, and use SLAM (Simultaneous Localization and Mapping) technology to understand the user's location and environment. After understanding the environment, the AR system will generate virtual objects and superimpose them onto the scene seen by the user. This requires computer graphics technology and strong processing power to render these objects in real time. Finally, in order to provide a better user experience, the AR system usually has some user interaction functions. For example, the user may need to interact with the superimposed virtual objects through the touch screen or gestures. All the above processes need to be carried out in real time so that the AR system can quickly adapt and update the virtual objects when the user moves or the environment changes. This requires a powerful processor and efficient algorithms. For example Figure 5, users can display a virtual 3D Earth model in the real environment through AR glasses or mobile phones.
[0100] In augmented reality applications, a core module is SLAM. SLAM (Simultaneous Localization and Mapping) is a key autonomous navigation technology that helps robots or AR devices to localize and map simultaneously in an unknown environment. Localization refers to determining the position and orientation of an object in an environment. In SLAM, localization usually uses the device's sensors (such as cameras, lidar, inertial measurement units, etc.) to estimate its own position and orientation. By continuously collecting and processing this data, the device can update its position in the environment in real time. At the same time, the device needs to create a map of the environment. It needs to identify and track feature points in the environment and store them in the map. In SLAM, the map can be 2D (such as grid or topology) or 3D (such as point cloud). The most critical part is that these two processes need to be done simultaneously. That is, the device needs to use the map to localize while building the map, and needs to update the map while localizing. SLAM technology has a wide range of applications in autonomous vehicles, drones, AR / VR, robots, etc. Figure 6 Demonstration is an example of SLAM building a map and localization.
[0101] Re-localization plays a crucial role in SLAM. Simply put, re-localization refers to finding the accurate position of a robot or device in a map when it has lost its way or detected a certain position deviation. Here are the detailed steps: 1. Position deviation occurs: In actual operation, due to various reasons (such as sensor noise, light changes, dynamic environment, temporary tracking failure, etc.), the estimated position of the device may deviate, or the device may completely lose its estimate of the position (such as the device rotates at a large angle in a short time, or is manually moved to a new position by a person), at which point re-localization is needed. 2. Re-localization process: When re-localization is needed, the device will try to find the map feature points in the current field of view and match them with the information previously stored in the map to determine its accurate position in the map. The re-localization algorithm usually requires fast and robust, because if the device cannot quickly find its own position, it may not be able to continue to perform tasks or cause a decline in user experience. It is worth noting that for SLAM systems with loop closure detection, re-localization and loop closure detection are almost similar. Loop closure detection is that the device recognizes that it has returned to an area it has visited before, and then closes the "loop" in the map, correcting the navigation error accumulated during movement.
[0102] In general, relocalization is a very important part of SLAM technology, and plays an important role in ensuring continuous and effective operation of the device and achieving high-precision positioning and mapping.
[0103] The core benefit of the present application is to improve the performance of relocalization in SLAM through the above method. The specific application scenario is generally embedded in other more mature product scheme modules, such as automatic driving based on SLAM, VR / AR applications, to improve the performance and effect of other product schemes. Because most products in practice use tracking capabilities, and actual tracking inevitably encounters failure situations, relocalization technology is needed to recover from the failure state.
[0104] As another specific embodiment, in the mega constellation operation and maintenance scenario, the mega constellation generally refers to a network composed of hundreds or thousands of satellites, which operate in the Earth's orbit to provide communication, navigation, remote sensing and other services for the world. The operation and maintenance management of these constellations is a complex and highly challenging task, involving satellite launch, on-orbit operation, fault detection, maintenance and final decommissioning, etc. Based on mixed reality technology, operation and maintenance personnel can view the running state of the satellite, communication state, fault handling and other common satellite operation and maintenance operations through gestures, voice or special equipment. As shown in Figure 7 The virtual 3D Earth model inside can be seen to be fixed at a certain position in the real space, and the position of the Earth 3D model display will not change when the user observes around the Earth model.
[0105] In the above product, the virtual Earth and satellite model is placed in a fixed position in the real scene, and is viewed and interacted with using a mixed reality device. Here, the mapping and positioning functions in SLAM are used. Thanks to the SLAM positioning algorithm, the virtual spatial coordinates are calculated based on the features in the real scene. The accuracy and stability of this coordinate determine whether the visual effect of the user observing the Earth model is stable or jittery. The improved scheme improves the SLAM positioning algorithm, which is equivalent to improving the stability and accuracy of the positioning coordinates of the final Earth model. It greatly improves the experience of mixed reality applications. There is no problem of dizziness caused by jitter. Considering that this operation and maintenance scenario is in an indoor environment, the lighting environment is weak, which can easily cause the loss of positioning view in the SLAM positioning process. In this scenario, the performance of the relocalization algorithm is highly dependent. The present application is also based on the poor performance of the relocalization algorithm in this application.
[0106] The above is only a preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which will not affect the effect of the implementation of the present application and the practicality of the patent.
Claims
1. An indoor relocation method based on virtual view synthesis, characterized in that, The method is as follows: Retrieve the query image and the actual database image, and generate a virtual view; The query image and real database images are processed by a retrieval network and a feature extraction network to generate global features G. q G r and local features L q L r These are used for image retrieval and image matching, respectively; the virtual view is processed using RenderNet to generate global features G. v and local features L v ; Calculate the global feature G q With global features G v The distance between them is used to find the k most similar real database images to the query image, and a 2D-3D correspondence is established between the query image and the found real database images. The initial camera pose is estimated based on the 2D-3D correspondence. The camera pose is optimized by using the virtual view under the initial camera pose as a relocalization reference.
2. The indoor relocation method based on virtual view synthesis according to claim 1, characterized in that, The RenderNet includes: Global features G used for image retrieval v RenderNet G Branches, and A set of local features L used to generate key points in the virtual view v RenderNet for sparse matching L Branch.
3. The indoor relocation method based on virtual view synthesis according to claim 2, characterized in that, The RenderNet G The branch employs a widely used generalized mean pooling layer to extract global features G. v The global feature G v It is an m-dimensional global feature vector.
4. The indoor relocation method based on virtual view synthesis according to claim 3, characterized in that, The global feature G is calculated using the following formula. q With global features G v Distance between: argmin φ ||RenderNet G (I p ,F p )-RetrievalNet(I)‖ 2 ; In the formula, Φ is RenderNet G The parameter, I p and F p These are the projection color and projection characteristics of the virtual view, respectively, and I is the real image captured from the virtual view's perspective.
5. The indoor relocation method based on virtual view synthesis according to claim 2, characterized in that, Based on the database information, key points of the database image near the virtual viewpoint are projected onto the virtual view, and then RenderNet is used. L An enhanced local descriptor is generated, which is an n-dimensional feature map from which local features of key points are interpolated.
6. The indoor relocation method based on virtual view synthesis according to claim 5, characterized in that, The local feature L r and local features L v The distance between them is calculated using the following formula: argmin ψ ||RenderNet L (I p ,F p )-DescNet(I)·M|| 2 ; In the formula, Ψ is RenderNet L The parameters are: M is the mask, I... p and F p These are the projection color and projection characteristics of the virtual view, respectively, and I is the real image captured from the virtual view's perspective.
7. The indoor relocation method based on virtual view synthesis according to claim 1, characterized in that, The method for optimizing the camera pose is as follows: Render a new virtual view based on the initial camera pose, and then use RenderNet. L The generated local features are then sparsely matched again, and the camera pose is estimated here.