Long-term visual localization method based on structure perception and multi-task distillation
By combining structural perception and multi-task distillation technology in visual positioning, a layered positioning network is built, which solves the problems of insufficient visual positioning accuracy and speed in the existing technology, and achieves an efficient and robust visual positioning effect.
Patent Information
- Application Number
- CN202310580112.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-05-23
AI Technical Summary
The prior art has problems of insufficient accuracy and speed in visual positioning, especially in terminal devices with limited computing resources, and the multi-task distillation method can lead to accuracy losses while increasing the speed.
A long-term visual positioning method based on structure perception and multi-task distillation is adopted to construct local and global structural perception through edge features, and combined with a multi-task distillation model, the robustness and real-time nature of the hierarchical positioning network are achieved.
It improves the speed and accuracy of visual positioning, can effectively deal with the problems of viewing angle, light or seasonal changes, and is suitable for application scenarios such as unmanned driving and large-scale visual positioning.
Smart Images

Figure CN116524028B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual positioning technology, and in particular to a long-term visual positioning method based on structure perception and multi-task distillation. Background Art
[0002] Long-term visual localization is a classic problem in computer vision and one of the key steps in autonomous driving. With the continuous maturity of autonomous driving technology in recent years, people have put forward higher requirements for the accuracy and speed of visual localization, especially in some challenging scenarios, such as changes in lighting, weather or seasons. The current method uses local features to construct Structure-from-Motion (SfM) point clouds to achieve 2D-3D matching, which can greatly improve the positioning accuracy. However, in terminal devices with limited computing resources, most network models are limited in application due to the huge number of parameters and the huge amount of calculation in the 2D-3D matching process. In addition, the generalization performance of the model is usually not strong. Edge features can accurately extract structural information in the scene. Although the current visual localization method has achieved good accuracy, its speed needs to be improved. The multi-task distillation method can greatly improve the feature extraction speed, but there will be a certain loss in accuracy.
[0003] The paper "Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Efficient & effective prioritized matching for large-scale image-based localization. IEEE PAMI, 39 (9): 1744–1756, 2017." uses the matching between image features and point clouds to achieve image positioning in point clouds and obtain very high accuracy. The paper "Sarlin PE, Cadena C, Siegwart R, et al. From coarse to fine: Robust hierarchical localization at large scale [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 12716-12725." proposes a hierarchical positioning method. By pre-building the SfM scene model, each 3D point in the model is associated with the corresponding descriptor in the database image. When the query image is input, the coarse positioning is first achieved through global image retrieval, and then the local features of the query image are extracted to achieve fine matching. The paper "Poma XS, Riba E, Sappa A. Dense extreme inception network: Towards a robust cnn model for edge detection [C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2020: 1923-1932." proposes a deep structure model for extracting image edge information and generating thin edges. This method outputs the encoder of each main block of the network to obtain output results containing edge information to varying degrees, and connects each sub-block output to ensure that each deep block can retain edge features well.
[0004] The paper "Torsten Sattler, Bastian Leibe, and Leif Kobbelt. Efficient & effective prioritized matching for large-scale image-based localization. IEEE PAMI, 39 (9): 1744–1756, 2017." can achieve high positioning accuracy, but when the environment becomes very large, the matching process takes a long time to process and cannot run in real time. The paper "Sarlin PE, Cadena C, Siegwart R, et al. From coarse to fine: Robust hierarchical localization at large scale [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 12716-12725." applies multi-task distillation to visual positioning, making the model lightweight enough for deployment on mobile devices. However, simply using knowledge distillation makes the network limited to the teacher model, and there will be a certain loss in accuracy. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a long-term visual positioning method based on structure perception and multi-task distillation, while improving the speed and accuracy of visual positioning.
[0006] A long-term visual localization method based on structure perception and multi-task distillation includes the following steps:
[0007] Step 1: Use edge features to build local structure perception and global structure perception, and extract the network to obtain structure perception constraints;
[0008] Use the edge feature extraction network Dexined. Dexined is a multi-channel output edge detection network that contains six sub-modules. The edge feature extraction network Dexined is divided into two parts: the first three sub-modules and the last three sub-modules. Insert the structure-aware block before the upsampling block of each side output port. The structure-aware module consists of a 1*1 convolution layer and an upsampling block. The first three layers of the edge feature extraction network Dexined and the last three layers of the side output ports are used as the output of the structure-aware constraint through 1*1 convolution, and then the local structure-aware and global structure-aware visualization images are obtained through the upsampling block.
[0009] Use X to represent the input image, Y to represent the true value of the edge, and the edge feature extraction network Dexined has 6 output layers. Each output layer generates a prediction through an upsampling block and uses w i represents the weight of the output of the i-th layer, and the total weight is expressed as W = (w 1 ,w 2 ,...,w 6 ); then the loss function of each side output layer is:
[0010]
[0011] The structure perception module uses cross entropy as the loss function of local and global structure perception. The loss function of the structure perception module is as follows:
[0012]
[0013] Where Y+ and Y- represent edge pixels and non-edge pixels respectively, ε=|Y - | / |Y|, represents the activation value of the sigmoid function at pixel j, is a learnable parameter, and Represent the prediction output of local structure perception and global structure perception respectively, α i ,β,γ are adjustment parameters, y j is the true value of the training image, i is the output layer, i=0, 1, ..., 6; CrossEntropy() is the cross entropy loss function, which is used to measure the difference between the predicted value output by the network and the true label.
[0014] Step 2: Hierarchical localization network implementation of structure perception and multi-task distillation model;
[0015] The local features and global features of the database image are extracted through the hierarchical positioning network SAMLoc, and then the local features are used to build the SfM map offline, and the global features of the database image are used to build the database index; when the query image is input, the global features and local features of the query image are obtained respectively through the structure perception module, and the global features are first used to perform a rough match with the database index through KNN matching, and then the local features obtained are used to perform 2D-3D matching with the SfM map through the random sampling consensus algorithm RANSAC and the matching algorithm PnP, and finally the 6-DOF pose of the query image is obtained;
[0016] The hierarchical positioning network uses MobilenetV3 as the backbone network. The local feature and global feature teacher models are Superpoint and NetVLAD respectively. Then, the structure perception module is introduced to constrain the distillation process. The 7th layer of MobilenetV3 is the local feature extraction branch, and the 18th layer of MobilenetV3 is the global feature extraction branch. The structure perception constraint is added to the 4th and 11th layers of MobilenetV3 to keep the structure of the backbone network unchanged.
[0017] The loss function of the structure perception and multi-task distillation model is divided into local loss and global loss. Formula (3) is the local loss function, formula (4) is the global loss function, and formula (5) is the total loss function of structure perception and multi-task distillation.
[0018]
[0019] in Represents the weight of each loss, d represents the descriptor, s represents the structure-aware constraint, and the subscript s i and t i They represent the student model and the teacher model respectively. The superscripts l and g represent local and global respectively. k represents key points. w i is the regularization term for each loss;
[0020] The local loss function (3) contains three terms: is the local descriptor distillation loss, is the local structure-aware constraint loss, Score key points;
[0021]
[0022] The global loss function (4) contains two terms: is the global descriptor distillation loss, is the global structure-aware constraint loss;
[0023] The total loss function of structure perception and multi-task distillation is as follows (5), where W i is the regularization term, and n is the number of targets that the model needs to learn;
[0024]
[0025] The beneficial effects of adopting the above technical solution are:
[0026] The present invention provides a long-term visual positioning method based on structural perception and multi-task distillation. In view of the poor robustness and real-time performance of previous 2D-3D visual positioning algorithms, the present invention proposes a structural perception module for scene feature extraction, and combines it with a hierarchical positioning network based on multi-task distillation to achieve real-time robust visual positioning. First, the edge detection algorithm is used to obtain robust structural information in the scene, and then the structural information is integrated into the network model of visual positioning in the form of knowledge distillation. At the same time, the main body of the network adopts a multi-task distillation method to use a lightweight model to learn the extraction of local features and global features, which greatly shortens the feature extraction time of hierarchical positioning while ensuring accuracy. The invention can effectively deal with problems such as perspective, light or seasonal changes that occur in the process of long-term visual positioning, so as to serve application scenarios such as unmanned driving and large-scale visual positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A network flow chart in an embodiment of the present invention;
[0028] Figure 2 It is a structure-aware network diagram in an embodiment of the present invention;
[0029] Figure 3 FIG. 4 is a multi-task distillation network diagram according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0031] A long-term visual localization method based on structure perception and multi-task distillation, such as Figure 1 Shown is the main flow chart of the technical solution of the present invention, comprising the following steps:
[0032] Step 1: Use edge features to build local structure perception and global structure perception, and extract the network to obtain structure perception constraints;
[0033] The edge feature extraction network Dexined is used. Dexined is a multi-channel output edge detection network, which contains six sub-modules. The edge feature extraction network Dexined is divided into two parts: the first three sub-modules and the last three sub-modules. The structure perception module Structure-Aware block is inserted before the upsampling block of each side output port. The structure perception module consists of a 1*1 convolution layer and an upsampling block. The first three layers of the edge feature extraction network Dexined and the last three layers of the side output port are used as the output of the structure perception constraint through 1*1 convolution, and then the visualization images of local structure perception and global structure perception are obtained respectively through the upsampling block. The network diagram of the structure perception module is shown in the figure Figure 2 As shown;
[0034] Use X to represent the input image, Y to represent the true value of the edge, and the edge feature extraction network Dexined has 6 output layers. Each output layer generates a prediction through an upsampling block and uses w i represents the weight of the output of the i-th layer, and the total weight is expressed as W = (w 1 ,w 2 ,...,w 6 ); then the loss function of each side output layer is:
[0035]
[0036] The structure perception module uses cross entropy as the loss function of local and global structure perception. The loss function of the structure perception module is as follows:
[0037]
[0038] Where Y+ and Y- represent edge pixels and non-edge pixels respectively, ε=|Y - | / |Y|, represents the activation value of the sigmoid function at pixel j, is a learnable parameter, and Represent the prediction output of local structure perception and global structure perception respectively, α i ,β,γ are adjustment parameters, and the local and global structure perception constraints of this system are Figure 2 The Local output layer and Global output layer in y j is the true value of the training image, this image is like Figure 2 The true value of this image is the combination of black and white pixels. The true value of this place is 1, which means y j , otherwise it is 0, and different weights are assigned to them through ε; i is the output layer, i = 0, 1, ..., 6; CrossEntropy() is the cross entropy loss function, which is used to measure the difference between the predicted value output by the network and the true label.
[0039] Step 2: Hierarchical localization network implementation of structure perception and multi-task distillation model;
[0040] The hierarchical positioning network SAMLoc of this system is divided into two parts: online visual positioning and offline map construction. The hierarchical positioning network SAMLoc extracts local features and global features of the database image, then uses the local features to build the SfM map offline, and uses the global features of the database image to build the database index; when the query image is input, the global features and local features of the query image are obtained respectively through the structure perception module, and the global features are first used to perform a rough match with the database index through KNN, and then the local features obtained are used to perform 2D-3D matching with the SfM map through the random sampling consensus algorithm RANSAC and the matching algorithm PnP, and finally the 6-DOF pose of the query image is obtained; the network framework of this system is as follows Figure 2 shown.
[0041] The database image is to divide the collected image into two parts: the database image and the query image. The database image is to build a database for image storage, and then match it with the database image when the query image is input.
[0042] Superpoint is a self-supervised key point and local descriptor extraction model, NetVLAD is a global feature extraction model, and MobilenetV3 is a lightweight network model with the advantages of small parameters and fast running speed. The hierarchical positioning network uses MobilenetV3 as the backbone network, and the local feature and global feature teacher models are selected as Superpoint and NetVLAD respectively. Then, the structure perception module is introduced to constrain the distillation process; the 7th layer of MobilenetV3 is the local feature extraction branch, and the 18th layer of MobilenetV3 is the global feature extraction. The structure perception constraint is added to the 4th and 11th layers of MobilenetV3, keeping the structure of the backbone network unchanged. The multi-task distillation part of this system is as follows Figure 3 shown.
[0043] The loss function of the structure perception and multi-task distillation model is divided into local loss and global loss. Formula (3) is the local loss function, formula (4) is the global loss function, and formula (5) is the total loss function of structure perception and multi-task distillation.
[0044]
[0045] in Represents the weight of each loss, d represents the descriptor, s represents the structure-aware constraint, and the subscript s i and t i They represent the student model and the teacher model respectively. The superscripts l and g represent local and global respectively. k represents key points. w i is the regularization term for each loss;
[0046] The local loss function (3) contains three terms: is the local descriptor distillation loss, is the local structure-aware constraint loss, Score key points;
[0047]
[0048] The global loss function (4) contains two terms: is the global descriptor distillation loss, is the global structure-aware constraint loss;
[0049] The total loss function of structure perception and multi-task distillation is as follows (5), where W i is the regularization term, and n is the number of targets that the model needs to learn;
[0050]
[0051] This system passes Figure 1 The flowchart shown in the figure is used for visual positioning, and the specific implementation is respectively achieved through Figure 2 and Figure 3 The network block diagram shown is implemented, and finally a 6-DoF pose is obtained to achieve fast and robust positioning.
[0052] In order to verify that the algorithm can achieve good positioning accuracy and can cope with various situations that arise in long-term visual positioning, the present invention is tested on the Aachen Day-Night, CMU Season and Robotcar Season datasets. In the night scene of Aachen Day-Night, the present invention achieves positioning accuracies of 61.2%, 77.6% and 84.7% at 0.25m / 2°, 0.50m / 5° and 5.0m / 10° respectively, which shows that our algorithm can cope well with scenes with day and night changes; in the seasonally changing CMU dataset, we achieve positioning accuracies of 91.2%, 94.0% and 97.3% in the Urban scene respectively, which shows that our method can cope well with scenes with seasonal changes; in the weather changing Robotcar dataset, we achieve positioning accuracies of 52.5%, 78.3% and 94.3% in the Day all scene respectively, which shows that our method is also robust to scenes with weather changes. In addition, in order to demonstrate the efficiency of multi-task distillation, the feature extraction part of the network was tested on the Aachen dataset. Our feature extraction took 20ms and the positioning process took 29ms, which can meet the real-time operation requirements of the actual system.
[0053] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.
Claims
1. A long-term visual localization method based on structure perception and multi-task distillation, characterized in that: The following steps are involved: Step 1: Use edge features to build local structure perception and global structure perception, and extract the network to obtain structure perception constraints; Use X to represent the input image, Y to represent the true value of the edge, and the edge feature extraction network Dexined has 6 output layers. Each output layer generates a prediction through an upsampling block and uses w i represents the weight of the output of the i-th layer, and the total weight is expressed as W = (w 1 ,w 2 ,...,w 6 ); then the loss function of each side output layer is: Where Y+ and Y- represent edge pixels and non-edge pixels respectively, ε=|Y-| / |Y|, represents the activation value of the sigmoid function at pixel j, is a learnable parameter; The edge features are extracted using an edge feature extraction network Dexined; The edge feature extraction network Dexined is a multi-channel output edge detection network, which includes six submodules. The edge feature extraction network Dexined is divided into two parts, namely, the first three submodules and the last three submodules. A structure-aware block is inserted before the upsampling block of each side output port. The structure-aware module consists of a 1*1 convolution layer and an upsampling block. The first three layers and the last three layers of the edge feature extraction network Dexined are used as the output of the structure-aware constraint through 1*1 convolution, and then the visualization images of local structure perception and global structure perception are obtained through the upsampling block respectively. The structure perception module uses cross entropy as the loss function of local and global structure perception. The loss function of the structure perception module is as follows: in and Represent the prediction output of local structure perception and global structure perception respectively, α i ,β,γ are adjustment parameters, y j is the true value of the training image, i is the output layer, i=0, 1, ..., 6; CrossEntropy() is the cross entropy loss function, which is used to measure the difference between the predicted value output by the network and the true label; Step 2: Hierarchical localization network implementation of structure perception and multi-task distillation model; The local features and global features of the database image are extracted through the hierarchical positioning network SAMLoc, and then the local features are used to build the SfM map offline, and the global features of the database image are used to build the database index; when the query image is input, the global features and local features of the query image are obtained respectively through the structure perception module, and the global features are first used to perform a rough match with the database index through KNN matching, and then the local features obtained are used to perform 2D-3D matching with the SfM map through the random sampling consensus algorithm RANSAC and the matching algorithm PnP, and finally the 6-DOF pose of the query image is obtained; The hierarchical positioning network uses MobilenetV3 as the backbone network, uses Superpoint and NetVLAD as the local feature and global feature teacher models respectively, and then introduces a structure-aware module to realize the constraints on the distillation process; the 7th layer of MobilenetV3 is the local feature extraction branch, and the 18th layer of MobilenetV3 is the global feature extraction. The structure-aware constraints are added to the 4th and 11th layers of MobilenetV3 to keep the structure of the backbone network unchanged.
2. The long-term visual positioning method based on structure perception and multi-task distillation according to claim 1 is characterized in that: The loss function of the structure-aware and multi-task distillation model is divided into a local loss function and a global loss function.
3. The long-term visual positioning method based on structure perception and multi-task distillation according to claim 2 is characterized in that: The local loss function is shown in formula (3), the global loss function is shown in formula (4), and formula (5) is the total loss function of structure perception and multi-task distillation; in Represents the weight of each loss, d represents the descriptor, s represents the structure-aware constraint, and the subscript s i and t i They represent the student model and the teacher model respectively. The superscripts l and g represent local and global respectively. k represents key points. w i is the regularization term for each loss; The local loss function (3) contains three terms: is the local descriptor distillation loss, is the local structure-aware constraint loss, Score key points; The global loss function (4) contains two terms: is the global descriptor distillation loss, is the global structure-aware constraint loss; The total loss function of structure perception and multi-task distillation is as follows (5), where W i is the regularization term, and n is the number of targets that the model needs to learn;