An intelligent vision positioning method
By introducing multi-level depth embedding Transformer module and depth-guided smooth constraints in visual positioning methods, combining color images and depth image features, the problem of poor visual positioning accuracy in aerial scenes is solved, and higher robustness and positioning accuracy are achieved.
Patent Information
- Application Number
- CN202411189144.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-08-28
AI Technical Summary
The existing visual positioning methods are poorly accurate in complex aviation scenarios and lack effective exploration of the global context information and spatial information of the scene, resulting in poor robustness.
The multi-level depth embedding Transformer module is adopted to combine color images and depth image features, and long-distance dependencies are modeled through Transformer to mine the global context information of aerial images, and train the network through depth-guided smooth constraints, regression losses and reprojection losses to improve the robustness and accuracy of visual positioning.
Effectively explore the global context information of aerial images, improve the robustness of the network to visual artifacts, improve the performance of visual positioning of aviation scenes, and obtain performance that is better than existing methods.
Smart Images

Figure CN119152031B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual positioning, and in particular to an intelligent visual positioning method. Background Art
[0002] Visual positioning, as one of the basic tasks of computer vision, has been widely applied in fields such as autonomous driving, simultaneous localization and mapping, and virtual reality. Visual positioning aims to estimate the 6-DoF pose of a camera in a known scene based on a query image, that is, the three-dimensional position coordinates and three-dimensional angular deflections of the camera in the world coordinate system. In recent years, relevant scholars have conducted in-depth research on visual positioning in indoor or outdoor street scenes and achieved excellent performance. However, there is little research on visual positioning methods for aviation scenes, which severely restricts the development of aviation systems relying on navigation. Therefore, exploring an intelligent visual positioning method to achieve precise positioning of cameras in aviation scenes has important research significance and application value.
[0003] Traditional visual positioning algorithms usually estimate the camera pose based on manually designed features, which can be mainly divided into geometric structure-based methods and image retrieval-based methods. Among them, the geometric structure-based method first extracts feature points in the query image, then matches the extracted 2D feature points with 3D coordinate points in the scene model, and finally calculates the camera pose based on the obtained 2D-3D matching relationship. The image retrieval-based method needs to first retrieve the nearest neighbor image of the query image by matching the global features of the images, then match the 2D feature points of the two images, and finally calculate the camera pose based on the obtained 2D-2D matching relationship. Although traditional methods have made great progress, limited by problems such as poor robustness and low generalization, in some complex scenes with drastic light changes and motion blur, positioning failures may occur.
[0004] In recent years, visual positioning algorithms based on deep learning have gradually shown more superior performance than traditional methods. Kendall et al. proposed a visual positioning algorithm based on a convolutional neural network, which learned the mapping relationship from the query image to the camera pose through the convolutional neural network to achieve visual positioning. However, this method only achieved good performance in indoor or outdoor street scenes and had poor accuracy when directly applied to complex aviation scenes. Recently, Yan et al. proposed an extensible aviation visual positioning algorithm based on multi-modal synthetic data, which improved the accuracy of visual positioning in aviation scenes by learning cross-modal visual representations. However, existing methods only use convolutional neural networks to extract features of aerial images and do not effectively explore the global context information of the scene. In addition, existing methods only use color images to extract features and lack explicit spatial information, so they are not robust enough to handle visual artifacts widely present in aerial images and are difficult to meet the requirements of practical applications. Summary of the Invention
[0005] The present invention provides an intelligent visual positioning method, aiming to effectively explore the global context information of aerial images, and by mining the spatial information contained in depth images, improve the robustness of the network to visual artifacts widely existing in aerial images, so as to achieve effective visual positioning in aerial scenes, as described in detail below:
[0006] An intelligent visual positioning method, the method comprising:
[0007] Obtaining a color image feature sequence and a depth image feature sequence through a feature sequence extraction module, and using them as inputs to a multi-level depth-embedded Transformer module;
[0008] The multi-level depth-embedded Transformer module is composed of multiple depth-embedded units and Transformer layers. Each depth-embedded unit takes a feature sequence as input and aims to output a spatially aware enhanced feature sequence;
[0009] Feeding the scene feature representation obtained by the multi-level depth-embedded Transformer module into a prediction head to obtain a scene coordinate prediction result. A horizontal gradient operator and a vertical gradient operator are respectively applied to the scene coordinate prediction result to generate a horizontal gradient and a vertical gradient of the scene coordinate prediction result; the horizontal gradient operator and the vertical gradient operator are also respectively applied to the depth image to generate a horizontal gradient and a vertical gradient of the depth image;
[0010] Training an intelligent visual positioning network using depth-guided smoothing constraints, regression loss, and reprojection loss, and constructing a pose solver for pose sampling and pose refinement.
[0011] Wherein, each depth-embedded unit takes a feature sequence as input, and the spatially aware enhanced feature sequence to be output is:
[0012] The input feature sequence and are respectively processed through a convolutional layer to generate low-dimensional latent embeddings and
[0013] The latent embedding is passed through a λ-smoothed spatial Softmax layer to generate a mask of class spatial attention. The generated mask g fovea is dot-multiplied with the latent embedding to generate a spatially enhanced embedding
[0014] The latent embedding and the spatially enhanced embedding After weighted fusion, a feature sequence with enhanced spatial perception is obtained through a convolutional layer The calculation formula is expressed as:
[0015]
[0016] Among them, g3(·) represents a convolutional layer, and α and β represent learnable weight parameters.
[0017] Among them, the depth-guided smooth constraint is: by and imposing an L1 penalty, and using and 's edge-aware terms to weight this penalty to achieve; as follows:
[0018]
[0019] Among them, d ij represents the depth value of the query image at the (i,j) position, and s ij represents the scene coordinate value predicted by the network at the (i,j) position of the query image.
[0020] Among them, the low-dimensional latent embedding and is:
[0021]
[0022] Among them, both g1(·) and g2(·) consist of a convolutional layer.
[0023] Among them, the spatial enhancement embedding is:
[0024]
[0025] Among them, represents 's feature vector at the (i,j) position, and γ l represents a learnable weight parameter, indicating the dot product operation.
[0026] The beneficial effects of the technical solution provided by the present invention are:
[0027] 1. The present invention utilizes the ability of Transformer to model long-range dependencies to mine the global context information of aerial images; at the same time, considering the characteristics of depth images in describing the spatial positions of objects, depth cues are introduced into the network to explicitly perceive spatial information, thereby enhancing the robustness of the network to visual artifacts and further improving the performance of aerial scene visual positioning;
[0028] 2. The present invention designs a multi - level deep - embedded Transformer module. By adaptively fusing depth - image features and color - image features, it enhances the network's perception ability of the spatial structure of the aerial scene, thereby improving the network's robustness to visual artifacts. In addition, a depth - guided smooth loss is designed. Under the guidance of depth information, the network is encouraged to learn the geometric characteristics of piece - wise smooth scene coordinates, thereby improving the prediction accuracy of scene coordinates.
[0029] 3. The present invention conducts experimental verification on two datasets for visual localization of aerial scenes and can obtain performance superior to existing methods for visual localization of aerial scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of an intelligent visual localization method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0032] The following illustrates the detailed embodiments of the intelligent visual localization method in the present invention through examples.
[0033] I. Construct a feature sequence extraction module
[0034] For an input color image and its corresponding depth image, the embodiment of the present invention constructs a feature sequence extraction module to obtain a color - image feature sequence T rgb and a depth - image feature sequence T depth . Specifically, the input color image and depth image first respectively pass through convolutional layers to obtain a color - image feature F r and a depth - image feature F d . Subsequently, the color - image feature F r and the depth - image feature F d are respectively added to a learnable position encoding P to introduce position information, and through a flattening operation, the color - image feature and depth - image feature with position information are respectively transformed into feature sequences T rgb and T depth .
[0035] The above - mentioned process is expressed by the following formula:
[0036] T rgb =[P1 + F r,1 , P2 + F r,2 ,..., P i +F r,i
[0037] T depth =[P1 + Fd,1 , P2 + F d,2 ,..., P j + F d,j
[0038] Among them, i represents the number of pixels in the color image feature map, j represents the number of pixels in the depth image feature map, and preferably i = j = 1350.
[0039] II. Design a multi-level depth-embedded Transformer module
[0040] Considering that color images tend to describe the visual appearance of objects, while depth images can provide explicit spatial information of the entire scene, the embodiments of the present invention design a multi-level depth-embedded Transformer module to fuse the information of depth images and color images and improve the robustness of the network to visual artifacts. Specifically, this module takes the color image feature sequence T rgb and the depth image feature sequence T depth as inputs, and enhances the network's perception ability of the spatial structure of the aerial scene by adaptively embedding the depth image features into multiple levels of the Transformer.
[0041] The designed multi-level depth-embedded Transformer module consists of multiple depth-embedded units and Transformer layers. Among them, the depth-embedded unit M l takes the feature sequences and as inputs, and aims to output a spatially enhanced feature sequence The Transformer layer E l takes the sum of the feature sequences and as an input, and aims to output a globally enhanced feature sequence
[0042] Taking the l-th depth-embedded unit M l as an example, first, the input feature sequences and are respectively processed by a convolutional layer to generate low-dimensional latent embeddings and The calculation formula is expressed as:
[0043]
[0044] Among them, both g1(·) and g2(·) consist of a convolutional layer.
[0045] Then, the latent embedding passes through a λ-smoothed spatial Softmax layer to generate a spatial attention mask gfovea Subsequently, the generated mask g fovea is dot-multiplied with the latent embedding to generate a spatially enhanced embedding The above process can be expressed as:
[0046]
[0047] where denotes the feature vector at the (i, j) position in l and γ represents a learnable weight parameter, indicating the dot-multiplication operation.
[0048] Finally, to adaptively fuse the depth information and the spatially enhanced embedding, the latent embedding and the spatially enhanced embedding after weighted fusion, pass through a convolutional layer to obtain a spatially aware enhanced feature sequence The calculation formula is expressed as:
[0049]
[0050] where g3(·) represents a convolutional layer, and α and β represent learnable weight parameters. Preferably, the multi-level depth embedding Transformer module consists of 12 depth embedding units and 12 Transformer layers.
[0051] Taking the l-th Transformer layer E l as an example, its input is the sum of the output of the l-th depth embedding unit and the output of the (l - 1)-th Transformer layer and the output is a globally aware enhanced feature sequence The calculation formula is expressed as follows:
[0052]
[0053] where LN(·) represents layer normalization. MSA(·) represents multi-head self-attention, and MLP(·) represents a multi-layer perceptron, used to enhance the expressive power of the model.
[0054] III. Design of Depth-Guided Smoothness Constraint
[0055] The embodiment of the present invention designs a depth-guided smoothness constraint, which guides the network to obtain a piecewise smooth scene coordinate prediction result by utilizing the geometric characteristics of piecewise smoothness of the depth image, thereby improving the prediction accuracy of the scene coordinates.
[0056] First, the scene features obtained by the multi-level depth embedding Transformer module It represents the input into the prediction head to obtain the scene coordinate prediction result s pred , and the calculation formula is expressed as:
[0057]
[0058] Among them, H(·) represents the prediction head, which consists of 4 convolutional layers and a rectified linear unit.
[0059] Then, the horizontal gradient operator and the vertical gradient operator act on the scene coordinate prediction result s respectively pred , generating the horizontal gradient of the scene coordinate prediction result and the vertical gradient Meanwhile, the horizontal gradient operator and the vertical gradient operator also act on the depth image respectively, generating the horizontal gradient of the depth image and the vertical gradient The designed depth-guided smooth constraint is achieved by imposing an L1 penalty on and , and using the edge-aware terms of and to weight this penalty.
[0060] The above process is expressed by the formula as follows:
[0061]
[0062] Among them, d ij represents the depth value of the query image at the position (i, j), and s ij represents the scene coordinate value predicted by the network at the position (i, j) of the query image.
[0063] IV. Training the intelligent visual positioning network
[0064] In order to obtain accurate scene coordinate prediction results, the embodiments of the present invention use multiple loss functions to train the network. Specifically, a depth-guided smooth constraint is designed to encourage the network to learn the geometric characteristics of piecewise smooth scene coordinates. Meanwhile, in order to effectively supervise the intelligent visual positioning network, the embodiments of the present invention adopt a regression loss L regress , and its formula is expressed as follows:
[0065]
[0066] Among them, represents the ground truth of the scene coordinates corresponding to the query image at the position (i, j), s ij represents the scene coordinate value predicted by the network at the position (i, j) of the query image, and u ijRepresents the uncertainty of the predicted result of the scene coordinates corresponding to the query image at the position (i, j).
[0067] In addition, the embodiment of the present invention also adopts the reprojection loss L reproject to constrain the prediction accuracy of the scene coordinates, and its formula is expressed as follows:
[0068]
[0069] where C represents the internal parameter matrix of the camera, h represents the true value of the camera pose, and P s represents the set of points in the predicted scene coordinates whose reprojection error is less than the set threshold, and P t represents the set of points in the predicted scene coordinates whose reprojection error is greater than the set threshold, and p ij represents the position coordinates of the query image at (i, j).
[0070] Finally, the total loss L smooth combining the depth-guided smooth constraint L regress , the regression loss L reproject and the reprojection loss L total is expressed as follows:
[0071] L total = L regress + L reproject + λ s L smooth
[0072] where λ s represents the weight of the depth-guided smooth constraint, which is set to 0.1 in the embodiment of the present invention.
[0073] The embodiment of the present invention uses the depth-guided smooth constraint, the regression loss and the reprojection loss to train the intelligent visual positioning network, which is divided into two stages. In the first stage, the neural network is trained with synthetic data for 150 epochs to obtain an initial intelligent visual positioning model; in the second stage, the initial model is fine-tuned with real-synthetic data to 4 scenarios, and finally an intelligent visual positioning model that can test the corresponding 4 different scenarios is obtained.
[0074] V. Constructing a pose solver
[0075] After training the intelligent visual positioning network, in order to further obtain the pose of the camera, the embodiment of the present invention constructs a pose solver for pose sampling and pose refinement. Specifically, the pose solver consists of a PnP operator, which optimizes the camera pose by minimizing the reprojection error of the scene coordinates.
[0076] In the embodiment of the present invention, the scene coordinates output by the intelligent vision positioning network and the internal parameter matrix of the camera are input into the PnP operator, and the 6-DoF pose of the camera can be obtained. Among them, the formula of the reprojection error is expressed as follows:
[0077]
[0078] Among them, r(·) represents the reprojection error, and p i represents the two-dimensional image coordinates associated with pixel i, K represents the internal parameter matrix of the camera, represents the pose hypothesis of the camera, and y i represents the scene coordinates associated with pixel i.
[0079] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0080] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent visual positioning method, characterized in that: The method comprises: The feature sequence extraction module is used to obtain the color image feature sequence and the depth image feature sequence, and these are used as the input of the multi-level deep embedding Transformer module. The multi-level deep embedding Transformer module is composed of multiple deep embedding units and Transformer layers, each deep embedding unit takes a feature sequence as input and aims to output a feature sequence with enhanced spatial perception; The scene feature representation obtained by the multi-level depth embedding Transformer module is sent to the prediction head to obtain the scene coordinate prediction result. The horizontal gradient operator and the vertical gradient operator act on the scene coordinate prediction result respectively to generate the horizontal gradient and vertical gradient of the scene coordinate prediction result; the horizontal gradient operator and the vertical gradient operator also act on the depth image respectively to generate the horizontal gradient and vertical gradient of the depth image; Use depth-guided smoothness constraints, regression loss, and reprojection loss to train an intelligent visual localization network, and build a pose solver for pose sampling and pose refinement; Each of the deep embedding units takes a feature sequence as input and aims to output a feature sequence with enhanced spatial perception as follows: The input feature sequence and After being processed by convolutional layers, low-dimensional potential embeddings are generated. and represents the feature sequence output by the l-1th Transformer layer, represents the feature sequence output by the l-1th deep embedding unit, represents the colored latent embedding, represents deep latent embedding; Embed the potential After a λ-smoothed spatial Softmax layer, a mask similar to spatial attention is generated, and the generated mask g fovea With latent embedding Dot product to generate spatially enhanced embedding λ represents a parameter in the smooth spatial Softmax layer; Potential Embedding and spatially enhanced embedding After weighted fusion, the feature sequence with enhanced spatial perception is obtained through the convolution layer. The calculation formula is expressed as: Among them, g3(·) represents a convolutional layer, and α and β represent learnable weight parameters.
2. The intelligent visual positioning method according to claim 1, characterized in that: The depth-guided smoothness constraint is: and Apply L1 penalty and use and The edge-aware term of is weighted to achieve this; as follows: Among them, d ij represents the depth value of the query image at position (i, j), sij represents the scene coordinate value predicted by the network at position (i, j) of the query image, N represents the number of pixels in the image, represents the horizontal gradient operator, Represents the vertical gradient operator.
3. The intelligent visual positioning method according to claim 1, characterized in that: The low-dimensional latent embedding and for: Among them, g1(·) and g2(·) are both composed of a convolutional layer.
4. The intelligent visual positioning method according to claim 1, characterized in that: The spatially enhanced embedding for: in, express The eigenvector at position (i, j) in l represents a learnable weight parameter, g fovea represents the generated mask and ⊙ represents the dot product operation.
Citation Information
Patent Citations
Monocular depth prediction method based on multi-level feature parallel interactive fusion
CN115578436A
RGB-D saliency target detection method based on three-branch multi-level Transform feature interaction
CN115908856A