Transform-based low-altitude unmanned aerial vehicle multi-view perception method
By employing a Transformer-based multi-view perception method, and utilizing viewpoint self-supervised depth correction and dynamic scale fusion, the scale uncertainty and geometric consistency problems in depth estimation of low-altitude UAVs are solved, achieving efficient sparse viewpoint BEV representation and improving the three-dimensional perception capability of low-altitude aircraft.
Patent Information
- Application Number
- CN202511508874.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Existing technologies for depth estimation of low-altitude UAVs suffer from scale uncertainty and poor geometric consistency. They are particularly difficult to meet the requirements for real-time performance and accuracy under complex attitudes, and their dense computation and real-time performance are insufficient.
A Transformer-based multi-view perception method is adopted, including the view self-supervised depth correction mechanism VAST, the dynamic scale fusion model TDSF, and the sparse view BEV fusion model DSSA. The depth map is corrected through a self-attention mechanism, the scale factor is dynamically estimated by combining navigation information, and BEV is represented under sparse view.
It achieves end-to-end optimization of visual perception tasks for low-altitude aircraft, improves the accuracy and real-time performance of depth estimation, enhances the robustness of 3D perception, and provides a solid foundation for autonomous obstacle avoidance and path planning.
Smart Images

Figure CN120976810A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a multi-view perception method for low-altitude unmanned aerial vehicles based on a Transformer. BACKGROUND
[0002] Low-altitude unmanned aerial vehicles are increasingly used in urban logistics, disaster rescue and other scenarios, and real-time perception of the surrounding environment is required.
[0003] Due to the limitations of on-board load and energy consumption, it is often not feasible to use multiple cameras or laser radars to build a dense three-dimensional map; a more common solution is to obtain images through monocular or a small number of cameras, then estimate the depth and convert it to a BEV representation. However, existing depth estimation models such as RSA, Hybrid Transformer and DynaDepth have problems of scale uncertainty and poor geometric consistency; the depth pre-training model Depth Anything V2 is trained based on large-scale synthetic data and pseudo-labels, which improves the prediction accuracy; DynaDepth improves the scale robustness by fusing image features and IMU, but these methods lack adaptive correction of body pose errors, making it difficult to meet the depth estimation needs of low-altitude aerial vehicles under complex attitudes. In terms of BEV representation, Lift-Splat-Shoot (LSS) projects image features to the bird's eye view plane, but the spatial resolution is low and cannot handle dynamic scenes; methods such as SDG-OCC and BEVFusion improve occupancy prediction accuracy through multi-modal fusion or three-dimensional convolution, but the computational overhead is large, and the above methods are effective in autonomous driving, but still have problems of dense computation and insufficient real-time performance in low-altitude unmanned aerial vehicle scenarios.
[0004] Therefore, how to design a multi-view perception framework that can adapt to the attitude changes of low-altitude aerial vehicles and also consider scale consistency and real-time performance is still a challenge. SUMMARY
[0005] To solve the above problems in the prior art, the present application provides a multi-view perception method for low-altitude unmanned aerial vehicles based on a Transformer, which solves the problem of large relative depth prediction error caused by the characteristics of low-altitude flight scenarios, scale ambiguity of monocular methods, and geometric deviation and low information utilization of bird's eye view (BEV) fusion under sparse viewing angles.
[0006] In order to achieve the above-mentioned purposes, the technical solution adopted by the present application is as follows: The present application provides a multi-view perception method for low-altitude unmanned aerial vehicles based on a Transformer, comprising: S1: obtaining images taken by multiple cameras; S2: Using Depth Anything as the basic monocular depth estimation model, the relative depth map corresponding to each image is obtained, and the view self-supervised depth correction mechanism VAST is used to correct the relative depth map. S3: Based on the corrected relative depth map, the Transformer-driven dynamic scale fusion model TDSF is used to fuse visual depth and navigation information through Transformer, dynamically estimate the optimal scale factor, and obtain the absolute depth map based on the optimal scale factor. S4: Based on the absolute depth map, the sparse view BEV fusion model DSSA with dynamic spatial self-attention is used to convert it into a BEV representation, and the final BEV feature map is obtained.
[0007] Furthermore, the VAST (View-Based Self-Supervised Depth Correction) mechanism for correcting the relative depth map includes: A1: Extract intermediate feature maps from the Depth Anything encoder, and use three independent convolutional mappings on the target pixel and its neighboring pixels to obtain the query vector, key vector, and value vector:
[0008] in, For query vector, For key vectors, For value vectors, , and These are the linear projection matrices used in the Transformer to generate the query, key, and value, respectively. For the Depth Anything encoder in the camera The intermediate feature map, Image pixel coordinates, for The neighborhood set, For neighboring pixels, For the first One camera, and Represents the pixel coordinates in the horizontal and vertical directions; A2: Calculate the similarity between the query vector of the target pixel and the key vector of each pixel in the neighborhood, and obtain the attention weights through normalization:
[0009] in, Attention weights, superscript For transpose, The key vector is obtained by linear projection of the neighboring pixels. The normalization factor for the feature dimension. Indicates in neighborhood Other pixels traversed in the middle; A3: The corrected relative depth value is obtained based on attention weight aggregation, and the corrected relative depth value is:
[0010] in, for Corrected relative depth value This represents the original relative depth. A4: Construct the photometric reprojection error as the photometric loss function, and introduce a gradient smoothing regularization term as the smoothing loss function. Smooth and suppress noise on the corrected relative depth values to obtain the total loss function. :
[0011]
[0012]
[0013] in, Let be the photometric loss function. Let slip loss function, The weight hyperparameters are used to smooth the loss. For the set of valid pixels, Indicates time The images in The value at that location, Indicates time An image is a function of pixel coordinates to pixel values. For the camera provided by the navigation sensor at any time arrive pose change matrix, This is a transformation function that backprojects pixels onto 3D and then onto the coordinate system of the next frame image according to the corrected depth. Describing the L1 norm, and These represent the difference operators in the horizontal and vertical directions, respectively.
[0014] Furthermore, based on the corrected relative depth map, the Transformer-driven Dynamic Scale Fusion Model (TDSF) is used to fuse visual depth and navigation information through a Transformer, dynamically estimating the optimal scale factor, and obtaining the absolute depth map based on the optimal scale factor, including: S301: Obtain the initial scale factor using the known absolute altitude of the aircraft:
[0015] in, This is the initial scale factor. This refers to the actual altitude of the aircraft. It is the set of pixels representing the ground region in a top-down camera image. This is the relative depth value corrected for a top-down camera view. Represents the average value; S302: Calculate the geometric scale factor based on the relative and absolute displacements of the navigation sensors. :
[0016] in, This is relative displacement. The absolute displacement vector provided for the navigation sensor. Let be the magnitude of the vector; S303: Based on the initial scale factor, the Transformer-driven dynamic scale fusion model TDSF is used to fuse multimodal temporal information, refine and smooth the geometric scale factor, and output the final optimal scale factor. S304: Multiply the relative depth map by the optimal scale factor by pixels to obtain the absolute depth map.
[0017] Furthermore, the relative displacement of the navigation sensor includes: B1: Acquire two consecutive frames of images, the corrected depth map, and the pose transformation and absolute displacement vector from the navigation sensor; B2: Assume that the set of matching feature points in two consecutive frames is... The corrected relative depth is and Using the intrinsic parameter matrix and attitude transformation The pixels are back-projected into three-dimensional space to obtain the projected pixels:
[0018] in, and This indicates that the same scene point matched in two consecutive frames is at time [time value missing]. and pixel coordinates, To match the number of feature pairs, and They are respectively and The corresponding corrected relative depth, and The coordinates are homogeneous pixels, with superscript. Represents the camera intrinsic parameter matrix The inverse matrix, and They are respectively and The corresponding projected pixels; B3: Relative displacement calculated using depth matching based on the projected pixels. It can be obtained from the difference in coordinates between the two points:
[0019] in, and It is the rotation and translation from the navigation sensor.
[0020] Furthermore, based on the initial scale factor, the Transformer-driven Dynamic Scale Fusion Model (TDSF) is used to fuse multimodal temporal information to refine and smooth the geometric scale factor, outputting the final optimal scale factor, including: C1: Constructing visual feature vectors and navigation feature vectors The two are then concatenated and projected onto a unified dimension through a fully connected layer to obtain the input labels for the Transformer:
[0021] in, and For learnable matrices, Input markers for the Transformer; C2: Employs the Transformer-driven dynamic scale fusion model TDSF's multi-head self-attention mechanism for historical data. Frame input token sequence Perform feature interaction:
[0022] in, For querying the matrix, The key matrix, For value matrices, For history The first frame in the input token sequence Frame input marker, This is the output of the multi-head self-attention operation; C3: Based on the input labeled sequence after feature interaction, the Transformer outputs the scale increment predicted by the feedforward neural network FFN. :
[0023] in, For the first Scale factor after frame fusion; C4: Based on geometric scale factor and scale increment Update the scale factor for the next time step. :
[0024] C5: Based on the scale factor, and using the reprojection error of multiple frames as supervision, a loss function is constructed to optimize the Transformer parameters, thereby obtaining the optimal scale factor. The loss function is:
[0025] in, For projection function, For loss function, Indicates time Three-dimensional rotation operator, For a moment The three-dimensional translation vector, For feature points at the th Pixel coordinates of a frame This represents the L2 norm.
[0026] Furthermore, the sparse-view BEV fusion model DSSA based on the absolute depth map and utilizing dynamic spatial self-attention is transformed into a BEV representation to obtain the final BEV feature map, including: S401: For the first The corrected pixel and absolute depth maps from each camera are used to calculate the corresponding 3D point coordinates based on the camera's intrinsic and extrinsic parameters. :
[0027] in, For the first The intrinsic parameter matrix of each camera. and For the first External parameters of a camera Absolute depth; S402: Convert 3D point coordinates Top-down projection yields BEV mesh elements. Network coordinates and intermediate feature maps By weight Aggregation yields the initial BEV feature representation ,in, and Represents network coordinates in the horizontal and vertical directions; S403: Define the BEV query vector To characterize the state of BEV network cells, the visible field of view is considered. The system determines whether a BEV network unit is visible based on the camera's view frustum, and calculates the position of the BEV network unit on the camera. Attention weights in :
[0028] in, For learnable matrices, For the first The feature vector of the camera mapped to the BEV spatial cell. Represents the set of all visible cameras. For the first Feature vectors of the camera on the BEV grid cell; S404: Define the geometric deviation term, taking into account the projection deviation between different cameras. :
[0029] in, The coordinates of the projection point, For the center of gravity of the grid, These are learnable bias weights; S405: Combines attention weights and geometric bias terms to fuse multi-camera features into a BEV query. :
[0030] S406: Query the BEV of all grids The final BEV feature map is formed.
[0031] The beneficial effects of this application are: This application presents a Transformer-based multi-view perception method for low-altitude UAVs, achieving end-to-end optimization of visual perception tasks for low-altitude aircraft. The VAST module utilizes the Transformer's self-attention mechanism to correct monocular depth under unlabeled conditions, and combines attitude compensation and viewpoint transformation to obtain accurate oblique depth. The TDSF module, based on geometric scale recovery, uses Transformer to fuse visual and navigation information, dynamically estimating the scale factor, alleviating the scale ambiguity problem of monocular depth and achieving real-time performance. The DSSA module uses depth-guided sparse mapping and spatial self-attention fusion to achieve efficient BEV representation under sparse viewpoints, overcoming the shortcomings of traditional LSS pipelines in depth accuracy and computational efficiency. The overall framework design fully utilizes navigation sensors and multi-view information, improving the accuracy and robustness of 3D perception while maintaining end-to-end trainability of the algorithm, laying a solid foundation for autonomous obstacle avoidance and path planning for low-altitude aircraft. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0033] Figure 1 This is a flowchart illustrating a multi-view perception method for low-altitude unmanned aerial vehicles based on Transformer, provided in an embodiment of this application.
[0034] Figure 2 The design framework diagram of the three algorithm models of depth estimation, scale fusion and spatial fusion provided in the embodiments of this application is shown.
[0035] Figure 3 A schematic diagram of rolling compensation provided for an embodiment of this application.
[0036] Figure 4 This is a structural diagram of the TDSF provided in an embodiment of this application. Detailed Implementation
[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0038] Example 1: Related information: Monocular depth estimation can be divided into self-supervised and supervised categories. RSA model uses language description to help solve scale ambiguity; Hybrid Transformer combines Vision Transformer and convolutional neural network to model features; Depth Anything V2 trains a depth estimation model using 595,000 accurately labeled synthetic images and 62 million pseudo-labeled real images, and outperforms the V1 version in detail prediction; DynaDepth integrates image features and IMU dynamic information to solve scale drift. However, these methods lack adaptive correction of body posture errors and are difficult to meet the depth estimation requirements of low-altitude aircraft in complex attitudes.
[0039] Lift-Splat-Shoot (LSS) projects two-dimensional features to three dimensions and then to the BEV plane, but 50% of the depth pixels are lost. SDG-OCC uses semantic and depth guidance to improve BEV occupancy prediction in the LSS framework; BEVFusion uses a unified multi-task multi-sensor fusion architecture to fuse LiDAR and multi-camera features in BEV; SurroundOcc improves BEV representation by 2D-3D pre-occupancy prediction and 3D reconstruction; OccTransformer improves 3D occupancy prediction by introducing Transformer based on BEVFormer (Bird's Eye View Transformer). The above methods have significant effects in autonomous driving, but still have problems of dense calculation and real-time performance in low-altitude unmanned aerial vehicle scenarios.
[0040] Therefore, the embodiment of the present application provides a low-altitude unmanned aerial vehicle multi-view perception method based on Transformer, which can be seen from Figure 1 , the method provides a new algorithm design space from the aspects of depth estimation, scale fusion and spatial fusion, which can be seen from Figure 2 , Figure 1 As shown in FIG. 1, the method provided by the embodiment of the present application is a low-altitude unmanned aerial vehicle multi-view perception method based on Transformer, which includes the following steps: S1: Obtain images captured by multiple cameras.
[0041] S2: Use Depth Anything as a basic monocular depth estimation model to obtain the relative depth map corresponding to each image, and use a view self-supervised depth correction mechanism VAST to correct the relative depth map.
[0042] In an embodiment of the present application, Depth Anything is used as a basic monocular depth estimation model. This model is pre-trained through large-scale real images and synthetic images, and has strong generalization ability across scenes. Compared with the V1 version, Depth Anything V2 significantly improves the accuracy and robustness of depth prediction by using synthetic data instead of real labeled data, expanding the capacity of the teacher model, and training the student model using large-scale pseudo-labeled data. Given the lack of large-scale labeled data in low-altitude flight scenarios, the model is used as the basic network, and further adaptation is made for flight perspectives: Camera setup: two opposite fisheye cameras and one overhead wide-angle camera capture images simultaneously. Assuming the image captured by the first camera is , and the output relative depth map is , where is the pixel coordinate, and represent the horizontal and vertical pixel coordinates.
[0043] Adaptation strategy: for images with large pitch angles and obvious fisheye distortion, simulated data is added during training, and different attitudes, angular velocities, dynamic lighting, etc. are simulated through data augmentation to further improve the adaptability of the model to flight scenarios.
[0044] Although Depth Anything has good generalization, it still produces systematic errors in scenes with large pitch angles and distortion. To this end, the present application proposes a view-adaptive self-supervised depth correction mechanism (View-Adaptive Self-supervised Transformer, VAST). This module uses the self-attention mechanism of Transformer to construct a self-supervised signal based on the photometric consistency between adjacent frames, and realizes dynamic correction of the depth map without additional ground truth depth labeling. Specifically, it includes the following: Feature representation: the intermediate features output by the Depth Anything encoder are denoted as . For pixel position , the features in its neighborhood are taken as key-value pairs. The query vector , key vector and value vector are obtained by linear projection, which are the last but one feature layer (channel number 256) of the Depth Anything encoder mapped by three independent convolutions, with feature dimension , as shown in the formula:
[0045] where, The height of the feature map, The width of the feature map. This represents the number of channels in the intermediate feature map. , and These are the linear projection matrices used in the Transformer to generate the query, key, and value, respectively. for Take one of the surrounding local neighborhood sets. The square window can capture local structure while controlling the amount of computation. The three-dimensional coordinates used to distinguish matching feature points in two frames of images For the first There are two cameras. The features output by the attention layer are mapped to depth increments through a feedforward network (two fully connected layers), and then added to the original depth to obtain the corrected depth.
[0046] Self-attention correction: for each pixel Calculate its attention weights:
[0047] in, Attention weights are used to measure neighboring pixel pairs. Contribution, superscript For transpose, These are the key vectors obtained by linear projection of the neighboring pixels, used to calculate and query the vector. similarity, The normalization factor for the feature dimension. Indicates in neighborhood The other pixels traversed are used to normalize the weights of all neighboring pixels.
[0048] The corrected depth value is obtained through attention convergence:
[0049] in, for Corrected relative depth value This represents the original relative depth.
[0050] Self-supervised loss: To avoid relying on ground truth depth, this study employs a photometric consistency constraint between two consecutive frames and designs photometric reprojection error as the photometric loss function.
[0051] in, For the set of valid pixels, Indicates time the value of the image at , denotes the time instant the function of the image from pixel coordinates to pixel values, is the pose change matrix of the camera provided by the navigation sensor from time instant to , is the transformation function that projects a pixel according to the corrected depth to 3D and projects it to the coordinate system of the next image, denotes the L1 norm.
[0052] Smoothness regularization: To alleviate the noise and ensure the local smoothness of the depth map, a gradient smoothness regularization term is added as the smoothness loss function:
[0053] where and denote the difference operators in the horizontal and vertical directions, respectively.
[0054] Total loss function: The training optimization goal of the VAST module is:
[0055] where is the weight hyperparameter of the smoothness loss, the self-supervised loss adopts the photometric re-projection error and the first-order gradient smoothness regularization term, and the weight ratio is , and the photometric consistency is calculated within a three-frame sliding window to enhance the temporal constraint. This loss encourages the corrected depth to maintain geometric consistency between consecutive frames.
[0056] In an embodiment of the present application, the low-altitude flying vehicle has significant roll and pitch movements, which will cause an angle deviation between the image coordinate system and the actual horizontal plane. Therefore, attitude compensation needs to be performed in the pixel plane, so that subsequent depth and viewing angle calculations are performed in a unified reference system.
[0057] Roll compensation: Let the original pixel coordinates (with the image center as the origin) be , and the roll angle of the flying vehicle be (the positive direction is the clockwise rotation of the body along the longitudinal axis, and the unit is radian), as shown in Figure 3 . The compensated coordinates are transformed by the rotation matrix :
[0058] Viewing angle offset calculation: For the compensated pixel , the horizontal offset angle and the vertical offset angle are defined as:
[0059] where, and denote the compensated pixel coordinates in horizontal and vertical directions, is the focal length of the camera (in pixel, can be converted from intrinsic matrix).
[0060] Real oblique depth calculation: depth output by the network denote the depth component parallel to the image plane, the real point-to-optical center distance The horizontal and vertical offset angles need to be considered at the same time:
[0061] where, is the real tangential depth.
[0062] S3: Based on the corrected relative depth map, a dynamic scale fusion model TDSF driven by a Transformer is used to fuse the visual depth and navigation information through the Transformer, to dynamically estimate an optimal scale factor, and based on the optimal scale factor, an absolute depth map is obtained.
[0063] In an embodiment of the present application, the inherent defect of monocular depth estimation is that the absolute scale cannot be directly obtained, that is, the predicted depth only reflects the relative distance relationship. The prior art points out that in the application of unmanned aerial vehicles, scale ambiguity and scale inconsistency seriously limit the application of depth estimation in tasks such as obstacle avoidance and autonomous landing. In order to obtain the absolute depth, the present application learns from the existing geometric scale recovery idea: using the correspondence relationship of features of two consecutive images and the pose information provided by the navigation sensor, the ratio of the relative displacement to the absolute displacement is calculated as a scale factor, and then the relative depth map is multiplied by the scale factor to obtain the absolute depth map. This method only needs a monocular camera and a common navigation sensor, does not depend on a calibration board or ground constraint, and verifies its robustness to sensor noise on the public dataset Mid-Air.
[0064] In order to make the scale recovery more real-time and robust, on the basis of geometric recovery, a Transformer-driven dynamic scale fusion module (Transformer-driven Dynamic Scale Fusion, TDSF) is designed, which fuses visual depth, image features and navigation information (IMU, GPS) through Transformer, and dynamically estimates the scale factor, such as Figure 4The visual feature, navigation feature after projection are shown as the input of the Transformer, and the output of the Transformer layer is the scale increment.
[0065] Visual feature vector It contains the mean, standard deviation, maximum value, etc. of the corrected depth map (4 dimensions in total) and a convolutional encoded image texture feature (32 dimensions); Navigation feature vector It is composed of the mean and standard deviation of the three-axis acceleration and angular velocity of the IMU (6 dimensions each) and the displacement information provided by the GPS (3 dimensions). The After splicing, a fully connected layer is used to project to 64 dimensions as the input token of the Transformer. A multi-head self-attention with 4 heads is adopted, and the dimension of each head is 16. The history sequence length is 5, which takes into account the recent scale changes and avoids the computational burden brought by too long sequences. The output scale increment is projected into a scalar through a feedforward network, and is added to the geometric scale to obtain the updated scale ; The smoothing factor is updated through a sliding average window (length 10) to reduce the influence of IMU noise.
[0066] Initial scale estimation: First, the known absolute height of the aircraft is used to obtain the initial scale factor. Set the pixel set of the ground area in the overhead camera image as , the corresponding corrected depth is , and the actual height of the aircraft is , then the initial scale factor is:
[0067] where represents the mean.
[0068] This operation overcomes the scale uncertainty of monocular depth by taking the average relative depth to true height ratio as the initial scale.
[0069] Geometric scale correction: Consider two consecutive images and . The absolute displacement vector provided by the navigation sensor is denoted as , and the relative displacement is obtained by matching the feature points in the depth map. Set the matching feature point set in the two images as , and the corrected depths are and . Use the intrinsic matrix and the pose transformation to project the pixel points back to the three-dimensional space, which has:
[0070] where, and represent the pixel coordinates of the same scene point in two consecutive frames at time and is the number of matched feature pairs, and are the corresponding corrected depths, and and are the homogeneous coordinates of the pixel, the superscript represents the inverse matrix of the camera intrinsic matrix and are the corresponding projected pixel points. The relative displacement calculated by depth matching can be obtained by the coordinate difference of two points:
[0071]
[0072] where, and are the rotation and translation from the navigation sensor.
[0073] The final geometric scale factor is obtained by the ratio of the absolute and relative displacements:
[0074] where, is the norm of the vector, is the relative displacement, is the absolute displacement vector provided by the navigation sensor, is the absolute displacement vector provided by the navigation sensor, and the scale factor is smoothed by a sliding window or exponential average to suppress the influence of noise.
[0075] Transformer fusion update: geometric methods provide a rough estimate of the scale, but in dynamic flight scenarios, navigation sensor noise and matching errors can cause the scale factor to be unstable. Drawing on the successful application of visual-inertial fusion Transformers in attitude estimation, the present application introduces a Transformer module for cross-modal fusion and dynamic updating of the scale.
[0076] Feature construction: let the scale estimation factor at time be , and define the visual feature vector composed of depth map statistics (e.g. mean, variance), image texture features and geometric scale factors; navigation feature vector including IMU three-axis acceleration, angular velocity, GPS displacement, etc. After splicing, it is projected to a unified dimension through a linear layer:
[0077] wherein, and are learnable matrices, is the input token of the Transformer, is a real set.
[0078] Cross-modal self-attention: adopt multi-head self-attention mechanism, and perform feature interaction on the input token sequence of the historical frame:
[0079] wherein, is the query matrix, is the key matrix, is the value matrix, obtained by different linear projections of the input token, is the input token of the historical frame, is the output of the multi-head self-attention operation.
[0080] The scale increment predicted by the Transformer output through the feedforward neural network FFN:
[0081] wherein, is the scale factor after fusion of the frame, is the scale increment predicted by the Transformer, is the feedforward neural network.
[0082] Scale update: integrate the geometric scale factor and the increment predicted by the Transformer to update the scale at the next time :
[0083] Training target: take the reprojection error of multiple frames of images as supervision, optimize the parameters of the Transformer, and make the estimated scale factor minimize the projection error of multiple frames of 3D points in each frame:
[0084] wherein, is a projection function, is a loss function, denotes the three-dimensional rotation operator at time is a three-dimensional translation vector at time is the pixel coordinate of the feature point in the th frame, denotes the L2 norm.
[0085] S4: Based on the absolute depth map, the sparse view BEV fusion model DSSA is converted into BEV representation using dynamic spatial self-attention, and the final BEV feature map is obtained.
[0086] In an embodiment of the present application, in multi-camera bird's eye view fusion, the common Lift-Splat-Shoot (LSS) pipeline projects the depth distribution of each pixel to the BEV space by discretization. This method is widely used in lightweight models, but the prior art points out that LSS has two main problems: first, the depth estimation error is large, and the geometric and semantic information cannot be fully utilized; second, the utilization rate of BEV space is low, only about 50% of the grid is effective, resulting in a large amount of redundant calculation. Especially for sparse view aircraft systems, excessive reliance on dense grids can significantly increase the computational overhead and reduce real-time performance.
[0087] Therefore, the present application proposes a dynamic spatial self-attention (DSSA) module, which realizes efficient BEV expression under sparse view through explicit depth-guided feature mapping and self-attention fusion under visible domain constraints.
[0088] 1. Depth-guided 2D-3D mapping Coordinate transformation: for the corrected pixel and the absolute depth in the th camera, the corresponding three-dimensional point coordinates are calculated according to the camera's intrinsic matrix and extrinsic :
[0089] wherein is the intrinsic matrix of the th camera, and is the extrinsic of the th camera, is the absolute depth, which is represented in the aircraft coordinate system. Its overhead projection For locating BEV mesh cells, a BEV mesh with a resolution of 0.5m is selected, where, and Represents the network coordinates in the horizontal and vertical directions; the weights of each pixel are adjusted according to the depth confidence during projection, and the confidence is given by the inverse variance of the deep network output.
[0090] Feature projection: Projecting intermediate feature maps Multiply by the depth-guided weight and project onto the corresponding BEV mesh cell. The weights are determined by depth confidence or semantic information to emphasize reliable spatial structure. For each grid cell, features from different camera projections are summed or averaged, and intermediate feature maps are used. By weight Aggregation yields the initial BEV feature representation. .
[0091] 2. Dynamic Spatial Self-Attention Fusion BEV Query Definition: Define the BEV query vector. Characterizing mesh cells The state. The feature vector of each camera mapped to the BEV space is denoted as... Define a query for each BEV mesh cell. (Dimension 64), Features of each camera projected onto this grid (Dimension 64) is mapped to keys and values through a linear layer.
[0092] Attention weights: considering the field of view (i.e., camera) (The set of meshes observable in BEV) determines whether a mesh is visible based on the camera's view frustum, and calculates the mesh. In the camera Attention weights in :
[0093] in, For learnable matrices, For the first The feature vectors mapped from the camera to the BEV space, Represents the set of all visible cameras. For the first The feature vector of the camera on the BEV grid cell.
[0094] Geometric deviation correction: Considering the projection deviation between different cameras, a geometric deviation term is defined. :
[0095] in, for the projection point coordinates, for the grid center, for the learnable bias weight.
[0096] Geometric bias term The distance between the camera projection point and the grid center is calculated by a layer of MLP (output 1-dimensional weight) as an additive bias of attention weight; the attention mechanism uses a single-head self-attention, and the output dimension is 64. This term is used to adjust the attention weight and suppress the mismatch caused by parallax.
[0097] Fusion update: integrate attention weight and geometric bias to fuse multi-camera features into BEV query:
[0098] where, is the normalized attention weight by softmax, and forms the final BEV feature map .
[0099] The learning process of self-attention is optimized by the loss function of downstream tasks (such as occupancy prediction and target detection).
[0100] Embodiment 2: In order to further explain the design scheme, an experiment is designed to elaborate the design method: 1. Experimental hardware and software configuration Hardware environment: all experiments are carried out on a workstation equipped with Intel Core i9-13900K processor, 32GB RAM and NVIDIA RTX 4090 graphics card, GPU memory 24GB, which can simultaneously train deep network and Transformer model. The sensor experiment uses a low-altitude aerial vehicle equipped with a bidirectional fisheye camera (resolution 1024x1024, 90FOV) and an IMU / GPS module, with an IMU sampling rate of 200Hz and a GPS frequency of 1Hz.
[0101] Software environment: the operating system is Ubuntu 22.04 LTS; the deep learning framework uses PyTorch 2.2; the Transformer layer uses the implementation provided by Hugging Face; tables and charts are drawn using Matplotlib and Seaborn; data generation and experimental code are implemented using Python 3.10.
[0102] 2. Dataset and preprocessing Public synthetic dataset: To verify the depth estimation and spatial perception capabilities of the algorithm in low-altitude flight scenarios, the Mid-Air dataset constructed by the AirSim platform, an aviation information and robot simulation platform, is selected. This dataset contains 54 low-altitude flight trajectories, each providing RGB images, camera internal and external parameters, depth maps, and navigation sensor data. The image resolution is 1024x1024 pixels, and the field of view is 90 degrees. In this experiment, 24 trajectories in the "PLE" environment are selected as the training set, and 5 trajectories are selected as the test set, and only four weather conditions, namely sunny, sunset, foggy, and cloudy, in spring are used.
[0103] Custom AirSim trajectories: To verify the applicability of the model under different speed conditions, the present application designs two additional flight trajectories in the AirSim open-source scene "Landscape Mountains": Constant speed trajectory: the speed is constant at 8 m / s, and a total of 845 frames of images are collected; Variable speed trajectory: the speed varies between 4-12 m / s, and a total of 856 frames of images are collected.
[0104] These two trajectories have the same resolution and field of view as the Mid-Air dataset, and are used to test the adaptability of the dynamic scale fusion module to motion changes.
[0105] Real UAV scenarios: In addition, the present application collects 3 real outdoor low-altitude flight videos, using a fisheye camera and IMU / GPS to record data, for verifying the generalization ability of the algorithm. The real data is processed for distortion removal and synchronization, and the true value depth is estimated according to the GPS altitude and flight log.
[0106] 3. Comparison algorithms To comprehensively evaluate the effectiveness of the proposed method, representative methods in depth estimation and BEV perception in the past three years are selected for comparison.
[0107] Table 1 Comparison algorithms
[0108] Continuation table
[0109] All comparison algorithms use the authors' open-source code and default settings to train and test on the Mid-Air dataset. If the original text supports, the present application also fine-tunes or directly evaluates the model on custom trajectories and real UAV data to ensure fair comparison.
[0110] 4. Evaluation indicators For the depth estimation task, the following indicators are used to measure error and accuracy: Absolute relative error (Abs Rel): .
[0111] Square Relative Error (Sq Rel): .
[0112] Root Mean Square Error (RMSE): .
[0113] Log Root Mean Square Error (RMSE log): .
[0114] Accuracy / / : the proportion of pixels satisfying , where , is the number of valid pixels participating in the evaluation, is the predicted depth at the th valid pixel, is the ground truth depth at the th valid pixel.
[0115] In the BEV occupancy prediction task, the mean Intersection over Union (mIoU), precision (Precision) and recall (Recall) are used to evaluate the accuracy of the spatial occupancy. In addition, the frame rate (FPS) of each method is reported to measure the real-time performance.
[0116] 5. Experimental results and analysis Depth estimation performance comparison: Table 2 gives the depth estimation results of the present application and eight comparison algorithms on the Mid-Air test set. In order to facilitate comparison, all relative depth predictions are converted to absolute scale by the standard "median scaling", and then compared with the absolute depth output of the perspective correction and dynamic scale fusion adopted in the present application. The smaller the number, the lower the error, the larger the number, the higher the accuracy.
[0117] Table 2 Depth estimation results of the present application and eight comparison algorithms on the Mid-Air test set
[0118] From Table 2, it can be seen that the method of the present application is superior to the comparison algorithms in all error indicators, especially the AbsRel and Sq Rel of DynaDepth which is closest in performance are reduced by about 14%-20%, and the accuracy The performance is improved to 0.98. This is due to: ① The VAST module significantly reduces the perspective distortion through self-supervised correction; ② The TDSF module dynamically estimates the scale factor using visual statistics and navigation information to solve the monocular scale ambiguity problem; ③ The attitude compensation and simulation data augmentation enhance the adaptability of the model in the scene with large pitch angle and fisheye distortion.
[0119] BEV occupancy prediction performance: To evaluate the advantages of DSSA in sparse perspective BEV representation, the present application compares it with the LSS pipeline and the deep semantic fusion scheme of SDG-OCC. All methods use the same backbone visual encoder to output the BEV occupancy map (80x80 grid). Table 3 shows the mIoU, Precision, Recall and frame rate on the Mid-Air test set.
[0120] Table 3 mIoU, Precision, Recall and frame rate on the Mid-Air test set
[0121] The results show that the DSSA module improves the mIoU of BEV occupancy prediction by about 9 percentage points after using depth confidence and perspective visibility constraints. At the same time, compared with the LSS method, the method of the present application only slightly reduces the frame rate (22FPS), but significantly improves the accuracy, which is suitable for real-time flight control.
[0122] To analyze the contribution of each sub-module, the present application conducts an ablation experiment: Table 4 Ablation experiment
[0123] It can be seen that removing any module will cause performance degradation, among which removing VAST has the greatest impact, indicating that perspective self-supervision is particularly critical in low-altitude flight scenarios.
[0124] Parameter and setting analysis: Frame interval: Change the interval between the reference frame and the current frame When frame and the Abs Rel result shows that when Abs Rel=0.12, when it increases to 0.14, when it is 0.18, verifying that the smaller the interval, the more sufficient the feature matching, and the more accurate the scale estimation.
[0125] Number of transformer heads and layers: Increasing the number of TDSF heads from 2 to 8 improves Abs Rel by less than 1%, but increases running time by about 15%; Therefore, the research adopts a balanced configuration of 4 heads and 2 layers.
[0126] Scale fusion sliding window: the window length increases from 5 frames to 15 frames, the scale estimation is smoother, but the real-time performance decreases significantly. The final selected length is 10.
[0127] The present application realizes end-to-end optimization of visual perception tasks of low-altitude aircraft by introducing three modules of VAST, TDSF and DSSA. The VAST module uses the self-attention mechanism of Transformer to correct monocular depth under unlabeled conditions, combines with attitude compensation and perspective conversion to obtain accurate oblique depth; the TDSF module dynamically estimates the scale factor by fusing visual and navigation information based on geometric scale recovery, which relieves the scale ambiguity problem of monocular depth and has real-time performance; the DSSA module realizes efficient BEV expression under sparse perspective by depth-guided sparse mapping and spatial self-attention fusion, which overcomes the shortcomings of traditional LSS pipeline in depth accuracy and computational efficiency. The overall framework design makes full use of navigation sensors and multi-view information, improves the accuracy and robustness of three-dimensional perception while keeping the algorithm end-to-end trainable, and lays a solid foundation for autonomous obstacle avoidance and path planning of low-altitude aircraft.
[0128] It should be noted that those skilled in the art will realize that the embodiments described herein are for the purpose of helping the reader to understand the principles of the present application and should be understood as not limiting the scope of protection of the present application to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations without departing from the essence of the present application according to the technical inspiration disclosed in the present application, and these modifications and combinations are still within the scope of protection of the present application.
Claims
1. A multi-view perception method for low-altitude unmanned aerial vehicles based on Transformer, characterized in that, include: S1: Acquire images captured by multiple cameras; S2: Using Depth Anything as the basic monocular depth estimation model, the relative depth map corresponding to each image is obtained, and the view self-supervised depth correction mechanism VAST is used to correct the relative depth map. S3: Based on the corrected relative depth map, the Transformer-driven dynamic scale fusion model TDSF is used to fuse visual depth and navigation information through Transformer, dynamically estimate the optimal scale factor, and obtain the absolute depth map based on the optimal scale factor. S4: Based on the absolute depth map, the sparse view BEV fusion model DSSA with dynamic spatial self-attention is used to convert it into a BEV representation, and the final BEV feature map is obtained.
2. The Transformer-based multi-view perception method for low-altitude UAVs according to claim 1, characterized in that, The method of using the viewpoint self-supervised depth correction mechanism (VAST) to correct the relative depth map includes: A1: Extract intermediate feature maps from the Depth Anything encoder, and use three independent convolutional mappings on the target pixel and its neighboring pixels to obtain the query vector, key vector, and value vector: in, For query vector, For key vectors, For value vectors, , and These are the linear projection matrices used in the Transformer to generate the query, key, and value, respectively. For the Depth Anything encoder in the camera The intermediate feature map, Image pixel coordinates, for The neighborhood set, For neighboring pixels, For the first One camera, and Represents the pixel coordinates in the horizontal and vertical directions; A2: Calculate the similarity between the query vector of the target pixel and the key vector of each pixel in the neighborhood, and obtain the attention weights through normalization: in, Attention weights, superscript For transpose, The key vector is obtained by linear projection of the neighboring pixels. The normalization factor for the feature dimension. Indicates in neighborhood Other pixels traversed in the middle; A3: The corrected relative depth value is obtained based on attention weight aggregation, and the corrected relative depth value is: in, for Corrected relative depth value This represents the original relative depth. A4: Construct the photometric reprojection error as the photometric loss function, and introduce a gradient smoothing regularization term as the smoothing loss function. Smooth and suppress noise on the corrected relative depth values to obtain the total loss function. : in, Let be the photometric loss function. Let slip loss function, The weight hyperparameters are used to smooth the loss. For the set of valid pixels, Indicates time The images in The value at that location, Indicates time An image is a function of pixel coordinates to pixel values. For the camera provided by the navigation sensor at any time arrive pose change matrix, This is a transformation function that backprojects pixels onto 3D and then onto the coordinate system of the next frame image according to the corrected depth. Describing the L1 norm, and These represent the difference operators in the horizontal and vertical directions, respectively.
3. The Transformer-based multi-view perception method for low-altitude UAVs according to claim 2, characterized in that, The method, based on the corrected relative depth map, utilizes the Transformer-driven Dynamic Scale Fusion Model (TDSF) to fuse visual depth and navigation information via Transformer, dynamically estimating the optimal scale factor, and then obtaining the absolute depth map based on the optimal scale factor, including: S301: Obtain the initial scale factor using the known absolute altitude of the aircraft: in, The initial scaling factor. This refers to the actual altitude of the aircraft. It is the set of pixels representing the ground region in a top-down camera image. This is the relative depth value corrected for a top-down camera view. Represents the average value; S302: Calculate the geometric scale factor based on the relative and absolute displacements of the navigation sensors. : in, This is relative displacement. The absolute displacement vector provided for the navigation sensor. Let be the magnitude of the vector; S303: Based on the initial scale factor, the Transformer-driven dynamic scale fusion model TDSF is used to fuse multimodal temporal information, refine and smooth the geometric scale factor, and output the final optimal scale factor. S304: Multiply the relative depth map by the optimal scale factor by pixels to obtain the absolute depth map.
4. The Transformer-based multi-view perception method for low-altitude UAVs according to claim 3, characterized in that, The relative displacement of the navigation sensor includes: B1: Acquire two consecutive frames of images, the corrected depth map, and the pose transformation and absolute displacement vector from the navigation sensor; B2: Assume that the set of matching feature points in two consecutive frames is... The corrected relative depth is and Using the intrinsic parameter matrix and attitude transformation The pixels are back-projected into three-dimensional space to obtain the projected pixels: in, and This indicates that the same scene point matched in two consecutive frames is at time [time value missing]. and pixel coordinates, To match the number of feature pairs, and They are respectively and The corresponding corrected relative depth, and The coordinates are homogeneous pixels, with superscript. Represents the camera intrinsic parameter matrix The inverse matrix, and They are respectively and The corresponding projected pixels; B3: Relative displacement calculated using depth matching based on the projected pixels. It can be obtained from the difference in coordinates between the two points: in, and It is the rotation and translation from the navigation sensor.
5. The Transformer-based multi-view perception method for low-altitude UAVs according to claim 4, characterized in that, The process involves refining and smoothing the geometric scale factor based on the initial scale factor, using the Transformer-driven Dynamic Scale Fusion Model (TDSF) to fuse multimodal temporal information, and outputting the final optimal scale factor. This includes: C1: Constructing visual feature vectors and navigation feature vectors The two are then concatenated and projected onto a unified dimension through a fully connected layer to obtain the input labels for the Transformer: in, and For learnable matrices, Input markers for the Transformer; C2: Employs the Transformer-driven dynamic scale fusion model TDSF's multi-head self-attention mechanism for historical data. Frame input token sequence Perform feature interaction: in, For querying the matrix, The key matrix, For value matrices, For history The first frame in the input token sequence Frame input marker, This is the output of the multi-head self-attention operation; C3: Based on the input labeled sequence after feature interaction, the Transformer outputs the scale increment predicted by the feedforward neural network FFN. : in, For the first Scale factor after frame fusion; C4: Based on geometric scale factor and scale increment Update the scale factor for the next time step. : C5: Based on the scale factor, and using the reprojection error of multiple frames as supervision, a loss function is constructed to optimize the Transformer parameters, thereby obtaining the optimal scale factor. The loss function is: in, For projection function, For loss function, Indicates time Three-dimensional rotation operator, For a moment The three-dimensional translation vector, For feature points at the th Pixel coordinates of a frame This represents the L2 norm.
6. The Transformer-based multi-view perception method for low-altitude UAVs according to claim 5, characterized in that, The absolute depth map-based sparse-view BEV fusion model DSSA, utilizing dynamic spatial self-attention, is used to transform the data into a BEV representation, resulting in the final BEV feature map, which includes: S401: For the first The corrected pixel and absolute depth maps from each camera are used to calculate the corresponding 3D point coordinates based on the camera's intrinsic and extrinsic parameters. : in, For the first The intrinsic parameter matrix of each camera. and For the first External parameters of a camera Absolute depth; S402: Convert 3D point coordinates Top-down projection yields BEV mesh elements. Network coordinates and intermediate feature maps By weight Aggregation yields the initial BEV feature representation ,in, and Represents network coordinates in the horizontal and vertical directions; S403: Define the BEV query vector To characterize the state of BEV network cells, the visible field of view is considered. The system determines whether a BEV network unit is visible based on the camera's view frustum, and calculates the position of the BEV network unit on the camera. Attention weights in : in, For learnable matrices, For the first The feature vector of the camera mapped to the BEV spatial cell. Represents the set of all visible cameras. For the first Feature vectors of the camera on the BEV grid cell; S404: Define the geometric deviation term, taking into account the projection deviation between different cameras. : in, The coordinates of the projection point, For the center of gravity of the grid, These are learnable bias weights; S405: Combines attention weights and geometric bias terms to fuse multi-camera features into a BEV query. : S406: Query the BEV of all grids The final BEV feature map is formed.
Citation Information
Patent Citations
Aircraft detection method in aerial image based on OfficientDet and Transformer
CN113283409A
Identification method and system for key power equipment in digital twin transformer area
CN114842363A
Urban low-altitude non-cooperative unmanned aerial vehicle digital risk assessment method
CN115359371A
Monocular depth estimation method based on CNN-Transform hybrid architecture
CN119478000A
Cited By
Streetscape ground object real-time semantic segmentation and geographic positioning method and system based on camera
CN121558064A