Unmanned aerial vehicle position heading regression method and device for cross-view scenes

By constructing a UAV position and heading regression method across perspective scenarios, and using two-dimensional remote sensing maps and three-dimensional view samples, the CVPHR model is adopted to predict the position and heading angle of the UAV. This solves the positioning and navigation problem of UAVs when GNSS signals are lost in extreme environments, and realizes high-precision autonomous navigation and lightweight model design.

CN121297867BActive Publication Date: 2026-03-31ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

When drones cannot receive GNSS signals in extreme environments, existing pure vision positioning and navigation algorithms suffer from problems such as short navigation distance, image-text alignment deviation, large model size, difficulty in deployment, and susceptibility to communication interference. Furthermore, the positioning accuracy of remote sensing maps is not high, and it is impossible to obtain absolute latitude and longitude coordinates and heading angles.

Method used

A method for UAV position and heading regression in cross-view scenarios is constructed. A cross-view dataset is built by using gridded continuous 2D remote sensing maps and 3D UAV view samples. The CVPHR model is used to predict position and heading angle. By utilizing cross-attention mechanism and feature extraction technology, the dependence on sensors is reduced and the positioning accuracy is improved.

Benefits of technology

It achieves accurate autonomous positioning and navigation of UAVs in the absence of GNSS, reduces cumulative errors, improves survivability in extreme environments, and the lightweight model is suitable for UAVs with limited computing power. It also has cross-view feature fusion capabilities and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121297867B_ABST
    Figure CN121297867B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for unmanned aerial vehicle autonomous positioning and navigation based on a cross-view direction regression network, which comprises the following steps: step 1, constructing a cross-view dataset containing direction information for a cross-view scene; step 2, constructing an unmanned aerial vehicle position and heading regression model for the cross-view scene; step 3, training and verifying the unmanned aerial vehicle position and heading regression model for the cross-view scene by using a training set and a verification set; and step 4, performing regression performance test on the unmanned aerial vehicle position and heading regression model for the cross-view scene by using a test set. The application only relies on visual images, can reduce the dependence of unmanned aerial vehicle positioning and navigation on sensors, improve the survival ability in extreme environments, use a cross-view remote sensing map to construct simulation training data and generate absolute direction regression capability, can eliminate the drift problem, and has a natural advantage in the aspect of migration from simulation to a real scene, and the model is light in weight and more suitable for the unmanned aerial vehicle continuous positioning and directional navigation scene with limited computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) navigation technology, and more specifically to a method and apparatus for UAV position and heading regression. Background Technology

[0002] Drones are widely used in security inspections, disaster relief, high-altitude operations, freight delivery, and agricultural monitoring, permeating various industries and becoming closely intertwined with human life. With the support of the future low-altitude economy and artificial intelligence technology, they have even broader development and application prospects. Drones primarily rely on global positioning systems such as GNSS and GPS for positioning and navigation. However, in certain special scenarios, such as disaster relief in mountainous areas, when obstructed by terrain, or when there is signal interference, as well as in some extreme environments, drones may fail to receive stable positioning signals, leading to mission failure and even significant safety hazards.

[0003] To address the aforementioned issues, UAV positioning technology under GNSS signal-free conditions has been extensively studied. This technology aims to acquire one or more types of information through airborne sensors (such as visual odometry, inertial measurement units, cameras, and lidar), and achieve UAV aerial positioning and navigation through methods such as motion estimation, map building, visual feature matching, and multimodal guidance. This ensures that UAVs can autonomously locate and navigate using their own onboard sensors when there are no external positioning and navigation signals, perform obstacle avoidance and landing maneuvers to reduce safety risks, and even continue to autonomously complete their missions.

[0004] Visual positioning is one of the key technologies for UAV autonomous navigation algorithms in complex environments and under GNSS denial conditions. It mainly relies on visual perception of the environment, acquires image information through airborne cameras, and uses various algorithms and techniques based on knowledge from multiple fields such as computer vision, machine learning, and cybernetics to estimate the position, attitude, and motion state of the UAV in three-dimensional space. It is the foundation of the aerial perception capability of embodied intelligent agents.

[0005] Currently, vision-based positioning and navigation algorithms are commonly combined with voice commands or remote sensing images. Visual-language navigation, by fusing the visual perception of UAVs with natural language commands and leveraging multimodal large model technology, provides a new paradigm for autonomous navigation in complex environments. However, it suffers from problems such as short navigation distance, image-text alignment deviation, weak spatial reasoning ability, large model size making it difficult to deploy on the edge, and susceptibility to communication interference.

[0006] Unmanned aerial vehicles (UAVs) rely solely on visual images for autonomous positioning and orientation, resulting in enhanced survivability in extreme environments. However, research in this field is still insufficient, and the potential of pure visual UAV positioning and navigation remains to be further explored. Due to the difficulty in obtaining real-world scene data, many current aerial visual positioning algorithms use 3D virtual environments for simulation, leading to performance deviations when migrating from simulation to real-world scenarios. One type of algorithm uses remote sensing maps for UAV positioning and navigation research, achieving positioning through cross-view image retrieval or matching. However, the positioning accuracy of retrieval and matching is low, and most algorithms cannot simultaneously obtain absolute latitude and longitude coordinates and heading angles. Most algorithms use discrete remote sensing image samples and do not support continuous point-to-point navigation. Some pure visual navigation algorithms achieve precise positioning or navigation, but they use 2D UAV view images, which differ significantly from real-world scenes. Summary of the Invention

[0007] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a method and apparatus for UAV position and heading regression in cross-view scenarios.

[0008] This invention constructs a cross-view pose regression dataset and a 3D view positioning and navigation simulation environment that are closer to real-world scenarios by using a gridded continuous two-dimensional remote sensing map and three-dimensional UAV view samples collected from the remote sensing map. It proposes a UAV position and heading regression method for cross-view scenarios.

[0009] The first aspect of this invention relates to a method for UAV position and heading regression in cross-view scenarios, comprising the following steps:

[0010] Step 1: Collect adjacent environmental tile samples using a 2D remote sensing map and collect 3D UAV view samples using a 3D map to construct a cross-view dataset containing orientation information. Divide the dataset into training set, validation set and test set according to a set ratio.

[0011] Step 2: Construct a UAV position and heading regression model (CVPHR) for cross-view scenarios. The CVPHR model predicts the absolute position and heading angle of the UAV from a set of square-arranged 2D environment tiles and a 3D UAV view of the same area. The structure of the CVPHR model is as follows: starting from the input, there are two parallel branches, namely the environment base map coordinate encoder and the VGG backbone network connected to the CSMG feature extractor. Both branches are connected to the cross attention module and the block similarity guidance module to obtain the fused feature vector. Finally, the position regression module and the orientation regression module output the position and orientation respectively.

[0012] Step 3: Use the training set and validation set to periodically train and validate the CVPHR model, and obtain the final CVPHR model weight file;

[0013] Step 4: Input the test set into the final CVPHR model and test the regression performance of the CVPHR model using different test samples in the test set.

[0014] The steps for constructing the cross-perspective dataset in step 1 are as follows:

[0015] Step 1-1: Download the side length based on the latitude and longitude range of the task area. Square remote sensing map of pixels Simultaneously record the latitude and longitude values ​​of its center pixel. unit pixel distance value ;

[0016] Step 1-2: [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require Divided into grids with side lengths of Small square environment tiles of pixels Any group of spatially adjacent environment tile matrices forms a square environment block. Each environment block has a unique row index and column index number;

[0017] Steps 1-3: For each environmental block in Steps 1-2, establish a system with the block center as the origin and horizontal... Axis and Vertical If the axes are relative to the Cartesian coordinate system, then the line connecting the center points of the outermost environment tile in each environment block forms a square effective interval. All environment blocks have such a square effective interval and can be seamlessly connected.

[0018] Steps 1-4: Define the due east direction on the remote sensing map as... And corresponding In the positive direction of the axis, within the effective coordinate range of each environment block, generate A set of samples with random latitude and longitude coordinates and random heading angles is used to collect data in a three-dimensional remote sensing map. A sample of drone views ;

[0019] Steps 1-5: Following the left-to-right and top-to-bottom process, repeat steps 1-3 and 1-4 for each environmental block until the remote sensing map has been traversed. All environmental blocks are used to ultimately obtain the CVPHR cross-view dataset CVPHR_Dataset.

[0020] In the CVPHR model described in step 2, the environment base map coordinate encoder encodes the relative coordinates of a set of spatially adjacent two-dimensional environment tiles into a high-dimensional vector; the VGG backbone network is used to extract feature maps of environment tiles or UAV views; the CSMG feature extractor is used to extract feature vectors from the feature maps; the cross-attention module is used to extract cross-view correlation features between the UAV view and environment tiles; the block similarity guidance module is used to estimate the similarity between the UAV view and a set of environment tiles and convert it into a weighted position estimation vector; the position regression module and the orientation regression module predict the UAV's position and heading angle from the fused features, respectively.

[0021] Furthermore, the structure of the CVPHR model described in step 2 also includes: the environment base map coordinate encoder is composed of cascaded fully connected layers, each using the ReLU activation function; in the cross-attention module, the UAV view features are processed through fully connected layers to obtain a query vector, and the base map features are processed through fully connected layers and transposed to obtain a key vector. The two are multiplied, scaled, and then processed through softmax to obtain attention weights, which are then multiplied by the value matrix generated by the fully connected layers to output the result. The module employs a cross-attention feature; the block similarity guidance module consists of a cosine similarity calculation unit, a softmax unit, a column broadcasting unit, a matrix element multiplier, and a column vector summation unit connected in sequence, with the matrix element multiplier connected to the relative coordinate matrix input of the base block; the position regression module consists of fully connected layers connected in series, with the last fully connected layer outputting the predicted coordinates; the direction regression module consists of fully connected layers connected in series, with the last fully connected layer outputting the predicted direction.

[0022] Furthermore, the process of constructing the CVPHR model also includes: designing the input and output dimensions, kernel size, stride, and padding size of all convolutional layers, pooling layers, and normalization layers in the VGG backbone network and CSMG feature extractor, as well as the input and output signal dimensions of other sub-modules and the input and output dimensions of fully connected layers, based on the dimensions of the input image and its coordinates of the CVPHR model input image.

[0023] Step 2 specifically includes:

[0024] Step 2-1: For a specific environmental block, collect a 3D view of the drone within that environmental block. And all environment tiles and their relative coordinates as input;

[0025] Step 2-2: Set the number of clusters and the dimension of each cluster feature in the CSMG feature extractor, and then use the VGG backbone network and the CSMG feature extractor to extract the UAV view feature vectors respectively. and environmental tile feature vectors ;

[0026] Step 2-3: Use the environment base map coordinate encoder to encode the relative coordinates of a set of environment map tiles from Step 2-1 to obtain coordinate encoding vectors. The enhanced neighborhood feature vector is obtained by summing the feature vectors with the corresponding environmental tile feature vectors. ;

[0027] Step 2-4: Set the compression ratio and output dimension of the cross-attention module. Input the UAV view feature vector from Step 2-1 and the enhanced neighborhood feature vector obtained in Step 2-3 into the cross-attention module to obtain the cross-feature vector. ;

[0028] Step 2-5: The UAV view feature vector from Step 2-2 and the enhanced neighborhood feature vector from Step 2-3 are input into the block similarity guidance module to obtain a 2D weighted position estimation vector. ;

[0029] Step 2-6: Concatenate the UAV view feature vector from Step 2-2, the cross feature vector from Step 2-4, and the weighted position estimation vector from Step 2-5 to obtain the fused feature vector. ;

[0030] Steps 2-7: Input the fused feature vector from Step 2-6 into the position regression module and the heading regression module respectively to obtain the predicted relative position. and predicted heading angle Then, the absolute latitude and longitude of the UAV are calculated using the parameters recorded in step 1-1, and finally the latitude and longitude position and heading angle of the UAV are obtained.

[0031] Furthermore, in step 2, the VGG backbone network and CSMG feature extractor network included in the core functional module of the CVPHR model are open source models. The VGG backbone network in the CVPHR model adopts a parameter freezing method, and the parameters of the VGG backbone network, the parameters of the CSMG feature extractor network, and other sub-modules are all set according to the designed structure and parameters.

[0032] Step 3 specifically includes:

[0033] Step 3-1: Divide the dataset into training, validation and test sets according to the proportions, and set the hyperparameters of the training model, such as batch size, initial learning rate and number of training rounds;

[0034] Step 3-2: Periodically train and validate the CVPHR model using the training and validation sets. Input the samples into the CVPHR model in batches, and calculate the predicted location of the UAV view according to steps 2-1 to 2-7. and predicted course ;

[0035] Step 3-3: Use the predicted position obtained in step 3-2 and predicted course with position truth value True value of heading angle The position error and heading error are calculated separately. The Smooth L1 Loss function is selected as the loss function to calculate the loss value. The Adam adaptive moment estimator optimizer is used as the gradient descent algorithm to realize backpropagation and iterative training of the CVPHR model.

[0036] Steps 3-4: The learning rate optimization method uses the ReduceLROnPlateau learning rate decay strategy during the plateau period. In each training round, the training set is used for training and the validation set is used for validation. The validation results are used to adjust the variable learning rate. The training is completed within the total training rounds, and the trained CVPHR model weight file is obtained. After the training is completed, the trained model parameters are saved to the weight file.

[0037] Step 4 specifically includes:

[0038] When testing using the CVPHR model, the weight file described in steps 3-4 is loaded into the CVPHR model. Then, samples from the test set are input into the CVPHR model one by one, and the predicted position of the UAV view is calculated according to steps 2-1 to 2-7. and predicted course The prediction results are compared with the true values ​​of position and heading labeled for each sample in the test set, and the regression prediction accuracy of position and heading is calculated.

[0039] This invention proposes an autonomous localization and navigation method for unmanned aerial vehicles (UAVs) based on a cross-view orientation regression network. The method includes: step 1, constructing a cross-view dataset containing orientation information for cross-view scenarios; step 2, constructing a UAV position and heading regression model for cross-view scenarios; step 3, training and validating the UAV position and heading regression model for cross-view scenarios using training and validation sets; and step 4, testing the regression performance of the UAV position and heading regression model for cross-view scenarios using a test set.

[0040] A second aspect of the present invention relates to a UAV position and heading regression device for cross-view scenarios, comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the UAV position and heading regression method for cross-view scenarios of the present invention.

[0041] A third aspect of the invention relates to a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the UAV position and heading regression method for cross-view scenarios of the present invention.

[0042] The working principle of this invention is:

[0043] This invention divides remote sensing maps into two modes: two-dimensional and three-dimensional perspectives. It collects 2D environmental tile samples and 3D drone view samples respectively, constructing a cross-perspective dataset for training and testing regression models. The 3D drone view has significant visual differences from its corresponding 2D image; the former is a drone flight overhead view that closely resembles the real-world scene and texture, while the latter was captured by a near-Earth orbit remote sensing satellite.

[0044] To enable UAVs to perform continuous navigation using remote sensing maps, they should possess the ability to convert visual images into accurate position and heading. This invention designs and implements a cross-view position and heading regression model for UAVs based on deep convolutional neural networks. The position and heading angle of a drone can be predicted from a set of four 2D environmental base maps and a 3D drone view of the same area.

[0045] To address the challenge of autonomous localization and orientation for vision-based UAVs in GNSS-free environments, existing methods often rely on multiple modalities, such as multi-sensor data fusion or voice guidance, for positioning. While these methods can achieve relative positioning of the UAV in GNSS-free conditions, they are highly dependent on multiple sensors or computational resources. This invention relies solely on visual images, reducing the UAV's sensor dependence on positioning and navigation and improving its survivability in extreme environments. Using cross-view remote sensing maps to construct simulation training data and generate absolute orientation regression capabilities can eliminate drift problems to some extent and also has a natural advantage in transferring from simulation to real-world scenarios. The lightweight model design improves inference efficiency and is more suitable for continuous localization, orientation, and navigation scenarios for UAVs with limited computational resources.

[0046] The innovation of this invention is:

[0047] (1) Using pure visual information to regress absolute position: Unlike existing technologies that only use remote sensing map retrieval and matching for coarse positioning, this invention focuses on extracting and utilizing global and local features of the image through cross-attention mechanism, and makes fuller use of pure visual information to obtain accurate absolute position and heading angle information.

[0048] (2) Continuous remote sensing map and cross-view simulation environment: This invention uses a two-dimensional remote sensing map with geographic information to construct a benchmark environment for positioning and orientation, uses a three-dimensional remote sensing map as the UAV positioning and orientation simulation environment, and uses a cross-view dataset to train a regression model. This design makes the algorithm performance closer to the real scene.

[0049] (3) Flexibility and scalability: The technical solution of the present invention can be flexibly applied to various terrains and landforms, and the task granularity can be expanded by increasing the area of ​​the remote sensing map and adjusting the specific parameters of the model according to the task requirements.

[0050] The present invention has the following advantages:

[0051] (1) Absolute positioning and orientation capability of cross-view orientation regression model. The established orientation regression network model can directly and accurately predict the latitude and longitude coordinates and heading angle of UAV, which is more accurate than the coarse positioning achieved by the general remote sensing map retrieval and matching method, and can reduce the cumulative error of pure visual positioning and orientation navigation scheme. This provides important technical support for the continuous positioning and orientation of UAV.

[0052] (2) Regression network structure design with cross-view feature fusion capability. Considering the significant visual errors and feature differences in cross-view images, the CSMG module is first used to extract information containing global and local features from the feature map output by the backbone network, which is more conducive to localization and orientation of cross-domain images, especially improving localization and orientation accuracy when the intersection and union of cross-domain images is small. Secondly, the designed cross-attention module can more efficiently and accurately extract the intrinsic connections between features of cross-view images. Finally, the cross-feature vectors are integrated. Weighted position estimation vector Obtain the fused feature vector This provides sufficient feature information for position and heading regression, further improving cross-view positioning accuracy.

[0053] (3) The dataset and simulation environment are closer to the real scene. Using continuous remote sensing maps as cross-view datasets and 3D remote sensing maps as simulation environment, this dataset and simulation environment are more realistic than general three-dimensional virtual scenes in terms of visual effects and ground texture. Moreover, the use of continuous remote sensing maps as the reference environment can better restore the positioning and orientation process of UAVs in the real environment and reduce the performance deviation when migrating from the simulation environment to the real environment.

[0054] (4) The cross-view orientation regression model has a small number of parameters and the inference time can reach near real-time. Attached Figure Description

[0055] Figure 1 This is a schematic diagram of the steps of the method of the present invention.

[0056] Figure 2 This is a schematic diagram illustrating the method for constructing a cross-view remote sensing map dataset according to the present invention.

[0057] Figure 3 This is a flowchart of the method of the present invention.

[0058] Figure 4 This is a flowchart of the core module of the method of the present invention.

[0059] Figure 5a and Figure 5b This is a screenshot of the test results of the method of the present invention, wherein, Figure 5a To provide a distribution map and statistical results of the distance error between the predicted and actual locations, Figure 5b This document presents a diagram showing the distribution of the angle error between the predicted heading and the actual heading, along with statistical results. Detailed Implementation

[0060] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0061] Example 1

[0062] like Figure 1 , Figure 3 As shown, this embodiment relates to a UAV position and heading regression method for cross-view scenarios, including the following steps:

[0063] Step 1: Collect adjacent environmental tile samples using a 2D remote sensing map and collect 3D UAV view samples using a 3D map to construct a cross-view dataset containing orientation information. Divide the dataset into training set, validation set and test set according to a set ratio.

[0064] Step 2: Construct a UAV position and heading regression model (CVPHR) for cross-view scenarios. The CVPHR model predicts the absolute position and heading angle of the UAV from a set of square-arranged 2D environment tiles and a 3D UAV view of the same area. The structure of the CVPHR model is as follows: starting from the input, there are two parallel branches, namely the environment base map coordinate encoder and the VGG backbone network connected to the CSMG feature extractor. Both branches are connected to the cross attention module and the block similarity guidance module to obtain the fused feature vector. Finally, the position regression module and the orientation regression module output the position and orientation respectively.

[0065] Step 3: Use the training set and validation set to periodically train and validate the CVPHR model, and obtain the final CVPHR model weight file;

[0066] Step 4: Input the test set into the final CVPHR model and test the regression performance of the CVPHR model using different test samples in the test set.

[0067] The steps for constructing the cross-perspective dataset in step 1 are as follows:

[0068] Step 1-1: Determine the latitude and longitude range corresponding to the 2D remote sensing map based on the task area and pixel resolution, and download the map with a side length of [missing information - likely pixels]. Square remote sensing map Simultaneously, the latitude and longitude values ​​corresponding to the center pixel of the remote sensing map are recorded. Distance value corresponding to each pixel ;

[0069] Step 1-2: Divide the remote sensing map downloaded in Step 1-1 into a grid, following the direction consistent with latitude and longitude, using a side length of... The remote sensing map is divided into square grids of pixels, with each small square representing an environmental tile. Remote sensing map Any group of spatially adjacent tiles can be combined to form a square region, which constitutes an environment block. Each environment block is assigned a serial number using a sequence of natural numbers from left to right and from top to bottom, along with row and column index numbers.

[0070] Steps 1-3: For each environmental block described in Steps 1-2 Establish a horizontal axis with the center point of the environmental block as the origin and parallel to latitude and longitude respectively. Axis and longitudinal Using a Cartesian coordinate system with axes as coordinate axes, and taking half the pixel value of the side length of the environment tile as the basic unit 1 of the Cartesian coordinate system, each center point of the environment tile in the environment block has a unique relative coordinate. Connecting the center points of the outermost environment tile with black dashed lines forms a square area, which is the effective range of the environment block. All environment blocks have such a square effective range and can slide according to the corresponding step size and connect seamlessly with each other.

[0071] Steps 1-4: Within the valid coordinate range of each environment block, generate A set of random relative coordinates and random heading angle samples, which are specified here. Corresponding coordinate axes The positive axis corresponds to due east on the remote sensing map. Then, based on this environmental block, the remote sensing map... The absolute latitude and longitude coordinates of each sample are calculated using the latitude and longitude coordinates in the data. Finally, the data is collected in a three-dimensional remote sensing map based on the absolute latitude and longitude coordinates and the heading angle. A sample of drone views ;

[0072] Steps 1-5: A sample 3D UAV view acquired in Steps 1-4 Together with its corresponding relative coordinates, heading angle, and all environmental tiles in the environmental block, they constitute the first environmental tile in that environmental block. Basic information for each sample, and each environmental block is obtained. One such sample;

[0073] Steps 1-6: Following the process from left to right and from top to bottom, repeat steps 1-3, 1-4, and 1-5 for each environmental block until the remote sensing map has been traversed. All environmental blocks are used to ultimately obtain the CVPHR cross-view dataset CVPHR_Dataset.

[0074] In the CVPHR model described in step 2, the environment base map coordinate encoder encodes the relative coordinates of a set of spatially adjacent two-dimensional environment tiles into a high-dimensional vector; the VGG backbone network is used to extract feature maps of environment tiles or UAV views; the CSMG feature extractor is used to extract feature vectors from the feature maps; the cross-attention module is used to extract cross-view correlation features between the UAV view and environment tiles; the block similarity guidance module is used to estimate the similarity between the UAV view and four environment tiles and convert it into a weighted position estimation vector; the position regression module and the orientation regression module predict the UAV's position and heading angle from the fused features, respectively.

[0075] Furthermore, the structure of the CVPHR model described in step 2 also includes: the environment base map coordinate encoder is composed of... The system consists of three cascaded fully connected layers, each using the ReLU activation function. The cross-attention module comprises three fully connected layers, two matrix multipliers, one scaling unit, and one softmax unit. These three fully connected layers are designated as fully connected layer 4, fully connected layer 5, and fully connected layer 6. The cross-attention feature is obtained by multiplying the output of fully connected layer 4 with the output of fully connected layer 5 (plus matrix transpose), performing scaling and softmax processing, and then multiplying it with the output of fully connected layer 6. The block similarity guidance module consists of a cosine similarity calculation unit, a softmax unit, a column broadcasting unit, a matrix element multiplier, and a column vector summation unit connected in sequence. The matrix element multiplier is connected to the relative coordinate matrix input of the base block. The position regression module consists of... The system consists of four fully connected layers connected in series, designated as fully connected layer 7, fully connected layer 8, fully connected layer 9, and fully connected layer 10. Fully connected layers 7, 8, and 9 all use the ReLU activation function, while fully connected layer 10 outputs the predicted coordinates. The orientation regression module comprises... The system consists of four fully connected layers connected in series, referred to as fully connected layer 11, fully connected layer 12, fully connected layer 13 and fully connected layer 14. Fully connected layers 11, 12 and 13 all use the ReLU activation function, and fully connected layer 14 outputs the prediction direction.

[0076] Furthermore, the process of constructing the CVPHR model also includes: designing the input and output dimensions, kernel size, stride, and padding size of all convolutional layers, pooling layers, and normalization layers in the VGG backbone network and CSMG feature extractor, as well as the input and output signal dimensions of other sub-modules and the input and output dimensions of fully connected layers, based on the dimensions of the input image and its coordinates of the CVPHR model input image.

[0077] Step 2 specifically includes:

[0078] Step 2-1: Input including remote sensing map center point latitude and longitude values unit pixel distance value Row and column index numbers of environmental blocks, a 3D drone view. and its corresponding set of environmental tiles; in the Cartesian coordinate system of the environmental region established in steps 1-3, collect the indexes and relative position coordinates of all environmental tiles in the current environmental block, and the three-dimensional view of the UAV. The center coordinates and the heading angle in sine and cosine form;

[0079] Step 2-2: The VGG backbone network in the CVPHR model is frozen. The VGG backbone network and the CSMG feature extractor share weights when processing each environment tile or UAV view. The number of clusters, the feature dimension of each cluster, and the output CSMG feature dimension of the CSMG feature extractor are set. The UAV view and the corresponding set of environment tiles are processed by the VGG backbone network and the CSMG feature extractor respectively to extract their corresponding UAV view feature vectors. and a set of environmental tile feature vectors ;

[0080] Steps 2-3: Use the environment base map coordinate encoder to encode the relative coordinates of this group of environment tiles in the current environment block to obtain coordinate encoding vectors. The enhanced neighborhood feature vector is obtained by summing this set of coordinate encoding vectors with the corresponding environmental tile feature vectors. ;

[0081] Step 2-4: Set the compression ratio of the cross-attention module, determine the output dimension of the cross-attention module, and convert the UAV view feature vector from Step 2-1 into a single value. And the enhanced neighborhood feature vector obtained in steps 2-3 Input the cross-attention module to obtain the cross-feature vector. ;

[0082] Step 2-5: Simultaneously, the UAV view feature vector from Step 2-2 is... And the enhanced neighborhood feature vector in steps 2-3 The input block similarity guidance module yields a 2D weighted position estimation vector. ;

[0083] Step 2-6: Convert the UAV view feature vector from Step 2-2 into... The cross feature vectors in steps 2-4 and the weighted position estimation vector in steps 2-5 The fused feature vector is obtained by performing the connection. ;

[0084] Step 2-7: Combine the fused feature vectors from Step 2-6 The predicted relative positions are obtained by inputting the data into the position regression module and the heading regression module, respectively. and predicted heading angle Using remote sensing maps Center point latitude and longitude values, unit pixel distance values The latitude and longitude of the center point of the environmental block can be calculated from the index of the environmental block, and then the absolute latitude and longitude of the UAV can be calculated.

[0085] Steps 2-8: Combining the input image size and coordinate dimensions of the CVPHR model, as well as the functions of each sub-module of the model, design the input and output dimensions of each sub-module, the size and parameters of the fully connected layer, convolutional layer, pooling layer, and normalization layer inside the sub-module, obtain the feature vectors and fused features of the UAV view and environment tiles, and finally regress the latitude and longitude position and heading angle of the UAV.

[0086] In steps 2-8, the VGG backbone network and CSMG feature extractor network included in the core functional module of the CVPHR model are open source models. The VGG backbone network in the CVPHR model adopts the parameter freezing method. The parameters of the VGG backbone network, the parameters of the CSMG feature extractor network, and other sub-modules are all set according to the designed structure and parameters.

[0087] Step 3 specifically includes:

[0088] Step 3-1: Divide the dataset into training, validation and test sets according to the proportions, and set the hyperparameters of the training model, such as batch size, initial learning rate and number of training rounds;

[0089] Step 3-2: Periodically train and validate the CVPHR model using the training and validation sets. Input the samples into the CVPHR model in batches, and calculate the predicted location of the UAV view according to steps 2-1 to 2-6. and predicted course ;

[0090] Step 3-3: Use the predicted position obtained in step 3-2 and predicted course with position truth value True value of heading angle The position error and heading error are calculated separately. The Smooth L1 Loss function is selected as the loss function to calculate the loss value. The Adam adaptive moment estimator optimizer is used as the gradient descent algorithm to realize backpropagation and iterative training of the CVPHR model.

[0091] Steps 3-4: The learning rate optimization method uses the ReduceLROnPlateau learning rate decay strategy during the plateau period. In each training round, the training set is used for training and the validation set is used for validation. The validation results are used to adjust the variable learning rate. The training is completed within the total training rounds, and the trained CVPHR model weight file is obtained. After the training is completed, the trained model parameters are saved to the weight file.

[0092] Step 4 specifically includes:

[0093] When testing using the CVPHR model, the weight file described in steps 3-4 is loaded into the CVPHR model. Then, samples from the test set are input into the CVPHR model one by one, and the predicted position of the UAV view is calculated according to steps 2-1 to 2-7. and predicted course The prediction results are compared with the true values ​​of position and heading labeled for each sample in the test set, and the regression prediction accuracy of position and heading is calculated.

[0094] Example 2

[0095] like Figure 1 , Figure 3 As shown, this embodiment relates to a UAV position and heading regression method for cross-view scenarios, including the following steps:

[0096] Step 1: Collect adjacent environmental remote sensing tile samples using Google 2D remote sensing maps. In this embodiment, the 2D and 3D remote sensing maps are preferably Google 2D and Google 3D remote sensing maps. Collect 3D UAV view samples using Google 3D maps to construct a cross-view dataset containing orientation information. Divide the dataset into training set, validation set and test set according to a set ratio.

[0097] Step 2: Construct a Cross-View PositionHeading-angle Regression (CVPHR) model for cross-view scenarios. The CVPHR model can predict the absolute position and heading angle of a UAV from a set of matrix-arranged 2D environment tiles and a 3D UAV view of the same area. In this embodiment, the matrix-arranged 2D environment tiles used are 2*2 squares, containing four spatially adjacent 2D environment tiles. The structure of the CVPHR model is as follows: starting from the input, there are two parallel branches, namely the environment base map coordinate encoder and the VGG backbone network connected to the CSMG feature extractor. Both branches are connected to the cross attention module and the block similarity guidance module to obtain the fused feature vector. Finally, the position regression module and the orientation regression module output the position and orientation, respectively.

[0098] The environment base map coordinate encoder encodes the relative coordinates of the four base maps into a high-dimensional vector; the VGG backbone network is used to extract feature maps from the environment base map or the UAV view; the CSMG feature extractor is used to extract feature vectors from the feature maps; the cross-attention module is used to extract cross-view correlation features between the UAV view and the environment base map; the block similarity guidance module is used to estimate the similarity between the UAV view and the four environment base maps and convert it into a weighted position estimation vector; the position regression module and the orientation regression module predict the UAV's position and heading angle from the fused features, respectively.

[0099] Step 3: Use the training set and validation set to periodically train and validate the CVPHR model, and obtain the final CVPHR model weight file;

[0100] Step 4: Input the test set into the final CVPHR model and test the regression performance of the CVPHR model using different test samples in the test set.

[0101] like Figure 2 The steps for constructing the cross-perspective dataset in step 1 are as follows:

[0102] Step 1-1: Determine the latitude and longitude range corresponding to the task area on the Google Remote Sensing Map and download the remote sensing map. Set appropriate resolution parameters to make The side length has 256 pixels. times ( (for natural numbers), it is best to make times ( (For natural numbers), and record at the same time Latitude and longitude values ​​corresponding to the center pixel and the distance value corresponding to each pixel ;

[0103] Steps 1-2: Divide the task area map using a square grid with a side length of 256 pixels, and obtain... A 2D environment tile (Base patch, Bpatch) with a size of 256×256 pixels; remote sensing map In this context, all four adjacent tiles can be combined to form a base block (Bblock). We start with the first block in the top left corner, using... This represents the block's sequence number, assigning row and column index numbers to all blocks. ;

[0104] Steps 1-3: For each block Establish with center The point is the origin, and the horizontal lines at the intersection of the four environmental base maps are... Axis and longitudinal Using a Cartesian coordinate system with axes as the coordinate axes and a unit of 128 pixels, the four adjacent base map indices (top left, bottom left, top right, and bottom right) can be represented as follows: The relative coordinates of the four base map center points in this coordinate system are [-1,1], [-1,-1], [1,1], and [1,-1], respectively. They are connected by a black dashed line to form a... The square area is the block. The effective range of all blocks is a square coordinate system region that is seamlessly connected to each other;

[0105] Steps 1-4: In the block The relative coordinates of the center point of a drone view (256×256 pixels) are randomly generated within the valid coordinate range. and heading angle The heading angle converted to sine and cosine form is ,in This stipulates Corresponding coordinate axes The positive direction of the axis corresponds to due east on the remote sensing map. This indicates the sequence number of the drone view sample collected in this block; then according to exist Latitude and longitude coordinates (can be used) and (Calculated) , Calculate the latitude and longitude coordinates and heading angle of the 3D drone view sample in Google 3D remote sensing map, and collect and save it as a sample image. ;

[0106] Steps 1-5: Data collection Then, its corresponding , ,as well as View, base map in this block The paths of a total of 5 images constitute... The first in The basic information of each sample is collected to complete the collection of a single sample. This indicates the sequence number of the four base maps in the sample;

[0107] Steps 1-6: Following the row-first, column-later process, select the next adjacent block. Repeat steps 1-3, 1-4, and 1-5 until all have been traversed. The final result was the CVPHR cross-perspective dataset CVPHR_Dataset, which collected a total of [number missing] data. One sample, including entity images A 2D environment base map and A 3D drone view.

[0108] like Figure 4 The structure of the CVPHR model described in step 2 further includes: the environment base map coordinate encoder consists of three cascaded fully connected layers, denoted as fully connected layer 1, fully connected layer 2, and fully connected layer 3, each using the ReLU activation function; the cross-attention module consists of three fully connected layers, two matrix multipliers, one scaling unit, and one softmax unit, the three fully connected layers being denoted as fully connected layer 4, fully connected layer 5, and fully connected layer 6, wherein fully connected layer 4 is multiplied by the output of fully connected layer 5 plus matrix transpose, then scaled and softmax processed, and finally multiplied by the output of fully connected layer 6 to obtain the cross-attention feature; the block similarity guidance module consists of sequentially connected cosine similarity calculation units. The system consists of a softmax unit, a column broadcasting unit, a matrix element multiplier, and a column vector summing unit. The matrix element multiplier is connected to the relative coordinate matrix input of the basis block. The position regression module consists of four fully connected layers connected in series, denoted as fully connected layer 7, fully connected layer 8, fully connected layer 9, and fully connected layer 10. Fully connected layers 7, 8, and 9 all use the ReLU activation function, and fully connected layer 10 outputs the predicted coordinates. The direction regression module consists of four fully connected layers connected in series, denoted as fully connected layer 11, fully connected layer 12, fully connected layer 13, and fully connected layer 14. Fully connected layers 11, 12, and 13 all use the ReLU activation function, and fully connected layer 14 outputs the predicted direction.

[0109] The process of constructing the CVPHR model also includes: designing the input and output dimensions, kernel size, stride, and padding size of all convolutional layers, pooling layers, and normalization layers in the VGG backbone network and CSMG feature extractor, as well as the input and output signal dimensions of other sub-modules and the input and output dimensions of fully connected layers, based on the dimensions of the input image and its coordinates of the CVPHR model input image.

[0110] The CVPHR model can predict the absolute position and heading angle of a UAV from a set of four 2D environmental base maps and a 3D UAV view of the same area.

[0111] Step 2 includes the following specific steps:

[0112] Step 2-1: Input including remote sensing map center point latitude and longitude values unit pixel distance value Block index A 3D drone view and its corresponding set of four tiles In the training mode The sample number is indicated by a different subscript symbol than that used in steps 1-5, but the essence is the same. Here, the emphasis is on independent samples that already exist in the dataset, while the subscript used when collecting samples is different. Emphasis on regional division of remote sensing images. Indicates the first The sequence numbers of the four base maps in each sample;

[0113] Step 2-2: In the relative Cartesian coordinate system of the region established in Steps 1-3, the relative position coordinates of the four base maps are [-1,1], [-1,-1], [1,1], and [1,-1], respectively. (UAV view) The true values ​​of the center coordinates and the heading angle in sine and cosine form are respectively and ,in , heading angle ;

[0114] Steps 2-3: The VGG backbone network is in a frozen state and shares weights when processing each environmental base map or UAV view. The clustering number of the CSMG feature extractor is set to a specific value. Each cluster feature dimension takes a value The output CSMG feature dimension Drone view and four environmental base maps D-dimensional feature vectors are extracted from the backbone network respectively. and 4 D-dimensional eigenvectors The latter is a 4×D dimensional matrix;

[0115] Steps 2-4: Encode the coordinates [-1,1], [-1,-1], [1,1], and [1,-1] of the four base maps using an environment base map coordinate encoder to obtain four D-dimensional coordinate encoding vectors. The four D-dimensional enhanced neighborhood feature vectors are obtained by summing these features with the corresponding four base graph feature vectors. Assuming and They represent the first The weight and bias matrices of each fully connected layer, where and Let represent the input and output dimensions of the fully connected layer, respectively, and set... and Let represent the input and output vectors of the fully connected layer, respectively. Then, the calculation process of adding the ReLU function to the fully connected layer in this step can be written as follows: ;

[0116] Steps 2-5: Convert the UAV view feature vector (Equivalent to query Q) and enhanced neighborhood feature vectors (Equivalent to key values ​​K and V) are input into the cross-attention module to obtain... Cross-feature vectors of dimension ,in The compression ratio can be set to 1, 2, or 4 to adjust the output dimension of the cross-attention module; generally, a value of 1, 2, or 4 is chosen. At the same time and 4 base map feature vectors The input block similarity guidance module yields a 2D weighted position estimation vector. Consistent with the assumptions made in steps 2-4, the calculation process for the fully connected layer without the ReLU function in this step can be written as follows: ;

[0117] Steps 2-6: Convert the UAV view feature vector , Cross-feature vectors of dimension and a 2D weighted position estimation vector Get the connection dimensional fusion feature vector ; The predicted positions are obtained by inputting the data into the position regression module and the heading regression module, respectively. and predicted heading angle The calculation of the fully connected layer is consistent with steps 2-4 or 2-5.

[0118] Steps 2-7: Using remote sensing maps center point latitude and longitude values unit pixel distance value and blocks index The latitude and longitude of the block's center point can be calculated. And then according to and Calculate the absolute latitude and longitude of the drone ,according to Calculate the heading angle of the drone ;

[0119] Steps 2-8: Combining the input image size and coordinate dimensions of the CVPHR model, and the functions of each sub-module, reasonably design the input and output dimensions of each sub-module, as well as the size and parameters of the fully connected layers, convolutional layers, pooling layers, and normalization layers within the sub-module. This allows for the acquisition of feature vectors and fused features from the UAV view and environmental base map, ultimately regressing the UAV's latitude and longitude position and heading angle. Specific parameters are shown in Table 1, where bias=True indicates that the fully connected layer contains a bias. The VGG backbone network and CSMG feature extractor network are open-source models. The VGG backbone network in the CVPHR model uses a frozen parameter method. Table 3 shows the default network parameter settings, and Table 3 also shows the CSMG feature extractor network parameters. The inplace setting of the ReLU function indicates that the calculation is performed in-place, without adding extra memory to store the calculation results.

[0120] Table 1

[0121]

[0122] Table 2

[0123]

[0124] Table 3

[0125]

[0126] Step 3 includes:

[0127] Step 3-1: Set the ratio of training set, validation set, and test set to 7:2:1, and set the batch hyperparameters for training the model. Initial learning rate Training rounds ;

[0128] Step 3-2: Periodically train and validate the CVPHR model using the training and validation sets, and distribute the samples according to... Input the data into the CVPHR model in batches, and calculate the predicted location of the UAV view according to steps 2-1 to 2-6. and predicted course ;

[0129] Step 3-3: Use the predicted position obtained in step 3-2 and predicted course with position truth value True value of heading angle The position error and heading error are calculated separately. The Smooth L1 Loss function is selected as the loss function to calculate the loss value. The Adaptive Moment Estimator Optimizer (Adam) is used as the gradient descent algorithm to realize backpropagation and iterative training of the CVPHR model.

[0130] Steps 3-4: The learning rate optimization method uses a plateau-period learning rate decay strategy (ReduceLROnPlateau). Each training epoch uses the training set for training and the validation set for validation. The validation results are used to adjust the variable learning rate. Training can be stopped when the loss function value does not exceed a set range for 10 consecutive epochs, or when training is completed within the total number of training epochs. The trained CVPHR model weight file is then obtained. After training, the trained model parameters are saved to the weight file. .

[0131] Step 4 includes:

[0132] When testing using the CVPHR model, the weight file described in steps 3-4 will be used. Load the data into the CVPHR model, then input each sample from the test set into the CVPHR model one by one, and calculate the predicted location of the UAV view according to steps 2-1 to 2-6. and predicted course The predicted results are compared with the true values ​​of position and heading labeled for each sample in the test set, and the regression prediction accuracy of position and heading is calculated. These samples are unseen samples of the model, which can be used to test the model's generalization ability. Using a remote sensing map with a map resolution of 0.14 meters / pixel, a coverage area of ​​720 meters × 720 meters, and a pixel resolution of 5120 × 5120, the UAV 3D view acquisition altitude is 115 meters, and the average area sampling size K is 50, the dataset is created according to step 1, with a total sample size of 18050, and 1805 test samples (accounting for 10% of the total sample size). The position regression prediction error and heading regression prediction error are 0.1297 and 0.719, respectively. Using steps 2-7, the corresponding absolute position error and heading angle error are calculated to be 3.68 meters and 6.21 degrees, respectively. Figure 5aTo predict the absolute position error distribution map and statistical results, Figure 5b To predict the distribution map and statistical results of heading angle errors. Due to the acquisition of remote sensing data... Figure 2 When the base map sample is used, it inherently contains latitude and longitude information. Therefore, the CVPHR model can obtain the latitude and longitude position and heading angle of the UAV with high accuracy at the same time.

[0133] The test results of the UAV position and heading regression model for cross-view scenarios are attached. Figure 5a and Figure 5b As shown. The model parameters are approximately 130MB, and the single run time is 90.93ms. The time consumption test of the algorithm of this invention is based on the following configuration:

[0134] a) Hardware environment: Intel Core i9-9880H CPU@2.30GHz, NVIDIA Quadro RTX 4000 GPU (8GB VRAM, 2560 CUDA cores), 64GB DDR4-2667 memory, 1TB SCSI HDD;

[0135] b) Software environment: Windows 10 64bit, MSVC (compilation parameter -O3), PyTorch 2.0.1, CUDA Toolkit 11.8;

[0136] c) Test data: Download a 5120×5120 pixel 2D remote sensing map from Google Remote Sensing Maps, collect 18050 256×256 pixel 3D view drone views in the same area to build a dataset, complete model training, and conduct drone cross-view position and heading regression tests.

[0137] d) Time measurement: Using Python time.time() to time the operation, after 5 warm-up runs, the average time for a single orientation regression operation of the UAV was 90.93ms after 70 positioning and orientation regression operations.

[0138] The working principle of this invention is:

[0139] First, this invention collects 2D environmental tile samples and 3D drone view samples from Google remote sensing maps in both 2D and 3D perspectives to construct a dataset for training and testing a regression model, such as... Figure 2 The dataset and regression model design module are shown below. Figure 2 It is evident that the 3D view of the drone has a significant visual difference from its corresponding 2D image. The former is a drone flight view that closely resembles the real-world scene and texture, while the latter was taken by a remote sensing satellite in low Earth orbit. This is the key to this cross-view dataset.

[0140] Secondly, to improve the perception capabilities of UAVs, the global and local structural texture features contained in remote sensing maps are fully utilized to convert visual images into the accurate position and heading of the UAV. This invention designs and implements a cross-viewpoint position and heading regression model for UAVs based on a deep convolutional neural network. ,like Figure 3 As shown, the algorithm can predict the drone's position and heading angle from a set of four 2D environmental base maps and a 3D drone view of the same area. This design, which utilizes an adjacency gridded map to regress the orientation, can better adapt to drone view localization and orientation under different intersection-union ratios.

[0141] Example 3

[0142] This embodiment relates to a UAV position and heading regression device for cross-view scenarios, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the UAV position and heading regression method for cross-view scenarios in Embodiment 1.

[0143] Example 4

[0144] This embodiment relates to a computer-readable storage medium storing a program that, when executed by a processor, implements the UAV position and heading regression method for cross-view scenarios described in Embodiment 1.

[0145] The embodiments described in this specification are merely examples of implementations of the inventive concept. The scope of protection of this invention should not be considered as limited to the specific forms stated in the embodiments. The scope of protection of this invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A UAV position heading regression method for cross-view scenes, characterized in that, Comprising the following steps: Step 1: Collect adjacent environment tile samples using two-dimensional perspective remote sensing maps, and collect three-dimensional unmanned aerial vehicle view samples using three-dimensional perspective maps, construct a cross-perspective dataset containing bearing information, and divide the dataset into a training set, a validation set and a test set according to a set proportion; Step 2: Construct a cross-perspective scene-oriented unmanned aerial vehicle position and heading regression model CVPHR, the CVPHR model predicts the absolute position and heading angle of the unmanned aerial vehicle from a group of two-dimensional environment tiles arranged in the form of a square matrix and a three-dimensional unmanned aerial vehicle view of the same region, the structure of the CVPHR model is: from the input, there are two parallel branches, which are environment base map coordinate encoder and VGG backbone network connected CSMG feature extractor, both branches are connected cross attention module and block similarity guide module and obtain fusion feature vector, finally output position and direction by position regression module and direction regression module respectively; Step 3: periodically train and validate the CVPHR model using the training set and the validation set, and obtain the final CVPHR model weight file; Step 4: input the test set into the final CVPHR model, and test the regression performance of the CVPHR model with different test samples in the test set.

2. The drone position heading regression method for cross-view scene according to claim 1, wherein, The cross-perspective dataset construction step in step 1 is as follows: Step 1-1: Download the square side length of 1 pixel according to the latitude and longitude range of the task area Square remote sensing map of pixels At the same time, record the latitude and longitude values of the center pixel point , the unit pixel distance value ; Steps 1-2: The are divided into a grid of squares with side length Small square environment tiles of pixels where any square array of spatially adjacent environment tiles constitutes a square environment block Each environment block has a unique row and column index number; Step 1-3: For each environmental block in step 1-2, establish a relative Cartesian coordinate system with the center of the block as the origin, the horizontal axis and the vertical axis as the coordinate axes, then the center points of the outermost environmental tiles in each environmental block form a square effective interval, and all environmental blocks have such a square effective interval and can seamlessly connect; ​​ Step 1-4: define the east direction of the remote sensing map as and the corresponding axis positive direction, in the valid coordinate interval of each environmental block, generate a group of samples with random latitude and longitude coordinates and random heading angles, and collect a UAV view sample in the three-dimensional perspective remote sensing map; Step 1-5: Repeat Step 1-3, Step 1-4 for each environmental block following the flow from left to right, top to bottom until all environmental blocks in the remote sensing map are traversed, and finally obtain the CVPHR cross-view dataset CVPHR_Dataset. ​ 3. The drone position heading regression method for cross-view scene according to claim 1, wherein, In the CVPHR model of step 2, the environment base map coordinate encoder encodes the relative coordinates of a group of spatially adjacent two-dimensional environment tiles into a high-dimensional vector; the VGG backbone network is used to extract the feature map of the environment tile or the unmanned aerial vehicle view; the CSMG feature extractor is used to extract the feature vector from the feature map; the cross attention module is used to extract the cross-perspective correlation features of the unmanned aerial vehicle view and the environment block; The block similarity guide module is used to estimate the similarity of the unmanned aerial vehicle view and a group of environment tiles and convert it into a weighted position estimation vector; the position regression module and the direction regression module respectively predict the position and heading angle of the unmanned aerial vehicle from the fusion features.

4. The drone position heading regression method for cross-view scene according to claim 3, wherein, The structure of the CVPHR model described in step 2 further comprises: the environment base map coordinate encoder is composed of full connection layers in series, each full connection layer uses a ReLU activation function; in the cross attention module, the UAV view feature obtains a query vector through a full connection layer, the base map feature obtains a key vector through a full connection layer and transposition, the two are multiplied, scaled, and subjected to softmax to obtain an attention weight, and then multiplied with a value matrix generated by a full connection layer to output cross attention features; the block similarity guiding module is composed of a cosine similarity calculation unit, a softmax unit, a column broadcast unit, a matrix element multiplier, and a column vector summation unit connected in sequence, and the matrix element multiplier is connected with the base block relative coordinate matrix input; the position regression module is composed of full connection layers in series, and the last full connection layer outputs a predicted coordinate; the direction regression module is composed of full connection layers in series, and the last full connection layer outputs a predicted direction.

5. The drone position heading regression method for cross-view scene according to claim 4, wherein, The process of constructing the CVPHR model further comprises: designing the input and output sizes, convolution kernel size, step size, padding size of all convolution layers, pooling layers and normalization layers in the VGG backbone network and CSMG feature extractor according to the dimension size of the CVPHR model input image and its coordinates, and the input and output signal dimension of other submodules, the input and output size of the full connection layer.

6. The drone position heading regression method for cross-view scene according to claim 5, wherein, Step 2 specifically includes: Step 2-1 : For a certain environment tile, collect a three-dimensional perspective drone view of the environment tile and all environment tiles and their relative coordinates as input; Step 2-2: Set the number of clusters of the CSMG feature extractor, the dimension of each cluster feature, and then extract the UAV view feature vector and the environmental tile feature vector respectively using the VGG backbone network and the CSMG feature extractor and environmental tile feature vector ; Step 2-3: encode the relative coordinates of the set of environmental tiles in step 2-1 respectively using the environmental base map coordinate encoder to obtain coordinate encoding vectors , and sum up the corresponding environmental tile feature vectors to obtain an enhanced neighborhood feature vector ; Step 2-4: Set the compression rate and output dimension of the cross attention module, input the UAV view feature vector in step 2-1 and the enhanced neighborhood feature vector obtained in step 2-3 into the cross attention module to obtain a cross feature vector ; Step 2-5: The UAV view feature vector in step 2-2 and the enhanced neighborhood feature vector input block similarity guiding module in step 2-3 are input into the block similarity guiding module to obtain a 2-dimensional weighted position estimation vector ; Step 2-6: Concatenate the UAV view feature vector in step 2-2, the cross feature vector in step 2-4, and the weighted position estimation vector in step 2-5 to obtain a fusion feature vector ; Step 2-7: The fusion feature vector in step 2-6 is input into the position regression module and the heading regression module respectively to obtain the predicted relative position and the predicted heading angle , and the absolute longitude and latitude of the UAV are calculated using the parameters recorded in step 1-1, and finally the longitude and latitude position and the heading angle of the UAV are obtained.

7. The drone position heading regression method for cross-view scene in claim 6, wherein, In step 2, the VGG backbone network and CSMG feature extractor network contained in the core functional module of the CVPHR model are open source models, the VGG backbone network in the CVPHR model adopts the parameter freezing mode, and the VGG backbone network parameters and CSMG feature extractor network parameters and other submodules are set according to the designed structure and parameters.

8. The drone position heading regression method for cross-view scene according to claim 7, wherein, Step 3 specifically includes: Step 3-1: divide the dataset into a training set, a validation set and a test set according to a proportion, and set the training model hyperparameters batch, initial learning rate, and training rounds; Step 3-2: Periodically train and validate the CVPHR model using the training set and validation set, input the samples into the CVPHR model in batches, and calculate the predicted position of the UAV view according to steps 2-1 to 2-7 and the predicted heading ; Step 3-3: using the predicted position obtained in step 3-2 and the predicted heading and the true value of the position and the true value of the heading angle Calculate the position error and the heading error respectively, select the smooth average absolute error loss function Smooth L1 Loss as the loss function to calculate the loss value, use the adaptive matrix estimation optimizer Adam as the gradient descent algorithm, realize the back propagation and iterative training of the CVPHR model; Step 3-4: The learning rate optimization method uses a plateau learning rate decay strategy ReduceLROnPlateau, each training round is trained using the training set and validated using the validation set, the validation result is used for the adjustment of the variable learning rate, the training is completed within the total training rounds, and the trained CVPHR model weight file is obtained, and the trained model parameters are saved to the weight file after the training is completed.

9. The drone position heading regression method for cross-view scene in claim 8, wherein, Step 4 specifically comprises: In the test using the CVPHR model, the weight file in step 3-4 is loaded into the CVPHR model, and then the samples in the test set are input into the CVPHR model one by one, and the predicted positions of the UAV view are calculated according to steps 2-1 to 2-7 and the predicted heading The prediction results are compared with the true values of the positions and headings labeled in each sample in the test set, and the regression prediction accuracy of the positions and headings is calculated.

10. A drone position heading regression apparatus for cross-view scenes, characterized by, The device comprises a memory and one or more processors, the memory stores executable code, and the one or more processors execute the executable code to implement the cross-view scene oriented unmanned aerial vehicle position heading regression method in any one of claims 1-9.

Citation Information

Patent Citations

  • Geometric feature-based cross-view visual positioning method and apparatus, and computer device

    CN116030136A

  • Cross-view-angle scene matching method for unmanned aerial vehicle image and satellite image

    CN116797948A