Unmanned aerial vehicle visual positioning method based on multi-level feature pyramid fusion architecture

Through the visual positioning method of the drone with a multi-level feature pyramid fusion architecture, the positioning accuracy problem in the environment of restricted GNSS signal is solved, and the autonomous positioning of the drone with high accuracy and robustness is achieved, which is suitable for task execution in complex environments.

CN120339392APending Publication Date: 2025-07-18CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510442674.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing drone visual positioning technology has reduced positioning accuracy in the environment of GNSS signal restricted, and traditional database image retrieval methods are difficult to meet the real-time requirements, affecting the task execution of drones in complex environments.

Method used

The visual positioning method of drone based on a multi-level feature pyramid fusion architecture is adopted. By introducing an improved attention mechanism and a symmetric pyramid feature fusion module, the multi-scale feature interaction and fusion of drone images and satellite images are realized to generate high-precision positioning results.

Benefits of technology

It improves the positioning accuracy and robustness of the drone in complex environments, and realizes efficient and autonomous positioning in the case of GNSS signal loss. It is suitable for a variety of small onboard computers and supports real-time positioning and task execution of the drone in denial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339392A_ABST
    Figure CN120339392A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle visual positioning method based on a multilevel feature pyramid fusion architecture, and the method comprises the steps: inputting an air-to-ground view angle image shot by an unmanned aerial vehicle and a regional high-resolution satellite image containing the position of the unmanned aerial vehicle; a feature extraction trunk with a self-attention mechanism and a cross attention mechanism is adopted to optimize feature interaction between an unmanned aerial vehicle image and a satellite image. Meanwhile, cross-level feature fusion is carried out by designing a feature fusion module of a symmetrical pyramid structure to generate a feature thermodynamic diagram, and the position of the unmanned aerial vehicle is accurately determined according to heat value distribution in the thermodynamic diagram. According to the method, feature information of different scales is effectively integrated, the calculation amount is reduced, meanwhile, the positioning precision is improved, the method is particularly suitable for the complex environment with GNSS signal loss or interference, high robustness and high efficiency are achieved, meanwhile, the method can be deployed on various small onboard computers, and the method is suitable for popularization and application. The method is widely applicable to autonomous positioning of the unmanned aerial vehicle under denial conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image processing, deep learning technology, and geolocation, and particularly relates to a UAV vision positioning method based on a multi-level feature pyramid fusion architecture. Background Art

[0002] In recent years, UAV technology has developed rapidly, and its applications have expanded from initial military uses to multiple civilian fields such as border patrol, search and rescue, forest fire prevention, precision agricultural mapping, and aerial photography. In these application scenarios, the Global Navigation Satellite System (GNSS) is usually used as the main positioning means. However, in complex scenarios such as urban canyons, forest-covered areas, or tactical environments, GNSS signals are extremely vulnerable to occlusion, interference, or even spoofing, resulting in a decrease in positioning accuracy, affecting the mission execution of UAVs, and even posing safety hazards. To solve this problem, researchers have proposed various auxiliary or alternative positioning technologies, such as Ultra-Wideband (UWB) positioning and Inertial Navigation System (INS). Although UWB technology can achieve high-precision positioning within a short distance, it is limited by the environmental topology and signal propagation range, and its positioning performance will significantly decrease in areas with dense high-rise buildings or severe multipath effects. INS can provide continuous position information without relying on external signals, but due to the existence of cumulative errors, the accuracy will decrease during long-term operation. In contrast, vision-based positioning methods, especially vision algorithms combined with deep learning, can infer accurate geographical location information by matching and analyzing the real-time captured overhead images with remote sensing satellite images, thereby providing more robust UAV positioning capabilities in complex environments.

[0003] With the continuous development of technology, vision-based positioning technology has gradually matured and become an important research direction for UAV autonomous positioning. Currently, the research in the field of vision navigation mainly focuses on database image retrieval. The database image retrieval method relies on a pre-constructed satellite image library that contains satellite images of the areas where UAVs may operate, and each image is attached with known position information. When a UAV flies in a GNSS-constrained environment, the image captured by its downward-facing camera will be feature-matched with the satellite images in the image library, and the optimal matching point will be determined through strategies such as global search or nearest neighbor search to achieve positioning. However, this method has certain limitations: First, satellite images need to be cropped to match the UAV's perspective, and then a database for retrieval is formed. At the same time, feature extraction needs to be performed on the cropped images. Second, the cumbersome retrieval and matching process is difficult to meet the real-time requirements of UAV positioning. Therefore, how to improve the positioning accuracy while optimizing the matching process to reduce its complexity has become an important problem that urgently needs to be solved in vision-based positioning technology. Summary of the Invention

[0004] To solve the problems existing in the background technology, a UAV vision positioning method based on a multi-level feature pyramid fusion architecture is proposed. This method introduces an improved attention mechanism in the pyramid backbone network with conditional position encoding, enabling full interaction between UAV images and satellite images during the feature extraction and encoding stages, thereby enhancing the accuracy of feature mapping. In addition, to more effectively fuse multi-scale features, a symmetric pyramid feature fusion strategy is adopted to further process and optimize the multi-level feature maps extracted at different stages of the backbone network, enhancing the model's adaptability to complex environments and positioning accuracy. Finally, the feature map is used to display the distribution probability of each point in the form of a heatmap to determine the final positioning result.

[0005] To achieve the above objectives, the technical solution of this application includes the following steps:

[0006] S1: Use the perception module of the UAV to capture air-to-ground perspective images, and obtain regional high-resolution satellite images corresponding to the UAV's position from the satellite. The position information of the UAV is marked in the satellite images.

[0007] S2: Construct a dataset for model training through the obtained images.

[0008] S3: Put the dataset into a neural network for training to obtain a neural network model for vision positioning.

[0009] S4: Deploy the trained model to the on-board computing module and process the top-down images of unknown positions captured by the UAV in real time.

[0010] The images will be used as inputs and fed into the trained neural network model for positioning prediction.

[0011] The neural network outputs the predicted UAV position information in real time and transmits it to the navigation and control module of the UAV system to assist flight control and mission execution.

[0012] Among them, in S1, the UAV images are captured vertically to the ground by the pan-tilt camera carried by the UAV, and the captured images record the geographical location information at the time of image capture through the built-in GPS of the UAV. The corresponding satellite images are obtained from Google Satellite Maps, and the satellite image range is centered on the UAV's capture position and covers an area 3-5 times the UAV's field of view angle.

[0013] Among them, in S2, a dataset is constructed using the UAV's down-looking images and high-resolution satellite images of the corresponding area.

[0014] Each set of training data contains a UAV image and a matching satellite image.

[0015] By reading the position information, determine the corresponding position of the UAV image in the satellite image, and crop an appropriate area as a matching sample;

[0016] Finally, perform normalization processing on the images, unify the sizes, and form a complete data set.

[0017] Among them, the specific content of S3 is as follows:

[0018] S3.1: Each data set contains a UAV image and a satellite image. After image enhancement, they are sent to the backbone of the neural network for feature extraction and interaction;

[0019] S3.2: Different stages of the network backbone will output different feature maps. After weighted fusion of the different feature maps, a probability map with high-level abstract features and low-level contour features is output. After processing, the point with the highest probability on the probability map is the UAV position predicted by the neural network.

[0020] Among them, in S4, the on-board computing module needs to pre-store a satellite map of a large area in advance; during flight, the UAV continuously acquires real-time images through a downward-looking gimbal camera and transmits them to the on-board computing module carried on the UAV; the trained neural network model has been deployed on the on-board computing module, which can perform feature extraction and fusion on the UAV image and the pre-stored satellite image in real time, so as to assist the accurate positioning of the UAV.

[0021] When the network of the present invention extracts the features of the UAV image and the satellite image, it uses a common backbone, and interacts and fuses the UAV image features and the satellite image features in the early stage.

[0022] In the solution of the present invention, a gamma image enhancement module is proposed. Before sending the UAV image and the satellite image into the network backbone for feature extraction, the images are processed to make the image quality better, and to avoid the impact of low-quality imaging in bad weather on UAV positioning.

[0023] In the solution of the present invention, a feature fusion module with a symmetric pyramid structure is proposed. It generates three feature maps output by the network backbone in a symmetric manner, and then connects these feature maps with residuals. The special symmetric structure enables better fusion of features at different levels and avoids missing some important features in feature fusion.

[0024] The beneficial effects of the present invention are as follows:

[0025] The present invention proposes a method for high-altitude and low-altitude positioning of UAVs applicable to a denied environment.

[0026] The solution of the present invention can replace the positioning source of the UAV when the traditional GPS fails. By inputting the downward-looking image of the UAV and the satellite image into the model for feature fusion and learning, the longitude and latitude information of the UAV can be calculated in real time. At the same time, the present invention also introduces a symmetric pyramid feature fusion structure. Compared with the traditional feature fusion structure, the symmetric pyramid structure has a stronger ability to capture features at different levels and has stronger robustness in the face of more complex geographical environments and weather conditions.

[0027] The model proposed by the present invention achieves an accuracy of 84.55% on the low-altitude dataset UL14 and 79.87% on the high-altitude dataset Crossview9, both of which are at the SOTA level on their respective datasets. In addition, by performing data augmentation and blurring on the traditional dataset, the robustness and adaptability of the model under occlusion and blur conditions are verified. It is worth mentioning that the model has been successfully deployed on the NVIDIA Jetson TX2 on-board computer, providing strong support for the real-time positioning application of UAVs. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the system framework diagram of the method of the present invention;

[0029] Figure 2 is the flow schematic diagram of the method of the present invention;

[0030] Figure 3 is the schematic diagram of randomly cutting the satellite map containing the UAV position by the method of the present invention;

[0031] Figure 4 is the schematic diagram of the feature extraction backbone of the method of the present invention;

[0032] Figure 5 is the schematic diagram of the process of the backbone cutting the image and sending it into the network in the method of the present invention;

[0033] Figure 6 is the schematic diagram of the symmetric pyramid feature fusion structure of the method of the present invention;

[0034] Figure 7 is the heatmap visualization of the positioning of the method of the present invention on two datasets;

[0035] Figure 8 is the physical diagram of the on-board computer deployed by the method of the present invention;

[0036] In the figure, 1, sensing module; 2, on-board computing module; 3, navigation and control module. DETAILED DESCRIPTION OF THE INVENTION

[0037] To make the objectives, technical solutions, and advantages of the present invention more clear, the following will describe in more detail the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. In the drawings, the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The described embodiments are some, but not all, of the embodiments of the present invention. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation to the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. The following will describe the embodiments of the present invention in detail with reference to the accompanying drawings.

[0038] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the scope of protection of the present invention.

[0039] The following will specifically describe the present invention with reference to the accompanying drawings, and further elaborate on a method for visual positioning of an unmanned aerial vehicle based on a multi-level feature pyramid fusion architecture proposed by the present invention.

[0040] As Figure 1 shown, the present invention is applied to an unmanned aerial vehicle, and a sensing module 1 needs to be carried below the unmanned aerial vehicle, and an on-board computing module 2 and a navigation and control module 3 are carried.

[0041] As Figure 2 shown, the positioning method steps of the present invention are as follows:

[0042] S1: During the process of the unmanned aerial vehicle performing a flight mission, use the high-precision gimbal camera carried to collect images of the target search area at multiple heights and multiple flight routes from a vertically downward perspective. The unmanned aerial vehicle autonomously flies according to a preset flight route and obtains high-resolution ground images at different altitude layers (such as 100m, 200m, 400m) to improve the spatial coverage and image matching accuracy. Each collected ground image is automatically recorded by the unmanned aerial vehicle with its own GPS coordinates (including longitude, latitude, and altitude), and flight attitude information such as heading angle, pitch angle, and roll angle is synchronously recorded to ensure that the image data has accurate spatial position information for subsequent georegistration and analysis.

[0043] Based on the central GPS coordinates of the images captured by the drone, high-resolution satellite imagery of the corresponding location is intercepted on Google Satellite Maps. To ensure the accuracy of the spatial correspondence, the geographical coordinates (latitude and longitude) of the four vertices of the intercepted satellite image are known values provided by GIS software. Through these vertex coordinates, the precise latitude and longitude of each pixel point on the satellite imagery can be deduced using methods such as proportional conversion.

[0044] S2: High-resolution satellite imagery with a spatial resolution of 0.297718 meters per pixel (20-level zoom) is selected. This resolution can clearly capture ground details, enabling high-precision matching between the ground images captured by the drone and the satellite imagery, and providing reliable training data for subsequent deep learning and computer vision tasks.

[0045] During the construction of the dataset, each set of training data includes a downward-looking drone image and a satellite image that spatially corresponds to it. First, the GPS information of the drone image is read, including the latitude and longitude (Lat, Lon) of the shooting center point and the height (H) of the drone during shooting, to ensure that the image data has accurate spatial positioning information. Then, the area corresponding to the drone image is extracted from the large-scale satellite imagery. For this purpose, the latitude and longitude of the satellite imagery boundary of this area are read, including the upper right corner (TE, TN) and the lower left corner (BE, BN), and the position coordinates (X, Y) of the drone in the satellite map are calculated based on these coordinates.

[0046] Subsequently, a local area of 5000×5000 pixels is cropped on the satellite image with this pixel position as the center to ensure the spatial correspondence between the drone image and the satellite image. During the cropping process, the boundary area is checked. If the cropping range exceeds the satellite image boundary, it is filled with the average color to ensure data integrity.

[0047] In terms of image size standardization, the cropped drone images are uniformly adjusted to 512×512 pixels, while the satellite images are uniformly adjusted to 1280×1280 pixels to ensure that all data samples have the same input format. Through this method, a high-quality dataset is successfully constructed, providing reliable data support for the matching, fusion, and intelligent analysis of drone images and satellite images.

[0048] S3: The dataset is put into the neural network for training to obtain a neural network model for visual positioning.

[0049] The specific content of S3 is as follows:

[0050] S3.1: First is data preprocessing. Image enhancement is performed on the original drone images and satellite images to improve the robustness and generalization ability of the model. The drone images first undergo gamma transformation to enhance the adaptability of illumination and contrast, and then are scaled to 96×96 pixels to ensure input consistency. The implementation of gamma transformation specifically refers to the following formula:

[0051]

[0052] where γ is a non-linear mapping parameter used to adjust the brightness response curve of the image, and α is a brightness adjustment factor to optimize the overall illumination condition of the image through linear scaling.

[0053] As Figure 3 shown, the satellite images adopt a random cropping strategy to increase data diversity while ensuring that the drone position is included; subsequently, the images also undergo gamma transformation to optimize visual features and are finally scaled to 256×256 pixels. The two processed images will be used as inputs to the backbone network to learn the spatial correspondence between the drone and satellite images.

[0054] As Figure 4 shown, S3.2: The proposed backbone network based on the multi-level feature pyramid fusion architecture contains three feature extraction stages. In each stage, the spatial scale of the feature map is gradually reduced (halved in width and height) to obtain semantic information at different levels. In each stage, the input data first passes through a Patch Embedding layer to divide the original image into small patches and map them to a high-dimensional feature space, and then enters an Attention Transformer Block layer to extract long-range dependencies using the self-attention mechanism and enhance the expressive ability of the features. After these three feature extraction stages, the network finally generates three feature maps of different scales, which contain local detail information, medium-scale features, and global context information respectively.

[0055] As Figure 5 shown, when cutting the large image in the Patch Embedding layer, it is not simply to evenly cut the large image, but each adjacent patch has partial overlap. The advantage of doing this is to reduce the information discontinuity between patches, make the features smoother, and help the model better understand the changes in local areas; the block unit of the Attention Transformer Block is mainly composed of layer normalization (LayerNorm), attention mechanism, and a multi-layer perceptron (MLP) module.

[0056] To fully integrate these multi-level features, a symmetric pyramid feature fusion module is designed. The three feature maps are input into this module for feature fusion, and finally a feature map containing multi-level information is generated.

[0057] As Figure 6 shown, the upper half of the symmetric pyramid structure is composed of three feature maps output by the three stages of the backbone, with sizes of 1 / 4, 1 / 8, and 1 / 16 of the original image respectively, which are denoted as {C1, C2, C3}. To better improve the model performance, a residual structure is designed to further process the feature maps output by the backbone. First, C3 is processed with a 1×1 convolution to reduce the number of channels, then the output is upsampled to align its dimension with C2 and the two are added together. Then, a 3×3 convolution is used to further extract features from the tensor after channel reduction, upsampling, and weighting, and then it is sent to the next layer. The advantage of this design is that after upsampling, due to the scale transformation, some high-level features are blurred, and the convolution can pay attention to these features again. The processing of C2 and C1 is similar. The processing of the lower half of the symmetric pyramid structure is as follows: First, C3 is respectively downsampled and upsampled twice to obtain C4 and C5 with sizes of 1 / 8 and 1 / 4 of the original Figure 1 image. Compared with C1 and C2, C4 and C5 are directly transformed from C3 and contain more pure abstract semantic information. At the same time, similar to the upper half layer, the idea of residual connection is also introduced, and the result of upsampling and fusion of each layer is passed to the next layer after passing through a 3×3 convolution. Such processing minimizes the impact of information loss caused by upsampling. Finally, the two feature maps obtained from the two paths up to C3 and down to C3 are fused to generate a feature map for prediction.

[0058] To ensure the effective comparison of this feature map with the satellite image, bilinear interpolation is used to restore it to the same size as the input satellite image, and it is defined as an equal feature image. Subsequently, based on normalization processing, the equal feature image is converted into a heat map, where the region with the highest heat value corresponds to the position point of the drone predicted by the model, realizing high-precision spatial positioning;

[0059] As Figure 7 shown, where Figure 7 (a) and Figure 7 (b) are respectively the visualization test results of the proposed model on different datasets. Among them, the first row represents the downward-looking image of the drone, the second row represents the satellite image, and the third row represents the heat map display of the position of the drone in the satellite image. The number in the upper left corner of the heat map represents the distance between the predicted position and the actual position.

[0060] S4 specifically is as follows: Before the drone executes a flight mission, first, satellite maps of a large area are pre-stored in the on-board computer carried by the drone. These satellite maps contain high-resolution images of the target area for the drone to compare with the images obtained in real time during flight. During flight, the drone continuously obtains drone images through a downward-looking gimbal camera and sends the image data captured in real time to the small on-board computer carried by the drone through a wired transmission method. The neural network model trained in S3 is pre-deployed on the on-board computer, which has powerful image processing and real-time computing capabilities. Whenever it receives real-time drone images, the computer immediately processes the images, extracts features and fuses the images using the neural network model. By comparing and fusing the real-time drone images with the pre-stored satellite images, the computer can extract ground features at different scales and angles and accurately calculate the current geographical location information of the drone.

[0061] As Figure 8 (a) shows, NVIDIA Jetson TX2 is equipped with a gimbal camera and a 12V lithium battery. Figure 8 (b) shows another smaller-sized NVIDIA Jetson ORIN NX. Both of these on-board computers have successfully deployed the model.

[0062] Finally, the on-board computer sends the calculated geographical location information to the control system of the drone through a data transmission link to ensure that the drone can continuously obtain its accurate real-time position. This process not only realizes the self-positioning function of the drone during flight but also enables the drone to quickly and accurately perform path planning, obstacle avoidance, and task execution in a complex environment.

[0063] In summary, the drone vision positioning method based on the multi-level feature pyramid fusion architecture provided by the embodiments of this application focuses on solving the drone cross-view geolocation task and improving the positioning accuracy and robustness of the drone in a complex environment. Through multi-stage feature extraction and fusion, this method can effectively capture the multi-scale information of drone images and satellite images and achieve more accurate spatial matching. In terms of data processing, this method uses gamma transformation to enhance the image quality and combines a random cropping strategy to optimize the matching robustness of satellite images. The backbone network performs deep feature extraction through Patch Embedding and Attention Transformer Block and uses a symmetric pyramid feature fusion module to fuse multi-level information, and finally generates an equal feature image with the same size as the satellite map. Subsequently, it is converted into a heat map through standardization to achieve high-precision prediction of the drone's position, and the predicted geographical location information is transmitted to the drone navigation system in real time.

[0064] The above technical solutions only reflect the preferred technical solutions of the technical solutions of the present invention. Some changes that those skilled in the art of the present technology may make to some parts thereof all reflect the principles of the present invention and fall within the protection scope of the present invention.

Claims

1. A method for visual positioning of unmanned aerial vehicles based on a multi-level feature pyramid fusion architecture, characterized in that The method includes the following steps: S1: Use the sensing module of the drone to capture an air-to-ground perspective image, and obtain a regional high-resolution satellite image corresponding to the position of the drone from a satellite. The position information of the drone is marked in the satellite image; S2: Construct a dataset for model training from the obtained images; S3: Put the dataset into a neural network for training to obtain a neural network model for visual positioning; S4: Deploy the trained model to an on-board computing module and process the top-down image of the unknown position captured by the drone in real time; The image will be input into the trained neural network model for positioning prediction; The neural network outputs the predicted drone position information in real time and transmits it to the navigation and control module of the drone system to assist in flight control and mission execution.

2. The drone vision positioning method based on the multi-level feature pyramid fusion architecture according to claim 1, characterized in that: In S1, the drone image is captured vertically to the ground by a gimbal camera carried by the drone, and the captured image records the geographical location information at the time of image capture through the built-in GPS of the drone. The corresponding satellite image is obtained from Google Satellite Maps. The satellite image range is centered on the drone shooting position and covers an area 3-5 times the field of view angle of the drone.

3. The UAV vision positioning method based on the multi-level feature pyramid fusion architecture according to claim 1, characterized in that: In S2, a dataset is constructed using the down-looking image of the drone and the high-resolution satellite image of the corresponding area; Each group of training data contains a drone image and a matching satellite image; By reading the position information, determine the corresponding position of the drone image in the satellite map, and crop an appropriate area as a matching sample; Finally, perform normalization processing on the images, unify the sizes, and form a complete dataset.

4. The drone vision positioning method based on the multi-level feature pyramid fusion architecture according to claim 1, wherein: S3 is specifically as follows: S3.1: Each group of datasets contains a drone image and a satellite image. After image enhancement, they are sent to the backbone of the neural network for feature extraction and interaction; S3.2: Different stages of the network backbone will output different feature maps. The different feature maps are weighted and fused to output a probability map with high-level abstract features and low-level contour features. After processing, the point with the highest probability on the probability map is the drone position predicted by the neural network.

5. The UAV vision positioning method based on the multi-level feature pyramid fusion architecture according to claim 4, wherein: In S4, the on-board computing module needs to pre-store a large-area satellite map in advance; during flight, the drone continuously obtains real-time images through the down-looking gimbal camera and transmits them to the on-board computing module carried on the drone; the trained neural network model has been deployed on the on-board computing module, which can perform feature extraction and fusion on the drone image and the pre-stored satellite map in real time, thereby assisting in the precise positioning of the drone.