Unmanned aerial vehicle positioning method and device and unmanned aerial vehicle

By using a multi-source information fusion method that integrates ground images, satellite images, and IMU data, the positioning problem of UAVs in environments with weak or failed GPS signals was solved, achieving high-precision and robust autonomous navigation.

CN120846339APending Publication Date: 2025-10-28XIAN CHENHANG EXCELLENCE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511014835.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Drones struggle to achieve high-precision and robust autonomous positioning and navigation in environments with weak or unavailable GPS signals, limiting their application in complex mission scenarios.

Method used

By fusing real-time ground images captured by UAVs, pre-stored satellite reference images, and IMU data, and utilizing a pre-trained UAV positioning model for multi-source information fusion, including feature extraction, feature matching, and relative positioning, high-precision autonomous positioning of UAVs is achieved.

Benefits of technology

In environments with insufficient or no GPS signal, high-precision positioning of UAVs was achieved, with an average positioning error of less than 20 meters, improving the navigation and operation capabilities of UAVs in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120846339A_ABST
    Figure CN120846339A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of unmanned aerial vehicle positioning, and particularly discloses an unmanned aerial vehicle positioning method and device and an unmanned aerial vehicle, and the method comprises the steps: obtaining a ground image shot by the unmanned aerial vehicle in real time and a pre-stored satellite reference image, and obtaining continuous frame images and IMU data shot by the unmanned aerial vehicle; preprocessing the ground image and the continuous frame image which are shot in real time; inputting the ground image and the continuous frame image which are shot in real time after preprocessing, and a pre-stored satellite reference image and IMU data into a pre-trained unmanned aerial vehicle positioning model to respectively obtain a first positioning result and a second positioning result; and fusing the first positioning result and the second positioning result to obtain a final unmanned aerial vehicle positioning result. According to the invention, the autonomous positioning capability of the unmanned aerial vehicle in a GPS signal failure environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of unmanned aerial vehicle (UAV) positioning technology, specifically relating to a UAV positioning method, device, and UAV. Background Technology

[0002] Traditional UAV positioning methods primarily rely on GPS systems for geolocation. However, in environments with weak or completely unavailable GPS signals, such as urban canyons, forests, and tunnels, UAVs cannot accurately locate and navigate, severely limiting their application capabilities in complex mission scenarios. To overcome this problem, some research has begun exploring assisted positioning methods such as visual inertial odometry (VIO), structured light, and lidar. However, limited by sensor accuracy, computational resources, and environmental adaptability, their actual positioning accuracy and robustness still have significant room for improvement.

[0003] Therefore, there is an urgent need for a UAV positioning technology that integrates multi-source information and possesses high precision and robustness to achieve stable navigation and autonomous flight in complex environments. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this application is to provide a UAV positioning method, device, and UAV. This application aims to improve the high-precision, robust, and continuous autonomous positioning capability of UAVs in environments where GPS signals are unavailable.

[0005] To achieve the above objectives, this application provides the following technical solution: A method for locating a drone, the method comprising: acquiring ground images captured in real time by the drone and pre-stored satellite reference images, and acquiring continuous frame images captured by the drone and IMU data, wherein the IMU data includes acceleration and angular velocity; The real-time captured ground images and continuous frame images are preprocessed; the preprocessed real-time captured ground images and continuous frame images, along with pre-stored satellite reference images and IMU data, are input into a pre-trained UAV positioning model to obtain a first positioning result and a second positioning result, respectively; the first positioning result and the second positioning result are fused to complement the positioning information in the real-time captured ground images, pre-stored satellite reference images, continuous frame images captured by the UAV, and IMU data, to obtain the final UAV positioning result.

[0006] Optionally, preprocessing is performed on the real-time captured ground image and the continuous frame image, including: converting the real-time captured ground image and the continuous frame image to grayscale; enhancing the contrast of the grayscale captured ground image and the continuous frame image; removing distortion from the contrast-enhanced real-time captured ground image and the continuous frame image; and normalizing the distortion-removed real-time captured ground image and the continuous frame image.

[0007] Optionally, the UAV positioning model includes: a feature extraction module, a feature matching module, an absolute positioning module, and a relative positioning module. The feature extraction module extracts multi-scale features from preprocessed real-time captured ground images, continuous frame images, and satellite reference images. The feature matching module performs confidence matching on the multi-scale features. The absolute positioning module estimates the UAV's pose in the geographic coordinate system based on the confidence matching results. The relative positioning module obtains the UAV's continuous trajectory and motion state based on the fusion data of preprocessed continuous frame images and IMU data.

[0008] Optionally, the feature extraction module includes an encoder and a decoder, wherein the encoder is used to perform multi-scale feature extraction on preprocessed real-time captured ground images, continuous frame images, and pre-stored satellite reference images; the decoder is used to decode the multi-scale features and generate high-dimensional feature vectors for feature matching and positioning calculations.

[0009] Optionally, the decoder includes a detection head and a descriptor head, wherein the detection head is used to decode the multi-scale features output by the encoder and detect key feature points; the descriptor head is used to generate high-dimensional feature descriptors for the key feature points detected by the detection head, so as to perform semantic description and unique encoding of the key feature points.

[0010] Optionally, the feature matching module includes: an attention graph neural network, an optimal matching layer, and a filtering layer, wherein the attention graph neural network is used to encode the location and confidence of multi-scale features and their descriptors to obtain feature matching vectors; the optimal matching layer is used to calculate the matching degree score matrix between feature matching vectors; and the filtering layer is used to filter the matching degree score matrix to obtain a set of matching point pairs with high confidence.

[0011] Optionally, the absolute positioning module includes a geometric estimation module and a latitude and longitude pose recovery module, wherein the geometric estimation module is used to estimate the relative pose of the UAV through feature matching and geometric constraints; and the latitude and longitude pose recovery module is used to convert the estimated phase pose into latitude and longitude pose in a geographic coordinate system.

[0012] Optionally, the relative positioning module includes an initial state estimation submodule and a sliding window optimization submodule, wherein the initial state estimation submodule is used to recover the initial pose, scale, velocity and gravity direction of the UAV under GPS time-sensitive conditions; the sliding window optimization submodule is used to continuously update and improve the state estimation accuracy and stability of the UAV through tight coupling optimization of visual and IMU data.

[0013] This application also provides a UAV positioning device, the device comprising: an acquisition module for acquiring real-time ground images and pre-stored satellite reference images captured by the UAV, and acquiring continuous frame images and IMU data captured by the UAV, wherein the IMU data includes acceleration and angular velocity; a preprocessing module for preprocessing the real-time ground images and continuous frame images; a positioning module for inputting the preprocessed real-time ground images and continuous frame images, as well as the pre-stored satellite reference images and IMU data, into a pre-trained UAV positioning model to obtain a first positioning result and a second positioning result respectively; and a fusion module for fusing the first positioning result and the second positioning result to complement the positioning information in the real-time ground images captured by the UAV, the pre-stored satellite reference images, the continuous frame images captured by the UAV, and the IMU data, to obtain a final UAV positioning result.

[0014] This application also provides a drone, the drone including a controller that performs a drone positioning method as described in any of the preceding claims.

[0015] Compared with the prior art, the beneficial effects of this application are as follows: This application constructs a multi-module collaborative system integrating feature extraction, feature matching, absolute positioning, and relative positioning by fusing information from multiple sources such as vision, inertial measurement unit (IMU), and satellite imagery. This system enables high-precision and stable autonomous positioning of unmanned aerial vehicles (UAVs) even in environments with weak or unavailable GPS signals. Experimental results show that this scheme achieves an average positioning error of less than 20 meters in mid-to-high altitude flight missions, exhibiting good real-time performance, robustness, and environmental adaptability, significantly improving the navigation and operational capabilities of UAVs in complex scenarios. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a drone positioning method according to an embodiment of this application; Figure 2 This is a schematic diagram of the neural network structure of a feature extraction module provided in another embodiment of this application; Figure 3 This is a flowchart illustrating the generation of feature points and descriptors in a feature extraction module provided in another embodiment of this application; Figure 4This is a schematic diagram illustrating the actual effect of a feature extraction module provided in another embodiment of this application; Figure 5 This is a schematic diagram of the network structure of a feature matching module provided in another embodiment of this application; Figure 6 This is a schematic diagram illustrating the actual effect of the feature matching module provided in another embodiment of this application; Figure 7 This is a schematic diagram illustrating the actual test results of the relative positioning module provided in another embodiment of this application; Figure 8 This is a schematic diagram of the positioning test trajectory effect of the relative positioning module provided in another embodiment of this application; Figure 9 This is a schematic diagram illustrating the positioning accuracy of a relative positioning module provided in another embodiment of this application. Figure 10 This is a schematic diagram of the structure of a drone positioning device provided in another embodiment of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0018] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0019] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0020] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0021] Figure 1 This is a flowchart illustrating a drone positioning method according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps: S100: Acquires real-time ground images captured by the UAV (single-frame images, referring to images captured at every moment during the UAV's flight) and pre-stored satellite reference images (referring to image data from satellites that are stored in advance, usually obtained through satellite remote sensing technology, reflecting the geographical information of a specific area), as well as continuous frame images captured by the UAV (referring to a sequence of ground images continuously captured by the camera at certain frame intervals during the UAV's flight to obtain visual information about the ground, with frame intervals such as 30ms to 100ms (i.e., 10 to 30 frames per second)) and IMU data (such as acceleration and angular velocity). S200: Preprocess the real-time captured ground images and continuous frame images; S300: Input the pre-processed real-time captured ground images and continuous frame images, as well as the pre-stored satellite reference images and IMU data into the pre-trained UAV positioning model to obtain the first positioning result and the second positioning result, respectively. S400: The first positioning result and the second positioning result are fused to complement the positioning information in the real-time ground images captured by the UAV, the pre-stored satellite reference images, the continuous frame images captured by the UAV, and the IMU data, so as to obtain the final UAV positioning result.

[0022] This embodiment achieves high-precision UAV positioning by combining real-time captured ground images, pre-stored satellite reference images, continuous frame images, and IMU data, employing a multi-source information fusion approach. Real-time captured ground images provide an immediate ground view, facilitating rapid correction of positioning errors; continuous frame images and IMU data enhance the stability and robustness of positioning through temporally continuous image sequences and motion information. After preprocessing, these image data and pre-stored satellite reference images are input into the positioning model to obtain positioning results based on vision and inertial measurement, respectively. By fusing these two positioning results, positioning accuracy and reliability are significantly improved, especially in complex environments with insufficient or failed GPS signals, thereby ensuring precise navigation and autonomous flight of the UAV.

[0023] In another embodiment, step S200 involves preprocessing the real-time captured ground images and consecutive frame images, including the following steps: S201: Perform grayscale processing on the real-time captured ground images and continuous frame images; In this step, the real-time captured ground images and consecutive frame images are converted into grayscale images. A weighted average method (e.g., Gray = 0.299R + 0.587G + 0.114B) can be used to remove redundant color information while preserving the structural and brightness features of the image, which helps to improve the efficiency and stability of subsequent feature point detection.

[0024] S202: Enhance the contrast of the grayscale captured ground image and the continuous frame image; In this step, histogram equalization can be used to adjust the contrast of the grayscale-processed real-time captured ground image and consecutive frame images to enhance image detail and recognizability. Contrast adjustment aims to stretch the pixel intensity distribution in the image, improving detail in dark areas and overall image recognizability. This processing enhances the visual saliency of feature edges and textured regions in the image, helping to improve the uniformity of feature point distribution and matching robustness.

[0025] S203: Dedistort the real-time captured ground image and the consecutive frame images after contrast enhancement; In this step, pre-calibrated camera-intrinsic distortion parameters can be used to correct radial and tangential distortions in the image. By inverse mapping of the distortion model, distorted edges and lines in the image can be restored to the true geometric state of the physical scene, ensuring consistency between image coordinates and actual spatial positions.

[0026] S204: Normalize the real-time captured ground image and the consecutive frame image after distortion correction.

[0027] In this step, the distortion-reduced image is cropped or scaled to a uniform resolution (e.g., 640×480) to meet the input size requirements of subsequent deep learning models. Normalization not only ensures consistent input image size but also helps improve the network model's generalization ability across different image sources.

[0028] In another exemplary embodiment, the UAV positioning model includes a feature extraction module, a feature matching module, an absolute positioning module, and a relative positioning module.

[0029] In this embodiment, the feature extraction module is used to extract multi-scale features from preprocessed real-time captured ground images, continuous frame images, and satellite reference images. For example... Figure 2 As shown, the feature extraction module includes an encoder and a decoder. The encoder is used to extract multi-scale features (e.g., corner points, edges, textures, etc.) from pre-processed real-time captured ground images, continuous frame images, and pre-stored satellite reference images. The decoder is used to decode the multi-scale features extracted by the encoder and generate high-dimensional feature vectors for feature matching and localization calculations. The decoder includes a feature point decoding branch and a descriptor decoding branch. The feature point decoding branch is used to decode the multi-scale features output by the encoder and detect key feature points. Specifically, in the feature point decoding branch, the encoded image is first compressed to a size of (W, H, 1) (W, H, 1). Indicates the height of the image. The width of the image is represented by 1, and the number of channels in the image is 1. For grayscale images, each pixel has only one grayscale value, so the number of channels is 1. Then, the maximum probability distribution of feature points (distinguished key point information extracted from the image) is calculated, and image reconstruction is performed so that the feature points can be restored to the original image size and position. The descriptor branch is used to generate high-dimensional feature descriptors for the key feature points detected by the feature point decoding branch. The size of the feature descriptor is (H, W, D) (H represents the height of the image, W represents the width of the image, and D represents the dimension of the feature descriptor), thereby providing semantic description and unique encoding for the key feature points. In the descriptor extraction branch, after the image is compressed, a bicubic interpolation algorithm is applied to obtain the complete descriptor. Finally, by using L2 normalization, the feature descriptor can be normalized to a unit length. Further, such as Figure 3As shown, the encoder comprises a first CR layer (a two-dimensional convolutional layer with Conv2d and ReLU activation function), a second CR layer, a first max-pooling layer, a third CR layer, a fourth CR layer, a second max-pooling layer, a fifth CR layer, a sixth CR layer, a third max-pooling layer, a seventh CR layer, and an eighth CR layer connected in sequence. The encoding process of the encoder for preprocessed real-time captured ground images and pre-stored satellite reference images is as follows: The input grayscale image is first fed into a module consisting of two consecutive CR layers. This stage mainly extracts low-level local features from the grayscale image, such as edges, corners, and textures, which are the basic information constituting feature points. Next, the feature map size is downsampled for the first time through the first max-pooling layer, typically by a factor of 2, such as compressing the image from 1600×1200 to 800×600. This not only significantly reduces the subsequent computational load but also enhances the model's translation and scale invariance.

[0030] Next, the feature map enters the second set of CR layers, where deeper convolutional kernels extract mid-level structural features from the image, such as local contours, target shapes, and texture combinations. Then, a max-pooling layer further reduces the dimensionality of the feature map to 400×300, gradually focusing on key information regions in the image. Following this, the third set of CR layers further deepens the receptive field, extracting high-level semantic features from the image, such as large-scale regional features like feature boundaries, road network structures, and building layouts. This level of representation is particularly important for cross-scale and cross-angle image matching. A third max-pooling operation further reduces the feature map to 200×150, forming a condensed but semantically rich image representation.

[0031] Finally, through the fourth module consisting of two CR layers, the model further integrates multi-layer features and completes global context modeling through a larger receptive field. This ensures that the final output feature map not only retains the spatial structure information of the image but also possesses good discriminative power. The vector at each position in these output feature maps represents a high-dimensional feature representation of a local region in the input image, which can be used in the subsequent feature matching module for feature point detection and descriptor generation.

[0032] In summary, the encoder processes the original image layer by layer, starting with the extraction of shallow, detailed features. It then gradually abstracts features at different scales through multi-layer convolution and pooling operations, ultimately fusing these multi-scale features to generate a stable and robust high-dimensional feature space. These multi-scale features not only effectively capture local and global information in the image but also exhibit strong adaptability to changes in viewpoint and illumination, thus providing reliable input for subsequent feature matching and localization calculations.

[0033] Please continue to refer to this. Figure 3The decoder comprises a detection head and a descriptor head. The detection head is responsible for decoding the multi-scale features output from the encoder and detecting key feature points, assigning a confidence score to each key feature point to indicate its reliability as a key feature point. The detection head consists of a ninth CR layer, a two-dimensional convolutional layer (Conv2d), and a softmax function connected in sequence. First, the multi-scale features output from the encoder are input into the ninth CR layer to further enhance the feature response and suppress noise. Then, a two-dimensional convolutional layer (Conv2d) is connected to reduce the dimension of the multi-scale features to a single-channel heatmap or a multi-class channel map representing the probability distribution of feature points. Finally, the response values ​​of all candidate points are normalized using the softmax function to obtain the probability distribution of each candidate point as a key feature point. The final output is a location confidence map, where high-response regions correspond to the most representative key points in the image. This process can accurately locate high-information regions such as corners and edge intersections in the image, providing a high-quality set of key points for subsequent matching.

[0034] Furthermore, during the model training phase, in order to supervise the detection head to learn accurate feature point detection capabilities, the model introduces the following detection loss function after outputting the confidence map: (1) in, The cross-entropy loss represents the feature point. , , representing the pixels in the 1 / 8 region of the feature map at height h and width w, respectively. Represents the actual value. This represents a feature vector of length 65, where each value in the vector represents the response value of the feature point at the corresponding pixel.

[0035] The loss function employs cross-entropy loss, comparing the output feature point probability map pixel-by-pixel with the ground truth feature point mask to calculate the error between the predicted and true values. This loss function guides the network to optimize the response intensity of the region of interest, ensuring that the detector can accurately identify discriminative keypoint regions in the image. This loss function operates throughout the entire detection path during backpropagation, including the CR layer and the encoder layer, thus achieving end-to-end feature point detection training.

[0036] The task of the descriptor head is to generate high-dimensional feature vectors for detected feature points to facilitate matching and to provide semantic descriptions and unique encodings for key features. The descriptor head consists of a tenth CR layer, a two-dimensional convolutional layer (Conv2d), and L2 norm (used to normalize vectors to unit length). First, the descriptor head extracts contextual information related to each feature point from the feature map output by the encoder through the tenth CR layer to enhance the discriminative power of the features. Next, the two-dimensional convolutional layer (Conv2d) compresses the feature map into a fixed-dimensional feature representation (e.g., a 128-dimensional or 256-dimensional vector for each feature point). To ensure scale uniformity and numerical stability during the matching process, the final step of the descriptor head applies L2 normalization to normalize each descriptor vector to unit length. This process ensures that descriptors can be stably compared using Euclidean distance or cosine similarity, ultimately outputting a set of high-dimensional feature vectors for each feature point, which are then used for subsequent high-confidence feature point matching and pose estimation.

[0037] Furthermore, in order to train the descriptor head to generate feature vectors with good discriminative power, the model introduces a descriptor matching loss function after outputting the normalized descriptor, as shown below: (2) in, This represents the cross-entropy loss of pixels in a certain region; This is expressed as the cumulative sum of the response values ​​of 65 feature points in the region; Represents the actual value; Represents pixel value; Indicates width as , length Pixel area; This represents a specific pixel within that region; This represents all pixels within the specified area.

[0038] The descriptor matching loss function employs either contrastive loss or triplet loss to ensure that the distances between matching point pairs are as close as possible, while maintaining a sufficiently large gap between non-matching point pairs. In actual training, the model collects real matching image pairs and constructs positive and negative sample pairs. It evaluates the matching quality by calculating the distance between descriptors and uses this information for backpropagation to optimize the entire descriptor generation path. By introducing a descriptor matching loss function, this application enables the model not only to detect feature points but also to learn how to encode their semantic information, ensuring that the descriptors for the same feature point remain consistent across different images, thereby improving the robustness and accuracy of feature matching.

[0039] The descriptor matching loss function is applied to all descriptor units in the first pixel region. and all points in the second pixel region The correspondence between two pixel regions is represented as follows: (3) in, This indicates the correspondence between the (h,w) pixel region and the (h′,w′) pixel region, representing whether the two regions match; Indicates the coordinates of the center point of the corresponding area; Represents the coordinates of a pixel within the corresponding region; Indicates to Perform homography transformation.

[0040] The descriptor loss is defined as: (4) (5) Due to computational constraints, the descriptor loss is calculated on low-resolution Hc and Wc.

[0041] in, The loss function is expressed in its expanded form as follows: , indicating that in two features and The loss function at the point; and These are represented as two pixel regions respectively; and This represents two feature descriptors; S and s are the corresponding relationships between two pixel regions in the above formula. Indicates the weight term. Indicates positive margin. Indicates negative margin. It is represented as the product of two feature descriptors. This indicates the search for the maximum value. The actual performance of the feature extraction module is shown below. Figure 4 As shown.

[0042] The feature matching module is used to perform confidence matching on the multi-scale features output by the feature extraction module. For example... Figure 5As shown, the feature matching module includes an attention graph neural network (GNN), an optimal matching layer, and a filtering layer. The attention graph neural network encodes the location, confidence, and descriptors of multi-scale features to obtain feature matching vectors, and enhances these vectors through self-attention and cross-attention mechanisms. The optimal matching layer calculates the matching score matrix between feature matching vectors based on the inner product, and then uses the Sinkhorn algorithm (iteration T times) to solve for the optimal feature assignment matrix. Specifically, assuming there are two images, A and B, with M and N feature points detected respectively, denoted as... and Each feature point is represented by (p,d), where, For the first i The normalized position and confidence of each feature point For the first i The eigenvectors of each feature point.

[0043] The feature matching module works as follows: First, the feature points and feature vectors of the input network are encoded based on the following formula: (6) in, This represents the initial representation for each feature point. A descriptor representing each feature point. This refers to a Multilayer Perceptron (MLP), used here to increase the dimensionality of low-dimensional features. Represented in encoded form, this allows the perception mechanism to fully consider the appearance and positional similarity of features. This represents the position of each feature point.

[0044] Through encoding, the location of a feature point and its feature vector can be encoded into the same feature vector. This allows the network to consider both feature descriptions and location similarity during matching. Next, the features... Input attention map neural network, there are two types of graphs here, one connecting feature points. i An undirected graph of other feature points within the same image and connecting feature points i An undirected graph of feature points within another image Calculate feature points using Transformer i With Figure Information The information The calculation process is as follows: (7) in, This indicates that all feature points have been aggregated. The result afterwards; and These respectively represent belonging to the neural graph network. Feature points; This represents the summation over all feature points; The value of a feature point is its key. The attention weight information is represented by the Softmax of the similarity between the key values ​​of the queried and retrieved objects.

[0045] The calculation is as follows: (8) in, Represents the key vector; This represents the query vector. The calculation process is as follows:

[0046] (9) In the above formula, all source feature points are located in the image. S superior, , To query a specific feature point on image Q i The mapping representation, and All of these represent feature points from the recalled image. j The mapping, , and All are represented as weight information. , and These are all corresponding balance coefficients. and These are represented as feature points located on image Q and image S, respectively.

[0047] Receive information Further updates to features: (10) in, Indicates a connection. Indicates the first Feature points on layer image A i Corresponding features Indicates the connection of feature points and feature information ,when When the number of points is odd, an undirected graph with feature points is calculated. The information, and when When the number is even, the undirected graph of feature points is calculated. Information; MLP This refers to the perceptron.

[0048] After completing the above calculations, perform an MLP layer on the features of all feature points to obtain the final Score matrix for matching. :

[0049] (11)

[0050] in, and These are represented as the features matched on graphs A and B, respectively. i The corresponding descriptor, This indicates that the feature point belongs to graph A. This indicates that the feature point belongs to graph B. Represents weight information, and The feature points are represented as shown in Figure A and Figure B, respectively. i ; For the dot product operation of vectors, Representing feature points i and j The plane located at the cross product of pixel region A and pixel region B; This represents an image of a certain layer. This represents a constant term.

[0051] If we disregard the case of feature point matching failure, then the resulting assignment matrix... It should satisfy: (12) Among them, the vector to be assigned and Let M and N represent vectors of length M and N respectively, with each element having a value of 1. To obtain the assignment matrix P, the score matrix is ​​calculated. get, This represents the transpose of matrix P. However, due to occlusion or noise in the real environment, there will always be cases where feature points cannot be matched. Therefore, it is necessary to define a transpose matrix in the scoring matrix. Add a row and a column as a fault tolerance mechanism (Dustbin): (13) in, This represents the increase in score. This indicates adding a new column of scores. This indicates adding a new row of scores. This indicates the addition of a row and a column to the score; the area after the newly added row and column remains within the original image area. .

[0052] The expansion vector to be assigned at this time is and And the allocation matrix must satisfy: (14) in, This represents the expected number of Dustbin matches for feature points in pixel region A. This represents the allocation matrix after adding Dustbin. express The transpose of the matrix, This represents the expected number of matches between the feature points in pixel region B and their Dustbin.

[0053] The loss function of the feature matching module It is expressed as follows: (15) Among them, images It is a pixel matrix, represented as , representing the actual value of the matching point. For any pixel within the pixel region, and These represent feature points in images A and B that did not match. Represents the allocation matrix Solving for the logarithm, Indicates the image The cumulative sum of the assignment matrix of all feature points. This represents the cumulative sum of the assignment matrix for the unmatched feature points in image A. This represents the cumulative sum of the assignment matrix for unmatched feature points in image B, where M and N are numerical values ​​representing the lengths of the vectors, respectively.

[0054] The filtering layer is used to filter the matching score matrix output by the optimal matching layer and retain point pairs with matching scores higher than a threshold (e.g., 0.8). The feature matching module ultimately outputs a set of high-confidence matching point pairs, with the actual effect as follows: Figure 6 As shown.

[0055] The absolute positioning module includes a geometric estimation module and a latitude and longitude pose recovery module. The geometric estimation module is used to estimate the relative pose of the UAV through feature matching and geometric constraints. The latitude and longitude pose recovery module is used to convert the estimated phase pose into latitude and longitude pose in the geographic coordinate system.

[0056] After the aforementioned matching is completed, the geometric estimation module estimates the geometric relationship between the images using feature points between them, specifically including: The homography matrix H is calculated for the feature points with high confidence. The homography matrix H is defined as follows: (16) in, , , , , , , , , There are 8 unknowns.

[0057] Then the pixels in normalized pixel region 1 and pixels in normalized pixel region 2 The relationship between them can be represented as: (17) To solve for the homography matrix H, we need to resolve the above 8 unknowns, requiring at least four sets of matching points. The process is as follows: (18) in, , , and Representing 4 sets of matching points, the homography matrix H can be finally solved to realize the perspective transformation of the two image components.

[0058] The latitude and longitude pose recovery module utilizes the output of the geometric estimation module to convert the estimated pose information in the image into latitude and longitude pose in the geographic coordinate system. Specifically, this includes: first, performing a point set transformation, where the height of the image captured by the camera is h, the width is w, and the normalized pixel coordinates of the four vertices are... , , and Then a point set can be formed. .

[0059] A new set of points is obtained by performing homography transformation H on the point set. : (19) The above transformations can map points in an image to a new coordinate system.

[0060] Secondly, the image coordinate system is converted to the geographic coordinate system.

[0061] Get point set Then, the geometric center of the fitted point set is... ,in, This represents the coordinates of the center pixel in the pixel coordinate system.

[0062] The top left corner of the known base satellite tile map Coordinates are The coordinates of the lower right corner are Then we can get: (20) (twenty one) and That is, the drone obtained in the actual geographic coordinate system Absolute positioning coordinates, where Table 1 shows the positioning accuracy of the absolute positioning module: Table 1

[0063] In summary, the geometric estimation module is responsible for estimating the geometric relationships between images by matching image feature points, generating the homography matrix H, and realizing coordinate transformation between images. The latitude and longitude pose recovery module uses the transformation information provided by the geometric estimation module to convert points in the image coordinate system to latitude and longitude coordinates in the geographic coordinate system, ultimately recovering the UAV's position in the actual environment.

[0064] In another exemplary embodiment, the relative positioning module includes an initial state estimation submodule and a sliding window optimization submodule. The initial state estimation submodule is used to recover the initial pose, scale, velocity, and gravity direction of the UAV in a GPS-deficient environment, providing an accurate initial state for VIO. The sliding window optimization submodule is used to continuously update and improve the accuracy and stability of the UAV's state estimation through tight coupling optimization of visual and IMU data. Specifically, the initial state estimation submodule includes an SfM initialization unit and an IMU (Inertial Measurement Unit) and vision fusion unit, aiming to recover the initial pose, scale, velocity, and gravity direction of the UAV in a GPS-denied environment, providing a stable initial state for the subsequent visual inertial odometry (VIO) system. The SfM initialization unit performs preliminary pose estimation on each frame of the image within the sliding window using Structure from Motion (SfM) and aligns the IMU and visual information at the scale. During initialization, if specific conditions are met (such as finding stable feature tracking (e.g., more than 30 tracked features) and sufficient disparity (e.g., more than 20 rotation-compensated pixels)), the five-point method is used to recover the relative rotation and translation pose between two frames, and the co-view feature points observed in the two frames are triangulated to obtain the 3D point cloud structure. Based on the triangulated point cloud structure, the PnP method is used to combine the pose relationship between the 3D point cloud and the 2D points in the image to estimate the initial pose of the remaining image frames in the sliding window. Finally, to improve the estimation accuracy, global BA (Bundle Adjustment) is further used to minimize the reprojection error in all image frames, thereby forming a consistent, scale-free camera trajectory and 3D scene structure.

[0065] Based on this, the IMU and vision fusion unit calculates the pose increment between visual keyframes using IMU pre-integration technology, including position increment, velocity increment, and rotation increment. Furthermore, the IMU integration results are scale-aligned with the visual trajectory output by the SfM, and by minimizing the position and orientation residuals between the two, the true physical scale factor, gravity direction, initial velocity, and the IMU acceleration and gyroscope bias are jointly estimated.

[0066] This fusion process employs a nonlinear optimization method to transform the visual estimation from a scale-free structure into a world coordinate system representation with absolute physical meaning. Ultimately, this enables the initialization module to output a precise set of initial state parameters, including aligned pose, velocity, gravity vector, and sensor bias. This provides accurate prior information for the sliding window VIO optimization module, ensuring stable autonomous navigation and positioning even in GPS-denied environments.

[0067] During the sliding window optimization phase, the relative positioning module continuously optimizes and updates the UAV state using a tightly coupled monocular visual inertial odometry system (VINS). The sliding window optimization submodule includes a feature management unit, a pose estimation unit, a tightly coupled monocular VIO unit, and an edge detection and loop closure detection unit. First, the feature management unit adds feature point pairs with high confidence extracted and matched by the feature extraction and matching modules to the global feature manager of the VINS to establish inter-frame visual constraints. As new image frames are continuously input, once the number of camera frames within the sliding window reaches the set maximum window capacity, the pose estimation unit performs nonlinear optimization to update the current system state. This includes: predicting the camera frame pose based on IMU pre-integration, i.e., estimating the initial pose of the current frame using the IMU pre-integration results (position, velocity, rotation) between two adjacent frames. The pre-integration residual is defined as follows: (twenty two) in, This represents the pre-integration error of the IMU. express k Time's up k The observation at time +1, Represents state variables. and Indicates from k The state at time k changes to the state at time k+1. express k Time's up k Position error at time +1 express k Time's up k The velocity error at time +1 Indicates attitude error. The bias error represents the acceleration. This indicates the gyroscope bias error.

[0068] Next, feature point triangulation is performed. This involves using the estimated pose of the current frame to spatially triangulate the feature points observed across multiple frames within the sliding window, thus recovering their three-dimensional coordinates. The triangulation process is as follows: Assuming 3D feature points Observed by M camera frames, it is known from The pose of each camera frame is Camera internal parameters are The observation of feature points in each camera frame is Based on the projection relationship, 3D coordinates can be represented as: (twenty three) in, , Indicates the index of the frame. This represents the rotation of each camera frame. This indicates the transition from the world coordinate system to the camera coordinate system. in This indicates that the transformation includes rotation and translation. Using the observed values ​​of the feature points, and through this projection relationship, the 3D feature points can ultimately be obtained. The value of .

[0069] The reprojection error is calculated and expressed as follows: (twenty four) The above formula is the least squares error function, which transforms the coordinates of a 3D point from the world coordinate system to the corresponding camera coordinate system. Here, it is written in least squares notation. Wherein, This represents the least squares error function. Represents the residual. express i The cumulative summation from 1 to n gives the covariance matrix of the errors. For the observation point, For true scale, These are the 3D points obtained after triangulation. This is the transformation matrix from the world coordinate system to the camera coordinate system. This is the intrinsic parameter matrix of the camera.

[0070] Finally, a sliding window is dynamically updated. The system deletes the oldest frame and its observation information, while inserting the current frame as a new keyframe into the sliding window to update the optimization variables. After nonlinear optimization, the latest UAV pose, velocity, and bias estimation results are obtained.

[0071] After initialization, the system officially enters the tightly coupled monocular VIO phase. In this phase, the sliding window continuously updates the current frame state by jointly optimizing the IMU pre-integration residual and the visual reprojection error. The optimization objective function is a maximum a posteriori estimation problem, in the following form: (25) Part One This represents prior information indicating marginalization. To represent a certain feature point, Indicates the state of the feature point. This represents the latest residual after marginalization; Indicates the marginalization information of the previous frame; Part Two This represents the IMU measurement residual. Represents a specific pixel region; This represents the cumulative summation of least squares errors; Indicates after the IMU (Internet Protocol version) k Frame and the k The residual between the observed and estimated values ​​obtained from +1 frame; and Indicates from k The state changes at time 1 to 2 k The state at time +1; Indicates that the IMU frame is from the first k Frame and the k +1 frame begins; Part 3 This refers to the reprojection residual of the camera; Represents a camera frame sequence, camera frame and All belong to camera frame sequences. Table 1 Frame camera frame; express Camera reprojection error of frames; Indicates the The observation point of the frame; This indicates that the reprojection error here is... Frames and camera frames It is obtained through calculation.

[0072] The positioning output of the relative positioning module during actual flight is as follows: Figure 7 As shown, Figure 7 The image shows the sparse mapping of ground features by the relative positioning module at high altitude and the flight trajectory in the local coordinate system.

[0073] In another exemplary embodiment, step S400, fusing the first positioning result and the second positioning result, includes the following steps: S401: Obtain incremental multi-source positioning information; In this step, after processing each frame of image, the system first extracts the pose change information between the current frame and the previous frame from multiple subsystems in the front end. Specifically, it obtains the latitude and longitude changes between the current frame and the previous frame from the absolute positioning module and converts them into Geo incremental displacement in the world coordinate system. Obtain the relative pose increments provided by the visual odometry from the relative positioning module (VIO). Simultaneously, the motion increment calculated by the sensor through integration within this time interval is also obtained from the IMU pre-integration module. These three sets of incremental information constitute the multi-source observation input required for back-end filtering.

[0074] S402: Establish a state prediction model using an IMU; In this step, the system uses an Extended State Kalman Filter (ESKF) to fuse the three types of incremental information to output a stable, high-frequency positioning result. Under normal conditions, only the output of the non-degraded VIO module (i.e., the visual inertial odometry increment) is used as the input source for fusion.

[0075] When the VIO module degrades, the system will automatically switch to using only the IMU pre-integral increment and absolute positioning increment as the fusion input, and adjust the residual equation accordingly to ensure the stable operation of the filtering system and data continuity.

[0076] S403: Calculate the observation residuals and build an updated model; In this step, to achieve the filter update, the system first models the observations. Suppose a sensor generates observations of the system's state variables, and its observation model is a nonlinear function: (26) in, Indicates the actual observed value; The nonlinear observation function is represented by x; the state variable is represented by x. Indicates observation noise. This indicates that the noise conforms to a standard normal distribution.

[0077] In a Kalman filter, the goal is to update the error state, therefore it is necessary to calculate the Jacobian matrix H of the observation function with respect to the error state, i.e.: (27) in, Represents the observation matrix State variables in Increment Solve for the partial derivatives.

[0078] S404: Calculate Kalman gain and error state update; In this step, based on the observation residuals and the linearized Jacobian matrix, the Kalman gain is calculated systematically. : (28) in, This indicates the prediction result for the previous frame. Representing the Jacobian matrix transpose, Indicates noise.

[0079] Then, the error status is updated: (29) in, Indicates Kalman gain, Let z represent the error, and z represent the observed value. This indicates a speculative state.

[0080] Next, the nominal state is corrected: (30) in, This indicates the state result of the previous frame.

[0081] Finally, update the covariance matrix: (31) in, Indicates the correction factor. Represents the identity matrix.

[0082] S405: Output the fused positioning results.

[0083] In this step, after the filter update is completed, the system outputs fused high-frequency positioning information, including geoordinates (latitude and longitude), attitude angles, and flight speed. This information is fed back to the flight control or navigation planning module in real time for control and trajectory tracking. The final result is that the system provides continuous, robust, and accurate high-frequency attitude solutions, supporting the navigation flight of UAVs in complex environments. To verify the performance of the solution described in this application in a real-world UAV flight environment, two sets of typical test experiments were conducted, corresponding to the positioning error assessment and trajectory output effect analysis under different flight altitude and speed conditions. The test data were obtained by aligning the real-time positioning results output by the system with the high-precision RTK values ​​before error calculation and trajectory comparison, ensuring the authenticity and representativeness of the assessment.

[0084] In Test 1, the aircraft cruised at an altitude of approximately 500 meters and a speed of approximately 20 meters per second. The system recorded the positioning results for each frame and saved them as a log file, including: the frame index, the actual latitude and longitude provided by RTK, the latitude and longitude output by the relative positioning module, and the error between the two (in meters). In addition, the success or failure of the match for each frame was also recorded, as shown in Table 2. Table 2

[0085] The test data shown in Table 2 indicates that, in 40 consecutive frames of sampling, the maximum positioning error is 30 meters, the minimum error is 2 meters, and the average error is stable at around 20 meters, reflecting that the visual positioning system has good real-time performance and error controllability.

[0086] In Test 2, the aircraft flew at an altitude of approximately 1,000 meters, with a ground speed between 15 and 30 meters per second. The flight range was approximately 6 km x 1 km, and the test environment was a typical rural-urban landscape.

[0087] The positioning test trajectory effect of the relative positioning module is as follows: Figure 8 As shown, the red trajectory represents the actual value, i.e., the GPS trajectory obtained by the flight controller, while the green trajectory is the output result of the fusion positioning.

[0088] The positioning accuracy test results of the relative positioning module are as follows: Figure 9 As shown, the average error of the relative positioning module was about 30 meters throughout the test, and there was no tendency for it to diverge as the flight distance increased.

[0089] Based on the combined results of the two sets of tests, it can be concluded that the technical solution described in this application can continuously provide positioning results with an average error of less than 50 meters under medium-to-high altitude flight conditions (500 meters to 1000 meters), and can effectively replace or supplement navigation needs when GPS signals fail.

[0090] In another exemplary embodiment, this application also provides a drone positioning device, such as... Figure 10 As shown, the device includes: an acquisition module 100, used to acquire real-time ground images and pre-stored satellite reference images captured by the UAV, as well as continuous frame images and IMU data captured by the UAV, wherein the IMU data includes acceleration and angular velocity; a preprocessing module 200, used to preprocess the real-time ground images and continuous frame images; a positioning module 300, used to input the preprocessed real-time ground images and continuous frame images, as well as the pre-stored satellite reference images and IMU data, into a pre-trained UAV positioning model to obtain a first positioning result and a second positioning result respectively; and a fusion module 400, used to fuse the first positioning result and the second positioning result to complement the positioning information in the real-time ground images captured by the UAV, the pre-stored satellite reference images, the continuous frame images captured by the UAV, and the IMU data, to obtain a final UAV positioning result.

[0091] In another exemplary embodiment, this application also provides a drone, the drone including a controller that performs a drone positioning method as described in any of the preceding embodiments.

[0092] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for locating unmanned aerial vehicles (UAVs), characterized in that, The method includes: The system acquires real-time ground images captured by the UAV and pre-stored satellite reference images, as well as continuous frame images and IMU data captured by the UAV, wherein the IMU data includes acceleration and angular velocity. The real-time captured ground images and continuous frame images are preprocessed; The pre-processed real-time captured ground images and continuous frame images, along with pre-stored satellite reference images and IMU data, are input into the pre-trained UAV localization model to obtain the first localization result and the second localization result, respectively. The first and second positioning results are fused to complement the positioning information in the real-time ground images captured by the UAV, the pre-stored satellite reference images, the continuous frame images captured by the UAV, and the IMU data, so as to obtain the final UAV positioning result.

2. The UAV positioning method according to claim 1, characterized in that, Preprocessing of the real-time captured ground images and the continuous frame images includes: The real-time captured ground images and the continuous frame images are processed into grayscale; Contrast enhancement is performed on the grayscale-converted real-time captured ground image and the consecutive frame images; Distortion is removed from the real-time captured ground images and the consecutive frame images after contrast enhancement; The distortion-free real-time captured ground images and the consecutive frame images are normalized.

3. The UAV positioning method according to claim 1, characterized in that, The UAV positioning model includes: The module includes a feature extraction module, a feature matching module, an absolute positioning module, and a relative positioning module. The feature extraction module is used to extract multi-scale features from preprocessed real-time captured ground images, consecutive frame images, and satellite reference images; The feature matching module is used to perform confidence matching on multi-scale features; The absolute positioning module is used to estimate the pose of the UAV in the geographic coordinate system based on the confidence matching results; The relative positioning module is used to obtain the continuous trajectory and motion state of the UAV based on the fusion data of preprocessed continuous frame images and IMU data.

4. The UAV positioning method according to claim 3, characterized in that, The feature extraction module includes an encoder and a decoder, wherein, The encoder is used to extract multi-scale features from preprocessed real-time captured ground images, continuous frame images, and pre-stored satellite reference images. The decoder is used to decode multi-scale features and generate high-dimensional feature vectors for feature matching and localization calculations.

5. The UAV positioning method according to claim 4, characterized in that, The decoder includes a detection header and a descriptor header, wherein, The detection head is used to decode the multi-scale features output by the encoder and detect key feature points; The descriptor head is used to generate high-dimensional feature descriptors for key feature points detected by the detection head, so as to semantically describe and uniquely encode the key feature points.

6. The UAV positioning method according to claim 3, characterized in that, The feature matching module includes: Attention graph neural network, optimal matching layer and filtering layer, in, Attention map neural networks are used to encode the location and confidence of multi-scale features and their descriptors to obtain feature matching vectors; The optimal matching layer is used to calculate the matching degree score matrix between feature matching vectors; The filtering layer is used to filter the matching score matrix to obtain a set of matching point pairs with high confidence.

7. A UAV positioning method according to claim 3, characterized in that, The absolute positioning module includes a geometric estimation module and a latitude and longitude pose recovery module, wherein... The geometric estimation module is used to estimate the relative pose of the UAV through feature matching and geometric constraints. The latitude and longitude pose recovery module is used to convert the estimated phase pose into latitude and longitude pose in the geographic coordinate system.

8. A UAV positioning method according to claim 3, characterized in that, The relative positioning module includes: The initial state estimation submodule and the sliding window optimization submodule, wherein, The initial state estimation submodule is used to recover the initial pose, scale, velocity and gravity direction of the UAV under GPS-enabled conditions. The sliding window optimization submodule is used to continuously update and improve the accuracy and stability of UAV state estimation through tight coupling of vision and IMU data.

9. A drone positioning device, characterized in that, The device includes: The acquisition module is used to acquire real-time ground images and pre-stored satellite reference images captured by the UAV, as well as continuous frame images and IMU data captured by the UAV, wherein the IMU data includes acceleration and angular velocity. The preprocessing module is used to preprocess real-time captured ground images and continuous frame images; The positioning module inputs pre-processed real-time captured ground images and continuous frame images, as well as pre-stored satellite reference images and IMU data into the pre-trained UAV positioning model to obtain the first positioning result and the second positioning result, respectively. The fusion module is used to fuse the first positioning result and the second positioning result to complement the positioning information in the real-time ground images captured by the UAV, the pre-stored satellite reference images, the continuous frame images captured by the UAV, and the IMU data, so as to obtain the final UAV positioning result.

10. A drone, characterized in that, The drone includes a controller that performs a drone positioning method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Scene matching method based on unmanned aerial vehicle image and satellite map in denial environment

    CN118053010A

  • Key point representation matching and unmanned aerial vehicle positioning method

    CN118262250A

  • Cooperative positioning and mapping method for unmanned aerial vehicle cluster

    CN118747772A

  • Unmanned aerial vehicle multi-mode visual positioning method and system in GNSS denial environment

    CN119511332A

  • Air-ground integrated unmanned aerial vehicle countering precise positioning method

    CN119826833A