An unmanned aerial vehicle image pixel-level geographic positioning method, device and equipment
Patent Information
- Application Number
- CN202610999331.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-22
AI Technical Summary
在实际应用中,无人机通常通过搭载POS设备获取位置、姿态数据),结合相机参数实现影像定位,但存在两大关键问题:一是倾斜拍摄导致影像存在严重透视变形,与卫星地图(垂直视角)的几何差异显著,传统几何校正方法难以完全消除偏差;二是异源影像(无人机与卫星、可见光与红外)存在传感器、模态、光照、季节差异,传统特征匹配算法(如SIFT、ORB)鲁棒性不足,易出现匹配失效
本发明通过正射投影消除无人机倾斜影像与卫星地图的几何差异,消除透视变形与尺度变化,将两者间的几何关系从复杂透视变换简化为平面相似变换或仿射变换,为异源特征匹配奠定基础。
Smart Images

Figure CN122793085A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) navigation and geolocation technology, and more specifically to a method, apparatus, and device for pixel-level geolocation of UAV images. Background Technology
[0002] UAV geolocation is a core supporting technology for UAVs to perform precise operations. In practical applications, UAVs typically acquire position and attitude data by carrying POS devices, and combine this with camera parameters to achieve image positioning. However, there are two major problems: First, tilted shooting causes severe perspective distortion in the images, resulting in significant geometric differences from satellite maps (vertical view), and traditional geometric correction methods are unable to completely eliminate the deviation. Second, heterogeneous images (UAV and satellite, visible light and infrared) have differences in sensors, modalities, lighting, and seasons, and traditional feature matching algorithms (such as SIFT and ORB) are not robust enough and are prone to matching failures.
[0003] In existing technologies, some solutions employ pure end-to-end deep learning models for heterogeneous image registration, but these require massive amounts of labeled data and have large model sizes, making them difficult to deploy on embedded low-computing-power platforms. Other solutions rely solely on geometric correction, failing to optimize for the visual differences in heterogeneous images, resulting in positioning accuracy often limited to the ten-meter level, which cannot meet pixel-level positioning requirements. For example, traditional orthorectification methods can only eliminate some perspective distortion, and still cannot achieve accurate matching with satellite maps when faced with terrain undulations or sensor differences. While large-scale deep learning models (such as D2-Net) can improve matching accuracy, their high computational complexity and slow inference speed on embedded platforms make them unsuitable for real-time positioning.
[0004] Therefore, proposing a pixel-level geolocation method for UAV images based on orthophoto projection and deep learning to achieve accurate registration between UAV aerial images and satellite maps is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of the above problems, the present invention is proposed to provide a method, apparatus and device for pixel-level geolocation of UAV images to overcome or at least partially solve the above problems. It utilizes tilted images captured by UAVs and their associated position and attitude data, and through a series of digital image processing and deep learning algorithms, effectively processes the registration and positioning of multimodal UAV tilted-view aerial images and satellite maps.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, embodiments of the present invention provide a pixel-level geolocation method for UAV images, comprising the following steps: S1. Acquire aerial images taken by the UAV and satellite maps of the corresponding areas, and perform orthophoto projection on the aerial images; S2. Perform image preprocessing on the orthophoto-projected aerial images and satellite maps respectively, and input them into the lightweight XFeat deep learning model to simultaneously extract heterogeneous features and perform matching; S3. Use the RANSAC algorithm to perform geometric verification on the matching point pairs and solve the homography matrix; S4. Use the homography matrix to transfer the geographic coordinates of the satellite map to the orthophoto; combine the inverse transformation of orthophoto to map any pixel in the UAV aerial image to the corresponding satellite map positioning coordinates.
[0008] Furthermore, in step S1, orthographic projection is performed on the aerial image, specifically including: The UAV pose data and camera intrinsic parameters are verified. After verification, the extrinsic parameter matrix from the camera coordinate system to the world coordinate system is calculated. Based on the extrinsic parameter matrix, the aerial image is transformed to generate an orthophoto. The orthophoto satisfies a planar similarity transformation or affine transformation relationship with the satellite map of the corresponding area in terms of geometric structure.
[0009] Furthermore, the verification process for the UAV pose data includes: The sliding window method is used to detect parameter jumps in UAV pose data. Pose data in which the changes in latitude, longitude, altitude or attitude angle exceed the corresponding preset values are marked as abnormal data. The abnormal data and their corresponding acquired image frames are then removed.
[0010] Furthermore, the verification process for the camera intrinsic parameters includes: For each frame of image, the camera intrinsic parameters associated with it are verified for focal length and principal point, and the image frames that meet the verification conditions are retained.
[0011] Furthermore, the data processing procedure for the lightweight XFeat deep learning model includes: S21. Simultaneously extract the keypoint heatmap K, feature descriptor F, and reliability weight map R using three parallel heads; wherein, the three parallel heads include a keypoint head, a descriptor head, and a reliability head; the keypoint head, based on a semi-dense keypoint detection module, predicts the probability of each spatial location becoming a stable keypoint; the descriptor head generates a discriminative 64-dimensional vector for each potential keypoint; the reliability head evaluates the confidence of feature matching in each local region; S22. Based on the reliability weight map R, the feature descriptor F is weighted and filtered, and combined with the stable key points obtained by the key point heat map K after non-maximum suppression, a coarse matching is performed between the two images. S23. The coarse matching results are optimized at the sub-pixel level through the semi-dense matching refinement module to obtain fine matching point pairs.
[0012] Furthermore, step S3 specifically includes: S31. Set the maximum number of iterations, the inlier determination threshold, and the inlier ratio termination condition for the RANSAC algorithm; perform a no-replacement selection iteration on the fine-matching point pairs; the fine-matching point pairs include pixel coordinates in the orthophoto and corresponding pixel coordinates in the satellite map. S32. In each iteration of selection without replacement, a candidate homography matrix is solved using direct linear transformation and then normalized. S33. Calculate the bidirectional reprojection error of all fine matching point pairs under the candidate homography matrix, and mark the matching point pairs with error values less than the inlier determination threshold as inliers. S34. Record the number of interior points corresponding to the current iteration, and take the candidate homography matrix with the most interior points as the current optimal homography matrix; during the iteration process, when the proportion of interior points exceeds the termination condition of the proportion of interior points or reaches the maximum number of iterations, the iteration loop is terminated and the optimal homography matrix is obtained.
[0013] Furthermore, the iterative optimization process of the lightweight XFeat deep learning model also includes: The test set is used as input, which includes oblique images taken by the UAV and the corresponding real coordinate data pairs in the satellite map of the area; Calculate the Haversine distance between the output positioning coordinates and the actual coordinates; If the value is less than the preset threshold, it is considered a valid location. If the value is greater than or equal to a preset threshold, the feature matching threshold of the XFeat deep learning model is optimized; until effective localization is achieved.
[0014] Secondly, embodiments of the present invention provide a pixel-level geolocation device for UAV images, comprising the following units: Acquisition and orthophoto projection unit: used to acquire UAV aerial images and satellite maps of the corresponding areas, and to perform orthophoto projection on the aerial images; Heterogeneous matching unit: used to preprocess the aerial images and satellite maps after orthophoto projection, and input them into the lightweight XFeat deep learning model to simultaneously extract heterogeneous features and perform matching; Mapping matrix calculation unit: used to perform geometric verification of matching point pairs using the RANSAC algorithm and solve the homography matrix; Coordinate transformation and positioning output unit: uses the homography matrix to transfer the geographic coordinates of the satellite map to the orthophoto; and combines the inverse transformation of orthophoto to map any pixel in the UAV aerial image to the corresponding satellite map positioning coordinates.
[0015] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a pixel-level geolocation method for UAV images as described in any of the first aspects.
[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a pixel-level geolocation method for UAV images as described in any of the first aspects.
[0017] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method, apparatus, and device for pixel-level geolocation of UAV images, which has the following beneficial effects: This invention eliminates the geometric differences between UAV tilted images and satellite maps through orthophoto projection, eliminating perspective distortion and scale changes, and simplifies the geometric relationship between the two from complex perspective transformation to planar similarity transformation or affine transformation, laying the foundation for heterogeneous feature matching.
[0018] Secondly, by using deep learning-based heterogeneous feature matching, the problem of feature matching between heterogeneous images such as drone images and satellite maps is solved, and the pixels selected by the user on the original tilted image are accurately converted into corresponding GPS coordinates, realizing pixel-level geolocation that can be determined by "pointing to the target". Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 This is a flowchart of a pixel-level geolocation method for UAV images provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the lightweight XFeat deep learning model provided in this embodiment of the invention. Figure 3 This is a structural diagram of a drone image pixel-level geolocation device provided in an embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1 This invention discloses a pixel-level geolocation method for UAV images, referring to... Figure 1 As shown, it includes the following steps: S1. Acquire drone aerial images and corresponding satellite maps of the area, and perform orthophoto projection on the aerial images; S2. Perform image preprocessing on the orthophoto-projected aerial images and satellite maps respectively, and input them into the lightweight XFeat deep learning model to simultaneously extract heterogeneous features and perform matching. S3. Use the RANSAC algorithm to perform geometric verification on the matching point pairs and solve the homography matrix; S4. Use the homography matrix to transfer the geographic coordinates of the satellite map to the orthophoto; combine the inverse transformation of orthophoto to map any pixel in the UAV aerial image to the corresponding satellite map positioning coordinates.
[0023] This invention can be applied to scenarios such as drone inspection, surveying and exploration, and emergency rescue. This embodiment is applied to the precise location of fault points during drone inspections of power lines. In the daily inspection of high-voltage transmission lines, maintenance personnel operate multi-rotor drones equipped with visible light cameras to fly along the line, automatically collecting high-resolution oblique images of key components such as towers, insulators, and conductors. When the AI algorithm detects damage or flashover marks on an insulator in real time, the operator needs to immediately obtain the precise geographical location of the defect in order to dispatch a maintenance team. At this time, the pixel-level geolocation method described in this invention is triggered: The system automatically acquires the current frame's oblique image and its corresponding pose (POS) data and camera intrinsic parameters, and simultaneously loads the latest high-resolution satellite map of the area. After POS verification and orthophoto projection processing, an orthophoto image geometrically aligned with the satellite image is generated. Robust matching between the orthophoto image and the satellite image is achieved using a lightweight XFeat model, and the homography matrix H is solved using RANSAC. When a user clicks on a suspected fault area (such as a damaged insulator) on the original oblique image, the pixel is mapped to the orthophoto image through inverse orthophoto transformation, and then converted to WGS84 latitude and longitude coordinates using the H matrix. Finally, the GPS coordinates of the defect point (e.g., 30.256789°N, 120.123456°E) are output, which can be directly imported into a GIS platform or navigation equipment to guide ground personnel to quickly reach the site. This process does not rely on RTK / PPK high-precision positioning equipment; it only relies on ordinary consumer-grade drones and publicly available satellite imagery to achieve "what you see is what you get" centimeter-meter-level geolocation, significantly improving the response efficiency and operational intelligence level of power line inspection.
[0024] This embodiment takes the two-step technical approach of "orthophoto projection + heterogeneous feature matching based on deep learning" as its core, and combines it with embedded platform deployment optimization to form a complete design system. The design of each module closely serves the pixel-level geolocation target of UAV imagery.
[0025] In the first step of orthophoto projection design, the core is to eliminate the geometric differences between UAV oblique imagery and satellite maps. Specifically, using the UAV's built-in POS data (latitude, longitude, altitude, roll, pitch, and yaw) and camera intrinsic parameters, the oblique imagery is converted into orthophotos with a vertical perspective based on the collinearity equation or digital projection transformation in photogrammetry. During the design process, the validity of the POS data (e.g., no abnormal jump values) and the accuracy of the camera intrinsic parameters (focal length, principal point coordinates) must first be verified to avoid the impact of original data errors on subsequent correction effects. Next, based on the POS data and the true GPS coordinates of the image center point, the initial extrinsic parameter matrix for each image is calculated from the camera coordinate system to the world coordinate system. Finally, relying on the extrinsic parameter matrix, the mapping from oblique imagery to orthophotos is completed, completely eliminating perspective distortion and scale changes, simplifying the geometric relationship between the two from complex perspective transformations to planar similarity transformations or affine transformations, laying the foundation for heterogeneous feature matching.
[0026] The implementation steps of the first step in this embodiment are described in detail below: S1. Acquire drone aerial images and corresponding satellite maps of the area.
[0027] This embodiment inputs UAV aerial images, POS data, camera intrinsic parameters, and corresponding regional satellite maps, parses image metadata, and verifies the validity of POS data and camera intrinsic parameters.
[0028] The drone aerial images are in visible light / infrared mode and in JPG format; the POS data includes the longitude, latitude, altitude, roll angle, pitch angle, and yaw angle at the time of shooting; and the camera intrinsic parameters include focal length and principal point coordinates.
[0029] First, the EXIF or XMP metadata of the tilted image is parsed to extract the embedded timestamps, GPS location, attitude information, and camera parameters as the initial data source; then, the validity of the acquired POS data and camera intrinsic parameters is verified.
[0030] The specific parsing process is as follows: Extract image metadata, including shooting time, sensor model, and image resolution, using the ExifTool tool; read the POS data file (supporting .txt / .csv format) using a custom parsing function, and extract the latitude, longitude, height, yaw, pitch, and roll of each frame of the image; camera intrinsic parameters are obtained through calibration files provided by the drone manufacturer or on-site calibration and stored in matrix form.
[0031] The verification process includes: POS data verification: A sliding window method is used. In this embodiment, the window size is 5 frames. The rate of change of each parameter is calculated. If the difference between a frame of data and the adjacent frame exceeds the threshold, such as lat / lon change > 0.0001°, height change > 5m, angle change > 5° in this embodiment, it is judged as an abnormal jump and an error feedback is triggered. At the same time, the integrity of the data is verified. Images that are missing key parameters such as lon or pitch are directly marked as invalid.
[0032] Camera intrinsic parameter verification: Verify the reasonableness of the focal length. Determine the normal range based on the drone camera model, such as 5-20mm. The principal point coordinates must be close to the center of the image resolution. If the deviation exceeds 100 pixels, it is considered invalid, and you will be prompted to re-enter or calibrate.
[0033] Verification result processing: If the data is abnormal, the system output includes the abnormality type. For example, feedback information such as POS jump, invalid internal parameters, and corresponding image number. This embodiment supports manual correction or automatic removal of invalid images.
[0034] Secondly, based on the verified UAV pose data and camera intrinsic parameters, the extrinsic parameter matrix from the camera coordinate system to the world coordinate system is calculated, and the UAV aerial images are transformed into orthophotos based on the extrinsic parameter matrix. After processing the abnormal data, this embodiment calculates the extrinsic parameter matrix. Based on the POS data and the true GPS coordinates of the image center point, an initial extrinsic parameter matrix from the camera coordinate system to the world coordinate system is established using photogrammetry formulas, laying the foundation for coordinate mapping. Subsequently, an orthophoto is generated. Based on the collinearity equation or digital projection transformation, the extrinsic parameter matrix is used to map the tilted image pixels to a plane perpendicular to the ground, eliminating perspective distortion and scale changes, making the orthophoto consistent with the perspective of the satellite map, and simplifying the geometric relationship to a plane similarity transformation or affine transformation. Finally, image preprocessing is performed, using contrast stretching or CLAHE to enhance ground features, and scaling the image to the input size of the XFeat model to achieve data standardization.
[0035] Regarding the calculation of the extrinsic parameter matrix: Based on the POS data and the true GPS coordinates of the image center point, an initial extrinsic parameter matrix M=[R|t] is established from the camera coordinate system to the world coordinate system using photogrammetry formulas. Here, R is the rotation matrix, obtained through yaw, pitch, and roll transformations, and t is the translation vector, calculated from latitude, longitude, and altitude to the origin of the world coordinate system. Specifically, the latitude and longitude are first converted to geodetic coordinates (X, Y, Z). Then, combining the camera height and attitude angles, R is calculated using the Euler angle rotation formula, and the translation vector t=(X0, Y0, Z0), where (X0, Y0, Z0) are the geodetic coordinates of the image center point.
[0036] In the second step, the design of heterogeneous feature matching based on deep learning, this embodiment focuses on solving the feature matching problem of heterogeneous images (UAVs and satellites, visible light and infrared). The lightweight XFea deep learning model is selected as the core, and its architecture is specifically designed to adapt to low computing power scenarios and heterogeneous matching requirements.
[0037] The second step in this embodiment first performs XFeat feature extraction. The preprocessed orthophoto and satellite map are input into the XFeat model and processed by a lightweight backbone network—the initial layer receives the image with 4 channels, gradually increasing to 128 channels as the spatial resolution decreases. Simultaneously, 1 / 8, 1 / 16, and 1 / 32 scale feature maps are fused. Then, a dual-branch feature extractor outputs sub-pixel-level key points (filtered by "dustbin") and a 64-dimensional dense feature map (combined with a reliability heatmap for selection). Next, coarse matching is performed, using nearest neighbor search based on feature descriptors. The initial matching pairs are quickly obtained to narrow down the scope of subsequent processing. Then, the fine matching stage is entered, and the XFeat semi-dense matching refinement module is enabled. Depending on the scene, either sparse (4096 key points) or semi-dense (10000 feature regions) mode is selected. The 8×8 offset probability distribution is predicted by a lightweight MLP, and the original resolution matching points are reversed according to the formula to correct minor deviations. Finally, RANSAC geometric verification is performed to iteratively filter inliers and remove outliers. The homography matrix H is calculated to establish pixel-level geometric mapping relationships. If the outlier ratio is too high, coarse and fine matching are re-executed.
[0038] In the coordinate transformation and positioning output stage, GPS coordinates are first transferred by using the homography matrix H to transfer the WGS84 coordinates of the satellite map to the UAV orthophoto. Then, arbitrary point coordinate extraction is achieved. After the user selects a pixel on the original oblique image, it is mapped to the orthophoto through inverse orthophoto transformation. Combined with the coordinate mapping relationship, the accurate GPS coordinates are obtained, realizing "point-to-point positioning". Finally, the accuracy is evaluated by calculating the distance between the positioning point and the true value. If the distance is less than 10m, it is considered a valid positioning and the success rate is calculated. The average accuracy of the valid positioning points is calculated. If the expected accuracy is not achieved, the model or orthophoto parameters are optimized.
[0039] The implementation steps of the second step in this embodiment are described in detail below: According to S2, the orthophotos and satellite maps are preprocessed separately and then input into the lightweight XFeat deep learning model to simultaneously extract heterogeneous features and perform matching.
[0040] Orthophotos and satellite maps are preprocessed in the same or different ways, such as size normalization, modality adaptation, contrast enhancement, and data format standardization.
[0041] The input is fed into the lightweight XFeat deep learning model, which employs a shared backbone network plus a three-way parallel head structure. (Refer to...) Figure 2 As shown, the feature map is extracted into a low-resolution feature tensor by a lightweight backbone network. Its format is H / 8×W / 8×64, where H and W represent the height and width of the original input image, H / 8×W / 8 means that the size of the feature map is only 1 / 64 of the original image, and 64 is the number of feature channels.
[0042] Subsequent processing is all based on this low-resolution feature map, significantly reducing memory usage and computational complexity. The three-way parallel header structure, based on the shared feature map, deploys three lightweight header modules to achieve simultaneous output of keypoint detection, feature descriptor generation, and reliability assessment. (1) The H / 8×W / 8×64 feature map is expanded into 8×8 spatial blocks to form a local feature vector sequence; a “1×1 convolution × 4” operation is performed on each 8×8 block to extract high-dimensional semantic information; the result is then folded back into the original image grid to output a key point heatmap K of size H×W×1, representing the probability that each pixel position becomes a key point; stable key point coordinates are extracted from K through non-maximum suppression (NMS).
[0043] (2) The descriptor header upsamples the H / 8×W / 8×64 feature map to H / 4×W / 4×64; after the number of channels is compressed by a 1×1 convolutional layer, it is upsampled to H×W×64; the output is a feature descriptor F of H×W×64, which is used for matching similarity calculation; this path introduces a "fusion block" structure, which combines multi-scale context information to improve the descriptor discrimination ability.
[0044] (3) The reliability header directly performs a 1×1 convolution on the H / 8×W / 8×64 feature map and outputs the H / 8×W / 8×1 reliability weight map R; R represents the confidence of the local region feature quality, which is used for subsequent weighted matching and screening to improve robustness.
[0045] This embodiment synchronously outputs three tensors of the same size (H / 8 × W / 8) as the input feature tensor through three parallel heads, which together support the heterogeneous image feature matching task. These are a keypoint heatmap K, a feature descriptor F, and a reliability weight map R. The keypoint heatmap K characterizes the confidence of each location in the image as a matching keypoint; a higher value indicates more stable features at that location and a greater suitability as a matching point. The feature descriptor F generates a highly distinctive feature vector for each location and possesses cross-modal (visible light-infrared) and cross-resolution (UAV-satellite) invariance, which can be used to calculate the similarity of features between different images. The reliability weight map R is responsible for evaluating the reliability of features at the corresponding location, filtering out low-quality features caused by noise and texture differences in heterogeneous images, thereby improving matching accuracy.
[0046] Figure 2 All intermediate feature tensors (such as H / 8×W / 8×64 gray squares) maintain a small size of H / 8×W / 8 (only 8x downsampling), which controls the memory usage of tensors while preserving sufficient spatial details; at the same time, the head modules all adopt lightweight operations—the key point is that the header is "8×8 block expansion ( Figure 2 (in the dashed box "expanded into 8×8 blocks") + 1×1 convolution (×4) + folding ( Figure 2 The descriptor header is "1×1 convolution + upsampling ("Fold back image")". Figure 2 The "upsampling" arrow indicates that there is no redundant computation of complex convolutional blocks, and the three headers (keypoint / descriptor / reliability, corresponding to...) are... Figure 2 The "Key Point Header" in the upper middle and the "Descriptor Header" and "Reliability Header" in the lower middle are arranged in parallel, which further reduces the inference time and accurately adapts to the performance limitations of low computing power devices.
[0047] To address the robustness requirements of heterogeneous matching, Figure 2 The multi-task branching and feature fusion module in the middle form a targeted design: Figure 2The model simultaneously outputs three branches: a key point heatmap K, a descriptor F, and a reliability map R. K is responsible for locating stable feature points in heterogeneous images (such as UAVs / satellites, visible light / infrared), F extracts cross-modal invariant feature representations, and R filters low-quality features caused by heterogeneous noise. Multi-task joint learning allows the model to simultaneously cover the "localization-representation-denoising" requirements of heterogeneous matching. The "upsampling + element-wise summation" fusion block in the descriptor header integrates multi-scale feature information, which can better cope with the resolution and texture differences of heterogeneous images and improve matching robustness.
[0048] These three outputs are not the final results. They will be used to complete a simplified heterogeneous image matching process: First, reliable features in the feature descriptor F are filtered using the reliability weight map R. Then, stable keypoints selected by the keypoint heatmap K are optimized using R. Finally, the keypoints and corresponding descriptors F of the two heterogeneous images are matched, and incorrect matches are eliminated to obtain accurate matching point pairs, providing support for subsequent image registration, stitching, and other tasks. The specific process of heterogeneous image matching is as follows: First, the feature descriptor F is weighted and filtered using a reliability weight map R to retain reliable features and remove low-quality features, thereby reducing the false match rate in subsequent matching. Next, non-maximum suppression (NMS) is applied to the keypoint heatmap K to select the peak points with the highest local responses as the final keypoints. At the same time, the keypoint set is further optimized by combining the weights of R to ensure the stability and repeatability of keypoints in heterogeneous images. Finally, for the keypoints and corresponding feature descriptors F extracted from the two heterogeneous images, algorithms such as K-nearest neighbor matching (KNN) or brute force matching (BF) are used to calculate the similarity between the descriptors. Then, erroneous matching pairs are removed by cross-validation or distance thresholding, finally obtaining accurate matching point pairs between the two heterogeneous images, providing core support for subsequent image registration, stitching and other tasks.
[0049] Then, following S3, the RANSAC algorithm is used to perform geometric verification on the matching point pairs and solve for the homography matrix.
[0050] The set of finely matched point pairs output by the model in this embodiment is {(x i ,y i ) (x i ′,y i ′)} M i=1 , where (x i ,y i (x) represents the pixel coordinates in the orthophoto, and (x) represents the pixel coordinates in the orthophoto. i ′,y i ′) represents the corresponding pixel coordinates in the satellite map. Set the RANSAC parameters, including the maximum number of iterations N. max=1000, the inlier determination threshold τ=3 pixels, and the inlier ratio termination condition η=70%; in this embodiment, in each iteration, 4 pairs of non-collinear matching points are randomly selected without replacement, and a candidate homography matrix H is solved using the direct linear transformation method. cand ∈R 3×3 And normalize it.
[0051] Then, for all M pairs of matching points, calculate their position in H. cand The bidirectional reprojection error under the given conditions is expressed by the formula:
[0052]
[0053]
[0054]
[0055]
[0056] Among them, P i P represents the pixel coordinates of the i-th matching point in the orthophoto. i In satellite maps and P i The corresponding pixel coordinates of the matching point This represents the candidate homography matrix after normalization. This indicates that the orthophoto point P is... i Predicted locations by projecting candidate matrices onto satellite maps This indicates that point P on the satellite map will be... i The predicted position is projected back to the orthophoto image through inverse transformation. Let represent the symmetric reprojection error of the i-th matching point pair, and represent the total distance deviation of the bidirectional projection. If e i If the value is less than τ, then mark the point pair as an interior point; count the number of interior points in the current model.
[0057] Record the candidate matrix with the most interior points. If the proportion of interior points obtained in a certain iteration exceeds a preset threshold (such as 70%), or the maximum number of iterations is reached, the iteration will be terminated early.
[0058] Use all Matching point pairs determined to be interior points are then analyzed using the nonlinear least squares method. Global optimization is performed to further minimize reprojection errors, yielding the final homography matrix H. The optimized homography matrix H is then used for subsequent geographic coordinate transfer, achieving geometric alignment between orthophotos and satellite maps. This process effectively eliminates interference from mismatched point pairs, ensuring that the obtained homography matrix H possesses high geometric consistency and robustness, providing a reliable spatial transformation basis for pixel-level geolocation.
[0059] Finally, according to S4, the geographic coordinates of the satellite map are transferred to the orthophoto using the homography matrix; combined with the inverse transformation of orthophoto, any pixel in the UAV aerial image is mapped to the corresponding satellite map positioning coordinates.
[0060] After obtaining the optimal homography matrix H verified by RANSAC geometry, the following two-stage coordinate mapping is performed to achieve pixel-level geolocation.
[0061] 1. GPS coordinate transfer to establish orthophoto geographic reference.
[0062] Using the geographic metadata of the satellite map (read from the GDAL library), obtain the WGS84 latitude and longitude coordinates (lat) corresponding to each pixel. s ,lon s Combined with the inverse transformation H of the homography matrix H. 1 This transfers the geographic coordinate system of the satellite map to the orthophoto from the UAV. For any pixel coordinate (x, y) in the orthophoto... o ,y o ), its corresponding WGS84 coordinates (lat o ,lon o The following method is used to determine (x) o ,y o ) is considered as the projected location on the satellite map, and H is applied in reverse. 1 Obtain the corresponding pixels in the satellite image; using the affine transformation parameters or geocoding information provided by GDAL, convert the satellite image pixels into WGS84 latitude and longitude; thus, the entire orthophoto is given a precise geographic coordinate reference.
[0063] 2. Extract coordinates of any point, "point-to-location" positioning.
[0064] When a user clicks on the target pixel (x) in the original tilted aerial imagery t ,y t When: Based on the collinear equation or digital projection transformation in step S2, through the inverse transformation of orthographic projection (i.e., the reverse process of finding the intersection of the camera ray and the projection elevation surface), (x) t ,y t ) mapped to the corresponding pixel (x) in the orthophotoo ,y o Based on the established orthophoto geographic mapping relationship, the value of (x) can be found. o ,y o The corresponding WGS84 coordinates (lat) t ,lon t Finally, the system outputs the precise GPS coordinates of the clicked point, achieving pixel-level geolocation that is "what you see is what you get".
[0065] To quantify positioning performance and ensure system reliability, this embodiment introduces the following accuracy evaluation process: Distance calculation. During the iterative optimization of the lightweight XFeat deep learning model, a test set is used as input. This test set includes oblique images taken by the UAV and corresponding real coordinate data pairs from satellite maps of the area. For each positioning result (lat... t ,lon t ) and true coordinates (lat) true ,lon true The spherical distance d (in meters) is calculated using the Haversine formula:
[0066] Where R is the Earth's radius, R = 6371 km; φ1 represents the latitude of the positioning result, φ2 represents the latitude of the actual coordinates, Δφ represents the latitude difference between the two locations, and Δλ represents the longitude difference between the two locations.
[0067] If d < 10m, it is considered a valid location and the success rate (number of valid locations / total number of locations) is calculated. The average accuracy (average of all d values) is calculated for all valid locations. If the average accuracy does not meet expectations (e.g., > 3m), the optimization of orthophoto projection parameters (e.g., re-verification of the extrinsic parameter matrix) or the XFeat model is returned, and the feature matching threshold is adjusted. The feature matching thresholds in this embodiment include: First, the response probability threshold of the key point heatmap K, which is the probability standard set before non-maximum suppression. Only potential stable key points with heatmap pixel values meeting the standard (0~1 range) are retained as pre-matching screening. Second, the confidence screening threshold of the reliability weight map R. Based on the matching confidence score of R, high-weight feature descriptors F are retained, and low-confidence features are filtered to reduce the coarse matching mismatch rate. Third, the similarity matching threshold of feature descriptors F. The matching result is determined by calculating the 64-dimensional vector similarity (e.g., cosine similarity). Only those meeting the standard form coarse matching point pairs, which is the final screening threshold for matching.
[0068] Example 2 This invention discloses a pixel-level geolocation device for UAV images, referring to... Figure 3 As shown, it includes the following units: Acquisition and orthophoto projection unit: used to acquire UAV aerial images and satellite maps of the corresponding areas, and to perform orthophoto projection on the aerial images; Heterogeneous matching unit: used to preprocess the aerial images and satellite maps after orthophoto projection, and input them into the lightweight XFeat deep learning model to simultaneously extract heterogeneous features and perform matching; Mapping matrix calculation unit: used to perform geometric verification of matching point pairs using the RANSAC algorithm and solve the homography matrix; Coordinate transformation and positioning output unit: uses the homography matrix to transfer the geographic coordinates of the satellite map to the orthophoto; and combines the inverse transformation of orthophoto to map any pixel in the UAV aerial image to the corresponding satellite map positioning coordinates.
[0069] This embodiment is designed for embedded platform deployment and is adapted to low computing power. The XFeat model is converted to .rknn format using the rknn inference framework, ensuring that all operators are compatible with the NPU and avoiding CPU fallback caused by operator incompatibility, thus improving inference speed. Regarding hardware load distribution, orthophoto projection and image preprocessing (enhancement, normalization) are performed by the CPU, while XFeat feature extraction and inference are handled by the NPU. Feature matching post-processing (KDTree construction, RANSAC) is implemented through CPU multi-threaded scheduling to balance the hardware load. For memory management, a dedicated memory pool is designed to store large blocks of data such as high-resolution images and dense feature maps output by XFeat, avoiding memory fragmentation and overflow (OOM) risks caused by frequent allocation and release, while also considering chip power consumption and thermal management to prevent frequency throttling issues caused by continuous high load. Furthermore, software libraries such as OpenCV, Eigen, and Gdal are integrated and cross-compiled and linked in the RK3588 embedded environment. Algorithms not supported by the NPU are implemented using a CPU+GPU OpenCL hybrid approach to ensure stable system operation.
[0070] By employing acquisition units, orthorectification units, heterogeneous matching units, mapping matrix calculation units, and coordinate transformation and positioning output units to perform orthorectification and deep learning-based heterogeneous feature matching, the core pain points in the existing UAV image geolocation field are accurately addressed. Furthermore, it possesses multi-dimensional technological advantages, as detailed below: Regarding the technical challenges addressed, the invention first tackles the geometric differences between UAV and satellite imagery. Traditional geometric correction can only eliminate some perspective distortion, while the orthophoto projection of the orthophoto transformation unit in this invention completely eliminates perspective distortion and scale changes in tilted images through collinearity equations, simplifying complex perspective transformations into planar similarity transformations or affine transformations. This lays a reliable geometric foundation for subsequent matching and avoids matching failures caused by differences in viewpoint and scale. Secondly, it addresses the problem of cross-difference matching between heterogeneous images. UAV and satellite imagery differ in sensor and resolution, while visible light and infrared imagery differ in modality. Traditional algorithms struggle to extract robust features. The heterogeneous matching unit in this invention is based on the XFeat model. Its lightweight backbone network's multi-scale fusion can adapt to resolution differences, the reliability heatmap of the dual-branch feature extractor can screen cross-modal robust features, and semi-dense matching refinement can correct subtle deviations, achieving accurate matching of cross-modal features such as infrared heat sources and visible light contours. Furthermore, this invention addresses the deployment challenge on low-computing-power platforms. Pure end-to-end models, with their numerous parameters and high computational demands, are unsuitable for embedded platforms. XFeat's lightweight model design (early implementation of fewer channels and multi-scale scaling to reduce computational load) and RKNN format conversion adapt to the NPU while ensuring real-time inference. Combined with CPU / NPU load balancing and memory pool management, resource bottlenecks are avoided, resolving the contradiction between the difficulty of deploying large models and the low accuracy of simple algorithms. Finally, it addresses the reliability issue of geometric verification. Existing solutions assume homography between images, failing in terrain-undulating areas, and the high ratio of outliers in heterogeneous matching leads to RANSAC verification failure. This invention first simplifies geometric relationships through orthophoto projection, then reduces the number of outliers through XFeat's precise matching, ensuring efficient RANSAC filtering of inliers and improving positioning reliability.
[0071] In terms of technological advantages, firstly, it boasts strong robustness, compatible with both visible light and infrared modes, and can cope with differences in lighting, seasons, and sensors. The positioning success rate of both visible light and infrared sequences is superior to traditional algorithms (such as the ORB algorithm), and it can still work stably in complex environments such as nighttime infrared inspections and winter snow mapping. Secondly, it offers high accuracy. Orthographic projection eliminates geometric deviations, and XFeat semi-dense matching achieves sub-pixel-level optimization, with positioning accuracy far exceeding the ten-meter accuracy of traditional geometric correction. It can accurately obtain the GPS coordinates of any pixel, meeting detailed positioning requirements. Thirdly, it offers flexible deployment and low cost, adapting to embedded platform deployment and integrating into UAV onboard systems without relying on ground computing power. It also reduces reliance on high-precision POS equipment, achieving accurate positioning through low-cost GPS combined with a two-step route. Hardware and data costs are far lower than pure end-to-end solutions. Fourthly, it boasts high engineering value. Its two-step approach addresses real-world pain points: the engineering implementation of multimodal heterogeneous matching, the precision-efficiency-power balance of the embedded platform, and the collaborative optimization of geometry and deep learning. All these aspects are geared towards industry application needs and can be directly applied to scenarios such as inspection and rescue, transforming laboratory technology into practical products. Fifthly, it offers excellent interpretability. Orthographic projection is based on well-defined photogrammetric principles, and XFeat features can be traced back to the structure of ground features (such as building shapes and road orientations). Compared to pure end-to-end black-box models, it is easier to troubleshoot matching failures, facilitate parameter tuning and technology iteration, and reduce subsequent maintenance costs.
[0072] Example 3 This invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a UAV image pixel-level geolocation method as described in any one of Embodiments 1.
[0073] Example 4 This invention discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements a pixel-level geolocation method for UAV images as described in any one of Embodiments 1.
[0074] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0075] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A pixel-level geolocation method for UAV imagery, characterized in that, Includes the following steps: S1. Acquire aerial images taken by the UAV and satellite maps of the corresponding areas, and perform orthophoto projection on the aerial images; S2. Perform image preprocessing on the orthophoto-projected aerial images and satellite maps respectively, and input them into the lightweight XFeat deep learning model to simultaneously extract heterogeneous features and perform matching; The data processing procedure for the lightweight XFeat deep learning model includes: S21. Simultaneously extract the keypoint heatmap K, feature descriptor F, and reliability weight map R using three parallel heads; wherein, the three parallel heads include a keypoint head, a descriptor head, and a reliability head; the keypoint head, based on a semi-dense keypoint detection module, predicts the probability of each spatial location becoming a stable keypoint; the descriptor head generates a discriminative 64-dimensional vector for each potential keypoint; the reliability head evaluates the confidence of feature matching in each local region; S22. Based on the reliability weight map R, the feature descriptor F is weighted and filtered, and combined with the stable key points obtained by the key point heat map K after non-maximum suppression, a coarse matching is performed between the two images. S23. The coarse matching results are optimized at the sub-pixel level through the semi-dense matching refinement module to obtain fine matching point pairs; S3. Use the RANSAC algorithm to perform geometric verification on the matching point pairs and solve the homography matrix; S4. Use the homography matrix to transfer the geographic coordinates of the satellite map to the orthophoto; combine the inverse transformation of orthophoto to map any pixel in the UAV aerial image to the corresponding satellite map positioning coordinates.
2. The method as described in claim 1, characterized in that, In step S1, orthophoto projection is performed on the aerial image, specifically including: The UAV pose data and camera intrinsic parameters are verified. After verification, the extrinsic parameter matrix from the camera coordinate system to the world coordinate system is calculated. Based on the extrinsic parameter matrix, the aerial image is transformed to generate an orthophoto. The orthophoto satisfies a planar similarity transformation or affine transformation relationship with the satellite map of the corresponding area in terms of geometric structure.
3. The method as described in claim 2, characterized in that, The verification process for the UAV pose data includes: The sliding window method is used to detect parameter jumps in UAV pose data. Pose data in which the changes in latitude, longitude, altitude or attitude angle exceed the corresponding preset values are marked as abnormal data. The abnormal data and their corresponding acquired image frames are then removed.
4. The method as described in claim 3, characterized in that, The camera intrinsic parameter verification process includes: For each frame of image, the camera intrinsic parameters associated with it are verified for focal length and principal point, and the image frames that meet the verification conditions are retained.
5. The method as described in claim 1, characterized in that, Step S3 specifically includes: S31. Set the maximum number of iterations, the inlier determination threshold, and the inlier ratio threshold for the RANSAC algorithm; perform iterative selection without replacement on the fine-matching point pairs; the fine-matching point pairs include pixel coordinates in the orthophoto and corresponding pixel coordinates in the satellite map. S32. In each iteration of selection without replacement, a candidate homography matrix is solved using direct linear transformation and then normalized. S33. Calculate the bidirectional reprojection error of all fine matching point pairs under the candidate homography matrix, and mark the matching point pairs with error values less than the inlier determination threshold as inliers. S34. Record the number of interior points corresponding to the current iteration, and take the candidate homography matrix with the most interior points as the current optimal homography matrix; during the iteration process, when the proportion of interior points exceeds the interior point proportion threshold or the number of iterations reaches the maximum number of iterations, terminate the iteration loop and obtain the optimal homography matrix.
6. The method as described in claim 1, characterized in that, The iterative optimization process of the lightweight XFeat deep learning model also includes: The test set is used as input, which includes oblique images taken by the UAV and the corresponding real coordinate data pairs in the satellite map of the area; Calculate the Haversine distance between the output positioning coordinates and the actual coordinates; If the value is less than the preset threshold, it is considered a valid location. If the value is greater than or equal to a preset threshold, the feature matching threshold of the XFeat deep learning model is optimized; until effective localization is achieved.
7. A pixel-level geolocation device for UAV images, characterized in that, Includes the following units: Acquisition and orthophoto projection unit: used to acquire UAV aerial images and satellite maps of the corresponding areas, and to perform orthophoto projection on the aerial images; Heterogeneous matching unit: used to preprocess the aerial images and satellite maps after orthophoto projection, and input them into the lightweight XFeat deep learning model to simultaneously extract heterogeneous features and perform matching; Mapping matrix calculation unit: used to perform geometric verification of matching point pairs using the RANSAC algorithm and solve the homography matrix; Coordinate transformation and positioning output unit: uses the homography matrix to transfer the geographic coordinates of the satellite map to the orthophoto; and combines the inverse transformation of orthophoto to map any pixel in the UAV aerial image to the corresponding satellite map positioning coordinates.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a pixel-level geolocation method for UAV images as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a pixel-level geolocation method for UAV images as described in any one of claims 1 to 6.