Adaptive variable structure unmanned aerial vehicle multi-modal scene matching navigation positioning method and device

By using coordinate transformation and image matching techniques from UAV aerial images, errors are corrected, and perspective transformation and homography transformation matrices are utilized to solve the problems of viewpoint difference compensation and computational resource constraints in UAV image localization, thus achieving high-precision multimodal image matching and localization.

CN121048628APending Publication Date: 2025-12-02BEIHANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511259883.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing UAV image positioning technology suffers from significantly reduced positioning accuracy and reliability when GPS signals are limited, camera calibration accuracy is insufficient, or flight attitude measurement errors are large. Furthermore, it faces challenges such as difficulty in compensating for viewpoint differences, limitations in multimodal image processing, and constraints on computing resources.

Method used

By acquiring the drone pose information from drone aerial images and performing coordinate transformation, the drone's geographical location error is corrected. Image matching is then performed using perspective transformation and homography transformation matrices to determine the latitude and longitude of points of interest, achieving multimodal adaptive processing and sub-pixel level accuracy.

Benefits of technology

It significantly improves the system versatility, robustness, and positioning accuracy of UAV image localization, solves the problems of viewpoint difference compensation and computational resource constraints, and achieves high-precision multimodal image matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121048628A_ABST
    Figure CN121048628A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of scene matching navigation, and provides a multi-mode scene matching navigation positioning method and device for a self-adaptive variable-structure unmanned aerial vehicle. According to the method, geographic coordinates corresponding to all pixel points in an aerial image are determined according to pose information in the aerial image of the unmanned aerial vehicle, and the geographic position of the unmanned aerial vehicle with errors is corrected based on difference information between the pixel coordinates of the central point of the aerial image of the unmanned aerial vehicle and the geographic position corresponding to the center of a camera. Based on the corrected geographic position of the unmanned aerial vehicle, determining a reference image pixel coordinate corresponding to the angular point pixel coordinate of the aerial image, and obtaining a simulated aerial image under the view angle of the unmanned aerial vehicle through perspective transformation so as to determine a homography transformation matrix; according to the method, the satellite reference image pixel coordinates of the interest point in the aerial image are determined based on the homography transformation matrix and the perspective transformation matrix, and finally the longitude and latitude of the interest point are determined, so that the limitation of the traditional method in view angle difference compensation, computing resource constraint and cross-modal processing is solved, and the universality, robustness and positioning accuracy of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of scene matching navigation technology, and in particular to an adaptive variable structure UAV multimodal scene matching navigation and positioning method and device. Background Technology

[0002] The rapid development of UAV technology has led to its increasingly widespread application in fields such as military reconnaissance, civilian surveying and mapping, environmental monitoring, and target recognition. In these applications, accurately determining the true geographical location of target objects or points of interest within UAV imagery is a critical technological requirement.

[0003] Traditional target localization methods primarily rely on the UAV's own GPS positioning information combined with camera calibration parameters for geometric projection calculations. However, under conditions of limited GPS signal, insufficient camera calibration accuracy, or large flight attitude measurement errors, positioning accuracy and reliability significantly decrease. Image-matching-based spatial positioning technology establishes a correspondence between pixels and real geographical locations by matching real-time UAV images with satellite reference images possessing precise geographic coordinates, offering advantages such as high positioning accuracy and strong anti-interference capabilities. However, existing image-matching-based positioning technologies still face challenges such as difficulties in compensating for viewpoint differences, limitations in multimodal image processing, and computational resource constraints.

[0004] Therefore, how to provide an image spatial positioning system that integrates multimodal processing, adaptive structural optimization, and high-precision matching, so that the image spatial positioning system has multimodal adaptive processing capabilities, hardware platform adaptability, and sub-pixel level matching accuracy, is a technical problem that needs to be solved. Summary of the Invention

[0005] In view of this, embodiments of this application provide an adaptive variable structure UAV multimodal scene matching navigation and positioning method and apparatus to solve the problems of weak image spatial positioning adaptive capability and insufficient matching accuracy in the prior art.

[0006] A first aspect of this application provides an adaptive variable structure unmanned aerial vehicle (UAV) multimodal scene matching navigation and positioning method, comprising:

[0007] Acquire drone aerial images, and use the drone pose information in the drone aerial images to perform coordinate transformation to obtain the geographic coordinates corresponding to each pixel in the aerial images;

[0008] Determine the difference between the pixel coordinates of the center point of the drone aerial image and the geographical location corresponding to the camera center, and use the difference information to correct the drone geographical location that contains errors obtained through low-precision inertial navigation.

[0009] Based on the corrected UAV geographical location, the corresponding reference image pixel coordinates of the corner pixel coordinates of the aerial image are determined. Perspective transformation is performed on the satellite reference image to obtain the simulated aerial image under the specific UAV attitude view obtained by the inertial navigation.

[0010] Determine target feature extraction and image matching models based on simulated aerial images and drone aerial images;

[0011] The homography transformation matrix is ​​determined based on the target feature extraction and image matching model;

[0012] Obtain the pixel coordinates of points of interest in the drone aerial image, and transform the pixel coordinates of the points of interest to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix;

[0013] The latitude and longitude of points of interest are determined based on the pixel coordinates of satellite reference images.

[0014] A second aspect of this application provides an adaptive variable structure unmanned aerial vehicle (UAV) multimodal scene matching navigation and positioning device, comprising:

[0015] The acquisition module is configured to acquire drone aerial images, and use the drone pose information in the drone aerial images to perform coordinate transformation to obtain the geographic coordinates corresponding to each pixel in the aerial images.

[0016] The correction module is configured to determine the difference information between the pixel coordinates of the center point of the drone aerial image and the geographical location corresponding to the camera center, and use the difference information to correct the drone geographical location that contains errors obtained through low-precision inertial navigation.

[0017] The simulation module is configured to determine the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image based on the corrected UAV geographical location, and perform perspective transformation on the satellite reference image to obtain the simulated aerial image under the specific UAV perspective obtained by the inertial navigation system.

[0018] The selection module is configured to determine the target feature extraction and image matching model based on simulated aerial images and drone aerial images;

[0019] The determination module is configured to determine the homography transformation matrix based on the target feature extraction and image matching model;

[0020] The inverse transformation module is configured to obtain the pixel coordinates of the point of interest in the drone aerial image and transform the pixel coordinates of the point of interest to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix.

[0021] The positioning module is configured to determine the latitude and longitude of points of interest based on the pixel coordinates of a satellite reference image.

[0022] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0023] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0024] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment first performs coordinate transformation on the drone pose information in the drone aerial image to obtain the geographic coordinates corresponding to each pixel in the aerial image. Then, it corrects the drone geographic location with errors based on the difference information between the pixel coordinates of the center point of the drone aerial image and the geographic location corresponding to the camera center. Then, it determines the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image based on the corrected drone geographic location. Then, it obtains a simulated aerial image from a specific drone perspective through perspective transformation. Next, it determines the homography transformation matrix based at least on the simulated aerial image. Based on the homography transformation matrix and the perspective transformation matrix, it determines the satellite reference image pixel coordinates of the points of interest in the aerial image. Finally, it determines the latitude and longitude of the points of interest. This solves the limitations of traditional methods in terms of perspective difference compensation, computational resource constraints, and cross-modal processing, and significantly improves the system's versatility, robustness, and positioning accuracy. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating an adaptive variable structure UAV multimodal scene matching navigation and positioning method provided in an embodiment of this application.

[0027] Figure 2 This is a flowchart illustrating the method for coordinate transformation using drone pose information from drone aerial photographs provided in this application.

[0028] Figure 3 This is a flowchart illustrating the method for determining the difference between the pixel coordinates of the center point of an aerial photograph taken by a drone and the geographical location corresponding to the camera center, as provided in an embodiment of this application.

[0029] Figure 4This is a flowchart illustrating a method for correcting the geographical location of a UAV obtained through low-precision inertial navigation using difference information, as provided in an embodiment of this application.

[0030] Figure 5 This is a flowchart illustrating the method for determining the reference image pixel coordinates corresponding to the corner pixel coordinates of an aerial photograph based on the corrected UAV geographical location, as provided in this application embodiment.

[0031] Figure 6 This is a flowchart illustrating the method provided in this application for obtaining a simulated aerial photograph from a specific UAV perspective by performing perspective transformation on a satellite reference image, corresponding to the UAV attitude view obtained by inertial navigation.

[0032] Figure 7 This is a flowchart illustrating the method for determining target feature extraction and image matching models based on simulated aerial photographs and UAV aerial photographs provided in this application embodiment.

[0033] Figure 8 This is a flowchart illustrating the method for cross-modal matching of UAV aerial images and simulated aerial images using a robust dense matching model, as provided in this application embodiment.

[0034] Figure 9 This is a flowchart illustrating the method for cross-modal matching of UAV aerial images and simulated aerial images using a semi-dense local feature matching model, as provided in this application embodiment.

[0035] Figure 10 This is a flowchart illustrating the method for determining the homography transformation matrix based on target feature extraction and image matching model provided in this application embodiment.

[0036] Figure 11 This is a flowchart illustrating a method for transforming the pixel coordinates of a point of interest to the pixel coordinates of a satellite reference image based on a homography transformation matrix and a perspective transformation matrix, as provided in an embodiment of this application.

[0037] Figure 12 This is a flowchart illustrating another adaptive variable structure UAV multimodal scene matching navigation and positioning method provided in this application embodiment.

[0038] Figure 13 This is a schematic diagram of a matched image completed using the method provided in the embodiments of this application.

[0039] Figure 14 This is a schematic diagram of an adaptive variable structure UAV multimodal scene matching navigation and positioning device provided in an embodiment of this application.

[0040] Figure 15 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0041] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0042] The following will describe in detail, with reference to the accompanying drawings, an adaptive variable structure unmanned aerial vehicle (UAV) multimodal scene matching navigation and positioning method and apparatus according to embodiments of this application.

[0043] As mentioned above, traditional target localization methods mainly rely on the UAV's own GPS positioning information combined with camera calibration parameters for geometric projection calculations. However, under conditions of limited GPS signal, insufficient camera calibration accuracy, or large flight attitude measurement errors, the positioning accuracy and reliability are significantly reduced. Image-matching-based spatial positioning technology establishes a correspondence between pixels and real geographical locations by matching real-time UAV images with satellite reference images with precise geographic coordinates, offering advantages such as high positioning accuracy and strong anti-interference capabilities. However, existing image-matching-based positioning technologies still face challenges such as difficulties in compensating for viewpoint differences, limitations in multimodal image processing, and computational resource constraints.

[0044] In this context, the technical challenge is how to provide an image spatial positioning system that integrates multimodal processing, adaptive structural optimization, and high-precision matching, so that the image spatial positioning system has multimodal adaptive processing capabilities, hardware platform adaptability, and sub-pixel-level matching accuracy.

[0045] In view of this, this application provides an adaptive variable structure UAV multimodal scene matching navigation and positioning method. First, coordinate transformation is performed on the UAV pose information in the UAV aerial image to obtain the geographic coordinates corresponding to each pixel in the aerial image. Then, the UAV's geographic location with errors is corrected based on the difference between the pixel coordinates of the center point of the UAV aerial image and the geographic location corresponding to the camera center. Next, the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image are determined based on the corrected UAV geographic location. Then, a simulated aerial image from a specific UAV perspective is obtained through perspective transformation. Next, at least based on the simulated aerial image, a homography transformation matrix is ​​determined. Based on the homography transformation matrix and the perspective transformation matrix, the satellite reference image pixel coordinates of the points of interest in the aerial image are determined. Finally, the latitude and longitude of the points of interest are determined. This method overcomes the limitations of traditional methods in terms of perspective difference compensation, computational resource constraints, and cross-modal processing, significantly improving the system's versatility, robustness, and positioning accuracy.

[0046] Figure 1This is a flowchart illustrating an adaptive variable structure UAV multimodal scene matching navigation and positioning method provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps:

[0047] In step S101, an aerial photograph of the UAV is acquired, and the UAV pose information in the aerial photograph is used to perform coordinate transformation to obtain the geographic coordinates corresponding to each pixel in the aerial photograph.

[0048] In step S102, the difference information between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center is determined, and the difference information is used to correct the UAV geographical location that contains errors obtained through low-precision inertial navigation.

[0049] In step S103, the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image are determined based on the corrected UAV geographical location. Perspective transformation is performed on the satellite reference image to obtain the simulated aerial image under the specific UAV perspective obtained by the inertial navigation system.

[0050] In step S104, a target feature extraction and image matching model is determined based on simulated aerial images and drone aerial images.

[0051] In step S105, the homography transformation matrix is ​​determined based on the target feature extraction and image matching model.

[0052] In step S106, the pixel coordinates of the point of interest in the drone aerial image are obtained, and the pixel coordinates of the point of interest are transformed to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix.

[0053] In step S107, the latitude and longitude of the point of interest are determined based on the pixel coordinates of the satellite reference image.

[0054] In some embodiments of this application, the method can be executed by a server or by a terminal device with certain processing capabilities. It is used to correct the latitude and longitude of a drone with errors by using the geographical location information corresponding to the center pixel coordinates of the drone aerial image. It utilizes prior drone pose information to solve for the transformation matrix, performs a first-level perspective transformation on the satellite reference image to simulate aerial scenes from a specific drone perspective, determines the modality of the drone aerial image, and enters image matching subprocesses based on different models to achieve multimodal processing. It automatically detects the processor type and selects the corresponding neural network structure processing module to perform multi-level feature extraction and matching operations to achieve adaptive variable structure functionality. It refines the matching accuracy step by step to obtain sub-pixel-level matching and a second-level homography transformation matrix between the real and simulated aerial scenes. Finally, it uses two-level inverse transformations to convert pixels in the real-time image to the satellite reference image, thereby using the latitude and longitude information of the satellite image to obtain the geographic location of the pixels for navigation and positioning.

[0055] In some embodiments of this application, an aerial photograph of the drone can be obtained first, and the drone pose information in the aerial photograph can be used to perform coordinate transformation to obtain the geographic coordinates corresponding to each pixel in the aerial photograph.

[0056] In some embodiments of this application, the difference information between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center can be determined, and the difference information can be used to correct the UAV geographical location that contains errors obtained through low-precision inertial navigation.

[0057] In some embodiments of this application, the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial photograph can be determined based on the corrected UAV geographical location, and perspective transformation can be performed on the satellite reference image to obtain a simulated aerial photograph from a specific UAV perspective obtained by the inertial navigation system.

[0058] In some embodiments of this application, a target feature extraction and image matching model can be determined based on simulated aerial images and UAV aerial images, and a homography transformation matrix can be determined based on the target feature extraction and image matching model.

[0059] In some embodiments of this application, the pixel coordinates of points of interest in the drone aerial image can be obtained, and the pixel coordinates of the points of interest can be transformed to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix. Then, the latitude and longitude of the points of interest can be determined based on the pixel coordinates of the satellite reference image.

[0060] According to the technical solution provided in the embodiments of this application, the geographical coordinates of each pixel in the aerial image are obtained by first performing coordinate transformation on the pose information of the UAV in the aerial image. Then, the geographical location of the UAV with errors is corrected based on the difference between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center. Then, the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image are determined based on the corrected UAV geographical location. Then, a simulated aerial image from a specific UAV perspective is obtained through perspective transformation. Next, at least based on the simulated aerial image, the homography transformation matrix is ​​determined. Based on the homography transformation matrix and the perspective transformation matrix, the satellite reference image pixel coordinates of the points of interest in the aerial image are determined. Finally, the latitude and longitude of the points of interest are determined. This solves the limitations of traditional methods in terms of perspective difference compensation, computational resource constraints, and cross-modal processing, and significantly improves the versatility, robustness, and positioning accuracy of the system.

[0061] Figure 2 This is a flowchart illustrating a method for coordinate transformation using drone pose information from drone aerial photographs, as provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps:

[0062] In step S201, the pixel coordinates of each pixel in the drone aerial image are converted into two-dimensional image coordinates through affine transformation.

[0063] In step S202, the principle of similar triangles is applied to the first and second planes of the two-dimensional image coordinates, respectively, and perspective projection is used to convert the two-dimensional image coordinates of each pixel into three-dimensional camera coordinates.

[0064] In step S203, a rotation matrix from the camera coordinate system to the geographic coordinate system is constructed using the UAV attitude angle in the UAV pose information. The rotation matrix is ​​then used to convert the three-dimensional camera coordinates of each pixel into geographic coordinates through rigid body transformation.

[0065] In some embodiments of this application, coordinate transformation using UAV pose information from UAV aerial images can be performed by first converting each pixel in the UAV aerial image from pixel coordinates to two-dimensional image coordinates through affine transformation. Then, applying the principle of similar triangles to the first and second planes of the two-dimensional image coordinates, perspective projection is used to convert the two-dimensional image coordinates of each pixel into three-dimensional camera coordinates. Finally, a rotation matrix from the camera coordinate system to the geographic coordinate system is constructed using the UAV attitude angles in the UAV pose information. This rotation matrix is ​​then used to convert the three-dimensional camera coordinates of each pixel into geographic coordinates through rigid body transformation.

[0066] In other words, pixel coordinates in an aerial photograph can first be converted to image coordinates using affine transformation. This conversion involves not only numerical values ​​but also unit conversions. The conversion from pixel coordinates to image coordinates can be expressed by the formula... It means that, among them, These are the pixel coordinates in the aerial image. The coordinates of the transformed image. , They represent and The distance corresponding to a unit pixel in the direction. , These represent the pixel coordinates of the principal point.

[0067] Then, the principle of similar triangles is applied to the first and second planes of the two-dimensional image coordinates, respectively, and perspective projection is used to transform the two-dimensional image coordinates into three-dimensional camera coordinates. The first and second planes of these two-dimensional image coordinates can be, for example, the XZ and YZ planes.

[0068] In other words, perspective projection can be used to transform two-dimensional image coordinates into three-dimensional camera coordinates. By applying the principle of similar triangles in the XZ and YZ planes respectively, the perspective transformation is obtained as follows:

[0069] ,in, These are the transformed camera coordinates. This refers to the camera's focal length.

[0070] Finally, a rotation matrix from the camera coordinate system to the geographic coordinate system is constructed using the UAV attitude angles, and the camera coordinates are transformed into geographic coordinates through rigid body transformation. The formula for rigid body transformation is as follows:

[0071] ,in, The converted geographic coordinates , and These are the drone's heading angle, pitch angle, and roll angle, respectively.

[0072] Figure 3 This is a flowchart illustrating the method for determining the difference between the pixel coordinates of the center point of an aerial photograph taken by a drone and the geographical location corresponding to the camera center, as provided in an embodiment of this application. Figure 3 As shown, the method includes the following steps:

[0073] In step S301, the point P corresponding to the center pixel of the UAV aerial image is obtained after being converted to geographic coordinates.

[0074] In step S302, the vector formed by the center point of the UAV aerial camera and point P is determined as the target vector.

[0075] In step S303, the angle between the target vector and the vertical direction of the geographic coordinate system is determined.

[0076] In step S304, the difference information is determined by multiplying the cosine of the height and angle of the UAV relative to the ground with the unit vector of the target vector direction.

[0077] This difference information includes at least the distance difference between the pixel coordinates of the center point of the drone aerial image and the geographical location corresponding to the camera center in the three directions of north, east, and ground.

[0078] In some embodiments of this application, when determining the difference between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center, the point P corresponding to the center pixel of the UAV aerial image after conversion to geographical coordinates can be obtained first, and the vector formed by the UAV aerial camera center point and point P can be determined as the target vector. Then, the angle between the target vector and the vertical direction of the geographical coordinate system is determined. Finally, the product of the quotient of the height of the UAV relative to the ground and the cosine of the angle with the target vector direction is determined as the difference between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center.

[0079] In other words, the drone's position can be reversed by using the difference between the center pixel coordinates of the drone aerial image and the geographical location corresponding to the camera center. This can be achieved by converting the center pixel coordinates of the drone aerial image to geographic coordinates using the aforementioned coordinate transformation. Then, assuming that the camera center point C and the corresponding point P in the geographic coordinate system formed a vector... The angle between the vector and the vertical direction is:

[0080] ,in, , Let be a norm function, representing The unit vector of direction. Unit vector in the vertical direction .

[0081] Since the distance between the drone's center point and the camera's center point is negligible compared to their distance relative to the ground, the drone's center point and the camera's center point can be considered the same point in computational space. The difference in geographic coordinates between the real ground point corresponding to the center pixel of the aerial image and the drone's center point can be obtained using the direction vector. ,in This represents the distance difference in the three directions: north, east, and ground. This refers to the altitude of the drone relative to the ground.

[0082] Figure 4 This is a flowchart illustrating a method for correcting the geographical location of a UAV obtained through low-precision inertial navigation using difference information, as provided in an embodiment of this application. Figure 4 As shown, the method includes the following steps:

[0083] In step S401, the change in geographic coordinates of the ground point corresponding to the center pixel of the aerial image relative to the center of the UAV is determined.

[0084] In step S402, the latitude and longitude differences in the Earth coordinate system are calculated based on the changes.

[0085] In step S403, the calculated latitude and longitude of the ground point corresponding to the center pixel of the aerial image is determined based on the latitude and longitude coordinates of the UAV with errors.

[0086] In step S404, a satellite reference image is acquired. Based on the pixel coordinates of the center pixel of the aerial image on the satellite reference image and the prior latitude and longitude information of the satellite reference image, the true latitude and longitude of the center pixel of the aerial image is determined by linear interpolation.

[0087] In step S405, the error between the actual latitude and longitude and the calculated latitude and longitude is corrected to obtain the actual latitude and longitude of the UAV.

[0088] In some embodiments of this application, correcting the erroneous UAV geographical location obtained through low-precision inertial navigation using difference information can be achieved by first determining the change in geographic coordinates of the ground point corresponding to the center pixel of the aerial image relative to the center of the UAV. Then, based on the change, the corresponding latitude and longitude difference in the Earth coordinate system is calculated, and the calculated latitude and longitude of the ground point corresponding to the center pixel of the aerial image is determined based on the erroneous UAV latitude and longitude coordinates. Next, a satellite reference image is acquired, and the true latitude and longitude of the center pixel of the aerial image is determined through linear interpolation based on the pixel coordinates of the center pixel on the satellite reference image and the prior latitude and longitude information of the satellite reference image. Finally, the error between the true latitude and longitude and the calculated latitude and longitude is corrected to obtain the true latitude and longitude of the UAV.

[0089] In other words, the difference in latitude and longitude in the Earth coordinate system can be calculated by using the change in geographic coordinates between the real ground point corresponding to the center pixel of the aerial image and the center of the drone. ,in The radius of curvature of the Earth's meridian. , The radius of curvature of the Earth's geoid. , The radius of the Earth's equator. =6,378,137 meters For the Earth's eccentricity, =0.0818191908426, The latitude is the local latitude. This refers to the local altitude.

[0090] Then, based on the known errors in the drone's latitude and longitude coordinates, the latitude and longitude coordinates of the real ground point corresponding to the center pixel of the aerial image are calculated. ,in and These are the latitude and longitude coordinates of the drone, which currently contain errors.

[0091] Using the pixel coordinates of the center point of the real-time image on the reference image and the prior latitude and longitude information of the satellite reference image, the true latitude and longitude of the center point are obtained through linear interpolation. ,in and The satellite reference image is located at and Pixel resolution in the direction, , These refer to the longitude and latitude ranges covered by the satellite reference image, respectively. The minimum longitude included in the satellite reference map. This represents the maximum latitude included in the satellite baseline map.

[0092] The difference between the actual latitude and longitude of the ground point corresponding to the center pixel of the aerial image and the calculated latitude and longitude represents the error contained in the prior position information of the UAV, which can be used to correct the actual latitude and longitude of the UAV. ,in , which represents the error contained in the prior position information of the UAV.

[0093] Figure 5 This is a flowchart illustrating the method for determining the reference image pixel coordinates corresponding to the corner pixel coordinates of an aerial photograph based on the corrected geographic location of the UAV, as provided in this application embodiment. Figure 5 As shown, the method includes the following steps:

[0094] In step S501, the geographic coordinates of the corner pixels of the aerial image are obtained.

[0095] In step S502, the latitude and longitude coordinates of the ground points corresponding to the corner pixels of the aerial image are determined based on the actual latitude and longitude of the UAV.

[0096] In step S503, the pixel coordinates of the corner pixels of the aerial photograph are obtained in the satellite reference image by linear interpolation, based on the prior latitude and longitude information of the satellite reference image, and are used as the reference image pixel coordinates of the corner pixels of the reference aerial photograph.

[0097] In some embodiments of this application, determining the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image based on the corrected UAV geographical location can be achieved by first obtaining the geographic coordinates of the corner pixel in the aerial image, then determining the latitude and longitude coordinates of the ground point corresponding to the corner pixel in the aerial image based on the actual latitude and longitude of the UAV, and finally obtaining the pixel coordinates of the corner pixel in the satellite reference image through linear interpolation using the prior latitude and longitude information of the satellite reference image. These coordinates are then used as the reference image pixel coordinates of the corner pixel in the reference aerial image.

[0098] In other words, calculating the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image based on the corrected UAV geographic location can be done by first using the aforementioned coordinate transformation to convert the aerial image... Figure 4 The pixel coordinates of each corner point are converted to a geographic coordinate system. Then, using a vector method similar to that used to calculate the latitude and longitude deviation between the pixel center point and the UAV center point, the latitude and longitude coordinates of the corresponding ground point are calculated based on the corrected UAV latitude and longitude. .

[0099] By combining the latitude and longitude information of the satellite reference image, the pixel coordinates of corner points in the real-time aerial image are obtained from the pixel coordinates of the reference image using linear interpolation. .

[0100] Figure 6This is a flowchart illustrating a method for obtaining a simulated aerial photograph from a specific UAV perspective by performing perspective transformation on a satellite reference image according to an embodiment of this application, based on the UAV attitude viewpoint acquired by inertial navigation. For example... Figure 6 As shown, the method includes the following steps:

[0101] In step S601, the reference image pixel coordinates of the corner pixel of the aerial image are combined with the pixel coordinates of the corner pixel of the aerial image to form a corner pixel pair.

[0102] In step S602, the perspective transformation matrix from the satellite reference image to the simulated aerial image is determined based on the corner pixel pairs.

[0103] In step S603, the reference image pixel coordinates of the corner pixels of the aerial image are converted into the corner pixel coordinates of the simulated aerial image using a perspective transformation matrix, thus obtaining a simulated aerial image from the perspective of the drone.

[0104] In some embodiments of this application, obtaining a simulated aerial image from a specific UAV perspective obtained by inertial navigation by performing perspective transformation on a satellite reference image can be achieved by first forming corner pixel pairs with the reference image pixel coordinates of the corner pixels in the aerial image and the pixel coordinates of the corner pixels in the aerial image. Then, a perspective transformation matrix from the satellite reference image to the simulated aerial image is determined based on the corner pixel pairs. Finally, the perspective transformation matrix is ​​used to convert the reference image pixel coordinates of the corner pixels in the aerial image to the corner pixel coordinates of the simulated aerial image, thus obtaining the simulated aerial image from the UAV perspective.

[0105] In other words, aerial photography can be used. Figure 4 The pixel coordinates of each corner point on the baseline map, combined with aerial photography. Figure 4 The original pixel coordinates of each corner point are used to form four point pairs. Then, these four point pairs are used to solve for the perspective transformation matrix of a specific area on the base map that is transformed into an aerial image containing eight unknowns. Applying this perspective transformation matrix to the satellite base map can generate a simulated aerial image from a specific drone perspective.

[0106] In one example, this perspective transformation matrix can be used to convert the coordinates of the corner pixels in an aerial photograph into the four corner points on a simulated aerial photograph. Since the simulated aerial photograph is an image of a specified size, this achieves the projection of an irregular quadrilateral region in the aerial photograph into an image of a specified size.

[0107] Figure 7 This is a flowchart illustrating the method for determining target feature extraction and image matching models based on simulated aerial photographs and UAV aerial photographs provided in this application. Figure 7 As shown, the method includes the following steps:

[0108] In step S701, the aerial image channel of the drone aerial image is detected.

[0109] In step S702, in response to determining that the drone aerial image includes one channel or three channels with the same value, the drone aerial image is determined to be an infrared image.

[0110] In step S703, a robust dense matching model is used to perform cross-modal matching between the UAV aerial image and the simulated aerial image.

[0111] In step S704, in response to determining that the drone aerial image includes three channels and that at least two of the three channels include different values, the drone aerial image is determined to be a visible light image.

[0112] In step S705, a semi-dense local feature matching model is used to perform cross-modal matching between the UAV aerial image and the simulated aerial image.

[0113] In some embodiments of this application, the target feature extraction and image matching model based on simulated aerial images and drone aerial images may first detect the aerial image channels of the drone aerial image. If the drone aerial image includes one channel, or includes three channels with the same value, then the drone aerial image can be determined to be an infrared image. In this case, a robust dense matching model can be used to perform cross-modal matching between the drone aerial image and the simulated aerial image.

[0114] Conversely, if a drone aerial image includes three channels, and at least two of the three channels contain different values, then the drone aerial image can be determined to be a visible light image. In this case, a semi-dense local feature matching model can be used to perform cross-modal matching between the drone aerial image and the simulated aerial image.

[0115] In other words, it can detect the channels of real aerial images to determine the image modality, and then use corresponding matching models according to different image modalities to match real aerial images taken by drones with simulated aerial images generated based on satellite reference images. By detecting the channels of aerial images, it distinguishes between infrared images and visible light images of different modalities. If an image has only one channel or three channels with the same value, it is determined to be an infrared image, and a robust dense matching model is used to perform cross-modal matching with the visible light simulated aerial image. If an image has three channels with different values, it is determined to be an infrared image, and a high-efficiency semi-dense local feature matching model is used to perform same-modal matching with the visible light simulated aerial image.

[0116] Figure 8 This is a flowchart illustrating the method for cross-modal matching of UAV aerial images and simulated aerial images using a robust dense matching model, as provided in an embodiment of this application. Figure 8 As shown, the method includes the following steps:

[0117] In step S801, the encoder-decoder system is obtained.

[0118] The encoder integrates a Convolutional Neural Network (CNN) and a DINOv2 (Dual-Stage Implicit Object-Oriented Network) feature extractor to construct a multi-scale feature pyramid, which includes resolutions from 1x to 16x. The Gaussian process module of the decoder calculates the feature similarity matrix through a cosine kernel function and solves the Bayesian posterior distribution to obtain the mean and covariance information of the matching.

[0119] In step S802, a coarse-to-fine multi-scale matching strategy is adopted. Starting from 16x downsampling, a coarse matching stream is generated at each scale through an embedding decoder. Then, a convolutional refiner is used to perform pixel-level optimization by combining local correlation calculation, displacement embedding, and bilinear interpolation.

[0120] In step S803, the optimized features are subjected to layer-by-layer upsampling and feature fusion processing, and threshold balanced sampling and kernel density estimation are used to obtain spatially uniform matching points as cross-modal matching results.

[0121] The cross-modal matching results also include sub-pixel-level dense correspondences and confidence levels.

[0122] In some embodiments of this application, cross-modal matching of UAV aerial images and simulated aerial images using a robust dense matching model can be achieved by first obtaining an encoder-decoder system. The encoder integrates a CNN and a DINOv2 feature extractor to construct a multi-scale feature pyramid, which includes resolutions from 1x to 16x. The Gaussian process module of the decoder calculates the feature similarity matrix using a cosine kernel function and solves the Bayesian posterior distribution to obtain the mean and covariance information of the matching.

[0123] Then, a coarse-to-fine multi-scale matching strategy is adopted. Starting from 16x downsampling, a coarse matching stream is generated at each scale through an embedding decoder. Then, a convolutional refiner is used to perform pixel-level optimization by combining local correlation calculation, displacement embedding and bilinear interpolation.

[0124] Finally, the optimized features are subjected to layer-by-layer upsampling and feature fusion processing, and threshold balanced sampling and kernel density estimation are used to obtain spatially uniform matching points as cross-modal matching results.

[0125] In other words, the robust dense cross-modal matching model can employ an encoder-decoder architecture to achieve dense image matching. The encoder integrates CNN and DINOv2 feature extractors to construct a multi-scale feature pyramid with resolutions ranging from 1× to 16×. The Gaussian process module of the decoder calculates the feature similarity matrix using a cosine kernel function and solves for the Bayesian posterior distribution to obtain the mean and covariance information of the matches. A coarse-to-fine multi-scale matching strategy can be adopted. Starting with 16x downsampling, a coarse matching stream is first generated at each scale through an embedding decoder, followed by pixel-level optimization using a convolutional refiner combined with local correlation calculation, shift embedding, and bilinear interpolation. Layer-by-layer upsampling and feature fusion, along with threshold-balanced sampling and kernel density estimation, ensure a uniform spatial distribution of matching points, ultimately outputting sub-pixel-level dense correspondences and confidence scores.

[0126] Figure 9 This is a flowchart illustrating a method for cross-modal matching of UAV aerial images and simulated aerial images using a semi-dense local feature matching model, as provided in an embodiment of this application. Figure 9 As shown, the method includes the following steps:

[0127] In step S901, multi-scale feature extraction at 1 / 2, 1 / 4, and 1 / 8 resolution is performed on the UAV aerial photograph and the simulated aerial photograph.

[0128] Among them, the 1 / 8 resolution features are used as coarse matching input, while the 1 / 2 and 1 / 4 resolution features are retained for subsequent fine processing.

[0129] In step S902, local aggregation is performed on the extracted 1 / 8 resolution coarse-grained feature map to remove redundant information, and global context enhancement is performed through alternating self-attention and cross-attention modules to obtain enhanced UAV aerial image features and enhanced simulated aerial image features.

[0130] In step S903, the similarity matrix between the enhanced UAV aerial image features and the enhanced simulated aerial image features is calculated. An initial matching relationship is established through a multi-level screening strategy. The initial matching relationship is used to quickly identify possible matching points as coarse matching points in the global scope.

[0131] In step S904, a feature pyramid structure is used to gradually fuse features at 1 / 8, 1 / 4, and 1 / 2 resolutions to the original resolution through residual connection and upsampling operations.

[0132] In step S905, pixel-level localization is performed within the local window extracted around the coarse matching point, and the confidence is calculated by bidirectional softmax to find the optimal matching position.

[0133] In step S906, subpixel regression is performed within a preset area around the optimal matching position to convert the heatmap into continuous subpixel coordinate offsets to obtain cross-modal matching results.

[0134] The cross-modal matching results also include sub-pixel level coordinate information and confidence level.

[0135] In some embodiments of this application, cross-modal matching of UAV aerial images and simulated aerial images using a semi-dense local feature matching model can be performed by first extracting multi-scale features at 1 / 2, 1 / 4, and 1 / 8 resolutions from the UAV aerial images and simulated aerial images. The 1 / 8 resolution features are used as coarse matching inputs, while the 1 / 2 and 1 / 4 resolution features are retained for subsequent fine processing.

[0136] Then, local aggregation is performed on the extracted 1 / 8 resolution coarse-grained feature map to remove redundant information, and global context enhancement is performed through alternating self-attention and cross-attention modules to obtain enhanced UAV aerial image features and enhanced simulated aerial image features.

[0137] Next, the similarity matrix between the enhanced drone aerial image features and the enhanced simulated aerial image features is calculated. An initial matching relationship is established through a multi-level screening strategy. The initial matching relationship is used to quickly identify possible matching points as coarse matching points in the global scope.

[0138] In one example, a multi-level filtering strategy may include confidence threshold filtering, boundary mask filtering, nearest neighbor filtering, and effective match extraction. Confidence threshold filtering filters out candidate matches with too low confidence; boundary mask filtering removes unreliable matches near image boundaries; nearest neighbor filtering ensures bidirectional consistency of matches, excluding unstable one-to-many or many-to-one matches; and effective match extraction extracts the final matching point pairs from the filtered masks. The goal of this multi-level filtering strategy is to extract sparse but high-quality matching point pairs from a dense similarity matrix.

[0139] Next, a feature pyramid structure is used to gradually fuse features at 1 / 8, 1 / 4, and 1 / 2 resolutions into the original resolution through residual connections and upsampling operations. Pixel-level localization is then performed within the local window extracted around the coarse matching point. Confidence is calculated using bidirectional softmax to find the optimal matching position.

[0140] Finally, subpixel regression is performed within a preset area around the optimal matching position to convert the heatmap into continuous subpixel coordinate offsets to obtain cross-modal matching results.

[0141] In other words, the high-efficiency semi-dense local feature matching model for same-modality matching can use CNN to extract features from two input images at multiple scales. On the extracted coarse-grained feature maps, local aggregation is performed to remove redundant information, and then global context enhancement is performed through alternating self-attention and cross-attention modules. In the coarse matching stage, the similarity matrix between the features of the two images is calculated and an initial matching relationship is established. In the fine preprocessing stage, the FPN feature pyramid structure is used to gradually fuse features of different resolutions to the original resolution through residual connections and upsampling operations. In the fine matching stage, a two-stage optimization strategy is adopted. In the first stage, pixel-level localization is performed within the local window extracted around the coarse matching point. In the second stage, sub-pixel regression is performed in the 3×3 region around the optimal point. The final output matching result contains sub-pixel level coordinate information and confidence.

[0142] Figure 10 This is a flowchart illustrating the method for determining the homography transformation matrix based on target feature extraction and image matching models provided in this application. Figure 10 As shown, the method includes the following steps:

[0143] In step S1001, the processor architecture is detected to determine the processor type.

[0144] In step S1002, a preset weight file is loaded.

[0145] The weight file includes at least one of the following: coarse feature extraction weight file, fine-grained feature extraction weight file, coarse feature matching weight file, and layer-by-layer network refinement weight file.

[0146] In step S1003, the weight file is inferred and matched based on the processor type to obtain the scene matching result.

[0147] In step S1004, for each matching point in the cross-modal matching results, the homography transformation matrix from the simulated aerial image to the UAV aerial image is determined using the random sample consensus algorithm based on the scene matching results.

[0148] In some embodiments of this application, determining the homography transformation matrix based on the target feature extraction and image matching model may involve first detecting the processor architecture and determining the processor type, then loading a preset weight file. Inference matching is then performed on the weight file based on the processor type to obtain scene matching results. Finally, for each matching point in the cross-modal matching results, the random sample consensus (RANSAC) algorithm is used to determine the homography transformation matrix from the simulated aerial image to the UAV aerial image based on the scene matching results.

[0149] In other words, the processor architecture can be automatically detected to modify the network structure of feature extraction and image matching in the scene matching system to obtain the homography transformation matrix for scene matching. The scene matching model is compatible with multi-processor architectures. By detecting the processor type, the corresponding weight file is loaded for inference matching. Each step of coarse feature extraction, fine-grained feature extraction, coarse feature matching, and layer-by-layer network refinement corresponds to multiple types of weight files. Among them, RKNN is for NPU processors, ONNX is for high-efficiency CPU processors, PyTorch-CPU is for CPU processors, and PyTorch-GPU is for GPU processors. After the scene matching process is completed, the matching point pairs are obtained, and the homography transformation matrix from the simulated aerial image to the aerial image is obtained through the RANSAC algorithm.

[0150] Figure 11 This is a flowchart illustrating a method for transforming the pixel coordinates of a point of interest to the pixel coordinates of a satellite reference image based on a homography transformation matrix and a perspective transformation matrix, as provided in an embodiment of this application. Figure 11 As shown, the method includes the following steps:

[0151] In step S1101, the coordinates of the points of interest in the UAV aerial photograph are mapped to the simulated aerial photograph based on the inverse matrix of the homography transformation matrix.

[0152] In step S1102, the coordinates of the point of interest in the simulated aerial image are converted to coordinates in the satellite reference image using a perspective transformation matrix.

[0153] The perspective transformation matrix is ​​determined when performing perspective transformation on the satellite reference image.

[0154] In step S1103, the latitude and longitude of the point of interest are obtained by linear interpolation based on the latitude and longitude information of the corner points in the known information of the satellite reference image.

[0155] In some embodiments of this application, transforming the pixel coordinates of a point of interest (POI) to the pixel coordinates of a satellite reference image based on a homography transformation matrix and a perspective transformation matrix can be achieved by first mapping the POI coordinates in the UAV aerial image to a simulated aerial image using the inverse of the homography transformation matrix; then, using the perspective transformation matrix to convert the coordinates of the POI in the simulated aerial image to the coordinates in the satellite reference image; and finally, obtaining the latitude and longitude of the POI through linear interpolation based on the corner latitude and longitude information from the known information in the satellite reference image.

[0156] In other words, aerial image pixels can be matched to a satellite reference image through a two-stage inverse transformation, and the Earth coordinate system position of the pixels can be calculated using the prior latitude and longitude of the reference image. In one example, a first-stage inverse transformation can be performed using the solved homography inverse matrix from the simulated aerial image to the aerial image to map the coordinates of the point of interest in the aerial image to the coordinates in the simulated aerial image. Then, a second-stage inverse transformation can be performed using the solved perspective transformation inverse matrix from the reference image to the simulated aerial image to transform the coordinates in the simulated aerial image to the coordinates in the reference image. Finally, the latitude and longitude of the point of interest can be obtained by linear interpolation using the corner latitude and longitude information in the prior information of the reference image.

[0157] Figure 12 This is a flowchart illustrating another adaptive variable structure UAV multimodal scene matching navigation and positioning method provided in this application embodiment. Figure 12 As shown, the drone pose with errors can be obtained, the geographical location can be obtained by transforming the coordinates of the center pixel of the aerial image, and the drone's latitude and longitude with errors can be corrected by using the geographical location information corresponding to the center pixel coordinates of the drone aerial image. After completing the drone's latitude and longitude correction, the corner pixel coordinates of the aerial image are transformed to obtain the corresponding latitude and longitude, and then the perspective transformation matrix M is calculated using the obtained latitude and longitude.

[0158] A satellite reference image can be acquired, and the transformation matrix can be solved using the prior information of the UAV's pose to perform a first-order perspective transformation on the satellite reference image, simulating an aerial scene from a specific UAV perspective, thus obtaining a simulated aerial image. Based on this simulated aerial image, the modality of the actual UAV aerial image is determined, and accordingly, image matching subprocesses based on different models are initiated to achieve multimodal processing.

[0159] On the one hand, if the real and simulated aerial images have the same modality, a high-efficiency semi-dense local feature matching model is used for processing. On the other hand, if the real and simulated aerial images have different modalities, a robust dense matching model is used for processing.

[0160] It can automatically detect the processor type and select the corresponding neural network structure processing module to perform multi-level feature extraction and matching operations to achieve adaptive variable structure functionality. For example, if the processor is a Neural Network Processing Unit (NPU), a dedicated neural network (Rockchip Neural Network, RKNN) model can be selected; if the processor is not an NPU but a Graphics Processing Unit (GPU), a PyTorch-GPU model can be selected; if the processor is neither an NPU nor a GPU, an Open Neural Network Exchange (ONNX) model and a PyTorch-CPU model can be selected. All of these models can first perform coarse feature extraction, then fine-grained feature extraction, followed by coarse feature matching, and finally achieve layer-by-layer network refinement.

[0161] Subpixel-level matching and a second-level homography transformation matrix H between real and simulated aerial images can be obtained by progressively refining the matching accuracy. For points of interest in the aerial image, their pixel coordinates can be acquired. Based on the perspective transformation matrix M and the homography transformation matrix H, the pixel coordinates of the point of interest are transformed to the pixel coordinates of the satellite reference image through two-level inverse transformations. Then, based on the location information of the satellite image, the latitude and longitude corresponding to the point of interest, i.e., the geographical location of the point of interest, can be calculated to achieve navigation and positioning.

[0162] Figure 13 This is a schematic diagram of a matched image completed using the method provided in the embodiments of this application. Figure 13 As shown, image (a) is a real-time scene image taken by a UAV in a specific pose. Image (b) is a satellite reference map, with the latitude and longitude information of its lower left and upper right corners known. Image (c) is a simulated aerial scene image from the UAV's perspective, obtained by perspective transformation of the satellite reference map. Image (d) is a benchmark aerial image obtained by homography transformation of the simulated aerial scene image after matching with a dense or semi-dense image matching model.

[0163] The technical solution provided in this application uses UAV pose information to perform perspective transformation on satellite images to generate simulated aerial images. It autonomously detects image modalities and selects corresponding matching algorithms. Based on the processor type, it adaptively adjusts the neural network structure to perform feature extraction and matching, achieving sub-pixel-level image matching and homography transformation. Finally, through two-stage inverse transformation, it accurately maps the pixels in the UAV image to the geographic coordinate system of the satellite reference image, thereby obtaining the precise geographic location information of the pixels.

[0164] The technical solution provided in this application can automatically identify image modalities and select corresponding matching algorithms to achieve multimodal adaptive processing. It can dynamically adjust the neural network structure according to the processor type to achieve hardware adaptive optimization. It adopts a two-level transformation mechanism and a multi-level feature extraction strategy to achieve sub-pixel-level matching accuracy. It effectively solves the limitations of traditional methods in terms of viewpoint difference compensation, computational resource constraints and cross-modal processing, and significantly improves the system's versatility, robustness and positioning accuracy.

[0165] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0166] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0167] Figure 14 This is a schematic diagram of an adaptive variable structure UAV multimodal scene matching navigation and positioning device provided in an embodiment of this application. Figure 14 As shown, the device includes:

[0168] The acquisition module 1401 is configured to acquire drone aerial images, and use the drone pose information in the drone aerial images to perform coordinate transformation to obtain the geographic coordinates corresponding to each pixel in the aerial images.

[0169] The correction module 1402 is configured to determine the difference information between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center, and use the difference information to correct the UAV geographical location that has errors obtained by low-precision inertial navigation.

[0170] The simulation module 1403 is configured to determine the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image based on the corrected UAV geographical location, and perform perspective transformation on the satellite reference image to obtain a simulated aerial image under the specific UAV perspective obtained by the inertial navigation system.

[0171] Select module 1404 is configured to determine a target feature extraction and image matching model based on simulated aerial images and drone aerial images.

[0172] The determination module 1405 is configured to determine the homography transformation matrix based on the target feature extraction and image matching model.

[0173] The inverse transformation module 1406 is configured to acquire the pixel coordinates of the point of interest in the drone aerial image and transform the pixel coordinates of the point of interest to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix.

[0174] The positioning module 1407 is configured to determine the latitude and longitude of a point of interest based on the pixel coordinates of a satellite reference image.

[0175] According to the technical solution provided in the embodiments of this application, the geographical coordinates of each pixel in the aerial image are obtained by first performing coordinate transformation on the pose information of the UAV in the aerial image. Then, the geographical location of the UAV with errors is corrected based on the difference between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center. Then, the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image are determined based on the corrected UAV geographical location. Then, a simulated aerial image from a specific UAV perspective is obtained through perspective transformation. Next, at least based on the simulated aerial image, the homography transformation matrix is ​​determined. Based on the homography transformation matrix and the perspective transformation matrix, the satellite reference image pixel coordinates of the points of interest in the aerial image are determined. Finally, the latitude and longitude of the points of interest are determined. This solves the limitations of traditional methods in terms of perspective difference compensation, computational resource constraints, and cross-modal processing, and significantly improves the versatility, robustness, and positioning accuracy of the system.

[0176] In some implementations, coordinate transformation is performed using the drone pose information in the drone aerial image, including: converting each pixel in the drone aerial image from pixel coordinates to two-dimensional image coordinates through affine transformation; applying the principle of similar triangles to the first and second planes of the two-dimensional image coordinates respectively, and using perspective projection to convert the two-dimensional image coordinates of each pixel into three-dimensional camera coordinates; constructing a rotation matrix from the camera coordinate system to the geographic coordinate system using the drone attitude angle in the drone pose information, and using the rotation matrix to convert the three-dimensional camera coordinates of each pixel into geographic coordinates through rigid body transformation.

[0177] In some implementations, determining the difference information between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center includes: obtaining the point P corresponding to the center pixel of the UAV aerial image after conversion to geographical coordinates; determining that the vector formed by the UAV aerial camera center point and point P is the target vector; determining the angle between the target vector and the vertical direction of the geographical coordinate system; determining the product of the quotient of the height of the UAV relative to the ground and the cosine of the angle with respect to the target vector and the unit vector of the target vector direction as the difference information; wherein, the difference information includes at least the distance difference between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center in the north, east, and ground directions.

[0178] In some implementations, the difference information is used to correct the erroneous UAV geographical location obtained through low-precision inertial navigation, including: determining the change in geographic coordinates of the ground point corresponding to the center pixel of the aerial image relative to the center of the UAV; calculating the corresponding latitude and longitude difference in the Earth coordinate system based on the change; determining the calculated latitude and longitude of the ground point corresponding to the center pixel of the aerial image based on the erroneous UAV latitude and longitude coordinates; acquiring a satellite reference image, determining the true latitude and longitude of the center pixel of the aerial image through linear interpolation based on the pixel coordinates of the center pixel of the aerial image on the satellite reference image and the prior latitude and longitude information of the satellite reference image; correcting the error between the true latitude and longitude and the calculated latitude and longitude to obtain the true latitude and longitude of the UAV.

[0179] In some implementations, the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image are determined based on the corrected UAV geographical location. This includes: obtaining the geographic coordinates of the corner pixel of the aerial image; determining the latitude and longitude coordinates of the ground point corresponding to the corner pixel of the aerial image based on the actual latitude and longitude of the UAV; and obtaining the pixel coordinates of the corner pixel of the aerial image in the satellite reference image through linear interpolation by combining the prior latitude and longitude information of the satellite reference image, which are then used as the reference image pixel coordinates of the corner pixel of the reference aerial image.

[0180] In some implementations, a perspective transformation is performed on a satellite reference image to obtain a simulated aerial image from a specific UAV perspective, corresponding to the UAV attitude view acquired by the inertial navigation system. This includes: forming corner pixel pairs by combining the reference image pixel coordinates of the corner pixels of the aerial image with the pixel coordinates of the corner pixels of the aerial image; determining a perspective transformation matrix from the satellite reference image to the simulated aerial image based on the corner pixel pairs; and using the perspective transformation matrix to convert the reference image pixel coordinates of the corner pixels of the aerial image into the corner pixel coordinates of the simulated aerial image, thereby obtaining a simulated aerial image from the UAV perspective.

[0181] In some implementations, the target feature extraction and image matching model is determined based on simulated aerial images and drone aerial images, including: detecting the aerial image channels of the drone aerial image; determining that the drone aerial image is an infrared image in response to determining that the drone aerial image includes one channel or three channels with the same value; performing cross-modal matching between the drone aerial image and the simulated aerial image using a robust dense matching model; determining that the drone aerial image is a visible light image in response to determining that the drone aerial image includes three channels, and that at least two of the three channels include different values; and performing cross-modal matching between the drone aerial image and the simulated aerial image using a semi-dense local feature matching model.

[0182] In some implementations, a robust dense matching model is used to perform cross-modal matching between UAV aerial images and simulated aerial images. This includes: acquiring an encoder-decoder system; the encoder fusing CNN and DINOv2 feature extractors to construct a multi-scale feature pyramid, which includes resolutions from 1x to 16x; the decoder's Gaussian process module calculating the feature similarity matrix using a cosine kernel function to solve for the Bayesian posterior distribution and obtain the mean and covariance information of the matches; employing a coarse-to-fine multi-scale matching strategy, starting with 16x downsampling, generating a coarse matching stream at each scale through embedding the decoder, and then using a convolutional refiner combined with local correlation calculation, displacement embedding, and bilinear interpolation for pixel-level optimization; performing layer-by-layer upsampling and feature fusion processing on the optimized features, and using threshold balanced sampling and kernel density estimation to obtain spatially uniform matching points as the cross-modal matching results; the cross-modal matching results also include sub-pixel-level dense correspondences and confidence scores.

[0183] In some implementations, a semi-dense local feature matching model is used to perform cross-modal matching between UAV aerial images and simulated aerial images. This includes: extracting multi-scale features at 1 / 2, 1 / 4, and 1 / 8 resolutions from both UAV and simulated aerial images, where the 1 / 8 resolution features are used as coarse matching input, and the 1 / 2 and 1 / 4 resolution features are retained for subsequent fine-grained processing; performing local aggregation on the extracted 1 / 8 resolution coarse-grained feature maps to remove redundant information, and performing global context enhancement through alternating self-attention and cross-attention modules to obtain enhanced UAV aerial image features and enhanced simulated aerial image features; and calculating the enhanced UAV aerial image features and enhanced simulated aerial image features. The similarity matrix between image features is used to establish initial matching relationships through a multi-level screening strategy. These initial matching relationships are then used to quickly identify possible matching points globally as coarse matching points. A feature pyramid structure is employed to gradually fuse features at 1 / 8, 1 / 4, and 1 / 2 resolutions to the original resolution through residual connections and upsampling operations. Pixel-level localization is performed within local windows extracted around the coarse matching points, and confidence is calculated using bidirectional softmax to find the optimal matching position. Sub-pixel regression is performed within a preset region around the optimal matching position, converting the heatmap into continuous sub-pixel coordinate offsets to obtain cross-modal matching results. The cross-modal matching results also include sub-pixel coordinate information and confidence.

[0184] In some implementations, determining the homography transformation matrix based on the target feature extraction and image matching model includes: detecting the processor architecture and determining the processor type; loading a preset weight file; the weight file includes at least one of the following: a coarse feature extraction weight file, a fine-grained feature extraction weight file, a coarse feature matching weight file, and a layer-by-layer network refinement weight file; performing inference matching on the weight file based on the processor type to obtain scene matching results; for each matching point in the cross-modal matching results, using the Random Sample Consensus (RANSAC) algorithm based on the scene matching results to determine the homography transformation matrix from the simulated aerial image to the UAV aerial image.

[0185] In some implementations, the pixel coordinates of the point of interest are transformed to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix. This includes: mapping the coordinates of the point of interest in the UAV aerial image to the simulated aerial image based on the inverse matrix of the homography transformation matrix; converting the coordinates of the point of interest in the simulated aerial image to the coordinates in the satellite reference image using the perspective transformation matrix; the perspective transformation matrix is ​​determined when performing perspective transformation on the satellite reference image; and obtaining the latitude and longitude of the point of interest by linear interpolation based on the latitude and longitude information of corner points in the known information of the satellite reference image.

[0186] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0187] Figure 15 This is a schematic diagram of the electronic device provided in an embodiment of this application. For example... Figure 15 As shown, the electronic device 15 of this embodiment includes: a processor 1501, a memory 1502, and a computer program 1503 stored in the memory 1502 and executable on the processor 1501. When the processor 1501 executes the computer program 1503, it implements the steps in the various method embodiments described above. Alternatively, when the processor 1501 executes the computer program 1503, it implements the functions of each module / unit in the various device embodiments described above.

[0188] Electronic device 15 may be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 15 may include, but is not limited to, processor 1501 and memory 1502. Those skilled in the art will understand that... Figure 15 This is merely an example of electronic device 15 and does not constitute a limitation on electronic device 15. It may include more or fewer components than shown, or different components.

[0189] The processor 1501 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0190] The memory 1502 can be an internal storage unit of the electronic device 15, such as a hard disk or RAM of the electronic device 15. The memory 1502 can also be an external storage device of the electronic device 15, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, FlashCard, etc., equipped on the electronic device 15. The memory 1502 can also include both internal and external storage units of the electronic device 15. The memory 1502 is used to store computer programs and other programs and data required by the electronic device.

[0191] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0192] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0193] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. An adaptive variable structure UAV multimodal scene matching navigation and positioning method, characterized in that, include: Aerial images of drones are acquired, and the drone pose information in the aerial images is used to perform coordinate transformation to obtain the geographic coordinates of each pixel in the aerial images. Determine the difference information between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center, and use the difference information to correct the UAV geographical location that contains errors obtained through low-precision inertial navigation. Based on the corrected UAV geographical location, the corresponding reference image pixel coordinates of the corner pixel coordinates of the aerial image are determined. Perspective transformation is performed on the satellite reference image to obtain the simulated aerial image under the specific UAV attitude view obtained by the inertial navigation. Based on the simulated aerial images and UAV aerial images, a target feature extraction and image matching model is determined; The homography transformation matrix is ​​determined based on the target feature extraction and image matching model. Obtain the pixel coordinates of the point of interest in the drone aerial image, and transform the pixel coordinates of the point of interest to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix; The latitude and longitude of the point of interest are determined based on the pixel coordinates of the satellite reference image.

2. The method according to claim 1, characterized in that, Coordinate transformation is performed using the drone pose information from the drone aerial image, including: The pixel coordinates of each pixel in the drone aerial image are converted into two-dimensional image coordinates through affine transformation. The principle of similar triangles is applied to the first and second planes of the two-dimensional image coordinates, and perspective projection is used to convert the two-dimensional image coordinates of each pixel into three-dimensional camera coordinates. A rotation matrix from the camera coordinate system to the geographic coordinate system is constructed using the UAV attitude angles in the UAV pose information. The rotation matrix is ​​then used to convert the 3D camera coordinates of each pixel into geographic coordinates through rigid body transformation.

3. The method according to claim 1, characterized in that, Determining the difference between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center includes: Obtain the point P corresponding to the center pixel of the drone aerial image after converting it to geographic coordinates; The vector formed by the center point of the UAV aerial camera and the point P is determined as the target vector; Determine the angle between the target vector and the vertical direction of the geographic coordinate system; The difference information is determined by multiplying the quotient of the drone's altitude relative to the ground and the cosine of the included angle with the unit vector of the target vector direction. The difference information includes at least the distance difference between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center in the three directions of north, east, and ground.

4. The method according to claim 1, characterized in that, Correcting the UAV's geographic location, which contains errors, using the aforementioned difference information includes: Determine the change in geographic coordinates of the ground point corresponding to the center pixel of the aerial image relative to the center of the drone; Calculate the corresponding latitude and longitude differences in the Earth coordinate system based on the changes; The calculated latitude and longitude of the ground point corresponding to the center pixel of the aerial image is determined based on the latitude and longitude coordinates of the UAV, which contain errors. A satellite reference image is acquired, and the true latitude and longitude of the center pixel of the aerial image is determined by linear interpolation based on the pixel coordinates of the center pixel of the aerial image on the satellite reference image and the prior latitude and longitude information of the satellite reference image. The error between the actual latitude and longitude and the calculated latitude and longitude is corrected to obtain the actual latitude and longitude of the UAV.

5. The method according to claim 1, characterized in that, Based on the corrected UAV geographic location, the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image are determined, including: Obtain the geographic coordinates of corner pixels in an aerial photograph; The latitude and longitude coordinates of the ground points corresponding to the corner pixels of the aerial photograph are determined based on the actual latitude and longitude of the UAV. By combining the prior latitude and longitude information of the satellite reference image, the pixel coordinates of the corner pixels of the aerial photograph in the satellite reference image are obtained through linear interpolation, and these coordinates are used as the reference image pixel coordinates of the corner pixels of the reference aerial photograph.

6. The method according to claim 1, characterized in that, Perspective transformation is performed on the satellite reference image to obtain a simulated aerial photograph from a specific UAV perspective, corresponding to the attitude view of the UAV acquired by the inertial navigation system, including: The reference image pixel coordinates of the corner pixel in the aerial image are combined with the pixel coordinates of the corner pixel in the aerial image to form a corner pixel pair; The perspective transformation matrix from the satellite reference image to the simulated aerial image is determined based on the corner pixel pairs; The perspective transformation matrix is ​​used to convert the reference image pixel coordinates of the corner pixels of the aerial image into the corner pixel coordinates of the simulated aerial image, thus obtaining a simulated aerial image from the perspective of the drone.

7. The method according to claim 1, characterized in that, Based on the simulated aerial images and UAV aerial images, a target feature extraction and image matching model is determined, including: Detect the aerial image channel of the drone aerial image; In response to determining that the drone aerial image includes one channel or three channels with the same value, the drone aerial image is determined to be an infrared image; A robust dense matching model is used to perform cross-modal matching between the UAV aerial images and the simulated aerial images; In response to determining that the drone aerial image includes three channels, and that at least two of the three channels include different values, the drone aerial image is determined to be a visible light image; A semi-dense local feature matching model is used to perform cross-modal matching between the UAV aerial image and the simulated aerial image.

8. The method according to claim 7, characterized in that, Using a robust dense matching model, cross-modal matching is performed between the UAV aerial images and the simulated aerial images, including: An encoder-decoder system is obtained; the encoder integrates CNN and DINOv2 feature extractors to construct a multi-scale feature pyramid, which includes a resolution of 1x to 16x; the Gaussian process module of the decoder calculates the feature similarity matrix through the cosine kernel function and solves the Bayesian posterior distribution to obtain the mean and covariance information of the matching. A multi-scale matching strategy from coarse to fine is adopted. Starting from 16x downsampling, a coarse matching stream is generated at each scale through an embedding decoder. Then, a convolutional refiner is used to combine local correlation calculation, displacement embedding and bilinear interpolation for pixel-level optimization. The optimized features are subjected to layer-by-layer upsampling and feature fusion processing, and the spatially uniform matching points are obtained by threshold balanced sampling and kernel density estimation as cross-modal matching results; the cross-modal matching results also include sub-pixel level dense correspondences and confidence scores.

9. The method according to claim 7, characterized in that, The UAV aerial imagery and the simulated aerial imagery are matched across modes using a semi-dense local feature matching model, including: Multi-scale feature extraction at 1 / 2, 1 / 4, and 1 / 8 resolutions is performed on the drone aerial images and the simulated aerial images. The 1 / 8 resolution features are used as coarse matching inputs, while the 1 / 2 and 1 / 4 resolution features are retained for subsequent fine processing. Local aggregation is performed on the extracted 1 / 8 resolution coarse-grained feature map to remove redundant information, and global context enhancement is performed through alternating self-attention and cross-attention modules to obtain enhanced UAV aerial image features and enhanced simulated aerial image features. Calculate the similarity matrix between the enhanced UAV aerial image features and the enhanced simulated aerial image features, establish an initial matching relationship through a multi-level filtering strategy, and use the initial matching relationship to quickly identify possible matching points as coarse matching points in the global scope; A feature pyramid structure is used to gradually fuse features at 1 / 8, 1 / 4, and 1 / 2 resolutions to the original resolution through residual connection and upsampling operations. Pixel-level localization is performed within a local window extracted around the coarse matching point, and the confidence is calculated through bidirectional softmax to find the optimal matching position; Subpixel regression is performed within a preset area around the optimal matching position to convert the heatmap into continuous subpixel coordinate offsets to obtain cross-modal matching results; the cross-modal matching results also include subpixel-level coordinate information and confidence level.

10. The method according to claim 1, characterized in that, Determining the homography transformation matrix based on the target feature extraction and image matching model includes: Detect the processor architecture to determine the processor type; Load a preset weight file; the weight file includes at least one of the following: a coarse feature extraction weight file, a fine-grained feature extraction weight file, a coarse feature matching weight file, and a layer-by-layer network refinement weight file; Based on the processor type, the weight file is inferred and matched to obtain scene matching results; For each matching point in the cross-modal matching results, the homography transformation matrix from the simulated aerial image to the UAV aerial image is determined using the Random Sample Consensus (RANSAC) algorithm based on the scene matching results.

11. The method according to claim 1, characterized in that, Transforming the pixel coordinates of the point of interest to the pixel coordinates of the satellite reference image based on the homography transformation matrix and perspective transformation matrix includes: The coordinates of the points of interest in the UAV aerial photograph are mapped to the simulated aerial photograph based on the inverse matrix of the homography transformation matrix. The perspective transformation matrix is ​​used to convert the coordinates of the point of interest in the simulated aerial image to the coordinates in the satellite reference image; the perspective transformation matrix is ​​determined when performing perspective transformation on the satellite reference image; The latitude and longitude of the point of interest are obtained by linear interpolation of the latitude and longitude information of the corner points in the known information of the satellite reference map.

12. An adaptive variable structure UAV multimodal scene matching navigation and positioning device, characterized in that, include: The acquisition module is configured to acquire drone aerial images, and use the drone pose information in the drone aerial images to perform coordinate transformation to obtain the geographic coordinates corresponding to each pixel in the aerial images. The correction module is configured to determine the difference information between the pixel coordinates of the center point of the UAV aerial image and the geographical location corresponding to the camera center, and use the difference information to correct the UAV geographical location that contains errors obtained through low-precision inertial navigation. The simulation module is configured to determine the reference image pixel coordinates corresponding to the corner pixel coordinates of the aerial image based on the corrected UAV geographical location, and perform perspective transformation on the satellite reference image to obtain the simulated aerial image under the specific UAV perspective obtained by the inertial navigation system. The selection module is configured to determine a target feature extraction and image matching model based on the simulated aerial images and drone aerial images; The determination module is configured to determine the homography transformation matrix based on the target feature extraction and image matching model; The inverse transformation module is configured to obtain the pixel coordinates of the point of interest in the drone aerial image and transform the pixel coordinates of the point of interest to the pixel coordinates of the satellite reference image based on the homography transformation matrix and the perspective transformation matrix. The positioning module is configured to determine the latitude and longitude of the point of interest based on the pixel coordinates of the satellite reference image.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 11.

14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 11.

Citation Information

Cited By

  • Multi-source fusion navigation positioning method and system for unmanned vehicle in complex terrain

    CN121916865A

  • Open-ground amphibious robot cross-domain autonomous positioning method and device based on view cone transformation

    CN122041849A

  • Airport runway multi-sensor cooperative target intrusion identification method and device

    CN122200373A