A method and system for visual navigation of a drone flight

By segmenting satellite reference images and real-world images to be matched into blocks and performing pixel-level feature matching, the modal compatibility problem in visual navigation technology was solved, enabling high-precision positioning of UAVs in complex environments.

CN120778119BActive Publication Date: 2025-11-21JIANGXI DONGQI AVIATION EQUIP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511243070.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-21
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing visual navigation technologies neglect modal compatibility when processing multimodal image registration, resulting in low matching accuracy and poor computational efficiency, which limits the real-time positioning capabilities of UAVs in complex environments.

Method used

By dividing satellite reference images into several reference blocks and dynamically splitting real-world images to be matched, pixel-level feature matching within overlapping areas is determined. Combined with lens distortion correction and real-time flight attitude data, matching accuracy and computational efficiency are improved.

Benefits of technology

This ensures that drones obtain reliable positioning output even when GPS is unavailable, improving matching accuracy and computational efficiency, and reducing navigation interruptions caused by data loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120778119B_ABST
    Figure CN120778119B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of visual navigation, in particular to a UAV flight visual navigation method and system. The method comprises the following steps: acquiring a satellite reference image and a real-time photographed real photograph to-be-matched image photographed by a UAV; dividing the satellite reference image into a plurality of reference blocks, and dynamically splitting the real photograph to-be-matched image into a plurality of real photograph blocks; determining an overlapping area of the satellite reference image and the real photograph to-be-matched image according to the plurality of reference blocks and the plurality of real photograph blocks; performing pixel-level feature matching in the overlapping area to determine a feature point offset; and calculating the current position of the UAV according to the feature point offset. The pixel-level feature matching is performed in the overlapping area to determine the feature point offset, the matching is performed in a limited area, the search range is reduced, the matching precision is improved, and the feature point offset is stable and reliable. The conversion from the image matching result to the geographical position is realized, and reliable positioning output of the UAV in a GPS failure environment is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of visual navigation, in particular to a UAV flight visual navigation method and system. BACKGROUND

[0002] Traditional UAV navigation mainly relies on the global positioning system (GPS) and inertial navigation system (INS), but in complex environments, GPS signals are prone to failure or data links are blocked or disconnected, resulting in navigation interruption. Therefore, visual navigation methods are introduced as a complementary solution, the core idea of which is to compare preloaded satellite-acquired images such as optical or synthetic aperture radar (SAR) images with images such as visible light or infrared images taken during UAV flight to achieve autonomous positioning and navigation.

[0003] However, existing visual navigation techniques often ignore modality compatibility when processing multi-modal image registration, resulting in low matching accuracy and poor computational efficiency, thereby limiting the real-time positioning capability of UAVs in complex environments. SUMMARY

[0004] The present application provides a UAV flight visual navigation method and system to solve the above problems.

[0005] In a first aspect, the present application provides a UAV flight visual navigation method, the method comprising:

[0006] acquiring a satellite reference image and a real-time taken image to be matched taken by a UAV;

[0007] dividing the satellite reference image into a plurality of reference blocks and dynamically splitting the real-time taken image to be matched into a plurality of real-time taken blocks;

[0008] determining an overlapping area of the satellite reference image and the real-time taken image to be matched according to the plurality of reference blocks and the plurality of real-time taken blocks;

[0009] performing pixel-level feature matching in the overlapping area to determine a feature point offset;

[0010] solving the current position of the UAV according to the feature point offset.

[0011] The satellite reference image and the real-time photographed image to be matched photographed by the unmanned aerial vehicle are acquired, so that the satellite reference image and the real-time photographed image to be matched are simultaneously available, and navigation interruption caused by data loss is avoided. The satellite reference image is divided into a plurality of reference blocks, and the real-time photographed image to be matched is dynamically divided into a plurality of real-time photographed blocks, so that the satellite reference image and the real-time photographed image to be matched are geometrically consistent, alignment failure caused by attitude change and lens distortion is reduced, and the efficiency and accuracy of the overlapping area identification are improved. The overlapping area of the satellite reference image and the real-time photographed image to be matched is determined according to the plurality of reference blocks and the plurality of real-time photographed blocks, so that the extraction error of the similarity peak value caused by the size difference and inefficient traversal is avoided, and the matching reliability and the calculation efficiency are improved. The pixel-level feature matching is performed in the overlapping area, the feature point offset is determined, the matching is performed in a limited area, the search range is reduced, the matching accuracy is improved, and the feature point offset is robust and reliable. The current position of the unmanned aerial vehicle is calculated according to the feature point offset, the conversion from the image matching result to the geographical position is realized, and reliable positioning output of the unmanned aerial vehicle in a GPS failure environment is ensured.

[0012] Optionally, the satellite reference image is divided into a plurality of reference blocks, and the real-time photographed image to be matched is dynamically divided into a plurality of real-time photographed blocks, and the method comprises the following steps.

[0013] The lens distortion coefficient of the camera of the unmanned aerial vehicle is acquired.

[0014] The satellite reference image and the real-time photographed image to be matched are respectively subjected to distortion elimination according to the lens distortion coefficient, so as to generate a reference corrected image and a real-time photographed corrected image.

[0015] The reference corrected image is divided into a plurality of reference blocks, and the real-time photographed corrected image is dynamically divided into a plurality of real-time photographed blocks.

[0016] The lens distortion coefficient of the camera of the unmanned aerial vehicle is acquired, so that the geometric structure distortion caused by the inherent characteristics of the lens is not corrected in the image matching. The satellite reference image and the real-time photographed image to be matched are respectively subjected to distortion elimination according to the lens distortion coefficient, so as to generate a reference corrected image and a real-time photographed corrected image, the consistency of the image geometric structure is realized, the foundation for the reference block division and the dynamic block division is laid, and the matching process is not disturbed by the inherent error of the lens. The reference corrected image is divided into a plurality of reference blocks, and the real-time photographed corrected image is dynamically divided into a plurality of real-time photographed blocks, so that the block alignment optimization is realized, the geometric distortion caused by the attitude change is reduced, and the similarity matrix is efficiently constructed and the overlapping area determination accuracy is improved.

[0017] Optionally, the real-time photographed corrected image is dynamically divided into a plurality of real-time photographed blocks, and the method comprises the following steps.

[0018] Obtaining real-time flight attitude data of the UAV and shooting parameters of a camera of the UAV;

[0019] According to the shooting parameters, determining an image size of the real-shot corrected image;

[0020] According to the real-time flight attitude data and the image size, generating a plurality of real-shot blocks consistent with a reference block size.

[0021] By the scheme, the real-time flight attitude data of the UAV and the shooting parameters of the camera of the UAV are obtained, the influence of geometric distortion caused by attitude change on block alignment is reduced, the size calculation deviation caused by inconsistent parameters is avoided, and the geometric compatibility and robustness of block generation are improved. According to the shooting parameters, the image size of the real-shot corrected image is determined to ensure that the block splitting process is efficient and has no overlap. According to the real-time flight attitude data and the image size, a plurality of real-shot blocks consistent with the reference block size are generated to eliminate geometric distortion caused by attitude change and reduce the false matching rate.

[0022] Optionally, the determining of the overlapping area of the satellite reference image and the real-shot image to be matched according to the plurality of reference blocks and the plurality of real-shot blocks comprises:

[0023] Analyzing the plurality of reference blocks and the plurality of real-shot blocks to obtain a reference pixel gray value set and a real-shot pixel gray value set;

[0024] Analyzing the reference pixel gray value set to determine a reference histogram;

[0025] Analyzing the real-shot pixel gray value set to determine a real-shot histogram;

[0026] According to the reference histogram and the real-shot histogram, calculating a normalized cross-correlation coefficient;

[0027] According to the normalized cross-correlation coefficient, constructing a similarity matrix;

[0028] Analyzing the similarity matrix to determine a plurality of similarity peak values;

[0029] Comparing the plurality of similarity peak values with a similarity threshold value respectively, and determining a block with a similarity peak value higher than the similarity threshold value as the overlapping area of the satellite reference image and the real-shot image to be matched.

[0030] By the scheme, a plurality of reference blocks and a plurality of actually photographed blocks are analyzed to obtain a reference pixel gray value set and an actually photographed pixel gray value set, ensuring that the pixel-level information of the blocks is fully captured and avoiding data omission. The reference pixel gray value set is analyzed to determine a reference histogram, reflecting the overall gray characteristics of the reference blocks and providing a structured statistical basis for the similarity comparison. The actually photographed pixel gray value set is analyzed to determine an actually photographed histogram, converting the gray data of the actually photographed blocks into a comparable statistical form for direct alignment with the reference histogram. According to the reference histogram and the actually photographed histogram, a normalized cross-correlation coefficient is calculated, which is used to objectively measure the matching degree of the block pair and eliminate the influence of the gray scale difference, providing reliable input for constructing a similarity matrix. According to the normalized cross-correlation coefficient, a similarity matrix is constructed, facilitating efficient access and scanning. The similarity matrix is analyzed to determine a plurality of similarity peaks, reducing noise interference, focusing on the most possible overlap candidates, and improving the accuracy of the similarity threshold comparison. The plurality of similarity peaks are compared with the similarity threshold, and the blocks with a similarity peak higher than the similarity threshold are determined as the overlap area of the satellite reference image and the actually photographed image to be matched, ensuring that only high-confidence matches are selected and avoiding false matches.

[0031] Optionally, the pixel-level feature matching in the overlap area is performed to determine the feature point offset, including:

[0032] Obtaining a multi-modal compatible feature descriptor generation rule;

[0033] According to the feature descriptor generation rule, feature points are extracted in the reference block and the actually photographed block of the overlap area, respectively;

[0034] The feature points are analyzed to determine a descriptor vector distance;

[0035] According to the descriptor vector distance, a distance variance of the matching point pair is determined;

[0036] The matching point pair with the largest variance within a preset percentage is removed;

[0037] According to the remaining descriptor vector distance, a homography transformation matrix is calculated;

[0038] The homography transformation matrix is analyzed to determine a projection error, and the feature point offset is determined according to the projection error.

[0039] By the scheme, the multi-modal compatible feature descriptor generation rule is obtained, the nonlinear intensity mismatch problem caused by the image modal difference is eliminated, and the distance variance problem caused by the nonlinear intensity mismatch is avoided. According to the feature descriptor generation rule, the feature points are extracted in the reference block and the real shot block in the overlapping area, and the geometric structure inconsistency problem caused by the lens distortion or the attitude change is solved. The feature points are analyzed, the descriptor vector distance is determined, and the distance calculation deviation caused by the modal difference is overcome. According to the descriptor vector distance, the distance variance of the matching point pair is determined, and the dispersion degree of the matching result is reflected. The matching point pair with the largest variance in the preset percentage is removed, the false matching rate is reduced, the error is avoided to be propagated to the pose solving link, and the reliability of the matching point pair can be improved. According to the remaining descriptor vector distance, the homography transformation matrix is calculated, the image distortion caused by the attitude change of the unmanned aerial vehicle is corrected, and the block-level geometric alignment is realized. The homography transformation matrix is analyzed, the projection error is determined, the feature point offset is determined according to the projection error, the relative attitude change of the unmanned aerial vehicle is reflected, and the cumulative deviation caused by the angular velocity error out of limit is avoided.

[0040] Optionally, the solving the current position of the unmanned aerial vehicle according to the feature point offset includes:

[0041] The geographic coordinate metadata of the satellite reference image and the camera projection model of the unmanned aerial vehicle are obtained.

[0042] The pose deviation of the unmanned aerial vehicle relative to the reference image is calculated according to the feature point offset and the camera projection model.

[0043] The longitude, latitude and altitude of the unmanned aerial vehicle are solved according to the pose deviation and the geographic coordinate metadata, and the current position of the unmanned aerial vehicle is determined according to the longitude, latitude and altitude.

[0044] By the scheme, the geographic coordinate metadata of the satellite reference image and the camera projection model of the unmanned aerial vehicle are obtained, the calculation efficiency and real-time performance are improved, the processing delay is reduced, the image size difference problem is eliminated, and the time-consuming operation of traversing the entire image is avoided. The pose deviation of the unmanned aerial vehicle relative to the reference image is calculated according to the feature point offset and the camera projection model, and the pose solving error propagation problem is avoided. The longitude, latitude and altitude of the unmanned aerial vehicle are solved according to the pose deviation and the geographic coordinate metadata, and the current position of the unmanned aerial vehicle is determined according to the longitude, latitude and altitude, so that the efficiency of the processing procedure is maintained, and the calculation efficiency problem is avoided.

[0045] Optionally, the calculating the pose deviation of the unmanned aerial vehicle relative to the reference image according to the feature point offset and the camera projection model includes:

[0046] An initial pose equation is established according to the camera projection model.

[0047] Randomly select a feature point offset, substitute it into the initial pose equation to obtain a rotation matrix and a translation vector;

[0048] Determine the rotation matrix and the translation vector as the pose deviation of the UAV relative to the reference image.

[0049] According to the camera projection model, the initial pose equation is established to provide a mathematical basis for the pose deviation calculation. Randomly selecting a feature point offset, substituting it into the initial pose equation to obtain a rotation matrix and a translation vector, avoids the error accumulation caused by the multi-matching point error, and reduces the cumulative deviation of the pose solution. Determining the rotation matrix and the translation vector as the pose deviation of the UAV relative to the reference image avoids the error caused by the attitude angle exceeding the threshold.

[0050] Optionally, the generating a plurality of real shot blocks consistent with the reference block size according to the real-time flight attitude data and the image size comprises:

[0051] Analyzing the real-time flight attitude data to determine the pitch angle and the roll angle;

[0052] According to the pitch angle and the roll angle, the rotation angle of the image plane relative to the horizontal plane is calculated;

[0053] According to the shooting parameter, the focal length of the camera is determined;

[0054] Obtaining the real-time flight height of the UAV;

[0055] According to the focal length of the camera, the real-time flight height and the image size, the edge length scaling factor of the block is calculated;

[0056] According to the rotation angle and the edge length scaling factor, affine transformation is performed on the real shot corrected image to generate a plurality of real shot blocks consistent with the reference block size.

[0057] According to the camera projection model, the initial pose equation is established to provide a mathematical basis for the pose deviation calculation. Randomly selecting a feature point offset, substituting it into the initial pose equation to obtain a rotation matrix and a translation vector, avoids the error accumulation caused by the multi-matching point error, and reduces the cumulative deviation of the pose solution. Determining the rotation matrix and the translation vector as the pose deviation of the UAV relative to the reference image avoids the error caused by the attitude angle exceeding the threshold.

[0058] Optionally, the calculating the longitude, latitude and altitude of the UAV according to the pose deviation and the geographic coordinate metadata comprises:

[0059] obtaining an angular velocity error threshold of the inertial measurement unit;

[0060] analyzing the pose deviation to determine a rotation matrix component;

[0061] determining a pitch angle deviation and a roll angle deviation according to the rotation matrix component;

[0062] if the pitch angle deviation or the roll angle deviation exceeds the angular velocity error threshold, performing weighted fusion on the rotation matrix according to the attitude angle output by the inertial measurement unit to generate a corrected rotation matrix;

[0063] calculating the longitude, latitude and altitude of the UAV according to the corrected rotation matrix, the translation vector and the geographic coordinate metadata.

[0064] According to the scheme, the angular velocity error threshold of the inertial measurement unit is obtained to ensure that the correction mechanism is triggered only when the deviation is significant, thereby avoiding unnecessary calculation overhead. The pose deviation is analyzed to determine the rotation matrix component, and the mathematical expression form of the rotation error is determined. The pitch angle deviation and the roll angle deviation are determined according to the rotation matrix component to reflect the attitude error amplitude of the UAV in the pitch and roll directions. If the pitch angle deviation or the roll angle deviation exceeds the angular velocity error threshold, the rotation matrix is weighted and fused according to the attitude angle output by the inertial measurement unit to generate a corrected rotation matrix, thereby significantly suppressing the rotation component error caused by the pose deviation exceeding the limit. The longitude, latitude and altitude of the UAV are calculated according to the corrected rotation matrix, the translation vector and the geographic coordinate metadata to ensure that the navigation and positioning result meets the accuracy requirements of visual navigation.

[0065] In a second aspect, the present application provides a UAV flight visual navigation system, which comprises:

[0066] an image acquisition module configured to acquire a satellite reference image and a real-time photographed image to be matched photographed by the UAV;

[0067] a block division module configured to divide the satellite reference image into a plurality of reference blocks and dynamically divide the real-time photographed image to be matched into a plurality of real-time photographed blocks;

[0068] an overlap determination module configured to determine an overlapping area of the satellite reference image and the real-time photographed image to be matched according to the plurality of reference blocks and the plurality of real-time photographed blocks;

[0069] an offset determination module configured to perform pixel-level feature matching in the overlapping area to determine a feature point offset;

[0070] A position determining module is configured to calculate the current position of the UAV according to the feature point offset. BRIEF DESCRIPTION OF DRAWINGS

[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0072] Figure 1 An application scenario diagram provided by an embodiment of the present application;

[0073] Figure 2 A flowchart of a UAV flight visual navigation method provided by an embodiment of the present application;

[0074] Figure 3 A UAV flight visual navigation system structure diagram provided by an embodiment of the present application. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions of the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0076] In addition, the term "and / or" in this paper is only to describe the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper, unless otherwise specified, generally represents an "or" relationship between the associated objects before and after it.

[0077] The embodiments of the present application will be further described in detail below with reference to the drawings of the specification.

[0078] When processing multi-modal image registration, visual navigation technology often ignores modality compatibility, resulting in low matching accuracy and poor computing efficiency, thereby limiting the real-time positioning ability of the UAV in complex environments.

[0079] Based on this, the application provides a UAV flight visual navigation method and system, satellite reference images and real-time shooting to-be-matched images shot by a UAV are acquired, the satellite reference images and the real-time shooting to-be-matched images are ensured to be available at the same time, and navigation interruption caused by data loss is avoided. The satellite reference images are divided into a plurality of reference blocks, and the real-time shooting to-be-matched images are dynamically block split to obtain a plurality of real-time shooting blocks, so that the satellite reference images and the real-time shooting to-be-matched images are consistent in geometry, alignment failure caused by attitude changes and lens distortion is reduced, and the efficiency and accuracy of overlapping area identification are improved. According to the plurality of reference blocks and the plurality of real-time shooting blocks, an overlapping area of the satellite reference images and the real-time shooting to-be-matched images is determined, similarity peak extraction errors caused by scale differences and inefficient traversal are avoided, and matching reliability and calculation efficiency are improved. Pixel-level feature matching is performed in the overlapping area, feature point offset is determined, matching is performed in a limited area, a search range is reduced, matching accuracy is improved, and the feature point offset is robust and reliable. According to the feature point offset, a current position of the UAV is solved, conversion from image matching results to a geographical position is realized, and reliable positioning output of the UAV in a GPS failure environment is ensured.

[0080] Figure 1 An application scenario provided by the application is provided, and the method provided by the application is applied when UAV flight visual navigation is performed.

[0081] Specifically, the method provided by the application is applied in any server, the server interacts with a satellite platform and a UAV on-board camera, satellite reference images are acquired through the satellite platform, and real-time shooting to-be-matched images are acquired through the UAV on-board camera. The satellite reference images are divided into a plurality of reference blocks, and the real-time shooting to-be-matched images are dynamically block split to obtain a plurality of real-time shooting blocks. According to the plurality of reference blocks and the plurality of real-time shooting blocks, an overlapping area of the satellite reference images and the real-time shooting to-be-matched images is determined. Pixel-level feature matching is performed in the overlapping area, feature point offset is determined, matching is performed in a limited area, a search range is reduced, matching accuracy is improved, and the feature point offset is robust and reliable. According to the feature point offset, a current position of the UAV is solved, conversion from image matching results to a geographical position is realized, and reliable positioning output of the UAV in a GPS failure environment is ensured. The specific implementation mode can be referred to the following embodiments.

[0082] Figure 2 A flowchart of a UAV flight visual navigation method provided by an embodiment of the application, the method of the embodiment can be applied to the server in the above scenario. As shown in the figure, the method includes: Figure 2

[0083] S201, acquiring satellite reference images and real-time shooting to-be-matched images shot by a UAV;

[0084] ​Satellite reference images can be multimodal image data acquired through a satellite platform.

[0085] A drone can be an unmanned aerial vehicle used to capture images and perform navigation tasks in real time during flight.

[0086] The real-time images to be matched can be images captured by drones in real time.

[0087] Specifically, satellite reference images are acquired through a satellite platform; the current environment is captured in real time by an airborne camera on a drone to generate real-world images to be matched. Each real-world image to be matched is accompanied by a timestamp and preliminary attitude data (including parameters such as pitch and roll angles).

[0088] S202. Divide the satellite reference image into several reference blocks, and dynamically split the real-shot image to be matched into several real-shot blocks.

[0089] A reference block can be a fixed-size rectangular area defined by a satellite reference image.

[0090] The real-shot block can be a shape region obtained by dynamically segmenting the real-shot image to be matched.

[0091] Specifically, based on the geographical coverage of the satellite reference image (the spatial range of the actual land surface area corresponding to the satellite reference image), it is divided into multiple equal-sized rectangular blocks, each of which represents a reference block, thus generating a dry reference block.

[0092] The real-world image to be matched is divided into an initial grid (the same size as the reference block), but the grid shape is dynamically distorted according to the preliminary attitude data (for example, when the pitch angle changes, the block changes from a rectangle to a trapezoid). After splitting, each adjusted grid cell is a real-world block, thus obtaining several real-world blocks.

[0093] S203. Based on several reference blocks and several real-shot blocks, determine the overlapping area between the satellite reference image and the real-shot image to be matched;

[0094] The overlapping area can be a geographical area that is covered by both the satellite reference image and the real-world image to be matched.

[0095] Specifically, the algorithm iterates through combinations of several reference blocks and several real-shot blocks, and uses a normalized cross-correlation coefficient algorithm to normalize the pixel values ​​of the reference blocks and real-shot blocks in the combination (subtracting the mean and dividing by the standard deviation); then, it calculates the similarity based on each normalized combination.

[0096] When the similarity exceeds a preset similarity threshold set according to experimental calibration (for filtering low-similarity combinations), it is determined that the satellite reference image and the real-shot image to be matched have an overlapping area (a single continuous area, avoiding multi-area conflicts).

[0097] S204, performing pixel-level feature matching in the overlapping area to determine a feature point offset;

[0098] The feature point offset can be a coordinate difference vector between matched feature points.

[0099] Specifically, key points (pixel positions with significant local structures, such as corner points, edge intersection points, and texture high-variation areas) are identified in the overlapping area using a feature detection algorithm; for different image modalities (such as SAR data (single-channel grayscale image) of the satellite reference image and visible light data (three-channel or multi-channel image) of the real-shot image to be matched), a descriptor algorithm is adaptively adjusted to handle nonlinear intensity differences (the pixel intensity mapping relationship of the same ground object in SAR data and visible light data does not conform to linear scaling (such as a cement pavement appearing bright white (high scattering) in SAR and gray (medium reflectivity) in visible light), to generate feature descriptors (quantitative representation of the visual characteristics of the local image area centered on the key points).

[0100] The feature descriptors in the overlapping area are compared, the distances between the descriptors are calculated (the smaller the distance, the higher the similarity), and the matching point pairs (the point pairs with the smallest distance) are selected; for each matching point pair, the coordinate offset in the image coordinate system (representing the imaging position deviation of the same place) is calculated, and the feature point offset is determined.

[0101] S205, calculating the current position of the unmanned aerial vehicle according to the feature point offset.

[0102] The current position of the unmanned aerial vehicle can be the real-time latitude, longitude, and altitude of the unmanned aerial vehicle.

[0103] Specifically, the camera internal and external parameters (intrinsic parameters describing the internal optical characteristics of the camera (focal length, principal point, etc.) and the pose (position and direction) of the camera in the physical space) obtained through camera calibration are applied to perspective transformation to map the feature point offset to the geographic spatial offset (displacement of the unmanned aerial vehicle in the geographic coordinates); at the same time, inertial measurement unit data (such as angular velocity) are fused to correct the attitude deviation (compensate for the pitch and roll angle error of the unmanned aerial vehicle); finally, the above data is iteratively optimized to calculate the current position of the unmanned aerial vehicle.

[0104] The satellite reference image and the real-time photographed image to be matched are acquired, so that the satellite reference image and the real-time photographed image to be matched are available at the same time, and navigation interruption caused by data loss is avoided. The satellite reference image is divided into a plurality of reference blocks, and the real-time photographed image to be matched is dynamically divided into a plurality of real-time photographed blocks, so that the satellite reference image and the real-time photographed image to be matched are consistent in geometry, alignment failure caused by attitude change and lens distortion is reduced, and the efficiency and accuracy of the overlapping area identification are improved. The overlapping area of the satellite reference image and the real-time photographed image to be matched is determined according to the plurality of reference blocks and the plurality of real-time photographed blocks, so that the extraction error of the similarity peak value caused by the size difference and the inefficient traversal is avoided, and the matching reliability and the calculation efficiency are improved. The pixel-level feature matching is performed in the overlapping area, the feature point offset is determined, the matching is performed in a limited area, the search range is reduced, the matching accuracy is improved, and the feature point offset is robust and reliable. The current position of the unmanned aerial vehicle is calculated according to the feature point offset, the conversion from the image matching result to the geographic position is realized, and reliable positioning output of the unmanned aerial vehicle in the GPS failure environment is ensured.

[0105] In some embodiments, a lens distortion coefficient of the unmanned aerial vehicle camera is acquired; the satellite reference image and the real-time photographed image to be matched are respectively subjected to distortion elimination according to the lens distortion coefficient, to generate a reference corrected image and a real-time photographed corrected image; the reference corrected image is divided into a plurality of reference blocks, and the real-time photographed corrected image is dynamically divided into a plurality of real-time photographed blocks.

[0106] The unmanned aerial vehicle camera can be an imaging device installed on the unmanned aerial vehicle.

[0107] The lens distortion coefficient can be a parameter set acquired through a calibration process of the unmanned aerial vehicle camera.

[0108] The reference corrected image can be an output image generated by applying the lens distortion coefficient to the satellite reference image for distortion elimination.

[0109] The real-time photographed corrected image can be an output image generated by applying the lens distortion coefficient to the real-time photographed image to be matched for distortion elimination.

[0110] Specifically, a set of images (covering different angles of view) are collected at a plurality of preset postures (such as rotation angles and translation positions) using a calibration board (such as a checkerboard pattern); the feature point coordinates (such as the corner points of the checkerboard) of the calibration board in each set of images are detected, and two-dimensional pixel positions (coordinate values of the feature points in the image coordinate system) are extracted by using an image processing algorithm; and then a calibration algorithm is applied to acquire the lens distortion coefficient of the unmanned aerial vehicle camera.

[0111] According to the lens distortion coefficient and the camera intrinsic parameter, a correction offset of each pixel is calculated; each pixel in the satellite reference image is transformed to a new position according to the correction offset by using a bilinear interpolation method, so as to generate a non-distorted reference correction image; each pixel in the actually-shot image to be matched is transformed to a new position according to the correction offset by using the bilinear interpolation method, so as to generate a non-distorted actually-shot correction image.

[0112] The reference correction image is cut according to the row and column order, each block covers a continuous pixel area; the image data of each pixel area is converted into an independent unit, so as to form a plurality of reference blocks. The actually-shot correction image is analyzed, the block size and position are adaptively determined, and dynamic block splitting is performed, so as to obtain a plurality of actually-shot blocks, for example, a smaller block is used in a high-texture area, a larger block is used in a low-texture area, and the block boundary is adjusted through a sliding window mechanism, so as to ensure that the block covers the entire image and there is no overlap.

[0113] According to the lens distortion coefficient, the satellite reference image and the actually-shot image to be matched are respectively subjected to distortion elimination, so as to generate the reference correction image and the actually-shot correction image, the consistency of the image geometric structure is realized, a foundation is laid for reference block division and dynamic block splitting, and the matching process is ensured not to be disturbed by the inherent error of the lens. The reference correction image is divided into a plurality of reference blocks, and the actually-shot correction image is subjected to dynamic block splitting, so as to obtain a plurality of actually-shot blocks, block alignment optimization is realized, geometric distortion caused by attitude change is reduced, so as to efficiently construct a similarity matrix and improve the determination accuracy of the overlapping area.

[0114] In some embodiments, real-time flight attitude data of the unmanned aerial vehicle and shooting parameters of the camera of the unmanned aerial vehicle are acquired; according to the shooting parameters, the image size of the actually-shot correction image is determined; according to the real-time flight attitude data and the image size, a plurality of actually-shot blocks with the same size as the reference block size are generated.

[0115] The real-time flight attitude data can be attitude information of the unmanned aerial vehicle collected in real time by an inertial measurement unit in the flight process.

[0116] The shooting parameters can be configuration parameters of the camera of the unmanned aerial vehicle, including image resolution (i.e. the width pixel number and the height pixel number of the actually-shot image), focal length, pixel size and the like.

[0117] The image size can be the geometric size of the actually-shot correction image.

[0118] Specifically, the flight attitude data is collected in real time by an inertial measurement unit sensor carried by the unmanned aerial vehicle; meanwhile, the shooting parameters are read from the hardware interface of the camera of the unmanned aerial vehicle. The image resolution (i.e. the width and height pixel values of the actually taken image) is extracted from the shooting parameters and directly applied to the actually taken corrected image; the width and height pixel numbers in the shooting parameters are used to determine the image size of the actually taken corrected image.

[0119] The fixed block size is obtained from the reference block definition; the geometric correction is applied using the real-time flight attitude data to compensate for the image rotation or scaling effect caused by the change in the attitude of the unmanned aerial vehicle (for example, the pitch angle affects the vertical direction deformation, and the roll angle affects the horizontal direction deformation); the actually taken corrected image is divided into a rectangular grid according to the image size (width and height); and then each actually taken block corresponds to a cell in the rectangular grid, thereby generating a plurality of actually taken blocks consistent with the reference block size.

[0120] According to the present solution, the real-time flight attitude data of the unmanned aerial vehicle and the shooting parameters of the camera of the unmanned aerial vehicle are obtained, the influence of geometric distortion caused by the change in attitude on block alignment is reduced, the size calculation deviation caused by the inconsistency of parameters is avoided, and the geometric compatibility and robustness of block generation are improved. According to the shooting parameters, the image size of the actually taken corrected image is determined to ensure that the block splitting process is efficient and has no overlap. According to the real-time flight attitude data and the image size, a plurality of actually taken blocks consistent with the reference block size are generated to eliminate geometric distortion caused by the change in attitude and reduce the false matching rate.

[0121] In some embodiments, a plurality of reference blocks and a plurality of actually taken blocks are analyzed to obtain a reference pixel gray value set and an actually taken pixel gray value set; the reference pixel gray value set is analyzed to determine a reference histogram; the actually taken pixel gray value set is analyzed to determine an actually taken histogram; the reference histogram and the actually taken histogram are used to calculate a normalized cross-correlation coefficient; the normalized cross-correlation coefficient is used to construct a similarity matrix; the similarity matrix is analyzed to determine a plurality of similarity peaks; and the plurality of similarity peaks are compared with a similarity threshold, and the blocks with a similarity peak higher than the similarity threshold are determined as the overlapping area of the satellite reference image and the actually taken image to be matched.

[0122] The reference pixel gray value set can be a set composed of all pixel gray values extracted from each reference block of the satellite reference image.

[0123] The actually taken pixel gray value set can be a set composed of all pixel gray values extracted from each actually taken block of the actually taken image to be matched.

[0124] The reference histogram can be a gray value distribution histogram generated based on the reference pixel gray value set.

[0125] The real-time histogram can be a gray value distribution histogram generated based on a set of real-time pixel gray values.

[0126] The normalized cross-correlation coefficient can be a normalized similarity index calculated by comparing the reference histogram and the real-time histogram.

[0127] The similarity matrix can be a two-dimensional matrix for the similarity relationship of a plurality of reference blocks and a plurality of real-time blocks.

[0128] The similarity peak value can be a local maximum point in the similarity matrix.

[0129] The similarity threshold value can be a numerical threshold value for comparison with the similarity peak value.

[0130] The block can refer to a sub-region block into which a satellite reference image or a real-time image to be matched is divided.

[0131] Specifically, for each reference block of the satellite reference image and each real-time block of the real-time image to be matched, all pixels in the block are traversed; for each reference block, the gray values of all pixels thereof are extracted to form a set of reference pixel gray values; at the same time, for each real-time block, the gray values of all pixels thereof are extracted to form a set of real-time pixel gray values.

[0132] According to the set of reference pixel gray values, the distribution of the gray values is counted; the gray range is determined by the minimum gray value and the maximum gray value in the distribution, the gray range (used to calculate the occurrence frequency of each interval) is divided into a plurality of intervals, and the occurrence frequency of the pixel gray values in each interval is calculated; the occurrence frequency values are organized into a histogram as the reference histogram.

[0133] According to the set of real-time pixel gray values, the distribution of the gray values is counted; the gray range is determined by the minimum gray value and the maximum gray value in the distribution, the gray range (used to calculate the occurrence frequency of each interval) is divided into a plurality of intervals, and the occurrence frequency of the pixel gray values in each interval is calculated; the occurrence frequency values are organized into a histogram as the real-time histogram.

[0134] The covariance value (reflecting the correlation) between the reference histogram and the real-time histogram is calculated; then, the standard deviations of the reference histogram and the real-time histogram are calculated to normalize the covariance value, thereby obtaining the normalized cross-correlation coefficient (used to quantify the similarity degree of the reference histogram and the real-time histogram).

[0135] The normalized cross-correlation coefficient set is organized into a similarity matrix, wherein the rows of the similarity matrix correspond to the reference blocks and the columns correspond to the real blocks. According to the similarity matrix, the normalized cross-correlation coefficient is compared with the values of adjacent elements (such as up, down, left, right or diagonal elements). If the similarity matrix is greater than the values of all adjacent elements, it is marked as a number of similarity peaks.

[0136] The similarity threshold is set through experimental tests (such as simulating the flight of a UAV in a GPS failure scenario). Then the normalized cross-correlation coefficients of a number of similarity peaks are extracted and compared with the similarity threshold. If the normalized cross-correlation coefficient is higher than the similarity threshold, the reference block and the real block corresponding to the peak are marked as the overlapping area of the satellite reference image and the real image to be matched.

[0137] According to the scheme, a number of reference blocks and a number of real blocks are analyzed to obtain a reference pixel gray value set and a real pixel gray value set, ensuring that the pixel-level information of the blocks is fully captured and avoiding data omission. The reference pixel gray value set is analyzed to determine the reference histogram, which reflects the overall gray scale characteristics of the reference block and provides a structured statistical basis for similarity comparison. The real pixel gray value set is analyzed to determine the real histogram, which converts the gray scale data of the real block into a comparable statistical form, facilitating direct alignment with the reference histogram. According to the reference histogram and the real histogram, the normalized cross-correlation coefficient is calculated to objectively measure the matching degree of the block pair and eliminate the influence of gray scale differences, providing reliable input for constructing the similarity matrix. According to the normalized cross-correlation coefficient, the similarity matrix is constructed to facilitate efficient access and scanning. The similarity matrix is analyzed to determine a number of similarity peaks, reducing noise interference and focusing on the most likely overlapping candidates to improve the accuracy of similarity threshold comparison. The similarity peaks are compared with the similarity threshold, and the blocks with similarity peaks higher than the similarity threshold are determined as the overlapping area of the satellite reference image and the real image to be matched, ensuring that only high-confidence matches are selected and avoiding false matches.

[0138] In some embodiments, a multi-modal compatible feature descriptor generation rule is obtained; according to the feature descriptor generation rule, feature points are extracted in the reference block and the real block of the overlapping area respectively; the feature points are analyzed to determine the descriptor vector distance; according to the descriptor vector distance, the distance variance of the matching point pair is determined; the matching point pair with the largest variance within a preset percentage is eliminated; according to the remaining descriptor vector distance, a homography transformation matrix is calculated; the homography transformation matrix is analyzed to determine the projection error, and according to the projection error, the feature point offset is determined.

[0139] The multi-modal compatibility can be that the feature descriptor generation rule can adapt to different image modalities of different imaging principles.

[0140] The feature descriptor generation rule can be a preset rule set for generating feature point descriptor vectors.

[0141] A feature point can be a key position point with a significant local feature identified by a feature descriptor generation rule in a reference block or a real shot block in the overlapping region.

[0142] A descriptor vector distance can be a difference quantization value between a reference block descriptor vector and a real shot block descriptor vector of the same feature point.

[0143] A matching point pair can be a point pair composed of a reference block feature point and a real shot block feature point.

[0144] A distance variance can be a statistical variance of the descriptor vector distance values of all matching point pairs.

[0145] A preset percentage can be a pre-set rejection ratio parameter for screening matching point pairs, pre-stored in the server and called when used.

[0146] A homography transformation matrix can be a matrix used to model the geometric projection transformation relationship from the reference block to the real shot block.

[0147] A projection error can be a pixel-level position deviation between the coordinates of a reference block feature point and the coordinates of a real shot block feature point.

[0148] Specifically, a multi-modal compatible feature descriptor generation rule (compatible with different image modalities, such as optical images or synthetic aperture radar images) is obtained from the embedded system of the unmanned aerial vehicle.

[0149] For each reference block in the overlapping region, all pixels in the reference block are traversed, key point positions (such as corner points or edge points) are identified using the feature descriptor generation rule, and corresponding feature points are generated. For each real shot block in the overlapping region, all pixels in the real shot block are traversed, key point positions (such as corner points or edge points) are identified using the feature descriptor generation rule, and corresponding feature points are generated.

[0150] According to the feature points of the reference block and the feature points of the real shot block, the descriptor vector distance is calculated, and the smaller the distance is, the higher the feature point matching degree is. For each reference block feature point, the feature point with the smallest descriptor vector distance in the real shot block is selected to form a matching point pair (each point pair contains one reference feature point and one real shot feature point); the descriptor vector distances of all matching point pairs are collected, and the statistical variance of the descriptor vector distance is calculated as the distance variance of the matching point pairs.

[0151] According to the distance variance, the matching point pairs are sorted (from high variance to low variance); then a preset percentage is set according to the embedded system of the unmanned aerial vehicle, the number of matching point pairs corresponding to the preset percentage is calculated, and the matching point pairs with the largest variance in the preset percentage are removed. According to the coordinate positions (pixel coordinates of the feature points of the reference block and the feature points of the photographed block in the image) and the descriptor vector distance of the remaining matching point pairs, the least square method is applied to calculate the homographic transformation matrix (used to describe the projection transformation relationship between the blocks).

[0152] The homographic transformation matrix is applied to project the reference block feature point coordinates into the photographed block feature point coordinates, and the projection error is calculated; according to the projection error, the average projection error is calculated as the feature point offset.

[0153] According to the feature descriptor generation rule, the feature points are extracted in the reference block and the photographed block in the overlapping area, and the geometric structure inconsistency caused by lens distortion or attitude change is solved. The feature points are analyzed, the descriptor vector distance is determined, and the distance calculation deviation caused by the modal difference is overcome. According to the descriptor vector distance, the distance variance of the matching point pairs is determined, and the dispersion degree of the matching result is reflected. The matching point pairs with the largest variance in the preset percentage are removed, the false matching rate is reduced, the error is avoided to be propagated to the pose solution link, and the reliability of the matching point pairs can be improved. According to the remaining descriptor vector distance, the homographic transformation matrix is calculated, the image distortion caused by the attitude change of the unmanned aerial vehicle is corrected, and the block-level geometric alignment is realized. The homographic transformation matrix is analyzed, the projection error is determined, the feature point offset is determined according to the projection error, the relative attitude change of the unmanned aerial vehicle is reflected, and the cumulative deviation caused by the over-limit angular velocity error is avoided.

[0154] In some embodiments, the geographic coordinate metadata of the satellite reference image and the camera projection model of the unmanned aerial vehicle are obtained; according to the feature point offset and the camera projection model, the pose deviation of the unmanned aerial vehicle relative to the reference image is calculated; according to the pose deviation and the geographic coordinate metadata, the longitude, latitude and altitude of the unmanned aerial vehicle are calculated, and the current position of the unmanned aerial vehicle is determined according to the longitude, latitude and altitude.

[0155] The geographic coordinate metadata can be the geographic reference information of the center point of the satellite reference image, including longitude, latitude and altitude information.

[0156] The camera projection model of the unmanned aerial vehicle can be the internal parameter data set of the pre-calibration of the camera of the unmanned aerial vehicle, including focal length, principal point coordinates and distortion coefficient data.

[0157] The relative reference image of the unmanned aerial vehicle can be the position relationship of the unmanned aerial vehicle relative to the satellite reference image in the world coordinate system.

[0158] The pose deviation can be a position deviation component of the UAV relative to the reference image.

[0159] The longitude and latitude can be longitude and latitude in geographic coordinates.

[0160] The altitude can be a height value of the UAV relative to sea level.

[0161] Specifically, the pre-stored geographic coordinate metadata of the satellite reference image is read from a non-volatile memory of a UAV embedded system (a hardware component for persistently storing geographic coordinate metadata); and a projection model of a UAV camera is read from the UAV embedded system.

[0162] According to the intrinsic parameters (such as focal length and principal point) of the camera projection model, the feature point offset is back-projected into the camera coordinate system (a three-dimensional coordinate system with the center of the UAV camera lens as the origin), and the relative position deviation (the position deviation of the UAV relative to the reference image) is calculated; then, combined with the real-time attitude data (such as the pitch angle and roll angle) of the UAV, the relative position deviation in the camera coordinate system is converted to the world coordinate system (used to represent the absolute physical position of the UAV), and the pose deviation is obtained.

[0163] According to the geographic coordinate metadata, the pose deviation is added to the reference point coordinates by applying coordinate system conversion to obtain the preliminary three-dimensional position of the UAV; according to the attitude deviation, the preliminary three-dimensional position is adjusted to compensate for the influence of the angle change (for example, by updating the position coordinates to ensure alignment with the geographic coordinate system); and the corrected three-dimensional position is converted into the longitude and latitude and altitude of the UAV. The longitude and latitude and altitude are integrated to determine the current position coordinates of the UAV.

[0164] By this scheme, the geographic coordinate metadata of the satellite reference image and the projection model of the UAV camera are obtained, the calculation efficiency and real-time performance are improved, the processing delay is reduced, the image scale difference problem is eliminated, and the time-consuming operation of traversing the entire image is avoided. According to the feature point offset and the camera projection model, the pose deviation of the UAV relative to the reference image is calculated, and the pose solution error propagation problem is avoided. According to the pose deviation and the geographic coordinate metadata, the longitude and latitude and altitude of the UAV are calculated, and according to the longitude and latitude and altitude, the current position of the UAV is determined, the processing efficiency is maintained, and the calculation efficiency problem is avoided.

[0165] In some embodiments, an initial pose equation is established according to the camera projection model; a feature point offset is randomly selected and substituted into the initial pose equation to obtain a rotation matrix and a translation vector; and the rotation matrix and the translation vector are determined as the pose deviation of the UAV relative to the reference image.

[0166] The initial pose equation can be a mathematical relationship for establishing a mapping relationship between the feature point offset and the pose deviation of the UAV relative to the reference image.

[0167] The rotation matrix can be a matrix representing a rotation deviation of the UAV relative to the reference image.

[0168] The translation vector can be a vector representing a translation deviation of the UAV relative to the reference image.

[0169] Specifically, an initial pose equation is directly constructed by using intrinsic parameters (including focal length and principal point coordinates) of a camera projection model, wherein the focal length is used to convert pixel displacement into physical scale displacement (e.g., how many meters of displacement corresponds to one pixel), and the principal point coordinates are used to correct the image center point offset (position deviation between the image center point (geometric center of the image plane) and the principal point).

[0170] A uniform random sampling algorithm is used to randomly extract feature point offset data; the extracted feature point offset data is inserted into the initial pose equation as an input variable; then, an iterative optimization algorithm is used to solve the initial pose equation to determine the optimal rotation matrix and translation vector. The rotation matrix and the translation vector are directly determined as the pose deviation of the UAV relative to the reference image.

[0171] According to the camera projection model, the initial pose equation is established to provide a mathematical basis for pose deviation calculation. The feature point offset is randomly selected and substituted into the initial pose equation to obtain the rotation matrix and the translation vector, which avoids the accumulation of errors of multiple matching points and reduces the cumulative deviation of the pose solution. The rotation matrix and the translation vector are determined as the pose deviation of the UAV relative to the reference image, which avoids errors caused by attitude angle exceeding the threshold.

[0172] In some embodiments, real-time flight attitude data is analyzed to determine the pitch angle and the roll angle; the rotation angle of the image plane relative to the horizontal plane is calculated according to the pitch angle and the roll angle; the camera focal length is determined according to the shooting parameters; the real-time flight height of the UAV is obtained; the block length scaling factor is calculated according to the camera focal length, the real-time flight height, and the image size; and the real-shot corrected image is subjected to affine transformation according to the rotation angle and the block length scaling factor to generate a plurality of real-shot blocks consistent with the reference block size.

[0173] The pitch angle can be an angle of rotation of the UAV about its lateral axis.

[0174] The roll angle can be an angle of rotation of the UAV about its longitudinal axis.

[0175] The image plane can be a physical projection plane of the UAV camera capturing the real-shot image.

[0176] The horizontal plane can be a reference plane of the rotation angle of the image plane.

[0177] The rotation angle can be a comprehensive deflection angle of the image plane relative to the horizontal plane.

[0178] The camera focal length can be the optical focal length of the camera lens.

[0179] The real-time flight height can be the vertical height of the UAV relative to the sea level or ground level.

[0180] The side length scaling factor can be a scale coefficient for adjusting the side length of the real-shot corrected image.

[0181] Specifically, the pitch angle field and the roll angle field in the real-time flight attitude data are read to directly obtain the pitch angle and the roll angle. For example, the pitch angle represents the front-to-back tilt angle of the UAV, and the roll angle represents the left-to-right tilt angle of the UAV. According to the pitch angle and the roll angle, the rotation angle of the image plane relative to the horizontal plane is calculated using a combination of trigonometric functions.

[0182] The focal length field is extracted from the camera shooting parameters to determine the camera focal length. The real-time flight height of the UAV is read from the GPS altimeter installed on the top of the UAV body. The quotient of the real-time flight height and the camera focal length is calculated, and then the product of the quotient and the image size (such as the pixel value of the image width or height) is calculated. The product is used as the side length scaling factor of the block.

[0183] The real-shot corrected image is rotated and transformed (to eliminate the image tilt caused by the attitude change) using the rotation angle, and the image after the rotation transformation is scaled and transformed using the side length scaling factor to adjust the image size to match the reference block size (uniformly adjust the image width and height). The scaled and transformed image is divided into grid-shaped blocks, wherein the size (such as the side length pixel number) of each grid-shaped block is set to be the same as the reference block size, and a plurality of real-shot blocks are generated.

[0184] Through the scheme, the real-time flight attitude data is analyzed to determine the pitch angle and the roll angle, ensuring that the parameters of the attitude change are quantitatively captured to provide input for calculating the image rotation angle. According to the pitch angle and the roll angle, the rotation angle of the image plane relative to the horizontal plane is calculated to eliminate the image tilt distortion. According to the shooting parameters, the camera focal length is determined to ensure that the image size adjustment is based on the camera hardware characteristics. The real-time flight height of the UAV is obtained to reflect the distance from the ground, ensuring that the image size is associated with the height change of the physical world. According to the camera focal length, the real-time flight height, and the image size, the side length scaling factor of the block is calculated to represent the scaling ratio of the image caused by the height difference. According to the rotation angle and the side length scaling factor, the real-shot corrected image is subjected to affine transformation to generate a plurality of real-shot blocks with the same reference block size, effectively eliminating the geometric distortion caused by the attitude.

[0185] In some embodiments, an angular velocity error threshold of the inertial measurement unit is acquired; a pose deviation is analyzed to determine a rotation matrix component; based on the rotation matrix component, a pitch angle deviation amount and a roll angle deviation amount are determined; if the pitch angle deviation amount or the roll angle deviation amount exceeds the angular velocity error threshold, a rotation matrix is fused by weighting based on a posture angle output by the inertial measurement unit to generate a corrected rotation matrix; and based on the corrected rotation matrix, a translation vector and geographic coordinate metadata, a longitude, a latitude and an altitude of the unmanned aerial vehicle are calculated.

[0186] The inertial measurement unit can be a sensor device for outputting a real-time posture angle of the unmanned aerial vehicle.

[0187] The angular velocity error threshold can be a constant value preset for comparison to determine whether the pitch angle deviation amount or the roll angle deviation amount exceeds the threshold.

[0188] The rotation matrix component can be a rotation matrix element separated from the pose deviation.

[0189] The pitch angle deviation amount can be a rotation angle deviation around the Y-axis calculated from the rotation matrix component.

[0190] The roll angle deviation amount can be a rotation angle deviation around the X-axis calculated from the rotation matrix component.

[0191] Specifically, an angular velocity error threshold of the inertial measurement unit is acquired from an embedded system of the unmanned aerial vehicle (for comparison to determine whether the pitch angle deviation amount or the roll angle deviation amount exceeds the threshold). A pose deviation is analyzed to separate a rotation part therefrom and extract a rotation matrix component. The rotation matrix component is inversely calculated to determine a pitch angle deviation amount based on a rotation deviation of the unmanned aerial vehicle around a transverse axis (Y-axis) and determine a roll angle deviation amount based on a rotation deviation of the unmanned aerial vehicle around a longitudinal axis (X-axis).

[0192] The pitch angle deviation amount and the roll angle deviation amount are respectively compared with the angular velocity error threshold by scalar comparison. If the pitch angle deviation amount or the roll angle deviation amount exceeds the angular velocity error threshold (for example, the pitch angle deviation amount exceeds the angular velocity error threshold or the roll angle deviation amount exceeds the angular velocity error threshold), a weighted fusion is triggered. A posture angle (including a pitch angle and a roll angle) is read in real time based on the inertial measurement unit; the rotation matrix component is linearly fused with the posture angle to generate a corrected rotation matrix.

[0193] The translation vector and the geographic coordinate metadata are integrated; then, the pose (position and posture) is rotationally transformed by the corrected rotation matrix to calculate the longitude, the latitude and the altitude of the unmanned aerial vehicle through coordinate mapping.

[0194] By the scheme, the angular velocity error threshold of the inertial measurement unit is acquired, so that the correction mechanism is triggered only when the deviation is significant, and unnecessary calculation overhead is avoided. The pose deviation is analyzed, the rotation matrix component is determined, and the mathematical expression form of the rotation error is determined. According to the rotation matrix component, the pitch angle deviation amount and the roll angle deviation amount are determined, and the attitude error amplitude of the unmanned aerial vehicle in the pitch and roll directions is reflected. If the pitch angle deviation amount or the roll angle deviation amount exceeds the angular velocity error threshold, the rotation matrix is fused according to the attitude angle output by the inertial measurement unit, a corrected rotation matrix is generated, and the rotation component error caused by the pose deviation exceeding the limit is significantly suppressed. According to the corrected rotation matrix, the translation vector and the geographic coordinate metadata, the longitude, latitude and altitude of the unmanned aerial vehicle are calculated, and the navigation and positioning result is ensured to meet the accuracy requirements of visual navigation.

[0195] Figure 3 A structural schematic diagram of an unmanned aerial vehicle flight visual navigation system provided by an embodiment of the present application is shown in the figure. Figure 3 The unmanned aerial vehicle flight visual navigation system 300 of the embodiment includes an image acquisition module 301, a block division module 302, an overlap determination module 303, an offset determination module 304 and a position determination module 305.

[0196] The image acquisition module 301 is configured to acquire a satellite reference image and a real-time shot to-be-matched image shot by an unmanned aerial vehicle.

[0197] The block division module 302 is configured to divide the satellite reference image into a plurality of reference blocks, and dynamically divide the real-time shot to-be-matched image into a plurality of real-time shot blocks.

[0198] The overlap determination module 303 is configured to determine an overlap area of the satellite reference image and the real-time shot to-be-matched image according to the plurality of reference blocks and the plurality of real-time shot blocks.

[0199] The offset determination module 304 is configured to perform pixel-level feature matching in the overlap area to determine a feature point offset amount.

[0200] The position determination module 305 is configured to calculate the current position of the unmanned aerial vehicle according to the feature point offset amount.

[0201] Optionally, when the block division module 302 divides the satellite reference image into a plurality of reference blocks, and dynamically divides the real-time shot to-be-matched image into a plurality of real-time shot blocks, the block division module 302 is configured to:

[0202] acquire a lens distortion coefficient of the camera of the unmanned aerial vehicle;

[0203] eliminate distortion of the satellite reference image and the real-time shot to-be-matched image according to the lens distortion coefficient, respectively, to generate a reference corrected image and a real-time shot corrected image;

[0204] The reference correction image is divided into a plurality of reference blocks, and the real shot correction image is dynamically block split to obtain a plurality of real shot blocks.

[0205] Optionally, when the block division module 302 dynamically block splits the real shot correction image to obtain a plurality of real shot blocks, the block division module 302 is configured to:

[0206] Obtain real-time flight attitude data of the unmanned aerial vehicle and shooting parameters of the camera of the unmanned aerial vehicle;

[0207] According to the shooting parameters, determine the image size of the real shot correction image;

[0208] According to the real-time flight attitude data and the image size, generate a plurality of real shot blocks consistent with the reference block size.

[0209] Optionally, when the overlap determination module 303 determines the overlapping area of the satellite reference image and the real shot image to be matched according to the plurality of reference blocks and the plurality of real shot blocks, the overlap determination module 303 is configured to:

[0210] Analyze the plurality of reference blocks and the plurality of real shot blocks to obtain a reference pixel gray value set and a real shot pixel gray value set;

[0211] Analyze the reference pixel gray value set to determine a reference histogram;

[0212] Analyze the real shot pixel gray value set to determine a real shot histogram;

[0213] According to the reference histogram and the real shot histogram, calculate a normalized cross-correlation coefficient;

[0214] According to the normalized cross-correlation coefficient, construct a similarity matrix;

[0215] Parse the similarity matrix to determine a plurality of similarity peak values;

[0216] Compare the plurality of similarity peak values with a similarity threshold value respectively, and determine the blocks with similarity peak values higher than the similarity threshold value as the overlapping area of the satellite reference image and the real shot image to be matched.

[0217] Optionally, when the offset determination module 304 performs pixel-level feature matching in the overlapping area to determine a feature point offset, the offset determination module 304 is configured to:

[0218] Obtain a multi-modal compatible feature descriptor generation rule;

[0219] According to the feature descriptor generation rule, extract feature points in the reference blocks and real shot blocks of the overlapping area respectively;

[0220] analyzing the feature points to determine descriptor vector distances;

[0221] determining a distance variance of matching point pairs according to the descriptor vector distances;

[0222] eliminating matching point pairs with a maximum variance within a preset percentage;

[0223] calculating a homography transformation matrix according to the remaining descriptor vector distances;

[0224] analyzing the homography transformation matrix to determine a projection error, and determining a feature point offset according to the projection error.

[0225] Optionally, when the position determining module 305 calculates the current position of the UAV according to the feature point offset, it is used for:

[0226] obtaining geographical coordinate metadata of a satellite reference image and a UAV camera projection model;

[0227] calculating a pose deviation of the UAV relative to the reference image according to the feature point offset and the camera projection model;

[0228] calculating the longitude, latitude and altitude of the UAV according to the pose deviation and the geographical coordinate metadata, and determining the current position of the UAV according to the longitude, latitude and altitude.

[0229] Optionally, when the position determining module 305 calculates the pose deviation of the UAV relative to the reference image according to the feature point offset and the camera projection model, it is used for:

[0230] establishing an initial pose equation according to the camera projection model;

[0231] randomly selecting a feature point offset and substituting it into the initial pose equation to obtain a rotation matrix and a translation vector;

[0232] determining the rotation matrix and the translation vector as the pose deviation of the UAV relative to the reference image.

[0233] Optionally, when the block dividing module 302 generates a plurality of real shot blocks consistent with the reference block size according to the real-time flight attitude data and the image size, it is used for:

[0234] analyzing the real-time flight attitude data to determine a pitch angle and a roll angle;

[0235] calculating a rotation angle of an image plane relative to a horizontal plane according to the pitch angle and the roll angle;

[0236] determining a camera focal length according to the shooting parameters;

[0237] acquire a real-time flight height of the UAV;

[0238] calculate a side length scaling factor of a block according to the camera focal length, the real-time flight height and the image size;

[0239] perform affine transformation on the real-time corrected image according to the rotation angle and the side length scaling factor, to generate a plurality of real-time blocks consistent with a reference block size.

[0240] Optionally, when the position determining module 305 calculates the longitude, latitude and altitude of the UAV according to the pose deviation and the geographic coordinate metadata, it is used for:

[0241] acquire an angular velocity error threshold of the inertial measurement unit;

[0242] analyze the pose deviation to determine a rotation matrix component;

[0243] determine a pitch angle deviation and a roll angle deviation according to the rotation matrix component;

[0244] if the pitch angle deviation or the roll angle deviation exceeds the angular velocity error threshold, then perform weighted fusion on the rotation matrix according to the attitude angle output by the inertial measurement unit, to generate a corrected rotation matrix;

[0245] calculate the longitude, latitude and altitude of the UAV according to the corrected rotation matrix, the translation vector and the geographic coordinate metadata.

[0246] The system of the embodiment can be used to execute the method of any of the above embodiments, and has similar implementation principles and technical effects, which will not be described here again.

Claims

1. A visual navigation method for unmanned aerial vehicles (UAVs), characterized in that, include: Acquire satellite reference images and real-time images of the target image captured by drones; The satellite reference image is divided into several reference blocks, and the real-shot image to be matched is dynamically divided into blocks to obtain several real-shot blocks; Based on the aforementioned reference blocks and the aforementioned real-shot blocks, the overlapping area between the satellite reference image and the real-shot image to be matched is determined; Perform pixel-level feature matching within the overlapping region to determine feature point offsets, including: Obtain multimodal compatible feature descriptor generation rules; Based on the feature descriptor generation rules, feature points are extracted from the reference block and the actual shooting block in the overlapping region, respectively. Analyze the feature points to determine the descriptor vector distance; Based on the descriptor vector distance, determine the distance variance of the matching point pair; Remove matching pairs with the largest variance within a preset percentage; Calculate the homography transformation matrix based on the remaining descriptor vector distances; The homography transformation matrix is ​​analyzed to determine the projection error, and the feature point offset is determined based on the projection error. The current position of the drone is calculated based on the offset of the feature points.

2. The method according to claim 1, characterized in that, The process involves dividing the satellite reference image into several reference blocks and dynamically segmenting the real-shot image to be matched to obtain several real-shot blocks, including: Obtain the lens distortion coefficient of the drone camera; Based on the lens distortion coefficient, distortion is eliminated in the satellite reference image and the actual captured image to be matched, respectively, to generate a reference corrected image and an actual captured corrected image. The reference correction image is divided into several reference blocks, and the real-shot correction image is dynamically segmented into several real-shot blocks.

3. The method according to claim 2, characterized in that, The process of dynamically segmenting the real-shot corrected image into several real-shot blocks includes: Acquire real-time flight attitude data of the drone and shooting parameters of the drone camera; Based on the shooting parameters, determine the image size of the actual shot correction image; Based on the real-time flight attitude data and the image size, several real-shot blocks with the same size as the reference block are generated.

4. The method according to claim 1, characterized in that, The step of determining the overlapping region between the satellite reference image and the real-shot image to be matched based on the plurality of reference blocks and the plurality of real-shot blocks includes: By analyzing the aforementioned reference blocks and the aforementioned real-shot blocks, a set of reference pixel grayscale values ​​and a set of real-shot pixel grayscale values ​​are obtained. Analyze the set of reference pixel gray values ​​to determine the reference histogram; Analyze the set of grayscale values ​​of the actual captured pixels to determine the actual captured histogram; Calculate the normalized cross-correlation coefficient based on the baseline histogram and the actual image histogram; Construct a similarity matrix based on the normalized cross-correlation coefficients; Analyze the similarity matrix to determine several similarity peaks; The similarity peaks are compared with a similarity threshold, and the blocks with similarity peaks higher than the similarity threshold are determined as the overlapping areas of the satellite reference image and the real-shot image to be matched.

5. The method according to claim 4, characterized in that, The step of calculating the current position of the UAV based on the feature point offset includes: Acquire geographic coordinate metadata of satellite reference images and UAV camera projection models; Based on the feature point offset and the camera projection model, calculate the pose deviation of the UAV relative to the reference image; Based on the pose deviation and the geographic coordinate metadata, the latitude, longitude, and altitude of the UAV are calculated, and the current position of the UAV is determined based on the latitude, longitude, and altitude.

6. The method according to claim 5, characterized in that, The step of calculating the pose deviation of the UAV relative to the reference image based on the feature point offset and the camera projection model includes: Based on the camera projection model, establish the initial pose equation; Randomly select the offset of the feature point and substitute it into the initial pose equation to obtain the rotation matrix and translation vector; The rotation matrix and the translation vector are determined as the pose deviation of the UAV relative to the reference image.

7. The method according to claim 3, characterized in that, The step of generating several real-shot blocks with the same size as the reference block based on the real-time flight attitude data and the image size includes: Analyze the real-time flight attitude data to determine the pitch and roll angles; Calculate the rotation angle of the image plane relative to the horizontal plane based on the pitch angle and the roll angle; Determine the camera focal length based on the shooting parameters; Obtain the real-time flight altitude of the drone; Calculate the block's side length scaling factor based on the camera focal length, the real-time flight altitude, and the image size; Based on the rotation angle and the side length scaling factor, an affine transformation is performed on the real-shot correction image to generate several real-shot blocks with the same size as the reference block.

8. The method according to claim 6, characterized in that, The step of calculating the latitude, longitude, and altitude of the UAV based on the pose deviation and the geographic coordinate metadata includes: Obtain the angular velocity error threshold of the inertial measurement unit; Analyze the pose deviation to determine the rotation matrix components; Based on the rotation matrix components, determine the pitch angle deviation and roll angle deviation; If the pitch angle deviation or the roll angle deviation exceeds the angular velocity error threshold, the rotation matrix is ​​weighted and fused based on the attitude angle output by the inertial measurement unit to generate a corrected rotation matrix. Based on the corrected rotation matrix, the translation vector, and the geographic coordinate metadata, the latitude, longitude, and altitude of the UAV are calculated.

9. A visual navigation system for unmanned aerial vehicles (UAVs), characterized in that, The method applied to any one of claims 1-8 includes: The image acquisition module is used to acquire satellite reference images and real-time images taken by drones for matching. The block division module is used to divide the satellite reference image into several reference blocks and to dynamically split the real-shot image to be matched into several real-shot blocks. An overlap determination module is used to determine the overlap area between the satellite reference image and the real-shot image to be matched based on the plurality of reference blocks and the plurality of real-shot blocks; The offset determination module is used to perform pixel-level feature matching within the overlapping region and determine the offset of feature points; The position determination module is used to calculate the current position of the UAV based on the offset of feature points.

Citation Information

Patent Citations

  • Aircraft visual navigation method based on deep learning matching and Kalman filtering

    CN116518981A

  • Unmanned aerial vehicle image matching and positioning method based on data preprocessing

    CN118262127A