Unmanned aerial vehicle visual positioning method and system based on satellite image map matching

Through the drone visual positioning method based on satellite image map matching, the feature points of the drone image and satellite image are extracted and matched, and combined with homography matrix transformation, the positioning problem of the drone in the GNSS denial environment is solved, and the high-precision and robust positioning effect is achieved.

CN119992140APending Publication Date: 2025-05-13CHANGAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510063783.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing drone visual positioning methods are difficult to achieve high-precision positioning in GNSS denial environments, especially in weak texture areas and complex terrain occlusion environments. They are poorly robust and difficult to deal with lighting changes and viewing angle changes.

Method used

The visual positioning method of drone based on satellite image map matching is adopted. By extracting feature points of drone images and satellite images, and performing feature matching, the texture features of the image are learned, the sensitivity of light changes is reduced, and the robustness of the algorithm is improved. Using homography matrix transformation, we process image angle changes to achieve accurate positioning of the drone.

Benefits of technology

It improves the accuracy and robustness of drone positioning, and can achieve long-distance and high-precision positioning in complex environments, breaking through the limitations of traditional visual positioning methods in weak texture areas and complex terrain occlusion environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992140A_ABST
    Figure CN119992140A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle visual positioning method and system based on satellite image map matching, and belongs to the technical field of unmanned aerial vehicle positioning and image processing. A plurality of image pairs are formed by a plurality of obtained satellite image maps and unmanned aerial vehicle images to be matched; determining a matching number corresponding to the feature points in each image pair, and obtaining a satellite image map corresponding to the maximum matching number; obtaining a matching error of the feature points of the satellite image map corresponding to the maximum matching number, and when the obtained matching error of the feature points of the satellite image map corresponding to the maximum matching number is smaller than a preset matching error threshold value, determining the satellite image map after matching error screening; and according to the satellite image map and the unmanned aerial vehicle image after matching error screening, obtaining corresponding pixel coordinates and latitude and longitude coordinates of the unmanned aerial vehicle image in the satellite image map by determining a homography matrix between the two images. Through the method, accurate positioning of the unmanned aerial vehicle can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more specifically to a method and system for unmanned aerial vehicle visual positioning based on satellite image map matching. Background Art

[0002] UAV positioning is usually achieved using satellite navigation. However, as a passive signal reception method, navigation signals are easily interfered with in special scenarios. When the signal is lost, the cumulative error of the inertial measurement unit will become larger and larger over time. Computer vision processes and analyzes visual information through a computer system to achieve functions such as target detection, recognition, tracking, and positioning, and has strong anti-interference capabilities. Therefore, UAV positioning based on visual matching can well solve the problem of UAV positioning under satellite denial conditions.

[0003] Traditional UAV visual positioning methods can be roughly divided into three types: map-free positioning methods (such as visual odometry), positioning methods based on building maps (such as simultaneous positioning and composition methods), and positioning methods based on existing maps (such as image matching methods). These three visual positioning methods each have their own advantages and disadvantages and scope of application.

[0004] Among them, the methods based on building maps and mapless positioning only require cameras installed on UAVs (Unmanned Aerial Vehicles), but the estimation errors of inter-frame motion will accumulate seriously. The image matching method requires an additional pre-recorded georeferenced image library, but the absolute position of the UAV can be obtained without accumulating errors. The image matching methods are mainly divided into traditional methods and deep learning methods; the traditional method extracts features based on manually designed descriptors to achieve remote sensing image matching, mainly seeking the correspondence between local features (regions, lines, points) through descriptor similarity and / or spatial geometric relationships. The use of local salient features enables such methods to run quickly and be robust to noise, complex geometric deformations, and significant radiometric differences. However, due to the popularity of higher resolution and larger size data, the manually designed descriptor method cannot meet the requirements of more correspondence, higher accuracy, and more flexible application. With the introduction of a large number of labeled data sets, deep learning methods, especially convolutional neural networks (CNNs), have achieved very good results in the field of image matching. The main advantage of CNN is that it can automatically learn features that are conducive to image matching under the guidance of labeled data. Compared with manually designed descriptors, deep learning-based features contain not only low-level spatial information but also high-level semantic information. Due to its powerful ability to automatically extract features, deep learning methods can achieve higher matching accuracy. However, the existing deep learning UAV visual positioning solution is implemented by a method that involves and trains a unified network for feature extraction and feature matching. This method has poor robustness, is difficult to handle scenes with scarce ground information, and has poor stability.

[0005] Domestic research on UAV visual positioning technology in GNSS-denied environments started late. In recent years, with the rapid development of UAV visual positioning technology, research has mainly focused on the following aspects: Visual inertial fusion navigation, which often uses solutions that integrate inertial navigation systems (INS) and visual information. Feature extraction and matching, which is committed to improving the efficiency and accuracy of feature matching. Anti-interference research, some studies focus on positioning problems in complex environments such as low illumination and sparse features.

[0006] Existing drone positioning technologies mainly include:

[0007] 1. GNSS positioning: This method determines the position and attitude of the drone by receiving satellite signals and is the most commonly used drone positioning technology. GNSS positioning has high accuracy, but is easily affected by factors such as terrain obstruction, signal interference, and urban canyons. The positioning accuracy will drop significantly in complex environments.

[0008] 2. Visual odometry (VO), which uses feature matching between adjacent images to calculate camera motion and estimate the position and attitude of the drone. VO methods usually require strong texture information and do not work well in weak texture areas and complex terrain occlusion environments.

[0009] 3. Simultaneous Localization and Mapping (SLAM), which achieves positioning and navigation by building an environmental map. SLAM methods usually require a long running time and are prone to cumulative errors in complex environments, making it difficult to meet real-time positioning requirements.

[0010] 4. Inertial Navigation System (INS), which uses sensors such as accelerometers and gyroscopes to measure the acceleration and angular velocity of the drone, and then infer the position and attitude of the drone. However, INS is susceptible to drift, and long-term use will cause large positioning errors.

[0011] 5. Fusion positioning technology: Fusion of the above positioning technologies to improve positioning accuracy and robustness. For example, fusion of GNSS and INS can effectively reduce the impact of INS drift; fusion of VO and SLAM can improve positioning accuracy and anti-interference ability.

[0012] In summary, since traditional drone positioning for large outdoor scenes mainly focuses on positioning with GNSS information, and in indoor and outdoor areas without GNSS information, the existing visual positioning methods are sensitive to environmental factors such as lighting changes and perspective changes, resulting in reduced positioning accuracy and poor robustness, making it difficult to extract feature points, and making it impossible to perform long-distance, high-precision positioning, affecting matching accuracy. Summary of the invention

[0013] In response to the problems existing in the above-mentioned fields, the present invention proposes a UAV visual positioning method and system based on satellite image map matching. By extracting the feature points of the UAV images and satellite images and performing feature matching, the texture features of the images are learned, the sensitivity to illumination changes is reduced, and the robustness of the algorithm is improved. By calculating the homography matrix transformation of the UAV images and satellite images, the angle of the image taken by the UAV is changed, and the position and posture of the UAV can be accurately calculated, thereby realizing the precise positioning of the UAV.

[0014] In order to solve the above technical problems, the present invention discloses a UAV visual positioning method based on satellite image map matching, comprising the following steps:

[0015] Extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively;

[0016] Multiple satellite image maps are respectively matched with the drone images to be matched to form multiple image pairs; the matching number corresponding to the feature points in each image pair is determined, and the satellite image map corresponding to the maximum matching number is obtained;

[0017] Obtaining the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the obtained maximum matching number is less than a preset matching error threshold, determining the satellite image map after matching error screening;

[0018] According to the satellite image map and UAV image after matching error screening, the pixel coordinates and longitude and latitude coordinates of the UAV image on the satellite image map are obtained by determining the homography matrix between the two images.

[0019] Preferably, the step of extracting feature points of the drone image to be matched; obtaining multiple satellite image maps and extracting feature points of the satellite image maps respectively, specifically includes:

[0020] Obtain multiple satellite images of the drone’s flight area and use a camera with a certain number of frames at intervals to obtain drone images;

[0021] Rotate the drone imagery to match the north-facing orientation of the satellite imagery map profile by providing image metadata for the drone and camera gimbal orientation;

[0022] Each image in the satellite image map is looped through, and the feature points of the drone image and each satellite image map are extracted through the SuperPoint feature extraction network.

[0023] Preferably, obtaining the satellite image map corresponding to the maximum number of matches specifically includes:

[0024] By looping through each photo in the satellite image map, the LightGlue feature matching algorithm is used to obtain the matching number of feature points between the drone image and each satellite image map.

[0025] By comparison, the satellite image map with the largest number of matches corresponding to the feature points of the drone image and each satellite image map is selected as the satellite image map corresponding to the maximum number of matches.

[0026] Preferably, the determining of the satellite image map after matching error screening specifically includes:

[0027] The matching error of the feature points of the satellite image map corresponding to the maximum matching number is determined by the RANSAC algorithm;

[0028] A matching error threshold is preset, and a satellite image map corresponding to a maximum matching number whose matching error of a feature point is smaller than a satellite image map corresponding to the preset matching error threshold is regarded as a successfully matched satellite image map.

[0029] Preferably, obtaining the pixel coordinates and longitude and latitude coordinates corresponding to the drone image on the satellite image map includes the following steps:

[0030] The homography matrix between the satellite image map and the UAV image after matching error screening is determined by the RANSAC algorithm, and the UAV image is perspective transformed to calculate the pixel coordinates of the center point of the UAV image after perspective transformation. The coordinate transformation relationship between the UAV image and the satellite image map is used to convert the pixel coordinates of the center point of the UAV image into the corresponding longitude and latitude coordinates, and obtain the pixel coordinates and longitude and latitude coordinates corresponding to the center point of the UAV image on the satellite image map.

[0031] Preferably, the drone image to be matched is an RGB color image or a full-color image taken vertically or obliquely by a drone in real time.

[0032] Preferably, it also includes a UAV visual positioning system based on satellite image map matching, including:

[0033] The image feature extraction module is used to extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively;

[0034] The image feature matching module is used to form multiple image pairs with multiple satellite image maps and drone images to be matched; determine the matching number corresponding to the feature points in each image pair, and obtain the satellite image map corresponding to the maximum matching number; obtain the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the maximum matching number is less than a preset matching error threshold, determine the satellite image map after matching error screening;

[0035] The positioning module is used to obtain the pixel coordinates and longitude and latitude coordinates of the drone image on the satellite image map according to the satellite image map and drone image after matching error screening by determining the homography matrix between the two images.

[0036] Preferably, a computer device is also included, the computer device comprising a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, the processor executes the following steps:

[0037] Extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively;

[0038] Multiple satellite image maps are respectively matched with the drone images to be matched to form multiple image pairs; the matching number corresponding to the feature points in each image pair is determined, and the satellite image map corresponding to the maximum matching number is obtained;

[0039] Obtaining the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the obtained maximum matching number is less than a preset matching error threshold, determining the satellite image map after matching error screening;

[0040] According to the satellite image map and UAV image after matching error screening, the pixel coordinates and longitude and latitude coordinates of the UAV image on the satellite image map are obtained by determining the homography matrix between the two images.

[0041] Preferably, it further comprises a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the following steps:

[0042] Extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively;

[0043] Multiple satellite image maps are respectively matched with the drone images to be matched to form multiple image pairs; the matching number corresponding to the feature points in each image pair is determined, and the satellite image map corresponding to the maximum matching number is obtained;

[0044] Obtaining the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the obtained maximum matching number is less than a preset matching error threshold, determining the satellite image map after matching error screening;

[0045] According to the satellite image map and UAV image after matching error screening, the pixel coordinates and longitude and latitude coordinates of the UAV image on the satellite image map are obtained by determining the homography matrix between the two images.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] The drone visual positioning method based on satellite image map matching proposed in the present invention, according to the acquired drone image to be registered, firstly extracts the feature points of the drone image to be matched and multiple satellite image maps respectively, uses the drone image to loop through each image in the satellite map, determines the maximum number of matches corresponding to the feature points in multiple image pairs formed by the drone image and the satellite image map, and then obtains the matching error of the feature points of the satellite image map corresponding to the maximum matching number. When the matching error of the feature points of the satellite image map corresponding to the maximum matching number is less than the preset matching error threshold, the satellite image map after matching error screening is determined and used as the successfully matched image. Unlike conventional visual positioning algorithms that are suitable for areas with many feature points such as urban areas, this method extracts the feature points of drone images and satellite images and performs feature matching, breaking through the limitation that traditional visual positioning methods are difficult to accurately position in weak texture areas and complex terrain occlusion environments. In the feature extraction process, by learning the texture features of the image, the sensitivity to illumination changes is reduced and the robustness of the algorithm is improved. By calculating the homography matrix transformation of drone images and satellite images, the coordinate system of the image taken by the drone is converted with the coordinate system of the satellite image map. The homography matrix can handle perspective transformation. Even if the angle of the image taken by the drone changes, the position and posture of the drone can be accurately calculated, thereby achieving precise positioning of the drone. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a flow chart of the UAV visual positioning method based on satellite image map matching proposed by the present invention;

[0049] Figure 2 The model architecture of the SuperPoint feature extraction network provided by the present invention;

[0050] Figure 3 The encoder structure of the SuperPoint feature extraction network provided by the present invention;

[0051] Figure 4 The decoder structure of the SuperPoint feature extraction network provided by the present invention;

[0052] Figure 5 A comparison chart of matching algorithms in the prior art provided by the present invention;

[0053] Figure 6 The image matching framework diagram provided by the present invention;

[0054] Figure 7 A schematic diagram of the homography matrix H of an image provided by the present invention;

[0055] Figure 8A schematic diagram of the translation process of an object in an image in the homography matrix provided by the present invention;

[0056] Fig. 9 A schematic diagram of the translation process of an object in an image in the homography matrix provided by the present invention;

[0057] Fig.10 A schematic diagram of the rotation process of an object in an image in the homography matrix provided by the present invention;

[0058] Fig.11 A schematic diagram of the affine transformation process of an object in an image in the homography matrix provided by the present invention;

[0059] Fig.12 A schematic diagram of the perspective transformation process of an object in an image in the homography matrix provided by the present invention;

[0060] Fig.13 The drone image obtained by the embodiment of the present invention;

[0061] Fig.14 A plurality of satellite image maps obtained by an embodiment of the present invention;

[0062] Fig.15 The feature point extraction result provided by the embodiment of the present invention;

[0063] Fig.16 The feature point matching result obtained by the embodiment of the present invention;

[0064] Fig.17 A schematic diagram of perspective transformation of drone images provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0065] The following will be combined with the attached embodiment of the present invention Figure 1-Figure 17 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms described in the present invention are only used to describe specific implementation methods and are not used to limit the present invention.

[0066] like Figure 1 As shown, the present invention proposes a UAV visual positioning method based on satellite image map matching, comprising the following steps:

[0067] S1: Take the collected image data as input

[0068] By using a camera with a certain number of frames at intervals to capture a drone image, and collect N satellite images of the area related to the drone flight

[0069] Install corresponding sensors on the drone, such as cameras, inertial measurement units (IMUs), etc. These sensors will be used to obtain real-time flight data. The TIF image taken by the drone is a lossless compression format that stores the precise data of each pixel of the image and can preserve high image quality and details. Images in TIF format preserve the precise data of each pixel, and the registration accuracy is usually high. When performing TIF image registration, the ground control point (GCP) and digital elevation model (DEM) can be combined to improve the registration accuracy.

[0070] Obtain high-resolution satellite image map data of the flight area and analyze the landmarks, buildings, roads and other features. Preprocess the acquired satellite image map data, including distortion removal, image overlap adjustment, background removal, etc., to ensure the consistency of the coordinate system of pixel coordinates and spatial coordinates. In the process of digital image registration, there are often problems such as image format and resolution, so resampling of the image is essential.

[0071] Resampling refers to processing multiple images with different resolutions, sizes or pixel spacing to align them on the same scale. The main resampling methods include point sampling, regional average sampling, regional nearest neighbor sampling, etc. Point sampling is fast, but some detail information will be lost; regional average sampling can well preserve the detail information of the image, but it will cause some parts to be blurred; regional nearest neighbor sampling can preserve edge information, but it will cause some parts to be too sharp. The application of different resampling techniques has a strong impact on the registration results. The use of appropriate resampling methods can improve the accuracy of digital image registration.

[0072] S2: Feature extraction, feature matching and positioning process

[0073] Using image metadata that provides the orientation of the drone and camera gimbal, the drone imagery is rotated to match the orientation of the map profile (always facing north).

[0074] Each image in the satellite map is looped through, and the feature points of each pair of drone images and satellite images are extracted through the SuperPoint feature extraction network.

[0075] The LightGlue feature matching algorithm is used to obtain the matching number of feature points between the drone image and each satellite image map. By comparison, the satellite image map with the largest matching number between the feature points of the drone image and each satellite image map is selected as the satellite image map corresponding to the largest matching number, and the matching results are stored.

[0076] Use drone images to loop through each image in the satellite image map, determine the satellite image map corresponding to the maximum matching number according to the maximum matching number, obtain the matching error of the feature points of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature points of the satellite image map corresponding to the obtained maximum matching number is less than a preset matching error threshold, determine the satellite image map after matching error screening.

[0077] The RANSAC algorithm is used to determine the matching error of the feature points of the satellite image map corresponding to the maximum matching number. A matching error threshold is preset, and the satellite image corresponding to the feature point of the satellite image map corresponding to the maximum matching number whose matching error is less than the preset matching error threshold is regarded as a successfully matched satellite image map.

[0078] S3: Based on the successfully matched satellite image map and drone image, by calculating the homography matrix between the two images, using the homography matrix to perform perspective transformation on the drone image, the drone image is matched with the satellite image, and the pixel coordinates and longitude and latitude coordinates corresponding to the drone image in the satellite image map are obtained as output.

[0079] Specifically, in step S1, the collected drone image is an RGB color image or a full-color image taken vertically or obliquely by a drone in real time.

[0080] The SuperPoint feature extraction network in step S2 is a fully convolutional network structure used to generate feature points and feature point descriptors. Figure 2 As shown in the figure, the model architecture of the feature extraction network has a shared encoder to process the input image to reduce the dimension of the input image. After the encoder, there are two decoders, one for feature point detection and the other for generating feature point descriptors.

[0081] Unlike the traditional feature extraction algorithm that detects interest points first and then calculates descriptors, most of the parameters of this feature extraction network are shared between the two tasks. The following will introduce the specific working principles of the encoder, feature point decoder, and descriptor decoder of this feature extraction network.

[0082] 1) Encoder

[0083] The encoder consists of convolutional layers and maximum pooling layers, and each convolutional layer must perform Relu nonlinear activation and BatchNorm (BN) normalization operations. Figure 3As shown in the figure, the encoder has 8 3×3 convolutional layers, the number of channels of the first 4 convolutional layers is 64, the number of channels of the last 4 convolutional layers is 128, and there is a 2×2 maximum pooling layer after every two convolutional layers. The encoder uses 3 maximum pooling layers, and performs 3 2×2 maximum pooling operations in the encoder to produce 8×8 pixel cells (cell refers to the pixel of the low-dimensional output), reducing the image size of h×w to one-eighth of the original, with 128 channels.

[0084] 2) Decoder

[0085] The function of this part of the network is to output each pixel corresponding to the probability that the pixel is a feature point in the input. The principle is as follows Figure 4 As shown in the figure, the input image is first upgraded to (h / 8)×(w / 8)×256 using a 3×3 convolution kernel with 256 channels, and then reduced to (h / 8)×(w / 8)×65 using a 1×1 convolution kernel with 65 channels. Then, a softmax operation is performed, the first 64 channels are taken, and one is sampled every 8 pixels. 64 ultra-low-resolution images can be combined into a high-resolution large image. The score of each pixel in the image is then subjected to non-maximum suppression (NMS) processing, so that each local area of ​​the image obtains a quasi-feature point. Subsequently, the score threshold is used for screening to remove redundant feature points that are very close and retain the scores of feature points with good quality. The output is still the score of the pixel, the score of the quasi-feature point position is retained, and the remaining unknown scores are set to zero.

[0086] like Figure 5 As shown, the horizontal axis Image Pairs Per Second represents the number of matched image pairs per second, and the vertical axis Relative PoseAccuracy represents the relative pose accuracy. By comparing multiple image matching algorithms in the prior art, the present invention finally selects the SuperPoint feature extraction network, which is relatively more suitable for aerial image matching, to extract the features of aerial images, and then selects the LightGlue matching algorithm, a deep network that is more accurate, more efficient, and easier to train than SuperGlue, to match the extracted feature points. Because these two algorithms not only have high accuracy, but also have better adaptability to image scale transformation and rotation transformation than other algorithms.

[0087] The image matching scheme of the present invention is divided into two parts, namely, the feature extraction network and the matching network. The function of the feature extraction network is to extract feature points and feature descriptors, and the function of the image matching network is to use the feature points and descriptor information extracted by the feature extraction network to perform feature matching. The matching process proposed by the present invention is as follows: Figure 6As shown, the aerial drone images and the pre-stored satellite image maps are first input into the SuperPoint feature extraction network to obtain feature points and descriptors, and then the extracted feature points and descriptors are input into the LightGlue image matching algorithm for matching.

[0088] like Figure 7 The figure shows the network architecture of the LightGlue image matching algorithm. Given a pair of input local features d and p, each layer enhances the visual descriptor with the context of self-attention and cross-attention units based on position encoding (⊙). The confidence classifier c helps decide whether to stop reasoning. If few points are confidence, the reasoning continues to the next layer and prunes the points that are confidently mismatched. Once the confidence state is reached, LightGlue predicts the distribution between points based on the local similarity and unary matching of the points.

[0089] In step S3, the present invention matches the feature points of the collected drone images with the satellite image map pre-stored in the aircraft. The regional position information of the drone image taken by the aircraft on the satellite image map can be obtained through image matching, and then the exterior orientation elements of the real-time image are solved by using the homography matrix transformation algorithm through these spatial coordinates. Therefore, in the positioning and solving process, two images are required to participate, one of which is the real-time image taken by the drone, and the other is the satellite image map with known coordinate information, and both images need to have the same overlapping area.

[0090] (1) Definition of homography transformation

[0091] The homography of an image refers to the projection transformation process between two two-dimensional planes. The mapping relationship between the two is the homography of the image. For example, a two-dimensional ground and its imaging plane in the camera is an example of a homography.

[0092] Now define that there is a point in three-dimensional space And the projection point of this point on the imaging device is The mapping between the two can be expressed by matrix multiplication. Using homogeneous coordinates to represent these two points, the following formula holds:

[0093]

[0094] in, represents the homogeneous coordinates in three-dimensional space, Represents the homogeneous coordinates of the two-dimensional image plane. The mapping relationship between the two can be expressed as the homography matrix ( The Z axis is not used when mapping, and the Z axis indicates that the depth of field does not affect the coordinate position during mapping):

[0095]

[0096] Among them, s represents the scaling factor of any scale, and H represents the homography matrix.

[0097] like Figure 8 As shown in Figure 1, the image can be projected onto any plane through the transformation of the homography matrix H. The translation, rotation, and perspective transformation of the target in the image can all be mapped through the homography matrix H.

[0098] In the homography matrix, the translation of the object in the image is Fig. 9 As shown, translating the figure from the origin to another coordinate in the coordinate system can be described by a mathematical expression:

[0099] x′=x+t x ,y′=y+t y

[0100] It can also be expressed using homogeneous coordinates and matrix multiplication:

[0101]

[0102] like Fig.10 As shown in the figure, the target in the image is rotated counterclockwise around the coordinate (0,0). The formula can be expressed as:

[0103] x′=x·cosθ-y·sinθ, y′=x·sinθ+y·cosθ

[0104] It can also be written in the form of matrix multiplication as follows:

[0105]

[0106] like Fig.11 As shown in the figure, it is the affine transformation of the target in the image. Compared with translation and rotation (which do not change the shape of the object), affine transformation not only changes the coordinate position of the target, but also changes the shape of the target, but still maintains the straightness of the target itself, that is, the originally parallel lines remain parallel after affine transformation.

[0107] like Fig.12As shown in the figure, it is the perspective transformation of the target in the image. The perspective transformation not only changes the shape of the target, but also makes the originally parallel lines in the target no longer maintain a parallel relationship, which completely changes the shape of the target. This transformation is similar to observing the plane of a building at different angles. The plane of the building presents different shapes at different angles. In addition, it can be concluded that the perspective transformation (equivalent to the homography transformation) includes the image affine transformation, which has a broader meaning than the affine transformation, can describe more complex image geometric transformations, and can be applied to more complex visual tasks.

[0108] (2) Solving the homography matrix

[0109] According to the basic theory of the image homography matrix mentioned above, the midpoint of the image is now defined to represent the two coordinates that match each other in the two images, and to represent the homography matrix. According to the principle of homography transformation, we have:

[0110]

[0111] Expanding the matrix multiplication, we can get the system of equations:

[0112]

[0113] After transforming and shifting the above formula, we get:

[0114]

[0115] According to linear algebra knowledge, the equation can be transformed into matrix and vector multiplication to obtain:

[0116]

[0117] Where h = [h 11 ,h 12 ,h 13 ,h 21 ,h 22 ,h 23 ,h 31 ,h 32 ,h 33 ] T is a 9-dimensional column vector;

[0118] If:

[0119]

[0120] Then the above equation can be written as:

[0121] Ah=0

[0122] Here A∈R 2×9 , is the matrix equation obtained from a pair of points.

[0123] Since the calculation of the homography matrix uses a homogeneous coordinate system, that is, there is an arbitrary scale scaling factor, when the matrix elements are multiplied by any non-zero constant k, the mapping relationship between the coordinates remains unchanged, so the 3×3 homography matrix has 9 parameters, but in fact has only 8 degrees of freedom. Since the homography matrix H has 8 degrees of freedom, and a pair of points provides two linear equations, at least 4 pairs of points are required to provide 8 linear equations to solve the value of the homography matrix H. Assuming that there are now n ≥ 4 point pairs, we get the matrix A∈R 2×9 , we can directly perform SVD (Singular Value Decomposition) on A to get U*∑*V T , and then take the last column of V as the value of the homography matrix H.

[0124] From the above homography matrix solving process, it can be seen that at least 4 pairs of key point coordinates, any three of which are not collinear, are required to calculate the homography matrix between the two images.

[0125] In traditional algorithms, a large number of coordinate mapping relationships are obtained by matching using local feature descriptors, but these mapping relationships are not all correct. Often, some coordinate matches are erroneous or incorrect, i.e., noise data. In this case, the RANSAC (Random sample consensus) algorithm can be used to eliminate the noise data and estimate the parameters of the optimal homography matrix H.

[0126] The present invention also proposes a UAV visual positioning system based on satellite image map matching, comprising:

[0127] The image feature extraction module is used to extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively;

[0128] The image feature matching module is used to form multiple image pairs with multiple satellite image maps and drone images to be matched; determine the matching number corresponding to the feature points in each image pair, and obtain the satellite image map corresponding to the maximum matching number; obtain the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the maximum matching number is less than a preset matching error threshold, determine the satellite image map after matching error screening;

[0129] The positioning module is used to obtain the pixel coordinates and longitude and latitude coordinates of the drone image on the satellite image map according to the satellite image map and drone image after matching error screening by determining the homography matrix between the two images.

[0130] In summary, the method proposed in the present invention has the following advantages:

[0131] (1) The present invention uses high-resolution satellite image maps as a reference to match the real-time images taken by the drone with the satellite image maps, thereby achieving high-precision positioning of the drone.

[0132] (2) Unlike conventional visual positioning algorithms that are applicable to areas with many feature points such as urban areas, the method proposed in the present invention breaks through the limitation that traditional visual positioning methods are difficult to accurately position in weak texture areas and complex terrain occlusion environments.

[0133] (3) The present invention adopts SuperPoint feature extraction network extraction and LightGlue feature matching algorithm to realize the matching process of drone images and satellite image maps, realizes efficient and accurate image feature extraction and matching, improves the recognition accuracy of feature points, and reduces the sensitivity to environmental factors.

[0134] (4) The coordinate system of the image taken by the UAV is converted with the coordinate system of the satellite image map by using the homography matrix transformation, thereby achieving accurate positioning of the UAV and avoiding complex coordinate system conversion steps.

[0135] (5) The present invention combines deep learning, map matching and homography matrix transformation technology to improve the robustness of the algorithm, enabling it to run stably in a variety of complex environments, especially in weak texture areas and complex terrain occlusion environments.

[0136] Example

[0137] like Fig.13 As shown in FIG. , the drone image to be matched is obtained in this embodiment, such as Fig.14 Figures (a) to (c) in the figure show the three satellite image maps obtained.

[0138] The drone image is input with multiple satellite image maps with longitude and latitude information to form multiple image pairs.

[0139] The drone images used were taken by a drone in a field environment at an altitude of about 120 meters. The drone used was a DJI Matrice 300 equipped with an RGB camera sensor.

[0140] Each image in the three satellite image maps is looped through, and the feature points of the drone image and each satellite image map are extracted through the SuperPoint feature extraction network. The feature point extraction results are as follows: Fig.15 As shown in Figures (a) to (c).

[0141] By looping through each image in the satellite image map, the LightGlue feature matching algorithm is used to obtain the matching numbers of the feature points of the drone image and each satellite image map, which are 119, 4, and 1, respectively. Fig.16 As shown in Figures (a) to (c).

[0142] By comparison, the satellite image map with the largest number of matches between the feature points of the drone image and each satellite image map is selected, that is, Fig.16 Figure (a) in the figure is the satellite image map corresponding to the maximum number of matches.

[0143] Furthermore, by judging whether the matching error of the feature point of the satellite image map corresponding to the maximum matching number is less than the preset matching error threshold, after verification, the matching error of the feature point of the satellite image map corresponding to the maximum matching number is less than the preset matching error threshold, so Fig.16 Figure (a) is the satellite image map that is successfully matched.

[0144] The homography matrix between the satellite image map and the drone image after matching error screening is calculated by the RANSAC algorithm, such as Fig.17 As shown in the figure, the perspective transformation of the drone image is performed. The pixel coordinates of the center point of the drone image after the perspective transformation are calculated. The coordinate transformation relationship between the drone image and the satellite image map is used to convert the pixel coordinates of the center point of the drone image into the corresponding longitude and latitude coordinates, and the pixel coordinates and longitude and latitude coordinates corresponding to the center point of the drone image in the satellite image map are obtained, that is, the drone coordinates.

[0145] In order to improve the accuracy, the random sample consensus (RANSAC) algorithm is used to screen the effective feature points, and only the inner points that satisfy the perspective transformation relationship are used for calculation, and the outer points with large deviations are eliminated. Finally, through these steps, the drone photos can be accurately associated with the geographic coordinates to achieve high-precision positioning.

[0146] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

[0147] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.

Claims

1. A UAV visual positioning method based on satellite image map matching, characterized in that: The following steps are involved: Extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively; Multiple satellite image maps are respectively matched with the drone images to be matched to form multiple image pairs; the matching number corresponding to the feature points in each image pair is determined, and the satellite image map corresponding to the maximum matching number is obtained; Obtaining the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the obtained maximum matching number is less than a preset matching error threshold, determining the satellite image map after matching error screening; According to the satellite image map and UAV image after matching error screening, the pixel coordinates and longitude and latitude coordinates of the UAV image on the satellite image map are obtained by determining the homography matrix between the two images.

2. The method for unmanned aerial vehicle visual positioning based on satellite image map matching according to claim 1 is characterized in that: The step of extracting feature points of the drone image to be matched and obtaining multiple satellite image maps and extracting feature points of the satellite image maps respectively includes: Obtain multiple satellite images of the drone’s flight area and use a camera with a certain number of frames at intervals to obtain drone images; Rotate the drone imagery to match the north-facing orientation of the satellite imagery map profile by providing image metadata for the drone and camera gimbal orientation; Each image in the satellite image map is looped through, and the feature points of the drone image and each satellite image map are extracted through the SuperPoint feature extraction network.

3. The method for unmanned aerial vehicle visual positioning based on satellite image map matching according to claim 2 is characterized in that: The obtaining of the satellite image map corresponding to the maximum number of matches specifically includes: By looping through each photo in the satellite image map, the LightGlue feature matching algorithm is used to obtain the matching number of feature points between the drone image and each satellite image map. By comparison, the satellite image map with the largest number of matches corresponding to the feature points of the drone image and each satellite image map is selected as the satellite image map corresponding to the maximum number of matches.

4. The method for unmanned aerial vehicle visual positioning based on satellite image map matching according to claim 3 is characterized in that: The determining of the satellite image map after matching error screening specifically includes: The matching error of the feature points of the satellite image map corresponding to the maximum matching number is determined by the RANSAC algorithm; A matching error threshold is preset, and a satellite image map corresponding to a maximum matching number whose matching error of a feature point is smaller than a satellite image map corresponding to the preset matching error threshold is regarded as a successfully matched satellite image map.

5. The method for unmanned aerial vehicle visual positioning based on satellite image map matching according to claim 4 is characterized in that: The method of obtaining pixel coordinates and longitude and latitude coordinates corresponding to the drone image on the satellite image map includes the following steps: The homography matrix between the satellite image map and the UAV image after matching error screening is determined by the RANSAC algorithm, and the UAV image is perspective transformed to calculate the pixel coordinates of the center point of the UAV image after perspective transformation. The coordinate transformation relationship between the UAV image and the satellite image map is used to convert the pixel coordinates of the center point of the UAV image into the corresponding longitude and latitude coordinates, and obtain the pixel coordinates and longitude and latitude coordinates corresponding to the center point of the UAV image on the satellite image map.

6. The method for unmanned aerial vehicle visual positioning based on satellite image map matching according to claim 5 is characterized in that: The drone image to be matched is an RGB color image or a full-color image taken vertically or obliquely by a drone in real time.

7. A UAV visual positioning system based on satellite image map matching, characterized in that: include: Image feature extraction module, used to extract feature points of the drone image to be matched; Obtain multiple satellite image maps, and extract feature points of the satellite image maps respectively; The image feature matching module is used to form multiple image pairs with multiple satellite image maps and drone images to be matched; determine the matching number corresponding to the feature points in each image pair, and obtain the satellite image map corresponding to the maximum matching number; obtain the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the maximum matching number is less than a preset matching error threshold, determine the satellite image map after matching error screening; The positioning module is used to obtain the pixel coordinates and longitude and latitude coordinates of the drone image on the satellite image map according to the satellite image map and drone image after matching error screening by determining the homography matrix between the two images.

8. A computer device, characterized in that: The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: Extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively; Multiple satellite image maps are respectively matched with the drone images to be matched to form multiple image pairs; the matching number corresponding to the feature points in each image pair is determined, and the satellite image map corresponding to the maximum matching number is obtained; Obtaining the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the obtained maximum matching number is less than a preset matching error threshold, determining the satellite image map after matching error screening; According to the satellite image map and UAV image after matching error screening, the pixel coordinates and longitude and latitude coordinates of the UAV image on the satellite image map are obtained by determining the homography matrix between the two images.

9. A computer-readable storage medium, characterized in that: The computer readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor performs the following steps: Extract the feature points of the drone image to be matched; obtain multiple satellite image maps and extract the feature points of the satellite image maps respectively; Multiple satellite image maps are respectively matched with the drone images to be matched to form multiple image pairs; the matching number corresponding to the feature points in each image pair is determined, and the satellite image map corresponding to the maximum matching number is obtained; Obtaining the matching error of the feature point of the satellite image map corresponding to the maximum matching number, and when the matching error of the feature point of the satellite image map corresponding to the obtained maximum matching number is less than a preset matching error threshold, determining the satellite image map after matching error screening; According to the satellite image map and UAV image after matching error screening, the pixel coordinates and longitude and latitude coordinates of the UAV image on the satellite image map are obtained by determining the homography matrix between the two images.

Citation Information

Cited By

  • Intelligent patrol method, device and equipment for ground feature risk of power transmission channel and storage medium

    CN120218637A

  • Unmanned aerial vehicle visual navigation positioning method and device based on multi-modal feature fusion

    CN120576778A

  • Device and method for obtaining target latitude and longitude of unmanned vehicle

    TWI936928B