Unmanned aerial vehicle image alignment method based on unsupervised learning

Through the unsupervised learning method, the image data set is constructed and the image alignment of the drone is aligned using an encoder and decoder combined with a differentiable transformation module, which solves the problem of insufficient robustness and generalization capabilities in the prior art, and achieves high accuracy and low cost image alignment.

CN120374925APending Publication Date: 2025-07-25NORTHWEST ELECTROMECHANICAL ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340795.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing drone image alignment methods are poorly robust on images with poor texture information, the deep learning-based methods are not generalized, and require manual annotation or forgery of labels, which is costly and insufficient accuracy.

Method used

The unsupervised learning method is adopted to construct the image data set and expand and enhance it. The encoder and decoder combine differentiable linear transformation module and differentiable single-strain transformation module are used to align images through an unsupervised learning network to avoid manual annotation and forge labels. The ConvNeXt convolution module is used to extract features and optimize the model using smooth-l1 loss function.

Benefits of technology

End-to-end drone image alignment is realized, improving alignment accuracy and generalization capabilities, reducing the impact of the model on outliers, reducing computational costs, and adapting to image alignment under different scenarios and weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374925A_ABST
    Figure CN120374925A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle image alignment method based on unsupervised learning. The method comprises the following steps: constructing an image data set; samples in the image data set are image pairs, and the image pairs comprise to-be-aligned images and target images; carrying out expansion and enhancement on the image data set, and carrying out size unification on the expanded and enhanced image pair; constructing input of an unmanned aerial vehicle image alignment model; constructing an unmanned aerial vehicle image alignment model by using an unsupervised learning network; the unmanned aerial vehicle image alignment model comprises an encoder, a decoder, a differentiable linear transformation module and a differentiable homography transformation module; and training the unmanned aerial vehicle image alignment model by using the image data set, and storing the trained unmanned aerial vehicle image alignment model for alignment of the to-be-aligned image pair. According to the method, manual marking is not needed, and the unmanned aerial vehicle images can be aligned robustly under different scenes, weather and image noise interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image processing and deep learning, and particularly relates to an unmanned aerial vehicle (UAV) image alignment method based on unsupervised learning. Background Art

[0002] Using UAVs helps to survey targets, confirm target locations, or conduct mission evaluations. UAVs are convenient to operate, have strong data collection capabilities, and are flexible to use. For UAV navigation, the Beidou satellite navigation plays a very important role. However, the GNSS system is very prone to failure in some scenarios, such as when approaching obstacles or encountering signal interference from the other party. At this time, navigation completely relies on the internal inertial navigation device of the UAV for its own state estimation. However, the accumulated time drift will bring large positioning errors, making UAV positioning completely unavailable.

[0003] Thus, it is very important to match and align the images captured by the UAV camera with satellite images. The existing image alignment methods can be divided into two categories. The first category belongs to the traditional method, which calculates the homography transformation matrix by finding matching feature points on two images. The second category belongs to the method based on deep learning, which enables the network to have the ability to learn the homography transformation by forging homography transformation labels. The former will have matching errors and poor robustness on images with insufficient texture information, and there is a certain gap between the forged labels and the real homography transformation in the latter, resulting in poor generalization ability. Summary of the Invention

[0004] The purpose of the present invention is to provide an unmanned aerial vehicle (UAV) image alignment method based on unsupervised learning, which can achieve the alignment task end-to-end without any annotation, and improve the alignment accuracy and the generalization ability of the network.

[0005] To achieve the above task, the present invention adopts the following technical solutions:

[0006] An unmanned aerial vehicle (UAV) image alignment method based on unsupervised learning, comprising:

[0007] Step 1, constructing an image data set; the samples in the image data set are image pairs, and the image pairs include the image to be aligned and the target image;

[0008] Step 2, expanding and enhancing the image data set, and unifying the sizes of the expanded and enhanced image pairs;

[0009] Step 3, constructing the input of the UAV image alignment model; the first part of the input is the image to be aligned in the image pair; the second part of the input is to randomly generate a square block at the same position of the image to be aligned and the target image, and use the square block to intercept the image to be aligned and the target image respectively to obtain the corresponding square sub-images, and use the square sub-images as the input; the third part of the input is the coordinates of the four corner points of the square block.

[0010] Step 4: Construct a UAV image alignment model using an unsupervised learning network; the UAV image alignment model includes an encoder, a decoder, a differentiable linear transformation module, and a differentiable homography transformation module, where:

[0011] The encoder takes the square sub - graph as input and outputs an output feature map containing spatial position information and image feature information;

[0012] The decoder is used to recover the image feature information contained in the output feature map of the encoder, and at the same time obtain a perspective transformation field output by the decoder according to the spatial position information learned from the encoder;

[0013] The differentiable linear transformation module determines the position transformation differences corresponding to the coordinates of the four corner points according to the perspective transformation field, and estimates a homography transformation matrix based on the position change differences;

[0014] The differentiable homography transformation module is used to perform differentiable processing on the homography transformation matrix, invert the differentiable - processed homography transformation matrix to obtain an inverse matrix; use the inverse matrix to transform the first - part input image to be aligned, and then intercept the positions corresponding to the coordinates of the four corner points of the third - part input on the transformed image to obtain a predicted image, and calculate the pixel loss through the predicted image and the square sub - graph corresponding to the target image, so as to regress the entire unsupervised learning network;

[0015] Step 5: Train the UAV image alignment model using an image data set, and save the trained UAV image alignment model for aligning image pairs to be aligned.

[0016] Furthermore, there are two methods for constructing the image pair:

[0017] The first is an image pair composed of the captured image of the UAV camera and the corresponding satellite image, where the captured image is used as the image to be aligned; first, calculate the rotation matrix of the UAV camera using the attitude information of the UAV itself, and based on the principle of pinhole imaging using the rotation matrix, obtain the position of the pixels in the captured image of the UAV camera on the satellite image, and intercept the satellite image using this position. The intercepted image and the captured image form the image pair;

[0018] The second is on the video captured by the UAV camera, forming an image pair with the current - frame video image and the video image at a preset number of frames interval.

[0019] Furthermore, the augmentation enhancement includes vertical mirroring, horizontal mirroring, reducing brightness, or increasing brightness; the horizontal mirroring is to swap the left and right halves of the image to be aligned and the target image in the image pair with the vertical central axis of the image as the central axis, while the vertical mirroring is to swap the upper and lower halves of the image with the horizontal central axis of the image as the central axis; increasing the brightness is to add the weights of the image to be aligned and the target image in the image pair with an image whose pixel values are all 0; reducing the brightness is to multiply the pixel values of the RGB three channels of the image to be aligned and the target image by the weight.

[0020] Furthermore, the square sub-images P A corresponding to the image I to be aligned B and the target image I A are respectively grayscale processed into single-channel images and then spliced into a grayscale image of H×W×2 as the input of the encoder; where H and W respectively represent the height and width of P B , P A . B

[0021] The encoder is based on the ConvNeXt convolution module and has a four-layer network structure; the first layer is a convolutional layer; the second to fourth layers are all composed of three cascaded ConvNeXt convolution modules; in the ConvNeXt convolution module, the input feature map is first processed by a depthwise separable convolution, and then passes through two convolutional layers and the output after the activation function GELU is fused with the input feature map to obtain the output feature map of the ConvNeXt convolution module.

[0022] Furthermore, the perspective transformation field is the position transformation difference of all pixel points on the horizontal and vertical coordinates of the square sub-images P A , P B ; two corresponding two-dimensional feature matrices are respectively constructed using the position transformation differences on the horizontal and vertical coordinates and and and are spliced to form the perspective transformation field.

[0023] Furthermore, the decoder has five network layers, where the first layer is composed of nine cascaded ConvNeXt convolution modules, and the second to fourth layers are all composed of three cascaded ConvNeXt convolution modules; different from the encoder, in all ConvNeXt convolution modules of the decoder, the convolution modules used for downsampling are replaced with transposed convolution modules; the fifth layer of the decoder is two convolutional layers.

[0024] Furthermore, the two two-dimensional feature matrices and In it, the coordinates P of the four corner points in the third part of the input are respectively selected 4point The position transformation differences on the corresponding abscissa and ordinate are denoted as (Δx 4point , Δy 4point );

[0025] The differentiable linear transformation module first restores the coordinates according to the position transformation differences (Δx 4point , Δy 4point ) and the coordinates P of the four corner points in the third part of the input 4point ;

[0026] {x′ i , y′ i} = {x i - Δx i , y i - Δy i}

[0027] where x i , y i respectively represent the coordinates of the i-th corner point, and x′ i , y′ i represent the restored coordinates of the i-th corner point, and Δx i , Δy i are the position transformation differences on the abscissa and ordinate corresponding to the i-th corner point;

[0028] Subsequently, it is necessary to solve:

[0029]

[0030] where α is the scale factor, and h k (k = 1, 2,..., 9) are the elements in the homography transformation matrix;

[0031] After eliminating the scale factor in the above formula, it is written in the form of matrix multiplication where:

[0032]

[0033] where the superscript T represents the transpose; the input for solving the homography transformation matrix using the DLT algorithm is the coordinate point pair {x i , x′ i}, and the output is the homography transformation matrix that satisfies

[0034] For each of the four corner points with coordinates P 4point , a matrix A i can be constructed, and the matrices A of the four corner points i ​Perform splicing to obtain a matrix A; perform SVD decomposition on matrix A, and the smallest eigenvalue vector obtained by solving is

[0035] Furthermore, the differentiable homography transformation module transforms the homography matrix The process of performing differentiable processing is as follows:

[0036] Normalize the height and width coordinates of the image I to be aligned A and the target image I B to a certain range; subsequently, it is necessary to construct a coordinate grid G = {G B} of the same size as the target image I i , and each element G i in the coordinate grid G corresponds to a pixel coordinate in I B , that is

[0037] Sample in the coordinate grid G = {G i}, and denote the sampled transformed image as an image V with C channels, width W′, and height H′:

[0038]

[0039] where H and W are the height and width of the image I to be aligned A respectively, is the value at the horizontal and vertical coordinate positions (n, m) in the c-th channel of the image I to be aligned A , V i c is the output pixel value at the i-th pixel position (u i , v i ) in the c-th channel, and the V i c of all channels constitute the image V; k(·) is a differentiable function, and Φ p , Φ q are the parameters of the differentiable function; choosing to use the bilinear interpolation function max(·) as the differentiable function, it can be written as:

[0040]

[0041] Then there is the following partial differential calculation formula:

[0042]

[0043] That is, the derivative can be calculated through the chain rule, thus constructing a differentiable sampling mechanism.

[0044] Furthermore, use smooth-l1 as the loss function of the UAV image alignment model, that is:

[0045]

[0046] A terminal device includes a processor, a memory, and a computer program stored in the memory; when the processor executes the computer program, the method for aligning drone images based on unsupervised learning is implemented.

[0047] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the method for aligning drone images based on unsupervised learning is implemented.

[0048] Compared with the prior art, the present invention has the following technical features:

[0049] 1. The present invention belongs to unsupervised learning. It neither requires manual annotation of the homography transformation matrix nor the construction of a labeled dataset by forging labels, avoiding the drawback that forged labels cannot truly reflect the homography transformation, and enabling end-to-end training without any post-processing.

[0050] 2. When estimating the homography matrix transformation, the present invention adopts the method of sampling vertices from the perspective transformation field. Predicting the perspective transformation field contains richer information than directly predicting the regression of four points, which can make the prediction more accurate. Sampling vertices can avoid the huge video memory overhead brought by the RANSAC algorithm.

[0051] 3. In the embodiment of the present invention, the encoding-decoding structure adopts the ConvNeXt network structure, which uses separable convolutions for grouped feature extraction conducive to feature fusion, and uses large convolution kernels to help obtain a larger receptive field. Description of the Drawings

[0052] Figure 1 It is a schematic diagram of the ConvNeXt convolution module.

[0053] Figure 2 It is a schematic diagram of the network architecture of the drone image alignment model.

[0054] Figure 3 It is a schematic diagram of using the trained model to align the images captured by the drone. Detailed Embodiments

[0055] The inventor has reviewed a large number of captured images and analyzed that the context relationship of the global image content in the captured images is crucial for inferring the perspective transformation matrix. However, traditional image alignment techniques, such as SIFT, SURF, ORB, etc., first find the most matching feature points based on features and then obtain the perspective transformation matrix. In this way, it is prone to failure in images with less texture features and noise. Some deep learning-based key point matching methods must label the corresponding points and require a large data scale, which undoubtedly brings huge cost overhead. Some perspective matrix estimation methods directly regress the coordinate differences of four points in the image and estimate the perspective matrix through these four sampled points. This approach is greatly affected by outliers in the prediction results, and moreover, the correlation between channels is ignored in these methods, thus unable to achieve the best prediction results. Although some methods use the method of forging datasets, there are still certain differences between the forging method and the real homography transformation, and the homography transformation cannot be accurately predicted.

[0056] The present invention adopts an unsupervised method to estimate the perspective transformation matrix. Compared with using supervised learning, it truly realizes end-to-end training without any post-processing process. Moreover, by sampling the perspective transformation field, the problem of gradient explosion caused by RANSAC is eliminated, the influence of the model on outliers is reduced, thereby improving the accuracy of estimating the perspective matrix and making the alignment effect of the model more accurate.

[0057] The present invention provides an unmanned aerial vehicle (UAV) image alignment method based on unsupervised learning, which can robustly align UAV images under different scenarios, weather conditions, and image noise interferences. The method includes the following steps:

[0058] Step 1, construct an image dataset; the samples in the image dataset are image pairs, and each image pair includes an image to be aligned and a target image; there are two methods for constructing the image pairs:

[0059] (1) The first one is the image pair composed of the captured image of the UAV camera and the corresponding satellite image, where the captured image is used as the image to be aligned; the UAV image alignment model is trained through the image pair to achieve the alignment of the captured image and the satellite image. Specifically as follows:

[0060] The UAV locates its position through inertial navigation information, including attitude solution and position solution. The attitude solution is output according to the parameters of the three-axis accelerometer and the three-axis magnetometer when the UAV is stationary. Based on this information, the initial attitude angle of the UAV is deduced. Subsequently, when the UAV is flying, the pitch angle, roll angle, and yaw angle are obtained through the integration of the gyroscope and the three-axis accelerometer over time, and thus the rotation matrix R of the UAV camera can be obtained. i ,

[0061]

[0062] Among them, α, β, and γ are the pitch angle, roll angle, and yaw angle respectively, and ω x , ω y , ω z respectively represent the angular accelerations of the UAV in the x-axis, y-axis, and z-axis directions, and α′, β′, and γ′ are the obtained camera matrix parameters.

[0063] Among them, the rotation matrix can be R i which can be expressed in the following form:

[0064]

[0065] Among them, a1, a2, a3, b1, b2, b3, c1, c2, and c3 are the elements of the first to third rows of the rotation matrix respectively.

[0066] After obtaining the rotation matrix R i position calculation can be performed. According to the principle of pinhole imaging, the position of the pixels in the captured image of the UAV camera on the satellite image can be obtained:

[0067]

[0068] Among them, f represents the focal length of the UAV camera, X i , Y i represents the coordinate of a pixel point on the captured image, X A , Y A , Z A represents the three-dimensional position coordinates of the pixel point coordinate on the satellite image, X S , Y S , Z S represents the position of the three-dimensional projection plane selected during conversion.

[0069] The position area of the pixels in the captured image on the satellite image is intercepted, and the intercepted image and the captured image form an image pair.

[0070] (2) The second method is on the video captured by the UAV camera. The current frame video image f c and the video image at an interval of a preset number of frames form an image pair; since when the UAV is flying, the viewing angle of the UAV will change, but the content of the video image captured is generally unchanged, so the video image after an interval of a preset number of frames can be used as the image to be aligned and paired with the target image f c to train the UAV image alignment model. In this embodiment, the preset number of frames is 50 frames.

[0071] Step 2, augment and enhance the image dataset; the augmentation and enhancement include vertical mirroring and horizontal mirroring, reducing brightness or increasing brightness; and unifying the sizes of the augmented and enhanced image pairs.

[0072] The horizontal mirroring is to swap the left and right halves of the image to be aligned and the target image in the image pair with the vertical central axis of the image as the central axis, while the vertical mirroring is to swap the upper and lower halves of the image with the horizontal central axis of the image as the central axis.

[0073] The brightness increase is to add the weights of the image to be aligned and the target image in the image pair and an image with all pixel values of 0, which can be described by the following formula:

[0074] dst = src1·α + src2·β + γ

[0075] Where src1 represents the image to be aligned or the target image before processing, src2 represents the image with all pixel values of 0, α represents the weight of the image to be aligned or the target image, β represents the weight of the image with all pixel values of 0, and γ is a scalar added to the sum, which is equivalent to the brightness adjustment; where the value of α is greater than 1 to achieve the brightness increase effect.

[0076] The brightness reduction is to multiply the pixel values of the three RGB channels of the image to be aligned and the target image by the weight ω, where the value of ω needs to be less than 1, and the smaller the value, the darker the image and the lower the brightness.

[0077] New image pairs are generated through augmentation and enhancement, realizing sample augmentation and data enhancement of the image dataset.

[0078] The unified size is to adjust the images to be aligned and the target images in all augmented and enhanced image pairs to a unified size of 240×320, which is convenient for the subsequent training of the UAV image alignment model and the forgery of perspective transformation.

[0079] Step 3, construct the input of the UAV image alignment model; the input includes three parts:

[0080] The first part of the input is the image to be aligned I in the image pair A .

[0081] The second part of the input is to randomly generate a square block of size 128x128 at the same position in the image to be aligned I A and the target image I B , and use this square block to intercept the image to be aligned I A and the target image I B respectively to obtain square sub-images P A , P B , and P A, P B as the input;

[0082] The third part of the input is the coordinates P of the four corner points of the square block 4point ; among which the coordinates of the upper left corner of the square block are and the coordinates of the remaining three vertices are which together serve as P 4point .

[0083] Step 4, construct a UAV image alignment model using an unsupervised learning network; the UAV image alignment model includes an encoder, a decoder, a differentiable linear transformation module, and a differentiable homography transformation module.

[0084] (1) Encoder.

[0085] Gray-process the square subgraphs P A , P B into single-channel images and splice them into a grayscale image of H×W×2 as the input of the encoder. Use the ConvNeXt convolution module in the encoder to calculate the correlation between channels through a series of depthwise separable convolution modules, and adopt large convolution kernels to increase the receptive field for forming the final output feature map.

[0086] The encoder is based on the ConvNeXt convolution module and has a four-layer network structure; the first layer is a 4×4 convolutional layer with 96 convolution kernels and a convolution stride of 4; the second to fourth layers are all composed of three cascaded ConvNeXt convolution modules; the structure of the ConvNeXt convolution module is as Figure 1 shown. After the input feature map is first processed by a depthwise separable convolution, it sequentially passes through two convolutional layers, and the output after the activation function GELU is fused with the input feature map to obtain the output feature map of the ConvNeXt convolution module. The specific number of features in each network layer of the encoder is shown in Table 1.

[0087] For the network layer m = {3, 4}, the network layer output is a feature map in the form of W m ×H m ×C m , C m = 3×2 m+2 ; while when the network layer m = {1, 2}, the output feature map dimensions are both 96; where H and W respectively represent the height and width of P A , P B , and C represents the dimension of the feature map; W m , H m , C m respectively represent the width, height, and number of channels of the feature map output by the m-th layer.

[0088] Table 1 Encoder Design Details

[0089]

[0090]

[0091] where n×n represents the size of the convolution kernel of the convolutional layer used, dn×n represents the depthwise separable convolution with a size of n, stride represents the stride of the convolution, the number after the convolution kernel represents the number of convolution kernels, and one [·] represents a ConvNeXt structure.

[0092] (2) Decoder.

[0093] The decoder is used to restore the image feature information contained in the output feature map of the encoder, and at the same time, according to the spatial position information learned from the encoder, so as to obtain the perspective transformation field (PF) output by the decoder, that is, the square sub-image P A 、P B The position transformation differences of all pixel points on the horizontal and vertical coordinates of ; the corresponding two-dimensional feature matrices are respectively constructed using the position transformation differences on the horizontal and vertical coordinates and and and are concatenated to form the perspective transformation field (PF); specifically as follows:

[0094]

[0095] where, respectively represent the position transformation differences on the horizontal and vertical coordinates of the pixel at the i-th row and j-th column of the square sub-image P A 、P B i = 1, 2..., H; j = 1, 2,..., W.

[0096] Therefore, the decoder needs to gradually restore the size of the input image from the downsampled feature map. To achieve this goal, the decoder has a symmetric structure similar to the encoder, with five network layers. The first layer consists of 9 cascaded ConvNeXt convolution modules, and the second to fourth layers each consist of three cascaded ConvNeXt convolution modules; different from the encoder, in all ConvNeXt convolution modules of the decoder, the convolution modules used for downsampling are replaced with transposed convolution modules; then in each network layer of the decoder, the low-resolution feature map can be passed to the high-resolution feature map through the transposed convolution module; the fifth layer of the decoder is two convolutional layers.

[0097] The purpose of the fifth layer of the design is to output a perspective transformation field with the same size as the input, that is, the shape of the vector of the final prediction result should also be H×W×2. Based on the fourth layer, an additional convolutional block is added. After the output 64-dimensional features are upsampled to 512, 1x1 convolution is used to downsample them to 2.

[0098] The decoder structure in the embodiments of the present invention is shown in Table 2. For the nth layer of the decoder, the size of the output feature map of each layer is W n ×H n ×C n , where n = {1,....,4}; C n = 3×2 8-n . Among them, H n , W n , C n respectively represent the height, width, and feature dimension of the output feature map.

[0099] Table 2 Decoder Design Details

[0100]

[0101] Different from the currently popular encoder-decoder model, the structures of the encoder and decoder networks provided by the present invention do not, like in solving the segmentation task, such as adding long skip connections between the encoder and decoder in UNet to transfer pixel information from the shallow network to the deep network. This is because the images before and after the homography transformation have large position differences, and the position of a pixel often changes between the input and output, so the skip connection has little effect.

[0102] (3) Differentiable linear transformation module.

[0103] From the two two-dimensional feature matrices and contained in the perspective transformation field PF output from the decoder, the position transformation differences in the horizontal and vertical coordinates corresponding to the coordinates P 4point of the four corner points in the third part of the input are respectively denoted as (Δx 4point , Δy 4point ).

[0104] The input of the differentiable linear transformation module is (Δx 4point , Δy 4point ), and the homography transformation matrix is estimated through differentiable linear transformation. This process involves differentiable linear transformation. For the gradient to propagate, this process needs to be differentiable:

[0105] First, according to the position transformation differences (Δx 4point , Δy 4point) and the coordinates P of the four corner points in the third part of the input 4point , restore the coordinates;

[0106] {x′ i , y′ i} = {x i - Δx i , y i - Δy i}

[0107] where x i , y i respectively represent the coordinates of the i-th corner point, x′ i , y′ i represent the restored coordinates of the i-th corner point, Δx i , Δy i are the position transformation differences on the abscissa and ordinate corresponding to the i-th corner point.

[0108] Subsequently, it is necessary to solve:

[0109]

[0110] where α is the scale factor, h k (k = 1, 2,..., 9) are the elements in the homography transformation matrix.

[0111] After eliminating the scale factor from the above formula, it is written in the form of matrix multiplication where:

[0112]

[0113] where the superscript T represents the transpose; obviously h9 = 1, there are 8 unknowns in the above formula, and four pairs of points are needed to solve this equation, which can be transformed into a homogeneous linear least squares problem. The general method to solve such problems is singular value decomposition.

[0114] Using the DLT algorithm to solve the homography transformation matrix, the input is the coordinate point pair {x i , x′ i}, and the output is the homography transformation matrix that satisfies Using the previous method, for each corner point in the coordinates P 4point , a matrix A i can be constructed. The matrices A i of the four corner points are concatenated to obtain a matrix A; perform SVD decomposition on matrix A, and the eigenvector corresponding to the smallest eigenvalue is The matrix size is 3x3.

[0115] (4) Differentiable homography transformation module.

[0116] A differentiable homography transformation module for performing a homography transformation matrix to perform differentiable processing, and for the homography transformation matrix after differentiable processing to find the inverse to obtain the inverse matrix H inv ; using the inverse matrix H inv to transform the to-be-aligned image I input in the first part A , and then intercepting at the positions corresponding to the four corner point coordinates P input in the third part on the transformed image 4point to obtain the predicted image Through the predicted image and the corresponding square sub-image P of the target image I B to calculate the pixel loss, thereby regressing the entire unsupervised learning network. B

[0117] To make the homography transformation matrix solved by the differentiable linear transformation module differentiable, it is necessary to establish a mapping between positions and pixel values, and this mapping must be differentiable:

[0118] Normalize the height and width coordinates of the to-be-aligned image I A and the target image I B to a range [-1, 1]; subsequently, it is necessary to construct a coordinate grid G = {G B} of the same size as the target image I i , and each element G in the coordinate grid G i corresponds to a pixel coordinate in I B , that is

[0119] Therefore, it is necessary to sample in the coordinate grid G = {G i}, and denote the sampled transformed image as an image V with C channels, width and height of W′, H′:

[0120]

[0121] where H and W are respectively the height and width of the to-be-aligned image I A , is the value at the horizontal and vertical coordinate positions (n, m) in the channel c of the to-be-aligned image I A , V i c is the output pixel value at the i-th pixel position (u i , v i ) in the channel c, and all channels' V i c i.e., constitute the image V; k(·) is a differentiable function, Φ p , Φ qis a parameter of a differentiable function; in this solution, the bilinear interpolation function max(·) is selected as the differentiable function, and it can be written as:

[0122]

[0123] As can be seen from the above formula, after passing through the max(·) function, the pixel coordinates with a distance less than 1 from the output will be output, and the pixel coordinates closer to the distance are assigned higher weights, that is, the eigenvalues of the four points around the pixel coordinates are used to calculate the final score. Because the max(·) function is differentiable, the following partial differential calculation formula can be obtained:

[0124]

[0125] That is, the derivative can be calculated through the chain rule, thus constructing a differentiable sampling mechanism. The gradient loss can not only flow to the feature map but also to the sampled pixel coordinates; in this way, the process of homography transformation is differentiable.

[0126] (5) Deep learning calculates the loss function of PF features.

[0127] For dense prediction problems, the commonly used loss function is the l2 loss function, which is used to calculate the Euclidean distance between the prediction result and the true label; but for this solution, since the outliers have too much influence on the loss value, l2 loss is not very suitable for use in this problem. Therefore, smooth-l1 is used as the loss function of the UAV image alignment model, that is:

[0128]

[0129] Use PyTorch to build the UAV image alignment model and train it using the image dataset; use Adamw as the optimizer, and the parameters in the Adamw optimizer are the default values, that is, β1 = 0.9, β2 = 0.99, momentum = 0.9. Use the initial learning rate of 1×10 -4 as the initial learning rate, and decay the learning rate to 0.1 after every 40 epochs. The size of each batch is 32, and a total of 200 epochs are trained.

[0130] Step 5, in practical applications, for a set of image pairs to be aligned, denote the image to be aligned as I ori and the target image as I target . Use the trained UAV image alignment model to implement the alignment process of the image to be aligned I ori and the target image I target , specifically as follows:

[0131] (1) According to the dimension unification method in step 2, adjust the image I to be aligned ori and the target image I target to be grayscale images after being converted to the unified size, and construct their corresponding square sub-images P A 、P B and obtain the coordinates of the four corner points P 4point , thereby constructing the three-part input of the model.

[0132] (2) Use the trained UAV image alignment model to calculate the homography transformation matrix and obtain its inverse matrix H inv , use the inverse matrix H inv to perform perspective transformation on the image I to be aligned ori (multiply each pixel point in the image I to be aligned ori by the inverse matrix), that is, align the image I to be aligned ori to the target image I target .

[0133] Based on the use of the ConvNeXt network, the present invention designs a brand-new decoder structure. In order to learn the transformation between channels, the separable convolution method is adopted, which can better learn the correlation between channels, so that the prediction of the homography transformation can be more accurate.

[0134] The smooth-l1 loss function used in the present invention overcomes the non-differentiability of the l1 loss function at 0, which may affect convergence, and overcomes the shortcoming that the l2 loss function is too sensitive to outlier eigenvalues. With an observable convergence speed, it can still achieve an excellent alignment effect.

[0135] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. An unmanned aerial vehicle image alignment method based on unsupervised learning, characterized in that Including: Step 1: Construct an image dataset; The samples in the image dataset are image pairs, and each image pair contains an image to be aligned and a target image; Step 2: Augment and enhance the image dataset, and unify the sizes of the augmented and enhanced image pairs; Step 3: Construct the input of the UAV image alignment model; The first part of the input is the image to be aligned in the image pair; The second part of the input is to randomly generate a square block at the same position of the image to be aligned and the target image, and use this square block to intercept the image to be aligned and the target image respectively to obtain the corresponding square sub-images, and use the square sub-images as the input; The third part of the input is the coordinates of the four corner points of the square block; Step 4: Use an unsupervised learning network to construct a UAV image alignment model; The UAV image alignment model includes an encoder, a decoder, a differentiable linear transformation module, and a differentiable homography transformation module, where: The encoder takes the square sub-image as the input and outputs an output feature map containing spatial position information and image feature information; The decoder is used to restore the image feature information contained in the output feature map of the encoder, and at the same time obtain the perspective transformation field output by the decoder according to the spatial position information learned from the encoder; The differentiable linear transformation module determines the position transformation difference corresponding to the coordinates of the four corner points according to the perspective transformation field, and estimates the homography transformation matrix based on the position change difference; The differentiable homography transformation module is used to perform differentiable processing on the homography transformation matrix, invert the differentiable processed homography transformation matrix to obtain an inverse matrix; Use the inverse matrix to transform the image to be aligned in the first part of the input, and then intercept the position corresponding to the coordinates of the four corner points in the third part of the input on the transformed image to obtain a predicted image, and calculate the pixel loss through the predicted image and the corresponding square sub-image of the target image, so as to regress the entire unsupervised learning network; Step 5: Use the image dataset to train the UAV image alignment model, and save the trained UAV image alignment model for aligning the image pairs to be aligned.

2. The method for aligning drone images based on unsupervised learning according to claim 1, wherein, There are two methods for constructing the image pair: The first one is the image pair composed of the captured image of the UAV camera and the corresponding satellite image, where the captured image is used as the image to be aligned; First, use the attitude information of the UAV itself to calculate the rotation matrix of the UAV camera, and use the rotation matrix based on the principle of pinhole imaging to obtain the position of the pixels in the captured image of the UAV camera on the satellite image, and use this position to intercept the satellite image, and the intercepted image and the captured image form the image pair; The second one is on the video captured by the UAV camera, and the current frame video image and the video image at a preset number of frames interval form an image pair.

3. The method for aligning UAV images based on unsupervised learning according to claim 1, wherein The augmentation enhancements include vertical mirroring, horizontal mirroring, reducing brightness, or increasing brightness; the horizontal mirroring is to swap the left and right halves of the image to be aligned and the target image in the image pair with the vertical central axis of the image as the central axis, while the vertical mirroring is to swap the upper and lower halves of the image with the horizontal central axis of the image as the central axis; the increasing brightness is to add the weights of the image to be aligned and the target image in the image pair with an image whose pixel values are all 0. The reducing brightness is to multiply the pixel values of the RGB three channels of the image to be aligned and the target image by weights.

4. The method for aligning UAV images based on unsupervised learning according to claim 1, wherein The image I to be aligned A and the target image I B are respectively corresponding square sub-images P A and P B After gray-scale processing into single-channel images, they are spliced into a gray-scale image of H×W×2 as the input of the encoder; where H and W respectively represent the height and width of P A and P B ; The encoder is based on the ConvNeXt convolutional module and has a four-layer network structure; the first layer is a convolutional layer; the second to fourth layers are all composed of three cascaded ConvNeXt convolutional modules; in the ConvNeXt convolutional module, the input feature map is first processed by a depthwise separable convolution, and then passes through two convolutional layers in sequence. After the output of the activation function GELU is fused with the input feature map, the output feature map of the ConvNeXt convolutional module is obtained.

5. The method for aligning UAV images based on unsupervised learning according to claim 1, wherein The perspective transformation field is the square sub-graph P A and P B The position transformation differences of all pixel points on the abscissa and ordinate; respectively construct the corresponding two-dimensional feature matrices using the position transformation differences on the abscissa and ordinate and and then and are spliced to form the perspective transformation field.

6. The method for aligning UAV images based on unsupervised learning according to claim 1, wherein The decoder has five network layers, where the first layer is composed of 9 cascaded ConvNeXt convolutional modules, and the second to fourth layers are all composed of three cascaded ConvNeXt convolutional modules; different from the encoder, in all ConvNeXt convolutional modules of the decoder, the convolutional modules used for downsampling are replaced with transposed convolutional modules; the fifth layer of the decoder is two convolutional layers.

7. The method for aligning drone images based on unsupervised learning according to claim 1, wherein Two two-dimensional feature matrices included in the perspective transformation field PF output from the decoder and respectively select the position transformation differences on the abscissa and ordinate corresponding to the coordinates P of the four corner points in the third part of the input, denoted as (Δx 4point , Δy 4point , 4point ); The differentiable linear transformation module first restores the coordinates according to the position transformation differences (Δx 4point , Δy 4point ) and the coordinates P 4point of the four corner points in the third part of the input; {x′ i ,y′ i} = {x i -Δx i ,y i -Δy i} where x i , y i respectively represent the coordinates of the i-th corner point, x′ i , y′ i represent the coordinates of the i-th corner point after restoration, and Δx i , Δy i is the position transformation difference on the abscissa and ordinate corresponding to the i-th corner point; Subsequently, it is necessary to solve: where α is a scale factor, h k (k = 1, 2, ..., 9) are the elements in the homography transformation matrix; After eliminating the scale factor, the above formula is written in the form of matrix multiplication where: where the superscript T represents transpose; the DLT algorithm is used to solve the homography transformation matrix. The input is the coordinate point pairs {x i , x i '}, and the output is the homography transformation matrix that satisfies ​ For the coordinates P of the four corner points 4point For each of the corner points, a matrix A can be constructed i , and the matrices A of the four corner points i are concatenated to obtain a matrix A; perform SVD decomposition on matrix A to solve the minimum The eigenvector is 8. The method for aligning drone images based on unsupervised learning according to claim 1, wherein, The differentiable homography transformation module performs differentiable processing on the homography transformation matrix The process of performing differentiable processing is as follows: Normalize the height and width coordinates of the image I to be aligned A and the target image I B to a certain range; Subsequently, it is necessary to construct a coordinate grid G = {G B} of the same size as the target image I i , and each element G i in the coordinate grid G corresponds to a pixel coordinate in I B , that is Sampling is performed in the coordinate grid G = {G i}, and the transformed image obtained by sampling is denoted as an image V with C channels, width W', and height H': where H and W are the height and width of the image I to be aligned, A respectively, and the value at the position (n, m) in the horizontal and vertical coordinates of the c-th channel of the image I to be aligned is A , and V i c is the value of the output pixel at the i-th pixel position (u i , v i ) in the c-th channel. The V values of all channels i c constitute the image V. k(·) is a differentiable function, and Φ p , Φ q are the parameters of the differentiable function. By choosing to use the bilinear interpolation function max(·) as the differentiable function, it can be written as: Then there is the following partial differential calculation formula: That is, the derivative calculation can be performed through the chain rule, thus constructing a differentiable sampling mechanism.

9. The method for aligning UAV images based on unsupervised learning according to claim 1, wherein, Use smooth-l1 as the loss function of the UAV image alignment model, that is:

10. A terminal device, comprising a processor, a memory, and a computer program stored in the memory; characterized in that, When the processor executes the computer program, it implements the UAV image alignment method based on unsupervised learning according to any one of claims 1-9.