Self-supervised remote sensing image registration method based on global information robust feature description

Through a self-supervised method based on the robust feature description of global information, a synthetic data set is generated and a neural network is trained, which solves the problem of complex distortion and noise problems in remote sensing image registration, and achieves high-precision and robust image registration effect.

CN120219446APending Publication Date: 2025-06-27XIDIAN UNIV

Patent Information

Application Number
CN202510172308.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, it is difficult to achieve accurate registration of remote sensing images when dealing with problems such as complex geometric distortion, spatial resolution gap, time difference, light radiation difference, and occlusion and noise.

Method used

A self-supervised remote sensing image registration method based on robust feature description of global information is adopted. By generating a synthetic homography data set, a key point extraction network is trained, an encoder-decoder neural network with a hybrid architecture is built, and a loss function is constructed to extract robust descriptors and match key points.

Benefits of technology

The accuracy of remote sensing image matching of visible light band under extreme viewing angle changes is improved, and the accurate registration of remote sensing images is achieved, which can more effectively extract and match image features, enhancing the robustness of registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219446A_ABST
    Figure CN120219446A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised remote sensing image registration method based on global information robust feature description, and mainly solves the problems that an existing method is weak in recognition capability in the face of complex deformation and extreme appearance change and needs a large amount of manual annotation data. The implementation scheme comprises the following steps: generating a synthetic homography data set, and training a key point extraction network by using the synthetic homography data set; performing random transformation on the large-scale training data set to generate a plurality of versions of the same picture; obtaining key point information of the large-scale training data set after random transformation by using the trained key point extraction network, and combining key points of different versions of the same picture; building a neural network, training the neural network by using the combined key points, and exporting a matching point set; and calculating a homography matrix by using the matching point set to carry out image registration. According to the method, manual labeling is avoided, more robust matching key points and descriptors can be obtained, the accuracy of visible light band remote sensing image matching under the extreme view angle change is improved, and the method can be used for image data preprocessing of remote sensing tasks such as environment monitoring, disaster assessment, land classification and city expansion analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a self-supervised remote sensing image registration method, which can be used for environmental monitoring, disaster assessment, land use classification, urban expansion analysis, and change detection. Background Art

[0002] Remote sensing image registration is a fundamental and key technology, which occupies an important position in remote sensing image analysis. Its main task is to spatially align multi-source images from different times, sensors, or viewpoints, so that they have the same geographic reference system. This process is of great value for applications such as environmental monitoring, disaster assessment, land use classification, urban expansion analysis, and change detection. Due to the diversity of remote sensing image acquisition sources and the complexity of landscape features, such as problems like translation, rotation, occlusion, scale distortion, viewpoint change, and illumination diversity, remote sensing image registration still poses certain challenges.

[0003] Traditional image registration methods can be roughly divided into two categories: region-based methods and feature-based methods. Region-based methods usually use the original image intensity values to solve the image registration problem, such as the class correlation method, mutual information method, etc. These methods do not require significant structural features in the image, but are very sensitive to different illuminations and distortions, and perform poorly in image pairs with high noise and extreme illumination changes.

[0004] With the development of remote sensing technology, remote sensing images from different platforms, different sensors, and different times are becoming increasingly rich. In order to perform effective data analysis on these remote sensing images, it is very important to use reliable remote sensing image registration technology. Due to the large differences in the acquisition of different remote sensing images, different remote sensing images may be deformed due to reasons such as sensors, shooting angles, and ground object changes, as well as possible radiation differences, and there is generally a certain noise effect in remote sensing images. Traditional feature-based remote sensing images are difficult to handle these complex situations. In contrast, deep learning-based feature point detection methods have been widely used in image registration, such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and deep features.

[0005] Quan et al. proposed a new neural network Cnet for multimodal image registration tasks in "Deep feature correlation learning for multi-modal remotesensing image registration[J].IEEE Transactions on Geoscience and RemoteSensing, 2022, 60:1-16." It uses spatial and channel attention to enhance the representation of key features in the image, introduces a new feature correlation loss function, uses a scale factor to accelerate training and improve stability, uses the nearest neighbor search for key point matching in the matching stage, and uses NNDR and GMS methods to remove false matches to achieve high-precision image registration. However, the hyperparameter search of this method is relatively complex, the engineering implementation is relatively cumbersome, the registration speed is limited, and the hardware facilities are required to be high.

[0006] The patent document with application number CN202210799973.0 discloses "an end-to-end deep learning-based image stitching method", and its implementation scheme includes: using a deep homography transformation network to calculate the homography matrix between two images, and aligning the images through a spatial transformer layer to obtain an aligned image; using a codec network to downsample and reconstruct the aligned image, learn the deformation rules of image stitching, and output a stitched image; using an encoder to extract image features, and using a grid motion regressor and a residual progressive regressor to predict the horizontal and vertical motion of each vertex, and finally warping the stitched image into a rectangular image. Although this method can achieve a high level of image stitching, due to the use of a relatively basic convolutional network as a baseline model in the feature extraction process, the subsequent stitching process is prone to distortion and distortion, and robustness cannot be guaranteed; at the same time, since this method only uses MS-COCO training without migration to the target distribution, it is directly applied to the stitching of remote sensing images, which is prone to problems such as mismatching and loss of key information.

[0007] The patent document with the application number CN202010446227.4 discloses an "Image Registration Method Based on Convolutional Neural Network and Local Homography Transformation", and its implementation scheme includes: randomly selecting a target image from a large number of image datasets and randomly perturbing its corner points; randomly cropping image patches in the target image and performing uniform grid division; using the global homography matrix to transform the grid block corner points, automatically generating a large number of effective samples and their corresponding labels for training the local homography matrix estimation model required for image registration; constructing a local homography matrix estimation model for image registration based on a convolutional neural network, and then realizing image registration. This method eliminates the high-cost manual annotation by automatically generating labels, but due to calculating the homography matrix using local image matrices respectively, there are problems of large computational complexity and limited global perception ability of images; at the same time, since the loss function used mainly measures the distance between the extracted key points, it cannot well distinguish between matching and non-matching key points, and the training accuracy is limited.

[0008] Due to the fact that remote sensing images may have complex geometric distortions, such as atmospheric refraction, perspective shift, and local complex terrain, and are prone to non-linear deformations; different remote sensing images may have a huge gap in spatial resolution, as well as time differences, illumination radiation differences, etc.; and there may be phenomena such as occlusion and image noise. Therefore, the above method has accuracy limitations and often it is difficult to achieve precise image registration for remote sensing images. Summary of the Invention

[0009] The purpose of the present invention is to propose a self-supervised remote sensing image registration method based on global information robust feature description in view of the above deficiencies of the prior art, so as to improve the accuracy of visible light band remote sensing image matching under extreme perspective changes and achieve precise registration of remote sensing images.

[0010] To achieve the above purpose, the self-supervised remote sensing image registration method based on global information robust feature description of the present invention includes:

[0011] (1) Generating a synthetic homography dataset including various shapes, various texture features, and various homography transformations;

[0012] (2) Selecting a key point extraction network and training it with the synthetic homography dataset until the loss function converges;

[0013] (3) Obtaining a large-scale target detection dataset and dividing it into a training set and a validation set;

[0014] (4) Randomly transforming the training set and validation set pictures for data augmentation to generate multiple versions of the same picture;

[0015] (5) Use the trained key point extraction network to extract the key point information of the training set and validation set images and save them, and merge the key points of different versions of the same image to obtain comprehensive key point information as the label.

[0016] (6) Build a neural network including an encoder-decoder and construct its loss function, where:

[0017] The encoder adopts a hybrid architecture of a convolutional neural network and a Transformer network;

[0018] The decoder adopts a dual-head structure including a key point head and a descriptor head;

[0019] (7) Use the comprehensive key point information, that is, use the label to iteratively train the neural network until the loss function converges;

[0020] (8) Use the trained model to export the key points and descriptors of the two images to be matched respectively and match them;

[0021] Preferably, the encoder of the hybrid architecture alternately connects every two convolutional neural networks and Transformer networks, and inserts a max pooling module with a stride of 2 between every two layers of convolutional neural networks to obtain an encoder of the hybrid architecture containing 8 convolutional neural networks and 4 Transformer networks. The convolutional neural network includes a two-dimensional convolutional Conv2d module, a batch normalization module BN, and a ReLU activation function module connected in sequence; the Transformer network contains a pre-coding module, an additive attention module, and a multi-layer perception module connected in sequence.

[0022] Preferably, the dual-head structure of the key point head and the descriptor head constituting the decoder. The key point head includes a Conv2d module, a BN module, a ReLU module, a Conv2d module, and a BN module connected in sequence, and the output channel number is 65, which is used to output the heat map containing key point information of each batch of images; its descriptor head contains the same modules as the key point head, and the output channel number is adjusted to 256, which is used to output the descriptor information of each batch of images.

[0023] Preferably, the loss function of the neural network constructed by the present invention is expressed as:

[0024] Loss(X,Y,X',Y',D1,D2)=L det (X,Y)+L det (X',Y')+λL desc (D1,D2)

[0025] Among them, X and Y are the predicted and ground truth image pairs of the original image, and X' and Y' are the predicted and ground truth image pairs of the distorted image.

[0026] The present invention has the following advantages compared with the prior art:

[0027] First, the present invention can automatically detect the features of remote sensing images, avoiding the overly costly manual annotation.

[0028] Through the production of synthetic datasets and the training of the basic network, the present invention can obtain the key point pseudo-ground truth of large-scale datasets. By using basic shapes for image synthesis, automatically annotating the corner points of each shape, then performing random homography transformation on the same image to obtain multiple versions of the image, continuing to extract key points on these images, and then synthesizing the key points extracted from different versions of the same image, a composite key point is obtained. Through this method, the key point pseudo-ground truth data can be made as comprehensive and real as possible, closer to the potential key points that people notice when observing objects in real life, and can avoid the overly costly manual annotation.

[0029] Second, by integrating the advantages of the CNN and Transformer architectures, while achieving network lightweight, it takes into account both local features and global features.

[0030] The present invention adds an encoder of the SwiftFormer structure based on the VGG-like architecture, taking it as a variant of the Transformer. It uses addition instead of multiplication to calculate self-attention, uses linear element-wise matrix multiplication instead of expensive matrix multiplication, and replaces the explicit key-value pair interaction with a linear layer, having multiple advantages: one is to simplify the calculation for lightweight and improve the inference speed; the second is that SwiftFormer can have better performance among many variants of the Transformer and can achieve high accuracy in multiple tasks such as image classification, object detection, and semantic segmentation; the third is to utilize the characteristics of the SwiftFormer hierarchical encoder, making it easy to be integrated into the networks of other CNN architectures and quickly adapted to multiple downstream tasks.

[0031] Third, the loss function used can effectively calculate the loss between descriptors.

[0032] The loss function used in the present invention includes two parts: matching loss and non-matching loss. Both use the Euclidean distance between descriptors to calculate, which can simultaneously minimize the distance between matching pairs and maximize the distance between non-matching pairs; and generate matching and non-matching descriptor pairs through geometric transformation between images, which is not only convenient for the network to learn from unlabeled data, flexibly control the weights and distance boundaries between matching and non-matching pairs, but also has the advantages of high computational efficiency and rapid convergence.

[0033] The simulation results show that, compared with other existing remote sensing image registration methods, the present invention can be applied to the registration of complex remote sensing images and low-altitude UAV images, can make full use of the features of remote sensing images, extract more robust descriptors, and obtain more accurate image matching using homography transformation. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is the implementation flowchart of the present invention;

[0035] Figure 2 is the neural network structure diagram constructed in the present invention;

[0036] Figure 3 Simulation result diagram of matching two SUIRD near-ground UAV satellite images using the present invention and three existing image registration methods;

[0037] Figure 4 Simulation result diagram of matching two AID visible light remote sensing images using the present invention and three existing image registration methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] In recent years, feature extraction and matching methods based on neural networks have been widely used in remote sensing image matching. These methods can automatically extract the features of remote sensing images and have good adaptability to the unique noise and complex deformations in remote sensing images. Since the existing technologies mainly use feature detectors to identify key points and extract feature descriptors of the image patches around these key points, their recognition ability is weak in the face of complex deformations and extreme appearance changes. The features they recognize are at a lower level, and the key points detected are often not effective enough in the matching stage, which limits the robustness of the matching. Moreover, most of these methods require a large amount of manually labeled data for training, with a high cost. There is a higher time consumption and memory occupancy in the matching stage. Therefore, the present invention proposes a self-supervised remote sensing image registration method based on globally informed robust feature description.

[0039] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] See Figure 1 , the implementation steps of this example include the following:

[0041] Step 1, construct a synthetic homography dataset and divide it into a training set and a validation set.

[0042] 1.1) Select basic geometric shapes including triangles, quadrilaterals, line segments, ellipses, and cubes, and specify the key point positions of each shape, including the corner points of triangles and quadrilaterals, the midpoints of line segments, the center and vertices of ellipses, and the vertices of cubes;

[0043] 1.2) Draw the specified shape pictures and perform background rendering with texture addition and Gaussian blur, that is, fill the background of the pictures with colors of random depths, and then use the Gaussian blur function in cv2 and a Gaussian kernel of size 11 to perform Gaussian blur processing on the background;

[0044] 1.3) Perform homography transformations of translation, scaling, and rotation on the pictures processed in step 1.2) to generate multiple homography transformation versions of the same picture to simulate rich perspective changes: that is, first translate the picture by [-100, 100] pixels on the horizontal and vertical axes, then scale the picture by 0.9 to 1.2 times the original size, and then perform positive and negative angle rotation transformations of 0 to 60° on the picture to obtain a synthetic homography dataset;

[0045] 1.4) Divide the images in the synthetic homography dataset in 1.3) into a first training set and a first validation set according to a ratio of 7:3.

[0046] Step 2, select a key point extraction network that includes an encoder and a decoder.

[0047] The encoders used in existing key point extraction networks include VGGNet, AlexNet, ResNet, DenseNet, etc. In this example, VGGNet is selected but not limited to it, which includes 8 convolutional layers of 3x3, and a 2x2 max pooling layer is used for downsampling between every two layers. After each convolutional layer, a ReLU activation function layer and a BatchNorm layer are connected;

[0048] In this example, according to the research experience in the field of key point extraction, a decoder with a convolutional structure is selected but not limited to it, which includes a 3x3 and a 1x1 convolutional layer, and the number of output channels is 65.

[0049] Step 3, train the key point extraction network.

[0050] 3.1) Set the cross-entropy loss function as the loss function for the key point extraction network in the training and validation stages;

[0051] 3.2) Set the training parameters:

[0052] Set the image size to 240×320, the batches for training and validation are both 16, and the learning rate is 10 -4, the optimizer uses the Adam optimizer, the number of training rounds is 200,000 iterations, the validation stage is entered after every 1,000 iterations, and all the parameters of the model are saved every 2,000 iterations;

[0053] 3.3) Divide the synthesized homography data images into different batches according to the batch size;

[0054] 3.4) Input the first training set image samples of a batch into the keypoint extraction network to obtain the output results. Calculate the difference between the probability of each pixel in each image in the output results becoming a keypoint and the true label as the training loss, and calculate the gradient of the training loss with respect to each model parameter through backpropagation. Then use the Adam optimizer to update the network parameters according to the gradients calculated by backpropagation;

[0055] 3.5) Repeat step (3.4) until the number of training rounds is an integer multiple of 1,000 and enter the validation stage. At this time, freeze the network parameters, input the first validation set image samples of a batch into the keypoint extraction network to obtain the output results, and calculate the difference between the probability of each pixel in each image in the output results becoming a keypoint and the true label as the validation loss; when the number of training rounds reaches an integer multiple of 2,000, save all the parameters of the model;

[0056] 3.6) Repeat steps (3.4) and (3.5) until the number of iterations reaches the maximum number of training rounds and the training stops. Select the model with the lowest validation loss among all the saved models as the optimal model to obtain the trained keypoint extraction network.

[0057] Step 4, obtain a large-scale object detection dataset and divide it into a training set and a validation set.

[0058] Download the object detection dataset COCO2014 and the remote sensing target recognition dataset AID from public websites;

[0059] Divide the object detection dataset COCO2014 into a second training set and a second validation set at a ratio of 2:1;

[0060] Divide the remote sensing target recognition dataset AID into a third training set and a third validation set at a ratio of 7:3.

[0061] Step 5, use the trained keypoint extraction network to extract the dataset pictures in step 4 to obtain comprehensive keypoint information as the label.

[0062] 5.1) Perform random transformations on the training set and validation set pictures divided in step 4:

[0063] 5.1.1) Crop the pictures in the second training set, the second validation set, the third training set, and the third validation set to a size of 240×320;

[0064] 5.1.2) Perform photometric changes on the images in (5.1.1), including brightness adjustment, contrast adjustment, and saturation adjustment. Among them, for brightness adjustment, multiply the pixel values of the image by a random value between 1 and 5; for contrast adjustment, adjust the contrast of the image to 0.5 to 1.5 times that of the original image; for saturation adjustment, adjust the saturation of the image to 1 to 5 times that of the original image;

[0065] 5.1.3) Perform random homography transformation on the images in step (5.1.2), including size adjustment, rotation, and translation: Size adjustment includes scaling the image to 0.9 to 1.2 times its original size, rotation includes rotating the image by positive and negative angles from 0 to 60°, and translation includes translating the image along the horizontal and vertical axes within the range of [-100, 100] pixels;

[0066] 5.2) For the multiple transformed versions of the images generated in step (5.1), use the key point extraction network to perform key point extraction respectively to obtain the key point information of each image;

[0067] 5.3) Superimpose the key point information of different versions of the same image after extraction, and use NMS (Non-Maximum Suppression) to remove duplicate key points within a radius of 4 to obtain more comprehensive key point information of the same image.

[0068] Step 6: Build a neural network including an encoder-decoder and construct its loss function.

[0069] Refer to Figure 2 , the implementation of this step includes the following:

[0070] 6.1) The main structure of the neural network is constructed to include an encoder using a hybrid architecture of a convolutional neural network and a Transformer network:

[0071] 6.1.1) Establish a convolutional neural network composed of a two-dimensional convolutional Conv2d module, a batch normalization module BN, and a ReLU activation function module connected in sequence;

[0072] 6.1.2) Establish a Transformer network composed of a pre-coding module, an additive attention module, and a multi-layer perception module connected in sequence;

[0073] 6.1.3) Alternately connect the convolutional neural network and the Transformer network, and insert a max pooling module with a stride of 2 between every two layers of the convolutional neural network to obtain an encoder with a hybrid architecture containing 8 convolutional neural networks and 4 Transformer networks.

[0074] 6.2) Construct a dual - head structure decoder including a key - point head and a descriptor head;

[0075] 6.2.1) Establish a key - point head composed of sequentially connected Conv2d modules, BN modules, ReLU modules, Conv2d modules, and BN modules. The output channel number of the last Conv2d module is set to 65, which is used to output the heatmap containing key - point information for each batch of images;

[0076] 6.2.2) Establish a descriptor head composed of sequentially connected Conv2d modules, BN modules, ReLU modules, Conv2d modules, and BN modules. The output channel number of the last Conv2d module is adjusted to 256, which is used to output the descriptor information for each batch of images;

[0077] 6.2.3) Parallel - connect the key - point head and the descriptor head to form a dual - head structure decoder;

[0078] 6.3) Cascade the above - constructed encoder and decoder to obtain a neural network including an encoder - decoder;

[0079] 6.4) Construct the loss function of the neural network:

[0080] 6.4.1) Select an existing key - point loss function, that is, the cross - entropy loss function L det (X,Y):

[0081]

[0082] where N is the number of image pixels, X and Y are the predicted key - point heatmap and the pseudo - ground - truth key - point label respectively,

[0083] is the cross - entropy output of a single image pixel in the two images, x i ,y i are the pixel values of each point respectively;

[0084] 6.4.2) Define the descriptor distance function as the weighted sum of the matching loss L match and the non - matching loss L nonmatch The expression is:

[0085] L desc (D1,D2) = λ d *L match +L nonmatch

[0086] where D1 and D2 represent the descriptors extracted from the original image and the distorted image respectively,

[0087] λ dTo balance the weights of the two loss functions,

[0088] N match To match the number of descriptors, P1 represents the set of all matching descriptors in the two images, and d i and d j are the descriptors extracted by the neural network from the original image and the distorted image respectively;

[0089] N nonmatch To be the number of non-matching descriptors, P2 represents the set of all non-matching descriptors in the two images, and m is the boundary value to control the maximum distance between non-matching descriptors;

[0090] 6.4.3) According to the key point loss function L det (X, Y) and the descriptor distance function L desc (D1, D2), the total loss function of the neural network is defined as:

[0091] Loss(X, Y, X', Y', D1, D2) = L det (X, Y) + L det (X', Y') + λL desc (D1, D2)

[0092] where X, Y are the predicted and ground truth image pairs of the original image, X', Y' are the predicted and ground truth image pairs of the distorted image, and D1, D2 represent the descriptor information extracted by the neural network from the original image and the distorted image respectively.

[0093] Step 7: Use the comprehensive key point information, that is, the labels, to iteratively train the neural network to obtain a trained neural network.

[0094] 7.1) Set the training parameters:

[0095] Set the learning rate to 10 -4 , the batch size to 8, the maximum number of iterations to 170000, and use the Adam algorithm for optimization during the optimization process. Enter the validation stage every 200 iterations and save all the parameters of the model at the same time;

[0096] 7.2) Divide all the COCO2014 dataset images in the second training set and the second validation set into different batches according to the batch size;

[0097] 7.3) Input the training set image samples of one batch into the neural network to obtain the output results. Calculate the difference between the network output and the pseudo ground truth labels as the training loss value through the loss function Loss(X, Y, X', Y', D1, D2), and calculate the gradient of the training loss value with respect to each model parameter through backpropagation. Then, use the Adam optimizer to update the network parameters according to the gradients calculated by backpropagation;

[0098] 7.4) Repeat step (7.3) until the number of training epochs is an integer multiple of 200, then enter the validation stage. At this time, freeze the network parameters, input the validation set image samples of one batch into the key point extraction network to obtain the output results, calculate the difference between the network output and the pseudo ground truth labels as the validation loss through the loss function Loss(X, Y, X', Y', D1, D2), and save all the parameters of the model;

[0099] 7.5) Repeat steps (7.3) to (7.4) until the number of iterations reaches the maximum number of training epochs and the training stops. Select the model with the lowest validation loss among all the saved models as the optimal model to obtain the neural network trained in the first stage;

[0100] 7.6) Divide the third training set and the third validation set obtained from the AID dataset images into different batches of size 8;

[0101] 7.7) Set the learning rate to 10 -5 , the maximum number of iterations to 100000, use the Adam algorithm for optimization, enter the validation stage every 200 iterations, and save all the parameters of the model at the same time;

[0102] 7.8) Load the parameters of the neural network trained in the first stage in step (7.5) as the starting point for the model training in this stage. Input the training set image samples of one batch into the neural network to obtain the output results. Calculate the difference between the network output and the pseudo ground truth labels as the training loss through the loss function Loss(X, Y, X', Y', D1, D2), and calculate the gradient of the training loss with respect to each model parameter through backpropagation. Then, use the Adam optimizer to update the network parameters according to the gradients calculated by backpropagation;

[0103] 7.9) Repeat step (7.8) until the number of training epochs is an integer multiple of 200, then enter the validation stage. At this time, freeze the network parameters, input the validation set image samples of one batch into the key point extraction network to obtain the output results, calculate the difference between the network output and the pseudo ground truth labels as the validation loss using the loss function Loss(X, Y, X', Y', D1, D2), and save all the parameters of the model;

[0104] 7.10) Repeat steps (7.8) to (7.9) until the number of iterations reaches the maximum number of training epochs, then stop training. Select the model with the lowest validation loss among all the saved models as the optimal model to obtain the finally trained neural network.

[0105] Step 8: Use the trained model to export the key points and descriptors of the two images to be matched respectively, and match the two images to be matched.

[0106] 8.1) Input the two images to be matched into the trained neural network to obtain the heat map of key point information and descriptor information;

[0107] 8.2) Set a sampling threshold, and sample and filter the heat map of key point information, that is, regard the pixel points greater than the sampling threshold of 0.015 as key points, and use NMS (Non-Maximum Suppression) to remove the repeated key points within a radius of 4;

[0108] 8.3) Perform bilinear interpolation sampling on the descriptor information according to the key points obtained in step (8.2) to obtain the descriptor vector of each key point;

[0109] 8.4) According to the obtained key points and descriptors, use the BFMatcher matcher to match the key points:

[0110] For each pair of images, use the BFMatcher matcher to calculate the distance between each pair of descriptor vectors in the two images respectively;

[0111] For each key point, select the key point corresponding to the descriptor with the smallest distance as the matching point to obtain the matching result of the key points of the two images;

[0112] 8.5) According to the key point matching result, use the following coordinate transformation formula to calculate the homography matrix H:

[0113]

[0114] Where h1 to h8 are the 8 degrees of freedom parameters of the projective transformation, x', y' are the coordinates of the key points of the original image, and x, y are the coordinates of the key points of the distorted image;

[0115] 8.6) Use the RANSAC (Random Sample Consensus) algorithm to iteratively remove the wrong matching points and calculate the homography matrix to generate the final homography matrix H';

[0116] 8.6.1) Set the reprojection error threshold. In this example, the threshold is set to 5 (not limited to this value), and the maximum number of iterations is 2000;

[0117] 8.6.2) Randomly select 4 pairs of points from the matching point pairs obtained in step (8.4) as the minimum sample set, and use this minimum sample set to calculate the equations by the DLT direct linear transformation method to generate the homography matrix H for this iteration;

[0118] 8.6.3) Calculate the projection error of all matching points using the homography matrix H for this iteration, and compare the error result with the set threshold:

[0119] If the error is less than the threshold 5.0, mark this group of matching points as inliers;

[0120] Otherwise, mark this group of matching points as outliers;

[0121] 8.6.4) Determine whether the number of matching points for this time is greater than the maximum value of the historical inlier number:

[0122] If it is greater than the historical maximum value, update the optimal parameters and the inlier set;

[0123] Otherwise, do not update the parameters and inliers;

[0124] 8.6.5) Repeat steps (8.6.2) to (8.6.4) until the maximum number of iterations 2000 is reached to obtain the final homography matrix H';

[0125] 8.7) Use the final homography matrix H' to perform coordinate transformation on the distorted image to obtain a registered image that is similar to the original image and has accurate ground object information.

[0126] The effects of the present invention will be further described below in conjunction with simulation experiments.

[0127] I. Simulation conditions:

[0128] The simulation experiment of the present invention is carried out in a hardware environment with a CPU of AMD Ryzen 9 5900X 12-Core Processor, a GPU of Nvidia RTX3090, and a memory of 31GB and a software environment of Pycharm;

[0129] Simulation data: Select the SUIRD low-altitude UAV dataset and the AID remote sensing image target detection dataset.

[0130] II. Simulation content and results:

[0131] Simulation 1, for the SUIRD low-altitude UAV dataset, use the method of the present invention and the existing Superpoint method, SIFT method, and ORB method to perform key point detection and matching simulation on two images of the same scene and different shooting angles. The results are as Figure 3 , among which:

[0132] Figure 3 (a) is a pair of real target images;

[0133] Figure 3 (b) is the matching result of the method of the present invention;

[0134] Figure 3 (c) is the matching result of the existing Superpoint method;

[0135] Figure 3 (d) is the matching result of the existing SIFT method;

[0136] Figure 3 (e) is the matching result of the existing ORB method.

[0137] In each result diagram, green ① represents the correct matching result, i.e., inliers, after iteration by the RANSAC algorithm, and red ② represents the incorrect matching, i.e., outliers.

[0138] Comparison Figure 3 (b) and Figure 3 It can be seen from comparing (c) that the method of the present invention detects the key point information in the picture more comprehensively, and at the same time has relatively few false matching cases, and the effectiveness of the key points is relatively high.

[0139] Comparison Figure 3 (b) and Figure 3 It can be seen from comparing (d) that when the method of the present invention detects pictures with large perspective changes, the matching area is relatively comprehensive, and there is a good matching for the information of the field part, while the main matching key points of the SIFT method are concentrated in the residential area and water area, indicating that the method of the present invention has a good extraction accuracy for the global information of the picture.

[0140] Comparison Figure 3 It can be seen from comparing (b), (c), (d) and (e) that the descriptors extracted by the SIFT method and the ORB method are concentrated in the areas with large changes in image information, and the matching pairs extracted by the method of the present invention are more comprehensively distributed on the image, and the number of incorrect matching pairs is less, and it has a good matching accuracy.

[0141] Simulation 2: For the pictures in the AID remote sensing image target detection dataset and their deformed pictures, the method of the present invention and the existing Superpoint method, SIFT method, and ORB method are respectively used to conduct a simulation experiment on key point detection and matching of two images before and after deformation in the dataset. The results are as Figure 4 , where:

[0142] Figure 4 (a) is a pair of real target images;

[0143] Figure 4 (b) is the matching result of the method of the present invention;

[0144] Figure 4 (c) is the matching result of the Superpoint method;

[0145] Figure 4 (d) is the matching result of the SIFT method;

[0146] Figure 4 (e) is the matching result of the ORB method.

[0147] In each result graph, the green ① represents the correct matching results (i.e., inliers) after RANSAC algorithm iteration, and the red ② represents the incorrect matches (i.e., outliers).

[0148] Comparison Figure 4 From the result images of (b) and other existing methods, it can be seen that the present invention has good extraction ability for various details in optical remote sensing images. The Superpoint method has more outliers after RANSAC, but some of these outliers are actually correct matching points, which may be related to the threshold setting during RANSAC iteration. Some points are considered outliers during the iteration process, while the matching points extracted by the present invention can effectively distinguish correct and incorrect matching points after RANSAC iteration and maintain a high correct matching ratio even at high iteration times. Although both the SIFT and ORB methods can extract dense and effective matches, they will also generate more incorrect matching points.

[0149] Simulation 3: In order to objectively and quantitatively analyze the results of the present invention, the present invention method and the existing Superpoint method, SIFT method, ASLFeat method, and improved ResNet method are respectively used to calculate five matching metrics, namely, computational repeatability, localization error, accuracy of the homography matrix at a 3-pixel threshold, mean average precision of multiple data points, and maximum matching point ratio score, using the HPatches dataset. The results are shown in Table 1.

[0150] Table 1 Registration metrics of the present invention and four existing remote sensing image matching algorithms for the HPatches dataset

[0151]

[0152] As can be seen from Table 1, compared with other remote sensing image registration methods, the positioning error Local_err of the present invention is the smallest, the average accuracy MAP is the highest, and it has the highest matching point score Matching score. The repeatability Repeatability and performance are also ranked in the middle. It can fully extract effective features, effectively locate the features of one picture into another picture, and achieve good results in effectively and objectively extracting key points and descriptor features.

[0153] The above simulation results show that: compared with the existing methods, the present invention can obtain more robust matching key points and descriptors, and both the matching index and the visual effect have shown good performance.

[0154] It should be noted that the step numbers in the specification and claims of the present invention are only for a clear description of the embodiments of the present invention for easy understanding, and the order of their serial numbers is not limited.

Claims

1. A self-supervised remote sensing image registration method based on global information robust feature description, characterized in that: include: (1) Generate synthetic homography datasets including multiple shapes, multiple texture features, and multiple homography transformations; (2) Select a key point extraction network and train it using a synthetic homography dataset until the loss function converges; (3) Obtain a large-scale target detection dataset and divide it into a training set and a validation set; (4) Randomly transform the training set and validation set images for data augmentation to generate multiple versions of the same image; (5) Use the trained key point extraction network to extract and save the key point information of the training set and validation set images, and merge the key points of different versions of the same image to obtain comprehensive key point information as labels; (6) Build a neural network including an encoder-decoder and construct its loss function, where: The encoder adopts a hybrid architecture of a convolutional neural network and a Transformer network; The decoder adopts a dual-head structure including a key point head and a descriptor head; (7) Using comprehensive key point information, i.e., labels, the neural network is iteratively trained until the loss function converges; (8) Use the trained model to derive the key points and descriptors of the two images to be matched, and then match them.

2. The method according to claim 1, characterized in that Step (1) generates a synthetic homography dataset including multiple shapes, multiple texture features, and multiple homography transformations, and its implementation includes the following: (1a) Select basic geometric shapes including triangle, quadrilateral, line segment, ellipse, and cube, and specify the key point positions of each shape; (1b) Draw the specified shape image, add texture, and render the background with Gaussian blur; (1c) Perform homography transformations such as translation, scaling, rotation, and perspective distortion on the specified image processed by (1b) to generate multiple homography transformed versions of the same image to simulate rich perspective changes.

3. The method according to claim 1, characterized in that The key point extraction network selected in step (2) includes an encoder and a decoder: The encoder includes 8 3x3 convolutional layers, with a 2x2 maximum pooling layer used for downsampling between each two layers, and each convolutional layer is followed by a ReLU activation function layer and a BatchNorm layer; The decoder includes a 3x3 and a 1x1 convolutional layer, and the number of output channels is 65.

4. The method according to claim 1, characterized in that: Step (2) uses the synthetic homography dataset to train the key point extraction network, and its implementation includes the following: (2a) Setting the cross entropy loss function as the loss function of the key point extraction network in the training stage and the verification stage; (2b) Set training parameters: The image size is set to 240×320, the batch size is 16 in both the training and validation phases, and the learning rate is 10 -4 The optimizer uses the Adam optimizer, the number of training rounds is 200,000 iterations, the validation phase begins after every 1,000 iterations, and all model parameters are saved every 2,000 iterations; (2c) dividing all synthetic homography data images of the training set and the validation set into different batches according to the batch size; (2d) In the training phase, a batch of training set image samples is input into the key point extraction network to obtain the output results. The difference between the probability of each pixel in each image in the output results becoming a key point and the true label is calculated as the training loss. The gradient of the training loss to each model parameter is calculated through back propagation, and the Adam optimizer is used to update the network parameters according to the gradient calculated by back propagation. (2e) Repeat step (2d) until the number of training rounds is an integer multiple of 1000, then enter the verification phase. At this time, freeze the network parameters, input a batch of verification set image samples into the key point extraction network to obtain the output results, and calculate the difference between the probability of each pixel in each image in the output result becoming a key point and the true label as the verification loss; When the number of training rounds is an integer multiple of 2000, all parameters of the model are saved; (2f) Repeat steps (2d) and (2e) until the number of iterations reaches the maximum number of training rounds. The training stops and the model with the lowest verification loss is selected as the optimal model among all saved models to obtain the trained key point extraction network.

5. The method according to claim 1, characterized in that In step (4), the training set and validation set images are randomly transformed to perform data enhancement and generate multiple versions of the same image. The implementation includes the following: (4a) Resize the image to (240, 320); (4b) Performing photometric changes on the image, including brightness adjustment, contrast adjustment, and saturation adjustment; (4c) Perform random homography transformations on the image, including resizing, image cropping, rotation, and translation.

6. The method according to claim 1, characterized in that In step (5), the key points of different versions of the same image are merged, and the implementation includes the following: For the generated multiple transformed versions of the images, the key points are extracted respectively using the key point extraction network; After the extraction is completed, the key point information of different versions of the same image is superimposed, and the repeated key points at the same coordinates are ignored to obtain more comprehensive key point information of the image as the pseudo ground truth label.

7. The method according to claim 1, characterized in that In step (6), an encoder with a hybrid architecture of a convolutional neural network and a Transformer network is used, and its implementation includes the following: The convolutional neural network includes a two-dimensional convolution Conv2d module, a batch normalization module BN, and a ReLU activation function module connected in sequence; The Transformer network includes a precoding module, an additive attention module and a multi-layer perception module connected in sequence; Every two convolutional neural networks and Transformer networks are alternately connected, and a maximum pooling module with a step size of 2 is inserted between every two layers of convolutional neural networks to obtain an encoder with a hybrid architecture consisting of 8 convolutional neural networks and 4 Transformer networks.

8. The method according to claim 1, characterized in that: The double-head structure of the key point header and the descriptor header constituting the decoder in step (6) includes the following: The key point head includes a Conv2d module, a BN module, a ReLU module, a Conv2d module, and a BN module connected in sequence, and the number of output channels is 65, which is used to output a heat map containing key point information of each batch of images; The descriptor head includes the same modules as the key point head, and the number of output channels is adjusted to 256, which is used to output the descriptor information of each batch of images.

9. The method according to claim 1, characterized in that: In step (6), the loss function of the neural network is constructed, and its implementation includes the following: (6a) Select the existing key point loss function, namely the cross entropy loss function L det (X,Y): Where N is the number of image pixels, X and Y are the predicted key point heat map and pseudo ground truth key point labels, respectively. is the cross entropy output of a single image pixel in the two images, x i ,y i are the pixel values ​​of each point respectively; (6b) Define the descriptor distance function as the matching loss L match and mismatch loss L nonmatch The weighted sum of is expressed as: L desc (D1,D2)=λ d *L match +L nonmatch Where D1 and D2 represent the descriptors extracted from the original image and the distorted image respectively, and λ d To balance the weights of the two loss functions, d i ,d j are the descriptors extracted by the neural network from the original image and the distorted image respectively; m is the boundary value, which controls the maximum distance between non-matching descriptors; (6c) According to the key point loss function L det (X,Y) and the descriptor distance function L desc (D1,D2), the total loss function is defined as: Loss(X,Y,X',Y',D1,D2)=L det (X,Y)+L det (X',Y')+λL desc (D1,D2) Where X, Y are the predicted and true image pairs of the original image, and X', Y' are the predicted and true image pairs of the distorted image.

10. The method according to claim 1, characterized in that In step (7), comprehensive key point information, i.e., labels, is used to iteratively train the neural network, and its implementation includes the following: (7a) Set training parameters: Set the learning rate to 10 -4 , the batch size is 8, the maximum number of iterations is 170000, the optimization process uses the Adam algorithm for optimization, every 200 iterations enter the verification phase, and all the parameters of the model are saved at the same time; (7b) Divide all COCO2014 dataset images in the training set and validation set into different batches according to the batch size; (7c) In the training phase, a batch of training set image samples is input into the neural network to obtain the output results. The difference between the network output and the pseudo ground truth label is calculated as the training loss through the loss function, and the gradient of the training loss to each model parameter is calculated through back propagation. Then, the Adam optimizer is used to update the network parameters according to the gradient calculated by back propagation; (7d) Repeat step (7c) until the number of training rounds is an integer multiple of 200 and enter the validation phase. At this time, freeze the network parameters, input a batch of validation set image samples into the key point extraction network to obtain the output results, calculate the difference between the network output and the pseudo ground truth label as the validation loss through the loss function, and save all the parameters of the model; (7e) Repeat steps (7c) to (7d) until the number of iterations reaches the maximum number of training rounds and the training stops. The model with the lowest verification loss among all the saved models is selected as the optimal model to obtain the neural network trained in the first stage; (7f) Divide the AID dataset images into different batches according to the batch sizes of the training set and the validation set; (7g) Set the learning rate to 10 -5 , trained on the AID dataset and labels for 100,000 iterations, optimized using the Adam algorithm, entered the validation phase every 200 iterations, and saved all the model parameters at the same time; (7h) In the training phase, the parameters of the neural network trained in the first phase in (7e) are first loaded as the starting point of the model training in this phase, and a batch of training set image samples are input into the neural network to obtain the output results. The difference between the network output and the pseudo ground truth label is calculated by the loss function as the training loss, and the gradient of the training loss to each model parameter is calculated by back propagation. Then, the Adam optimizer is used to update the network parameters according to the gradient calculated by back propagation; (7i) Repeat step (7h) until the number of training rounds is an integer multiple of 200 and enter the validation phase. At this time, freeze the network parameters, input a batch of validation set image samples into the key point extraction network to obtain the output results, use the loss function to calculate the difference between the network output and the pseudo ground truth label as the validation loss, and save all the parameters of the model; (7j) Repeat steps (7h) to (7i) until the number of iterations reaches the maximum number of training rounds. The training stops and the model with the lowest verification loss is selected as the optimal model among all the saved models to obtain the final trained neural network.

11. The method according to claim 1, characterized in that: In step (8), the trained model is used to derive the key points and descriptors of the two images to be matched, and the two images are matched. The implementation includes the following: (8a) Input the two images to be matched into the trained neural network to obtain the key point information heat map and descriptor information; (8b) Setting a sampling threshold and performing sampling screening on the key point information heat map, that is, considering the pixels greater than the sampling threshold as key points; (8c) performing bilinear interpolation sampling on the key points obtained according to (8b) to obtain a descriptor vector for each key point; (8d) According to the obtained key points and descriptors, the BFMatcher matcher is used to match the key points: For each pair of images, the BFMatcher matcher is used to calculate the distance between each pair of descriptor vectors in the two images; For each key point, select the key point corresponding to the descriptor with the smallest distance as the matching point to obtain the matching results of the key points of the two images; (8e) According to the key point matching results, the homography matrix H is calculated using the following coordinate transformation formula: Among them, h1 to h8 are the 8 degrees of freedom parameters of the projection transformation; (8f) The calculated homography matrix H is used to transform the coordinates of the distorted image to obtain a registered image that is similar to the original image and has accurate ground object information.

Citation Information

Patent Citations

  • Image registration method based on convolutional neural network and local homography transformation

    CN111833237A

  • End-to-end image splicing method based on deep learning

    CN117437120A

Cited By

  • Unmanned aerial vehicle model identification method based on image processing

    CN121033398A