Double-branch multi-mode remote sensing image registration method based on multi-scale and multi-direction features

By constructing a dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features, the problem of accurate alignment in multi-modal remote sensing image registration is solved, and the accurate alignment and unified coordinate system of images are realized, and data utilization efficiency and analysis accuracy are improved.

CN120355758APending Publication Date: 2025-07-22XIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510392139.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing multimodal remote sensing image registration method is difficult to achieve accurate alignment when facing factors such as differences in imaging principle, noise, resolution differences and geometric distortion of different modal images, resulting in low data utilization efficiency and insufficient analysis accuracy.

Method used

The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features is adopted. By constructing a dual-branch feature extraction network, affine parameter regression network and spatial transformation network, combining multi-scale and multi-directional features, affine transformation parameters related to the registration task are learned to achieve accurate image alignment.

Benefits of technology

It effectively improves feature expression ability and adaptability to different transformations, can accurately obtain land object information, enhances the analysis and processing capabilities of image fusion and change detection, and improves the accuracy of surface change monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355758A_ABST
    Figure CN120355758A_ABST
Patent Text Reader

Abstract

The invention discloses a dual-branch multi-mode remote sensing image registration method based on multi-scale and multi-direction features, and the method comprises the steps: dividing an original image data set into a training set and a test set, and carrying out the preprocessing; then constructing a dual-branch multi-mode remote sensing image registration network model based on multi-scale and multi-direction features; the dual-branch multi-mode remote sensing image registration network model based on the multi-scale and multi-direction features is trained; and finally, registering test data by using the trained dual-branch multi-mode remote sensing image registration network model based on the multi-scale and multi-direction characteristics. According to the invention, images from different sensors, different time phases or different resolutions are accurately aligned in spatial positions. Through registration, geometric differences among the images can be eliminated, so that the images have a unified coordinate system, subsequent analysis processing such as image fusion, change detection and information extraction is facilitated, and ground feature information is more accurately acquired and ground surface change is more accurately monitored.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image registration in image processing, and particularly relates to a dual-branch multimodal remote sensing image registration method based on multi-scale and multi-direction features. Background Art

[0002] In the current rapid development of remote sensing technology, multimodal remote sensing image registration is extremely crucial but full of challenges. Various remote sensing sensors are constantly innovating, producing multimodal images covering optics, SAR, infrared, etc. Due to their different imaging principles and large differences in information expression characteristics, it brings both opportunities and problems to the registration technology. Multimodal remote sensing image registration aims to achieve precise alignment of the spatial dimensions of different modal images, so that the same ground object corresponds to the positions in each image, thereby improving the data utilization efficiency and analysis accuracy. However, the imaging principles and physical characteristics of different modal images result in very different expressions of ground object features and difficult-to-find matching features, and factors such as image noise, resolution differences, and geometric distortions increase the registration difficulty.

[0003] Traditional registration methods are mainly divided into two categories: region-based and feature-based. Region-based methods determine registration parameters by comparing regional features and maximizing similarity measures, while feature-based methods establish image correspondence through feature extraction, matching, parameter estimation, and spatial transformation. When these methods are used for multimodal registration, due to the non-linear differences, complex deformations, and radiation changes in the images, they have limitations in aspects such as regional similarity measurement and feature point extraction, and are also easily affected by noise interference.

[0004] In recent years, methods based on deep learning have emerged. They can automatically learn multimodal image features and effectively handle feature differences and noise problems. However, the commonly used siamese convolutional neural network is not end-to-end registration, and the efficiency is relatively low. Although there are end-to-end neural network methods that can better handle image transformations and have strong generality, they require a large number of image pairs for learning, and there are difficulties in feature extraction and parameter determination in scenarios such as complex textures and blurred boundaries. Summary of the Invention

[0005] The purpose of the present invention is to provide a dual-branch multimodal remote sensing image registration method based on multi-scale and multi-direction features, which precisely aligns images from different sensors, different time phases, or different resolutions in terms of spatial position. Through registration, the geometric differences between images can be eliminated, enabling them to have a unified coordinate system, facilitating subsequent analysis and processing such as image fusion, change detection, and information extraction, so as to more accurately obtain ground object information and monitor surface changes.

[0006] The technical solution adopted by the present invention is a dual-branch multimodal remote sensing image registration method based on multi-scale and multi-direction features, which is specifically implemented according to the following steps:

[0007] Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing;

[0008] Step 2: Construct a dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0009] Step 3: Train the dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0010] Step 4: Use the trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

[0011] The features of the present invention also lie in that,

[0012] Step 1 is specifically implemented according to the following steps:

[0013] Step 1.1: The elements in the panchromatic and multispectral datasets are the cropped images of three remote sensing satellites, namely GF-1, WorldView-2, and WorldView-4. In the panchromatic and multispectral datasets, the resolution of the multispectral image is 256×256 pixels, while the resolution of the panchromatic image is 1024×1024 pixels. The multispectral image is regarded as the moving image, and the panchromatic image is regarded as the fixed image. Conversely, the multispectral image is regarded as the fixed image, and the panchromatic image is regarded as the moving image;

[0014] Step 1.2: For the panchromatic and multispectral datasets, first downsample the panchromatic image to a size of 256×256 pixels, and randomly select 1104 pairs of images for training and 148 pairs of images for testing.

[0015] Combined with Figures 2 to 6 , in Step 2:

[0016] The structure of the dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features consists of three parts: a dual-branch feature extraction network, an affine parameter regression network, and a spatial transformation network. The panchromatic and multispectral remote sensing images obtained in Step 1 are respectively input into the same dual-branch feature extraction network to extract features containing global and local information; after the deep features obtained from the dual-branch feature extraction network for the panchromatic and multispectral modalities are concatenated, they are input into the affine parameter regression network. Through the affine parameter regression network, six affine transformation parameters related to the registration task, including rotation, translation, and scaling, are learned; then the six affine transformation parameters are input into the spatial transformation network and applied to the moving image to achieve the registration of the distorted image.

[0017] The processing process of the dual-branch feature extraction network in Step 2 is as follows:

[0018] For remote sensing images in both panchromatic and multispectral modalities, the upper branch of the dual-branch feature extraction network performs multi-scale feature extraction. The input image first undergoes preliminary feature extraction through a basic 3×3 convolution, and then is divided into 4 branches according to different scales to process the data separately. After extracting features, the information at different levels is fused. Then, the number of channels is adjusted through a 1×1 convolution, and non-linearity is added through the GELU activation function. Finally, the output feature map is further refined through two Swin modules. The lower branch focuses on multi-directional feature extraction. The input image first undergoes preliminary feature extraction through two basic 3×3 convolutions, and then captures feature information from four directions: horizontal, vertical, left diagonal, and right diagonal to achieve a more comprehensive feature representation. Then, the features are concatenated and normalized through a batch normalization layer, and the non-linearity is enhanced by a multi-layer perceptron. Finally, the features are weighted from the channel and spatial dimensions by a dual attention module, and finally, a feature map is output through a 3×3 convolution.

[0019] After completing the dual-branch feature extraction and obtaining the deep features of the two modalities, these features are concatenated and input into the affine parameter regression network.

[0020] The processing procedure of the affine parameter regression network in step 2 is as follows:

[0021] The fixed image and the moving image obtain deep features through the dual-branch feature extraction network. They are concatenated and input into the parameter regression network to learn 6 affine transformation parameters including rotation, translation, and scaling related to the registration task. The deep features are defined as:

[0022] Z F ,Z M =DBFE(I F ),DBFE(I M )

[0023] where I F ,I M ,Z F ,Z M represent the fixed image, the moving image, the deep features of the fixed image, and the deep features of the moving image respectively. DBFE represents the dual-branch feature extraction network;

[0024] The obtained 6 parameters are defined as:

[0025] [θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 =FC(Resnet(Cat(Z F ,Z M )))

[0026] where θ 11θ 12 θ 13 θ 21 θ 22 θ 23 respectively represent six affine transformation parameters, FC represents the fully connected layer, Resnet represents the parameter regression network, Cat(Z F ,Z M ) represents the concatenated deep features.

[0027] The processing process of the spatial transformation network in step 2 is as follows:

[0028] Input the obtained affine transformation parameters into the spatial transformation network. In this step, the spatial transformation network will reshape the six affine transformation parameters into the parameter matrix φ according to the received six affine transformation parameters. This reshaping process is expressed as:

[0029] φ = reshape([θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 )

[0030] where φ is the parameter matrix, and θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 respectively represent six affine transformation parameters, and reshape is the reshaping operation;

[0031] Apply the obtained matrix to the moving image through spatial transformation to obtain the final registration result. The process is expressed as:

[0032]

[0033] where φ is the parameter matrix, I M is the moving image, STN is the spatial transformation network, is the registered image;

[0034] Through spatial transformation operations, distorted images can be corrected, thus achieving precise registration of multi-modal remote sensing images and making different modal remote sensing images consistent in spatial position and geometric relationship.

[0035] Step 3 is specifically implemented according to the following steps:

[0036] Step 3.1: Use the training set generated in step 1 as the input of the network structure model in step 2 for supervised training. The total loss function used in the training process is:

[0037] Laff = αL1 + βL symaff + γL nmi (1)

[0038] where L1, L symaff , L nmi L aff are the difference loss, the symmetry loss based on inverse consistency, the normalization loss, and the total loss between the affine transformation parameters predicted by the network and the actual parameters, respectively; α, β, and γ are weight hyperparameters;

[0039] Step 3.2: The three loss terms are defined as follows:

[0040]

[0041] where N represents the number of training samples, φ prc φ gt and respectively represent the predicted affine transformation parameters and the ground truth; H1 represents the forward affine transformation matrix, H2 represents the inverse affine transformation matrix, E represents the identity matrix, represents matrix multiplication, i, j are the rows and columns of the matrix; A and B represent the registered image and the ground truth respectively, P AB (a, b) represents the joint probability distribution of the two images, P A (a) and P B (b) represent the marginal probability distributions of the images;

[0042] Step 3.3: Use the Adam method to update the model parameters, set the learning rate during the training process to 0.001, and stop training until the loss value converges or reaches the set maximum number of iterations, and obtain the trained dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features.

[0043] Step 4 is specifically implemented according to the following steps:

[0044] Input the test set in Step 1 into the dual-branch multi-modal remote sensing image registration network model trained in Step 3 to register the test data, and bring in the trained weight parameters, and finally output the registered image.

[0045] The beneficial effects of the present invention are as follows. The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features first extracts features through dual branches, skillfully integrating the characteristics of the upper-branch multi-scale and the lower-branch multi-directional. It can not only obtain rich multi-scale texture information but also capture global semantics and long-range correlations, thereby more comprehensively and accurately extracting the features of two different modal remote sensing images, effectively enhancing the feature expression ability and adaptability to different transformations, and laying a high-quality feature foundation for subsequent registration. Secondly, in the affine parameter regression stage, different convolutional kernels are used for parallel operations, and the features extracted by them complement each other, reducing the dependence on a single feature type, and the combination of different scale features enhances the parameter regression accuracy, and six affine parameters can be accurately learned. Finally, in the spatial transformation stage, a learnable STN module is used, which can explicitly perform spatial transformation on the input image, enhancing the network's adaptability to geometric deformations, and automatically correcting the image with the learned parameters, which is particularly suitable for remote sensing image registration involving large-scale rigid transformations. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flowchart for the implementation of the present invention;

[0047] Figure 2 is a framework diagram of the network model adopted by the present invention;

[0048] Figure 3 is a framework diagram of the Swin module in the network model of the present invention;

[0049] Figure 4 is a framework diagram of the CBAM module in the network model of the present invention;

[0050] Figure 5 is a framework diagram of the PRN module in the network model of the present invention;

[0051] Figure 6 is a framework diagram of the STN module in the network model of the present invention;

[0052] Figure 7 are the results obtained from the qualitative analysis of the network model of the present invention and other comparative methods in the experiment. DETAILED DESCRIPTION OF THE INVENTION

[0053] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0054] The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0055] Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing;

[0056] Step 1 is specifically implemented according to the following steps:

[0057] Step 1.1: The elements in the panchromatic and multispectral datasets are the cropped images of three remote sensing satellites, namely GF-1, WorldView-2, and WorldView-4. The resolution of the multispectral images in the panchromatic and multispectral datasets is 256×256 pixels, while the resolution of the panchromatic images is higher, at 1024×1024 pixels. During the experiment, for the purpose of registration, the multispectral images are regarded as moving images, and the panchromatic images are regarded as fixed images. Conversely, the multispectral images can be regarded as fixed images, and the panchromatic images as moving images.

[0058] Step 1.2: For the panchromatic and multispectral datasets, first downsample the panchromatic images to a size of 256×256 pixels for subsequent registration work. Randomly select 1104 pairs of images for training and 148 pairs of images for testing.

[0059] Step 2: Construct a dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0060] Combined with Figures 2 to 6 , in Step 2:

[0061] The structure of the dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features consists of three parts: a dual-branch feature extraction network, an affine parameter regression network, and a spatial transformation network. The panchromatic and multispectral remote sensing images obtained in Step 1 are respectively input into the same dual-branch feature extraction network to extract features containing global and local information. After splicing the deep features obtained from the dual-branch feature extraction network for the panchromatic and multispectral modalities, input them into the affine parameter regression network. Through the affine parameter regression network, learn 6 affine transformation parameters related to the registration task, including rotation, translation, and scaling. Then input the 6 affine transformation parameters into the spatial transformation network and apply them to the moving image to achieve the registration of the distorted image. The detailed implementation steps are as follows:

[0062] Combined with Figures 2 to 6 , the processing process of the dual-branch feature extraction network in Step 2 is as follows:

[0063] For remote sensing images in both panchromatic and multispectral modalities, the dual-branch feature extraction network deeply mines features from multiple scales and directions through a unique mechanism. The upper branch conducts multi-scale feature extraction. The input image first undergoes preliminary feature extraction through a basic 3×3 convolution, and then is divided into 4 branches according to different scales to process the data separately. After extracting features, the information at different levels is fused. Then, the number of channels is adjusted through a 1×1 convolution, and non-linearity is added through the GELU activation function. Finally, the output feature map is further refined through two Swin modules. The lower branch focuses on multi-directional feature extraction. The input image first undergoes preliminary feature extraction through two basic 3×3 convolutions, and then captures feature information from four directions: horizontal, vertical, left diagonal, and right diagonal to achieve a more comprehensive feature representation. Then, the features are concatenated and the distribution is normalized through a batch normalization layer, and the non-linearity is enhanced by a multi-layer perceptron. Then, the features are weighted from the channel and spatial dimensions by a dual attention module. Finally, a feature map is output through a 3×3 convolution.

[0064] After completing the dual-branch feature extraction and obtaining the deep features of the two modalities, these features are concatenated and input into the affine parameter regression network.

[0065] Combined with Figures 2 to 6 , the processing process of the affine parameter regression network in step 2 is as follows:

[0066] In this network, the present invention innovatively adopts parallel operations with convolution kernels of different sizes. The smaller convolution kernel focuses on extracting fine texture features, while the larger convolution kernel focuses on obtaining macroscopic structural features. The two cooperate with each other to generate a very rich and comprehensive feature representation. The core task of the parameter regression network is to derive the corresponding affine transformation parameters between the two remote sensing images based on the concatenated features. These parameters cover various transformation information such as translation, rotation, and scaling. Accurate affine transformation parameters are the key to achieving accurate image registration.

[0067] The fixed image and the moving image obtain deep features through the dual-branch feature extraction network, and they are concatenated and input into the parameter regression network to learn 6 affine transformation parameters related to the registration task, including rotation, translation, and scaling. The deep features are defined as:

[0068] Z F ,Z M =DBFE(I F ),DBFE(I M )

[0069] where I F ,I M ,Z F ,Z Mrespectively represent a fixed image, a moving image, deep features of the fixed image, and deep features of the moving image, and DBFE represents a dual-branch feature extraction network;

[0070] The six obtained parameters are defined as:

[0071] [θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 = FC(Resnet(Cat(Z F ,Z M )))

[0072] where θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 respectively represent six affine transformation parameters, FC represents a fully connected layer, Resnet represents a parameter regression network, and Cat(Z F ,Z M ) represents the concatenated deep features.

[0073] Combined with Figures 2 to 6 , the processing process of the spatial transformation network in step 2 is as follows:

[0074] Input the obtained affine transformation parameters into the spatial transformation network. In this step, the spatial transformation network will reshape the six affine transformation parameters into a parameter matrix φ according to the received six affine transformation parameters. This reshaping process is expressed as:

[0075] φ = reshape([θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 )

[0076] where φ is the parameter matrix, and θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 respectively represent six affine transformation parameters, and reshape is the reshaping operation;

[0077] Apply the obtained matrix to the moving image through spatial transformation to obtain the final registration result. The process is expressed as:

[0078]

[0079] Among them, φ is the parameter matrix, I M is the moving image, STN is the spatial transformation network, is the registered image;

[0080] Through the spatial transformation operation, the distorted image can be corrected, so as to realize the accurate registration of multi-modal remote sensing images, and make the remote sensing images of different modalities consistent in spatial position and geometric relationship.

[0081] Step 3: Train the dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features;

[0082] Step 3 is specifically implemented according to the following steps:

[0083] Step 3.1: Use the training set generated in Step 1 as the input of the network structure model in Step 2 for supervised training. The total loss function used in the training process is:

[0084] L aff = αL1 + βL symaff + γL nmi (1)

[0085] Among them, L1, L symaff , L nmi L aff are the difference loss between the affine transformation parameters predicted by the network and the actual parameters, the symmetric loss based on inverse consistency, the normalization loss, and the total loss respectively; α, β, and γ are weight hyperparameters;

[0086] Step 3.2: The three loss terms are defined as follows respectively:

[0087]

[0088] Among them, N represents the number of training samples, φ prc φ gt and respectively represent the predicted affine transformation parameters and the ground truth; H1 represents the forward affine transformation matrix, H2 represents the inverse affine transformation matrix, E represents the identity matrix, represents matrix multiplication, i, j are the rows and columns of the matrix; A and B represent the registered image and the ground truth respectively, P AB (a, b) represents the joint probability distribution of the two images, P A (a) and P B (b) represent the marginal probability distributions of the images;

[0089] Step 3.3: Update the model parameters using the Adam method, set the learning rate during the training process to 0.001, and stop training until the loss value converges or reaches the set maximum number of iterations, obtaining a trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features.

[0090] Step 4: Use the trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

[0091] Step 4 is specifically implemented according to the following steps:

[0092] Input the test set in Step 1 into the dual-branch multimodal remote sensing image registration network model trained in Step 3 to register the test data, and bring in the trained weight parameters, and finally output the registered image.

[0093] Embodiment 1

[0094] The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0095] Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing;

[0096] Step 2: Construct a dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0097] Step 3: Train the dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0098] Step 4: Use the trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

[0099] Step 5: Evaluate the quality of the registered image.

[0100] The specific implementation of Step 5 is as follows:

[0101] For the image registered in Step 4, use the reprojection error (RE) and root mean square error (RMSE) commonly used in strict registration as quantitative indicators. These two indicators evaluate the difference between the registered image and the ground truth, as well as the difference between the predicted affine transformation parameters and the actual parameters. Compare the registered image with the original panchromatic and multispectral images to conduct a qualitative evaluation of the registered image.

[0102] Regarding the registration of remote sensing images based on convolutional neural networks, the present invention can obtain rich multi-scale texture information by referring to neural networks with multi-scale and multi-directional features, and can also capture global semantics and long-range correlations, which helps to more comprehensively and accurately extract the features of two different modal remote sensing images. At the same time, introducing dual attention can weight the features corresponding to the dimensions, improve the feature extraction ability, enhance the model's attention to key information, reduce the sensitivity to noise, and improve the robustness and generalization ability.

[0103] Embodiment 2

[0104] The dual-branch multi-modal remote sensing image registration method based on multi-scale and multi-directional features of the present invention, the flowchart is as Figure 1 shown, and is specifically implemented according to the following steps:

[0105] Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing;

[0106] Step 1 is specifically implemented according to the following steps:

[0107] Step 1.1: The elements in the panchromatic and multi-spectral datasets are cropped images of three remote sensing satellites, namely Gaofen-1, WorldView-2, and WorldView-4. The resolution of the multi-spectral images in the panchromatic and multi-spectral datasets is 256×256 pixels, while the resolution of the panchromatic images is higher, at 1024×1024 pixels. During the experiment, for the purpose of registration, the multi-spectral images are regarded as moving images, and the panchromatic images are regarded as fixed images, or vice versa, the multi-spectral images are regarded as fixed images, and the panchromatic images are regarded as moving images;

[0108] Step 1.2: For the panchromatic and multi-spectral datasets, first downsample the panchromatic images to a size of 256×256 pixels for subsequent registration work. Randomly select 1104 pairs of images for training and 148 pairs of images for testing.

[0109] Step 2: Construct a dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features;

[0110] Combined with Figures 2 to 6 , in the said Step 2:

[0111] The dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features consists of three parts: a dual-branch feature extraction network, an affine parameter regression network, and a spatial transformation network. The panchromatic and multi-spectral remote sensing images obtained in step 1 are input into the same dual-branch feature extraction network to extract features containing global and local information respectively. After concatenating the deep features obtained from the dual-branch feature extraction network for the panchromatic and multi-spectral modalities, they are input into the affine parameter regression network. Through the affine parameter regression network, six affine transformation parameters related to the registration task, including rotation, translation, and scaling, are learned. Then, the six affine transformation parameters are input into the spatial transformation network and applied to the moving image to achieve the registration of the distorted image. The detailed implementation steps are as follows:

[0112] Step 3: Train the dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features;

[0113] Step 4: Use the trained dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

[0114] Example 3

[0115] The dual-branch multi-modal remote sensing image registration method based on multi-scale and multi-directional features of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0116] Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing;

[0117] Step 1 is specifically implemented according to the following steps:

[0118] Step 1.1: The elements in the panchromatic and multi-spectral datasets are cropped images of three remote sensing satellites, namely GF-1, WorldView-2, and WorldView-4. The resolution of the multi-spectral images in the panchromatic and multi-spectral datasets is 256×256 pixels, while the resolution of the panchromatic images is higher, at 1024×1024 pixels. During the experiment, for the purpose of registration, the multi-spectral images are regarded as moving images, and the panchromatic images are regarded as fixed images, or vice versa, with the multi-spectral images regarded as fixed images and the panchromatic images regarded as moving images;

[0119] Step 1.2: For the panchromatic and multi-spectral datasets, first downsample the panchromatic images to a size of 256×256 pixels for subsequent registration work. Randomly select 1104 pairs of images for training and 148 pairs of images for testing.

[0120] Step 2: Construct a dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features;

[0121] Combined with Figures 2 to 6, in Step 2:

[0122] The dual-branch multimodal remote sensing image registration network model structure based on multi-scale and multi-directional features consists of three parts: a dual-branch feature extraction network, an affine parameter regression network, and a spatial transformation network. The panchromatic and multispectral remote sensing images obtained in Step 1 are respectively input into the same dual-branch feature extraction network to extract features containing global and local information. After splicing the deep features obtained from the dual-branch feature extraction network for the panchromatic and multispectral modalities, the result is input into the affine parameter regression network. Through the affine parameter regression network, six affine transformation parameters related to the registration task, including rotation, translation, and scaling, are learned. Then, the six affine transformation parameters are input into the spatial transformation network and applied to the moving image to achieve the registration of the distorted image. The detailed implementation steps are as follows:

[0123] Combined with Figures 2 to 6 , the processing process of the dual-branch feature extraction network in Step 2 is as follows:

[0124] For the panchromatic and multispectral remote sensing images, the dual-branch feature extraction network deeply mines features from multiple scales and directions through a unique mechanism. The upper branch performs multi-scale feature extraction. The input image first undergoes preliminary feature extraction through a basic 3×3 convolution, and then is divided into 4 branches according to different scales to process the data separately. After extracting features, the information at different levels is fused. Then, the number of channels is adjusted through a 1×1 convolution, and the GELU activation function is used to add nonlinearity. Finally, the output feature map is further refined through two Swin modules. The lower branch focuses on multi-directional feature extraction. The input image first undergoes preliminary feature extraction through two basic 3×3 convolutions, and then captures feature information from four directions: horizontal, vertical, left diagonal, and right diagonal to achieve a more comprehensive feature representation. Then, the features are spliced and the distribution is normalized through a batch normalization layer, and the multi-layer perceptron enhances the nonlinearity. Finally, the features are weighted from the channel and spatial dimensions by the dual attention module, and finally, a 3×3 convolution outputs the feature map.

[0125] After completing the dual-branch feature extraction and obtaining the deep features of the two modalities, these features are spliced and input into the affine parameter regression network.

[0126] Combined with Figures 2 to 6 , the processing process of the affine parameter regression network in Step 2 is as follows:

[0127] In this network, the present invention innovatively adopts parallel operations of convolutional kernels of different sizes. The smaller convolutional kernels focus on extracting fine texture features, while the larger convolutional kernels focus on obtaining macroscopic structural features. The two cooperate with each other to generate extremely rich and comprehensive feature representations. The core task of the parameter regression network is to derive the corresponding affine transformation parameters between two remote sensing images based on the concatenated features. These parameters cover various transformation information such as translation, rotation, and scaling. Accurate affine transformation parameters are the key to achieving accurate image registration.

[0128] The fixed image and the moving image obtain deep features through a dual-branch feature extraction network, and they are concatenated and input into the parameter regression network to learn 6 affine transformation parameters related to the registration task, including rotation, translation, and scaling. The deep features are defined as:

[0129] Z F ,Z M =DBFE(I F ),DBFE(I M )

[0130] where I F ,I M ,Z F ,Z M represent the fixed image, the moving image, the deep features of the fixed image, and the deep features of the moving image respectively, and DBFE represents the dual-branch feature extraction network;

[0131] The 6 obtained parameters are defined as:

[0132] [θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 =FC(Resnet(Cat(Z F ,Z M )))

[0133] where θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 represent the 6 affine transformation parameters respectively, FC represents the fully connected layer, Resnet represents the parameter regression network, and Cat(Z F ,Z M ) represents the concatenated deep features.

[0134] Step 3: Train the dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-direction features;

[0135] Step 4: Use the trained dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

[0136] Example 4

[0137] The dual-branch multi-modal remote sensing image registration method based on multi-scale and multi-directional features of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0138] Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing;

[0139] Step 1 is specifically implemented according to the following steps:

[0140] Step 1.1: The elements in the panchromatic and multi-spectral datasets are cropped images of three remote sensing satellites, namely Gaofen-1, WorldView-2, and WorldView-4. In the panchromatic and multi-spectral datasets, the resolution of the multi-spectral image is 256×256 pixels, while the resolution of the panchromatic image is higher, at 1024×1024 pixels. During the experiment, for the purpose of registration, the multi-spectral image is regarded as the moving image, and the panchromatic image is regarded as the fixed image, or vice versa, the multi-spectral image is regarded as the fixed image, and the panchromatic image is regarded as the moving image;

[0141] Step 1.2: For the panchromatic and multi-spectral datasets, first downsample the panchromatic image to a size of 256×256 pixels for subsequent registration work. Randomly select 1104 pairs of images for training and 148 pairs of images for testing.

[0142] Step 2: Construct a dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features;

[0143] Combined with Figures 2 to 6 , in the said Step 2:

[0144] The structure of the dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features consists of three parts: a dual-branch feature extraction network, an affine parameter regression network, and a spatial transformation network. The panchromatic and multi-spectral remote sensing images obtained in Step 1 are respectively input into the same dual-branch feature extraction network to extract features containing global and local information; after the deep features obtained from the dual-branch feature extraction network for the panchromatic and multi-spectral modalities are concatenated, they are input into the affine parameter regression network. Through the affine parameter regression network, 6 affine transformation parameters including rotation, translation, and scaling related to the registration task are learned; then the 6 affine transformation parameters are input into the spatial transformation network and applied to the moving image to achieve the registration of the distorted image. The detailed implementation steps are as follows:

[0145] Combined with Figures 2 to 6, the processing of the dual-branch feature extraction network in step 2 is as follows:

[0146] For remote sensing images in both panchromatic and multispectral modalities, the dual-branch feature extraction network deeply mines features from multiple scales and directions through a unique mechanism. The upper branch performs multi-scale feature extraction. The input image first undergoes preliminary feature extraction through a basic 3×3 convolution, and then is divided into 4 branches according to different scales to process the data separately. After extracting features, the information at different levels is fused. Then, the number of channels is adjusted through a 1×1 convolution, and the GELU activation function is used to add non-linearity. Finally, the output feature map is further refined through two Swin modules. The lower branch focuses on multi-directional feature extraction. The input image first undergoes two basic 3×3 convolutions to extract preliminary features, and then captures feature information from four directions: horizontal, vertical, left diagonal, and right diagonal to achieve a more comprehensive feature representation. Then, the features are concatenated and normalized through a batch normalization layer, and the multi-layer perceptron enhances non-linearity. Then, the dual attention module weights the features from both channel and spatial dimensions. Finally, a 3×3 convolution outputs the feature map.

[0147] After completing the dual-branch feature extraction and obtaining the deep features of the two modalities, these features are concatenated and input into the affine parameter regression network.

[0148] Combined with Figures 2 to 6 , the processing of the affine parameter regression network in step 2 is as follows:

[0149] In this network, the present invention innovatively uses parallel operations with convolutional kernels of different sizes. The smaller convolutional kernel focuses on extracting fine texture features, while the larger convolutional kernel focuses on obtaining macroscopic structural features. The two cooperate with each other to generate a very rich and comprehensive feature representation. The core task of the parameter regression network is to derive the corresponding affine transformation parameters between the two remote sensing images based on the concatenated features. These parameters cover various transformation information such as translation, rotation, and scaling. Accurate affine transformation parameters are the key to achieving accurate image registration.

[0150] The fixed image and the moving image obtain deep features through the dual-branch feature extraction network, and they are concatenated and input into the parameter regression network to learn 6 affine transformation parameters related to the registration task, including rotation, translation, and scaling. The deep features are defined as:

[0151] Z F ,Z M =DBFE(I F ),DBFE(I M )

[0152] where I F ,I M ,Z F ,ZM represent a fixed image, a moving image, deep features of the fixed image, and deep features of the moving image respectively, and DBFE represents a dual-branch feature extraction network;

[0153] The six obtained parameters are defined as:

[0154] [θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 = FC(Resnet(Cat(Z F ,Z M )))

[0155] where θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 represent six affine transformation parameters respectively, FC represents a fully connected layer, Resnet represents a parameter regression network, and Cat(Z F ,Z M ) represents the concatenated deep features.

[0156] Combined with Figures 2 to 6 , the processing process of the spatial transformation network in step 2 is as follows:

[0157] Input the obtained affine transformation parameters into the spatial transformation network. In this step, the spatial transformation network will reshape the six affine transformation parameters into a parameter matrix φ according to the received six affine transformation parameters. This reshaping process is expressed as:

[0158] φ = reshape([θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 )

[0159] where φ is the parameter matrix, and θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 represent six affine transformation parameters respectively, and reshape is the reshaping operation;

[0160] Apply the obtained matrix to the moving image through spatial transformation to obtain the final registration result. The process is expressed as:

[0161]

[0162] Among them, φ is a parameter matrix, and I M is a moving image, STN is a spatial transformation network, is the registered image;

[0163] Through spatial transformation operations, distorted images can be corrected, thereby achieving precise registration of multimodal remote sensing images and making different modal remote sensing images consistent in spatial position and geometric relationship.

[0164] Step 3: Train a dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0165] Step 4: Use the trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

[0166] Example 5

[0167] The multimodal remote sensing image registration method based on multi-scale and multi-directional features of the present invention has a flowchart as Figure 1 shown, and is specifically implemented according to the following steps:

[0168] Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing;

[0169] Step 1 is specifically implemented according to the following steps:

[0170] Step 1.1: The elements in the panchromatic and multispectral datasets are cropped images of three remote sensing satellites, namely GF-1, WorldView-2, and WorldView-4. The resolution of the multispectral images in the panchromatic and multispectral datasets is 256×256 pixels, while the resolution of the panchromatic images is higher, at 1024×1024 pixels. During the experiment, for the purpose of registration, the multispectral images are regarded as moving images, and the panchromatic images are regarded as fixed images, or vice versa, with the multispectral images regarded as fixed images and the panchromatic images regarded as moving images;

[0171] Step 1.2: For the panchromatic and multispectral datasets, first downsample the panchromatic images to a size of 256×256 pixels for subsequent registration work. Randomly select 1104 pairs of images for training and 148 pairs of images for testing.

[0172] Step 2: Construct a dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0173] Combined with Figures 2 to 6 In the said step 2:

[0174] The dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features consists of three parts: a dual-branch feature extraction network, an affine parameter regression network, and a spatial transformation network. The panchromatic and multispectral remote sensing images obtained in Step 1 are respectively input into the same dual-branch feature extraction network to extract features containing global and local information. After concatenating the deep features of the panchromatic and multispectral modalities obtained from the dual-branch feature extraction network, they are input into the affine parameter regression network. Through the affine parameter regression network, six affine transformation parameters related to the registration task, including rotation, translation, and scaling, are learned. Then, the six affine transformation parameters are input into the spatial transformation network and applied to the moving image to achieve the registration of the distorted image. The detailed implementation steps are as follows:

[0175] Combined with Figures 2 to 6 , the processing process of the dual-branch feature extraction network in Step 2 is as follows:

[0176] For the panchromatic and multispectral remote sensing images, the dual-branch feature extraction network deeply excavates features from multiple scales and directions through a unique mechanism. The upper branch performs multi-scale feature extraction. The input image first undergoes preliminary feature extraction through a basic 3×3 convolution, and then is divided into 4 branches according to different scales to process the data respectively. After extracting features, the information of different levels is fused. Then, the number of channels is adjusted through a 1×1 convolution, and the GELU activation function is used to add non-linearity. Finally, the output feature map is further refined through two Swin modules. The lower branch focuses on multi-directional feature extraction. The input image first undergoes preliminary feature extraction through two basic 3×3 convolutions, and then captures feature information from four directions: horizontal, vertical, left diagonal, and right diagonal to achieve a more comprehensive feature representation. Then, the features are concatenated and the distribution is normalized through a batch normalization layer, and the multi-layer perceptron enhances non-linearity. Then, the dual attention module weights the features from the channel and spatial dimensions. Finally, the feature map is output through a 3×3 convolution once.

[0177] After completing the dual-branch feature extraction and obtaining the deep features of the two modalities, these features are concatenated and input into the affine parameter regression network.

[0178] Step 3: Train the dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0179] Step 3 is specifically implemented according to the following steps:

[0180] Step 3.1: Use the training set generated in Step 1 as the input of the network structure model in Step 2 for supervised training. The total loss function used in the training process is:

[0181] L aff = αL1 + βL symaff + γL nmi(1)

[0182] Among them, L1, L symaff , L nmi L aff are respectively the difference loss, the symmetry loss based on inverse consistency, the normalization loss, and the total loss between the affine transformation parameters predicted by the network and the actual parameters; α, β, and γ are weight hyperparameters;

[0183] Step 3.2: The three loss terms are defined as follows:

[0184]

[0185] Among them, N represents the number of training samples, φ prc φ gt and respectively represent the predicted affine transformation parameters and the ground truth; H1 represents the forward affine transformation matrix, H2 represents the inverse affine transformation matrix, E represents the identity matrix, represents matrix multiplication, i, j are the rows and columns of the matrix; A and B respectively represent the registered image and the ground truth, P AB (a, b) represents the joint probability distribution of the two images, P A (a) and P B (b) represent the marginal probability distributions of the images;

[0186] Step 3.3: Use the Adam method to update the model parameters, set the learning rate during the training process to 0.001, and stop training until the loss value converges or reaches the set maximum number of iterations, and obtain the trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features.

[0187] Step 4: Use the trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

[0188] Step 4 is specifically implemented according to the following steps:

[0189] Input the test set in Step 1 into the trained dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features in Step 3 to register the test data, and bring in the trained weight parameters, and finally output the registered image.

[0190] Example 6

[0191] Execute Step 1. The data set is the cropped images from three remote sensing satellites, GF-1, WorldView-2, and WorldView-4. After dividing them into training data and test data, perform data preprocessing on them;

[0192] Execute step 2 to construct a dual-branch multimodal remote sensing image registration network model based on multi-scale and multi-directional features;

[0193] Execute step 3 to train the dual-branch multimodal remote sensing image registration network model designed by the present invention using the training set divided in step 1;

[0194] Execute step 4 to register the test data divided in step 1 using the network model trained in step 3,

[0195] Execute step 5 to evaluate the registered images output in step 4, compare the present invention with five existing remote sensing image registration methods, and qualitatively analyze the registration results of the proposed method. The specific comparison results are as Figure 7 shown. RIFT is a registration method based on feature matching, which shows global mismatching phenomena during the registration process. When facing a large number of affine transformations in multimodal images, it is difficult to extract sufficient point features or perform incorrect feature matching. TWMM is a reliable registration method, but there are many misaligned places in the image registration. TransMorph is a method in the medical field that integrates convolutional neural networks and transformers. It shows large-scale rigid distortion on remote sensing images. SuperFusion is a method that integrates tasks such as registration, fusion, and segmentation within a single framework for infrared and optical datasets. Although surface fusion achieves a certain degree of registration on distorted motion images, it does not obtain complete alignment between the two images. ADRNet is the current state-of-the-art multimodal remote sensing image registration method, which adopts a two-stage method that combines affine transformation with flow field prediction to enhance the overall registration performance. Nevertheless, this method inevitably increases the number of parameters in the network, which brings a great burden to computing resources. Generally speaking, the method of the present invention still maintains good registration effects when facing large-scale rigid transformations of images.

Claims

1. A dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features, characterized in that The implementation is carried out specifically according to the following steps: Step 1: Divide the original image dataset into a training set and a test set, and perform preprocessing; Step 2: Construct a dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features; Step 3: Train the dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features; Step 4: Use the trained dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features to register the test data.

2. The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features according to claim 1, characterized in that The specific implementation of Step 1 is as follows: Step 1.1: The elements in the panchromatic and multi-spectral datasets are the cropped images of three remote sensing satellites, namely GF-1, WorldView-2, and WorldView-4. In the panchromatic and multi-spectral datasets, the resolution of the multi-spectral image is 256×256 pixels, while the resolution of the panchromatic image is 1024×1024 pixels. The multi-spectral image is regarded as the moving image, and the panchromatic image is regarded as the fixed image. Conversely, the multi-spectral image is regarded as the fixed image, and the panchromatic image is regarded as the moving image; Step 1.2: For the panchromatic and multi-spectral datasets, first downsample the panchromatic image to a size of 256×256 pixels, and randomly select 1104 pairs of images for training and 148 pairs of images for testing.

3. The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features according to claim 2, wherein Combined with FIGS. 2 to 6, in Step 2: The structure of the dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-directional features consists of three parts: a dual-branch feature extraction network, an affine parameter regression network, and a spatial transformation network. The panchromatic and multi-spectral remote sensing images obtained in Step 1 are respectively input into the same dual-branch feature extraction network to extract features containing global and local information; after the deep features obtained from the dual-branch feature extraction network for the panchromatic and multi-spectral modalities are concatenated, they are input into the affine parameter regression network. Through the affine parameter regression network, 6 affine transformation parameters related to the registration task, including rotation, translation, and scaling, are learned; then the 6 affine transformation parameters are input into the spatial transformation network and applied to the moving image to achieve the registration of the distorted image.

4. The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features according to claim 3, characterized in that The processing process of the dual-branch feature extraction network in Step 2 is as follows: For the panchromatic and multi-spectral remote sensing images, the upper branch of the dual-branch feature extraction network performs multi-scale feature extraction. The input image first undergoes preliminary feature extraction through a basic 3×3 convolution, and then is divided into 4 branches according to different scales to process the data respectively. After extracting features, the information at different levels is fused. Then, the number of channels is adjusted through a 1×1 convolution, and the GELU activation function is used to add non-linearity. Finally, the output feature map is further refined through two Swin modules. The lower branch focuses on multi-directional feature extraction. The input image first undergoes preliminary feature extraction through two basic 3×3 convolutions, and then captures feature information from four directions: horizontal, vertical, left diagonal, and right diagonal to achieve a more comprehensive feature representation. Then the features are concatenated and the distribution is normalized through a batch normalization layer, and the multi-layer perceptron enhances non-linearity. Then, the dual attention module weights the features from the channel and spatial dimensions. Finally, the feature map is output through a 3×3 convolution; After completing the extraction of dual-branch features and obtaining the deep features of the two modalities, these features are concatenated and input into the affine parameter regression network.

5. The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features according to claim 4, characterized in that, The processing procedure of the affine parameter regression network in step 2 is as follows: The fixed image and the moving image obtain deep features through the dual-branch feature extraction network, and they are concatenated and input into the parameter regression network to learn 6 affine transformation parameters related to the registration task, including rotation, translation, and scaling. The deep features are defined as: Z F ,Z M = DBFE(I F ), DBFE(I M ) Among them, I F , I M , Z F , Z M respectively represent the fixed image, the moving image, the deep features of the fixed image, and the deep features of the moving image. DBFE represents the dual-branch feature extraction network; The 6 obtained parameters are defined as: [θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 = FC(Resnet(Cat(Z F , Z M ))) where θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 respectively represent six affine transformation parameters, FC represents the fully connected layer, Resnet represents the parameter regression network, Cat(Z F ,Z M ) represents the concatenated deep features.

6. The dual-branch multi-modal remote sensing image registration method based on multi-scale and multi-directional features according to claim 5, wherein The processing procedure of the spatial transformation network in step 2 is as follows: The obtained affine transformation parameters are input into the spatial transformation network. In this step, the spatial transformation network will reshape the 6 affine transformation parameters into a parameter matrix φ according to the received 6 affine transformation parameters. This reshaping process is expressed as: φ = reshape([θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 ) where φ is a parameter matrix, θ 11 θ 12 θ 13 θ 21 θ 22 θ 23 respectively represent 6 affine transformation parameters, and reshape is an integer operation; The obtained matrix is applied to the moving image through spatial transformation to obtain the final registration result. The process is expressed as: Among them, φ is a parameter matrix, I M is a moving image, STN is a spatial transformation network, is the registered image; Through the spatial transformation operation, the distorted image can be corrected, so as to achieve the precise registration of multi-modal remote sensing images, and make the different-modal remote sensing images consistent in spatial position and geometric relationship.

7. The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features according to claim 6, characterized in that Step 3 is specifically implemented according to the following steps: Step 3.1: Use the training set generated in step 1 as the input of the network structure model in step 2 for supervised training. The total loss function used in the training process is: L aff = αL1 + βL symaff + γL nmi (1) Among them, L1, L symaff , L nmi L aff are the difference loss, the symmetry loss based on inverse consistency, the normalization loss, and the total loss between the affine transformation parameters predicted by the network and the actual parameters, respectively; α, β, and γ are weight hyperparameters; The three loss terms are respectively defined as follows: where N represents the number of training samples, φ prc φ gt and respectively represent the predicted affine transformation parameters and the ground truth; H1 represents the positive affine transformation matrix, H2 represents the inverse affine transformation matrix, E represents the identity matrix, denotes matrix multiplication where i, j are the rows and columns of the matrix; A and B represent the registered image and the ground truth respectively, P AB (a, b) represents the joint probability distribution of the two images, P A (a) and P B (b) represent the marginal probability distributions of the images; Use the Adam method to update the model parameters, set the learning rate during the training process to 0.001, and stop training until the loss value converges or reaches the set maximum number of iterations, to obtain the trained dual-branch multi-modal remote sensing image registration network model based on multi-scale and multi-direction features.

8. The dual-branch multimodal remote sensing image registration method based on multi-scale and multi-directional features according to claim 4, characterized in that Step 4 is specifically implemented according to the following steps: Input the test set in step 1 into the dual-branch multi-modal remote sensing image registration network model trained in step 3 to register the test data, and bring in the trained weight parameters to finally output the registered image.

Citation Information

Cited By

  • Inspection unmanned aerial vehicle non-aligned two-time-phase image intelligent change detection method

    CN120913115A