Physical driving-based multi-modal image loop iteration registration method

Through the physically driven multimodal image loop iterative registration method, combined with the SAR imaging mechanism learning model and the multi-scale optical image feature extraction module, the limitations of texture feature extraction in the prior art are solved, and the multimodal remote sensing image registration is achieved with high precision and robustness.

CN120088301APending Publication Date: 2025-06-03NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510163135.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing remote sensing image registration methods mainly focus on texture-based feature extraction, making it difficult to effectively deal with nonlinear deformation and radiation differences between multimodal remote sensing images, resulting in insufficient registration accuracy.

Method used

The multimodal image loop iterative registration method based on physics-driven is adopted to learn the SAR image's own features through the SAR imaging mechanism learning model, and combine the multi-scale optical image feature extraction module to use the multimodal feature sharing learning model for feature fusion and alignment. The dual loss strategy and LSTM iterative optimization method are adopted to gradually optimize network parameters to improve matching accuracy.

Benefits of technology

It improves the accuracy and robustness of multimodal remote sensing image registration, can handle nonlinear deformation and radiation differences more effectively, and meets the requirements of modern remote sensing image registration for high accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088301A_ABST
    Figure CN120088301A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image loop iteration registration method based on physical driving. The method comprises the following steps: learning own characteristics of an SAR image by using an SAR imaging mechanism learning model; inputting the optical image into a multi-scale optical image feature extraction module, and capturing optical information from local to global; carrying out feature fusion by using a multi-modal feature sharing learning model, and introducing sharing parameters to carry out feature alignment and preliminary comparison so as to match the most suitable point pair; in the feature matching optimization stage, an RIFT2 method is used as an external supervision signal, global features extracted by deep learning are combined, a dual loss strategy is adopted, and two pieces of key information from the RIFT2 are integrated into a loss function; and final registration is realized through a checkerboard visualization result. According to the invention, through an efficient heterogeneous image registration technology, the capability of realizing heterogeneous image accurate registration under the influence of different differences of heterogeneous images is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a physically-driven multimodal image cyclic iterative registration method, belonging to the cross-field of computer vision, remote sensing technology and geographic information system, and is applied to the registration task of misaligned heterogeneous remote sensing image pairs with different resolutions. Background Technique

[0002] Currently, the global digital development is accelerating day by day. With the launch of numerous remote sensing satellites and remote sensing facilities, the integrated three-dimensional observation technology of sky-earth multi-sensors has been comprehensively developed, marking the arrival of the new infrastructure and intelligent photogrammetry era. Remote sensing image registration is the cornerstone of intelligent remote sensing and a basic step for many remote sensing image applications.

[0003] Image registration is a process of matching and superimposing two or more images obtained under different times, different sensors, different perspectives, and different imaging conditions of the same scene. Due to the differences in spatial geometric configuration and physical radiation mechanism among different sensors, and the different imaging platforms carried, there are usually one or several of the significant "five differences" (radiation difference, geometric difference, scale difference, perspective difference, and temporal difference) among multimodal remote sensing images.

[0004] With the continuous development of remote sensing technology, the resolution of satellite sensors has been gradually improved, and the improvement of dataset quality has laid a solid foundation for the registration task. At the same time, the enhancement of computer computing power and the progress of machine learning algorithms have made it possible to carry out remote sensing heterogeneous image registration using more powerful algorithms, greatly improving the registration accuracy. Currently, important research progress has been made in the technology of heterogeneous remote sensing image registration.

[0005] Early image registration methods mainly focused on traditional methods, which were usually divided into two categories: region-based methods and feature-based techniques. Later, with the rise of deep learning methods, researchers gradually turned their attention to this field. The methods using deep learning technology are mainly divided into three categories: 1) integrating deep learning into the traditional image registration process to form a new registration method; 2) using deep learning to convert one image modality into another, simplifying the registration of multi-source remote sensing images into the registration of single-modal remote sensing images; 3) directly regressing the transformation parameters between multi-modal images to achieve end-to-end multi-source remote sensing image registration from the image pair to the transformation metric.

[0006] Whether it is artificial feature extraction based on traditional algorithms such as SIFT (Scale-Invariant Feature Transform) and SURF (Speeded Up Robust Features), or automatic feature extraction in deep learning through CNN (Convolutional Neural Network), both methods need to extract representative image features and evaluate the matching degree between two images through similarity measurement. Common similarity measurement methods include Normalized Cross-Correlation (NCC) and Cosine Similarity in traditional methods, as well as Structural Similarity Index (SSIM) in deep learning methods. Finally, based on the results of similarity measurement, the optimal registration transformation is determined.

[0007] Traditional methods rely on artificial design and definition of specific image features. However, when dealing with complex scenes or images with large variations, these features may not be able to effectively capture all key image information, resulting in poor flexibility. In addition, traditional methods have certain limitations in dealing with non-linear deformation, radiation differences, and local optimization problems, and it is difficult to meet the requirements of high precision and robustness for modern remote sensing image registration. Although existing heterologous remote sensing image registration methods have made significant progress in high-dimensional deep feature extraction through deep learning, generating more effective feature descriptors or similarity measurements, and improving the resistance to non-linear radiation differences between multi-modal remote sensing images, most current deep learning methods still mainly focus on texture-based feature extraction and matching, and rarely consider the special effects of imaging mechanisms such as SAR. In addition, when dealing with complex scenes, there are still deficiencies in the matching accuracy of traditional matching methods and deep learning networks based on single-channel optimization, especially when there are significant non-linear deformations, this problem is particularly prominent, resulting in great challenges for high-precision image registration. Summary of the Invention

[0008] The purpose of the present invention is to provide a physically-driven multi-modal image cyclic iterative registration method to solve the limitation problem that existing methods mainly focus on texture-based feature extraction.

[0009] To achieve the above purpose, the present invention adopts the following technical solutions:

[0010] A physically-driven multi-modal image cyclic iterative registration method includes the following steps:

[0011] (1) Use the SAR imaging mechanism learning model to learn the inherent features of SAR images;

[0012] (2) Input the optical image into the multi-scale optical image feature extraction module to capture optical information from local to global, so as to obtain optical features at different levels;

[0013] (3) Use the multi-modal feature sharing learning model to perform feature fusion on the self-owned features of the SAR image obtained in step (1) and the optical features obtained in step (2), introduce shared parameters for feature alignment and preliminary comparison to match the most suitable point pairs;

[0014] (4) Optimize the matching of the features fused in step (3). In the feature matching optimization stage, use the RIFT2 method as an external supervision signal, and combine the global features extracted by deep learning. Adopt a dual loss strategy to integrate two key pieces of information from RIFT2: the rotation principal direction extracted offline and the aligned feature descriptors into the loss function; use the internal alternating multiple iteration method to gradually optimize the network parameters, improve the matching accuracy by setting thresholds and feeding back new loss functions, and achieve the final registration through checkerboard visualization results.

[0015] In step (1), the entire SAR imaging mechanism learning model is divided into three branches. First, the input image is processed through three convolutional layers to generate an initial feature map;

[0016] After obtaining the initial feature map, the physical feature maps generated by the SAR imaging mechanism learning model through formulas are integrated into additional input features and passed as additional feature channels to the subsequent layers of the network. Based on the radar equation and the parameters estimated deterministically and empirically, a feature map is generated for each pixel position in the SAR image; each pixel position corresponds to its scattering characteristics under physical constraints, and the generated image shows the theoretical scattering intensity of each pixel; the last branch uses the input feature F∈R H×W×C , where H and W are the spatial resolutions and C is the number of channels. Group the input feature map F according to the channel dimension, and then evenly divide the number of channels C into K groups to form multiple sub-region physical feature maps (F 1 , F 2 ,..., F k ), where the number of channels of each sub-region physical feature map is Next, in order to map each group of local features to the physical high-level semantic space, use independent grouped convolutional kernels W K Extract the target feature map for each group of local physical features F K respectively:

[0017] F k ' = GroupConv(F k , W k )

[0018] Among them, F k ' is the feature map after convolution. GroupConv is group convolution, and each convolution kernel corresponds to an independent sub-feature map;

[0019] Then, the feature map F k ' generated by the physical driver extracted by group convolution will be element-wise multiplied with the corresponding input feature map F m that is not guided by physical characteristics; among them, the preliminary input feature map F m ∈ R H×W×C is also grouped according to the channel dimension, and the grouping method is the same as that of F k ; the final input feature map is weighted by the physical feature map to emphasize the features of the regions with high physical responses, thereby realizing the physical-driven dynamic enhancement of the input feature map; finally, the refined feature groups are connected to generate the output physical attribute map:

[0020] M k = F k ' eF m

[0021] Among them, M k represents the physical importance weights at different channel positions, and M k is reconnected to form the complete output feature map;

[0022] The feature maps obtained from the three branches are weighted and fused to generate a feature map with enhanced physical characteristics.

[0023] In step (2), a multi-scale feature learning branch is designed for the optical image, and different-resolution feature maps are extracted using a three-layer convolution block; let the patch of the multi-modal image be I oi and I si ; I oi is the i th -th optical image block, and I si is the SAR image block corresponding to I oi , where i in the SAR image and the optical image is the same, and as long as i is equal, they will match. Then, the features of the multi-modal image obtained from the two branch networks are:

[0024] g oi = G O (I oi )

[0025] g si = G S (I si )

[0026] Among them, G O represents the feature extraction branch of the optical image, G SRepresents the feature extraction branch of the SAR image, g oi is the learned optical feature of I oi and g si is the learned SAR feature of I si .

[0027] In step (3), design a network structure with shared parameters to extract the shared features between the inherent features and optical features of the SAR image:

[0028] f oi = G C (g oi )

[0029] f si = G C (g si )

[0030] where f oi is the shared feature of I oi , f si is the shared feature of I si , and G C represents the shared feature network;

[0031] By comparing the similarity of the feature vectors, determine whether the two images correspond; after obtaining the similarity score, directly generate the matching line points according to the most similar points; calculate the correlation of the features between the optical image and the SAR image in the shared feature space, and use the following formula to represent it:

[0032]

[0033] where f opt and f sar correspond to the feature vectors of the optical and SAR images respectively, d represents the Euclidean distance, D represents the dimension of the feature vector, and j represents the dimension index of the feature vector, from 1 to D.

[0034] In step (4), set the RMSE threshold. If the RMSE obtained in the preliminary matching stage has met the matching condition, there is no need to enter the iterative optimization and RIFT2 supervision module; if the ideal RMSE has still not been achieved, discard the image pairs in this part;

[0035] During the iterative optimization process, compare the matching points generated by the deep learning network with the RIFT2 supervision signal obtained in the offline stage, and construct a loss function to guide the optimization process of the matching network;

[0036] LSTM is used as the feature vector f = [f 1 , f 2 ,..., f DThe input of

[0037] The iterative update formula of the LSTM is as follows:

[0038] h t = LSTM(h t-1 , x t )

[0039] where x t is the current input feature, and h t is the hidden state and output updated by the LSTM in the t-th iteration.

[0040] In step (4), design the first loss function scattering characteristic loss L scat for measuring and minimizing the difference between the scattering characteristics estimated by the network and the true scattering characteristics, so as to ensure that the model can accurately learn and identify the scattering mechanism of SAR images;

[0041] L scat = D KL (S true || S pred )

[0042]

[0043] where S pred represents the distribution of the scattering characteristics predicted by the network, S true represents the distribution of the true scattering characteristics, D KL represents the Kullback-Leibler divergence, and N represents the total number of true values.

[0044] In the SAR imaging mechanism learning model, there is a loss function for high-level physical semantic features. In order to ensure the consistency of the geometric positions of scattering centers in the registration task, a geometric constraint loss L geo is introduced:

[0045]

[0046] where is the geometric position of the scattering center predicted by the network, is the geometric shape of the actual scattering center;

[0047] The total scattering characteristic loss L scat is:

[0048] L scat = α 1 L dis + α2 L geo

[0049] Among them, α 1 and α 2 are weight parameters.

[0050] In the said step (4), the loss function applies the matching loss L match through gradual fusion and adaptive weight adjustment to improve the accuracy and stability of feature matching; this process divides the loss function into two stages: the first stage includes the contrastive loss L contrastive , which is used to distinguish positive and negative matching pairs, ensure that positive sample pairs are close and negative sample pairs are far apart, and further enhance the separability of the iterative optimization space;

[0051]

[0052] Among them, and are feature pairs from optical and SAR images respectively, y i is the label of the matching pair. When y i = 1, it means that the feature pair is a correctly matched point pair. If y i = 0, it means a non-matching point pair. represents the Euclidean distance between features, and P is the boundary value;

[0053] The main direction and alignment descriptor of RIFT2 are used as reference signals to guide the feature alignment and direction consistency of the matching module, thereby improving the rotational robustness of the matching, corresponding to the losses L rotation_match and L descriptor_match respectively;

[0054]

[0055] Among them, θ DL represents the direction information generated by the network during the matching process, and θ RIFT represents the main rotational direction extracted by RIFT2. represents the predicted matching point at the i-th iteration, f DL represents the feature vector generated by the network during the matching pair process, and f RIFT represents the direction alignment descriptor extracted by RIFT2;

[0056] The loss f RIFT supervised by RIFT2 includes:

[0057] L RIFT = L rotation_match +L descriptor_match

[0058] In the iterative optimization section, a loss function is required to guide the optimization process in a better direction. The iterative formula is as follows:

[0059]

[0060] where the gradient represents the gradient of the loss at the matching points, and γ represents the learning rate, which controls the step size of each iterative optimization update. Therefore, the matching loss L match is obtained as follows:

[0061] L match = L contrastive + α 3 L RIFT

[0062] where α 3 is the weight parameter;

[0063] The entire loss function L total is as follows:

[0064] L total = L scat + L contrastive + α 3 L RIFT

[0065] The registration result is continuously optimized through the loss function, and the transformation matrix corresponding to the optical image and the SAR image is estimated. Finally, the final registration result is presented through connection lines and checkerboard displays.

[0066] Advantageous effects: The present invention continuously optimizes the registration result through the loss function, estimates the transformation matrix corresponding to the optical image and the SAR image, and finally presents the final registration result through connection lines and checkerboard displays. The present invention improves the ability to achieve accurate registration of heterogeneous images under the influence of different differences in heterogeneous images through an efficient heterogeneous image registration technology, serving multiple important application fields such as multi-temporal change detection, multi-view image analysis, and multi-source information fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is the overall flowchart of the physics-driven multi-modal heterogeneous image cyclic iterative registration method;

[0068] Figure 2 is the input SAR image;

[0069] Figure 3 is the specific flowchart of the SAR imaging mechanism learning module;

[0070] Figure 4 is the specific framework diagram of the physics-driven mapping feature map;

[0071] Figure 5is the input optical image;

[0072] Figure 6 is the LSTM framework diagram;

[0073] Figure 7 is the specific flow chart of iterative optimization;

[0074] Figure 8 is the registration result of different methods for the SEN1_2 dataset comparison;

[0075] Figure 9 is the registration experiment result diagram of the SEN1_2 dataset;

[0076] Figure 10 is the registration experiment result diagram of the WHU-OPT-SAR dataset;

[0077] Figure 11 is the checkerboard experiment result diagram of the QXSLAB_SAROPT dataset;

[0078] Figure 12 is the result diagram of the influence of the existence of different loss functions on the registration accuracy;

[0079] Figure 13 is the registration experiment result diagram of the Dajinshan dataset. Detailed implementation mode

[0080] The following further explains the present invention in conjunction with the accompanying drawings.

[0081] As Figure 1 shown, a physical-driven multi-modal image cyclic iterative matching method of the present invention includes the following steps:

[0082] (1) First, use the SAR imaging mechanism learning model to learn the inherent features of the SAR image. The SAR image features are used as prior knowledge to guide this part of the learning process.

[0083] In step (1), the input image is as Figure 2 shown. The entire SAR mechanism learning module is divided into three branches. First, the input image is processed through three convolutional layers to generate initial feature maps, which usually contain basic information such as edges and textures. After preliminary convolutional feature extraction, the physical feature maps generated by the formula in SAR imaging are integrated as additional input features and passed as additional feature channels to the subsequent layers of the network. Specifically, based on the radar equation and the parameters estimated deterministically and empirically, feature maps are generated for each pixel position in the SAR image. The specific framework diagram of the SAR mechanism learning module is as Figure 3As shown, where each pixel position corresponds to its scattering characteristics under physical constraints, and the generated image shows the theoretical scattering intensity of each pixel. The last branch uses input features F ∈ R H×W×C , where H and W are the spatial resolutions and C is the number of channels. The input feature map F is grouped according to the channel dimension, and then the number of channels C is evenly divided into K groups to form multiple sub-region physical feature maps (F 1 , F 2 ,..., F k ), k ∈ K, where the number of channels of each sub-region physical feature map is Next, in order to map each group of local features to the physical high-level semantic space, independent grouped convolutional kernels W K are used to extract each group of target feature maps for the local physical feature F K respectively.

[0084] F k ' = GroupConv(F k , W k )

[0085] where F k ' is the convolutional feature map, GroupConv is grouped convolution, and each convolutional kernel corresponds to an independent sub-feature map.

[0086] Then, the feature map F k ' generated by the physical driver extracted by grouped convolution will be element-wise multiplied with the corresponding input feature map F m that is not guided by physical characteristics. Among them, the preliminary input feature map F m ∈ R H×W×C is also grouped according to the channel dimension, and the grouping method is the same as that of F k . The final input feature map is weighted by the physical feature map to emphasize the features of the regions with high physical responses, thereby realizing the physically driven dynamic enhancement of the input feature map. Finally, the refined feature groups are concatenated to generate the output physical attribute map.

[0087] M k = F k ' ⊙ F m

[0088] where M k represents the physical importance weights at different channel positions, and M k is reconnected to form the complete output feature map. The output feature map here not only contains the original information of the input feature map but also integrates the semantic enhancement information of the physical feature map. The framework diagram of this part is as shown in Figure 4 .

[0089] Finally, the feature maps obtained from the three branches are weighted and fused to generate a feature map with enhanced physical properties. Through this fusion method, the network can learn data-driven features while referring to the physical model constraints, providing a collaborative optimization framework for physical and data-driven features, making the training process more in line with physical principles.

[0090] (2) Input the optical image into the multi-scale optical image feature extraction module to capture optical information from local to global. Through preliminary feature extraction, ensure the acquisition of optical features at different levels.

[0091] In step (2), the input image is as Figure 5 shown. To address the multi-faceted and significant differences between SAR images and optical images, since the present invention designs a learning network for the SAR imaging mechanism. The present invention designs a multi-scale feature learning branch for optical images, using three-layer convolutional blocks to extract feature maps of different resolutions. This enables the adaptive learning of multi-modal image features and the extraction of their respective unique basic features from SAR and optical images respectively. Let the patch of the multi-modal image be I oi and I si . I oi is the i th th optical image patch, and I si is the corresponding SAR image patch to I oi . Where i in the SAR and optical images is the same, as long as i is equal, they will match. Then the features of the multi-modal images obtained from the two branch networks are:

[0092] g oi =G O (I oi )

[0093] g si =G S (I si )

[0094] Where G O represents the feature extraction branch of the optical image, G S represents the feature extraction branch of the SAR image, and g oi is the learned optical feature of I oi , and g si is the learned SAR feature of I si .

[0095] (3) Then use the multi-modal feature sharing learning model to fuse the self-owned features of the SAR image obtained in step (1) and the optical features obtained in step (2), introduce shared parameters for feature alignment and preliminary comparison to match the most suitable point pairs.

[0096] In step (3), after the extraction of respective features, the differences in multi-modal features are mainly reflected in their respective basic features, but they still contain the same semantic information. Therefore, after the extraction of proprietary features, it is necessary to obtain the shared features of the multi-modal images and map the features of SAR and optical images to the common feature space through shared weights. In this space, the features of two different modalities of images can be directly compared to preliminarily determine their similarity. The present invention designs a network structure with shared parameters to extract the shared features between the two.

[0097] f oi = G C (g oi )

[0098] f si = G C (g si )

[0099] Among them, f oi is the shared feature of I oi , and f si is the shared feature of I si . G C represents the shared feature network.

[0100] By comparing the similarity of feature vectors, the network can determine whether two images correspond. After obtaining the similarity score, it can directly generate the matching line points according to the most similar points. Calculate the correlation of features between the optical image and the SAR image in the shared feature space, and use the following formula to represent it:

[0101]

[0102] Among them, f opt and f sar correspond to the feature vectors of the optical and SAR images respectively, d represents the Euclidean distance, D represents the dimension of the feature vector, and j represents the dimension index of the feature vector, from 1 to D.

[0103] (4) Match and optimize the features fused in step (3). In the feature matching and optimization stage, use the traditional RIFT2 method as the external supervision signal, and combine the global features extracted by deep learning. Adopt a dual-loss strategy to integrate two key pieces of information from RIFT2: the rotation principal direction extracted offline and the aligned feature descriptors into the loss function. Use the internal alternating multiple iteration method to gradually optimize the network parameters, improve the matching accuracy by setting thresholds and feeding back the new loss function, and achieve the final registration through the visualization result of the checkerboard.

[0104] In step (4), the next crucial step is how to ensure more accurate and optimized matching features between these two modes under the premise of insufficient precision. The present invention proposes an innovative iterative optimization method for feature matching. The LSTM framework diagram is as shown in Figure 6 and the iterative optimization method is as shown in Figure 7 . This method is supervised by multiple optimization iterations and traditional methods to finally complete the matching. Specifically, although LSTM is mainly used to process time series data, it can also be used as the input of the feature vector f = [f 1 , f 2 ,..., f D , where D represents the dimension of the feature vector. LSTM processes these inputs in sequence and updates its hidden state, gradually accumulating an understanding of the overall features. In addition, the storage unit and gating mechanism of LSTM enable it to maintain important historical information in multiple iterations and make adjustments based on this information. During the iteration process, LSTM can also adaptively adjust the registration features through the error feedback mechanism, gradually reducing the RMSE.

[0105] The present invention sets different RMSE thresholds according to the conditions of different data sets in the experiment. If the RMSE obtained in the preliminary matching stage already meets the matching conditions, there is no need to enter the iterative optimization and RIFT2 supervision module. Considering the challenges of processing large-scale data sets in the registration task and aiming to significantly reduce the computational cost and minimize resource consumption, we limit the number of iterations to no more than 5 times. If the ideal RMSE is still not achieved, the image pairs in this part are discarded. During the iterative optimization process, the matching points generated by the deep learning network are compared with the RIFT2 supervision signals obtained in the offline stage, and a loss function is constructed to guide the optimization process of the matching network. The combination of this traditional method and modern deep learning improves the robustness of the matching while ensuring high precision of the results.

[0106] The formula for iterative update of LSTM is as follows:

[0107] h t = LSTM(h t-1 , x t )

[0108] where x t is the current input feature, and h t is the hidden state and output updated by LSTM in the t-th iteration.

[0109] In step (4), in addition to improving the structure of the registration network, the present invention also designs the first loss function scattering characteristic loss L scatIts purpose is to measure and minimize the difference between the scattered characteristics estimated by the network and the true scattered characteristics, so as to ensure that the model can accurately learn and identify the scattering mechanism of SAR images.

[0110] L scat = D KL (S true ||S pred )

[0111]

[0112] Wherein, S pred represents the distribution of the scattered characteristics predicted by the network, and S true represents the distribution of the true scattered characteristics, D KL represents the Kullback-Leibler scatter, and N represents the total number of true values.

[0113] According to the SAR imaging mechanism learning model proposed by the present invention, a loss function for high-level physical semantic features should also be designed. In order to ensure that the geometric positions of the scattering centers are consistent in the registration task, a geometric constraint loss L geo is introduced here:

[0114]

[0115] Wherein, is the geometric position of the scattering center predicted by the network, is the geometric shape of the actual scattering center.

[0116] The total scattered characteristic loss L scat is:

[0117] L scat = α 1 L dis + α 2 L geo

[0118] Wherein, α 1 and α 2 are weight parameters, and the sum of α 1 and α 2 is 1, and they are respectively set to 0.3 and 0.7 here.

[0119] Previously, in the feature matching network, methods of multiple optimization iterations and iterative optimization supervised by traditional methods were introduced. On this basis, the present invention also correspondingly proposes a loss function: by gradually fusing and adaptively adjusting weights, the matching loss L match is applied to improve the accuracy and stability of feature matching. This process divides the loss function into two stages. The first stage includes the contrast loss L contrastive, mainly used to distinguish positive and negative matching pairs, ensure that positive sample pairs are close and negative sample pairs are far apart, and further enhance the separability of the iterative optimization space;

[0120]

[0121] Among them and are feature pairs from optical and SAR images respectively, and y i belongs to the label of the matching pair. When y i = 1, it means that the feature pair is a correctly matched point pair. If y i = 0, it means a non-matching point pair. represents the Euclidean distance between features, and P is the boundary value, usually set to 1.

[0122] The main direction and alignment descriptor of RIFT2 are used as reference signals to guide the feature alignment and direction consistency of the matching module, thereby improving the rotational robustness of the matching, corresponding to losses L rotation_match and L descriptor_match .

[0123]

[0124] Among them, θ DL represents the direction information generated by the network during the matching process, and θ RIFT represents the main rotation direction extracted by RIFT2. represents the predicted matching point at the i-th iteration, f DL represents the feature vector generated by the network during the matching pair process, and f RIFT represents the direction alignment descriptor extracted by RIFT2.

[0125] The loss f RIFT generated by RIFT2 supervision includes:

[0126] L RIFT = L rotation_match + L descriptor_match

[0127] In the iterative optimization part, a loss function is needed to guide the optimization process in a better direction. The iterative formula is as follows:

[0128]

[0129] Among them, the gradient represents the gradient of the loss on the matching point, and γ represents the learning rate, which controls the step size of each iterative optimization update.

[0130] Therefore, the matching loss L match can be obtained as:

[0131] L match = L contrastive + α 3 L RIFT

[0132] where α 3 is the weight parameter, and α 3 takes 0.5 here.

[0133] The entire loss function L total is as follows:

[0134] L total = L scat + L contrastive + α 3 L RIFT

[0135] The present invention continuously optimizes the registration result through the loss function, estimates the transformation matrix corresponding to the optical image and the SAR image, and finally displays the final registration result through wire connection and checkerboard display.

[0136] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A multimodal image iterative registration method based on physics, characterized by: The following steps are involved: (1) Using the SAR imaging mechanism learning model to learn the inherent characteristics of SAR images; (2) Inputting the optical image into a multi-scale optical image feature extraction module to capture optical information from local to global to obtain optical features at different levels; (3) Using a multimodal feature sharing learning model, the inherent features of the SAR image obtained in step (1) and the optical features obtained in step (2) are fused, and shared parameters are introduced to perform feature alignment and preliminary comparison to match the most appropriate point pair; (4) The features fused in step (3) are matched and optimized. In the feature matching optimization stage, the RIFT2 method is used as the external supervision signal, combined with the global features extracted by deep learning, and a dual loss strategy is adopted to integrate two key information from RIFT2: the main rotation direction extracted offline and the aligned feature descriptor into the loss function; the network parameters are gradually optimized using an internal alternating multiple iteration method, and the matching accuracy is improved by threshold setting and feedback of the new loss function. The final alignment is achieved through checkerboard visualization results.

2. According to claim 1, a multi-modal image iterative registration method based on physical drive is characterized by: In the step (1), the entire SAR imaging mechanism learning model is divided into three branches. First, the input image is processed through three convolutional layers to generate an initial feature map; After obtaining the initial feature map, the physical feature map generated by the formula in the SAR imaging mechanism learning model is integrated as an additional input feature and passed to the subsequent layers of the network as an additional feature channel. Based on the radar equation and the parameters estimated by determinism and experience, a feature map is generated for each pixel position in the SAR image; each pixel position corresponds to its scattering characteristics under physical constraints, and the generated image shows the theoretical scattering intensity of each pixel; the last branch uses the input feature F∈R rich in physical information H×W×C , where H and W are spatial resolutions, C is the number of channels, the input feature map F is grouped according to the channel dimension, and then the number of channels C is evenly divided into K groups to form multiple sub-region physical feature maps (F1, F2, ..., F k ), where the number of channels of each sub-region physical feature map is Next, in order to map each group of local features to the physical high-level semantic space, it will be passed through an independent group convolution kernel W K is the local physical feature F K Extract each group of target feature maps separately: F k '=GroupConv(F k ,W k ) Among them, F k ' is the feature map after convolution, GroupConv is group convolution, and each convolution kernel corresponds to an independent sub-feature map; Then, the feature map F generated by the physical driver extracted by the group convolution k 'Compare with the corresponding input feature map F that is not guided by physical characteristics m Element-by-element multiplication; the input preliminary feature map is also grouped according to the channel dimension, and the grouping method is the same as F k consistent; the final input feature map is weighted by the physical feature map to emphasize the features of regions with high physical response, thereby achieving a physically driven dynamic enhancement of the input feature map; finally, the refined feature groups are concatenated to generate the output physical property map: M k =F k '⊙F m Among them, M k represents the physical importance weights of different channel positions, and M k Reconnect to form the complete output feature map; The feature maps obtained from the three branches are weighted and fused to generate a feature map with enhanced physical properties.

3. The method of multimodal image iterative registration based on physical drive according to claim 1, characterized in that: In the step (2), a multi-scale feature learning branch is designed for the optical image, and feature maps of different resolutions are extracted using three-layer convolution blocks; assuming that the patch of the multimodal image is I oi and I si ;I oi is the i th Optical image blocks, I si Yes and I oi The corresponding SAR image patches, where i in the SAR image and the optical image are the same, as long as i is equal, they will match, then the characteristics of the multimodal image obtained from the two branch networks are: g oi =G O (I oi ) g si =G S (I si ) Among them, G O represents the feature extraction branch of the optical image, G S represents the feature extraction branch of SAR image, g oi isI oi The optical characteristics of the study, g si isI si The learned SAR features.

4. The method of multimodal image iterative registration based on physical drive according to claim 1, characterized in that: In the step (3), a network structure with shared parameters is designed to extract the shared features between the inherent features and optical features of the SAR image: f oi =G C (g oi ) f si =G C (g si ) Among them, f oi isI oi The shared features of si isI si The shared features of G C represents a shared feature network; By comparing the similarity of feature vectors, determine whether the two images correspond; after obtaining the similarity score, directly generate matching line points based on the most similar points; calculate the correlation between the features of the optical image and the SAR image in the shared feature space, and use the following formula to express it: Among them, f opt and f sar They correspond to the feature vectors of optical and SAR images respectively, d represents the Euclidean distance, D represents the dimension of the feature vector, and j represents the dimension index of the feature vector, from 1 to D.

5. The method for multimodal image iterative registration based on physical drive according to claim 1, characterized in that: In the step (4), the RMSE threshold is set. If the RMSE obtained in the preliminary matching stage has satisfied the matching condition, there is no need to enter the iterative optimization and RIFT2 supervision modules; if the ideal RMSE is still not achieved, the image pairs in this part are discarded; During the iterative optimization process, the matching points generated by the deep learning network are compared with the RIFT2 supervision signal obtained in the offline stage, and a loss function is constructed to guide the optimization process of the matching network; LSTM as feature vector f=[f1,f2,...,f D ], where D represents the dimension of the feature vector; LSTM will process these inputs sequentially and update its hidden state, gradually accumulating an understanding of the overall features; The formula for LSTM iterative update is as follows: h t =LSTM(h t-1 ,x t ) Among them, x t is the current input feature, h t are the updated hidden states and outputs of the LSTM at the tth iteration.

6. The method of multimodal image iterative registration based on physical drive according to claim 1, characterized in that: In the step (4), a first loss function scattering characteristic loss L based on the distribution distance is designed. scat , which is used to measure and minimize the difference between the scattering characteristics estimated by the network and the true scattering characteristics to ensure that the model can accurately learn and identify the scattering mechanism of SAR images; L scat =D KL (S true ||S pred ) Among them, S pred represents the distribution of scattering properties predicted by the network, S true represents the distribution of the true scattering properties, D KL represents Kullback-Leibler scattering, and N represents the total number of true values.

7. The method of multimodal image iterative registration based on physical drive according to claim 1, characterized in that: The SAR imaging mechanism learning model contains a loss function of high-level physical semantic features. In order to ensure that the geometric position of the scattering center remains consistent in the registration task, a geometric constraint loss L is introduced. geo : in, is the geometric location of the scattering center predicted by the network, is the geometry of the actual scattering center; Total scattering loss L scat for: L scat =α1L dis +α2L geo Among them, α1 and α2 are weight parameters.

8. The method of multimodal image iterative registration based on physical drive according to claim 1, characterized in that: In step (4), the loss function is gradually integrated and adaptively adjusted to apply the matching loss L match To improve the accuracy and stability of feature matching; This process is to divide the loss function into two stages: the first stage includes the contrast loss L contrastive , used to distinguish positive and negative matching pairs, ensuring that positive sample pairs are close and negative sample pairs are far away, further enhancing the separability of the iterative optimization space; in, and Feature pairs from optical and SAR images, y i The label of the matching pair, when y i = 1, it means that the feature pair is a correctly matched point pair. If y i =0, it means non-matching point pair, represents the Euclidean distance between features, and P is the boundary value; The main direction and alignment descriptors of RIFT2 are used as reference signals to guide the feature alignment and direction consistency of the matching module, thereby improving the rotation robustness of the matching, corresponding to the loss L rotation_match and L descriptor_match ; Among them, θ DL represents the direction information generated by the network during the matching process, θ RIFT represents the main direction of rotation extracted by RIFT2, represents the predicted matching point of the i-th iteration, f DL represents the feature vector generated by the network during the pair matching process, f RIFT represents the direction-aligned descriptor extracted by RIFT2; The loss f generated by RIFT2 supervision RIFT include: L RIFT =L rotation_match +L descriptor_match In the iterative optimization part, a loss function is needed to guide the optimization process in a better direction. The iterative formula is as follows: Among them, the gradient represents the gradient of the loss at the matching point, γ represents the learning rate, and controls the step size of each iterative optimization update; therefore, the matching loss L is obtained match : L match =L contrastive +α3L RIFT Among them, α3 is the weight parameter; The entire loss function L total for: L total =L scat +L contrastive +α3L RIFT The registration results are continuously optimized through the loss function, the transformation matrix corresponding to the optical image and the SAR image is estimated, and finally the final registration results are displayed through connecting lines and checkerboard display.