Remote sensing image registration method, system and equipment based on enhanced feature fusion

By building an image registration network based on U-Net network and enhanced feature fusion module, combined with multiple loss functions, the problem of low accuracy in multimodal remote sensing image registration is solved, and high-precision registration of optical images and SAR images is achieved, improving the overall effect of feature fusion and image registration.

CN120339351AActive Publication Date: 2025-07-18CENT SOUTH UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510814052.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing remote sensing image registration methods have low matching accuracy between multimodal images, especially between optical images and SAR images. This is mainly due to the inability of traditional feature extraction methods to effectively capture the complex feature relationships across modal images, insufficient feature fusion capabilities, separation of spatial transformation and registration, and a single loss function design, making it difficult to adapt to the structural differences between cross-modal images.

Method used

The image registration network model based on U-Net network and enhanced feature fusion module is adopted, and the image registration network and spatial transformation network are generated by building a deformation field, combining mutual information, structural similarity, deformation regularization and structured distribution correlation loss functions to achieve cross-layer feature fusion and joint optimization, and improve image registration accuracy.

Benefits of technology

It improves the accuracy and robustness of multimodal remote sensing image registration, can better handle nonlinear radiation differences between optical images and SAR images, enhances feature fusion capabilities, and ensures high structural consistency and smoothness of image registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339351A_ABST
    Figure CN120339351A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image registration method, system and equipment based on enhanced feature fusion. The method comprises the following steps: constructing an image registration network model comprising a deformation field generation network and a spatial transformation network; inputting the optical remote sensing image and a synthetic aperture radar image to be registered into a deformation field generation network to obtain a deformation field; inputting the synthetic aperture radar image to be registered and the deformation field into the spatial variation network to obtain a pre-registered image; constructing a target loss function based on the deformation field, the to-be-registered synthetic aperture radar image, the pre-registered image and the optical remote sensing image; training the image registration network model according to the target loss function until the target loss function converges, and obtaining a trained image registration network model; and performing remote sensing image registration on the target optical remote sensing image and the target synthetic aperture radar image to be registered through the trained image registration network model. According to the invention, the accuracy of multi-modal remote sensing image registration can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to a remote sensing image registration method, system, and device based on enhanced feature fusion. Background Art

[0002] Image registration is a key task in computer vision and remote sensing image analysis, aiming to accurately align images acquired at different times, angles, or by different sensors, so as to achieve subsequent image fusion, change detection, and analysis. In the remote sensing scenario, the registration task is particularly complex because images often come from different platforms (such as satellites, drones), different types of sensors (such as SAR and optical), and differences in imaging conditions, resulting in scale changes, geometric distortions, illumination differences, and even essential differences between modalities in the images. Therefore, high-precision image registration technology is crucial for the comprehensive utilization of remote sensing data.

[0003] Existing feature-based methods extract significant features (such as points, lines, and regions) from two images, and then estimate the transformation between the images through matching algorithms (such as brute-force matching, RANSAC). Many representative feature-based methods, such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and Oriented FAST and Rotated BRIEF (ORB), are applicable to unimodal images with rotation and scale differences due to perspective or time changes (such as the matching between optical and optical, SAR and SAR), but are vulnerable to multimodal image matching. This is because they mainly rely on detecting highly repeatable features, and these features are usually affected by the intensity or texture differences between multimodal images. Therefore, their matching performance is very sensitive to feature differences. Due to the huge differences in intensity or texture, the features extracted from multimodal images are usually not repeatable, which greatly affects the matching accuracy.

[0004] Therefore, in existing remote sensing image registration, it is difficult for feature-based methods to extract common features between optical images and SAR images because they have significant non-linear radiation differences, which results in relatively low registration accuracy for multimodal remote sensing images. Summary of the Invention

[0005] This application aims to propose a remote sensing image registration method, system, and device based on enhanced feature fusion, which can improve the accuracy of multimodal remote sensing image registration.

[0006] In a first aspect, an embodiment of this application provides a remote sensing image registration method based on enhanced feature fusion, and the method includes: Obtain an optical remote sensing image for model training and a synthetic aperture radar image to be registered; Construct an image registration network model including a deformation field generation network and a spatial transformation network, where the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; Input the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field; Input the synthetic aperture radar image to be registered and the deformation field into the spatial transformation network to obtain a pre-registered image; Construct an objective loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; Train the image registration network model according to the objective loss function until the objective loss function converges to obtain a trained image registration network model; Perform remote sensing image registration on a target optical remote sensing image and a target synthetic aperture radar image to be registered through the trained image registration network model.

[0007] Compared with the prior art, the first aspect of the present application has the following beneficial effects: In this method, an optical remote sensing image for model training and a synthetic aperture radar image to be registered are obtained; an image registration network model including a deformation field generation network and a spatial transformation network is constructed, where the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; the optical remote sensing image and the synthetic aperture radar image to be registered are input into the deformation field generation network to obtain a deformation field; the synthetic aperture radar image to be registered and the deformation field are input into the spatial transformation network to obtain a pre-registered image; an objective loss function is constructed based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; the image registration network model is trained according to the objective loss function until the objective loss function converges to obtain a trained image registration network model; remote sensing image registration is performed on a target optical remote sensing image and a target synthetic aperture radar image to be registered through the trained image registration network model. In this way, by combining a U-Net network and an enhanced feature fusion module to construct a deformation field generation network, feature fusion can be enhanced, efficient cross-layer feature fusion can be achieved, and the accuracy of remote sensing image registration can be improved; by combining the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image to construct an objective loss function, and training and optimizing the network model through the constructed objective loss function, the registered image can have high structural consistency, while ensuring the smoothness of the deformation field, avoiding overfitting, and further improving the accuracy of multi-modal remote sensing image registration.

[0008] In some embodiments, the step of inputting the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field includes: Stitch the optical remote sensing image and the synthetic aperture radar image to be registered to obtain a stitched image; Use the stitched image as the input of the deformation field generation network. Through the encoder in the U-Net network, perform convolutional feature extraction and downsampling on the stitched image to obtain semantic features at multiple coding levels; Through the decoder in the U-Net network, upsample the semantic features extracted at the last coding level to obtain the current upsampled features; Through the enhanced feature fusion module, perform cross-modal feature fusion on the current upsampled features and the semantic features at the corresponding coding level of the next upsampling to obtain target fusion features; Through the decoder in the U-Net network, upsample the target fusion features, and continue to perform cross-modal feature fusion on the current upsampling result and the semantic features at the corresponding coding level of the next upsampling until the decoder decoding is completed to obtain a deformation field.

[0009] In some embodiments, the enhanced feature fusion module includes a feature enhancement sub-module, a channel attention sub-module, and a spatial attention sub-module. The step of performing cross-modal feature fusion on the current upsampled features and the semantic features at the corresponding coding level of the next upsampling through the enhanced feature fusion module to obtain target fusion features includes: Input the current upsampled features and the semantic features at the corresponding coding level of the next upsampling into the feature enhancement sub-module to obtain a first fusion feature; Perform a residual connection on the first fusion feature and the semantic features at the corresponding coding level of the next upsampling to obtain a second fusion feature; Input the second fusion feature into the channel attention sub-module to obtain a third fusion feature; Input the third fusion feature into the spatial attention sub-module to obtain target fusion features.

[0010] In some embodiments, based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image, construct a target loss function, including: Construct a mutual information loss function by calculating the mutual information between the synthetic aperture radar image to be registered and the optical remote sensing image; Construct a structural similarity loss function by calculating the structural similarity between the pre-registered image and the optical remote sensing image; Construct a deformation regularization loss function according to the deformation field; Construct a structured distribution correlation loss function by calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image. Perform a weighted sum of the mutual information loss function, the structural similarity loss function, the deformation regularization loss function, and the structured distribution correlation loss function to obtain a target loss function.

[0011] In some embodiments, the step of constructing a structured distribution correlation loss function by calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image includes: Construct the local pixels in the synthetic aperture radar image to be registered into a gamma distribution; Construct the local pixels in the optical remote sensing image into a Gaussian distribution; Calculate a first structure mapping value of the synthetic aperture radar image to be registered, and calculate a second structure mapping value of the optical remote sensing image; Calculate the correlation weight of the synthetic aperture radar image to be registered and the optical remote sensing image in the structural domain according to the first structure mapping value and the second structure mapping value; Calculate the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image according to the gamma distribution, the Gaussian distribution, and the correlation weight; Construct a structured distribution correlation loss function based on the structured distribution correlation.

[0012] In some embodiments, the step of calculating a first structure mapping value of the synthetic aperture radar image to be registered and calculating a second structure mapping value of the optical remote sensing image includes: ; where represents the first structure mapping value of the synthetic aperture radar image to be registered, represents the gradient of the synthetic aperture radar image to be registered at the point, represents the second structure mapping value of the optical remote sensing image, represents the gradient of the optical remote sensing image at the point, represents the vector differential operator.

[0013] In some embodiments, the step of calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image according to the gamma distribution, the Gaussian distribution, and the correlation weight includes: ; Among them, represents the structured distribution correlation, represents the sliding window size, represents the correlation weight, represents the probability density estimate of the current window calculated through the gamma distribution, represents the global reference distribution of the synthetic aperture radar image to be registered, represents the probability density estimate of the current window calculated through the Gaussian distribution, represents the global reference distribution of the optical remote sensing image.

[0014] In a second aspect, an embodiment of the present application further provides a remote sensing image registration system based on enhanced feature fusion. The system includes: A training data acquisition unit, configured to acquire an optical remote sensing image for model training and a synthetic aperture radar image to be registered; A network model construction unit, configured to construct an image registration network model including a deformation field generation network and a spatial transformation network, wherein the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; A deformation field obtaining unit, configured to input the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field; An image pre-registration unit, configured to input the synthetic aperture radar image to be registered and the deformation field into the spatial transformation network to obtain a pre-registered image; A loss function construction unit, configured to construct a target loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; A network model training unit, configured to train the image registration network model according to the target loss function until the target loss function converges to obtain a trained image registration network model; A target image registration unit, configured to perform remote sensing image registration on a target optical remote sensing image and a target synthetic aperture radar image to be registered through the trained image registration network model.

[0015] In a third aspect, an embodiment of the present application further provides an electronic device, including at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute a remote sensing image registration method based on enhanced feature fusion as described above.

[0016] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a remote sensing image registration method based on enhanced feature fusion as described above.

[0017] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For relevant descriptions, reference can be made to the relevant descriptions in the above first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where: Figure 1 is a schematic flowchart of an embodiment of a remote sensing image registration method based on enhanced feature fusion provided by the present application; Figure 2 is a schematic structural diagram of an image registration network model in the best embodiment of a remote sensing image registration method based on enhanced feature fusion provided by the present application; Figure 3 is a schematic structural diagram of a deformation field generation network in the best embodiment of a remote sensing image registration method based on enhanced feature fusion provided by the present application; Figure 4 is a schematic structural diagram of an enhanced feature fusion module in the best embodiment of a remote sensing image registration method based on enhanced feature fusion provided by the present application; Figure 5 is a schematic structural diagram of an embodiment of a remote sensing image registration system based on enhanced feature fusion provided by the present application; Figure 6 is a schematic structural diagram of an embodiment of an electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present application and should not be construed as limiting the present application.

[0020] In the description of the present application, if the first, second, etc. are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0021] In the description of this application, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, etc., it is based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application.

[0022] In the description of this application, it should be noted that unless otherwise clearly defined, terms such as setting, installation, connection, etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above terms in this application in combination with the specific content of the technical solution.

[0023] First, several nouns involved in this application are analyzed: Remote sensing image registration: The process of precisely aligning remote sensing images taken by different sensors, at different times, and from different perspectives into a unified spatial coordinate system, which is convenient for subsequent fusion and comparative analysis.

[0024] Enhanced Feature Fusion Module EFF (Enhanced Feature Fusion): A deep network module specifically used for cross-level and cross-modal feature fusion, which contains multiple sub-modules (EAG, SA, and ECA), and realizes efficient fusion through enhanced feature representation and attention mechanism.

[0025] Enhanced Attention Group Module EAG (Enhanced Attention Group): A sub-module that uses grouped convolution and residual connection to enhance the feature expression ability and effectively fuse cross-modal features.

[0026] Efficient Channel Attention Sub-module ECA (Efficient Channel Attention) Module: A sub-module that calculates channel attention weights through one-dimensional convolution, which is used to strengthen important feature channels, suppress interference from irrelevant information, and improve the network's attention to different channel features.

[0027] Spatial Attention Sub-module SA (Spatial Attention): Captures key regions in the spatial position through global average pooling and global maximum pooling, enhances the network's attention to important regions in the spatial position, and thus improves the effectiveness of feature expression.

[0028] Deformation Field: A two-dimensional vector field that describes the pixel-level correspondence between the image to be registered and the reference image. The vector at each pixel position represents the displacement that the corresponding pixel needs to undergo in space.

[0029] Existing feature-based methods extract significant features (e.g., points, lines, and regions) from two images and then estimate the transformation between the images through matching algorithms such as brute-force matching and RANSAC. Many representative feature-based methods, such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and Oriented FAST and Rotated BRIEF (ORB), are applicable to unimodal images with rotational and scale differences due to perspective or time changes (such as matching between optical-optical, SAR-SAR), but are vulnerable to multimodal image matching. This is because they mainly rely on detecting highly repeatable features, which are usually affected by intensity or texture differences between multimodal images. Therefore, their matching performance is very sensitive to feature differences. Due to the large differences in intensity or texture, the features extracted from multimodal images are usually not repeatable, which greatly affects the matching accuracy.

[0030] To solve the problem of relatively low registration accuracy of multimodal remote sensing images in the existing technology, this application proposes a remote sensing image registration method, system, and device based on enhanced feature fusion.

[0031] Refer to Figure 1 , the flowchart of the remote sensing image registration method based on enhanced feature fusion provided by the embodiments of this application. The remote sensing image registration method based on enhanced feature fusion is applied to an electronic device, which can be a server or a mobile terminal, etc. As Figure 1 shown, the remote sensing image registration method based on enhanced feature fusion may include the following steps: Step S100, obtain an optical remote sensing image for model training and a synthetic aperture radar image to be registered; Step S200, construct an image registration network model including a deformation field generation network and a spatial transformation network, where the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; Step S300, input the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field; Step S400, input the synthetic aperture radar image to be registered and the deformation field into the spatial transformation network to obtain a pre-registered image; Step S500, construct an objective loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; Step S600, train the image registration network model according to the objective loss function until the objective loss function converges to obtain a trained image registration network model; Step S700, perform remote sensing image registration on the target optical remote sensing image and the target synthetic aperture radar image to be registered through the trained image registration network model.

[0032] In this embodiment, an optical remote sensing image for model training and a synthetic aperture radar image to be registered are obtained; an image registration network model including a deformation field generation network and a spatial transformation network is constructed, wherein the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; the optical remote sensing image and the synthetic aperture radar image to be registered are input into the deformation field generation network to obtain a deformation field; the synthetic aperture radar image to be registered and the deformation field are input into the spatial transformation network to obtain a pre-registered image; a target loss function is constructed based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; the image registration network model is trained according to the target loss function until the target loss function converges to obtain a trained image registration network model; the trained image registration network model is used to perform remote sensing image registration on a target optical remote sensing image and a target synthetic aperture radar image to be registered. In this way, by combining the U-Net network and the enhanced feature fusion module to construct the deformation field generation network, feature fusion can be enhanced, efficient cross-layer feature fusion can be achieved, and the accuracy of remote sensing image registration can be improved; by combining the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image to construct the target loss function, and training and optimizing the network model through the constructed target loss function, the registered image can have high structural consistency, while ensuring the smoothness of the deformation field, avoiding overfitting, and further improving the accuracy of multi-modal remote sensing image registration.

[0033] The above enhanced feature fusion module can be used in the feature fusion process between the encoder and the decoder in the U-Net network. By strengthening the feature fusion between different levels and modalities, the low-level texture information and high-level semantic information in the image can be better combined.

[0034] The above constructing the target loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image can be to construct loss functions between each pair of the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image to obtain multiple loss functions, and then combine the multiple loss functions to construct the target loss function.

[0035] In some embodiments, inputting the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field includes: The optical remote sensing image and the synthetic aperture radar image to be registered are spliced to obtain a spliced image; The spliced image is used as the input of the deformation field generation network, and the spliced image is subjected to convolutional feature extraction and downsampling through the encoder in the U-Net network to obtain semantic features at multiple coding levels; The semantic features extracted from the last encoding level are upsampled by the decoder in the U-Net network to obtain the current upsampled features; The current upsampled features and the semantic features of the corresponding encoding level of the next upsampling are subjected to cross-modal feature fusion through the enhanced feature fusion module to obtain the target fusion features; The target fusion features are upsampled by the decoder in the U-Net network, and the current upsampling result is continuously subjected to cross-modal feature fusion with the semantic features of the corresponding encoding level of the next upsampling until the decoder finishes decoding to obtain the deformation field.

[0036] In this embodiment, by splicing the optical remote sensing image and the synthetic aperture radar image to be registered, the spliced image is obtained; the spliced image is used as the input of the deformation field generation network, and the spliced image is subjected to convolutional feature extraction and downsampling through the encoder in the U-Net network to obtain semantic features of multiple encoding levels; the semantic features extracted from the last encoding level are upsampled by the decoder in the U-Net network to obtain the current upsampled features; the current upsampled features and the semantic features of the corresponding encoding level of the next upsampling are subjected to cross-modal feature fusion through the enhanced feature fusion module to obtain the target fusion features; the target fusion features are upsampled by the decoder in the U-Net network, and the current upsampling result is continuously subjected to cross-modal feature fusion with the semantic features of the corresponding encoding level of the next upsampling until the decoder finishes decoding to obtain the deformation field. Thus, by using the symmetric structure of the encoder and decoder in the U-Net network for image feature extraction and reconstruction, and then through the enhanced feature fusion module to perform cross-modal feature fusion on the current upsampled features and the semantic features of the corresponding encoding level of the next upsampling, feature fusion can be enhanced, efficient cross-layer feature fusion can be achieved, and thus the accuracy of remote sensing image registration can be improved.

[0037] In some embodiments, the enhanced feature fusion module includes a feature enhancement sub-module, a channel attention sub-module, and a spatial attention sub-module. Subjecting the current upsampled features and the semantic features of the corresponding encoding level of the next upsampling to cross-modal feature fusion through the enhanced feature fusion module to obtain the target fusion features includes: Inputting the current upsampled features and the semantic features of the corresponding encoding level of the next upsampling into the feature enhancement sub-module to obtain the first fusion features; Performing residual connection on the first fusion features and the semantic features of the corresponding encoding level of the next upsampling to obtain the second fusion features; Inputting the second fusion features into the channel attention sub-module to obtain the third fusion features; Inputting the third fusion features into the spatial attention sub-module to obtain the target fusion features.

[0038] In this embodiment, by inputting the current upsampled feature and the semantic feature of the corresponding encoding layer of the next upsampling into the feature enhancement sub-module, a first fused feature is obtained; the first fused feature and the semantic feature of the corresponding encoding layer of the next upsampling are subjected to residual connection to obtain a second fused feature; the second fused feature is input into the channel attention sub-module to obtain a third fused feature; the third fused feature is input into the spatial attention sub-module to obtain the target fused feature. In this way, preliminary feature extraction and cross-layer fusion are realized through the feature enhancement sub-module; then, through the channel attention calculation of the channel attention sub-module, weights are assigned to each channel to highlight important feature channels; subsequently, it enters the spatial attention sub-module, where the feature map is weighted in the spatial dimension to focus on key regions and suppress background interference, thereby realizing efficient cross-layer feature fusion.

[0039] In some embodiments, based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image, a target loss function is constructed, including: By calculating the mutual information between the synthetic aperture radar image to be registered and the optical remote sensing image, a mutual information loss function is constructed; By calculating the structural similarity between the pre-registered image and the optical remote sensing image, a structural similarity loss function is constructed; According to the deformation field, a deformation regularization loss function is constructed; By calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image, a structured distribution correlation loss function is constructed; The mutual information loss function, the structural similarity loss function, the deformation regularization loss function, and the structured distribution correlation loss function are weighted and summed to obtain the target loss function.

[0040] In this embodiment, by calculating the mutual information between the synthetic aperture radar image to be registered and the optical remote sensing image, a mutual information loss function is constructed; by calculating the structural similarity between the pre-registered image and the optical remote sensing image, a structural similarity loss function is constructed; according to the deformation field, a deformation regularization loss function is constructed; by calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image, a structured distribution correlation loss function is constructed; the mutual information loss function, the structural similarity loss function, the deformation regularization loss function, and the structured distribution correlation loss function are weighted and summed to obtain the target loss function. In this way, by combining the mutual information loss function, the structural similarity loss function, the deformation regularization loss function, and the structured distribution correlation loss function for weighted summation, the target loss function is constructed, making the target loss function more flexible and refined, suitable for multi-modal remote sensing image registration, and enabling the image registration network model trained by the target loss function to perform remote sensing image registration more accurately.

[0041] In some embodiments, by calculating the structural distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image, a structural distribution correlation loss function is constructed, including: Construct the local pixels in the synthetic aperture radar image to be registered into a gamma distribution; Construct the local pixels in the optical remote sensing image into a Gaussian distribution; Calculate the first structure mapping value of the synthetic aperture radar image to be registered, and calculate the second structure mapping value of the optical remote sensing image; According to the first structure mapping value and the second structure mapping value, calculate the correlation weight of the synthetic aperture radar image to be registered and the optical remote sensing image in the structural domain; According to the gamma distribution, the Gaussian distribution and the correlation weight, calculate the structural distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image; Based on the structural distribution correlation, construct a structural distribution correlation loss function.

[0042] In this embodiment, by constructing the local pixels in the synthetic aperture radar image to be registered into a gamma distribution; constructing the local pixels in the optical remote sensing image into a Gaussian distribution; calculating the first structure mapping value of the synthetic aperture radar image to be registered, and calculating the second structure mapping value of the optical remote sensing image; according to the first structure mapping value and the second structure mapping value, calculate the correlation weight of the synthetic aperture radar image to be registered and the optical remote sensing image in the structural domain; according to the gamma distribution, the Gaussian distribution and the correlation weight, calculate the structural distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image; based on the structural distribution correlation, construct a structural distribution correlation loss function. In this way, through the constructed structural distribution correlation loss function, the structural consistency registration effect between the synthetic aperture radar image to be registered and the optical remote sensing image can be improved, and the problem of insufficient modeling of the image structure distribution difference by the traditional loss function when processing multi-modal remote sensing images can be solved.

[0043] In some embodiments, calculating the first structure mapping value of the synthetic aperture radar image to be registered, and calculating the second structure mapping value of the optical remote sensing image, includes: ; Wherein, represents the first structure mapping value of the synthetic aperture radar image to be registered, represents the synthetic aperture radar image to be registered at the gradient at point, represents the second structure mapping value of the optical remote sensing image, represents the optical remote sensing image at the gradient at point, represents a vector differential operator.

[0044] In some embodiments, calculating the structural distribution correlation between a synthetic aperture radar image and an optical remote sensing image to be registered according to a gamma distribution, a Gaussian distribution, and a correlation weight includes: ; wherein, represents the structural distribution correlation, represents the sliding window size, represents the correlation weight, represents the probability density estimate of the current window calculated through the gamma distribution, represents the global reference distribution of the synthetic aperture radar image to be registered, represents the probability density estimate of the current window calculated through the Gaussian distribution, represents the global reference distribution of the optical remote sensing image.

[0045] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments: Image registration is a key task in computer vision and remote sensing image analysis, aiming to accurately align images acquired at different times, angles, or by different sensors, so as to achieve subsequent image fusion, change detection, and analysis. In the remote sensing scenario, the registration task is particularly complex because images often come from different platforms (such as satellites, drones), different types of sensors (such as SAR and optical), and differences in imaging conditions, resulting in scale changes, geometric distortions, illumination differences, and even essential differences between modalities in the images. Therefore, high-precision image registration technology is crucial for the comprehensive utilization of remote sensing data.

[0046] Feature-based methods extract significant features (such as points, lines, and regions) from two images, and then estimate the transformation between the images through matching algorithms (such as brute-force matching, RANSAC). Many representative feature-based methods, such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and Oriented FAST and Rotated BRIEF (ORB), are applicable to unimodal images with rotation and scale differences due to perspective or time changes (such as the matching between optical and optical, SAR and SAR), but are vulnerable to multimodal image matching. This is because they mainly rely on detecting highly repeatable features, which are usually affected by the intensity or texture differences between multimodal images. Therefore, their matching performance is very sensitive to feature differences. Due to the large differences in intensity or texture, the features extracted from multimodal images are usually not repeatable, which greatly affects the matching accuracy. Therefore, in remote sensing image registration, it is difficult for feature-based methods to extract common features between optical images and SAR images because they have significant non-linear radiation differences.

[0047] SAR images are the result of radar wave reflection, mainly reflecting the backscattering characteristics of ground objects, and there are problems of strong noise (such as speckle noise) and lack of structural information. On the other hand, optical images are taken with visible light, have clear textures and rich color information, and are suitable for visual analysis. Due to the different modalities of these two types of images, traditional registration methods (such as SIFT) perform poorly in feature matching, resulting in low image registration accuracy. For existing deep learning-based methods, there is also a lack of cross-modal feature extraction, which further leads to low image registration accuracy.

[0048] Therefore, the existing remote sensing image registration methods have the following disadvantages: (1) Poor adaptability of traditional feature extraction methods in cross-modal registration.

[0049] Existing technologies usually rely on traditional feature extraction methods (such as SIFT, SURF, and ORB, etc.) to align images of different modalities. These methods are based on hand-designed feature extraction methods to extract local information (such as edges, corners, etc.) in images. However, SAR images (i.e., synthetic aperture radar images) and optical images (i.e., optical remote sensing images) have significant differences in texture, contrast, and structure. Traditional feature extraction methods often cannot fully capture the complex feature relationships of cross-modal images, resulting in low registration accuracy. The significant differences between the noise and blurriness of SAR images and the high resolution and clear textures of optical images make it difficult for handcrafted features to provide effective cross-modal matching.

[0050] (2) Poor feature fusion ability, that is, the feature extraction process lacks an effective multi-scale and cross-layer information fusion mechanism.

[0051] Many existing registration methods use simple convolutional neural networks (CNNs) for feature extraction, but lack an effective cross-layer feature fusion mechanism. In the image registration task, low-level features (such as edges, textures) and high-level features (such as objects, semantic information) have different expressions. Although existing feature extraction modules can handle certain context information, they are not sufficiently optimized for complex image registration tasks. Especially when dealing with cross-modal images, they may not be able to fully learn higher-level semantic information. As a result, it is impossible to effectively combine low-level and high-level features in the image, resulting in the failure to fully utilize important cross-modal information in the image registration process, affecting the accuracy of remote sensing image registration. Especially in the case of large image changes, the accuracy is even lower.

[0052] (3) Spatial transformation and registration are separated, that is, the image registration and spatial transformation modules are independent of each other and lack a joint optimization mechanism.

[0053] In the prior art, the image registration process is usually decomposed into several stages, such as feature extraction, transformation estimation, and transformation application. In this process, the transformation estimation (such as affine or perspective transformation) and the optimization of local deformation are usually carried out separately. Since the geometric deformation in the image is not only a global transformation, in many cases, non-rigid deformation may occur in the local area. Existing methods rely on the separate optimization of global transformation and local deformation, and cannot effectively combine the two for joint optimization. Separately optimizing the global transformation and local deformation will lead to a lack of effective synergy between these two stages in the registration process. There may be inconsistent global and local transformations in the image registration process, resulting in an unsatisfactory image registration effect, especially for remote sensing images with large terrain changes or the presence of occlusions.

[0054] (4) The loss function design is simple, that is, the loss function design is single, and it is difficult to adapt to the cross-modal image structure differences.

[0055] Many existing image registration methods adopt pixel-level error metrics (such as L2 loss) or traditional gradient-based loss functions. These loss functions usually only focus on the overall differences of the images and cannot fully capture the subtle structural differences (such as texture, edges, etc.) between multi-modal images. For cross-modal remote sensing images, especially between SAR images and optical images, the texture and structure may be very different. Relying only on pixel-level error metrics (such as L2 loss) cannot accurately reflect the structural differences between images, resulting in a reduction in the quality of the registration results. The structural information in the images cannot be accurately aligned during the registration process, resulting in low registration accuracy, especially in complex environments or under large lighting changes.

[0056] To solve the above-mentioned drawbacks of the prior art, referring to Figure 2 , the image registration network model of this embodiment effectively solves the problems of insufficient cross-modal feature extraction ability, insufficient feature fusion mechanism, separation of spatial transformation and registration tasks, and single loss function design in the existing remote sensing image registration technology by organically combining the U-Net structure with an efficient feature fusion module (EFF) and introducing a spatial transformation network (STN), thereby achieving the effect of significantly improving the registration accuracy and robustness between SAR images and optical images. The technical solution of this embodiment specifically includes the following contents: Step S1: Obtain a training image dataset containing multiple SAR images and optical images (i.e., optical remote sensing images) to be registered. Input the SAR images and optical images to be registered for model training into the registration network (i.e., the deformation field generation network), and obtain the deformation field through the registration network. The deformation field represents the pixel-level position correspondence between the image to be registered and the reference image. Referring to Figure 3 , the specific processing process of the registration network is as follows: The registration network of this embodiment adopts an encoder-decoder structure based on U-Net, and integrates an enhanced feature fusion module (EFF) to generate a high-precision deformation field. The encoder performs convolution feature extraction and downsampling on the input SAR image and optical image (joined in the channel dimension to form a multi-channel input) layer by layer to capture multi-scale low-level texture and high-level semantic features; then the decoder gradually upsamples and fuses the features of the corresponding layer in the encoding stage with the current upsampled features through jump connections. The EFF module introduced in the fusion process contains sub-modules such as EAG, SA, and ECA, which are used to enhance feature representation, focus on key spatial areas, and adjust channel weights, respectively, so as to fully combine feature information of different modalities and different levels. After encoding-decoding and multi-level feature fusion, the registration network finally outputs a two-dimensional deformation field of the same size as the input image through convolution (the two channels represent the displacement of pixels in the x and y directions, respectively), which is used to characterize the corresponding relationship between the SAR image to be registered and the reference optical image in pixel coordinates. The deformation field directly describes the distortion transformation required for the spatial alignment of the two images.

[0057] Specifically, the first step is the encoding stage. The input SAR image and the optical image are spliced as the input of the registration network. The input size is , by performing convolutional layers and activation functions to extract the features of the input image, the size is Next, the obtained feature map is pooled to retain the number of channels of the feature map. , reducing the image size to . Then, the convolutional layer is used to extract deeper semantic information, increasing the number of channels to , and the feature map size is The process continues: the maximum pooling operation is performed again, and the size is further reduced to , then convolution extracts features and expands the channel to , the output is Continuing to encode downward, after entering the deepest feature extraction stage, we get a feature map of size And send it to the EFF (Efficient Feature Fusion) module for multi-scale context fusion and feature compression.

[0058] The next step is the decoding stage. Through upsampling, the deep feature map is transformed from Upsample to , and the corresponding feature map in the encoding path At the same time, it is used as the input of the EFF module for feature fusion to obtain the fused features , and extract features through the convolutional layer and expand the channels to . The decoding path continues upward and is restored to a feature map with a size of (the same size as the other end of the encoder) through the second upsampling, and fuse the feature map in the encoding stage. The feature map passes through the convolutional EFF module to obtain a new fused feature. This process is repeated and gradually upsampled to At each level, the current feature map is concatenated with the skip connection feature in the encoding path, and fused and enhanced through the EFF module. Finally, the number of channels is compressed from to 2 through the convolutional layer, and a deformation field with a size of is output, which is used to estimate the corresponding offset at each pixel point between the SAR image and the optical image, realizing high-precision image registration.

[0059] Among them, the enhanced feature fusion (EFF) module is a key component in the registration network of this embodiment, which is used to improve the effect of cross-modal feature extraction and fusion during the feature fusion process of the encoder and decoder. As Figure 4 shown, the EFF module receives feature maps from different levels (on the one hand, the low-level texture features transmitted by the encoder, and on the other hand, the high-level semantic features transmitted by the decoder), and processes the features through its three internal sub-modules in sequence, and outputs the fused and enhanced feature map. First, the input features are processed by the grouped convolution of the EAG sub-module (feature enhancement sub-module) to achieve preliminary feature extraction and cross-layer fusion; then, through the channel attention calculation of the ECA sub-module (channel attention sub-module), weights are assigned to each channel to highlight important feature channels; subsequently, it enters the SA sub-module (spatial attention sub-module), and the feature map is weighted in the spatial dimension to focus on the key areas and suppress background interference. After the above three-stage processing, the features output by the EFF module fuse multi-scale and multi-modal information, retain both low-level detail textures and combine high-level semantic structures, thus providing a rich and consistent feature representation for the subsequent registration deformation field estimation.

[0060] The EAG submodule mainly realizes the enhanced fusion of cross-modal features through convolution operations and residual connections (i.e., represented by C). The EAG submodule uses grouped convolution (i.e., Group Conv(1x1)) operations on the input features to extract features in parallel after grouping the channels. This structure reduces the convolution parameters and the amount of calculation while enabling the convolution filter to specifically learn the features of different groups, thereby improving the ability to adapt to cross-modal feature differences. The features extracted by grouped convolution are batch normalized (BN) and nonlinearly activated (ReLU) to capture the local combination of feature patterns in the input. Subsequently, the parallel extracted features after batch normalization (BN) and nonlinear activation (ReLU) are added, and then the features extracted by convolution are obtained after nonlinear activation (ReLU), convolution (Conv(1x1)), and sigmoid activation function. Then, the EAG submodule uses residual connections to directly add and fuse the input features with the features extracted by convolution. Through this residual path, the original low-level detail information is directly integrated into the high-level abstract features, avoiding the loss of key information in the feature extraction process and enhancing the expressiveness of the features. The output of the EAG submodule is a feature map that combines the original input and the newly extracted information of the convolution. It contains both the enhanced cross-modal common features and the original important details, providing a good feature foundation for the subsequent attention module.

[0061] The ECA (channel attention) submodule is responsible for recalibrating the importance weights of each channel of the feature map in the channel dimension. The ECA submodule first processes the output feature map from EAG globally, typically using a global average pooling operation or an equivalent 1D convolution to compress the entire feature map in the spatial dimension and generate a global description for each channel. Next, ECA filters the channel description vector in the channel dimension through a 1D convolution. This 1D convolution operation can be regarded as an efficient way to achieve information interaction between channels and learn the correlation between different channels. The output of the convolution is converted into channel attention weights through an activation function, and each channel corresponds to a weight coefficient between 0 and 1. Subsequently, these weights are multiplied one by one with each channel of the original feature map to complete the remodulation of the feature map channels. Functional effect: Through the ECA submodule, the network can automatically highlight important feature channels and suppress irrelevant or interfering channel information in the registration. For example, in the SAR image and optical image registration scenario, some channels may correspond to the edge or structural features common to the two images, while other channels may mainly reflect the noise or artifacts of a single mode; ECA will give higher weights to the former and lower weights to the latter, so that the feature map after ECA is more focused on the key information required for alignment. The output result is a channel-weighted feature map that retains the spatial size of the original feature map, but has been reordered according to the importance of the channel strength, highlighting the signal that is most beneficial to the registration.

[0062] The SA (Spatial Attention) sub-module is used to generate attention weights for the feature map in the two-dimensional spatial dimension, enabling the network to focus on discriminative regions in the image and ignore irrelevant regions. The way the SA sub-module extracts spatial importance clues from the input feature map is as follows: globally average pooling (GAP) and globally max pooling (GMP) are applied in parallel to compress the feature map in the channel dimension. After these two pooling operations, we obtain two feature maps with only the spatial dimension (width × height): the average pooling map reflects the average activation degree of each channel at each spatial position, and the max pooling map reflects the strongest response at each position in some channels. After superimposing these two maps pixel by pixel in the spatial dimension, they are fed into a convolutional layer and a sigmoid activation function to generate a spatial attention weight map with the same size as the original feature map. Each pixel value (between 0 and 1) of this weight map represents the importance degree of the corresponding spatial position. Finally, the SA sub-module multiplies the weight map element-wise with the original input feature map to achieve spatial weighting of the input feature map. After this operation, regions in the feature map containing key ground objects will be given higher responses, while the responses of those background regions with more noise or irrelevant to registration will be reduced. Functional effect: The SA sub-module enables the network to adaptively focus on common and prominent target regions in cross-modal images, improves the cross-modal feature alignment ability, and at the same time reduces interference caused by background differences. The output result is a spatially weighted feature map, which has the same size and number of channels as the input, but the key regions are highlighted and the secondary regions are suppressed in the spatial response distribution.

[0063] Step S2: Input the SAR image to be registered and the deformation field into the spatial transformation network to obtain the registered image (i.e., the pre-registered image). Specifically, After obtaining the deformation field between the image to be registered (SAR image) and the reference image (optical image), the deformation field and the image to be registered are used as inputs and fed into the spatial transformation network (STN). The spatial transformation network usually consists of three parts: a localization network, a grid generator, and a sampler. In this embodiment, the registration network (the content described earlier) simultaneously serves as the localization network and the grid generator to generate the deformation field and clarify the coordinate mapping relationship of each output pixel in the original SAR image. Then, the sampler performs an interpolation transformation on the image to be registered according to this deformation field to obtain the final registration result. Specifically, As a differentiable image transformation module, STN applies the deformation field predicted by the previous registration network to the image to be registered, achieving spatial distortion alignment of the image. Specifically, STN consists of three parts: a localization network, a grid generator, and a sampler. The localization network predicts the corresponding spatial transformation parameters based on the input deformation field. The grid generator then generates a sampling coordinate grid accordingly, that is, the mapping of the corresponding positions of each pixel in the output image in the original SAR image. The sampler resamples the SAR image at these irregular coordinates using algorithms such as bilinear interpolation. Through the STN module, complex geometric corrections such as translation, rotation, scaling, and even projection transformation can be performed on the SAR image, and its pixels can be accurately mapped to the coordinate system of the reference optical image. Since each part of STN is trainable, in this embodiment, it is combined end-to-end with the U-Net backbone network to generate an alignment result in the forward propagation and jointly optimize the parameters based on the loss function in the backward propagation, thereby realizing the joint optimization of the image registration process and greatly improving the accuracy and robustness of registration.

[0064] Step S3: Calculate the loss between the obtained pre-registered image and the optical image to obtain a loss value, and update the network parameters through backpropagation. And continuously iterate the training until the loss function value converges to a stable level, and finally obtain a trained SAR and optical remote sensing image registration model (i.e., an image registration network model). Use the trained SAR and optical remote sensing image registration model to perform remote sensing image registration on the target optical remote sensing image and the target synthetic aperture radar image to be registered. The target optical remote sensing image and the target synthetic aperture radar image to be registered are the image pair that the user wants to register. Specifically, Construct the target loss function as: ; where, is the registered image (i.e., the SAR image after deformation field and STN transformation), is the reference image (i.e., the optical image). is the deformation field generated by the backbone part (U-net network) of the network. 、 、 、 are the weight coefficients of each loss function, used to adjust the contribution of different loss functions. The weight coefficient of each loss is adjusted accordingly through experimental results.

[0065] To maximize the statistical correlation between the SAR image and the optical image, especially in cross-modal image registration, mutual information can well measure the similarity between images, especially when there are illumination changes or texture differences in the images. The mutual information loss function is as follows: ; Among them, is the joint probability distribution of the SAR image and the optical image. and are the marginal probability distributions of the SAR image and the optical image respectively.

[0066] To optimize the local structural consistency of the image and ensure the similarity of the registered image and the reference image in local features such as structure, texture, and edges. The structural similarity loss function is as follows: ; Among them, measures the similarity of the brightness, contrast, and structure of the image. The result range is [0, 1], and the value closer to 1 indicates that the images are more similar.

[0067] To maintain the smoothness of the deformation field, avoid unnatural deformations, and ensure the geometric consistency of the registered image. The deformation regularization loss function is as follows: ; To further improve the registration effect of the structural consistency between the SAR image and the optical image and solve the problem of insufficient modeling of the structural distribution differences of the traditional loss function when dealing with multimodal remote sensing images, this embodiment innovatively designs a structured distribution correlation loss function (Structured Distribution Correlation Loss, abbreviated as SDC). This loss function combines local statistical modeling and a structure-sensitive weighting mechanism, and can significantly improve the structural alignment effect in the detail areas of the registration results, especially in the case where the SAR image has speckle noise and the optical image is affected by light interference.

[0068] On the one hand, the SAR image has multiplicative speckle noise, the texture structure is blurred, and the gray distribution usually follows a Gamma distribution; in addition, the optical image is sensitive to light, has sharp edges, and is rich in texture, and its gray distribution can be approximated as a Gaussian distribution or a Gaussian mixture model; traditional losses such as pixel-level L2, SSIM, or mutual information MI are all difficult to fully capture the structural similarities and differences and the statistical feature correspondences between cross-modalities. Therefore, to solve the above problems, this embodiment proposes a new loss function that combines distribution difference modeling and structure domain correlation quantification to evaluate the structured statistical consistency between the registration result and the reference image.

[0069] In this embodiment, a window with a fixed size is slid in the image , and within this local window, the following probability distribution models are respectively established for the SAR image and the optical image.

[0070] SAR image: The local pixel values follow a Gamma distribution (i.e., gamma distribution): ; where and are the shape and scale parameters of the Gamma distribution respectively, and can be calculated from local pixels through maximum likelihood estimation.

[0071] Optical image: The local pixel values follow a Gaussian distribution (i.e., Gaussian distribution): ; where is the local mean, is the local variance, which can also be estimated from local pixels.

[0072] To establish cross-modal correspondence in the structural domain, a structure mapping function is introduced, and its form is the local gradient magnitude of the image: ; where represents the structure mapping value of the SAR image, represents the gradient of the SAR image at point X, represents the structure mapping value of the optical image, represents the gradient of the optical image at point Y.

[0073] This mapping projects the features in the gray-scale domain to the structural space, enhances the response of edges and textures, and has stronger adaptability to modal differences.

[0074] where , The calculation process represents the calculation process of the above formula This The essence of the formula is the calculation of the gradient. The structure mapping function is constructed in the form of the gradient magnitude. By calculating the rate of change of the image in the horizontal and vertical directions, its local structural strength is extracted to measure the correspondence of multimodal images in the structural domain.

[0075] In the structural space, the correspondence degree between the SAR pixel value x and the optical pixel value y is controlled by the following weight function: ; where is a hyperparameter that controls the structural sensitivity (usually set to 2-3 times the standard deviation of the gradient). The design inspiration of this function comes from the Gaussian kernel similarity, which can assign higher fusion weights to regions with similar structures and suppress the interference of irrelevant regions.

[0076] Based on the above modeling, the present embodiment proposes the following structured distribution correlation (SDC) metric function: ; where and are the probability density estimates of the current windows of the SAR image and the optical image respectively, and the probability density estimates are calculated according to the Gamma distribution and the Gaussian distribution; and are the global reference distributions of the SAR image and the optical image respectively, and the global reference distributions are also calculated according to the Gamma distribution and the Gaussian distribution; is the correlation weight on the domain; the numerator is the weighted covariance of the distribution difference, and the denominator is the normalization factor to ensure that the output range of this function is [-1, 1], similar to the correlation coefficient. When the SDC value is closer to 1, it indicates that the structural distributions of the two images in this local area are more consistent, and vice versa. Therefore, the structured distribution correlation loss function is as follows: ; The specific registration steps of the SAR and optical images in this embodiment are as follows: Step 1: Data preprocessing. Obtain a training image dataset containing multiple SAR images and optical images (i.e., optical remote sensing images) to be registered. Preprocess the image pairs composed of the SAR image and the optical image to be registered, as well as the image pairs composed of the target optical remote sensing image and the target synthetic aperture radar image to be registered. The preprocessing includes image normalization, size resampling, denoising, and data augmentation operations to ensure that the spatial dimensions, resolutions, and number of channels of the SAR image and the optical image match each other to meet the consistency of the input dimensions of the neural network.

[0077] Step 2: Model initialization. Construct a deformation field generation network based on the U-net structure, and initialize the hyperparameters of the model, including the learning rate, Batch Size, weight initialization method, etc. At the same time, set the Adam optimizer and configure the optimizer parameters.

[0078] Step 3: Forward propagation and image registration. Concatenate the preprocessed SAR image and optical image in the channel dimension to form multi-channel input data, and send it into the U-Net network for forward propagation. The network outputs a two-dimensional deformation field. Then, use the spatial transformation network to generate a coordinate sampling grid according to the predicted deformation field, and use bilinear interpolation to resample and transform the SAR image to be registered, so as to obtain the registered SAR image aligned with the optical image.

[0079] Step 4: Calculation of the target loss function. Calculate the target loss function for the registered SAR image and the reference optical image. The target loss function is calculated using mutual information loss, structural similarity loss, deformation regularization loss, and structured distribution correlation loss to quantitatively evaluate the registration quality and obtain the loss value.

[0080] Step 5: Backpropagation and parameter optimization. Based on the loss value, calculate the network parameter gradients through the backpropagation algorithm. Use the selected Adam optimizer to update the network's weight parameters and further optimize the network parameter configuration to improve the registration performance of the model.

[0081] Step 6: Model iterative training and convergence judgment. Repeat Steps 3 to 5 for continuous iterative training until the loss function value converges to a stable level, and finally obtain the trained SAR and optical remote sensing image registration model.

[0082] Step 7: Use the trained SAR and optical remote sensing image registration model (i.e., the trained image registration network model) to perform remote sensing image registration on the target optical remote sensing image and the target synthetic aperture radar image to be registered.

[0083] The present embodiment makes the following solutions for the shortcomings of the existing technology, specifically: (1) Existing feature extraction methods have certain limitations when dealing with the modal differences between SAR images and optical images.

[0084] To address this issue, the present embodiment uses a deep learning-based feature extraction architecture to replace manual features. Specifically, a U-Net convolutional neural network is introduced as the basic structure for feature extraction. The encoder-decoder symmetric structure of U-Net can automatically learn multi-level and multi-scale feature representations from the input SAR and optical images. The encoder part extracts low-level texture details and high-level semantic information through layer-by-layer convolution and pooling; the decoder part gradually reconstructs the features through upsampling and transposed convolution and outputs the deformation field required for registration. This data-driven approach enables the extracted features to have the ability to adapt to modal differences, capture the correspondence between the backscattering patterns of SAR images and the texture structures of optical images, and no longer rely on artificial features that are vulnerable to modal changes. At the same time, the present embodiment incorporates channel attention and spatial attention mechanisms into U-Net, enabling the network to automatically focus on discriminative regions and features in cross-modal images, thereby further improving the adaptability of feature extraction to SAR / optical differences. In summary, the deep feature extraction module based on U-Net endows the present embodiment with the ability to robustly represent different modal image features and significantly improves the matching accuracy of cross-modal registration.

[0085] (2)In the existing methods, feature fusion is insufficient and fails to effectively combine features at different levels of the image.

[0086] To address the problem of insufficient feature fusion ability and inability to integrate multi-level feature information, in this embodiment, an Enhanced Feature Fusion module (EFF) is designed and embedded in the U-Net architecture to fully fuse features at different levels.

[0087] The schematic diagram of the registration network structure based on the attention U-Net and EFF module proposed in this embodiment is shown in Figure 3 . The features of each layer extracted by the encoder are input into the decoder through skip connections, and at each corresponding layer, the EFF module fuses with the upsampled features of the decoder, and the output fused features are used to reconstruct the deformation field layer by layer. The EFF module organically combines low-level texture details with high-level semantic information through a multi-scale and cross-modal feature fusion strategy. Inside the EFF, there are three sub-modules for feature enhancement and fusion in different dimensions: First, a Feature Enhancement sub-module (EAG) is introduced to extract cross-modal common features through grouped convolution and use residual connections to bridge low- and high-level features, improving the adaptability to modal differences while reducing information loss; Second, an Efficient Channel Attention sub-module (ECA) is adopted to calculate channel weights using one-dimensional convolution, highlighting the importance of feature channels in multi-modal images, automatically screening out the feature channels that play a key role in registration, and suppressing irrelevant or noisy information; Third, a Spatial Attention sub-module (SA) is integrated to obtain the significant spatial regions of the feature map through global average pooling and max pooling, focusing on the key positions corresponding to the structures in SAR and optical images, and avoiding wasting attention on background noise regions. Through the above mechanism, the EFF module fuses the features from the encoder and the current decoding layer layer by layer in the decoding stage. In each level of decoding, first, the deep features of the previous layer are upsampled and concatenated with the features generated by the corresponding encoder, and then through the attention fusion of the EFF module, the fused features containing both detailed textures and rich high-level semantics are obtained and input into the subsequent convolution to restore the image deformation field. This cross-level feature fusion design effectively makes up for the deficiency of traditional methods that only use single-layer features, enabling the registration network to simultaneously use multi-scale and multi-semantic information to match the corresponding regions of SAR and optical images, thereby improving the accuracy and robustness of registration.

[0088] (3)The existing technology separately optimizes the global transformation and local deformation, resulting in insufficient coordination in the registration process.

[0089] To address the issue of the disconnection between the global and local registration processes and the lack of coordination in optimization, this embodiment constructs an end-to-end joint optimization framework that integrates global geometric transformation and local fine alignment. In terms of the method structure, a Spatial Transformer Network (STN) is used as a differentiable image transformation module, which is tightly coupled with the U-Net feature extraction sub-network, enabling the global and local adjustments in the registration process to be optimized simultaneously within the same network. In terms of the execution mechanism, first, the SAR image to be registered and the reference optical image are concatenated in the channel dimension and then input into the U-Net network. Through forward propagation, a two-dimensional pixel-level deformation field (i.e., the displacement vector of each pixel) is obtained. This deformation field can represent the comprehensive transformation amount required for aligning the two images due to global rigid transformations (such as translation, rotation, and scale change) and local non-linear deformations (such as local offsets of ground objects). Then, the predicted deformation field is applied to the SAR image to be registered through the STN module, and geometric transformation and interpolation sampling are performed on it to obtain the registration result that is aligned with the optical image. In this process, the global transformation parameters and local fine deformations are jointly generated by the same network, avoiding the inconsistency that may occur in traditional methods where the global and local optimizations are carried out separately and independently. During training, end-to-end backpropagation is used: after calculating the comprehensive loss between the registration result and the reference image, the gradient is backpropagated through the STN to the deformation field generation network, enabling the network parameters representing global and local registration to be updated synchronously. Since the registration process is regarded as a single differentiable system, this embodiment can coordinately optimize the overall pose alignment and local detail matching of the images within a unified framework, significantly improving the coordination and accuracy of the registration process. At the same time, the end-to-end joint optimization avoids the cumbersome feature matching and step-by-step optimization processes, improving the computational efficiency and being more suitable for the registration requirements of large-scale data in actual remote sensing applications.

[0090] (4) Existing methods use simple pixel-level error metrics and do not fully capture the structural differences between cross-modal images.

[0091] To address the problem of a single loss function design that is difficult to capture the structural differences between cross-modalities, this embodiment constructs a multi-loss function that combines multiple metrics to evaluate and constrain the registration result from different aspects, in order to accurately capture the structural differences between SAR and optical images.

[0092] First, the Mutual Information (MI) loss is introduced to measure the statistical correlation between the two registered images. Mutual information can quantify the dependence relationship between the gray-scale distributions of SAR and optical images, and is robust for cross-modal registration with non-linear intensity differences, which can encourage the registered images to achieve the maximum degree of correlation in the histogram distribution.

[0093] Secondly, the Structural Similarity (SSIM) loss is added to measure the similarity between the registered image and the reference image in terms of local luminance, contrast, and structural patterns. SSIM focuses on the structural information of human visual perception, and the SSIM value increases when two images are more similar in local structures such as edges and textures. By minimizing the SSIM loss, the network is encouraged to produce a registration result that is more consistent with the reference optical image in terms of detailed structure, thus effectively capturing the structural differences between cross-modal images.

[0094] In addition, a deformation regularization loss is added to impose a smoothing constraint on the predicted deformation field. This regularization term encourages the estimated deformation field to vary smoothly in space by penalizing overly drastic or discontinuous deformations, avoiding unreasonable distortions, and ensuring the physical and geometric rationality of the registration transformation.

[0095] Finally, the Structured Distribution Correlation (SDC) loss models the SAR image as a Gamma distribution and the optical image as a Gaussian distribution, and performs weighted matching in the structure space, which has a stronger cross-modal adaptation ability; a structure-sensitive weight function is introduced in the structure domain to achieve precise control of the matching in the detail regions, effectively enhancing the edge and texture alignment effect. Compared with traditional Mutual Information (MI) or SSIM losses, the SDC loss is insensitive to speckle noise and more sensitive to structural misalignment, making it suitable for registration task scenarios with high noise and strong structural differences. The SDC loss function is in a differentiable form and can participate in backpropagation as part of an end-to-end training framework and be jointly optimized with registration modules such as the Spatial Transformer Network (STN) module.

[0096] The above four types of losses are combined with appropriate weights to form the final training objective function. Through the comprehensive optimization of multiple metrics, the design of the objective loss function in this embodiment overcomes the limitation that a single pixel difference metric (such as the L2 loss) cannot take into account the significant differences between modalities, enabling the model to simultaneously consider global statistical matching, local structure alignment, deformation field smoothing, and structured distribution correlation during the training process. This more refined and flexible loss function strategy ensures the quality of cross-modal registration: it not only maintains the consistency of the overall distributions between SAR and optical images but also preserves the local structural features of the ground objects, thereby significantly improving the accuracy and robustness of the registration result in complex scenarios. A novel loss function that combines distribution difference modeling and structure domain correlation quantification is proposed in this embodiment to evaluate the structured statistical consistency between the registration result and the reference image.

[0097] Compared with the prior art, in the process of feature extraction, cross-layer fusion, and image space transformation in this embodiment, innovative deep learning architectures and module combinations are adopted, significantly improving the registration accuracy, robustness, and computational efficiency.

[0098] Refer to Figure 5, The embodiment of the present application further provides a remote sensing image registration system based on enhanced feature fusion. The system includes a training data acquisition unit 100, a network model construction unit 200, a deformation field obtaining unit 300, an image pre-registration unit 400, a loss function construction unit 500, a network model training unit 600, and a target image registration unit 700, where: The training data acquisition unit 100 is configured to acquire an optical remote sensing image for model training and a synthetic aperture radar image to be registered; The network model construction unit 200 is configured to construct an image registration network model including a deformation field generation network and a spatial transformation network, where the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; The deformation field obtaining unit 300 is configured to input the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field; The image pre-registration unit 400 is configured to input the synthetic aperture radar image to be registered and the deformation field into the spatial transformation network to obtain a pre-registered image; The loss function construction unit 500 is configured to construct a target loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; The network model training unit 600 is configured to train the image registration network model according to the target loss function until the target loss function converges to obtain a trained image registration network model; The target image registration unit 700 is configured to perform remote sensing image registration on a target optical remote sensing image and a target synthetic aperture radar image to be registered through the trained image registration network model.

[0099] It should be noted that since the remote sensing image registration system based on enhanced feature fusion in this embodiment and the above-mentioned remote sensing image registration method based on enhanced feature fusion are based on the same inventive concept, the corresponding content in the method embodiment also applies to the system embodiment of the present application and will not be elaborated here.

[0100] Refer to Figure 6 , The embodiment of the present application further provides an electronic device. The electronic device includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the above-mentioned remote sensing image registration method based on enhanced feature fusion of the present disclosure.

[0101] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0102] The electronic device according to the embodiments of the present application will be introduced in detail below.

[0103] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure; The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute the remote sensing image registration method based on enhanced feature fusion according to the embodiments of the present disclosure.

[0104] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.

[0105] The embodiments of the present disclosure also provide a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above-mentioned remote sensing image registration method based on enhanced feature fusion.

[0106] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories that are remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0107] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0108] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0111] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0112] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the associated objects before and after. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", where a, b, and c can be single or multiple.

[0113] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0114] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, each functional unit in various embodiments of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0116] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. The embodiments of this application have been described in detail above in conjunction with the accompanying drawings, but this application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can also be made without departing from the purpose of this application.

[0117] The embodiments of this application have been described in detail above in conjunction with the accompanying drawings, but this application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can also be made without departing from the purpose of this application.

Claims

1. A remote sensing image registration method based on enhanced feature fusion, characterized in that, The method includes: Obtaining an optical remote sensing image for model training and a synthetic aperture radar image to be registered; Constructing an image registration network model including a deformation field generation network and a spatial transformation network, wherein the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; Inputting the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field; Inputting the synthetic aperture radar image to be registered and the deformation field into the spatial transformation network to obtain a pre-registered image; Constructing an objective loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; Training the image registration network model according to the objective loss function until the objective loss function converges to obtain a trained image registration network model; Performing remote sensing image registration on a target optical remote sensing image and a target synthetic aperture radar image to be registered through the trained image registration network model.

2. The remote sensing image registration method based on enhanced feature fusion according to claim 1, wherein The step of inputting the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field includes: Stitching the optical remote sensing image and the synthetic aperture radar image to be registered to obtain a stitched image; Taking the stitched image as the input of the deformation field generation network, and performing convolutional feature extraction and downsampling on the stitched image through the encoder in the U-Net network to obtain semantic features at multiple coding levels; Upsampling the semantic features extracted at the last coding level through the decoder in the U-Net network to obtain a current upsampled feature; Performing cross-modal feature fusion on the current upsampled feature and the semantic features at the corresponding coding level of the next upsampling through the enhanced feature fusion module to obtain a target fusion feature; Upsampling the target fusion feature through the decoder in the U-Net network, and continuously performing cross-modal feature fusion on the current upsampling result and the semantic features at the corresponding coding level of the next upsampling until the decoder decoding is completed to obtain a deformation field.

3. The remote sensing image registration method based on enhanced feature fusion according to claim 2, characterized in that, The enhanced feature fusion module includes a feature enhancement sub-module, a channel attention sub-module, and a spatial attention sub-module. The step of performing cross-modal feature fusion on the current upsampled feature and the semantic features at the corresponding coding level of the next upsampling through the enhanced feature fusion module to obtain a target fusion feature includes: Inputting the current upsampled feature and the semantic features at the corresponding coding level of the next upsampling into the feature enhancement sub-module to obtain a first fusion feature; Performing a residual connection on the first fusion feature and the semantic features at the corresponding coding level of the next upsampling to obtain a second fusion feature; Inputting the second fusion feature into the channel attention sub-module to obtain a third fusion feature; Inputting the third fusion feature into the spatial attention sub-module to obtain a target fusion feature.

4. The remote sensing image registration method based on enhanced feature fusion according to claim 1, characterized in that, Constructing an objective loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image includes: Construct a mutual information loss function by calculating the mutual information between the synthetic aperture radar image to be registered and the optical remote sensing image; Construct a structural similarity loss function by calculating the structural similarity between the pre-registered image and the optical remote sensing image; Construct a deformation regularization loss function according to the deformation field; Construct a structured distribution correlation loss function by calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image; Perform a weighted sum of the mutual information loss function, the structural similarity loss function, the deformation regularization loss function, and the structured distribution correlation loss function to obtain an objective loss function.

5. The remote sensing image registration method based on enhanced feature fusion according to claim 4, characterized in that, The step of constructing a structured distribution correlation loss function by calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image includes: Construct the local pixels in the synthetic aperture radar image to be registered into a gamma distribution; Construct the local pixels in the optical remote sensing image into a Gaussian distribution; Calculate the first structure mapping value of the synthetic aperture radar image to be registered and calculate the second structure mapping value of the optical remote sensing image; Calculate the correlation weight of the synthetic aperture radar image to be registered and the optical remote sensing image in the structural domain according to the first structure mapping value and the second structure mapping value; Calculate the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image according to the gamma distribution, the Gaussian distribution, and the correlation weight; Construct a structured distribution correlation loss function based on the structured distribution correlation.

6. The remote sensing image registration method based on enhanced feature fusion according to claim 5, wherein The step of calculating the first structure mapping value of the synthetic aperture radar image to be registered and calculating the second structure mapping value of the optical remote sensing image includes: ; Among them, represents the first structure mapping value of the synthetic aperture radar image to be registered, represents the synthetic aperture radar image to be registered at the gradient at the point, represents the second structure mapping value of the optical remote sensing image, represents the optical remote sensing image at the gradient at the point, represents the vector differential operator.

7. The remote sensing image registration method based on enhanced feature fusion according to claim 5, characterized in that The step of calculating the structured distribution correlation between the synthetic aperture radar image to be registered and the optical remote sensing image according to the gamma distribution, the Gaussian distribution, and the correlation weight includes: ; Among them, represents the structured distribution correlation, represents the sliding window size, represents the correlation weight, represents the probability density estimate of the current window calculated by the gamma distribution, represents the global reference distribution of the synthetic aperture radar image to be registered, represents the probability density estimate of the current window calculated by the Gaussian distribution, represents the global reference distribution of the optical remote sensing image.

8. A remote sensing image registration system based on enhanced feature fusion, characterized in that, The system includes: A training data acquisition unit for acquiring an optical remote sensing image and a synthetic aperture radar image to be registered for model training; A network model construction unit for constructing an image registration network model including a deformation field generation network and a spatial transformation network, wherein the deformation field generation network is constructed based on a U-Net network and an enhanced feature fusion module; A deformation field obtaining unit for inputting the optical remote sensing image and the synthetic aperture radar image to be registered into the deformation field generation network to obtain a deformation field; An image pre-registration unit for inputting the synthetic aperture radar image to be registered and the deformation field into the spatial transformation network to obtain a pre-registered image; A loss function construction unit for constructing an objective loss function based on the deformation field, the synthetic aperture radar image to be registered, the pre-registered image, and the optical remote sensing image; A network model training unit for training the image registration network model according to the objective loss function until the objective loss function converges to obtain a trained image registration network model; A target image registration unit for performing remote sensing image registration on a target optical remote sensing image and a target synthetic aperture radar image to be registered through the trained image registration network model.

9. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatingly connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the remote sensing image registration method based on enhanced feature fusion according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the remote sensing image registration method based on enhanced feature fusion according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Unsupervised MRI (Magnetic Resonance Imaging) image registration algorithm and system based on Transform and ConvNet fusion model

    CN117876445A

  • Image generation method and device based on network joint learning

    CN118037875A

  • Three-dimensional image splicing method and device, computer equipment and storage medium

    CN118195891A

  • Multi-source remote sensing optical image registration and fusion method based on space deformation field

    CN118941600A

  • Image registration method and model training method thereof

    US20210390716A1