Semantic position probability distribution method for SAR-optical image registration
By adopting the semantic position probability distribution method in remote sensing image registration, the pyramid geometric invariant feature extraction network and semantic position attention module are used, combined with the joint loss function, the universality, adaptability and computational complexity of multi-source remote sensing image registration in the prior art is solved, and efficient and accurate SAR-optical image matching is achieved.
Patent Information
- Application Number
- CN202411900653.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-27
AI Technical Summary
The existing multi-source remote sensing image registration methods have poor versatility and adaptability, low matching accuracy, and high computational complexity and time-consuming.
The semantic position probability distribution method for SAR-optical image registration is adopted, and the network and semantic position attention module are extracted by constructing a pyramid geometric invariant feature, combined with the joint loss function, the network weight is optimized, and the matching position probability distribution of each pixel in the target image is generated.
The accuracy and computing efficiency of multimodal SAR-optical image matching are improved, and efficient and accurate registration of multi-source heterogeneous remote sensing images is achieved, reducing the need for model retraining due to scene changes.
Smart Images

Figure CN120047702A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image processing, and relates to a deep learning framework with semantic probability distribution for SAR and optical image registration. Background Art
[0002] The rapid development of satellite sensor technology has provided diverse sources for remote sensing image acquisition. The joint use of synthetic aperture radar (SAR) and optical images under different imaging conditions allows for providing highly complementary information about the same observation scene. Effective registration between the two is the prerequisite and foundation for realizing the fusion and utilization of this information. The challenge lies in overcoming the heterogeneous characteristics between the images.
[0003] Traditional image matching methods include feature-based and region-based methods. Among them, feature-based methods first need to obtain significant feature descriptors in the image, such as corners, lines, surfaces, etc., and then measure the corresponding relationships between the feature descriptions, such as SIFT, RIFT. Such methods are usually only applicable to images that meet specific radiation constraints and have no geometric distortion. Different from feature-based matching methods, region-based methods use similarity metrics to find the matching relationships between image pairs, avoiding feature extraction. Similarity metrics include sum of squared differences SSD, normalized cross-correlation NCC, and normalized mutual information MI. However, these methods perform poorly when dealing with large geometric deformations and have a high computational complexity.
[0004] In recent years, deep learning-based matching methods have utilized the powerful feature extraction ability of convolutional neural networks to obtain key points and feature descriptions, improving the stability and efficiency of SAR-optical image matching and overcoming the limitations of handcrafted feature extraction methods in modeling complex non-linear relationships. However, there are still challenges in dealing with the non-linear radiation differences between SAR and optical images, especially the problem of non-repeatability of key points. Some studies use deep neural networks to generate candidate matching regions, but these methods still have limitations in terms of computational efficiency and adaptability. There are also some works that use semantic template matching frameworks to directly obtain image correspondences through feature fusion, but such methods rely on the fitting of the dataset and need to retrain the network for different scenarios in practical applications.
[0005] In summary, there are still two main problems in the current related research on SAR and optical image matching based on deep neural network frameworks. First, due to the complex diversity of scenarios in practical applications, the network needs to be retrained for different scenarios. Second, the strategy of searching pixel by pixel to obtain the similarity of the match is very time-consuming. Summary of the Invention
[0006] Aiming at the deficiencies existing in the prior art, the purpose of the present invention is to provide a semantic position probability distribution method for SAR-optical image registration, so as to solve the problems of poor generality and adaptability, low matching accuracy, high computational complexity and time consumption of existing multi-source remote sensing image registration methods. The present invention aims to improve the accuracy and computational efficiency of multi-modal SAR-optical image matching and achieve efficient and accurate registration of multi-source heterogeneous SAR-optical remote sensing images.
[0007] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:
[0008] A semantic position probability distribution method for SAR-optical image registration includes the following steps:
[0009] Step 1, construct a pyramid geometric invariant feature extraction network to extract multi-scale and high-level semantic features from the reference image, i.e., the optical image, and the sensed image, i.e., the SAR image to be registered; adopt a dual-branch ResNet50 backbone network with shared weights, combined with a Feature Pyramid Network (FPN), to generate feature maps with geometric invariance and high semantic expressiveness at different scales;
[0010] Step 2, construct a semantic position attention module to perform cross-image interaction and fusion on the same-level features in the multi-scale feature maps; this module captures the correlation between the reference image and the sensed image in terms of spatial position through a position attention mechanism, generates a position-dependent feature map, and strengthens the cross-modal adaptability of the model;
[0011] Step 3, design a training strategy for the image semantic position registration network to learn the matching position relationship between the input registration sample pair of images; adopt a joint loss function, comprehensively considering the semantic matching loss and the position matching loss, optimize the network weights, and construct a cross-modal matching model;
[0012] Step 4, apply the trained registration network to the input reference image and the image to be registered to generate the matching position probability distribution of each pixel in the target image and predict the spatial correspondence relationship between the two images.
[0013] The present invention further includes the following technical features:
[0014] Specifically, in the above-mentioned step 1, the pyramid geometric invariant feature extraction network adopts a dual-branch ResNet50+FPN architecture, including a bottom-up residual network layer, a top-down upsampling layer, and an intermediate connection structure.
[0015] Specifically, the residual network layer includes an initial CNN layer and a residual network; the residual network further includes four residual blocks connected in sequence, and the initial CNN layer is connected to the first residual block; the initial CNN layer uses a CNN structure with stride = 2 to reduce the size of the feature map and improve the computational efficiency in the position attention module; the four residual blocks are denoted as Res-1, Res-2, Res-3, and Res-4; their outputs are respectively corresponding to the inputs of the upsampling layers from bottom to top after passing through a 1×1 convolutional layer.
[0016] The upsampling layer is performed using the interpolation upsampling method to expand the upsampled feature map to the same dimension as the feature map of the next layer; after passing through a 3×3 convolution, the upsampling layer is fused with the output of the residual block at the same level.
[0017] The intermediate connection structure sequentially fuses the feature maps generated from bottom to top with the upsampled result map to generate four new feature maps; the intermediate connection structure also includes a 1×1 convolution and a 3×3 convolution; the 1×1 convolution is used to change the number of channels of the output feature map of the residual block and fuse it with the upsampled feature map; the 3×3 convolution is used to eliminate the aliasing effect of upsampling; the input of the intermediate connection structure corresponds to the output of the residual block at the same level, and its output corresponds to the input of the upsampling layer at the same level.
[0018] Specifically, in step 2, the multi-level feature maps of the extracted reference image and the perceptual image are sequentially input into the position attention module for semantic matching, which is used to capture the dependency relationship between any two positions in the reference image and the perceptual image, and upsample the fused features to obtain feature maps of the same size; the position attention module includes four position attention units, and each position attention unit includes multiple convolutional layers for feature extraction and a cross-attention calculation module for feature fusion.
[0019] Specifically, the calculation process of the position attention module capturing the correlation relationship between the reference image and the perceptual image in the spatial position and generating a feature map with position dependence is as follows:
[0020] Let s represent the feature map of the perceptual image, with a shape of (c, h 1 , w 1 ), and let o represent the feature map of the reference image, with a shape of (c, h 2 , w 2 ); the feature maps pass through the convolutional layer to obtain s 1 , s 2 and o 1 , o 2 , then s 1 , s 2 and o 1Reshape into (c, 1, n), (c, 1, n), and (c, 1, N), where n = h 1 × w 1 , N = h 2 × w 2 ; then, multiply s 1 by o 1 and apply the SoftMax operation to obtain the spatial attention map A(N, n), then multiply A(N, n) by s 2 and reshape it into a feature map of shape (c, h 2 , w 2 ), and finally add the result to o 2 to obtain the final output feature map.
[0021] Specifically, in step 3, the joint loss function is:
[0022] L = αL p + βL s
[0023] where L represents the total loss function, L p represents the centroid position loss, which is used to measure the deviation between the predicted centroid coordinates and the true coordinates to ensure the spatial position matching accuracy; L s represents the cross-entropy loss, which is used to measure the semantic classification error of the model to ensure that the output probability is close to the true label; α and β represent the relative importance for adjusting the two loss functions, and α + β = 1.
[0024] Specifically, the cross-entropy loss is:
[0025]
[0026] where L s represents the cross-entropy loss function, which is used to measure the error between the model prediction probability and the true label; N represents the total number of samples; i is the index of the sample; y i represents the true label of the i-th sample.
[0027] Specifically, the centroid position loss function L p is:
[0028] L p = (x - x true ) 2 + (y - y true )
[0029] where x true and y true represent the true coordinate positions of the template centroid in the reference image.
[0030] Specifically, the centroid position (x, y) is:
[0031]
[0032] Among them, C I and C J represent the weights of all output positions i and j respectively; C IJ represents the sum of semantic probabilities.
[0033] Specifically, the formula for the sum of the semantic probabilities is:
[0034]
[0035] Among them, P′ ij represents the semantic position probability value of the pixel (i, j); C IJ represents the sum of the semantic probabilities, which is obtained by summing the probability values P′ ij of all pixel positions; the superscript P of the summation represents the number of rows and the height of the feature map in the vertical direction; the superscript w of the summation represents the number of columns and the width of the feature map in the horizontal direction;
[0036] The weights C I and C J of all output positions i and j are calculated as follows:
[0037]
[0038] Among them, C I represents the weighted centroid coordinate in the vertical direction; the vertical coordinates i of all pixel positions are weighted with the corresponding semantic probability P′ ij and accumulated over the entire feature map; C J represents the weighted centroid coordinate in the horizontal direction; the horizontal coordinates j of all pixel positions are weighted with the corresponding semantic probability P′ ij , and accumulated over the entire feature map; i is the vertical direction coordinate index of the pixel in the feature map; j is the horizontal direction coordinate index of the pixel in the feature map; w represents the width and height of the feature map.
[0039] Compared with the prior art, the present invention has the following technical effects:
[0040] The present invention proposes a new deep neural network model for solving current SAR and optical image registration. On the one hand, a ResNet-50+FPN network with a pyramid structure is introduced to construct a pyramid-shaped feature for each local pixel to handle the geometric differences between SAR and optical images, enhancing the resistance to local distortion. Secondly, a SAR–optical image position attention module is designed to input the features containing high-level semantic information obtained from the reference image and the perceptual image to aggregate matching information for obtaining the dependency relationship between two positions of the SAR-optical images, reducing the need for retraining due to scene changes. On the other hand, different from the pixel-by-pixel search matching method, each pixel position of the SAR image is mapped onto the optical image to obtain the matching position probability distribution in this network. Based on this, a new loss function based on the output position probability distribution is proposed to perform weighted averaging on the output probability distribution, converting the position probability distribution boundary alignment into a point-to-point matching problem, optimizing the network from both semantic and matching position perspectives, and further reducing the computational complexity while offsetting some pixel-independent biases caused by the radiometric differences between SAR and optical images. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of a deep learning method for semantic position probability distribution for image registration.
[0042] Figure 2 It is a structural block diagram of an image registration network.
[0043] Figure 3 It is a schematic diagram of a feature extraction network.
[0044] Figure 4 It is a schematic diagram of a position attention module.
[0045] Figure 5 It is a schematic diagram of a centroid position joint loss function.
[0046] Figure 6 It is a registration result diagram obtained by the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following are specific embodiments of the present invention. It should be noted that the present invention is not limited to the following specific embodiments, and all equivalent transformations made on the basis of the technical solution of the present application fall within the protection scope of the present invention.
[0048] Embodiment:
[0049] An embodiment of the present invention provides a semantic position probability distribution method for SAR-optical image registration. The new deep neural network model proposed by the present invention for solving the current SAR and optical image registration focuses on the multi-scale feature extraction of the reference image and the image to be registered by the pyramid geometric invariant feature extraction network to handle the geometric differences between SAR and optical images. A position attention module is designed to capture the dependency relationship between two positions of the SAR-optical images, and the semantic position probability distribution of the network is obtained by aggregating the matching information. The problem of matching accuracy and efficiency is solved by jointly optimizing the network from two perspectives of semantic and centroid position matching. In this method, the multi-scale pyramid features include the outputs of different levels of the feature extraction network. The outputs of the same level are input into the position attention module for feature interaction and fusion to obtain feature maps with position dependency at different scales. The feature maps with position dependency are resampled to the same size and then coupled and fed into the final CNN layer to generate the matching position probability distribution result.
[0050] See Figure 1 , the method includes:
[0051] Step 1, construct a pyramid geometric invariant feature extraction network to extract multi-scale and high-level semantic features from the reference image (optical image) and the sensed image (SAR image to be registered); adopt a weight-sharing dual-branch ResNet50 backbone network, combined with a Feature Pyramid Network (FPN), to generate feature maps with geometric invariance and high semantic expressiveness at different scales.
[0052] Step 2, construct a semantic position attention module to perform cross-image interaction and fusion on the same-level features in the multi-scale feature maps; this module captures the correlation relationship between the reference image and the image to be registered in the spatial position through the position attention mechanism, generates feature maps with position dependency, and strengthens the cross-modal adaptability of the model.
[0053] Step 3, design the training strategy of the image semantic position registration network to learn the matching position relationship between the input registration sample pair of images; adopt a joint loss function, comprehensively consider the semantic matching loss and the position matching loss, optimize the network weights, and construct a cross-modal matching model.
[0054] Step 4, apply the trained registration network to the input reference image and the image to be registered to generate the matching position probability distribution of each pixel in the target image and predict the spatial correspondence relationship between the two images.
[0055] Specifically, the reference image and the image to be registered are input into the pyramid geometric invariant feature extraction network. The pyramid geometric invariant feature extraction network adopts a dual-branch Resnet50+FPN structure with shared weights, which is used to extract the pyramid-shaped high-level semantic information of the reference image and the sensed image respectively, and construct pyramid features for the reference image and the sensed image.
[0056] In this embodiment, aiming at the problems of poor generality and adaptability, low matching accuracy, high computational complexity and time consumption in the existing multi-source remote sensing image registration methods, it aims to improve the accuracy and computational efficiency of multi-modal SAR-optical image matching and achieve the efficient and accurate registration of multi-source heterogeneous remote sensing images. The core lies in using two ResNet50+FPN architectures with shared weights to construct pyramid features for the reference image and the sensed image, enhancing the resistance to local distortion. At the same time, a position attention module is designed to effectively obtain the matching position dependence relationship between features, reduce the need for retraining due to scene changes, and improve the generalization ability of the model. The response of the final output label represents the corresponding relationship of the semantic position probability of the sensed image in the reference image.
[0057] In one embodiment, Figure 2 The structural block diagram of the image registration network is shown, including a pyramid geometric invariant feature extraction network and a position attention module.
[0058] The pyramid geometric invariant feature extraction network adopts a dual-branch ResNet50+FPN architecture with shared weights to construct pyramid features for the reference image and the sensed image. Specifically, it includes a bottom-up residual network layer, a top-down upsampling layer and an intermediate connection structure. See Figure 3 .
[0059] The residual network layer includes an initial CNN layer and a residual network; the residual network also includes four residual blocks connected in sequence. The initial CNN layer is connected to the first residual block; the initial CNN layer uses a CNN structure with stride = 2 to reduce the size of the feature map and improve the computational efficiency in the position attention module. Specifically, the four residual blocks are denoted as Res-1, Res-2, Res-3 and Res-4; their outputs are respectively corresponding to the inputs of the bottom-up upsampling layer after passing through a 1×1 convolutional layer.
[0060] The upsampling layer is implemented using the interpolation upsampling method to expand the upsampled feature map to the same dimension as the feature map of the next layer. Specifically, the upsampling layer is fused with the output of the same-level residual block after passing through a 3×3 convolution.
[0061] The intermediate connection structure sequentially fuses the feature maps generated bottom-up with the upsampled result maps to generate four new feature maps p 1 , p 2 , p3 , p 4 ; The intermediate connection structure further includes a 1×1 convolution and a 3×3 convolution; the 1×1 convolution is used to change the number of channels of the feature map output by the residual block and fuse it with the upsampled feature map; the 3×3 convolution is used to eliminate the aliasing effect of upsampling. Specifically, the input of the intermediate connection structure corresponds to the output of the residual block at the same level, and its output corresponds to the input of the upsampling layer at the same level.
[0062] The position attention module includes four position attention units, see Figure 4 . The position attention module is used to establish a position-dependent reference image between two images. Specifically, the multi-scale high-level semantic information feature maps of the reference image and the perceptual image correspond level by level, and are jointly input into the position attention unit to obtain four fused feature maps with position dependence and perform upsampling to obtain feature maps of the same size. Finally, the upsampled semantic position feature maps are coupled and fed into the CNN to output the matching position probability distribution. Specifically, the position attention unit includes multiple convolutional layers for feature extraction and a cross-attention calculation module for feature fusion.
[0063] In the embodiment, the multi-level feature maps of the extracted reference image and the perceptual image are input into the position attention module for semantic matching level by level, which is used to capture the dependence between any two positions in the reference image and the perceptual image. The calculation process of the fused feature map with position dependence is as follows:
[0064] Let s represent the feature map of the perceptual image, with a shape of (c, h 1 , w 1 ), and let o represent the feature map of the reference image, with a shape of (c, h 2 , w 2 ); the feature maps pass through convolutional layers to obtain s 1 , s 2 and o 1 , o 2 , and then reshape s 1 , s 2 and o 1 into (c, 1, n), (c, 1, n) and (c, 1, N), where n = h 1 ×w 1 , N = h 2 ×w 2 ; then, multiply s 1 by o 1 and apply the SoftMax operation to obtain the spatial attention map A(N, n), and then multiply A(N, n) by s 2 and reshape it into the shape (c, h 2 , w 2The feature map of (), and finally add the result to o 2 Obtain the feature map of the final output; the calculation process is described as:
[0065]
[0066] Among them, A(N,n) is used to measure the influence of position i on position j. The more similar the feature representations of two positions are, the greater their contribution to the correlation between them.
[0067]
[0068] Among them, α is the learned weight and is initialized to 0. The result E at each position is expressed as the weighted sum of the feature domain optical image features of all positions in the SAR. Therefore, E has a global semantic position relationship, which allows further enhancing the matching degree and semantic consistency of similar features.
[0069] Upsample the fused features to obtain a feature map with the same size. Finally, couple and feed the upsampled semantic position feature map into the CNN to output the matching position probability distribution. The semantic position probability distribution describes the probability that each pixel in the SAR image is located in the reference image.
[0070] Since the output position probability distribution is unordered, they cannot be used to determine the coordinates and matching probabilities of pixels. In addition, the final matching result cannot be based on the correspondence of specific 1-pixel probability values in the SAR image. Therefore, use the calculated centroid to replace all matching pixels, and the output centroid is the weighted average of the feature map.
[0071] In this embodiment, a loss function is proposed to calculate the centroid position loss while considering the semantic loss between the label and the output.
[0072] Specifically, by including the centroid position loss function in the deep learning network, the features are semantically abstracted into high levels. The semantic loss function causes the network output to fit the label, while the centroid position loss function maximizes the matching from a single point without having to fully fit the training label. Combining the two loss functions can further improve the matching accuracy because the semantic loss function can guide the network output to fit the label, while the centroid position loss function can offset some pixel-independent biases caused by the radiation difference between the SAR and the optical image.
[0073] In one embodiment, Figure 5 Shows the loss function used for network optimization: while considering the semantic position matching loss between the label result and the prediction result, introduce the centroid position loss, and jointly optimize the process of training the image registration network. The specific calculation process includes:
[0074] Formula for the centroid of the feature map:
[0075]
[0076] Where P′ ij represents the semantic position probability value of pixel (i, j); C IJ represents the sum of semantic probabilities, obtained by summing the probability values P′ ij of all pixel positions; The summation superscript P: represents the number of rows (height) of the feature map in the vertical direction; The summation superscript w: represents the number of columns (width) of the feature map in the horizontal direction.
[0077]
[0078] Where C I represents the weighted centroid coordinate in the vertical direction (row); The vertical coordinates i of all pixel positions are weighted, with the weight being the corresponding semantic probability P′ ij and accumulated over the entire feature map; C J represents the weighted centroid coordinate in the horizontal direction (column); The horizontal coordinates j of all pixel positions are weighted, with the weight being the corresponding semantic probability P′ ij , and accumulated over the entire feature map; The vertical direction (row) coordinate index of pixel i in the feature map; The horizontal direction (column) coordinate index of pixel j in the feature map; w represents the width and height of the feature map (assuming a square feature map, width equals height).
[0079]
[0080] Where P′ ij represents the semantic position probability value of pixel (i, j); C I and C J represent the weights of all output positions i and j respectively; (x, y) represents the centroid position;
[0081] The centroid position loss function L p , used to guide the output to correspond to the center point:
[0082] L p =(x - x true ) 2 +(y - y true )
[0083] Where x true and y true represent the true coordinate positions of the template centroid in the reference image.
[0084] During training, the ground truth label is the position of each pixel in the perceptual image corresponding to the reference image. The semantic position loss function L sIt is defined as the cross-entropy loss between the position-dependent matching probability and the true value label, and is used to optimize the matching position of each pixel in the sensed image.
[0085]
[0086] Among them, L s represents the cross-entropy loss function, which is used to measure the error between the model prediction probability and the true label; N represents the total number of samples; i is the index of the sample; y i represents the true label of the i-th sample.
[0087] By jointly optimizing the two loss functions, the final loss function is expressed as:
[0088] L = αL p + βL s
[0089] Among them, L represents the total loss function, and L p represents the centroid position loss, which is used to measure the deviation between the predicted centroid coordinates and the true coordinates, and ensure the spatial position matching accuracy; L s represents the cross-entropy loss, which is used to measure the semantic classification error of the model and ensure that the output probability is close to the true label; among them, α and β represent the relative importance used to adjust the two loss functions; and α + β = 1.
[0090] In one embodiment, different weights are usually assigned to the two loss functions. Specifically, (α, β) = (0.1, 0.9) mainly emphasizes semantic information, and the generated features are close to the label. After experiments, the accuracy of its matching results is lower than that of the combination (0.3, 0.7). Similarly, for the combination (α, β) = (0.9, 0.1), it means that it mainly emphasizes the probability related to the centroid position, and the overall matching accuracy is also less than that of the combination (0.3, 0.7). The original intention of designing the network is to find the dependency or matching for each pixel of the sensed image according to the reference image. However, due to the non-linear radiation change between the SAR image and the optical image, it is impossible to achieve an exact match for each pixel on the reference image. Therefore, adjusting the semantic position dependency and the centroid position loss function of the weights (α, β) enables the network to obtain an accurate match.
[0091] The trained network uses the reference image and the sensed image as an image pair to predict the corresponding positions between the two images. Its core idea is to construct pyramid features for the reference image and the sensed image using two weight - shared ResNet - 50+FPN architectures, and then use the position attention mechanism to establish position dependencies between the two images. By mapping the matching position probabilities between the two images, two loss functions are combined to further improve the matching accuracy, because the semantic loss function can guide the network to output a fitting label, while the centroid position loss function can offset some pixel - independent biases caused by the radiometric differences between SAR and optical images.
[0092] To further verify the registration performance of the method of this application on the entire image, experiments were carried out on four groups of SAR and optical image pairs facing urban scenes. Figure 6 Four groups of qualitative matching results are shown. Specifically, (a) the SAR image is used as the sensed image to be registered, (b) the optical image is used as the reference image, and (c) the qualitative registration result map obtained after the experiment of the two images. For the registered image of the chessboard mosaic, the edges of its features are continuous, and in the SAR - optical matching results of the four groups of scenes, a good overlapping effect between regions is shown.
[0093] This method has successfully overcome the need for model retraining caused by scene changes and the high computational complexity brought by pixel - by - pixel search. This not only improves the accuracy and robustness of multimodal SAR - optical image matching, but also significantly improves the computational efficiency of the matching process, enhancing the applicability and popularization of the method in practical remote sensing applications.
[0094] The preferred embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above - mentioned embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.
[0095] In addition, it should be noted that, in the case of no contradiction, the various specific technical features described in the above - mentioned specific embodiments can be combined in any appropriate way. To avoid unnecessary repetition, the present invention does not separately describe various possible combination ways.
[0096] Furthermore, any combination can be made between various different embodiments of the present invention as long as it does not violate the idea of the present invention, and it should also be regarded as the content disclosed by the present invention.
Claims
1. A semantic position probability distribution method for SAR-optical image registration, characterized in that: The following steps are involved: Step 1: Construct a pyramid geometry-invariant feature extraction network to extract multi-scale and high-level semantic features from the reference image, i.e., the optical image, and the perception image, i.e., the SAR image to be registered. A weight-sharing dual-branch ResNet50 backbone network is used in combination with a feature pyramid network FPN to generate feature maps with geometric invariance and high semantic expression at different scales. Step 2: Construct a semantic position attention module to perform cross-image interaction and fusion for the same-level features in the multi-scale feature map. This module captures the spatial relationship between the reference image and the perceived image through the position attention mechanism, generates a feature map with position dependence, and enhances the cross-modal adaptability of the model. Step 3: Design a training strategy for the image semantic position registration network to learn the matching position relationship between the input registration sample pairs; use a joint loss function to combine the semantic matching loss and the position matching loss, optimize the network weights, and build a cross-modal matching model; Step 4: Apply the trained registration network to the input reference image and the image to be registered, generate the probability distribution of the matching position of each pixel in the target image, and predict the spatial correspondence between the two images.
2. The semantic position probability distribution method for SAR-optical image registration according to claim 1, characterized in that: In the step 1, the pyramid geometry invariant feature extraction network adopts a dual-branch weight-sharing ResNet50+FPN architecture, including a bottom-up residual network layer, a top-down upsampling layer, and an intermediate connection structure.
3. The semantic position probability distribution method for SAR-optical image registration according to claim 2, characterized in that: The residual network layer includes an initial CNN layer and a residual network; the residual network also includes four residual blocks connected in sequence, and the initial CNN layer is connected to the first residual block; the initial CNN layer uses a CNN structure with stride=2 to reduce the size of the feature map and improve the computational efficiency in the position attention module; the four residual blocks are represented as Res-1, Res-2, Res-3 and Res-4; their outputs correspond to the inputs of the bottom-up upsampling layers after passing through a 1×1 convolution layer; The upsampling layer is performed using an interpolation upsampling method to expand the upsampled feature map to the same dimension as the feature map of the next layer; The upsampling layer undergoes 3×3 convolution and is fused with the output of the residual block at the same level; The intermediate connection structure fuses the feature maps generated from bottom to top with the upsampled result maps in turn to generate four new feature maps; the intermediate connection structure also includes 1×1 convolution and 3×3 convolution; the 1×1 convolution is used to change the number of channels of the residual block output feature map and fuse it with the upsampled feature map; the 3×3 convolution is used to eliminate the confounding effect of upsampling; the input of the intermediate connection structure corresponds to the output of the residual block at the same level, and its output corresponds to the input of the upsampling layer at the same level.
4. The semantic position probability distribution method for SAR-optical image registration according to claim 3, characterized in that: In the step 2, the extracted multi-level feature maps of the reference image and the perceived image are input into the position attention module for semantic matching step by step, which is used to capture the dependency between any two positions in the reference image and the perceived image, and the fused features are upsampled to obtain feature maps of the same size; the position attention module includes four position attention units, and the position attention unit includes multiple convolutional layers for feature extraction and a cross-attention calculation module for feature fusion.
5. The semantic position probability distribution method for SAR-optical image registration according to claim 4, characterized in that: The position attention module captures the relationship between the reference image and the perceived image in spatial position, and the calculation process of generating a feature map with position dependence is: Denote s as the feature map of the perceived image with a shape of (c, h1, w1), and denote o as the feature map of the reference image with a shape of (c, h2, w2); the feature maps are passed through the convolutional layer to obtain s1, s2, o1, o2, and then s1, s2, and o1 are reshaped into (c, 1, n), (c, 1, n), and (c, 1, N), where n = h1 × w1, N = h2 × w2; then, s1 is multiplied by o1 and the SoftMax operation is applied to obtain the spatial attention map A(N, n), and then A(N, n) is multiplied by s2 and reshaped into a feature map of shape (c, h2, w2), and finally the result is added to o2 to obtain the feature map of the final output.
6. The semantic position probability distribution method for SAR-optical image registration according to claim 1, characterized in that: In step 3, the joint loss function is: L=αL p +βL s Among them, L represents the total loss function, L p represents the centroid position loss, which is used to measure the deviation between the predicted centroid coordinates and the true coordinates to ensure the accuracy of spatial position matching; L s represents the cross entropy loss, which is used to measure the semantic classification error of the model and ensure that the output probability is close to the true label; α and β represent the relative importance of adjusting the two loss functions, and α+β=1.
7. The semantic position probability distribution method for SAR-optical image registration according to claim 6, characterized in that: The cross entropy loss is: Among them, L s represents the cross entropy loss function, which is used to measure the error between the model's predicted probability and the true label; N represents the total number of samples; i is the index of the sample; y i represents the true label of the i-th sample.
8. The semantic position probability distribution method for SAR-optical image registration according to claim 6, characterized in that: The centroid position loss function L p for: L p (xx) true ) 2 +(yy true ) Among them, x true and true Represents the true coordinate position of the template centroid in the reference image.
9. The semantic position probability distribution method for SAR-optical image registration according to claim 8, characterized in that: The centroid position (x, y) is: Among them, C I and C J Represent the weights of all output positions i and j respectively; C IJ Represents the sum of semantic probabilities.
10. The semantic position probability distribution method for SAR-optical image registration according to claim 9, characterized in that: The summation formula of the semantic probability is: Among them, P′ ij Represents the semantic position probability value of pixel (i, j); C IJ Represents the sum of semantic probabilities, expressed by the probability values P′ for all pixel positions ij The sum is obtained; the sum superscript P: represents the number of rows and height of the feature map in the vertical direction; the sum superscript w: represents the number of columns or width of the feature map in the horizontal direction; The weights C of all output positions i and j I and C J The calculation formula is: Among them, C I Represents the weighted centroid coordinates in the vertical direction; the vertical coordinates i of all pixel positions are weighted, and the weight is the corresponding semantic probability P ij ′ and accumulates over the entire feature map; C J Represents the weighted centroid coordinates in the horizontal direction; the horizontal coordinates j of all pixel positions are weighted, and the weight is the corresponding semantic probability P ij ′, and accumulate it on the entire feature map; The i pixel is the vertical coordinate index in the feature map; the j pixel is the horizontal coordinate index in the feature map; w represents the width and height of the feature map.