Multi-modal remote sensing image registration model construction method based on double-branch parallel feature learning network, registration method, related device, electronic equipment and storage medium

By using a dual-branch parallel feature learning network in remote sensing image registration, the local and global features of optical images and SAR images are extracted and fused, the problem of limited registration accuracy in the prior art is solved, and more efficient and robust multimodal remote sensing image registration is achieved.

CN119941812APending Publication Date: 2025-05-06XINJIANG UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510023077.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing remote sensing image registration method based on deep learning can only extract local features, but cannot extract global and local features of optical images and SAR images at the same time, resulting in limited registration accuracy.

Method used

The multimodal remote sensing image registration model based on the dual-branch parallel feature learning network is adopted, and the local and global features of the image are extracted respectively through the optical branch network and the SAR branch network, and the modal adaptive encoder and hierarchical feature registration module are fused and enhanced.

Benefits of technology

It improves the accuracy and robustness of remote sensing image registration, can fully consider the characteristic differences between optical images and SAR images, and enhances the accuracy and stability of registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941812A_ABST
    Figure CN119941812A_ABST
Patent Text Reader

Abstract

A multi-modal remote sensing image registration model construction method based on a double-branch parallel feature learning network, a registration method, a related device, an electronic device and a storage medium, and the method comprises the steps: obtaining a plurality of samples, dividing the samples into a training set and a test set according to a proportion, training a model constructed based on the double-branch parallel feature learning network through the training set, and obtaining a multi-modal remote sensing image registration model; and obtaining a multi-modal remote sensing image registration model, testing the trained multi-modal remote sensing image registration model by using the test set, and optimizing model parameters. An optical branch network and an SAR branch network constructed based on a double-branch parallel feature learning network are introduced, each of the optical branch network and the SAR branch network comprises a modal adaptive encoder and a hierarchical feature registration module, and the modal adaptive encoders can extract global features and local features of an image; the model obtained by training is no longer limited by local features, and the hierarchical feature registration module can optimize and enhance features and improve the registration accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image registration, and is a multimodal remote sensing image registration model construction method based on a dual-branch parallel feature learning network, a registration method, a related device, an electronic device and a storage medium. Background Art

[0002] Remote sensing image registration is a preprocessing step in the field of remote sensing image analysis. It is the process of spatial and temporal registration of data acquired by different sensors at different viewing angles and at different times. The purpose is to match and superimpose two or more images acquired at different times, different sensors or under different conditions to ensure their spatial consistency. Among them, optical image and synthetic aperture radar (SAR) image registration plays a vital role in remote sensing image analysis and can provide high-quality data support for subsequent remote sensing image change detection.

[0003] Existing remote sensing image registration methods include:

[0004] (1) Region-based remote sensing image registration method. The limitations of this method are mainly reflected in two aspects. First, this method relies on some prior conditions to predict the approximate matching range to facilitate the subsequent template matching process. The prior conditions can be the geographic reference information of the multimodal remote sensing image (RPC, POS data, etc.), or a rough mapping model (for example, affine and projection transformation models), or obtained by manually selecting several evenly distributed checkpoints. Second, this method has requirements for the overlapping area between the input image and the reference image. If the overlapping area between the input image and the reference image is too small, the registration result will be greatly affected or even fail.

[0005] (2) Feature-based remote sensing image registration methods, which use the local spatial relationship between adjacent pixels to construct high-dimensional information feature vectors for each feature point. These high-dimensional feature vectors are crucial to the matching performance of the designed descriptor. Therefore, feature matching methods usually face a heavy computational burden compared to template matching methods, and are prone to inevitable serious outliers (false matches) in the assumed matches, especially in multimodal registration cases where scale, rotation, and radiation differences exist at the same time. Moreover, the registration robustness of feature-based methods is not as stable as that of region-based methods.

[0006] (3) Remote sensing image registration method based on deep learning. This method is completely driven by training data. Its ability to extract deep semantic information of images can effectively improve the performance of remote sensing image registration. For example, multimodal image registration method based on attention mechanism and multimodal remote sensing image registration method based on deep learning can use convolutional neural network (CNN) to extract image features, and then realize registration through feature matching and transformation parameter estimation, and use generative adversarial network (GAN) to generate more realistic registration images to improve registration accuracy.

[0007] However, due to factors such as diverse scenes, wide imaging range, and mixed noise, it is difficult to construct a sufficient and representative dataset for model training and testing. In addition, there are differences in the imaging mechanisms of optical images and SAR images in multimodal remote sensing image registration (for example, optical images are greatly affected by lighting conditions, while SAR images are affected by surface roughness and dielectric constant). Existing remote sensing image registration methods based on deep learning can usually only extract local features for registration, and cannot simultaneously extract global and local features of optical and SAR images. Therefore, the feature differences between the two images cannot be fully considered, and there are problems such as loss of discriminant information and degradation of matching performance, which can easily lead to limited registration accuracy. Summary of the invention

[0008] The present invention provides a multimodal remote sensing image registration model construction method, registration method, related device, electronic device and storage medium based on a dual-branch parallel feature learning network, which overcomes the shortcomings of the above-mentioned prior art and can effectively solve the problem that the existing remote sensing image registration method based on deep learning can only extract local features for registration, resulting in limited registration accuracy.

[0009] One of the technical solutions of the present invention is achieved by the following measures: a method for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network, comprising:

[0010] A number of samples are obtained and divided into a training set and a test set in proportion, wherein each sample includes an optical image and a SAR image of the same area and an image registration result identifier;

[0011] The model constructed based on the dual-branch parallel feature learning network is trained using the training set. A loss function of a weighted strategy is introduced during the training. When the value of the loss function is stable, the training is terminated to obtain a multimodal remote sensing image registration model, wherein the multimodal remote sensing image registration model includes an image feature extraction model and a feature registration model. The image feature extraction model includes an optical branch network and a SAR branch network constructed based on the dual-branch parallel feature learning network. Both include a modality adaptive encoder and a hierarchical feature registration module. The modality adaptive encoder extracts local features and global features of the image and fuses them to form a corresponding comprehensive feature vector. The hierarchical feature registration module enhances the comprehensive feature vector based on the attention mechanism and the residual structure and outputs the corresponding feature map.

[0012] The trained multimodal remote sensing image registration model is tested using the test set, the model parameters are optimized, and a multimodal remote sensing image registration model that meets the test requirements is output.

[0013] The following are further optimizations and / or improvements to the above technical solutions:

[0014] As mentioned above, the modal adaptive encoder includes an adaptive feature extraction module, a Swin Transformer encoding module and a multi-scale feature fusion module. The adaptive feature extraction module uses multiple basic convolution blocks to extract local features of the image. The Swin Transformer encoding module extracts global features of the image based on residual connection and attention mechanism. The multi-scale feature fusion module performs multi-scale fusion on the local features and global features of the image and outputs the corresponding feature map.

[0015] The above-mentioned hierarchical feature registration module performs dimension adjustment, average pooling and feature enhancement on the input feature map in sequence based on convolution blocks, attention mechanism and residual structure.

[0016] The above feature registration model performs feature registration based on feature point similarity evaluation and local block cosine similarity evaluation.

[0017] The loss function calculation process of the above weighted strategy includes:

[0018] Determine the edge loss and registration loss during sample training;

[0019] Determine the weight of each loss and calculate the overall loss;

[0020]

[0021] Among them, α is the weight of edge loss, β is the weight of registration loss, is the edge loss, is the registration loss.

[0022] After obtaining a number of samples, the process further includes preprocessing the samples, wherein the preprocessing includes random sharpening, pixel adjustment, denoising, and enhancement.

[0023] The second technical solution of the present invention is achieved by the following measures: a multimodal remote sensing image registration method based on a dual-branch parallel feature learning network, comprising:

[0024] Acquire a set of remote sensing images to be registered, wherein the set of remote sensing images to be registered includes an optical image and a SAR image of the same area;

[0025] The set of remote sensing images to be registered is input into a multimodal remote sensing image registration model to obtain an image registration result, wherein the multimodal remote sensing image registration model is constructed by a multimodal remote sensing image registration model construction method based on a dual-branch parallel feature learning network.

[0026] The third technical solution of the present invention is achieved by the following measures: a multimodal remote sensing image registration model construction device based on a dual-branch parallel feature learning network, comprising:

[0027] A sample acquisition unit, which acquires a number of samples and divides them into a training set and a test set in proportion, wherein each sample includes an optical image and a SAR image of the same area and an image registration result identifier;

[0028] A training unit, using a training set to train a model built based on a dual-branch parallel feature learning network, introducing a loss function of a weighted strategy during training, and ending the training when the value of the loss function is stable, to obtain a multimodal remote sensing image registration model, wherein the multimodal remote sensing image registration model includes an image feature extraction model and a feature registration model, the image feature extraction model includes an optical branch network and a SAR branch network built based on a dual-branch parallel feature learning network, both of which include a modality adaptive encoder and a hierarchical feature registration module, the modality adaptive encoder extracts local features and global features of an image, and fuses them to form a corresponding comprehensive feature vector, the hierarchical feature registration module enhances the comprehensive feature vector based on an attention mechanism and a residual structure, and outputs a corresponding feature map;

[0029] The testing unit uses the test set to test the trained multimodal remote sensing image registration model, optimizes the model parameters, and outputs a multimodal remote sensing image registration model that meets the test requirements.

[0030] The fourth technical solution of the present invention is achieved by the following measures: a multimodal remote sensing image registration device based on a dual-branch parallel feature learning network, comprising:

[0031] A data acquisition unit acquires a set of remote sensing images to be registered, wherein the set of remote sensing images to be registered includes an optical image and a SAR image of the same area;

[0032] The calibration unit inputs the remote sensing image set to be registered into the multimodal remote sensing image registration model to obtain the image registration result, wherein the multimodal remote sensing image registration model is constructed by a multimodal remote sensing image registration model construction method based on a dual-branch parallel feature learning network.

[0033] The fifth technical solution of the present invention is achieved through the following measures: an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the steps in the multimodal remote sensing image registration method based on a dual-branch parallel feature learning network.

[0034] The sixth technical solution of the present invention is achieved through the following measures: a storage medium, characterized in that a computer program that can be read by a computer is stored on the storage medium, and the computer program is configured to execute the steps in the multimodal remote sensing image registration method based on a dual-branch parallel feature learning network when running.

[0035] The present invention introduces an optical branch network and a SAR branch network constructed based on a dual-branch parallel feature learning network to form an image feature extraction model. Both the optical branch network and the SAR branch network include a modal adaptive encoder and a hierarchical feature registration module. The modal adaptive encoder can extract global features and local features of the image, so that the trained model is no longer limited to local features, thereby improving the accuracy of registration. Furthermore, a hierarchical feature registration module based on an attention mechanism is used to optimize and enhance the fused feature map, which can highlight important features and suppress irrelevant information, and gradually refine the feature matching process using a hierarchical structure to achieve precise alignment of features from coarse to fine. Furthermore, a weighted strategy of edge loss and registration loss is used, wherein the edge loss can ensure that the features of the edge area of ​​the image are fully learned, and the registration loss ensures the consistency and accuracy of the global features. The model training is completed by balancing the edge loss and the registration loss, thereby further improving the accuracy and robustness of the registration using the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Attached Figure 1 A schematic diagram of an implementation environment provided for an embodiment of the present invention.

[0037] Attached Figure 2 A schematic flow chart of a method for constructing a multimodal remote sensing image registration model provided by one embodiment of the present invention.

[0038] Attached Figure 3 A schematic diagram of a model training framework provided for an embodiment of the present invention.

[0039] Attached Figure 4A schematic diagram of a registration result of a group of optical images and SAR images provided by an embodiment of the present invention.

[0040] Attached Figure 5 A schematic diagram of a multimodal remote sensing image registration method flow chart provided for one embodiment of the present invention.

[0041] Attached Figure 6 A schematic diagram of the structure of a multimodal remote sensing image registration model building device provided by one embodiment of the present invention.

[0042] Attached Figure 7 A schematic diagram of the structure of a multimodal remote sensing image registration device provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0043] The present invention is not limited by the following embodiments, and specific implementation methods can be determined based on the technical solution of the present invention and actual conditions.

[0044] Those skilled in the art will understand that, unless otherwise stated, in the embodiments of the present invention, a "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or memory) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0045] In addition, the term “plurality” in the embodiments of the present invention refers to two or more than two, and “first” and “second” are used to distinguish descriptions and should not be understood as implying relative importance.

[0046] Currently, existing remote sensing image registration methods based on deep learning can usually only extract local features for registration, and cannot simultaneously extract the global and local features of optical images and SAR images. Therefore, they cannot fully consider the feature differences between the two images. There are problems such as loss of discriminant information and degradation of matching performance, which can easily lead to limited registration accuracy.

[0047] The embodiment of the present invention provides a multimodal remote sensing image registration method based on a dual-branch parallel feature learning network, obtaining a set of remote sensing images to be registered, wherein the remote sensing image set to be registered includes an optical image and a SAR image of the same area; the remote sensing image set to be registered is input into a multimodal remote sensing image registration model to obtain an image registration result, wherein the multimodal remote sensing image registration model is obtained by training a model constructed based on a dual-branch parallel feature learning network using a training set, wherein the multimodal remote sensing image registration model includes an image feature extraction model and a feature registration model, the image feature extraction model includes an optical branch network and a SAR branch network constructed based on the dual-branch parallel feature learning network, both of which include a modality adaptive encoder and a hierarchical feature registration module, the modality adaptive encoder extracts local features and global features of an image, and forms a corresponding comprehensive feature vector after fusion, and the hierarchical feature registration module enhances the comprehensive feature vector based on an attention mechanism and a residual structure, and outputs a corresponding feature map.

[0048] Among them, the method provided by the embodiment of the present invention may involve artificial intelligence (AI) technology, and can be implemented based on artificial intelligence technology, for example, using a deep learning approach and using sample training to obtain a corresponding model.

[0049] Among them, Machine Learning (ML) is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence.

[0050] Among them, Deep Learning (DL) specifically refers to machine learning based on deep neural network models and methods. It is developed based on statistical machine learning, artificial neural network and other algorithm models, combined with the development of contemporary big data and large computing power. The most important technical feature of deep learning is the ability to automatically extract features.

[0051] The above-mentioned machine learning and deep learning usually include technologies such as neural networks, belief networks, reinforcement learning, and transfer learning.

[0052] Among them, the loss function is in the process of training the neural network, because we hope that the output of the neural network is as close as possible to the value we really want to predict, so we can compare the predicted value of the current network with the target value, and then update the weight vector of each layer of the neural network according to the difference between the two (there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the neural network), until the neural network can predict the target value or a value very close to the target value. Therefore, in deep learning, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function.

[0053] As attached Figure 1 As shown, it shows a schematic diagram of an implementation environment provided by an embodiment of the present invention. The implementation environment may include: training equipment and use equipment.

[0054] Both the training device and the use device are computer devices; optionally, the computer device is a terminal device, such as a mobile phone, a tablet computer, a PC (Personal Computer) and other electronic devices; or, the computer device is a server, which can be a single server, a server cluster composed of multiple servers, or a cloud computing service center, which is not limited to the embodiments of the present invention.

[0055] The training device refers to a computer device that has the ability to train and learn a model built based on a dual-branch parallel feature learning network. Optionally, the training device has the ability to obtain the model, and train and learn it according to application requirements to obtain a multimodal remote sensing image registration model. For example, the training device obtains the model from other devices through the network, and then trains it through training samples according to application requirements so that the model has the ability to obtain image registration results; Optionally, the training device has the ability to build the model, and it can build the model by itself according to application requirements, and then train and learn it. For example, in order to achieve image registration results based on optical images and SAR images of the same area, the training device builds the model by itself, and then trains and learns it through samples according to application requirements.

[0056] The device used refers to a computer device that has the need to use the multimodal remote sensing image registration model. Optionally, the device uses the model according to application requirements and obtains the model from other devices through a network. For example, if the device uses the need to obtain image registration results, it can obtain the multimodal remote sensing image registration model that has been trained and learned to achieve image registration from other devices through a network, and use the multimodal remote sensing image registration model for image registration.

[0057] Based on this, the technical solution of the present invention will be introduced and explained in combination with several examples below.

[0058] Embodiment 1: As attached Figure 2 As shown, the embodiment of the present invention discloses a method for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network, comprising:

[0059] Step S110, obtaining a number of samples, and dividing them into a training set and a test set in proportion, wherein each sample includes an optical image and a SAR image of the same area and an image registration result identifier.

[0060] It should be noted that after obtaining a number of samples, the samples can also be preprocessed, where the preprocessing includes random sharpening, pixel adjustment, denoising, and enhancement, as follows:

[0061] The optical image and SAR image in each sample are randomly sharpened to enhance their detail contrast;

[0062] The size of the optical image and SAR image in each sample is uniformly adjusted to 256x256 pixels to meet the specific input requirements;

[0063] Gaussian blur algorithm is used to smooth the optical image and SAR image in each sample to reduce noise;

[0064] A random rotation operation is performed on the optical image and SAR image in each sample to increase the diversity of samples and improve the robustness of the model to changes in image orientation. The random rotation operation includes horizontal flipping and vertical flipping.

[0065] It should also be noted that the image registration result identifier may include a key point matching result map and a visual registration result map.

[0066] Step S120, using the training set to train the model constructed based on the dual-branch parallel feature learning network, introducing a loss function of a weighted strategy during training, and ending the training when the value of the loss function is stable, to obtain a multimodal remote sensing image registration model, wherein the multimodal remote sensing image registration model includes an image feature extraction model and a feature registration model, the image feature extraction model includes an optical branch network and a SAR branch network constructed based on the dual-branch parallel feature learning network, both of which include a modal adaptive encoder and a hierarchical feature registration module, the modal adaptive encoder extracts local features and global features of the image, and fuses them to form a corresponding comprehensive feature vector, the hierarchical feature registration module enhances the comprehensive feature vector based on the attention mechanism and the residual structure, and outputs the corresponding feature map.

[0067] Step S130, using the test set to test the trained multimodal remote sensing image registration model, optimizing the model parameters, and outputting a multimodal remote sensing image registration model that meets the test requirements.

[0068] It should be noted that in the above training process, the Adam optimizer is used to update the model parameters to accelerate the convergence process and improve the optimization efficiency. The cosine annealing learning rate scheduler is used to dynamically adjust the learning rate during the training process to avoid skipping the optimal solution or falling into the local optimum when approaching the optimal solution.

[0069] The present invention introduces an optical branch network and a SAR branch network constructed based on a dual-branch parallel feature learning network to form an image feature extraction model. The optical branch network and the SAR branch network both include a modal adaptive encoder and a hierarchical feature registration module. The modal adaptive encoder can extract global and local features of the image so that the trained model is no longer limited to local features. The hierarchical feature registration module can optimize and enhance features and improve the accuracy of registration.

[0070] Embodiment 2: As attached Figure 3 As shown, the embodiment of the present invention is a one-step optimization of the above embodiment, wherein the modal adaptive encoder includes an adaptive feature extraction module, a Swin Transformer encoding module and a multi-scale feature fusion module;

[0071] (I) Adaptive Feature Extraction Module (SAFM)

[0072] The adaptive feature extraction module (SAFM) extracts local features of the image, including multiple layers of basic convolution blocks (Blk) and introduces residual connections. The basic convolution block (Blk) is as follows:

[0073] Blk(x)=ReLU(BN(Conv(x))

[0074] Among them, Conv is the convolution operation, BN is batch normalization, and ReLU is the activation function.

[0075] The corresponding adaptive feature extraction module extracts local features as follows:

[0076]

[0077] Among them, i is the index value of the current layer, Blk i is the i-th convolutional block.

[0078] (II) Swin Transformer encoding module

[0079] The Swin Transformer encoding module extracts the global features of the image based on residual connections and attention mechanisms. Specifically, the input image is divided into multiple non-overlapping windows, and the self-attention mechanism is used to perform calculations in each window to extract local features. Then, the local features are fused into global features through connection operations between windows. The calculation formula is as follows:

[0080] x1 = norm(flatten(Conv(x)))

[0081] x1=x1+drop_out(attn(norm(x1)))

[0082] x′1=x1+drop_out(mlp(norm(x1)))

[0083] x″1=Conv(permute(x′1))

[0084] x2 = norm(flatten(Conv(x)))

[0085] x2=x2+drop_out(attn(norm(x2)))

[0086] x′2=x2+drop_out(mlp(norm(x2)))

[0087] x″2=Conv(permute(x′2))

[0088] G S / O =x fuse =concat(x″1,x″2)

[0089] Among them, Conv is the convolution operation, flatten is the flattening of the feature map, the norm operation is used for layer normalization, attn is used to calculate the attention information, drop_out is used for path discarding, and mlp is used to learn the feature information using a multi-layer perceptron.

[0090] (III) Multi-scale feature fusion module

[0091] The multi-scale feature fusion module performs multi-scale fusion on the local features and global features of the image and outputs the corresponding feature map. Specifically, feature maps of different scales are first generated in the Swin Transformer encoding module, and then the feature maps are integrated into the fused feature representation through the residual fusion operation, which not only retains the detailed information of the image, but also enhances the model's adaptability to complex scenes. The calculation formula is as follows:

[0092] f=linear_fuse(G S / O ,L S / O )

[0093] It should be noted that the parameters of the adaptive feature extraction modules in the optical branch network and the SAR branch network are not shared, but the parameters of the Swin Transformer encoding module are shared.

[0094] Embodiment 3: As attached Figure 3 As shown, the embodiment of the present invention is a one-step optimization of the above embodiment, wherein the hierarchical feature registration module sequentially performs dimension adjustment, average pooling, and feature enhancement processing on the input feature map based on the convolution block, attention mechanism, and residual structure, as follows:

[0095] (1) The convolution block (Conv) is used to adjust the dimension of the image feature map output by the modality adaptive encoder to facilitate subsequent calculations. The calculation process is as follows:

[0096] x = BN(Conv(f));

[0097] (2) A hierarchical global-local attention mechanism is used to extract local features. Then, a global attention score is calculated to measure the importance of each local feature. Global context information is captured by average pooling in the horizontal and vertical directions. The calculation process is as follows:

[0098] x=x+(GL_Attn(BN(x)))

[0099] x GL =x+(MLP(BN(x)));

[0100] (3) The residual structure is used to add the x″1 of the Swin Transformer encoding module to the feature information obtained in the first layer to further enhance the expressive power of the feature representation. The calculation process is as follows:

[0101]

[0102] x = ReLU(BN(Conv(x))

[0103] Embodiment 4: As attached Figure 3 As shown, the embodiment of the present invention is a one-step optimization of the above embodiment, wherein the feature registration model performs feature registration based on feature point similarity evaluation and local block cosine similarity evaluation, specifically including:

[0104] Step S410, for the optical feature map and the SAR feature map output by the image feature extraction model, calculate the feature point similarity score between the optical feature and the SAR feature, perform key point matching, and obtain a corresponding matching result map;

[0105] Step S420, segmenting the optical image and the SAR image based on the feature point similarity scores, and performing transposition and normalization operations on the segmented local blocks;

[0106] Step S430, calculating the cosine similarity between the local block of the optical image and the local block of the SAR image, performing local block registration, and obtaining a corresponding visual registration result image.

[0107] Embodiment 5: The embodiment of the present invention is a one-step optimization of the above embodiment, wherein the loss function calculation process of the weighted strategy includes:

[0108] Step S510, determining the edge loss and registration loss in the sample training process.

[0109] The specific edge loss and registration loss calculation process is as follows:

[0110] (1) Edge loss

[0111] First, a mask is generated based on the edge information of the feature map to emphasize or suppress the edge area in the image. Then, the edge loss is calculated to penalize the feature points with high weights at the edge of the feature map, and the mask mean is updated in the iterative process for normalization. The calculation of the edge loss is as follows:

[0112]

[0113] Among them, mean is the mean processing, x S With x O is the SAR feature map and optical feature map after boundary mask filtering, mask S With mask O They are masks generated using the edge information of the SAR feature map and the optical feature map respectively.

[0114] (2) Registration loss

[0115] For the optical feature map and SAR feature map output by the image feature extraction model, the feature point similarity score between the optical feature and the SAR feature is calculated. The calculation process is as follows:

[0116]

[0117] Among them, N is the number of feature points, D is the feature dimension, is the ith eigenvalue of the nth sample in the SAR feature, For Remapping was performed (nearest neighbor interpolation and boundary alignment), is the i-th eigenvalue of the n-th sample in the optical feature, and M is a mask.

[0118] The optical image and SAR image are segmented based on the feature point similarity score, and the segmented local blocks are transposed and normalized. The calculation process is as follows:

[0119] Patches s =norm(transpose(unfold(f S )))

[0120] Patches O =norm(transpose(unfold(f O )))

[0121] patches_simi=transpose(unfold(feat_wise_point_simi))

[0122] Among them, norm() is the normalization operation, transpose() is the transposition operation, and unfold() is the local block segmentation operation.

[0123] The cosine similarity between the local blocks of the optical image and the local blocks of the SAR image is calculated. The calculation process is as follows:

[0124]

[0125] patches_cos=cosim(patches S ,patches O )

[0126] Among them, cosim() is the cosine similarity operation, and clamp() is the numerical restriction operation.

[0127] Determine the registration loss, the calculation process is as follows:

[0128]

[0129] Step S520, determine the weight of each loss and calculate the overall loss;

[0130]

[0131] Among them, α is the weight of edge loss, β is the weight of registration loss, is the edge loss, is the registration loss.

[0132] Embodiment 6: As attached Figure 4 As shown, an optical image and a SAR image of the same area are obtained, and then a multimodal remote sensing image registration model constructed by a multimodal remote sensing image registration model construction method based on a dual-branch parallel feature learning network disclosed in the present invention is used to perform multimodal remote sensing image registration, and the attached Figure 4 From the key point matching result diagram and the visual registration result diagram, it can be seen that the registration result is accurate.

[0133] Embodiment 7: As attached Figure 5As shown, the embodiment of the present invention discloses a multimodal remote sensing image registration method based on a dual-branch parallel feature learning network, comprising:

[0134] Step S710, obtaining a set of remote sensing images to be registered, wherein the set of remote sensing images to be registered includes an optical image and a SAR image of the same area;

[0135] Step S720, inputting the remote sensing image set to be registered into a multimodal remote sensing image registration model to obtain an image registration result, wherein the multimodal remote sensing image registration model is constructed by the method described in Examples 1 to 6.

[0136] Embodiment 8: As attached Figure 6 As shown, the embodiment of the present invention discloses a multimodal remote sensing image registration model construction device based on a dual-branch parallel feature learning network, comprising:

[0137] A sample acquisition unit, which acquires a number of samples and divides them into a training set and a test set in proportion, wherein each sample includes an optical image and a SAR image of the same area and an image registration result identifier;

[0138] A training unit, using a training set to train a model built based on a dual-branch parallel feature learning network, introducing a loss function of a weighted strategy during training, and ending the training when the value of the loss function is stable, to obtain a multimodal remote sensing image registration model, wherein the multimodal remote sensing image registration model includes an image feature extraction model and a feature registration model, the image feature extraction model includes an optical branch network and a SAR branch network built based on a dual-branch parallel feature learning network, both of which include a modality adaptive encoder and a hierarchical feature registration module, the modality adaptive encoder extracts local features and global features of an image, and fuses them to form a corresponding comprehensive feature vector, the hierarchical feature registration module enhances the comprehensive feature vector based on an attention mechanism and a residual structure, and outputs a corresponding feature map;

[0139] The testing unit uses the test set to test the trained multimodal remote sensing image registration model, optimizes the model parameters, and outputs a multimodal remote sensing image registration model that meets the test requirements.

[0140] The specific implementation steps of each of the above units are the same as those described in Examples 1 to 6 and will not be repeated here.

[0141] Embodiment 9: As attached Figure 7 As shown, the embodiment of the present invention discloses a multimodal remote sensing image registration device based on a dual-branch parallel feature learning network, comprising:

[0142] A data acquisition unit acquires a set of remote sensing images to be registered, wherein the set of remote sensing images to be registered includes an optical image and a SAR image of the same area;

[0143] The calibration unit inputs the remote sensing image set to be registered into the multimodal remote sensing image registration model to obtain an image registration result, wherein the multimodal remote sensing image registration model is constructed by the method described in Examples 1 to 6.

[0144] Embodiment 10: The embodiment of the present invention discloses a storage medium, on which is stored a computer program that can be read by a computer, and the computer program is configured to execute a multimodal remote sensing image registration device based on a dual-branch parallel feature learning network when running.

[0145] The above storage medium may include, but is not limited to: a USB flash drive, a read-only memory, a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.

[0146] Embodiment 11: The embodiment of the present invention discloses an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement a multimodal remote sensing image registration device based on a dual-branch parallel feature learning network.

[0147] The processor may be a central processing unit (CPU), a general purpose processor, a digital signal processor (DSP), an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. It may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The memory may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory, a mobile hard disk, a magnetic disk or an optical disk.

[0148] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present invention may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0149] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0150] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0151] The above content is only a specific implementation mode of the present invention, which has strong adaptability and implementation effect, but the protection scope of the present invention is not limited to this. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered within the protection scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope covered by the present invention.

Claims

1. A method for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network, characterized in that: include: A number of samples are obtained and divided into a training set and a test set in proportion, wherein each sample includes an optical image and a SAR image of the same area and an image registration result identifier; The model constructed based on the dual-branch parallel feature learning network is trained using the training set. A loss function of a weighted strategy is introduced during the training. When the value of the loss function is stable, the training is terminated to obtain a multimodal remote sensing image registration model, wherein the multimodal remote sensing image registration model includes an image feature extraction model and a feature registration model. The image feature extraction model includes an optical branch network and a SAR branch network constructed based on the dual-branch parallel feature learning network. Both include a modality adaptive encoder and a hierarchical feature registration module. The modality adaptive encoder extracts local features and global features of the image and fuses them to form a corresponding comprehensive feature vector. The hierarchical feature registration module enhances the comprehensive feature vector based on the attention mechanism and the residual structure and outputs the corresponding feature map. The trained multimodal remote sensing image registration model is tested using the test set, the model parameters are optimized, and a multimodal remote sensing image registration model that meets the test requirements is output.

2. The method for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network according to claim 1, characterized in that: The modality adaptive encoder includes an adaptive feature extraction module, a Swin Transformer encoding module and a multi-scale feature fusion module. The adaptive feature extraction module uses multiple basic convolution blocks to extract local features of the image. The Swin Transformer encoding module extracts global features of the image based on residual connections and attention mechanisms. The multi-scale feature fusion module performs multi-scale fusion on local and global features of the image and outputs the corresponding feature map. or / and, The hierarchical feature registration module performs dimension adjustment, average pooling, and feature enhancement on the input feature map based on convolution blocks, attention mechanisms, and residual structures.

3. The method for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network according to claim 1 or 2, characterized in that: The feature registration model performs feature registration based on feature point similarity evaluation and local block cosine similarity evaluation.

4. The method for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network according to claim 1 or 2, characterized in that: The loss function calculation process of the weighted strategy includes: Determine the edge loss and registration loss during sample training; Determine the weight of each loss and calculate the overall loss; Among them, α is the weight of edge loss, β is the weight of registration loss, is the edge loss, is the registration loss.

5. The method for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network according to any one of claims 1 to 4, characterized in that: After obtaining a number of samples, the samples are preprocessed, wherein the preprocessing includes random sharpening, pixel adjustment, denoising, and enhancement.

6. A multimodal remote sensing image registration method based on a dual-branch parallel feature learning network, characterized in that: include: Acquire a set of remote sensing images to be registered, wherein the set of remote sensing images to be registered includes an optical image and a SAR image of the same area; The remote sensing image set to be registered is input into a multimodal remote sensing image registration model to obtain an image registration result, wherein the multimodal remote sensing image registration model is constructed by the method described in any one of claims 1 to 5.

7. A device for constructing a multimodal remote sensing image registration model based on a dual-branch parallel feature learning network using the method as described in any one of claims 1 to 5, characterized in that: include: A sample acquisition unit, which acquires a number of samples and divides them into a training set and a test set in proportion, wherein each sample includes an optical image and a SAR image of the same area and an image registration result identifier; A training unit, using a training set to train a model built based on a dual-branch parallel feature learning network, introducing a loss function of a weighted strategy during training, and ending the training when the value of the loss function is stable, to obtain a multimodal remote sensing image registration model, wherein the multimodal remote sensing image registration model includes an image feature extraction model and a feature registration model, the image feature extraction model includes an optical branch network and a SAR branch network built based on a dual-branch parallel feature learning network, both of which include a modality adaptive encoder and a hierarchical feature registration module, the modality adaptive encoder extracts local features and global features of an image, and fuses them to form a corresponding comprehensive feature vector, the hierarchical feature registration module enhances the comprehensive feature vector based on an attention mechanism and a residual structure, and outputs a corresponding feature map; The testing unit uses the test set to test the trained multimodal remote sensing image registration model, optimizes the model parameters, and outputs a multimodal remote sensing image registration model that meets the test requirements.

8. A multimodal remote sensing image registration device based on a dual-branch parallel feature learning network using the method as claimed in claim 6, characterized in that: include: A data acquisition unit acquires a set of remote sensing images to be registered, wherein the set of remote sensing images to be registered includes an optical image and a SAR image of the same area; The calibration unit inputs the remote sensing image set to be registered into a multimodal remote sensing image registration model to obtain an image registration result, wherein the multimodal remote sensing image registration model is constructed by the method described in any one of claims 1 to 5.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the steps in the method according to claim 6.

10. A storage medium, characterized in that: The storage medium stores a computer program that can be read by a computer, and the computer program is configured to execute the steps of the method according to claim 6 when running.

Citation Information

Cited By

  • Optical-SAR image automatic registration system and method

    CN121391944A

  • Automatic registration system and method for optical-SAR images

    CN121391944B