Large-scale multi-modal image registration method and system driven by structured enhancement

By introducing detector-free framework and structural enhancement module in multimodal image registration, a robust reference anchor point is established and a many-to-many matching strategy is adopted, the problems of nonlinear radiation differences and geometric distortion in multimodal image registration are solved, and high-precision and robust image registration effect are achieved.

CN120047505AActive Publication Date: 2025-05-27WUHAN UNIV

Patent Information

Application Number
CN202510533633.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Existing multimodal image registration methods are difficult to effectively match feature points when facing nonlinear radiation differences and geometric distortions, especially in scenarios where textures are scarce or scale differences, resulting in insufficient matching accuracy and robustness.

Method used

A large-scale multi-modal image registration method driven by structured enhancement is proposed. By introducing a detector-free framework, directly searching for corresponding pixels between images, establishing a robust reference anchor point to enhance structural information, and adopting a many-to-many matching strategy to make up for scale differences.

Benefits of technology

It effectively alleviates the problem of feature point offset, enhances the structural information of weak textures and large modal differences in multimodal images, and improves the accuracy and robustness of image registration, especially to achieve high-quality pixel-level matching in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047505A_ABST
    Figure CN120047505A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale multi-modal image registration method and system based on structured enhancement driving, and belongs to the field of image registration, and the method comprises the steps: collecting images of different modals, and obtaining a to-be-registered image data set which comprises a to-be-registered image and a reference image; randomly selecting two images with different modalities from the to-be-registered image data set, and extracting two feature map pairs with different downsampling resolutions; obtaining a rough feature map and a reference anchor point for the feature map pair with one resolution; performing structure enhancement on the rough feature map by using a reference anchor point to obtain an enhanced feature map; calculating a matching probability between the enhanced feature maps, and obtaining a rough matching result; obtaining a fine matching result by combining the rough matching result with the feature map pair of the other resolution; and calculating a transformation matrix of the image by using the fine matching result, transforming the to-be-registered image, and carrying out sub-pixel level alignment on the to-be-registered image and the reference image to obtain a final registration result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a structured enhancement driven large-scale multimodal image registration method and system. Background Art

[0002] Image registration is a core challenge in visual understanding and interpretation. It plays a vital role in the fields of remote sensing and computer vision and is widely used in tasks such as image stitching, image fusion, and change detection. However, single-modal data often lacks sufficient depth and is difficult to meet the needs of remote sensing and computer vision applications. Therefore, it is particularly important to use data from different modalities, which can not only give play to the complementary advantages of each modality, but also help achieve a comprehensive understanding of the image. Therefore, how to effectively integrate and deeply analyze multi-sensor, multi-resolution and multi-temporal data has become a key direction of current research. In this context, multimodal image matching has become one of the core issues that need to be solved urgently.

[0003] Multimodal image registration not only involves complex geometric differences (such as scale, rotation, etc.) that are common in single-modal matching, but its main difficulty lies in the nonlinear radiometric differences (NRD) between images. NRD usually originates from different imaging mechanisms or external environmental factors. For example, optical images and synthetic aperture radar (SAR) images present the same scene differently, resulting in significant radiometric differences between image pairs. These differences are usually more obvious in texture-rich areas, which are ideal locations for feature point extraction. Therefore, NRD hinders the effective matching of features between images and increases the complexity and challenge of the image matching task.

[0004] In recent years, many methods have been actively devoted to solving this problem, which can be divided into two categories: region-based methods and feature-based methods. These methods have been shown to be able to extract representative and sufficiently repeatable feature points in the face of geometric distortion and NRD. However, the study of these multimodal matching methods has revealed three major problems that cannot be ignored: (1) In order to obtain high-precision matching relationships, existing methods often impose a mandatory one-to-one constraint on the matching process. Although this constraint helps to improve the accuracy of matching, when the scale difference between the image pairs is large, it is often limited to the smaller scale image, resulting in a significant reduction in the number of matching points that can be extracted.

[0005] (2) Existing methods mainly rely on common structural features (such as object edges) between images of different modalities to extract feature points. However, in scenes with scarce textures or sparse common structures, effectively extracting feature points remains a huge challenge.

[0006] (3) When extracting feature points from a single image, the problem of feature point offset is likely to occur. Due to the differences in texture details of different modality images, the feature points in an image may not be precisely matched with those in another image. Instead, adjacent points with similar descriptions may be misidentified as matching points, leading to offset, and the existence of NRD further exacerbates this problem. Summary of the Invention

[0007] In view of the above problems, the present invention proposes a large-scale multi-modal image registration method driven by structural enhancement. First, a detector-free framework (LoFTR) is introduced to alleviate the feature point offset problem by directly finding corresponding pixels between images. Subsequently, to solve the NRD and geometric distortion problems in multi-modal images, a structural enhancement module is designed to establish robust reference anchors to enhance the structural information of regions with significant modal differences and enrich the corresponding pixels in such regions. In addition, to make up for the problem of fewer matching points between images with large scale differences, the present invention removes the one-to-one matching constraint in the original detector-free framework and adopts a many-to-many matching strategy. Finally, more precise positioning of highly correlated corresponding pixels is performed to obtain an optimized matching relationship, and an image transformation matrix is estimated to perform image registration on multi-modal images.

[0008] The purpose of the present invention is to propose a large-scale multi-modal image registration method driven by structural enhancement. This method includes the following steps: Collect images of different modalities to obtain a dataset of images to be registered, where the dataset of images to be registered includes images to be registered and reference images; Arbitrarily select two different modality images from the dataset of images to be registered and extract pairs of feature maps with two different downsampling resolutions; For a pair of feature maps at one resolution, obtain a rough feature map and reference anchors; Use the reference anchors to perform structural enhancement on the rough feature map to obtain an enhanced feature map; Calculate the matching probability between the enhanced feature maps and obtain a rough matching result; Combine the rough matching result and the pair of feature maps at the other resolution to obtain a fine matching result; Use the fine matching result to calculate the transformation matrix of the image, transform the image to be registered, and align it with the reference image at the sub-pixel level to obtain the final registration result.

[0009] Furthermore, the images of different modalities include: optical images, synthetic aperture radar images, map images, depth maps, and infrared images; a transformation matrix is randomly generated for each image pair and applied to one of the images to obtain a deformed dataset of images to be registered.

[0010] Furthermore, an existing ResNet network is used to obtain feature map pairs with different downsampling resolutions.

[0011] Furthermore, for a feature map pair with one of the resolutions, the specific implementation of obtaining a rough feature map and reference anchor points is as follows: For a given pair of multimodal images and , and refer to paired images in image registration, representing the sensed image and the reference image respectively; position encoding is added to the feature map at the 1 / 8 scale, and then the position-encoded features are processed through self-attention layers and cross-attention layers to achieve information interaction between different modalities:

[0012] Among them, in the case of the self-attention mechanism, represents a self-attention layer or a cross-attention layer, both have values of or , and in the case of the cross-attention mechanism, is , has a value of , represents a scale factor; the feature map obtains the corresponding after the above processing, , and by multiplying these two flattened feature maps , , a similarity matrix is obtained, where H and W respectively refer to the height and width of image A or B; a bidirectional softmax operation is performed on to obtain a confidence matrix , and pixels with high confidence are selected as reference anchor points: , , , where is a hyperparameter that controls the number of reference anchor points, is an intermediate quantity, MAX means taking the maximum, MIN means taking the minimum, represents the set of reference anchor points, n represents the number of reference anchor points, i represents the serial number of all pixel points in , and j represents the serial number of all pixel points in .

[0013] ​Further, the specific implementation of obtaining the enhanced feature map is as follows: By calculating the rough feature map to the Euclidean distance of the reference anchor points where includes and find the nearest reference anchor point for each pixel point and calculate the angle between it and the nearest reference anchor point :

[0014] where represents the relative position between the reference anchor point closest to point i and point i, represents the relative position between the position of the reference anchor point closest to point j and point j, represents the normalization operation; subsequently, the obtained structural information and the feature map are concatenated together, and a multi-layer perceptron network (MLP) is used to fuse the concatenated information to obtain the fused feature and a cross-attention layer is used to perform information interaction on the feature pairs:

[0015]

[0016] where respectively represent the reference anchor point, the angle and length between the pixel point and the reference anchor point, and the subscript c represents the serial number of the nearest reference anchor point, represents concatenation.

[0017] Further, for the enhanced feature map recalculate the similarity matrix between them and perform softmax operations on the two dimensions of respectively to obtain the matching probabilities at the rough stage: ,

[0018] In the matching probability matrix assuming is a pair of matching points, then in the confidence value is higher than the threshold and higher than other elements in the same row; similarly for , in the confidence value is higher than the threshold and higher than other elements in the same column, so the rough matching prediction is expressed as:

[0019] Among them, represents or the serial number corresponding to the maximum value in is a pair of rough matching points, is the rough matching result.

[0020] Furthermore, for each pair of rough matching points , first locate its corresponding position in the 1 / 2-scale feature map or , and then crop out two local windows with a size of ; then, use a self-attention layer and a cross-attention layer to perform information interaction on the features in each window to obtain two transformed local feature maps or or , centered on and respectively; subsequently, calculate the correlation between the central vector of and all vectors in to generate a heat map, and the values on the heat map represent the matching probability of each pixel around ; by calculating the expectation of the matching probability, obtain the final sub-pixel level accurate position , and aggregate all matching points to obtain the final fine matching result , is the number of matches.

[0021] Furthermore, use the Random Sample Consensus algorithm to estimate the transformation matrix of the image from the fine matching result.

[0022] Furthermore, use the transformation matrix H to transform the image to be registered, and align it with the reference image at the sub-pixel level:

[0023] Among them, is the fine matching result the coordinates of the potential inliers in the image to be registered in is after being transformed by the transformation matrix in the image coordinates.

[0024] The present invention also provides a large-scale multi-modal image registration system driven by structured enhancement, including: A processor and a memory, where the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a structured enhancement-driven large-scale multimodal image registration method as described in the above technical solution.

[0025] The present invention proposes a structured enhancement-driven large-scale multimodal image registration method, which weakens the feature point offset problem by establishing pixel-level correspondence relationships and enhances the correspondence relationships in weakly textured and regions with large texture differences affected by NRD and geometric distortions using reference anchor points. By introducing a detector-free framework, corresponding pixels are directly searched between images to alleviate the feature point offset problem. Subsequently, to solve the NRD and geometric distortion problems in multimodal images, a structure enhancement module is designed to establish robust reference anchor points to enhance the structural information of regions with significant modal differences and enrich the corresponding pixels in such regions. In addition, to make up for the problem of fewer matching points between images with large scale differences, the present invention removes the one-to-one matching constraint in the original detector-free framework and adopts a many-to-many matching strategy. Finally, we perform more precise positioning on highly correlated corresponding pixels to obtain optimized correspondence relationships, estimate the image transformation matrix, and perform image registration on multimodal images. The present invention effectively solves the challenges of feature point offset and NRD in multimodal images by enhancing the structural information of weakly textured and regions with large modal differences. Description of the Drawings

[0026] Figure 1 It is a flowchart of an embodiment of the present invention.

[0027] Figure 2 It is an effect diagram of an embodiment of the present invention. Detailed Embodiment

[0028] The technical solution of the present invention will be further described below with reference to the drawings.

[0029] The technical process of the embodiment of the present invention is shown in Figure 1 , a structured enhancement-driven large-scale multimodal image registration, including the following steps: 1) Multimodal data processing A total of five types of modal data are collected: optical images, SAR images, maps, depth images, and infrared images. In the training stage, we transform images of different modalities by randomly generating transformation matrices and use the transformation matrices as the ground truth of the method.

[0030] 2) Feature extraction We use a standard convolutional architecture with a Feature Pyramid Network (denoted as a local feature Convolutional Neural Network (CNN)) to extract rough features at 1 / 8 scale and fine features at 1 / 2 scale of the original image size from two different modality images. CNN has inductive biases of translational equivariance and locality, which are very suitable for local feature extraction. The downsampling introduced by CNN also reduces the input length of the feature extraction module, which is crucial for ensuring controllable computational costs.

[0031] 3) Coarse correspondence acquisition Position encoding is added to the feature map at 1 / 8 scale. The transformed features will become position-related, which is crucial for the present invention to generate sufficient matches in weakly textured regions. We process the transformed features through self-attention layers and cross-attention layers to achieve information interaction between different modalities. The self-attention layer can capture the dependencies between different positions within the same modality, while the cross-attention layer further enhances the fusion and collaboration of multimodal information by comparing features between different modalities. Through this information interaction mechanism, the model can accurately establish associations between feature points among multimodal images, thereby improving the accuracy and robustness of cross-modal matching.

[0032] On this basis, we further calculate the similarity of each pixel on the 1 / 8-scale feature map to evaluate the matching possibility between different pixels. To further improve the matching accuracy, we use a bi-directional - softmax operation to obtain the confidence matrix of the feature map. This matrix reflects the matching confidence between each pixel and other pixels, so that the most robust reference anchors can be selected according to the confidence. By selecting these high-confidence reference anchors, we can enhance the structural information in the weakly textured regions of the image, making the matching process more stable and accurate. Especially in those regions with scarce texture or large modality differences, the reliability of the reference anchors can effectively improve the overall matching quality.

[0033] 4) Structural enhancement In regions of an image where the texture is weak or there are differences, it is often challenging to confirm the corresponding positions solely through pixel scanning. First, we calculate the Euclidean distance between each pixel point in the 1 / 8 feature map and a pre-selected reference anchor point, and find the nearest reference anchor point for each pixel point based on the distance. In this way, we can determine the relative position of each pixel in the image and further calculate the angle between the pixel and its nearest reference anchor point. Next, we use a multi-layer perceptron (MLP) network to concatenate and fuse the features of each pixel point with the distance and angle information to its nearest reference anchor point. Subsequently, we perform information interaction on the feature map with structured enhancement through a cross-attention layer and obtain a coarse matching matrix. This approach not only enhances the feature description of pixels but also introduces more global information, enabling pixels in regions with weak texture or large modality differences to effectively rely on the relative position relationship with the reference anchor points for matching.

[0034] 5) Adaptive many-to-many matching strategy In the process of obtaining the coarse corresponding matches, we no longer rely on the traditional mutual nearest neighbor strategy to enforce a one-to-one matching constraint to improve accuracy. Instead, after obtaining the confidence matrix using a bidirectional-softmax operation, we consider all matches greater than a fixed threshold as correct matches. In this process, especially on the 1 / 8 scale feature map, we may encounter the situation where multiple pixels in the large-scale feature map correspond to the same pixel in the small-scale feature map. This situation is very common in image pairs with large scale differences, but it does not necessarily mean a mis-match. Therefore, we map this many-to-one matching relationship back to the 1 / 2 scale feature map and perform sub-pixel localization on the many-to-one match, thereby converting it into a one-to-one match. This strategy effectively maximally restores the matching relationship in the large-scale context.

[0035] 6) Obtaining optimized corresponding relationships After establishing the rough matches, we map these matches to the 1 / 2 scale feature map and crop out local windows of a fixed size around each match point. Next, we send these cropped windows into self-attention layers and cross-attention layers for information interaction and feature aggregation. Our goal is to find the exact matching relationship between two corresponding windows, that is, the specific corresponding position of the central pixel point of one window in the other window. Through this exact matching, we can ensure that the positions of the corresponding points reach sub-pixel level accuracy, significantly improving the accuracy of the matching.

[0036] 7) Multi-modal image registration Finally, after obtaining the exact matching relationships between the multi-modal images, we use the Random Sample Consensus (RANSAC) algorithm to estimate the transformation matrix of the images from these hypothesized correspondences. The transformation matrix is used to transform the image to be registered, aligning it with the reference image at the sub-pixel level.

[0037] The following is the specific implementation process of the above steps: 1) As is well known, having sufficient training data is a prerequisite for achieving good performance. However, obtaining the transformation parameters of real image pairs as the ground truth remains a major challenge in the image registration task. To address this issue, we adopted a simple synthetic image pair method to generate the transformation parameters of the pseudo labels. Specifically, we randomly generated a transformation matrix for each image pair and applied it to one of the images to obtain the deformed dataset of the image to be registered. We collected a total of 5 different modalities of images, including optical images, synthetic aperture radar images, map images, depth maps, and infrared images.

[0038] 2) For a given pair of multi-modal images and , and refer to the paired images in image registration, representing the image to be registered and the reference image respectively, and can be of any modality; the existing ResNet network is used to obtain multi-layer features with different downsampling resolutions . In this embodiment, the rough matching of the image is obtained on the feature map at the 1 / 8 scale , and the optimized matching of the image is obtained on the feature map at the 1 / 2 scale . Those skilled in the art can select appropriate scale feature maps for subsequent processing according to actual needs. 3) Add positional encoding to the feature map at the 1 / 8 scale : By applying sine and cosine functions with different frequencies to the pixel points on the feature map, the transformed features will become position-related. Subsequently, the transformed features are processed through the self-attention layer and the cross-attention layer to achieve information interaction between different modalities:

[0039] Among them, in the case of the self-attention mechanism, represents the self-attention layer or the cross-attention layer, values are both or , in the case of the cross-attention mechanism, is , value is , Denote the scale factor. The feature map After the above processing, the corresponding , , by multiplying these two flattened feature maps , , a similarity matrix can be obtained , where H and W respectively refer to the height and width of image A or B. Perform a bidirectional softmax operation on to obtain the confidence matrix , where we select the pixels with high confidence as the reference anchor points: , , , where is a hyperparameter that controls the number of reference anchor points, is an intermediate quantity, MAX means taking the maximum, MIN means taking the minimum, denotes the set of reference anchor points, n denotes the number of reference anchor points, the i-th dimension represents the serial numbers of all pixel points in and the j-th dimension is

[0040] 4) By calculating the Euclidean distance from the rough feature map (including and ) to the reference anchor points , the nearest reference anchor point for each pixel point can be found , and the angle between it and the nearest reference anchor point is calculated :

[0041] where represents the relative position between the reference anchor point closest to point i and point i, represents the relative position between the position of the reference anchor point closest to point j and point j, represents the normalization operation. Subsequently, we splice the obtained structural information and the feature map together, and use a multi-layer perceptron network (MLP) to fuse the spliced information to obtain the fused feature , and use cross-attention to perform information interaction on the feature pairs:

[0042]

[0043] where respectively represent the reference anchor points, the angle and length between the pixel points and the reference anchor points, c represents the serial number of the nearest reference anchor point, denotes stitching. Thus, with the help of the reference anchor points, the structure enhancement of the regions with weak texture or large modal differences is completed, and the corresponding relationships of such regions are restored.

[0044] 5) For the updated feature maps , recalculate the similarity matrix between them , and perform softmax operations on the two dimensions of respectively to obtain the matching probabilities at the coarse stage: ,

[0045] From the matching probability matrix , assuming is a pair of matching points, then 's confidence value in should be higher than the threshold and higher than other elements in the same row; in this matrix, the i dimension represents the serial numbers of all pixel points in , and the j dimension is the serial numbers of all pixel points in , so the correspondence between i and j is the correspondence of pixel points in and . Similarly, for , we select the index higher than the threshold in the column and the remaining elements as the matching prediction. We represent the coarse matching prediction as:

[0046] Among them, represents the serial number corresponding to the maximum value in or , is a pair of coarse matching points, and the coarse matching result lifts the restriction of the one-to-one matching strategy and adopts the many-to-many matching strategy to make up for the corresponding relationships of image pairs with large scale differences to the greatest extent.

[0047] 6) After establishing the coarse matching , these matches will be mapped back to the resolution of the original image through coarse-to-fine processing. For each pair of coarse matching points , we first locate their corresponding positions or in the 1 / 2 scale feature map , and then crop out two patches of size The local window. Then, an self-attention layer and a cross-attention layer are used to perform information interaction on the features in each window, obtaining two transformed local feature maps or , centered at and respectively. Subsequently, we calculate the correlation between the center vector of and all the vectors in to generate a heat map, where the values on the heat map represent the matching probability between each pixel around and . By calculating the expectation of the matching probability, we obtain the final sub-pixel accurate position , and aggregate all the matching points to get the final fine matching result , being the number of matches.

[0048] 7) Finally, after obtaining the fine matching between the multi-modal images , the random sample consensus algorithm is used to estimate the transformation matrix of the image from these assumed correspondences. The transformation matrix is used to transform the image to be registered, aligning it with the reference image at the sub-pixel level:

[0049] where is the coordinate of the potential inlier in the image to be registered in , while is the coordinate in the image after being transformed by the transformation matrix in .

[0050] The effects of the present invention are illustrated by specific data below. The multi-modal data used in the embodiments of the present invention all come from public data sets: Optical-SAR image registration: This data set contains pairs of aligned grayscale images and SAR images obtained from satellites, covering rural and urban landscapes. The data set includes 2011 pairs of images for training and 424 pairs of images for testing.

[0051] Optical-Infrared image registration: This data set mainly comes from the LLVIP data set, which is taken by a stereo camera at dusk and contains 26 different street locations. The image content mainly includes a large number of pedestrians, vehicles and cyclists.

[0052] Optical - Map Image Registration: This dataset is divided into two parts. The first part includes 200 pairs of image pairs, focusing on buildings and streets. The second part includes 330 pairs of image pairs, directly collected from OpenStreetMap (OSM), covering farmland and buildings. The image pairs from OSM are used for the training set, while the image pairs in the first part are used for the test set.

[0053] Optical - Depth Image Registration: This dataset is provided by the DIML / CVL RGB - D dataset, including indoor data built using Microsoft Kinect v2 and outdoor data built using stereo cameras. It mainly contains specific scenes, including classrooms, libraries, hospitals, parks, streets, etc., and was taken from 2015 to 2017.

[0054] Shown in Figure 2 are the visualization results of the present invention on 4 datasets. The first column is the image pairs to be registered in different modalities. The left image in the first column is the sensed image, i.e., the image to be registered, and the right image is the reference image. These image pairs have varying degrees of rotation and scale changes; the second column is the feature matching result, and the yellow line represents the fine matches estimated by the algorithm ; the third column is the multi - modal image registration result. The sensed image undergoes a transformation matrix estimated by fine matching to obtain the registered image.

[0055] In summary, by proposing a structured enhancement - driven large - scale multi - modal image registration method, the present invention effectively improves the accuracy and robustness of image registration, especially in regions with weak texture and large modal differences. By establishing reliable reference anchor points to provide a robust topological structure for pixel points in regions with weak texture and large modal differences, the present invention can achieve high - quality pixel - level matching in complex scenes, thus promoting the wide application of multi - modal image analysis technology. This technological progress not only reduces the dependence on traditional feature detectors, thereby reducing the cost of hardware and computing resources, but also provides more efficient and accurate solutions for industries such as remote sensing, medical imaging, and autonomous driving, promoting the development of related industries.

[0056] On the other hand, the embodiments of the present invention also provide a structured enhancement - driven large - scale multi - modal image registration system, including: A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a structured enhancement - driven large - scale multi - modal image registration method as described in the above technical solution.

[0057] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains may make various modifications or supplements to the described specific embodiments or use similar means for substitution, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A structured enhancement driven large-scale multimodal image registration method, characterized in that: include: Collecting images of different modalities to obtain a dataset of images to be registered, wherein the dataset of images to be registered includes a plurality of image pairs consisting of images to be registered and reference images; For any pair of images, extract two feature map pairs with different downsampling resolutions; For a pair of feature maps at one of the resolutions, obtain the corresponding coarse feature map and reference anchor point respectively; Use the reference anchor point to perform structural enhancement on the coarse feature map to obtain the enhanced feature map; Calculate the matching probability between the enhanced feature maps and obtain the rough matching result; Combine the coarse matching result with a feature map pair of another resolution to obtain a fine matching result; The image transformation matrix is ​​calculated using the fine matching results, the image to be registered is transformed, and the image is aligned with the reference image at the sub-pixel level to obtain the final registration result.

2. The structured enhancement driven large-scale multimodal image registration method according to claim 1, characterized in that: Images of different modalities include: optical images, synthetic aperture radar images, map images, depth maps, and infrared images. A transformation matrix is ​​randomly generated for each image pair and applied to one of the images to obtain a deformed image dataset to be registered.

3. The structured enhancement driven large-scale multimodal image registration method according to claim 1, characterized in that: An existing ResNet network is used to obtain feature map pairs with different downsampled resolutions.

4. The structured enhancement driven large-scale multimodal image registration method according to claim 1, characterized in that: For a feature map pair of one resolution, the specific implementation method of obtaining a coarse feature map and a reference anchor point is as follows: For a given multimodal image pair and , where A and B are the numbers of the image to be registered and the reference image respectively; in the feature map of 1 / 8 scale Position encoding is added to the image, and then the position encoded features are processed through the self-attention layer and the cross-attention layer to achieve information interaction between different modalities: in, Represents a self-attention layer or a cross-attention layer. In the case of a self-attention mechanism, The values ​​of or , in the case of the cross-attention mechanism, for , The value of , represents the scale factor; Feature Map After the above processing, the corresponding , , by flattening these two feature maps , Multiply them together to get the similarity matrix , H and W refer to the height and width of image A or B respectively; Perform a bidirectional softmax operation to obtain the confidence matrix , select pixels with high confidence as reference anchor points: , , , in is a hyperparameter that controls the number of reference anchors, It is the intermediate quantity. MAX means taking the maximum value, and MIN means taking the minimum value. represents the reference anchor point set, n represents the number of reference anchor points, i represents The serial numbers of all pixels in The serial numbers of all pixels in .

5. The structured enhancement driven large-scale multimodal image registration method according to claim 4, characterized in that: The specific implementation method of the enhanced feature map is as follows: By calculating the Euclidean distance from the coarse feature map to the reference anchor point , find the nearest reference anchor point for each pixel , and calculate the angle between the pixel and the nearest reference anchor point : in Indicates the relative position between the reference anchor point closest to point i and point i, Indicates the relative position between the reference anchor point closest to point j and point j. Represents the normalization operation; then the obtained structural information and feature map are spliced ​​together, and the spliced ​​information is fused using a multi-layer perceptron network to obtain the fused features , and use cross attention to exchange information between feature pairs: in Respectively represent the reference anchor point, the angle and length between the pixel point and the reference anchor point. The subscript c represents the serial number of the nearest reference anchor point. Indicates splicing.

6. The structured enhancement driven large-scale multimodal image registration method according to claim 1, characterized in that: For the enhanced feature maps, recalculate the similarity matrix between them , and The two dimensions of are softmaxed to obtain the matching probability of the coarse stage: , In the matching probability matrix In, assuming is a pair of matching points, then exist The confidence value in is higher than the threshold And higher than other elements in the same row; similarly for , exist The confidence value in is higher than the threshold And higher than other elements in the same column, so the rough matching prediction is expressed as: in, express or The sequence number corresponding to the maximum value in , is a pair of rough matching points, This is the rough matching result.

7. The structured enhancement driven large-scale multimodal image registration method according to claim 1, characterized in that: For each pair of coarse matching points in the coarse matching result , first in the 1 / 2 scale feature map or Locate its corresponding position , and then cut out two sizes of A and B are the numbers of the image to be registered and the reference image, respectively. w is the size of the local window; then, a self-attention layer and a cross-attention layer are used to interact with the features in each window to obtain two transformed local feature maps or , respectively and as the center; then, The center vector of All vectors in the graph are correlated to generate a heat map. The values ​​on the heat map represent Each pixel around The matching probability is calculated by performing expected calculation on the matching probability to obtain the final sub-pixel precision position. , and all matching points Aggregate to get the final fine matching result , is the number of matches.

8. The structured enhancement driven large-scale multimodal image registration method according to claim 1, characterized in that: A random sampling consensus algorithm is used to estimate the image transformation matrix from the refined matching results.

9. The structured enhancement driven large-scale multimodal image registration method according to claim 1, characterized in that: The image to be registered is transformed using the transformation matrix H to align it with the reference image at the sub-pixel level: in, Is the result of a precise match Image to be registered The coordinates of the potential interior points in yes After transformation matrix After transformation, the image The coordinates in .

10. A structured enhancement driven large-scale multimodal image registration system, characterized in that: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a structured enhancement driven large-scale multimodal image registration method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-modal image feature matching method based on multi-feature matching

    CN115496928A

  • Multi-modal image registration method based on prediction correction and convergence attention transformer

    CN117173226A

  • Multi-modal image registration method based on feature fusion and Transform

    CN118799366A

  • Systems and methods for multi-modal multi-dimensional image registration

    US20230281751A1

  • Method and system for automatic registration of images

    US9245201B1

Cited By

  • Remote sensing image feature matching and splicing method based on improved LoFTR algorithm

    CN120634880A

  • Image registration method and system based on common feature and imaging parameter joint guidance

    CN121921349A