A Structured Enhancement-Driven Large-Scale Multimodal Image Registration Method and System
Through the multimodal image registration method driven by structured enhancement, the detector-free framework and many-to-many matching strategy are used to solve the problems of NRD and geometric distortion in multimodal image registration, and high-precision and robust image registration are achieved, suitable for remote sensing and computer vision fields.
Patent Information
- Application Number
- CN202510533633.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing multimodal image registration methods are difficult to effectively match feature points when facing nonlinear radiation differences (NRD) and geometric distortions between images, especially in areas where textures are scarce or modal differences are large, resulting in insufficient matching accuracy and robustness.
The structured enhancement drive method is adopted to directly search the corresponding pixels between images through a detector-free framework, establish a robust reference anchor to enhance structural information, and adopt a many-to-many matching strategy, combining self-attention and cross-attention layers for information interaction, and optimize the matching relationship.
The accuracy and robustness of multimodal image registration are improved, especially in areas with large differences in weak textures and modalities, and precise matching at the sub-pixel level is achieved, reducing dependence on traditional feature detectors and reducing computing resource requirements.
Smart Images

Figure CN120047505B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a structured enhancement driven large-scale multimodal image registration method and system. Background Art
[0002] Image registration is a core challenge in visual understanding and interpretation. It plays a vital role in the fields of remote sensing and computer vision and is widely used in tasks such as image stitching, image fusion, and change detection. However, single-modal data often lacks sufficient depth and is difficult to meet the needs of remote sensing and computer vision applications. Therefore, it is particularly important to use data from different modalities, which can not only give play to the complementary advantages of each modality, but also help achieve a comprehensive understanding of the image. Therefore, how to effectively integrate and deeply analyze multi-sensor, multi-resolution and multi-temporal data has become a key direction of current research. In this context, multimodal image matching has become one of the core issues that need to be solved urgently.
[0003] Multimodal image registration not only involves complex geometric differences (such as scale, rotation, etc.) that are common in single-modal matching, but its main difficulty lies in the nonlinear radiometric differences (NRD) between images. NRD usually originates from different imaging mechanisms or external environmental factors. For example, optical images and synthetic aperture radar (SAR) images present the same scene differently, resulting in significant radiometric differences between image pairs. These differences are usually more obvious in texture-rich areas, which are ideal locations for feature point extraction. Therefore, NRD hinders the effective matching of features between images and increases the complexity and challenge of the image matching task.
[0004] In recent years, many methods have been actively devoted to solving this problem, which can be divided into two categories: region-based methods and feature-based methods. These methods have been shown to be able to extract representative and sufficiently repeatable feature points in the face of geometric distortion and NRD. However, the study of these multimodal matching methods has revealed three major problems that cannot be ignored:
[0005] (1) In order to obtain high-precision matching relationships, existing methods often impose a mandatory one-to-one constraint on the matching process. Although this constraint helps to improve the accuracy of matching, when the scale difference between the image pairs is large, it is often limited to the smaller scale image, resulting in a significant reduction in the number of matching points that can be extracted.
[0006] (2) Existing methods mainly rely on common structural features (such as object edges) between images of different modalities to extract feature points. However, in scenes with scarce textures or sparse common structures, effectively extracting feature points remains a huge challenge.
[0007] (3) When extracting feature points from a single image, the problem of feature point offset is likely to occur. Due to the differences in texture details of images in different modalities, the feature points in an image may not be precisely matched with those in another image. Instead, adjacent points with similar descriptions may be misidentified as matching points, leading to offset, and the existence of NRD further exacerbates this problem. Summary of the Invention
[0008] To address the above problems, the present invention proposes a structured enhancement-driven large-scale multi-modal image registration method. First, a detector-free framework (LoFTR) is introduced to alleviate the feature point offset problem by directly finding corresponding pixels between images. Subsequently, to solve the NRD and geometric distortion problems in multi-modal images, a structure enhancement module is designed to establish robust reference anchors to enhance the structural information of regions with significant modal differences and enrich the corresponding pixels in such regions. In addition, to make up for the problem of a small number of matching points between images with large scale differences, the present invention removes the one-to-one matching constraint in the original detector-free framework and adopts a many-to-many matching strategy. Finally, more accurate positioning of highly correlated corresponding pixels is performed to obtain an optimized matching relationship, and the image transformation matrix is estimated to perform image registration on multi-modal images.
[0009] The object of the present invention is to propose a structured enhancement-driven large-scale multi-modal image registration method, and this method includes the following steps:
[0010] Collect images in different modalities to obtain a dataset of images to be registered, where the dataset of images to be registered includes images to be registered and reference images;
[0011] Arbitrarily select two different-modal images from the dataset of images to be registered, and extract pairs of feature maps with two different downsampling resolutions;
[0012] For a pair of feature maps at one resolution, obtain a rough feature map and reference anchors;
[0013] Use the reference anchors to perform structure enhancement on the rough feature map to obtain an enhanced feature map;
[0014] Calculate the matching probability between the enhanced feature maps and obtain a rough matching result;
[0015] Combine the rough matching result and the pair of feature maps at the other resolution to obtain a fine matching result;
[0016] Use the fine matching result to calculate the transformation matrix of the image, transform the image to be registered, and perform sub-pixel level alignment with the reference image to obtain the final registration result.
[0017] Furthermore, the images of different modalities include: optical images, synthetic aperture radar images, map images, depth maps, and infrared images; a transformation matrix is randomly generated for each pair of images and applied to one of the images to obtain a dataset of deformed images to be registered.
[0018] Furthermore, an existing ResNet network is used to obtain pairs of feature maps with different downsampling resolutions.
[0019] Furthermore, the specific implementation of obtaining the rough feature map and the reference anchor points for a pair of feature maps of one resolution is as follows:
[0020] For a given pair of multimodal images and , and refer to the paired images in image registration, representing the sensed image and the reference image respectively; position encoding is added to the feature map at a scale of 1 / 8 , and then the position-encoded features are processed through the self-attention layer and the cross-attention layer to achieve information interaction between different modalities:
[0021]
[0022] Among them, in the case of the self-attention mechanism, represents the self-attention layer or the cross-attention layer, the values of or , in the case of the cross-attention mechanism, is , the value of is , represents the scale factor; after the above processing, the feature map obtains the corresponding , by multiplying these two flattened feature maps , , the similarity matrix is obtained, where H and W refer to the height and width of image A or B respectively; a bidirectional softmax operation is performed on to obtain the confidence matrix , and the pixels with high confidence are selected as the reference anchor points:
[0023] ,
[0024] ,
[0025] ,
[0026] Among them is a hyperparameter that controls the number of reference anchor points, is an intermediate quantity, MAX means taking the maximum, and MIN means taking the minimum, represents the set of reference anchor points, n represents the number of reference anchor points, and i represents the serial number of all pixel points in and j represents the serial number of all pixel points in
[0027] Furthermore, the specific implementation method for obtaining the enhanced feature map is as follows:
[0028] By calculating the Euclidean distance from the rough feature map to the reference anchor points where includes and find the nearest reference anchor point for each pixel point and calculate the angle between it and the nearest reference anchor point :
[0029]
[0030] where represents the relative position between the reference anchor point closest to point i and point i, represents the relative position between the position of the reference anchor point closest to point j and point j, represents the normalization operation; subsequently, the obtained structural information and the feature map are concatenated together, and a multi-layer perceptron network (MLP) is used to fuse the concatenated information to obtain the fused feature and a cross-attention layer is used to perform information interaction on the feature pairs:
[0031]
[0032]
[0033] where respectively represent the reference anchor point, the angle and length between the pixel point and the reference anchor point, and the subscript c represents the serial number of the nearest reference anchor point, represents concatenation.
[0034] Furthermore, for the enhanced feature map recalculate the similarity matrix between them and perform softmax operations on the two dimensions of respectively to obtain the matching probability at the coarse stage:
[0035] ,
[0036]
[0037] In the matching probability matrix assuming is a pair of matching points, then in the confidence value is higher than the threshold and higher than other elements in the same row; similarly for in the confidence value is higher than the threshold and higher than other elements in the same column. Therefore, the rough match prediction is expressed as:
[0038]
[0039] wherein represents or the serial number corresponding to the maximum value in is a pair of rough matching points, is the rough match result.
[0040] Furthermore, for each pair of rough matching points first, locate its corresponding position or in the 1 / 2-scale feature map then crop out two local windows of size ; then, use a self-attention layer and a cross-attention layer to perform information interaction on the features in each window to obtain two transformed local feature maps or centered on and respectively; subsequently, calculate the correlation between the center vector of and all vectors in to generate a heat map, and the values on the heat map represent the matching probability of each pixel around ; by calculating the expectation of the matching probability, the final sub-pixel accuracy position is obtained, and all matching points are aggregated to obtain the final fine match result is the number of matches.
[0041] Furthermore, the random sample consensus algorithm is used to estimate the transformation matrix of the image from the fine match result.
[0042] Further, the image to be registered is transformed using the transformation matrix H and aligned with the reference image at the sub-pixel level:
[0043]
[0044] Wherein, is the fine matching result of the coordinates of the potential inliers in the image to be registered while is the coordinates in the image after being transformed by the transformation matrix .
[0045] The present invention also provides a large-scale multi-modal image registration system driven by structured enhancement, including:
[0046] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a large-scale multi-modal image registration method driven by structured enhancement as described in the above technical solution.
[0047] The present invention proposes a large-scale multi-modal image registration method driven by structured enhancement. By establishing pixel-level correspondence relationships, the problem of feature point offset is weakened, and reference anchor points are used to enhance the correspondence relationships in weak texture and regions with large texture differences affected by NRD and geometric distortion. By introducing a detector-free framework, corresponding pixels are directly searched between images to alleviate the problem of feature point offset. Subsequently, to solve the problems of NRD and geometric distortion in multi-modal images, a structure enhancement module is designed to establish robust reference anchor points to enhance the structure information in regions with significant modal differences and enrich the corresponding pixels in such regions. In addition, to make up for the problem of fewer matching points between images with large scale differences, the present invention removes the one-to-one matching constraint in the original detector-free framework and adopts a many-to-many matching strategy. Finally, we perform more accurate positioning on highly correlated corresponding pixels to obtain optimized correspondence relationships, estimate the image transformation matrix, and perform image registration on multi-modal images. The present invention effectively solves the challenges of feature point offset and NRD in multi-modal images by enhancing the structure information in weak texture and regions with large modal differences. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flowchart of an embodiment of the present invention.
[0049] Figure 2 is an effect diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0051] The technical process of the embodiments of the present invention is shown in Figure 1 A large-scale multimodal image registration driven by structured enhancement, comprising the following steps:
[0052] 1) Multimodal data processing
[0053] A total of five modalities of data were collected: optical images, SAR images, maps, depth images, and infrared images. In the training stage, we transformed the images of different modalities by randomly generating transformation matrices, and used the transformation matrices as the ground truth of the method.
[0054] 2) Feature extraction
[0055] We used a standard convolutional architecture and a feature pyramid network (denoted as a local feature convolutional neural network (CNN)) to extract rough features at a scale of 1 / 8 of the original image size and fine features at a scale of 1 / 2 from two different modality images. CNN has inductive biases of translational equivariance and locality, which are very suitable for local feature extraction. The downsampling introduced by CNN also reduces the input length of the feature extraction module, which is crucial for ensuring controllable computational costs.
[0056] 3) Coarse correspondence acquisition
[0057] Position encoding is added to the feature map at a scale of 1 / 8. The transformed features will become position-related, which is crucial for the present invention to generate sufficient matches in weakly textured regions. We processed the transformed features through self-attention layers and cross-attention layers to achieve information interaction between different modalities. The self-attention layer can capture the dependencies between different positions within the same modality, while the cross-attention layer further enhances the fusion and cooperation of multimodal information by comparing the features between different modalities. Through this information interaction mechanism, the model can accurately establish the association between feature points among multimodal images, thereby improving the accuracy and robustness of cross-modal matching.
[0058] On this basis, we further calculated the similarity of each pixel on the 1 / 8-scale feature map to evaluate the matching possibility between different pixels. To further improve the matching accuracy, we used a bi-directional softmax operation to obtain the confidence matrix of the feature map. This matrix reflects the matching confidence between each pixel and other pixels, so that the most robust reference anchor points can be selected according to the confidence. By selecting these high-confidence reference anchor points, we can enhance the structural information in the weakly textured regions of the image, making the matching process more stable and accurate. Especially in those regions with scarce textures or large modality differences, the reliability of the reference anchor points can effectively improve the overall matching quality.
[0059] 4) Structured enhancement
[0060] In regions of an image where the texture is weak or there are differences, it is often challenging to confirm the corresponding positions solely through pixel scanning. First, we calculate the Euclidean distance between each pixel point in the 1 / 8 feature map and a pre-selected reference anchor point, and find the nearest reference anchor point for each pixel point based on the distance. In this way, we can determine the relative position of each pixel in the image and further calculate the angle between the pixel and its nearest reference anchor point. Next, we use a multi-layer perceptron (MLP) network to concatenate and fuse the features of each pixel point with the distance and angle information to its nearest reference anchor point. Subsequently, we pass the feature map with structured enhancement through a cross-attention layer for information interaction and obtain a rough matching matrix. This approach not only enhances the feature description of pixels but also introduces more global information, enabling pixels in regions with weak texture or large modal differences to effectively rely on the relative position relationship with the reference anchor points for matching.
[0061] 5) Adaptive many-to-many matching strategy
[0062] In the process of obtaining rough corresponding matches, we no longer rely on the traditional mutual nearest neighbor strategy to enforce a one-to-one matching constraint to improve accuracy. Instead, after obtaining the confidence matrix using a bi-directional - softmax operation, we consider all matches greater than a fixed threshold as correct matches. During this process, especially on the 1 / 8 scale feature map, we may encounter the situation where multiple pixels in the large-scale feature map correspond to the same pixel in the small-scale feature map. This situation is very common in image pairs with large scale differences, but it does not necessarily mean a mis-match. Therefore, we map this many-to-one matching relationship back to the 1 / 2 scale feature map and perform sub-pixel localization on the many-to-one match, thereby converting it into a one-to-one match. This strategy effectively maximally restores the matching relationship in the large-scale context.
[0063] 6) Obtaining optimized corresponding relationships
[0064] After establishing the rough matches, we map these matches to the 1 / 2 scale feature map and crop out local windows of a fixed size around each match point. Next, we send these cropped windows into self-attention layers and cross-attention layers for information interaction and feature aggregation. Our goal is to find the exact matching relationship between two corresponding windows, that is, the specific corresponding position of the central pixel point of one window in the other window. Through this exact matching, it is possible to ensure that the positions of the corresponding points reach sub-pixel level accuracy, significantly improving the accuracy of the matching.
[0065] 7) Multi-modal image registration
[0066] Finally, after obtaining the exact matching relationships between the multi-modal images, we use the Random Sample Consensus (RANSAC) algorithm to estimate the transformation matrix of the images from these hypothesized correspondences. The transformation matrix is used to transform the image to be registered, aligning it with the reference image at the sub-pixel level.
[0067] The following is the specific implementation process of the above steps:
[0068] 1) As is well known, having sufficient training data is a prerequisite for achieving good performance. However, obtaining the transformation parameters of real image pairs as ground truth remains a major challenge in the image registration task. To address this issue, we adopt a simple synthetic image pair method to generate the transformation parameters of pseudo-labels. Specifically, we randomly generate a transformation matrix for each image pair and apply it to one of the images to obtain the deformed dataset of the image to be registered. We collected a total of 5 different modalities of images, including optical images, synthetic aperture radar images, map images, depth maps, and infrared images.
[0069] 2) For a given multi-modal image pair and , and refer to the paired images in image registration, representing the image to be registered and the reference image respectively, and can be of any modality; the existing ResNet network is used to obtain multi-layer features with different downsampling resolutions . In this embodiment, the coarse matching of the image is obtained on the feature map at the 1 / 8 scale , and the optimized matching of the image is obtained on the feature map at the 1 / 2 scale . Those skilled in the art can select appropriate scale feature maps for subsequent processing according to actual needs.
[0070] 3) Add positional encoding to the feature map at the 1 / 8 scale : By applying sine and cosine functions with different frequencies to the pixel points on the feature map, the transformed features will become position-related. Subsequently, the transformed features are processed through the self-attention layer and the cross-attention layer to achieve information interaction between different modalities:
[0071]
[0072] Among them, in the case of the self-attention mechanism, represents the self-attention layer or the cross-attention layer, values are all or , in the case of the cross-attention mechanism, is , The value is , indicating the scale factor. The feature map after the above processing results in the corresponding , . By multiplying these two flattened feature maps , , a similarity matrix can be obtained, where H and W respectively refer to the height and width of image A or B. Perform a bidirectional softmax operation on to obtain the confidence matrix , where we select the pixels with high confidence as the reference anchor points:
[0073] ,
[0074] ,
[0075] ,
[0076] where is a hyperparameter that controls the number of reference anchor points, is an intermediate quantity, MAX represents taking the maximum, MIN represents taking the minimum, represents the set of reference anchor points, n represents the number of reference anchor points, the i - dimension represents the serial numbers of all pixel points in and the j - dimension is
[0077] 4) By calculating the Euclidean distance from the rough feature map (including and ) to the reference anchor points , the nearest reference anchor point for each pixel point can be found, and the angle between it and the nearest reference anchor point is calculated:
[0078]
[0079] where represents the relative position between the reference anchor point nearest to point i and point i, represents the relative position between the position of the reference anchor point nearest to point j and point j, represents the normalization operation. Subsequently, we splice the obtained structural information and the feature map together, use a multi - layer perceptron network (MLP) to fuse the spliced information, obtain the fused feature , and perform information interaction on the feature pairs using cross - attention:
[0080]
[0081]
[0082] Among them respectively represent the reference anchor point, the angle and length between the pixel point and the reference anchor point, c represents the serial number of the nearest reference anchor point, represents stitching. Thus, with the help of the reference anchor point, the structure enhancement of the weak texture or the region with large modal differences is completed, and the corresponding relationship of such regions is restored.
[0083] 5) For the updated feature map , recalculate the similarity matrix between them , and perform softmax operations on the two dimensions of respectively to obtain the matching probability at the coarse stage:
[0084] ,
[0085]
[0086] From the matching probability matrix , assuming is a pair of matching points, then the confidence value of in should be higher than the threshold and higher than other elements in the same row; in this matrix, the i dimension represents the serial numbers of all pixel points in , and the j dimension is the serial numbers of all pixel points in , so the correspondence between i and j is the correspondence of pixel points in and . Similarly, for , we select the index higher than the threshold in the column and the remaining elements as the matching prediction. We represent the coarse matching prediction as:
[0087]
[0088] Among them, represents the serial number corresponding to the maximum value in or , is a pair of coarse matching points, and the coarse matching result relieves the restriction of the one-to-one matching strategy and adopts the many-to-many matching strategy to make up for the corresponding relationship of image pairs with large scale differences to the greatest extent.
[0089] 6) When establishing the coarse matching After that, these matches are mapped back to the resolution of the original image through a coarse-to-fine process. For each pair of coarse matching points , we first locate their corresponding positions in the 1 / 2-scale feature map or , and then crop out two local windows of size . Next, an self-attention layer and a cross-attention layer are used to perform information interaction on the features in each window, obtaining two transformed local feature maps or , centered at and respectively. Subsequently, we calculate the correlation between the central vector of and all vectors in , thus generating a heat map, where the values on the heat map represent the matching probabilities of each pixel around with . By calculating the expectation of the matching probabilities, we obtain the final sub-pixel accurate position , and aggregate all matching points to obtain the final fine matching result , being the number of matches.
[0090] 7) Finally, after obtaining the fine matching between multi-modal images , a random sample consensus algorithm is used to estimate the transformation matrix of the image from these assumed correspondences. The transformation matrix is used to transform the image to be registered, aligning it with the reference image at the sub-pixel level:
[0091]
[0092] where is the coordinate of the potential inlier in the image to be registered in , while is the coordinate in the image after being transformed by the transformation matrix .
[0093] The effects of the present invention are illustrated by specific data below. The multi-modal data used in the embodiments of the present invention are all from public data sets:
[0094] Optical-SAR image registration: This data set contains pairs of aligned grayscale images and SAR images obtained from satellites, covering rural and urban landscapes. The data set includes 2011 pairs of images for training and 424 pairs of images for testing.
[0095] Optical - Infrared Image Registration: This dataset is mainly sourced from the LLVIP dataset, taken by a stereo camera at dusk, and contains 26 different street locations. The image content mainly includes a large number of pedestrians, vehicles, and cyclists.
[0096] Optical - Map Image Registration: This dataset is divided into two parts. The first part includes 200 pairs of image pairs, focusing on buildings and streets. The second part includes 330 pairs of image pairs, directly collected from OpenStreetMap (OSM), covering farmland and buildings. The image pairs from OSM are used for the training set, while the image pairs in the first part are used for the test set.
[0097] Optical - Depth Image Registration: This dataset is provided by the DIML / CVL RGB - D dataset, including indoor data built using a Microsoft Kinect v2 and outdoor data built using a stereo camera. It mainly contains specific scenes, including classrooms, libraries, hospitals, parks, streets, etc., taken from 2015 to 2017.
[0098] In Figure 2 shows the visualization results of the present invention on 4 datasets. The first column is the image pairs to be registered in different modalities, where the left image in the first column is the sensed image, i.e., the image to be registered, and the right image is the reference image. These image pairs have different degrees of rotation and scale changes; the second column is the feature matching result, and the yellow line represents the fine matches estimated by the algorithm. ; the third column is the multi - modal image registration result. The sensed image undergoes a fine match to estimate the transformation matrix and obtains the registered image.
[0099] In summary, by proposing a structured enhancement - driven large - scale multi - modal image registration method, the present invention effectively improves the accuracy and robustness of image registration, especially in regions with weak texture and large modal differences. By establishing reliable reference anchor points to provide a robust topological structure for pixel points in regions with weak texture and large modal differences, the present invention can achieve high - quality pixel - level matching in complex scenes, thus promoting the wide application of multi - modal image analysis technology. This technological advancement not only reduces the dependence on traditional feature detectors, thereby reducing the cost of hardware and computing resources, but also provides more efficient and accurate solutions for industries such as remote sensing, medical imaging, and autonomous driving, promoting the development of related industries.
[0100] On the other hand, the embodiment of the present invention also provides a structured enhancement - driven large - scale multi - modal image registration system, including:
[0101] A processor and a memory, where the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a large-scale multi-modal image registration method with a structured enhanced drive as described in the above technical solution.
[0102] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar ways of substitution, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A structured enhancement-driven large-scale multimodal image registration method, characterized in that Including: Collect images of different modalities to obtain a dataset of images to be registered, where the dataset of images to be registered includes several pairs of image pairs composed of images to be registered and reference images; For any pair of image pairs, extract pairs of feature maps with two different downsampling resolutions; For the pair of feature maps with one of the resolutions, obtain the corresponding rough feature map and reference anchor points respectively; The specific implementation method is as follows: For a given pair of multimodal images and , where A and B are the numbers of the image to be registered and the reference image respectively; add positional encoding to the feature map at 1 / 8 scale , and then process the feature with positional encoding through self-attention layer and cross-attention layer to achieve information interaction between different modalities: Among them, represents a self-attention layer or a cross-attention layer. In the case of the self-attention mechanism, the values are all or . In the case of the cross-attention mechanism, is , the value of is which represents a scaling factor; Feature map After the above processing, the corresponding , , by multiplying these two flattened feature maps , , a similarity matrix is obtained, where H and W respectively refer to the height and width of image A or B; for a two-way softmax operation is performed to obtain a confidence matrix , and the pixel points with high confidence are selected as reference anchor points: , , , Among them is a hyperparameter that controls the number of reference anchor points, is an intermediate quantity, MAX means taking the maximum, and MIN means taking the minimum, represents the set of reference anchor points, n represents the number of reference anchor points, and i represents the serial numbers of all pixel points in, and j represents the serial numbers of all pixel points in; Use the reference anchor points to enhance the structure of the rough feature map to obtain the enhanced feature map; Calculate the matching probability between the enhanced feature maps and obtain the rough matching result; Combine the rough matching result and the pair of feature maps with the other resolution to obtain the fine matching result; Use the fine matching result to calculate the transformation matrix of the image, transform the image to be registered, and align it with the reference image at the sub-pixel level to obtain the final registration result.
2. The large-scale multimodal image registration method with structured enhanced driving as claimed in claim 1, wherein: Images of different modalities include: optical images, synthetic aperture radar images, map images, depth maps, and infrared images; randomly generate a transformation matrix for each image pair and apply it to one of the images to obtain a deformed dataset of images to be registered.
3. A large-scale multimodal image registration method with a structured enhanced drive as described in claim 1, characterized in that: Use the ResNet network to obtain pairs of feature maps with different downsampling resolutions.
4. A large-scale multi-modal image registration method with structured enhanced driving as claimed in claim 1, wherein: The specific implementation method for obtaining the enhanced feature map is as follows: By calculating the Euclidean distance from the rough feature map to the reference anchor points , the nearest reference anchor point for each pixel is found , and the angle between the pixel and the nearest reference anchor point is calculated : Among them represents the relative position between the reference anchor point closest to point i and point i, represents the relative position between the position of the reference anchor point closest to point j and point j, represents a normalization operation; subsequently, the obtained structural information and the feature map are spliced together, and a multi-layer perceptron network is used to fuse the spliced information to obtain the fused feature , and cross-attention is used to perform information interaction on the feature pairs: wherein respectively represent the reference anchor point, the angle and length between the pixel point and the reference anchor point, and the subscript c represents the serial number of the nearest reference anchor point, represents splicing.
5. The large-scale multimodal image registration method with structured enhanced driving as claimed in claim 1, wherein: For the enhanced feature maps, recalculate the similarity matrix between them , and perform softmax operations on the two dimensions of respectively to obtain the matching probabilities at the coarse stage: , In the matching probability matrix assuming is a pair of matching points, then the confidence value of in is higher than the threshold and higher than other elements in the same row; similarly for , the confidence value of in is higher than the threshold and higher than other elements in the same column. Therefore, the rough matching prediction is expressed as: Among them, denotes or the serial number corresponding to the maximum value in is a pair of rough matching points, is the rough matching result.
6. The large-scale multimodal image registration method with structured enhanced driving according to claim 1, characterized in that: For each pair of coarsely matched points in the coarsely matched results , first locate its corresponding position in the 1 / 2 scale feature map or , and then crop out two local windows of size , where A and B are the numbers of the images to be registered and the reference image respectively, and w is the size of the local window; then, use a self-attention layer and a cross-attention layer to perform information interaction on the features in each window to obtain two transformed local feature maps or , centered at and respectively; subsequently, calculate the correlation between the center vector of and all vectors in to generate a heat map, where the values on the heat map represent the matching probability of each pixel around with ; by calculating the expectation of the matching probability, obtain the final sub-pixel level accurate position , and aggregate all the matched points to obtain the final fine matching result , is the number of matches.
7. The large-scale multimodal image registration method with structured enhanced driving according to claim 1, wherein: Adopt the random sample consensus algorithm to estimate the transformation matrix of the image from the fine matching result.
8. A large-scale multimodal image registration method driven by structured enhancement as claimed in claim 1, wherein: Use the transformation matrix H to transform the image to be registered and align it with the reference image at the sub-pixel level: Among them, is the fine matching result of the coordinates of the potential inliers in the image to be registered , while is the coordinates in the image after transformation by the transformation matrix in .
9. A large-scale multimodal image registration system driven by structured enhancement, characterized in that, Including: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a large-scale multi-modal image registration method driven by structural enhancement as described in any one of claims 1-8.