Multimodal image matching method and system based on saliency map structure enhancement

By constructing a saliency confidence map and a self-cross attention mechanism, the problems of feature offset and illumination difference caused by different imaging methods in multimodal image matching are solved, and high-precision and stable matching of multimodal images is achieved, which is suitable for remote sensing image registration and cross-sensor data fusion.

CN120543997BActive Publication Date: 2025-10-03WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511037884.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-03
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

Existing multimodal image matching technologies have difficulty achieving stable and accurate matching across scales, rotations, and perspectives when faced with feature expression offsets, geometric deformations, and lighting differences caused by different imaging methods.

Method used

A pixel-level saliency confidence map is constructed, and by fusing multi-scale structural features with semantic segmentation information, combined with self-attention and cross-attention mechanisms, the internal structural consistency of the image and cross-modal semantic alignment are enhanced to achieve semi-dense matching of multimodal images.

Benefits of technology

It significantly improves the accuracy and stability of multimodal image matching, especially under complex conditions, showing comprehensive performance that is better than traditional methods and existing deep matching frameworks. It is suitable for remote sensing image registration, cross-sensor data fusion and high-precision visual positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543997B_ABST
    Figure CN120543997B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal image matching method and system based on saliency graph structure enhancement, belonging to the field of image processing. First, the present invention innovatively constructs a pixel-level saliency confidence map to measure the matching potential of each region. Through this map, the attention mechanism is guided to dynamically focus on key areas in the graph structure. Second, multi-scale structural features are integrated with semantic segmentation information to enhance the semantic perception ability of feature expression. Finally, two types of heterogeneous graph structures, intra-image structural graph and inter-image semantic guidance graph, are constructed. By introducing saliency-modulated self-attention and cross-attention mechanisms, global-local information enhancement and cross-modal semantic alignment are achieved on the graph structure, thereby significantly improving the matching accuracy and stability, and realizing semi-dense matching of multimodal images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular relates to a multimodal image matching method and system based on saliency graph structure enhancement. Background Art

[0002] Multimodal image matching, a key technology for cross-modal data alignment, aims to establish pixel-level or feature-level correspondences between heterogeneous imaging data (such as visible light, SAR, and thermal infrared), enabling semantic alignment and geometric registration of cross-domain information. In recent years, with the rapid development of multi-sensor fusion technology, this research direction has achieved breakthroughs at the intersection of computer vision and remote sensing. Its core value lies not only in improving the consistency of cross-modal data representation but also in overcoming the inherent limitations of single sensors in illumination adaptability, texture sensitivity, and geometric invariance through modal complementarity. Multimodal matching has become a core technology for robust environmental perception in critical scenarios such as high-precision positioning for autonomous driving, 3D urban reconstruction, and disaster emergency response. However, compared to intramodal matching, it presents more significant challenges, primarily due to significant geometric, radiometric, and viewpoint differences between different imaging technologies, which make it difficult to extract common features across modalities.

[0003] To overcome these difficulties, experts and scholars have proposed a series of solutions in recent years. For example, early image matching methods primarily relied on the construction of feature points and feature descriptors. The most classic methods include SIFT and SURF, which perform matching based on scale-invariant feature transforms. For SIFT, see D.G. Lowe, “Distinctive image features from scale-invariant keypoints,” International journal of computer vision, vol. 60, pp. 91–110, 2004. For SURF, see H. Bay, T.Tuytelaars, and L. Van Gool, “Surf: Speeded up robust features,” in Computer Vision–ECCV 2006: 9th European Conference on Computer Vision, Graz, Austria, May 7-13, 2006. Proceedings, Part I 9, Springer, 2006, pp. 404–417. The Position-Scale-Oriented SIFT (PSO-SIFT) algorithm proposed by Ma et al. aims to overcome the problems of nonlinear brightness differences and rotation changes. See W. Ma et al., “Remote sensing image registration with modified SIFT and enhanced feature matching,” IEEE Geoscience and Remote Sensing Letters, vol. 14, no. 1, pp. 3–7, 2016.Xiong et al. proposed the oriented self-similarity (OSS) method and the adaptive scaling (ASS) method, which target the modal differences and rotational changes of multimodal remote sensing images. These methods gradually show stability in image matching with small nonlinear radiometric distortion. See X. Xiong, G. Jin, Q. Xu, and H. Zhang, “Self-similarity features for multimodal remote sensing image matching,” IEEE Journal of selected topics in applied earth Observations and remote sensing, vol. 14, pp. 12440–12454, 2021. To address images with large radiometric distortion, Li et al. proposed the RIFT and RIFT2 methods. RIFT uses a polar grid to design descriptors, reducing descriptor dimensionality to improve computational efficiency. However, it performs poorly with noise, illumination, and scale variations. For more information on the RIFT method, see: J. Li, Q. Hu, and M. Ai, “RIFT: Multi-modal image matching based on radiation-variation insensitive feature transform,” IEEE Transactions on Image Processing, vol. 29, pp. 3296–3310, 2019. RIFT2 enhances robustness and is particularly suitable for remote sensing images or images with complex textures. For more information on the RIFT2 method, see: J. Li, P. Shi, Q. Hu, and Y. Zhang, “RIFT2: Speeding-up RIFT with anew rotation-invariance technique,” ​​arXiv preprint arXiv:2303.00319, 2023.Yao et al. proposed a matching method based on the histogram of absolute phase consistency (HAPCG, reference: Y. Yao, Y.Zhang, Y. Wan, X. Liu, and H. Guo, “Heterologous images matching consideringanisotropic weighted moment and absolute phase orientation,” Geomatics andInformation Science of Wuhan University, vol. 46, no. 11, pp. 1727–1736, 2021), a diffusion tensor descriptor based on multi-orientation features (MoTIF, reference: Y. Yao, B. Zhang, Y.Wan, and Y. Zhang, “MOTIF: Multi-orientation tensor index feature descriptorfor SAR-optical image registration,” The International Archives of thePhotogrammetry, Remote Sensing and Spatial Information Sciences, vol. 43, pp.99–105, 2022) and a weighted phase orientation model (HOWP, reference: Y. Zhang et al., “Histogramof the orientation of the These three methods can overcome image scale, translation, and rotation differences. HOWP performs better in matching multi-source remote sensing images with large noise differences, but the stability of these three methods is still insufficient when dealing with large rotation differences.Recently, Liao et al. proposed a refinement method to extract repetitive features from multimodal images and reduce the impact of inter-modal differences on matching. (See Y. Liao et al., “Refining multi-modal remote sensing imagematching with repetitive feature optimization,” International Journal of Applied Earth Observation and Geoinformation, vol. 134, p. 104186, 2024.) They also proposed an AMSE method to adaptively adjust filter parameters, reducing the complexity of manual parameter adjustment and improving matching success rate and accuracy. (See Y. Liao, P. Tao, Q. Chen, L. Wang, and T. Ke, “Highly adaptive multi-modal image matching based on tuning-free filtering and enhanced sketch features,” Information Fusion, vol. 112, p. 102599, 2024.) The POS-GIFT method proposed by Hou et al. achieves good matching results on multimodal data by designing multi-layer circular point sampling, rotation-invariant feature descriptors, and a position-orientation-scale guided interior point recovery strategy. (See Z. Hou, Y. Liu, and L. Zhang, “POS-GIFT: A geometric and intensity-invariant feature transformation for multimodal images,” Information Fusion, vol. 102, p. 102027, 2024.) However, although these methods have continuously optimized and updated their algorithms to address modal differences such as radiometric, illumination, and geometry encountered in multimodal image matching, their ability to represent multimodal image features is limited, making it impossible to simultaneously address the matching challenges in complex cross-scale, cross-rotation, cross-viewpoint, and cross-modal scenarios. Summary of the Invention

[0004] In response to the problems existing in the prior art, the present invention proposes a multimodal image matching method and system based on saliency graph structure enhancement, aiming to improve the consistency and robustness of matching features from the perspective of image structure saliency and cross-modal semantic collaboration. First, the present invention innovatively constructs a pixel-level saliency confidence map to measure the matching potential of each region. Through this map, the attention mechanism is guided to dynamically focus on key areas in the graph structure; secondly, multi-scale structural features and semantic segmentation information are integrated to enhance the semantic perception ability of feature expression; finally, two types of heterogeneous graph structures, intra-image structure map and inter-image semantic guidance map, are constructed. By introducing saliency-modulated self-attention and cross-attention mechanisms, global-local information enhancement and cross-modal semantic alignment are achieved on the graph structure, thereby significantly improving the matching accuracy and stability, and realizing semi-dense matching of multimodal images.

[0005] The matching framework based on the saliency graph structure enhancement mechanism innovatively proposed in this invention effectively alleviates the problem of feature expression offset, geometric deformation and illumination difference caused by different imaging methods in multimodal images, and provides strong technical support for the precise alignment and collaborative optimization processing of multimodal images.

[0006] The present invention proposes a multimodal image matching method and system based on saliency graph structure enhancement to solve the problem of stable matching under complex modal conditions.

[0007] The technical solution adopted by the present invention comprises the following steps:

[0008] Step 1: extract multi-scale structural features of image pairs;

[0009] Step 2: extract multi-scale semantic features of the image pair;

[0010] Step 3: Fuse the multi-scale structural features and the multi-scale semantic features to obtain the fused features;

[0011] Step 4: Construct a pixel-level saliency confidence map;

[0012] Step 5: construct a multi-level feature enhancement network. Each layer of the multi-level feature enhancement network adopts a self-attention mechanism and a cross-attention mechanism, and uses a saliency confidence map to enhance the fusion feature to obtain the enhanced final fusion feature.

[0013] Step 6: Calculate the matching confidence score based on the final fusion feature, and filter the matching point set according to the matching confidence score;

[0014] Step 7: Perform geometric consistency verification based on the basic matrix, remove abnormal point pairs in the semi-dense matching point set, and output a stable matching result.

[0015] Furthermore, in step 1, multi-scale structural features of the image pairs are extracted through the multi-layer convolutional neural network Unet, and in step 2, the pre-trained lightweight semantic segmentation network MobileSAM is used to extract the multi-scale semantic features of the image pairs; in step 3, the multi-scale structural features extracted by Unet and the multi-scale semantic features extracted by the lightweight semantic segmentation network MobileSAM are spliced ​​in the channel dimension and then reduced to fusion features through 1×1 convolution.

[0016] Furthermore, the specific implementation of step 4 is as follows:

[0017] Sobel operator is used to calculate the horizontal and vertical direction The gradient component corresponding to each pixel , calculate the gradient amplitude :

[0018]

[0019] in, , W and H represent the width and height of the image;

[0020] Then, the global gradient mean of the entire image is calculated and standard deviation , used for normalization processing, combined with fusion features , calculate the global mean of the gradient map, which is the saliency confidence map:

[0021]

[0022] in, is the sigmoid function; is the contrast enhancement coefficient; A very small number to prevent the denominator from being 0; Represents pixel-by-pixel multiplication operation; Represents pixel coordinates The confidence level of significance.

[0023] Furthermore, the processing process of each layer in the multi-level feature enhancement network in step 5 is as follows:

[0024] Step 5.1: For each image, construct the image internal structure, establish the adjacent edge relationship set between pixels, and perform a saliency weighted self-attention mechanism on the adjacent edge relationship set to obtain the self-attention feature;

[0025] Step 5.2: Using the reference image as the source and the target image as the target, we construct a cross-modal guidance map from the reference image to the target image using the semantic similarity between the images, and perform saliency-weighted cross-modal feature fusion to obtain the cross-attention feature.

[0026] In step 5.3, the fusion features of each pixel are fused through the self-attention features and the cross-attention features to obtain the final fusion features.

[0027] Furthermore, the specific implementation of step 5.1 is as follows:

[0028] First, for each image, construct the image internal structure:

[0029]

[0030] in, is a set of pixels, Represents the set of adjacent edge relationships between pixels; , is the pixel position; the distance threshold Control the current layer The attentional proximity of

[0031] Structure within the image On the defined set of adjacent edge relationships, the following saliency-weighted self-attention mechanism is performed:

[0032] Basic self-attention score Calculated from the query and key vector:

[0033]

[0034] Among them, query ,key The vector is obtained by linear transformation, , and is a learnable parameter, For the The fused feature vector corresponding to the pixel points, is the rotation position encoding function;

[0035] Add the saliency weight modulation factor to obtain the saliency self-attention score :

[0036]

[0037] in, and Indicates the and The confidence level of the pixel's saliency, is a constant;

[0038] Use Softmax normalization and weighted aggregation of adjacent node features to obtain self-attention features:

[0039]

[0040] in, are learnable parameters, Indicates the i The set of adjacent edge relationships of pixels, For the The fused feature vector corresponding to each pixel point.

[0041] Furthermore, the distance threshold The calculation formula is as follows:

[0042]

[0043] in, is the initial global connection threshold, set to the maximum pixel spacing, for and The Euclidean distance between is the lower distance threshold, is the total number of layers, is the current level;

[0044] The rotation encoding function is defined as follows:

[0045]

[0046] in, Indicates the Pixels relative to the The direction angle of the pixel point, Indicates the The horizontal and vertical coordinates of the pixel points, Indicates the j The horizontal and vertical coordinates of the pixel points; is a small constant that prevents division by zero.

[0047] Furthermore, the specific implementation of step 5.2 is as follows:

[0048] Define a semantic guidance graph with reference image A as source and target image B as target :

[0049]

[0050] Among them, TopK means retaining the top K candidate points with the highest semantic similarity; Indicates the calculation of cosine similarity; and Represent the pixel sets of the reference image and the target image respectively; and Represents the reference image The first fusion feature vector and the target image fused feature vectors;

[0051] For each pair , calculate the cross attention score :

[0052]

[0053] in, and are the parameters that can be learned by the network, is the rotation position encoding function; and Represents the first i pixel points and the first pixel on the target image j pixels;

[0054] The saliency cross-attention score after combining saliency modulation is expressed as:

[0055]

[0056] in, and Represents the reference image The saliency confidence of the pixel and the target image The confidence level of the pixel's saliency, is a constant;

[0057] Use Softmax normalization and weighted aggregation of semantic information from the target graph to obtain cross-attention features:

[0058]

[0059] in, Representation semantic guidance graph Middle i The neighboring points of pixels, are learnable parameters.

[0060] Furthermore, the specific implementation of step 6 is as follows:

[0061] After mapping the final fusion features to the matching space, the matching score matrix is ​​constructed, and then the matching confidence score is obtained through bidirectional Softmax normalization, and the matching confidence score is calculated based on the preset threshold. Select matching point pairs that meet the conditions to form a semi-dense matching point set.

[0062] Furthermore, the specific implementation of step 7 is as follows:

[0063] Estimate the fundamental matrix using the RANSAC algorithm and the eight-point method , calculate the epipolar constraint error for each pair of matching points in the semi-dense matching point set as a measure of geometric consistency:

[0064]

[0065] in, and Respectively represent the first matching point and the first matching point in the target image B Homogeneous coordinates of matching points; represents the fundamental matrix; Represents the calculated epipolar error. Point pairs with errors less than the threshold are considered as geometric inliers to form the final set of matching points.

[0066] The present invention also provides a multimodal image matching system based on saliency map structure enhancement, comprising:

[0067] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the multimodal image matching method based on saliency map structure enhancement as described in the above technical solution.

[0068] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0069] The present invention proposes a multimodal image matching method and system based on saliency graph structural enhancement. The method extracts the structural features of the image through a multi-layer convolutional neural network, and introduces a saliency confidence map generated by gradient and semantic fusion, guiding the subsequent attention mechanism to focus on areas with higher matching potential, effectively suppressing redundant interference in weak texture areas, and improving the stability of the overall matching. At the same time, a pre-trained semantic segmentation network is used to extract a semantic response map, and the semantic information and structural features are spliced ​​and compressed at the channel level in the fusion module to form a robust cross-modal feature expression. Furthermore, the present invention constructs two types of graph structures: intra-image structural map and inter-image semantic map. By introducing saliency-modulated self-attention and cross-attention mechanisms in the graph structure, the internal structure consistency modeling of the image and the semantic alignment between images are achieved. The proposed saliency weight mechanism dynamically regulates the information flow path in the attention calculation, strengthens the dominant position of high semantic areas in feature propagation, and thus significantly improves the matching robustness under complex conditions such as geometric deformation, modality inconsistency and texture loss.

[0070] Experimental results show that the saliency graph structure enhancement method proposed in this paper exhibits better comprehensive performance than traditional methods and existing deep matching frameworks in difficult image matching tasks such as multimodality, weak texture and large parallax. It has wide engineering applicability, especially for practical application scenarios such as remote sensing image registration, cross-sensor data fusion and high-precision visual positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 Flow chart of the method of the present invention;

[0072] Figure 2 Schematic diagram of the fusion of structural features and semantic features of the present invention;

[0073] Figure 3 Schematic diagram of matching enhancement guided by weighted saliency graph structure of the present invention;

[0074] Figure 4 This is the qualitative comparison result of multimodal image matching in an embodiment of the present invention;

[0075] Figure 5 A physical picture of the binocular multimodal imaging mobile vehicle deployed with the method of the present invention;

[0076] Figure 6 This is the multimodal matching result diagram measured on the mobile terminal of the method of the present invention. DETAILED DESCRIPTION

[0077] In order to facilitate ordinary technicians in this field to understand and implement the present invention, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.

[0078] Please see Figure 1 Flowchart, the present invention provides a multimodal image matching method based on saliency map structure enhancement, comprising the following steps:

[0079] Step 1: Use convolutional neural networks to extract multi-scale structural features of image pairs, providing underlying feature support for subsequent semantic fusion and matching calculations;

[0080] As a preferred approach, a convolutional neural network with a U-Net structure is used to perform layer-by-layer encoding and decoding operations on the two input multimodal images, aiming to extract visual elements such as local edges and texture structures in the multimodal images through a deep convolutional network. The output structural features are recorded as:

[0081]

[0082] in, is the structural feature of the reference image A, is the structural feature of the target image B, H and W are the height and width of the input image, respectively, and D is the number of feature channels corresponding to each pixel position. This structural feature map serves as the initial descriptor for each pixel in the subsequent matching process. The Unet encoder uses multi-layer convolution + ReLU + max pooling to extract semantic information, while the decoder uses upsampling and convolution to restore resolution, while retaining skip connections to preserve edges and structural details.

[0083] Step 2: Use a lightweight semantic segmentation network to extract multi-scale semantic features of image pairs;

[0084] Cross-modal image matching relies on high-level semantic consistency. Semantic features can improve regional semantic understanding and alleviate the interference of modality differences on matching.

[0085] As a preference, a pre-trained lightweight semantic segmentation network (MobileSAM) is used to extract semantic feature maps of image pairs:

[0086]

[0087] is the semantic feature of the reference image A, is the semantic feature of the target image B. This feature map provides pixel-level category distribution information, which helps to identify semantically consistent areas such as buildings and roads, and is subsequently used for fusion with structural information.

[0088] Step 3: Fuse structural features with semantic features to perform semantically enhanced feature embedding to improve the robustness of feature matching under multimodal conditions;

[0089] A key challenge in multimodal image matching is that semantic regions can have significant differences in imaging style and radiometric response. Fusion of structural and semantic features can balance texture stability and category consistency, effectively improving the accuracy of cross-modal region matching.

[0090] As a preference, Figure 2 As shown, the structural feature map extracted by Unet Multi-scale semantic feature maps extracted by the lightweight semantic segmentation network MobileSAM After concatenation in the channel dimension, the dimensionality is reduced to fusion features through 1×1 convolution:

[0091]

[0092] This fusion feature retains structural edge information and high-level semantic information, which facilitates enhancing multimodal matching consistency.

[0093] Step 4: Identify regions with high matching potential in the image pair and construct a pixel-level saliency confidence map to guide the subsequent attention mechanism;

[0094] Matching accuracy depends on region saliency. Structurally salient (e.g., edges and corners) and semantically salient (e.g., roads and buildings) regions are more likely to be correctly matched across modalities. This step jointly models gradient structure and semantic confidence to generate a pixel-level saliency confidence map, which provides guidance for the subsequent attention mechanism.

[0095] As a preference, first use the Sobel operator to calculate the horizontal direction of the input image and vertical direction The gradient component is used to detect areas with drastic edge changes, corresponding to each pixel , calculate the gradient amplitude :

[0096]

[0097] in, , W and H represent the width and height of the image.

[0098] Then, the global mean gradient of the entire image is calculated and standard deviation , used for normalization processing, combined with fusion feature map (include and ), calculate the global mean of the gradient map, which is the saliency confidence map:

[0099]

[0100] in, is the sigmoid function; =1.5, is the contrast enhancement coefficient; A very small number to prevent the denominator from being 0; Represents pixel-by-pixel multiplication operation; Represents pixel coordinates The pixel-level saliency confidence map is used to modulate the attention weight in the graph structure, emphasizing the importance of salient regions in feature propagation.

[0101] Step 5: Construct a multi-level feature enhancement network. Each layer of the feature enhancement network uses a self-attention mechanism and a cross-attention mechanism, and uses a saliency confidence distribution map to enhance the fusion features. The processing process of each layer in the feature enhancement network is as follows:

[0102] Step 5.1: Construct the image internal structure and use the saliency-guided self-attention mechanism to model the internal structural contextual relationship of the image to obtain the self-attention feature;

[0103] The internal regional structures of an image have complex contextual dependencies. Local point matching depends not only on its own features but also on the consistency of its neighborhood structure. This step builds on the image's internal graph structure and implements a saliency-weighted self-attention mechanism, guiding feature points to aggregate contextual information within structurally stable regions.

[0104] As a preference, Figure 3 As shown, after completing the fusion feature map After the construction of , in order to enhance the contextual association between the internal features of the image, a sparse graph structure is constructed within the image and a saliency-guided self-attention mechanism is introduced

[0105] (1) Graph construction and node definition:

[0106] For each image, construct the image interior:

[0107]

[0108] in, is a set of pixels, Represents the set of adjacent edge relationships between pixels, and a connection is established only when the two pixels are close enough in space; , is the pixel position; the distance threshold Control the current layer The attention neighbor range of , whose value decreases layer by layer with the number of network layers, realizes the gradual transition strategy from shallow global connection to deep local connection. Its value is as follows:

[0109]

[0110] in, is the initial global connection threshold, which is set to the maximum pixel spacing to ensure that the shallow layer of the network can effectively cover the global features. for and The Euclidean distance between To avoid excessive sparseness in the deep layers of the network, set the lower limit of the distance threshold. (After a lot of experiments, it has been proved that Set to between 50 and 100); is the total number of layers, is the current layer number.

[0111] (2) Saliency-weighted attention scoring:

[0112] In the graph structure The defined set of adjacency edge relationships On top of it, the following saliency-weighted self-attention mechanism is implemented:

[0113] The query and key vectors are obtained through linear transformation:

[0114]

[0115] in, and are the network learnable parameters, That is the first The fused feature vector corresponding to the pixels.

[0116] Added rotation position encoding function , calculate the basic self-attention score:

[0117]

[0118] The rotation encoding is defined as follows:

[0119]

[0120] in, Indicates the Pixels relative to the The direction angle of the pixel point, Indicates the The horizontal and vertical coordinates of the pixel points, Indicates the j The horizontal and vertical coordinates of the pixel points; is a small constant that prevents division by zero.

[0121] Then introduce the saliency weight modulation (using the constructed in step 4 ), and obtain the significant self-attention score :

[0122]

[0123] in, and Indicates the and The confidence level of the saliency of each pixel.

[0124] (3) Significance normalization aggregation:

[0125] Use Softmax normalization and weighted aggregation of adjacent node features (only for ), and get the self-attention feature:

[0126]

[0127] in, are the parameters that can be learned by the network. for Middle The feature vector of each pixel.

[0128] This mechanism significantly enhances the influence of highly salient regions in self-attention propagation, effectively enhancing the local geometric consistency and semantic structure expression capabilities.

[0129] Step 5.2: Use the semantic similarity between images to construct a cross-modal guidance map, perform saliency-weighted cross-modal feature fusion, and obtain cross-attention features;

[0130] Matching multimodal images requires establishing cross-domain feature associations. Because semantic consistency between modalities outweighs texture consistency, this step constructs a semantically guided inter-image graph structure and introduces a saliency-modulated cross-attention mechanism to extract semantically common regions and achieve information alignment.

[0131] As a preference, Figure 3 As shown in the figure, in order to extract common semantic regions between cross-modal images, a directed semantic graph is constructed between image pairs, and a saliency-guided cross-attention mechanism is introduced to achieve semantic alignment between features. The specific steps are as follows:

[0132] (1) Construction of inter-image guidance graph:

[0133] Define a semantic guidance graph with image A as source and image B as target:

[0134]

[0135] Among them, TopK means retaining the top K candidate points with the highest semantic similarity; Indicates the calculation of cosine similarity; and Represent the pixel points of the image to be registered (i.e., the reference image) and the target image respectively; and Represents the image to be registered The first fusion feature vector and the target image fused feature vectors.

[0136] (2) Saliency-weighted cross-attention scoring:

[0137] For each pair , calculate the attention score,

[0138]

[0139] in, and are the parameters that can be learned by the network.

[0140] The saliency cross-attention score after combining saliency modulation is expressed as:

[0141]

[0142] in, and Represents the image to be registered The saliency confidence of the pixel and the target image The confidence level of the saliency of each pixel.

[0143] (3) Cross-graph feature information aggregation:

[0144] The final attention aggregation operation is only for the adjacent points in the guidance graph Execute and get the cross attention feature:

[0145]

[0146] This mechanism helps to compress the modality gap, enhance the representation ability of common semantic regions, and provide robust feature support for subsequent semi-dense matching.

[0147] Step 5.3, finally, the fusion feature vector of each pixel is obtained after the fusion of self-attention feature and cross-attention feature. .

[0148] Step 6: Matching score matrix calculation and semi-dense matching screening;

[0149] After obtaining the robust feature representation of structural-semantic fusion, the similarity of all point pairs between the two images is calculated, and the precise set of point pairs is selected based on the matching confidence. A bidirectional normalization mechanism is used to ensure matching consistency.

[0150] As a preference, after completing the feature enhancement guided by the internal and external image structures, the final fusion representation of all pixels in image A and image B is obtained. and This step completes the semi-dense matching result screening by linear mapping and matching score matrix construction. The specific steps are as follows:

[0151] (1) Feature mapping and similarity calculation

[0152] Project the features into the matching space and construct the similarity matrix:

[0153]

[0154] in is the network learnable parameter, matching score matrix Each term in is defined as the inner product:

[0155]

[0156] in Represents the vector dot product operation, which is used to measure the similarity between feature pairs.

[0157] (2) Bidirectional normalized matching strategy

[0158] In order to ensure that the matching point pairs are the best matches to each other in both the source and target images, bidirectional Softmax normalization is used:

[0159]

[0160] in, Indicates the normalization of each row, indicating the distribution of each source point in the target graph; Indicates the normalization of each column, indicating the distribution of each target point in the source image; Score the confidence of the match. Set the threshold , select all that satisfy Matching point pairs , forming a semi-dense matching point set .

[0161] Step 7: Verify the geometric consistency based on the basic matrix model, remove abnormal point pairs, and output a stable matching result;

[0162] Even if the matching points have high feature similarity, geometric inconsistencies may occur due to occlusion, false detection, etc. This step introduces geometric verification to evaluate whether the point pairs satisfy the epipolar geometric constraints through the basic matrix model to further eliminate false matches.

[0163] As an optimization, in order to further filter out mismatches, a geometric verification module is introduced. The RANSAC random sampling algorithm and the eight-point method are used to fit the basic matrix. , semi-dense matching point set For each pair of matching points in , the epipolar constraint error is calculated as a measure of geometric consistency:

[0164]

[0165] in, and Represents the first matching points and the first Homogeneous coordinates of matching points; Represents the basic matrix of RANSAC fitting; represents the calculated epipolar error.

[0166] Point pairs with smaller errors are considered geometric inliers and form the final set of matching points. This step effectively removes spatially inconsistent mismatched pairs, ensuring the stability and credibility of the output results in terms of structural transformation.

[0167] Step 8: Quantitatively evaluate the matching performance based on key point accuracy and geometric reprojection error;

[0168] In order to quantitatively evaluate the matching accuracy and geometric stability of the multimodal image matching method proposed in this paper, two types of indicators are used for comprehensive evaluation:

[0169] (1) Number of correct matching points within 5-pixel error: used to measure the matching accuracy of local point pairs, that is, the predicted matching points are considered correctly matched when the projection error is less than 5 pixels. The results are shown in Table 1.

[0170] (2) Reprojection error of the four image corner points under homography transformation: used to evaluate the geometric fitting ability of the homography matrix, that is, the average Euclidean distance of the four corner points reprojected into the target image. The results are shown in Table 2.

[0171] Table 1 Statistics of correct matching points within 5 pixel error for different methods in multimodal image registration tasks

[0172]

[0173] Table 2 Corner point reprojection errors of different methods in multimodal image registration tasks

[0174]

[0175] Note: The unit of corner point reprojection error is pixel. Errors exceeding 500 are recorded as 500.

[0176] The matching method of the present invention was compared with the current mainstream traditional algorithms (such as RIFT, POS-GIFT) and deep learning methods (such as LightGlue, XoFTR, RoMa, MINIMA, etc.) in terms of performance.

[0177] As shown in Tables 1 and 2, in typical multimodal image registration tasks, different methods perform significantly differently when faced with complex conditions such as cross-modality, strong distortion, and weak texture. Traditional methods such as RIFT and POS-GIFT have a small number of matching points and large reprojection errors in most scenes, indicating that their stability and geometric accuracy are limited when faced with modal differences. Deep learning methods trained solely on optical images (such as SuperGlue and LightGlue) have a large number of matching points in some scenes, but still show insufficient generalization ability in modal registration such as infrared and SAR. In particular, the corner reprojection error index is generally high, and there is a large geometric distortion.

[0178] In contrast, although the multimodal trained MINIMA series and RoMa achieve more matching points in some scenarios (such as infrared-visible light), they still suffer from high mismatch rates and large fluctuations in reprojection errors in some tasks, indicating that their stability is limited.

[0179] The proposed method achieves significantly better matching numbers and geometric accuracy than other methods in multiple multimodal combinations. In Table 1, the proposed method achieves a large number of valid matching points in all scenes, showing extremely high coverage and robustness; in Table 2, its corner reprojection error remains within the range of 2 to 6 pixels, which is significantly lower than that of most of the comparison methods, demonstrating excellent geometric consistency. More qualitative results can be seen. Figure 4 .

[0180] The RIFT method is cited in J. Li, Q. Hu, and M. Ai, “RIFT: Multi-modal image matching based on radiation-variation insensitive feature transform,” IEEE Transactions on Image Processing, vol. 29, pp. 3296–3310, 2019.

[0181] The POS-GIFT method is cited in Z. Hou, Y. Liu, and L. Zhang, “POS-GIFT: Ageometric and intensity-invariant feature transformation for multimodal images,” Information Fusion, vol. 102, p. 102027, 2024.

[0182] The SuperGlue method is cited in P.-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural networks,” in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2020, pp. 4938–4947.

[0183] The LightGlue method described above is cited in P. Lindenberger, P.-E. Sarlin, and M.Pollefeys, “Lightglue: Local feature matching at light speed,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, 2023, pp. 17627–17638.

[0184] The LoFTR method is cited in J. Sun, Z. Shen, Y. Wang, H. Bao, and X. Zhou, “LoFTR: Detector-free local feature matching with transformers,” in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2021, pp. 8922–8931.

[0185] The above XoFTR method is cited in Ö. Tuzcuoğlu, A. Köksal, B. Sofu, S. Kalkan, andA. A. Alatan, “Xoftr: Cross-modal feature matching transformer,” in Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition, 2024, pp. 4275–4286.

[0186] The RoMa method is cited in J. Edstedt, Q. Sun, G. Bökman, M. Wadenbäck, and M.Felsberg, “RoMa: Robust dense feature matching,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19790–19800.

[0187] The above MINIMALG MINIMA LoFTR MINIMA RoMa Method cited in X. Jiang, J. Ren, Z. Li, X.Zhou, D. Liang, and X. Bai, “MINIMA: Modality Invariant Image Matching,” arXiv preprint arXiv:2412.19412, 2024.

[0188] The "multimodal image matching method based on saliency graph structure enhancement" proposed in this invention has been successfully deployed in a mobile intelligent perception platform with binocular multimodal imaging capabilities, such as Figure 5 The platform consists of an AGX Orin industrial-grade controller, a color area array camera, a thermal infrared camera, a time synchronization board, a network switch, and an antenna module. It is integrated into a mobile chassis and has the capabilities of autonomous navigation and heterogeneous sensor fusion.

[0189] In actual deployment, the proposed algorithm processes multimodal image sequences captured by infrared and visible light cameras in real time, and performs accurate cross-modal feature matching and registration through algorithm modules, providing high-quality input for downstream tasks such as target detection, scene understanding, and map construction. Test results show that the proposed method maintains good robustness and matching accuracy in extreme environments such as weak texture and low light, and achieves stable operation on embedded platforms. Figure 6 shown.

[0190] On the other hand, an embodiment of the present invention further provides a multimodal image matching system based on saliency map structure enhancement, comprising:

[0191] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the multimodal image matching method based on saliency map structure enhancement as described in the above technical solution.

[0192] It should be understood that parts not elaborated in detail in this specification belong to the prior art.

[0193] It should be understood that the above description of the preferred embodiment is relatively detailed and cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the guidance of the present invention, ordinary technicians in this field can also make substitutions or modifications without departing from the scope of protection of the claims of the present invention, which all fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.

Claims

1. A multimodal image matching method based on saliency map structure enhancement, characterized in that: The following steps are involved: Step 1: extract multi-scale structural features of image pairs; Step 2: extract multi-scale semantic features of the image pair; Step 3: Fuse the multi-scale structural features and the multi-scale semantic features to obtain the fused features; Step 4: Construct a pixel-level saliency confidence map; The specific implementation of step 4 is as follows: Sobel operator is used to calculate the horizontal and vertical direction The gradient component corresponding to each pixel , calculate the gradient amplitude : in, , W and H represent the width and height of the image; Then, the global gradient mean of the entire image is calculated and standard deviation , used for normalization processing, combined with fusion features , calculate the global mean of the gradient map, which is the saliency confidence map: in, is the sigmoid function; is the contrast enhancement coefficient; A very small number to prevent the denominator from being 0; Represents pixel-by-pixel multiplication operation; Represents pixel coordinates Confidence level of significance at ; Step 5: construct a multi-level feature enhancement network. Each layer of the multi-level feature enhancement network adopts a self-attention mechanism and a cross-attention mechanism, and uses a saliency confidence map to enhance the fusion feature to obtain the enhanced final fusion feature. The processing process of each layer in the multi-level feature enhancement network in step 5 is as follows: Step 5.1: For each image, construct the image internal structure, establish the adjacent edge relationship set between pixels, and perform a saliency weighted self-attention mechanism on the adjacent edge relationship set to obtain the self-attention feature; Step 5.2: Using the reference image as the source and the target image as the target, we construct a cross-modal guidance map from the reference image to the target image using the semantic similarity between the images, and perform saliency-weighted cross-modal feature fusion to obtain the cross-attention feature. In step 5.3, the fusion feature of each pixel is fused with the self-attention feature and the cross-attention feature to obtain the final fusion feature; Step 6: Calculate the matching confidence score based on the final fusion feature, and filter the matching point set according to the matching confidence score; Step 7: Perform geometric consistency verification based on the basic matrix, remove abnormal point pairs in the semi-dense matching point set, and output a stable matching result.

2. The multimodal image matching method based on saliency map structure enhancement according to claim 1, characterized in that: In step 1, the multi-scale structural features of the image pairs are extracted through the multi-layer convolutional neural network Unet. In step 2, the pre-trained lightweight semantic segmentation network MobileSAM is used to extract the multi-scale semantic features of the image pairs. In step 3, the multi-scale structural features extracted by Unet and the multi-scale semantic features extracted by the lightweight semantic segmentation network MobileSAM are concatenated in the channel dimension and then reduced to fusion features through 1×1 convolution.

3. The multimodal image matching method based on saliency map structure enhancement according to claim 1, characterized in that: The specific implementation of step 5.1 is as follows: First, for each image, construct the image internal structure: in, is a set of pixels, Represents the set of adjacent edge relationships between pixels; , is the pixel position; the distance threshold Control the current layer The attentional proximity of Structure within the image On the defined set of adjacent edge relationships, the following saliency-weighted self-attention mechanism is performed: Basic self-attention score Calculated from the query and key vector: Among them, query ,key Vectors are obtained by linear transformation: , and is a learnable parameter, For the The fused feature vector corresponding to the pixel points, is the rotation position encoding function; Add the saliency weight modulation factor to obtain the saliency self-attention score : in, and Indicates the and The confidence level of the pixel's saliency, is a constant; Use Softmax normalization and weighted aggregation of adjacent node features to obtain self-attention features: in, are learnable parameters, Indicates the i The set of adjacent edge relationships of pixels, For the The fused feature vector corresponding to each pixel point.

4. The multimodal image matching method based on saliency map structure enhancement according to claim 3, characterized in that: Distance threshold The calculation formula is as follows: in, is the initial global connection threshold, set to the maximum pixel spacing, for and The Euclidean distance between is the lower distance threshold, is the total number of layers, is the current level; The rotation encoding function is defined as follows: in, Indicates the Pixels relative to the The direction angle of the pixel point, Indicates the The horizontal and vertical coordinates of the pixel points, Indicates the j The horizontal and vertical coordinates of the pixel points; is a small constant that prevents division by zero.

5. The multimodal image matching method based on saliency map structure enhancement according to claim 1, characterized in that: The specific implementation of step 5.2 is as follows: Define a semantic guidance graph with reference image A as source and target image B as target : Among them, TopK means retaining the top K candidate points with the highest semantic similarity; Indicates the calculation of cosine similarity; and Represent the pixel sets of the reference image and the target image respectively; and Represents the reference image The first fusion feature vector and the target image fused feature vectors; For each pair , calculate the cross attention score : in, and are the parameters that can be learned by the network, is the rotation position encoding function; and Represents the first i pixel points and the first pixel on the target image j pixels; The saliency cross-attention score after combining saliency modulation is expressed as: in, and Represents the reference image The saliency confidence of the pixel and the target image The confidence level of the pixel's saliency, is a constant; Use Softmax normalization and weighted aggregation of semantic information from the target graph to obtain cross-attention features: in, Representation semantic guidance graph Middle i The neighboring points of pixels, are learnable parameters.

6. The multimodal image matching method based on saliency map structure enhancement according to claim 1, characterized in that: The specific implementation of step 6 is as follows: After mapping the final fusion features to the matching space, the matching score matrix is ​​constructed, and then the matching confidence score is obtained through bidirectional Softmax normalization, and the matching confidence score is calculated based on the preset threshold. Select matching point pairs that meet the conditions to form a semi-dense matching point set.

7. The multimodal image matching method based on saliency map structure enhancement according to claim 1, characterized in that: The specific implementation of step 7 is as follows: Estimate the fundamental matrix using the RANSAC algorithm and the eight-point method , calculate the epipolar constraint error for each pair of matching points in the semi-dense matching point set as a measure of geometric consistency: in, and Respectively represent the first matching point and the first matching point in the target image B Homogeneous coordinates of matching points; represents the fundamental matrix; Represents the calculated epipolar error. Point pairs with errors less than the threshold are considered as geometric inliers to form the final set of matching points.

8. A multimodal image matching system based on saliency map structure enhancement, characterized in that: include: A processor and a memory, the memory being used to store program instructions, and the processor being used to call the stored instructions in the memory to execute the multimodal image matching method based on saliency map structure enhancement as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Optical remote sensing image saliency detection method and system based on contrast enhancement and hierarchical attention fusion, storage medium and electronic equipment

    CN119672556A

  • Image segmentation system via graph or multiscale cascaded attention decoding

    US20250139775A1