Multi-modal image matching method and system based on sub-pixel deviation estimation

By using subpixel deviation estimation technology in multimodal image matching, multiple scale features are extracted and fine-grained descriptors are generated, the problems of subpixel-level accuracy and position deviation in multimodal image matching are solved, and high-precision and robust image matching are achieved.

CN120088516AActive Publication Date: 2025-06-03HUNAN UNIV

Patent Information

Application Number
CN202510538440.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-06-03
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

When the existing multimodal image matching technology processes different sensor images, it is difficult to achieve subpixel-level accuracy matching, and there is a problem of position deviation.

Method used

Using a multimodal image matching method based on subpixel deviation estimation, a pre-trained multimodal image matching network is used to extract feature tensors of multiple scales for images of different modes, fuse the coarse-grained descriptors of image key points, and further extract the fine-grained descriptors. Then, the similarity between coarse-grained descriptors of the image keys is calculated to generate coarse-grained matching point pairs, and feature point blocks are cropped by the fine-grained descriptor, and the association heat map is calculated to estimate the subpixel-level offset between images.

Benefits of technology

Subpixel-level accuracy matching in multimodal image matching is realized, reducing position deviation and improving the accuracy and robustness of image matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088516A_ABST
    Figure CN120088516A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image matching method and system based on sub-pixel deviation estimation, and the method comprises the steps: inputting two input images of different modals into a pre-trained multi-modal image matching network, so as to obtain the offset between the two input images, comprising the steps that feature tensors of multiple scales are extracted from two input images respectively, coarse-grained descriptors of image key points are obtained through fusion, and fine-grained descriptors are further extracted; calculating the similarity between the coarseness descriptors of the image key points of the two input images to generate a coarseness matching point pair; and the coarse-grained matching point pairs of the two input images are combined with the fine-grained descriptors to cut the corresponding feature point blocks, an associated heat map between the feature point blocks is calculated, and the offset between the two input images is estimated. The method aims at solving the problems of modal change and geometric distortion of different sensor images, matching of sub-pixel-level precision is achieved, and the image matching precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to visual image processing technology in the field of computers, and particularly relates to a multi-modal image matching method and system based on sub-pixel deviation estimation. Background Art

[0002] Image matching is a basic technology in computer vision and is widely applied in fields such as image fusion, image restoration, and depth estimation. In each field, methods for accurately matching images from different sensors have important application values. For example, in power inspection, by combining visible light images and infrared images, information such as the shape, state, and working temperature of power equipment can be comprehensively obtained, providing more reliable data support for fault detection. However, how to accurately match images from different sensors has become a key problem in technical implementation. Multi-modal image matching has become an important research topic. Although significant progress has been made in multi-modal image matching in the past few decades, the position deviation between the estimated matching points and the ground truth points remains a problem to be solved.

[0003] The position deviation in multi-modal image matching is mainly caused by two reasons: the detector extracts non-corresponding key points, which may lead to the key points being matched to adjacent pixels of the ground truth, thus generating position deviation; the extracted descriptors may lack sufficient robustness to the geometric distortion and modal changes of the multi-modal image pair, resulting in many incorrect matches and generating position deviation. To solve this problem, there are two major categories of existing research works, each with different limitations: The first category of methods is based on manually designed methods. Manually designed methods mainly focus on designing feature detectors and descriptors that are robust to geometric distortion or modal changes to address the multi-modal image matching problem. Typical methods such as 3MRS fall into this category. This method refines the coarse-grained matching with position deviation. The refinement module then calculates the template features through the phase consistency (PC) map for matching to improve the matching performance. Although this method can significantly improve the matching accuracy in some cases, its robustness to modal changes is limited. And the MSFT method, which obtains the initial transformation matrix for coarse-grained matching through the fast sample consensus (FSC) method and uses a hard elimination threshold, and then adjusts the position of the coarse-grained matching to minimize the transformation error based on the initial matrix, thereby enhancing the transformation consistency of the adjusted matching and reducing the position deviation. Although this method can reduce the position deviation, it depends on the accuracy of the initial transformation matrix. If the initial matrix is inaccurate, the adjustment effect may be greatly reduced; The second category of methods is based on learning methods. Among them, detector-based methods such as Keypt2Subpx can refine any detector by predicting the offset vector of the detected key points to obtain sub-pixel-level matching. However, when used in conjunction with sub-pixel-level detectors (such as SIFT), the performance of the method will decline and may not achieve the best effect. Detector-free methods use the self-attention and cross-attention layers in the Transformer to obtain feature descriptors, thereby obtaining a global receptive field to generate dense matches. However, the dense matches and attention layers in the Transformer involve a large number of pixels, resulting in a significant increase in the computational requirements during training. Summary of the Invention

[0004] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, a multi-modal image matching method and system based on sub-pixel deviation estimation are provided. The present invention aims to achieve sub-pixel-level accuracy matching and improve the image matching accuracy for the problems of modal changes and geometric distortion of different sensor images.

[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows: A multi-modal image matching method based on sub-pixel deviation estimation, which includes inputting two input images of different modalities into a pre-trained multi-modal image matching network to obtain the offset between the two input images. The processing of the two input images by the multi-modal image matching network includes: extracting feature tensors of multiple scales from the two input images respectively, fusing to obtain a coarse-grained descriptor of image key points and further extracting a fine-grained descriptor; calculating the similarity between the coarse-grained descriptors of the image key points of the two input images to generate coarse-grained matching point pairs; combining the coarse-grained matching point pairs of the two input images with the fine-grained descriptors to crop the corresponding feature point blocks, calculating the correlation heat map between the feature point blocks and estimating the offset between the two input images.

[0006] Optionally, the functional expression for respectively extracting feature tensors of multiple scales from the two input images, fusing to obtain a coarse-grained descriptor of image key points and further extracting a fine-grained descriptor is: , , wherein, is the input image, is the encoder for extracting feature tensors of multiple scales, and the encoders of the two input images of different modalities do not share network parameters, is the operation of fusing to obtain a coarse-grained descriptor of image key points, is the coarse-grained descriptor of image key points, is the descriptor head composed of a multi-layer perceptron, is the fine-grained descriptor of image key points.

[0007] Optionally, the feature tensors of multiple scales respectively include three feature tensors with sizes of , , where and are the height and width of the input image respectively; the fusion to obtain a coarse-grained descriptor of image key points includes bilinearly upsampling the two feature tensors of , to the size of and then summing element-wise with the feature tensor with a size of to obtain the coarse-grained descriptor of image key points.

[0008] Optionally, the calculation of the similarity between the coarse-grained descriptors of the image key points of the two input images to generate coarse-grained matching point pairs includes: respectively using the detection head for the coarse-grained descriptors of the image key points of the two input images to obtain the key point detection score map , the detection head It is composed of a convolutional layer and an activation function; for the key point detection score map The top K image key points with higher scores are obtained through a non-maximum suppression detector, a confidence matrix is generated according to the coarse-grained descriptors of the top K image key points with the highest scores in the two input images, and the point pairs with the maximum mutual confidence in the confidence matrix are used as the coarse-grained matching point pairs.

[0009] Optionally, the functional expression for generating the confidence matrix is: , where is the confidence matrix, and are the coarse-grained descriptors of the input images and the input image respectively, is the transpose operation.

[0010] Optionally, the combining of the coarse-grained matching point pairs of the two input images with the fine-grained descriptors to crop the corresponding feature point blocks, calculate the correlation heat map between the feature point blocks, and estimate the offset between the two input images includes: for each coarse-grained matching point pair of the two input images , at one of the coarse-grained matching points a feature point block is cropped with a given radius r, and the fine-grained descriptor points are used to crop a feature point block in the area centered on the other coarse-grained matching point with a given radius r; the correlation heat map between the feature point block and the feature point block is calculated according to the following formula: , where is the softmax activation function, is the transpose operation; The offset between the two input images is estimated according to the following formula: , where is the offset between the two input images, is the feature point block area with radius r, is the coordinate in the feature point block area, is the correlation heat map and is the score value at the coordinate

[0011] Optionally, the functional expression of the loss function adopted by the multi-modal image matching network during training is: , , Among them, is the loss function is the coarse-grained loss function, is the detection loss function, is the fine-grained loss function, is the sufficiency loss, is the peak loss, is the coupling loss function, and the functional expression of the coarse-grained loss function is: , Among them, and respectively represent the calculated losses in the row and column directions, and represent the confidence matrix on the diagonal of row column, the similarity of the coarse-grained descriptions of the row column; the functional expression of the fine-grained loss function is: , Among them, is the number of coarse-grained matching point pairs of two input images, is the inverse of the variance of the associated heatmap corresponding to the j-th coarse-grained matching point pair of represents the operation of taking the average value, is the inverse of the variance of the associated heatmaps corresponding to all coarse-grained matching point pairs of is the true offset between two input images, is the offset between two input images, and there is: , , Among them, and are the image key points belonging to the input image and the input image respectively in a coarse-grained matching point pair, is the image key point in the input image in the input image corresponding point, is the input image and the input image ​The transformation matrix therebetween; the functional expression of the coupling loss function is as follows: , wherein, and are respectively the pixel sets of two input images, is the confidence matching probability, is the input image corresponding key point detection score map the score of pixel in, is the input image corresponding key point detection score map the score of pixel in, and there is: , wherein, and are respectively the maximum values calculated in the row and column directions.

[0012] In addition, the present invention further provides a multi-modal image matching system based on sub-pixel deviation estimation, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the multi-modal image matching method based on sub-pixel deviation estimation.

[0013] In addition, the present invention further provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the multi-modal image matching method based on sub-pixel deviation estimation through a processor.

[0014] In addition, the present invention further provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the multi-modal image matching method based on sub-pixel deviation estimation through a processor.

[0015] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: The present invention includes extracting feature tensors of multiple scales from two input images respectively to cope with modality changes, fusing the feature tensors of multiple scales to obtain a coarse-grained descriptor of image key points and further extracting a fine-grained descriptor to realize the processing of geometric distortion, calculating the similarity between the coarse-grained descriptors of the image key points of the two input images to generate coarse-grained matching point pairs; on the basis of the coarse-grained matching point pairs, a sub-pixel offset estimation method is designed, including: combining the coarse-grained matching point pairs of the two input images with the fine-grained descriptors to crop the corresponding feature point blocks, calculating the correlation heat map between the feature point blocks and estimating the offset between the two input images, so as to accurately adjust the position deviation in the coarse-grained matching, thereby improving the matching accuracy. The present invention can achieve accurate matching in multi-modal image matching, and helps to solve the position offset problem between the estimated matching points and the ground truth points. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 FIG. is a schematic diagram of the basic flow of the method according to the embodiment of the present invention.

[0017] Figure 2 FIG. is a schematic diagram of the network structure of the multi-modal image matching network in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0019] As Figure 1 shown, this embodiment provides a multi-modal image matching method based on sub-pixel deviation estimation, including inputting two input images of different modalities into a pre-trained multi-modal image matching network to obtain the offset between the two input images. The processing of the two input images by the multi-modal image matching network includes: extracting feature tensors of multiple scales from the two input images respectively, fusing them to obtain a coarse-grained descriptor of image key points and further extracting a fine-grained descriptor; calculating the similarity between the coarse-grained descriptors of the image key points of the two input images to generate coarse-grained matching point pairs; combining the coarse-grained matching point pairs of the two input images with the fine-grained descriptors to crop the corresponding feature point blocks, calculating the correlation heat map between the feature point blocks and estimating the offset between the two input images.

[0020] See Figure 2, in this embodiment, the multi-modal image matching network can be divided into three stages: multi-scale feature extraction, pixel-level feature matching, and sub-pixel level offset estimation. Among them, in the multi-scale feature extraction stage, feature tensors of multiple scales are extracted from two input images respectively, fused to obtain a coarse-grained descriptor of image key points, and further a fine-grained descriptor is extracted. In the pixel-level feature matching stage, the similarity between the coarse-grained descriptors of the image key points of the two input images is calculated to generate coarse-grained matching point pairs. In the sub-pixel level offset estimation stage, the coarse-grained matching point pairs of the two input images are combined with the fine-grained descriptors to crop the corresponding feature point blocks, calculate the correlation heat map between the feature point blocks, and estimate the offset between the two input images.

[0021] In this embodiment, the functional expression for extracting feature tensors of multiple scales from two input images respectively, fusing to obtain a coarse-grained descriptor of image key points, and further extracting a fine-grained descriptor is: , , Among them, is the input image, is the encoder for extracting feature tensors of multiple scales, and the encoders of two input images of different modalities do not share network parameters. is the operation of fusing to obtain a coarse-grained descriptor of image key points. is the coarse-grained descriptor of image key points. is the descriptor head composed of multi-layer perceptrons. is the fine-grained descriptor of image key points. The encoder for extracting feature tensors of multiple scales can adopt the required encoder type according to needs. For example, as an alternative implementation, refer to Figure 2 , in this embodiment, an adaptive encoder is adopted; the multi-layer perceptron that constitutes the descriptor head is composed of multiple fully connected layers plus batch normalization and activation functions.

[0022] Two input images of different modalities can be expressed as: , and their corresponding pixels are and , and The geometric transformation between is as follows: Among them, represents the geometric transformation matrix between image pairs.

[0023] As an alternative implementation, in this embodiment, the feature tensors of multiple scales respectively include sizes of , , Three feature tensors, where and are the height and width of the input image respectively; the fused coarse-grained descriptor of the image key points includes , The two feature tensors of are bilinearly upsampled to and then element-wise summed with the one of size Figure 2 to obtain the coarse-grained descriptor of the image key points. As an alternative implementation, see Figure 2 , in this embodiment, the fused coarse-grained descriptor of the image key points is specifically implemented through a multi-scale fusion module. For a bimodal input image, two non-shared-weight adaptive encoders are used to extract the common information of different-modal images, and then a multi-scale feature fusion module is adopted to improve the robustness of the descriptor to geometric distortion. The coarse-grained descriptor is obtained through the multi-scale feature fusion module and input into the descriptor head to obtain the fine-grained descriptor.

[0024] In this embodiment, calculating the similarity between the coarse-grained descriptors of the image key points of two input images to generate coarse-grained matching point pairs includes: respectively using the detection head for the coarse-grained descriptors of the image key points of the two input images to obtain the key point detection score map , the detection head is composed of a convolutional layer plus an activation function. The 128-dimensional coarse-grained descriptor becomes a one-dimensional score matrix after passing through the detection head ; for the key point detection score map , the top K image key points with higher scores are extracted through a non-maximum suppression detector, and a confidence matrix is generated according to the coarse-grained descriptors of the top K image key points with the highest scores in the two input images. The point pairs with the maximum mutual confidence in the confidence matrix are used as the coarse-grained matching point pairs. In this embodiment, the functional expression for generating the confidence matrix is: , where, is the confidence matrix, and are the coarse-grained descriptors of the input image and the input image respectively, is the transpose operation.

[0025] In this embodiment, a sub-pixel offset estimation method is designed based on the coarse-grained matching points and combined with the fine-grained descriptors to accurately adjust the position deviation in the coarse-grained matching to improve the matching accuracy. Specifically, for each pair of coarse-grained matching points of the two input images, the steps of combining the fine-grained descriptors to crop the corresponding feature point patches, calculating the correlation heat map between the feature point patches, and estimating the offset between the two input images include: , crop out a feature point patch with a given radius r at one of the coarse-grained matching points , and crop out a feature point patch with a given radius r in the area centered on the other coarse-grained matching point ; calculate the correlation heat map between the feature point patch and the feature point patch according to the following formula: , where, is the softmax activation function, is the transpose operation; Estimate the offset between the two input images according to the following formula: , where, is the offset between the two input images, is the feature point patch area (patch) with radius r, is the coordinate in the feature point patch area, is the correlation heat map and is the score value at the coordinate

[0026] In this embodiment, the function expression of the loss function adopted by the multi-modal image matching network during training is: , , where, is the loss function is the coarse-grained loss function, is the detection loss function, is the fine-grained loss function, is the sufficiency loss, is the peak loss, is the coupling loss function.

[0027] For any pair of bimodal images in the dataset, in order to obtain pixel-level coarse-grained matching, it is necessary to align the descriptors of the corresponding pixels. Even in the case of geometric deformation and modality change, the similarity of the descriptors of the corresponding pixels should be maximized. To ensure that the similarity of the descriptors of the corresponding pixels is mutually maximized, the system uses the dual-softmax loss in the confidence matrix. Therefore, the functional expression of the coarse-grained loss function in this embodiment is: , where and respectively represent the calculated losses in the row and column directions, and represent the confidence matrix on the diagonal of the th row th column, the th row th column of the similarity of the coarse-grained descriptions; the functional expression of the fine-grained loss function is: , where is the number of coarse-grained matching point pairs of the two input images, is the variance of the associated heatmap corresponding to the jth coarse-grained matching point pair inverse, represents the operation of taking the average value, is the variance of the associated heatmaps corresponding to all coarse-grained matching point pairs inverse, is the true offset between the two input images, is the offset between the two input images, and there is: , , where and are the image key points belonging to the input image and the input image respectively in a coarse-grained matching point pair, is the image key point in the input image corresponding point in the input image , is the input image and the input image transformation matrix between. In order to estimate the sub-pixel level position offset of the coarse-grained matching, it is necessary to pair the coarse-grained matching key points ​Sub-pixel alignment is established for supervision. Additionally, considering that there are some low-quality correspondences in the coarse-grained matching, the system introduces an interpretable uncertainty measure as the weight of the fine-grained loss function. For each coarse-grained match, the variance of the corresponding heatmap is calculated as the uncertainty measure. A high variance indicates the presence of multiple or scattered patterns, reflecting low-quality or ambiguous predictions. Therefore, the fine-grained loss function optimizes the distance between the estimated position offset and the ground truth offset to achieve sub-pixel accuracy, thereby improving the matching performance.

[0028] To detect repeatable key points at corresponding positions, the detection score maps of the input images and should be similar. Additionally, to ensure that key points within a local region can be detected, the scores of the key points should be significant. Therefore, the detection loss function includes a repeatability loss, a peak loss, and a coupling loss function. The functional expression of the repeatability loss is: , where is the number of corresponding pixel pairs , and are the detection score maps of the input images and respectively. are the feature point blocks centered at with a radius of respectively. The repeatability loss optimizes the scores of the corresponding patches to generate repeatable key points. The peak loss enhances the significance of the key points. The functional expression of the peak loss is: , where is the set of overlapping patches of size with a stride of 8 in the image, is the number of patches. The peak loss promotes the introduction of significant points within a local region by minimizing the mean of the patch scores and maximizing the maximum value.

[0029] Although repeatable key points are extracted by optimizing the repeatability loss and the peak loss, low-quality key points with a low matching success rate are also detected, which reduce the matching effect. Previous studies have shown that combining detection and description can suppress these low-quality key points. Therefore, the system designs a coupling loss function based on the confidence matrix and the detection score map to couple feature detection and description. The functional expression of this coupling loss function is: , Among them, and are respectively the pixel sets of two input images, is the confidence matching probability, is the input image corresponding key point detection score map the score of the pixel in it, is the input image corresponding key point detection score map the score of the pixel in it, and there is: and Among them, and respectively calculate the maximum values in the row and column directions. It should be noted that, The gradient of is only backpropagated through

[0030] The coupled loss function improves the detection scores of pixels with high confidence matching probabilities, which is helpful for suppressing low-quality key points and detecting key points with high matching success rates.To verify the multi-modal image matching method based on sub-pixel deviation estimation in this embodiment, the development language used in this embodiment is Python, and the deep learning framework is Pytorch. The model is trained on a machine equipped with an i9-11900k CPU (3.50GHz×16) and an NVIDIA GeForce RTX 3090 graphics card. A multi-modal image dataset collected by ReDFeat is used. This multi-modal image dataset includes three types of cross-modal images, namely visible light-synthetic aperture radar (VIS-SAR) images, visible light-infrared (VIS-IR) images, and visible light-near infrared (VIS-NIR) images: The VIS-SAR dataset includes 2011 pairs and 424 pairs of aligned images, with a unified size of 512×512, for training and testing. This dataset is captured by satellites and includes scenes from fields and urban areas. The VIS-IR images come from the RGB-LWIR dataset and the RoadScene dataset. The RGB-LWIR dataset provides 44 pairs of static images, mainly of buildings during the day. The RoadScene dataset provides 221 pairs of images, containing many objects such as people and cars, covering both day and night. The VIS-IR dataset has a total of 265 pairs of aligned images, with an average size of 533×321. 47 images are randomly selected as the test set, and the remaining image pairs are used for training. The VIS-NIR images are provided by the RGB-NIR scene dataset. This dataset contains 9 scenes, forming 477 pairs of aligned images, with an average size of 983×686, including scenes such as rural areas, fields, forests, indoors, mountains, old buildings, streets, cities, and waters. The VIS-NIR images are randomly divided into a training set and a test set according to a ratio of 3:1. The training set has 345 pairs, and the test set has 128 pairs. It should be noted that the test set is manually verified and screened to ensure more reliable ground truth. During training, the Adam optimizer is used for training, with a total of 120,000 steps, and the initial learning rate is 3.0×10 -4 , and the batch size is 16. In this embodiment, the matching success rate (SR), the number of correct matches (NCM), the mean error (ME), and the root mean square error (RMSE) are used as evaluation metrics, and the fine-grained matching point pairs are calculated residuals , and the residuals less than or equal to three pixels are regarded as correct matches. The expression for the residuals is: , where and are the image key points in the fine-grained matching point pair that belong to the input image and the input image respectively, is the input image and the input image is the transformation matrix between them. The number of correct matches is the number of correct matches NCM. If the number of correct matches NCM > 10, it is considered a successful match. The matching success rate SR represents the percentage of the number of successfully matched image pairs in the total number of image pairs. The average error ME and the root mean square error RMSE represent the matching error and the error fluctuation respectively, and their functional expressions are: , , where is the th residual .

[0031] In the sub-pixel level offset estimation stage, the radius r of the feature point block area (patch) is an important parameter that determines the offset estimation range, thus affecting the performance of the system. In this embodiment, training with different radii r was carried out on the aforementioned dataset, and the final results are shown in Table 1.

[0032] Table 1: Training results with different radii r

[0033] In Table 1, "↑" indicates that the larger the index, the better, and "↓" indicates that the smaller the index, the better. Referring to Table 1, it can be seen that when the radius r is 2px, the matching success rate (SR), the number of correct matches (NCM), the average error (ME), and the root mean square error (RMSE) all achieve the optimal results.

[0034] In order to verify the effectiveness of each module in the multi-modal image matching network of the method of this embodiment, in this embodiment, the method of this embodiment and the following variants of the method of this embodiment were trained and tested on the aforementioned dataset: the method without the multi-scale feature fusion block (W / O-FB), the method without the detector and descriptor coupling (W / O-CP), and the method without the sub-pixel level offset estimation (W / O-OE). Among them, the multi-scale feature fusion module performs feature fusion operations on the image after passing through the adaptive encoder, and fuses , , the feature tensors of three scales. Multi-scale feature fusion can improve the robustness and accuracy of matching when dealing with complex scenarios such as scale changes, insufficient texture, and occlusion. The method without the multi-scale feature fusion block (W / O-FB) directly The feature tensor passes through a weight - shared fully convolutional encoder to replace the multi - scale feature fusion module and outputs a coarse - grained descriptor, thereby reducing the adaptability to geometric distortion and resulting in a decline in results compared to the method of this example. The detector - descriptor coupling method generates a detection score map by performing feature detection on the coarse - grained descriptor. Through the detection score map, unreliable key points with a low matching success rate can be suppressed. The method without detector - descriptor coupling (W / O - CP) abandons the coupling of the detector and the descriptor, and the feature descriptor is not filtered by the score map generated by the detector. This means that all feature points will participate in the matching, which causes many key points with a low matching success rate to be detected, leading to a performance decline. The method without sub - pixel level offset estimation (W / O - OE) deletes the sub - pixel level offset estimation module and directly uses the pixel - level rough matching result for feature matching, making the obtained pixel - level rough matching result not fine enough, thus reducing the matching result. The final results are shown in Table 2.

[0035] Table 2: Training results of the method of this embodiment (this method) and its variants on the VIS - SAR dataset

[0036] In Table 2, "↑" indicates that the larger the index, the better, and "↓" indicates that the smaller the index, the better. Referring to Table 2, it can be seen that the method of this embodiment has achieved the best results in terms of matching success rate (SR), number of correct matches (NCM), mean error (ME), and root mean square error (RMSE) compared to its variants.

[0037] To quantify the improvement of sub - pixel level offset estimation on sub - pixel accuracy, in this embodiment, the performance of the method of this embodiment and the method without sub - pixel level offset estimation (W / O - OE) is compared on the VIS - SAR dataset, VIS - IR dataset, and VIS - NIR dataset. The number of matches representing sub - pixel accuracy, and the final results are shown in Table 3.

[0038] Table 3: Comparison results between the method of this embodiment (this method) and the method without sub - pixel level offset estimation

[0039] In Table 3, "↑" indicates that the larger the index, the better, and "↓" indicates that the smaller the index, the better. Referring to Table 3, it can be seen that the sub - pixel level offset estimation in the multi - modal image matching network of the method of this embodiment improves the results of all the indexes listed in Table 3, and the number of matches with sub - pixel accuracy also increases significantly, especially on the VIS - NIR dataset, increasing by approximately 125%.

[0040] In addition, this embodiment also provides a multimodal image matching system based on sub-pixel deviation estimation, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the multimodal image matching method based on sub-pixel deviation estimation.

[0041] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the multimodal image matching method based on sub-pixel deviation estimation through a processor.

[0042] In addition, this embodiment also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the multimodal image matching method based on sub-pixel deviation estimation through a processor.

[0043] Those skilled in the art should understand that the technical solution provided by the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide means for realizing the functions in the processFigure 1 One process or multiple processes and / or boxes Figure 1 Steps of functions specified in one box or multiple boxes.

[0044] The above is only the preferred embodiment of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.

Claims

1. A multimodal image matching method based on sub-pixel deviation estimation, characterized in that: The method includes inputting two input images of different modalities into a pre-trained multimodal image matching network to obtain the offset between the two input images. The multimodal image matching network processes the two input images by: extracting feature tensors of multiple scales from the two input images respectively, fusing to obtain coarse-grained descriptors of image key points and further extracting fine-grained descriptors; calculating the similarity between the coarse-grained descriptors of the image key points of the two input images to generate coarse-grained matching point pairs; combining the coarse-grained matching point pairs of the two input images with the fine-grained descriptors to crop corresponding feature point blocks, calculating the correlation heat map between the feature point blocks and estimating the offset between the two input images.

2. The multimodal image matching method based on sub-pixel deviation estimation according to claim 1, characterized in that: The function expression for extracting feature tensors of multiple scales from two input images, fusing them to obtain coarse-grained descriptors of image key points, and further extracting fine-grained descriptors is: , , in, is the input image, is an encoder for extracting feature tensors of multiple scales, and the encoders of two input images of different modalities do not share network parameters. To fuse the coarse-grained descriptors of the image key points, is a coarse-grained descriptor of image key points, The descriptor header for the multi-layer perceptron, A fine-grained descriptor for image key points.

3. The multimodal image matching method based on sub-pixel deviation estimation according to claim 1, characterized in that: The feature tensors of the multiple scales include , , The three feature tensors of and are the height and width of the input image respectively; The fusion method for obtaining the coarse-grained descriptor of the image key point includes: , The two feature tensors are bilinearly upsampled to Size and size The element-by-element summation yields a coarse-grained descriptor of the image key points.

4. The multimodal image matching method based on sub-pixel deviation estimation according to claim 1, characterized in that: The method of calculating the similarity between the coarse-grained descriptors of the image key points of the two input images to generate the coarse-grained matching point pairs comprises: respectively using the detection head to calculate the similarity between the coarse-grained descriptors of the image key points of the two input images Get the key point detection score map , the detection head It is composed of a convolutional layer plus an activation function; for key point detection score map The first K image key points with higher scores are extracted through the non-maximum suppression detector, and the confidence matrix is ​​generated according to the coarse-grained descriptors of the K image key points with the highest scores in the two input images. The point pairs with the largest confidence matrices are taken as coarse-grained matching point pairs.

5. The multimodal image matching method based on sub-pixel deviation estimation according to claim 4, characterized in that: The function expression for generating the confidence matrix is: , in, is the confidence matrix, and The input images are and the input image The coarse-grained descriptor of is the transpose operation.

6. The multimodal image matching method based on sub-pixel deviation estimation according to claim 4, characterized in that: The method of combining the coarse-grained matching point pairs of the two input images with the fine-grained descriptors to crop the corresponding feature point blocks, calculating the correlation heat map between the feature point blocks and estimating the offset between the two input images comprises: for each coarse-grained matching point pair of the two input images , at one of the coarse-grained matching points Cut out the feature point block with a given radius r , and place the fine-grained descriptor point on another coarse-grained matching point The feature point block is cut out with a given radius r in the center area ; Calculate the feature point block according to the following formula , feature point block Heatmap of the correlation between : , in, is the softmax activation function, is the transpose operation; The offset between the two input images is estimated according to the following formula: , in, is the offset between the two input images, is the feature point block area with radius r, is the coordinate in the feature point block area, Correlation heat map Central coordinates The score value at .

7. The multimodal image matching method based on sub-pixel deviation estimation according to claim 1, characterized in that: The functional expression of the loss function used by the multimodal image matching network during training is: , , in, is the loss function is the coarse-grained loss function, To detect the loss function, is the fine-grained loss function, For the loss of sufficiency, is the peak loss, is the coupling loss function, and the function expression of the coarse-grained loss function is: , in, and Respectively represent the calculation loss in the row and column directions, and Represents the confidence matrix The diagonal line of OK Column, No. OK The similarity of the coarse-grained description of the column; the functional expression of the fine-grained loss function is: , in, is the number of coarse-grained matching point pairs between the two input images, is the variance of the associated heat map corresponding to the jth coarse-grained matching point pair The inverse, represents the operation of taking the average value, is the variance of the associated heatmap corresponding to all coarse-grained matching point pairs The inverse, is the actual offset between the two input images, is the offset between the two input images, and: , , in, and For a coarse-grained matching point pair, and the input image The key points of the image, For the input image Image key points in On the input image The corresponding point in For the input image and the input image The transformation matrix between them; the functional expression of the coupling loss function is: , in, , are the pixel sets of the two input images respectively, is the confidence matching probability, For the input image Corresponding key point detection score map Medium Pixels The score, For the input image Corresponding key point detection score map Medium Pixels The score of , and there are: , in, and Calculate the maximum values ​​in row and column directions respectively.

8. A multimodal image matching system based on sub-pixel deviation estimation, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the multimodal image matching method based on sub-pixel deviation estimation as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the multimodal image matching method based on sub-pixel deviation estimation as claimed in any one of claims 1 to 7 through a processor.

10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the multimodal image matching method based on sub-pixel deviation estimation as claimed in any one of claims 1 to 7 through a processor.

Citation Information

Patent Citations

  • Multi-modal image feature matching method based on multi-feature matching

    CN115496928A

  • Coarse-to-fine different-source image matching method based on edge guidance

    CN118135256A

  • SAR and optical image matching method and device

    CN118470074A

  • Cross-modal external ear region segmentation and key point positioning method and system

    CN119624993A

  • Method and system for registering images acquired with different modalities for generating fusion images from registered images acquired with different modalities

    US20230281837A1

Cited By

  • Circuit pattern feature matching method based on dense matching

    CN121640106A