Heterogeneous image matching method based on cross-modal conversion network and optimal transfer theory
By constructing a cross-modal Transformer matching network and combining the optimal transmission theory, the problem of insufficient accuracy and speed in heterologous image matching is solved, and the effects of high accuracy and fast matching are achieved.
Patent Information
- Application Number
- CN202210998060.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-08-19
AI Technical Summary
The prior art has shortcomings in matching accuracy and matching speed in heterologous image matching, especially under the differences in grayscale distribution between different mode images and noise interference, making it difficult to achieve high accuracy and fast matching.
A heterologous image matching method based on cross-modal attention and optimal transmission theory is adopted. By constructing an end-to-end cross-modal Transformer matching network, the matching results are optimized in combination with the optimal transmission theory to improve the accuracy and speed of matching.
It achieves higher matching accuracy and smaller matching errors, significantly improves matching speed, and can better adapt to different land scenes and improves generalization capabilities.
Smart Images

Figure CN115331029B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer vision image processing, and in particular relates to a heterogeneous image matching method, which can be used for auxiliary guidance of aircraft. Background Art
[0002] With the development of technology, remote sensing information has shown the characteristics of multi-sensor, multi-modal and large data volume. Obtaining information from massive remote sensing images has become an important information channel. Different satellite-borne sensors can obtain remote sensing data of different modes. The remote sensing images obtained by traditional visible light remote sensing systems adopt passive imaging mode, receiving electromagnetic radiation reflected and scattered after sunlight hits the surface target. The semantics are clear and intuitive, and they are the most commonly used remote sensing image type. However, due to the limitations of passive sensors, the performance of optical remote sensing at night and under the cover of clouds and fog will be greatly affected. With the continuous development of synthetic aperture radar SAR technology, it has been widely used in geographic surveying and military reconnaissance. Compared with traditional visible light band remote sensing technology, SAR uses active sensors to emit microwave band radiation and receive echoes, so SAR has all-day and all-weather observation capabilities, and is not affected by atmospheric clouds. Traditional visible light remote sensing images can make up for the problem that SAR images are not intuitive in semantics, and SAR can supplement the observation capabilities of visible light sensors at night. Images of different modes contain different electromagnetic scattering characteristics and geometric spatial information of the same object, so combining heterogeneous SAR and visible light images is of great significance for practical applications. Template matching is an image processing technology that finds the exact position of a small-sized image in a given large-sized image, and is used in many scenarios. As for the matching of heterogeneous images, due to the difference in modality, the prominence of the same object in different modalities is different; secondly, due to the SAR imaging method, the SAR image itself has a large amount of multiplicative noise, which together increase the difficulty of matching heterogeneous images.
[0003] The existing multimodal image matching methods are mainly divided into traditional methods and neural network-based methods.
[0004] As for traditional methods, they are mainly divided into two categories:
[0005] One type of traditional method directly uses the pixel grayscale information of the image, and uses the normalized cross-correlation NCC and mutual information MI between the grayscale of the two images as similarity metrics to find the corresponding matching position according to the grayscale information of the images of different modalities. Liang et al. used the spatial mutual information method combined with the ant colony optimization algorithm to realize the local area similarity measurement between images; Patel et al. proposed a method based on maximum likelihood estimation to calculate mutual information in order to improve the speed of the matching method based on mutual information. The grayscale-based method has a simple starting point and is easy to implement, but since the grayscale distribution of the same area in images of different modalities may be greatly different, this type of method cannot adapt well to the matching between multi-modal images. On the one hand, the direct use of similarity measurement criteria needs to adapt to the changes caused by image grayscale distortion, and on the other hand, it needs to accurately distinguish the differences between different objects. There is a conflict between these two requirements. The changes caused by grayscale distortion and the differences between objects cannot be distinguished by grayscale values. Moreover, for heterogeneous images, the grayscale mapping between images cannot reflect stable regularity, so it has great limitations.
[0006] Another traditional method is based on manually designed image features. After extracting feature descriptors from two images, the similarity of feature descriptors is calculated, and the position with the greatest similarity is obtained as the matching position based on the calculated similarity measure. This type of method is widely used on homologous images, such as the widely used scale-invariant feature transform SIFT feature descriptor. In addition. Many scholars have developed feature descriptors for heterogeneous images. Ye et al. proposed the phase consistency histogram HOPC, which uses a phase consistency model with illumination and contrast invariance to construct a geometric structure feature descriptor, and matches based on the structural features between the heterogeneous images. Xiang et al. focused on solving the differences between modalities and used the modal specific gradient operator in the Harris scale space, which can better deal with the matching errors caused by the difference in radiation intensity of the same area in different modalities. The proposal of the manually designed feature descriptor has good mathematical interpretability and usually has high performance under the assumption of the descriptor, but the situation in the actual application scenario is complex and changeable, and the assumed prerequisites may not be guaranteed to be met. Especially in areas where the terrain scenes themselves are more complex, the image has a larger amount of information, the texture details are more complex, and the noise interference in the imaging and other factors together make it difficult for manual design methods to achieve ideal results in practical applications.
[0007] Matching methods based on deep learning have made great progress in recent years. In essence, deep learning is also a feature-based method, but unlike traditional methods, deep features are features that are abstracted and extracted from a large amount of training data during the training process, rather than manually designed. End-to-end training and end-to-end reasoning can be achieved based on deep learning. At the same time, due to the powerful feature extraction ability of deep models, the extracted deep features are usually more consistent with the actual data distribution than manually designed features. Han et al. proposed a matching network MatchNet, which extracts features through a convolutional neural network, and then uses the connection of several fully connected layers to use the output results as a measure of the degree of matching. Merkle et al. proposed a twin network structure, and the relative displacement between the template image and the source image is used to determine the matching position. Mou et al. defined matching as a binary classification problem and trained a pseudo-twin network to predict the central pixel correspondence between SAR and optical patches. Citak proposed an attention mechanism using SAR and optical visual saliency maps as the feature extraction arm of the twin matching network. Wang et al. used a self-learning deep neural network to directly learn the mapping between the source image and the reference image, with the aim of applying the mapping for remote sensing image registration. Hoffmann et al. trained a fully convolutional network (FCN) to learn a similarity metric that is invariant to small affine transformations between SAR and optical patch pairs. Ma et al. proposed an accurate registration method based on feature extraction from a fine-tuned VGG16 model.
[0008] Although the above-mentioned deep learning-based matching methods have greatly improved the matching accuracy, the disadvantage of this type of method is that if you want to find the position of the template image in the source image, you need to perform pixel-by-pixel sliding window calculations and find the matching position by judging whether each pair of image blocks matches. This approach is applied to large-size images, which will not only greatly increase the matching time, but also make it difficult to distinguish between the image blocks at the correct matching position and similar image blocks in the neighborhood around the correct position, resulting in large errors at the pixel level. Summary of the invention
[0009] The purpose of the present invention is to address the shortcomings of the above-mentioned prior art in matching accuracy and matching speed, and to propose a heterogeneous image matching method based on cross-modal attention and optimal transmission theory, so as to improve the matching speed and improve the matching accuracy.
[0010] The technical idea of the present invention is to improve the matching speed by building an end-to-end cross-modal Transformer matching network, so that the visible light and SAR modalities have better interaction, and obtain the similarity measurement of SAR images and visible light images; optimize the matching results through optimal transmission to improve the matching accuracy.
[0011] According to the above ideas, the implementation scheme of the heterogeneous image matching method based on the cross-modal conversion network and the optimal transmission theory of the present invention includes the following:
[0012] 1. A heterogeneous image matching method based on a cross-modal conversion network and optimal transmission theory, characterized by comprising:
[0013] (1) Constructing training data and test data for heterogeneous image matching:
[0014] (1a) Image pairs of size 512×512 are selected from the open source dataset OS Dataset as the selected dataset, which contains paired SAR and visible light images that have been registered;
[0015] (1b) The visible light image in each pair of images in the selected dataset is used as the search image, and pixels are randomly selected in each visible light SAR image as the upper left corner coordinates. A 256×256 image is cropped as the template image, and the upper left corner coordinates are saved as the true label of the image pair.
[0016] (1c) 80% of the image pairs of the cropped SAR images and the corresponding visible light images are used as training sets, and 20% of the image pairs are used as test sets;
[0017] (2) Construct a cross-modal Transformer matching network N1:
[0018] (2a) Setting up the Segformer feature extraction skeleton including the correlation graph constraints;
[0019] (2b) Establish a Transformer network N0 that includes cross-modal attention;
[0020] (2c) The Segformer feature extraction skeleton containing the correlation graph constraint and the Transformer network containing the cross-modal cross attention are cascaded in sequence to form a cross-modal Transformer matching network N1;
[0021] (3) Using the training data and optimal transmission theory, the Adam algorithm is used to iteratively train the matching network N1 to obtain the trained matching network N2;
[0022] (4) Use the optimal transfer theory and the trained matching network N2 to match the image pairs in the test set.
[0023] Compared with the prior art, the present invention has the following advantages:
[0024] 1. Higher accuracy and smaller matching error
[0025] The present invention constructs a Transformer-based matching network model, adds correlation graph constraints and cross-modal attention to the Segformer network structure to constrain the importance of features in the search graph, and performs matching optimization based on optimal transmission to improve the matching accuracy.
[0026] 2. Faster matching speed
[0027] The present invention uses cosine similarity to measure feature similarity, without the need for very time-consuming pixel-by-pixel cross-correlation operations, and the entire network performs end-to-end reasoning. The matching time is shorter than that of existing deep learning methods, thereby improving the matching speed.
[0028] 3. More adaptable to different terrain scenes
[0029] The present invention uses Segformer with strong feature extraction capability based on attention mechanism as the feature extraction skeleton. Facing complex and changeable terrain scenes, the network can extract more effective feature representation and obtain accurate matching results, thereby improving the generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flow chart of the implementation of the present invention;
[0031] Figure 2 It is a Segformer feature extraction skeleton structure diagram constructed by adding correlation graph constraints in the present invention;
[0032] Figure 3 It is a diagram of the cross-modal attention Tranformer network structure constructed in the present invention;
[0033] Figure 4 The comparison diagram is a result of matching a SAR image and a visible light image on an urban area image in the open source dataset OS Dataset using the present invention and eight existing algorithms;
[0034] Figure 5 This is a comparison chart of the results of matching a SAR image and a visible light image on an airport area image in the open source dataset OS Dataset using the present invention and eight existing algorithms. DETAILED DESCRIPTION
[0035] The embodiments and effects of the present invention are further described in detail below with reference to the accompanying drawings.
[0036] Reference Figure 1 , the implementation steps of the present invention are as follows:
[0037] Step 1. Construct training data and test data for heterogeneous image matching.
[0038] (1.1) Select image pairs of size 512×512 from the open source dataset OS Dataset as the selected dataset, which contains paired SAR and visible light images that have been registered;
[0039] (1.2) The visible light image in each pair of images in the selected dataset is used as the search image, and pixels are randomly selected in the SAR image corresponding to each visible light image as the upper left corner coordinates. A 256×256 image is cropped as the template image, and the upper left corner coordinates are saved as the true label of the image pair.
[0040] (1.3) 80% of the image pairs in the paired cropped SAR images and the corresponding visible light images are used as training sets, and 20% of the image pairs are used as test sets.
[0041] Step 2. Build a cross-modal Transformer matching network N1.
[0042] (2.1) Construct the Segformer feature extraction skeleton including the relevant graph constraints:
[0043] The specific implementation of this step is to improve the existing Segformer network, which contains 4 Transformer Blocks and two multi-layer perceptrons MLP; each Transformer Block contains a cascade structure of N efficient self-attention modules and a mixed feedforward neural network Mix-FFN, as well as the last Overlap Patch Merging module. The efficient self-attention module is a self-attention module with sequence compression, which calculates the attention score of the input feature; the mixed feedforward neural network Mix-FFN is a convolutional feedforward neural network with a convolution kernel size of 3 and zero padding; the overlapping block merging module is a convolution layer with a convolution kernel of 7, zero padding of 4, and a stride of 2. After receiving the input, the network passes through each Transformer Block, merges the features of each image block to obtain the output of the Transformer Block; the output features of multiple Transformer Blocks of different resolutions are fused through the first multi-layer perceptron to obtain the fused features, and then the fused features are input into the second multi-layer perceptron to obtain the final output features of the Segformer.
[0044] Reference Figure 2 , This step improves the existing Segformer network by adding relevant graph constraints to it, which is specifically implemented as follows:
[0045] (2.1.1) Calculate the cross-correlation matrices Cor1 and Cor3 of the SAR image features and visible light image features output by the first and third Transformer Blocks in the existing Segformer network respectively;
[0046] (2.1.2) Construct two sizes of and The zero matrix of and As the initial first correlation graph and second correlation graph;
[0047] (2.1.3) Iterate the first correlation graph and the second correlation graph respectively to obtain their respective final correlation graphs:
[0048] Set The total number of iterations is the number of elements in Cor1. In each iteration, a point (x, y) is selected from Cor1 without repetition to obtain the modification range of this iteration. Will Modify the value in the modification range to The maximum value between Cor1(x,y) and the final first correlation graph is obtained at the end of the iteration Where, Cor1(x,y) is the value of Cor1 at the point (x,y);
[0049] Set The total number of iterations is the number of elements in Cor3. In each iteration, a point (x, y) is selected from Cor3 without repetition to obtain the modification range of this iteration. Will Modify the value in the modification range to The maximum value between Cor3(x,y) and the final second correlation graph is obtained after the iteration. Where, Cor3(x,y) is the value of Cor3 at the point (x,y);
[0050] (2.1.4) The final first correlation graph The output of the first Transformer Block is multiplied by the visible light image feature as the input of the second Transformer Block. Multiply the visible light image features output by the third Transformer Block as the input of the fourth Transformer Block to complete the addition of the correlation graph constraint and obtain the Segformer feature extraction skeleton containing the correlation graph constraint;
[0051] (2.2) Construct a Transformer network N0 containing cross-modal attention:
[0052] Reference Figure 3 In this step, a Transformer network with cross-modal attention is established by improving the existing Segformer network. The specific implementation is as follows:
[0053] (2.2.1) Remove the 3rd and 4th Transformer Blocks in the existing Segformer network;
[0054] (2.2.2) Exchange the visible light image feature query in the first Transformer Block and SAR image feature query
[0055] (2.2.3) Exchange the visible light image feature query in the second Transformer Block and SAR image feature query Get the Transformer network N0 containing cross-modal cross attention;
[0056] (2.3) The Segformer feature extraction skeleton containing the correlation graph constraint and the Transformer network N0 with cross-modal cross-attention are cascaded in sequence to obtain the cross-modal Transformer matching network N1.
[0057] Step 3. Using the training data and optimal transmission theory, use the Adam algorithm to iteratively train the network N1 to obtain the trained matching network N2.
[0058] (3.1) Select a pair of SAR images and visible light images in the training set, and input the SAR image and visible light image into the cross-modal Transformer matching network N1 constructed in step 2, and obtain the SAR image feature map f s And the visible light image feature map f o ;
[0059] (3.2) For the SAR image feature map f s And the visible light image feature map f o Calculate the similarity matrix M:
[0060]
[0061] Where T represents the transpose of the matrix, || || represents modulus;
[0062] (3.3) Based on the similarity matrix M of the training set SAR image features and visible light image features, the optimal matching probability C is calculated using the optimal transmission * :
[0063] (3.3.1) Set a matrix C as the matching probability from SAR image to visible light image;
[0064] (3.3.2) In order to avoid trivial solutions, the class activation map CAM of the SAR image features and visible light image features output by the second Transformer Block of the cross-modal cross-attention Transformer matching network is used as the constraint conditions μ for optimal transmission. sar and μ opt ;
[0065] (3.3.3) The optimal transmission problem is solved by the Sinkhorn-Knopp algorithm to obtain the optimal matching probability C between the training set SAR image and the visible light image: * :
[0066]
[0067] Among them, C ij is the value of matrix C at (i, j), M ij represents the value of matrix M at (i,j), h s ,w s Respectively represent the height and width of the SAR image feature, h o ,w o Respectively represent the height and width of the visible light image feature; Indicates size h s w s The unit column vector of ; Indicates size h o w o The unit column vector of ; represents the sum of each row of matrix C, represents the sum of each column of the matrix C, and T represents the transpose of the matrix;
[0068] (3.4) Substitute the optimal matching probability C obtained from (3.3.3) * Multiplying the similarity matrix M obtained by (3.2) gives the optimized training set similarity measure matrix M opt :
[0069] M opt =C * ⊙M
[0070] Among them, ⊙ represents the multiplication of elements at corresponding positions in the matrix;
[0071] (3.5) M opt The coordinates of the maximum value point are used as matching points And calculate the loss function Loss between the matching point and the true label:
[0072]
[0073] Among them, (x t ,y t ) is the true label coordinate;
[0074] (3.6) Repeat (3.1) to (3.5), and update the parameters of each layer of the network according to the loss function value of each iteration until the set number of iterations E = 300 is reached, and the trained cross-modal Transformer matching network N2 is obtained.
[0075] Step 4. Use the optimal transfer theory and the trained matching network N2 to match the image pairs in the test set.
[0076] (4.1) Input the SAR image and visible light image in the test set into the trained matching network N2 to obtain the SAR image features f of the test image pair s ′ and visible light image feature f o ′;
[0077] (4.2) Calculate the similarity matrix M′ of the test image to the output features:
[0078]
[0079] Where T represents the transpose of the matrix, and || || represents modulo;
[0080] (4.3) Based on the similarity matrix M′ of the output features of the test image pair, the optimal matching probability C of the test image pair is calculated using the optimal transmission * ′:
[0081] (4.3.1) Set a matrix C′ as the matching probability between the test set SAR image and the visible light image;
[0082] (4.3.2) The class activation map CAM of the SAR image features and visible light image features output by the second Transformer Block of the cross-modal cross-attention Transformer matching network is used as the constraint condition μ′ for the optimal transmission of the test image sar and μ′ opt ;
[0083] (4.3.3) Solve the following problem using the Sinkhorn-Knopp algorithm to obtain the optimal matching probability C of the test image: * ′:
[0084]
[0085] Among them, C′ ijis the value of matrix C′ at (i,j), M′ ij represents the value of the matrix M′ at (i, j), h′ s ,w′ s Respectively represent the height and width of the SAR image features of the test set, h′ o ,w′ o Respectively represent the height and width of the visible light image features of the test set; Indicates size h′ s w′ s The unit column vector of ; Indicates size h′ o w′ o The unit column vector of ; represents the sum of each row of the matrix C′, represents the sum of each column of the matrix C′, and T represents the transpose of the matrix;
[0086] (4.4) The optimal matching probability C of the test image * ′ is multiplied by its similarity matrix M′ to obtain the optimized similarity measurement matrix M′ opt :
[0087] M′ opt =C * ′⊙M′
[0088] Among them, ⊙ represents the multiplication of elements at corresponding positions in the matrix;
[0089] (4.5) M′ opt The coordinates of the maximum value point are used as matching points This point is the corresponding matching position of the SAR image in the test set in the visible light image, completing the matching of heterogeneous images.
[0090] The effect of the present invention can be further illustrated by the following experiments:
[0091] 1. Experimental conditions
[0092] The server used in this experiment is configured with a 3.2GHz Intel Core i7-9700K CPU and a 12-GB NVIDIA GeForce RTX2080Ti GPU. The deep network model is implemented using the PyTorch 1.5.1 code framework, and the programming language is Python 3.7.
[0093] The dataset used in the experiment is the open source dataset OS Dataset, which includes 1,300 pairs of heterogeneous images and their labels. The size of the SAR image is 256×256, and the SAR image is collected from China's multi-polarization C-band SAR satellite Gaofen-3 with a resolution of 1 meter. The size of the visible light image is 512×512, and the image is collected from the Google Earth platform and resampled to a resolution of 1 meter.
[0094] In this example, 80% of the images are used as training sets and 20% of the images are used as test sets. The experiment is conducted on the test set with subjects’ error less than or equal to 5 pixels, the average error of correctly matched images, the average error of all images, and the matching time;
[0095] There are eight comparison methods used in the experiment, namely, the normalized cross-correlation algorithm NCC, the normalized mutual information algorithm NMI, the channel feature algorithm of oriented gradient CFOG, the phase consistency histogram HOPC, the radiation insensitive feature transform algorithm RIFT, the pseudo-twin convolutional neural network algorithm PSiam, the deep matching network based on visual saliency features VSMatch, and the step-by-step cascade matching network SCMNet.
[0096] 2. Experimental content
[0097] Experiment 1: Under the above experimental conditions, the present invention and the existing eight algorithms, NCC, NMI, HOPC, CFOG, RIFT, PSiam, VSMatch, and SCMNet, are used to match a pair of SAR images and visible light images of the urban area in the above test set. The results are as follows: Figure 4 As shown, where:
[0098] Figure 4 (a) is the SAR image template,
[0099] Figure 4 (b) is the true label,
[0100] Figure 4 (c) is the matching result of the NCC algorithm.
[0101] Figure 4 (d) is the matching result of the NMI algorithm.
[0102] Figure 4 (e) is the matching result of HOPC algorithm.
[0103] Figure 4 (f) is the matching result of CFOG algorithm.
[0104] Figure 4 (g) is the visible light image,
[0105] Figure 4(h) is the matching result of RIFT algorithm.
[0106] Figure 4 (i) is the matching result of PSiam algorithm,
[0107] Figure 4 (j) is the matching result of VSMatch algorithm.
[0108] Figure 4 (k) is the matching result of SCMNet algorithm.
[0109] Figure 4 (l) is the matching result of the method of the present invention.
[0110] The solid square box in each picture is the actual matching position, and the dotted square box is the predicted matching position obtained by each method. The closer the position of the dotted prediction box is to the actual matching position of the solid square box, the better the matching effect of the algorithm.
[0111] from Figure 4 The results show that the predicted position of the comparison method is offset from the actual position, while the corresponding Figure 4 (l) In urban areas where local feature differences are small, the predicted position and the actual position completely overlap, indicating that the present invention can achieve accurate matching in similar ground feature scenes.
[0112] Experiment 2: Under the above experimental conditions, the SAR image and visible light image of a pair of airport areas in the above test set are matched using the present invention and the existing eight algorithms, namely, NCC, NMI, HOPC, CFOG, RIFT, PSiam, VSMatch, and SCMNet. The results are as follows: Figure 5 As shown, where:
[0113] Figure 5 (a) is the SAR image template,
[0114] Figure 5 (b) is the true label,
[0115] Figure 5 (c) is the matching result of the NCC algorithm.
[0116] Figure 5 (d) is the matching result of the NMI algorithm.
[0117] Figure 5 (e) is the matching result of HOPC algorithm.
[0118] Figure 5 (f) is the matching result of CFOG algorithm.
[0119] Figure 5 (g) is the visible light image,
[0120] Figure 5 (h) is the matching result of RIFT algorithm.
[0121] Figure 5 (i) is the matching result of PSiam algorithm,
[0122] Figure 5 (j) is the matching result of VSMatch algorithm.
[0123] Figure 5 (k) is the matching result of SCMNet algorithm.
[0124] Figure 5 (l) is the matching result of the algorithm proposed in the present invention.
[0125] The solid square box in each picture is the actual matching position, and the dotted square box is the predicted matching position obtained by each method. The closer the position of the dotted prediction box is to the actual matching position of the solid square box, the better the matching effect of the algorithm.
[0126] from Figure 5 It can be seen from the results that the presence of the aircraft in the experimental image makes the local features of the scene differ greatly, and due to the imaging method of the SAR image, the aircraft produces more coherent speckle noise in the SAR image, making accurate matching more difficult. The matching results of all comparison methods have large errors. The predicted position of the present invention in this area is consistent with the actual position, achieving accurate matching.
[0127] In experiment 3, the SAR images and visible light images in the test set are matched, and the evaluation index is calculated based on all matching results and labels. The results are shown in Table 1:
[0128] Table 1 Evaluation indexes of the present invention and 8 existing methods
[0129]
[0130] It can be seen from the results in Table 1 that the accuracy of the present invention reached 81.67% in the experiment, which significantly improved the accuracy of heterogeneous image matching; compared with similar deep learning matching methods participating in the comparison, the time required for the present invention to complete the matching is significantly reduced, which greatly improves the matching speed, and in the experiment, the average error of the correctly matched image and the average error of all images of the present invention are the lowest, which improves the matching accuracy.
[0131] In summary, the heterogeneous image matching method based on cross-modal conversion network and optimal transmission theory constructed in the present invention can obtain better matching results compared with the existing NCC, NMI, CFOG, HOPC, RIFT, PSiam, VSMatch, and SCMNet algorithms. The results have higher matching accuracy and smaller average error, and the matching time is in a leading position among similar deep learning-based algorithms. It has good adaptability to different types of ground object scenes and has stronger generalization ability.
Claims
1. A heterogeneous image matching method based on a cross-modal conversion network and optimal transmission theory, characterized in that: include: (1) Constructing training data and test data for heterogeneous image matching: (1a) Image pairs of size 512×512 are selected from the open source dataset OSDataset as the selected dataset, which contains paired SAR and visible light images that have been registered; (1b) The visible light image in each pair of images in the selected dataset is used as the search image, and pixels are randomly selected in each visible light SAR image as the upper left corner coordinates. A 256×256 image is cropped as the template image, and the upper left corner coordinates are saved as the true label of the image pair. (1c) 80% of the image pairs of the cropped SAR images and the corresponding visible light images are used as training sets, and 20% of the image pairs are used as test sets; (2) Construct a cross-modal Transformer matching network N1: (2a) Setting up the Segformer feature extraction skeleton including the correlation graph constraints; (2b) Establish a Transformer network N0 that includes cross-modal attention; (2c) The Segformer feature extraction skeleton containing the correlation graph constraint and the Transformer network containing the cross-modal cross attention are cascaded to form a cross-modal Transformer matching network N1; (3) Using the training data and optimal transmission theory, the Adam algorithm is used to iteratively train the matching network N1 to obtain the trained matching network N2; (4) Use the optimal transmission and the trained matching network N2 to match the image pairs in the test set: (4a) Input the SAR image and visible light image in the test set into the trained matching network N2 to obtain the SAR image features f of the test image pair s ′ and visible light image feature f o ′; (4b) Calculate the similarity matrix M′ of the test image to the output features: Among them, T represents the transpose of the matrix, and |||| represents modulo; (4c) Based on the similarity matrix M′ of the output features of the test image pair, the optimal matching probability C of the test image pair is calculated by using the optimal transmission optimization * ′; (4d) The optimal matching probability C of the test image * ′ is multiplied by its similarity matrix M′ to obtain the optimized similarity measurement matrix M′ opt : M′ opt =C * ′⊙M′ Among them, ⊙ represents the multiplication of elements at corresponding positions in the matrix; (4e) M′ opt The coordinates of the maximum value point are taken as the matching point (x test ,y test ), which is the corresponding matching position of the SAR template image in the test set in the visible light image, completing the matching of heterogeneous images.
2. The method according to claim 1, characterized in that: The Segformer feature extraction skeleton including the correlation graph constraint is set in (2a) and implemented as follows: (2a1) In the existing Segformer network, let the output feature map size of the SAR image output by the first Transformer Block be The output feature map size of the visible light image is Construct a size The zero matrix of as the first correlogram to be corrected; (2a2) Calculate the cross-correlation matrix Cor1 of the SAR image and visible light features output by the first Transformer Block in the Segformer network. Correction is performed to obtain the corrected first correlation diagram and will Multiply it with the features output by the first Transformer Block as the input of the second Transformer Block; (2a3) In the existing Segformer network, let the output feature map size of the SAR image output by the third Transformer Block be The output feature map size of the visible light image is Construct a size The zero matrix of as the second correlogram to be corrected; (2a4) Calculate the cross-correlation matrix Cor3 of the SAR image and visible light features output by the third Transformer Block in the Segformer network. Correction is performed to obtain the corrected second correlation diagram Will Multiply it with the features output by the third Transformer Block as the input of the fourth Transformer Block.
3. The method according to claim 2, characterized in that In (2a2), the first correlation graph is calculated based on Cor1. To make corrections, we take each point (x, y) in Cor1 as the coordinate of the upper left corner of the correction range and compare the first correlation graph. Make corrections, namely: First, set the correction range corresponding to the upper left corner coordinates (x, y) for each correction: Then, the first correlation diagram is Change the value in to The corrected correlation diagram is obtained in: express The value at the midpoint (i, j), Cor1(x, y) is the value at the midpoint (x, y) of Cor1; Indicates taking The maximum value between and Cor1(x,y).
4. The method according to claim 2, characterized in that: In (2a4), the second correlation diagram is calculated based on Cor3. Make corrections and implement them as follows: First, set the correction range corresponding to each point (x, y) in Cor3 at each correction Then, the second correlation graph is converted according to the set correction range. Change the value in to The corrected correlation diagram is obtained in: express The value at the midpoint (i, j), Cor3(x, y) is the value at the midpoint (x, y) of Cor3; Indicates taking The maximum value between and Cor3(x,y).
5. The method according to claim 1, characterized in that The Transformer network N1 including cross-modal attention in (2b) is established by improving the existing Segformer network, and is specifically implemented as follows: First, remove the 3rd and 4th Transformer Blocks in the existing Segformer network; Then, swap the visible light image feature query in the first Transformer Block and SAR image feature query Finally, swap the visible light image feature query in the second Transformer Block and SAR image feature query The Transformer network N1 containing cross-modal cross attention is obtained.
6. The method according to claim 1, characterized in that In (3), the training data and optimal transmission are used to iteratively train the matching network N1 using the Adam algorithm, which is implemented as follows: (3a) Select a pair of SAR images and visible light images from the training set and input them into the cross-modal Transformer matching network N1 to obtain f o ; (3b) Calculate the SAR image features f of the training set s And the similarity matrix of visible light image features: Among them, T represents the transpose of the matrix, and |||| represents modulo; (3c) Based on the similarity matrix M of the training set SAR image features and visible light image features, the optimal matching probability is calculated using optimal transmission: (3c1) Set a matrix C as the matching probability from SAR image to visible light image; (3c2) The class activation map CAM of the SAR image features and visible light image features output by the second Transformer Block of the cross-modal cross-attention Transformer matching network is used as the constraint condition μ for optimal transmission. sar and μ opt ; (3c3) The following problem is solved by the Sinkhorn-Knopp algorithm to obtain the optimal matching probability C of the training set SAR image and visible light image: * : Among them, C ij is the value of matrix C at (i, j), M ij represents the value of matrix M at (i,j), h s ,w s Respectively represent the height and width of the SAR image feature, h o ,w o Respectively represent the height and width of the visible light image feature; Indicates size h s w s The unit column vector of ; Indicates size h o w o The unit column vector of ; represents the sum of each row of matrix C, represents the sum of each column of the matrix C, and T represents the transpose of the matrix; (3d) Substitute the optimal matching probability C obtained from (3c3) * Multiplying the similarity matrix M obtained by (3b) gives the optimized training set similarity measure matrix M opt : M opt =C * ⊙M Among them, ⊙ represents the multiplication of elements at corresponding positions in the matrix; (3e) M opt The coordinates of the maximum value point are used as matching points And calculate the loss function Loss between the matching point and the true label: Among them, (x t ,y t ) is the true label coordinate; (3f) Repeat (3a) to (3e), and update the parameters of each layer of the network according to the loss function value of each iteration until the set number of iterations E = 300 is reached, and the trained cross-modal Transformer matching network N2 is obtained.
7. The method according to claim 1, characterized in that In (4c), the optimal matching probability C is calculated based on the similarity matrix M′ using the optimal transmission * ′, and is implemented as follows: First, a matrix C′ is set as the matching probability between the SAR image and the visible light image; Then, the class activation map CAM of the SAR image features and visible light image features output by the second Transformer Block of the cross-modal cross-attention Transformer matching network is used as the constraint condition μ′ for optimal transmission. sar and μ′ opt ; Finally, the optimal matching probability C can be obtained by solving the following problem through the Sinkhorn-Knopp algorithm: * ′: Among them, C ij ′ is the value of matrix C′ at (i, j), M ij ′ represents the value of the matrix M′ at (i, j), h′ s ,w′ s Respectively represent the height and width of the SAR image features of the test set, h′ o ,w′ o Respectively represent the height and width of the visible light image features of the test set; Indicates size h′ s w′ s The unit column vector of ; Indicates size h′ o w′ o The unit column vector of ; represents the sum of each row of the matrix C′, represents the sum of each column of the matrix C′, and T represents the transpose of the matrix.
Citation Information
Patent Citations
Step-by-step different-source image template matching method based on cascade network
CN114140700A
SAR-visible light remote sensing image matching method
CN114358150A