Heterogeneous Image Matching Method and System Based on Contrastive Learning
Through the optimization of pseudo-twin neural network structure and loss function based on comparison learning, the performance degradation of heterologous image matching model in unknown test domains is solved, and the matching accuracy and robustness of the model under large distribution differences are improved.
Patent Information
- Application Number
- CN202410498186.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-04-24
AI Technical Summary
Existing heterologous image matching models will significantly reduce performance when the distribution of training data and test data is large, especially in unknown test domains.
Using a method based on contrast learning, the features of SAR images and optical images are extracted separately through the pseudo-twin neural network structure, the feature similarity is calculated using the cross-correlation operator, and the matching point position is estimated through the classification head and the regression head. Combined with the InfoNCE, Focal and GIoU loss function optimization model, the generalization performance of the model in unknown test domains is improved.
It improves the generalization performance of heterologous image matching model on unknown test domains, improves matching accuracy and robustness, and solves the problem of performance degradation of the model under large distribution differences.
Smart Images

Figure CN118334362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and specifically to a method and system for heterogeneous image matching based on contrast learning. Background Art
[0002] Due to the strong complementarity of remote sensing information between synthetic aperture radar and optical images, by matching SAR images and optical images, complementary of a large amount of surface feature information can be achieved, and tasks such as change detection, target recognition, and ground object classification can be better solved. Existing heterogeneous image matching methods can be divided into two categories according to traditional methods and deep learning-based methods. Traditional methods for heterogeneous image matching can be divided into template matching and feature-based. Template matching is a pixel-based matching method that obtains the matching result by comparing whether the template image and the corresponding large image area are similar at each position. Feature-based matching methods extract feature points from the image and generate descriptors of the scale and direction information of the feature points for image feature point matching.
[0003] Due to the huge differences between SAR images and optical images, most traditional methods for heterogeneous image matching need to combine the intensity and gradient information of the images, and these information are very sensitive to the non-linear radiation distortion of images from different sensors. Traditional intensity-based matching methods are difficult to process image pairs with non-linear radiation distortion, while traditional feature-based methods usually need to extract key points of the image and design complex artificial descriptors based on the key points, with large computational complexity and difficult to process in real time.
[0004] Deep learning-based heterogeneous image matching methods usually use pseudo-siamese neural networks to extract image features and predict image matching points. However, existing deep learning-based heterogeneous image matching models do not consider the domain generalization problem of the model. The goal of domain generalization research is to learn a model with strong generalization ability from multiple datasets with different data distributions, which can achieve better results on unknown test domains. In traditional machine learning models, it is usually assumed that the test data and the training data are independently and identically distributed, but this assumption does not always hold in actual situations. When there are obvious differences in the distributions of the test data and the training data, the model may be affected by the existence of the domain distribution gap, resulting in a decline in model performance. Deep learning-based heterogeneous image matching methods usually assume that the training data and the test data are independently and identically distributed. When the test data is unknown and has a large distribution difference from the training data, the matching performance of the model decreases. For example, in practical applications, the training data are image pairs of some landforms such as plains, lakes, and ports, but the test data are desert landforms with large differences. At this time, the performance of the heterogeneous image matching model will be greatly reduced. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the method and system for heterologous image matching based on contrast learning provided by the present invention solve the problem that when the distribution difference between training data and detection data is large, the performance of the model will be significantly reduced.
[0006] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is as follows:
[0007] Provide a method for heterologous image matching based on contrast learning, which includes the following steps:
[0008] S1. Obtain the SAR image and the optical image to be matched;
[0009] S2. Respectively obtain the image features of the SAR image and the image features of the optical image through the pseudo-twin neural network structure;
[0010] S3. Calculate the similarity between the image features of the SAR image and the image features of the optical image through the cross-correlation operator;
[0011] S4. Based on the similarity, estimate the position of the matching point through the classification head to obtain the rough matching point coordinates;
[0012] S5. Based on the similarity, estimate the position of the matching point through the regression head to obtain the offset between the predicted matching point and the true matching point;
[0013] S6. Based on the rough matching point coordinates and the offset, calculate the accurate matching point coordinates to obtain the matching result.
[0014] The beneficial effect of the present invention is that this method can solve the problem of the performance degradation of the heterologous image matching model on the unknown test domain and improve the generalization performance of the heterologous matching model on the unknown test data.
[0015] Further, in step S1, the original SAR image and the optical image are respectively standardized and normalized in sequence, so that the pixel ranges of the original SAR image and the optical image are constrained between 0 and 1 to obtain the SAR image and the optical image to be matched.
[0016] The beneficial effect of adopting the above further solution is that due to the large appearance difference between the optical image and the SAR image, the method of standardization and normalization is used to constrain the pixel range of the image between 0 and 1, which is convenient for subsequent processing.
[0017] Furthermore, the pseudo-twin neural network structure includes an SAR image branch and an optical image branch with the same structure but non-shared weights. Each branch includes a first convolutional layer and a second convolutional layer connected in sequence. The output ends of the second convolutional layer are respectively connected to the input ends of a third convolutional layer and a fourth convolutional layer. The output end of the fourth convolutional layer is connected to the input end of a fifth convolutional layer. The output end of the fifth convolutional layer is connected to the input end of a sixth convolutional layer. The output of the sixth convolutional layer is connected to the output of the fourth convolutional layer with a residual connection. The output of the third convolutional layer and the residual result are concatenated and then input into a seventh convolutional layer. The output end of the seventh convolutional layer is connected to the input end of an eighth convolutional layer. The output end of the eighth convolutional layer is the output end of the current branch, and the input end of the first convolutional layer is the input end of the current branch.
[0018] Furthermore, both the first convolutional layer and the second convolutional layer perform four-fold downsampling using a convolution with a stride of 2.
[0019] The beneficial effects of adopting the above further scheme are as follows: Only the first convolutional layer and the second convolutional layer perform four-fold downsampling using a convolution with a stride of 2, which can reduce the downsampling rate and retain the detailed features of the image. Moreover, a residual connection is adopted later, and the residual structure can effectively alleviate the training difficulties of deep networks, improving the feature extraction accuracy and efficiency of the pseudo-twin neural network structure.
[0020] Furthermore, the specific method of step S3 is as follows:
[0021] Through depthwise convolution, taking the image features of the SAR image as the convolution kernel, perform convolution operations with the image features of the optical image in each channel to calculate the feature similarity at each position in the search space.
[0022] Furthermore, both the classification head and the regression head include an input channel and an output channel connected in sequence; the input channel structures of the classification head and the regression head are the same and share parameters. The input channel includes four convolutional layers connected in sequence, where the output of the first convolutional layer and the output of the fourth convolutional layer are connected with a residual connection and then used as the input of the fifth convolutional layer. The output of the fifth convolutional layer is the output of the input channel;
[0023] In the classification head, the output channel is a convolutional layer with a channel number of 1;
[0024] In the regression head, the output channel is a convolutional layer with a channel number of 2.
[0025] Furthermore, the calculation expression of step S6 is:
[0026]
[0027]
[0028] Where To accurately match the point coordinates; is the coordinate of the rough matching point; is the offset between the predicted matching point and the actual matching point; S is the downsampling multiple.
[0029] Furthermore, the value of S is 4.
[0030] Furthermore, a training process is also included, and the training process includes the following steps:
[0031] A1. Randomly sample a picture from the image features of the SAR image as a template, and sample a picture matching the template from the image features of the optical image, and form a positive sample pair with the template;
[0032] A2. Sample N images that do not match the template from the image features of the optical image, and form N negative sample pairs with the template;
[0033] A3. With the purpose of shortening the distance between positive samples and increasing the distance between negative samples in the feature space, the loss value of the pseudo-twin neural network structure is obtained through the InfoNCE loss function, which is recorded as contrastive learning loss.
[0034] A4. Obtain the loss of the classification head through the Focal loss function; obtain the loss of the regression head through the GIoU loss function;
[0035] A5. The sum of the contrastive learning loss, the loss of the classification head, and the loss of the regression head is taken as the total loss. With the goal of minimizing the total loss, the parameters of the pseudo-twin neural network structure, the classification head, and the regression head are adjusted through back propagation until the training is completed.
[0036] The beneficial effects of adopting the above further scheme are: through contrast learning, the distance between matching points in the feature space is shortened and the unmatched points are pushed away, which can improve the generalization performance of the heterogeneous image matching model on unknown detection data, and further avoid the problem of performance degradation of heterogeneous images in unknown detection domains. The classification head uses the Focal loss function to solve the problem of imbalance between positive and negative samples, and can also be used to mine difficult negative samples.
[0037] A heterogeneous image matching system based on contrastive learning is provided, comprising:
[0038] A data acquisition module, used for acquiring SAR images and optical images to be matched;
[0039] An image feature extraction module, used to obtain the image features of the SAR image and the image features of the optical image respectively through a pseudo-twin neural network structure;
[0040] A feature similarity calculation module, used for calculating the similarity between the image features of the SAR image and the image features of the optical image through a cross-correlation operator;
[0041] A rough matching point coordinate calculation module, configured to estimate the position of the matching point through a classification head based on similarity, and obtain the rough matching point coordinates;
[0042] An offset calculation module, configured to estimate the position of the matching point through a regression head based on similarity, and obtain the offset between the predicted matching point and the true matching point;
[0043] An image matching module, configured to calculate the accurate matching point coordinates based on the rough matching point coordinates and the offset, and obtain the matching result.
[0044] The beneficial effects of the present invention are as follows: This method can solve the problem of the performance degradation of the heterologous image matching model on the unknown test domain, and improve the generalization performance of the heterologous matching model on the unknown test data. Description of the Drawings
[0045] Figure 1 It is a schematic flow diagram of this method;
[0046] Figure 2 It is a schematic structural diagram of a branch in the pseudo-twin neural network structure;
[0047] Figure 3 It is a schematic structural diagram of the input channels of the classification head and the regression head;
[0048] Figure 4 It is a schematic structural diagram of the training process;
[0049] Figure 5 It is the SAR image to be matched in the embodiment;
[0050] Figure 6 It is the optical image to be matched in the embodiment;
[0051] Figure 7 It is a schematic diagram of the matching result given by this method / system in the embodiment. Detailed Embodiments
[0052] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0053] As Figure 1 shown, the heterologous image matching method based on contrast learning includes the following steps:
[0054] S1. Obtain the SAR image and the optical image to be matched;
[0055] S2. Obtain the image features of the SAR image and the image features of the optical image respectively through the pseudo-twin neural network structure;
[0056] S3. Calculate the similarity between the image features of the SAR image and the image features of the optical image through the cross-correlation operator;
[0057] S4. Based on the similarity, estimate the position of the matching points through the classification head to obtain the rough matching point coordinates;
[0058] S5. Based on the similarity, estimate the position of the matching points through the regression head to obtain the offset between the predicted matching points and the true matching points;
[0059] S6. Based on the rough matching point coordinates and the offset, calculate the accurate matching point coordinates to obtain the matching result.
[0060] In step S1, the original SAR image and the optical image are obtained and sequentially normalized and standardized respectively, so that the pixel ranges of the original SAR image and the optical image are constrained between 0 and 1 to obtain the SAR image and the optical image to be matched.
[0061] The pseudo-twin neural network structure includes a SAR image branch and an optical image branch with the same structure but non-shared weights. As Figure 2 shown, each branch includes a first convolutional layer and a second convolutional layer connected in sequence. The output end of the second convolutional layer is respectively connected to the input end of the third convolutional layer and the input end of the fourth convolutional layer. The output end of the fourth convolutional layer is connected to the input end of the fifth convolutional layer. The output end of the fifth convolutional layer is connected to the input end of the sixth convolutional layer. The output of the sixth convolutional layer is connected to the output of the fourth convolutional layer in a residual connection. After the output of the third convolutional layer and the residual result are concatenated, they are input into the seventh convolutional layer. The output end of the seventh convolutional layer is connected to the input end of the eighth convolutional layer. The output end of the eighth convolutional layer is the output end of the current branch, and the input end of the first convolutional layer is the input end of the current branch.
[0062] Both the first convolutional layer and the second convolutional layer use convolution with a stride of 2 for four-fold downsampling.
[0063] The specific method of step S3 is:
[0064] Through depthwise convolution, use the image features of the SAR image as the convolution kernel to perform convolution operations with the image features of the optical image in each channel to calculate the feature similarity at each position in the search space.
[0065] As Figure 3As shown, both the classification head and the regression head include an input channel and an output channel connected in sequence; the input channel structures of the classification head and the regression head are the same and share parameters. The input channel includes four convolutional layers connected in sequence. Among them, the output of the first convolutional layer is connected to the output of the fourth convolutional layer by residual connection and then used as the input of the fifth convolutional layer. The output of the fifth convolutional layer is the output of the input channel.
[0066] In the classification head, the output channel is a convolutional layer with 1 channel number.
[0067] In the regression head, the output channel is a convolutional layer with 2 channel numbers.
[0068] The calculation expression of step S6 is:
[0069]
[0070]
[0071] where is the coordinate of the exact matching point; is the coordinate of the rough matching point; is the offset between the predicted matching point and the true matching point; S is the downsampling factor. The value of S is 4.
[0072] The heterologous image matching system based on contrastive learning includes:
[0073] A data acquisition module for acquiring the SAR image and the optical image to be matched;
[0074] An image feature extraction module for respectively obtaining the image features of the SAR image and the optical image through a pseudo-twin neural network structure;
[0075] A feature similarity calculation module for calculating the similarity between the image features of the SAR image and the optical image through a cross-correlation operator;
[0076] A rough matching point coordinate calculation module for estimating the matching point position through the classification head based on the similarity to obtain the rough matching point coordinate;
[0077] An offset calculation module for estimating the matching point position through the regression head based on the similarity to obtain the offset between the predicted matching point and the true matching point;
[0078] An image matching module for calculating the exact matching point coordinate based on the rough matching point coordinate and the offset to obtain the matching result.
[0079] As Figure 4 shown, in the specific implementation process, the method and system include the following steps during training:
[0080] A1. Randomly sample a graph from the image features of the SAR image as a template, and sample a graph that matches the template from the image features of the optical image, and form a positive sample pair with the template;
[0081] A2. Sample N graphs that do not match the template from the image features of the optical image, and form N negative sample pairs with the template;
[0082] A3. With the aim of pulling the distance between positive samples closer and pushing the distance between negative samples farther in the feature space, obtain the loss value of the pseudo-siamese neural network structure through the InfoNCE loss function, denoted as the contrastive learning loss;
[0083] A4. Obtain the loss of the classification head through the Focal loss function; obtain the loss of the regression head through the GIoU loss function;
[0084] A5. Take the sum of the contrastive learning loss, the loss of the classification head, and the loss of the regression head as the total loss. With the goal of minimizing the total loss, adjust the parameters of the pseudo-siamese neural network structure, the classification head, and the regression head through backpropagation until the training ends.
[0085] In the specific implementation process, since this method is insensitive to the image orientation, in order to increase the training data, data augmentation methods of random up-down flipping, left-right flipping, and transposition are applied; and since the images in the dataset have different brightnesses, a color jitter method is also used to randomly adjust the brightness and contrast of the training images to expand the training set.
[0086] In an embodiment of the present invention, in order to illustrate the effectiveness of the training method of the present invention, experiments are carried out on four publicly available datasets. The four publicly available datasets are the SEN dataset, the QXS dataset, the WHU dataset, and the SO dataset. In order to simulate the domain generalization problem, three publicly available datasets are selected for training in each group of experiments, and the remaining one publicly available dataset is used for testing. Two models are trained in each group of experiments, namely the training method using the contrastive learning loss and the training method without using the contrastive learning loss, and the matching performance of the models obtained by the two training methods is tested on the test data.
[0087] There are two metrics, Avg L2 and SR, for measuring the performance of the heterologous image matching model. Let the true coordinates of the matching points be , and the coordinates predicted by the model be , then the calculation formula for the L2 distance between the two coordinates is:
[0088] Avg L2 represents the average of the L2 distances of all image pairs on the test data, and its calculation formula is:
[0089]
[0090] Among them, d represents the L2 distance between each pair of images, and n represents the total number of pairs of images in the test data. The smaller the value of Avg L2, the better the matching performance of the model.
[0091] If the L2 distance between the coordinates of the matching points predicted by the model and the true coordinates is less than 3 pixels, it is regarded as a successful match. The number of successfully matched image pairs is counted, and the percentage of successfully matched image pairs among all test image pairs is calculated as another performance metric SR of the cross-source image matching model. The calculation formula of SR is:
[0092]
[0093] Among them represents the number of successfully matched image pairs, and n represents the total number of test images. The larger the SR value, the better the matching performance of the model.
[0094] For the two matching performance metrics, the smaller the value of Avg L2, the better, and the larger the value of SR, the better. From the experimental results shown in Table 1, it can be seen that the method using the contrastive learning loss has achieved better matching performance in all four groups of experiments and is better than the case without using the contrastive learning loss. This shows the effectiveness of the cross-source image matching method based on contrastive learning. In this embodiment, by adding the contrastive learning loss and constraining the feature consistency of the matching points, the generalization performance of the cross-source image matching method on the unknown test domain is improved.
[0095] Table 1: Experimental Results
[0096]
[0097] In another embodiment, the SAR image to be matched is as Figure 5 shown, and the optical image to be matched is as Figure 6 shown. The matching result given by the method / system of the present invention is as Figure 7 shown. Figure 7 The position where the SAR image is matched in the optical image is marked by the green frame in
[0098] In summary, the present invention can effectively perform cross-source image matching, solve the problem of performance degradation of the cross-source image matching model on the unknown test domain, and improve the generalization performance of the cross-source matching model on the unknown test data.
Claims
1. A method for cross-source image matching based on contrastive learning, characterized in that, The following steps are involved: S1. Acquire a SAR image and an optical image to be matched; S2, respectively obtaining the image features of the SAR image and the image features of the optical image through a pseudo-twin neural network structure; S3, calculating the similarity between the image features of the SAR image and the image features of the optical image through a cross-correlation operator; S4, based on the similarity, the matching point position is estimated through the classification head to obtain the rough matching point coordinates; S5. Based on the similarity, the position of the matching point is estimated through the regression head to obtain the offset between the predicted matching point and the actual matching point; S6. Based on the rough matching point coordinates and the offset, the precise matching point coordinates are calculated to obtain a matching result. It also includes a training process, which includes the following steps: A1. Randomly sample a picture from the image features of the SAR image as a template, and sample a picture matching the template from the image features of the optical image, and form a positive sample pair with the template; A2. Sample N images that do not match the template from the image features of the optical image, and form N negative sample pairs with the template; A3. With the purpose of shortening the distance between positive samples and increasing the distance between negative samples in the feature space, the loss value of the pseudo-twin neural network structure is obtained through the InfoNCE loss function, which is recorded as contrastive learning loss. A4. Obtain the loss of the classification head through the Focal loss function; obtain the loss of the regression head through the GIoU loss function; A5. The sum of the contrastive learning loss, the loss of the classification head, and the loss of the regression head is taken as the total loss. With the goal of minimizing the total loss, the pseudo-twin neural network structure, the parameters of the classification head, and the regression head are adjusted by back propagation until the training is completed. The pseudo-twin neural network structure includes a SAR image branch and an optical image branch with the same structure and non-shared weights. Each branch includes a first convolutional layer and a second convolutional layer connected in sequence. The output of the second convolutional layer is connected to the input of the third convolutional layer and the input of the fourth convolutional layer respectively, the output of the fourth convolutional layer is connected to the input of the fifth convolutional layer, the output of the fifth convolutional layer is connected to the input of the sixth convolutional layer, the output of the sixth convolutional layer and the output of the fourth convolutional layer are residually connected, the output of the third convolutional layer and the residual result are spliced and input into the seventh convolutional layer, the output of the seventh convolutional layer is connected to the input of the eighth convolutional layer, the output of the eighth convolutional layer is the output of the current branch, and the input of the first convolutional layer is the input of the current branch.
2. The method for cross-source image matching based on contrastive learning according to claim 1, wherein In step S1, the original SAR image and the optical image are acquired and standardized and normalized respectively, so that the pixel range of the original SAR image and the optical image is constrained to be between 0 and 1, thereby obtaining the SAR image and the optical image to be matched.
3. The method for heterologous image matching based on contrastive learning according to claim 1, characterized in that Both the first and second convolutional layers use convolution with a stride of 2 for four times downsampling.
4. The method for cross-source image matching based on contrastive learning according to claim 1, wherein The specific method of step S3 is: Through depthwise convolution, the image features of the SAR image are used as the convolution kernel, and the image features of the optical image are convolved in each channel to calculate the feature similarity of each position in the search space.
5. The method for heterologous image matching based on contrastive learning according to claim 1, wherein Both the classification head and the regression head include an input channel and an output channel connected in sequence; the input channel structures of the classification head and the regression head are the same and share parameters. The input channel includes four convolutional layers connected in sequence, where the output of the first convolutional layer is connected to the output of the fourth convolutional layer in a residual connection and serves as the input of the fifth convolutional layer, and the output of the fifth convolutional layer is the output of the input channel. In the classification head, the output channel is a convolutional layer with 1 channel number. In the regression head, the output channel is a convolutional layer with 2 channel numbers.
6. The method for heterologous image matching based on contrastive learning according to claim 1, wherein The calculation expression of step S6 is: ; ; wherein are the coordinates of the exact matching points; are the coordinates of the rough matching points; is the offset between the predicted matching point and the true matching point; S is the downsampling factor.
7. The method for heterologous image matching based on contrastive learning according to claim 6, wherein The value of S is 4.
8. A system for the method of cross-source image matching based on contrastive learning according to any one of claims 1 to 7, characterized in that, It includes: A data acquisition module for acquiring the SAR image and the optical image to be matched. An image feature extraction module for respectively obtaining the image features of the SAR image and the optical image through a pseudo-siamese neural network structure. A feature similarity calculation module for calculating the similarity between the image features of the SAR image and the optical image through a cross-correlation operator. A rough matching point coordinate calculation module for estimating the matching point position through the classification head based on the similarity to obtain the rough matching point coordinates. An offset calculation module for estimating the matching point position through the regression head based on the similarity to obtain the offset between the predicted matching point and the true matching point. An image matching module for calculating the accurate matching point coordinates based on the rough matching point coordinates and the offset to obtain the matching result.
Citation Information
Patent Citations
SAR and optical remote sensing image registration method based on pseudo twin convolutional neural network
CN111028277A