Heterogeneous Image Registration Method and Device Based on Feature Joint Driving

By constructing a matching network model driven by feature joint, and combining physical structural features and deep semantic features for heterologous image registration, the problem of insufficient feature description in the prior art is solved, and a more robust and efficient heterologous image registration effect is achieved.

CN116168065BActive Publication Date: 2025-06-10XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211572767.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-06-10
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

The prior art is difficult to fully mine image information in heterologous image registration, especially in the insufficient feature description of the image pixel domain, resulting in certain limitations in heterologous registration. At the same time, although deep learning methods can mine deep semantic information, their interpretability is weak and the description of image structure details and texture information is insufficient.

Method used

The heterologous image registration method based on feature joint drive is adopted, and the matching network model driven by feature joint drive is constructed, combining physical structural features and deep semantic features for registration. The model includes feature extraction module, feature fusion module and matching module. It uses RIFT and HardNet models to extract physical structure features and deep semantic features respectively, and splicing them through feature fusion modules to generate a more comprehensive feature vector.

Benefits of technology

This achieves a more comprehensive description of image information, enhances the robustness of heterologous image registration, and improves matching performance through difficult sample sampling strategies, allowing more accurate image matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116168065B_ABST
    Figure CN116168065B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for heterologous image registration based on feature joint driving, which relates to the technical field of image processing, and includes: obtaining paired optical scene graphs and synthetic aperture radar scene graphs; obtaining feature points in the scene graphs, cropping with the feature points as the center to obtain scene graph slices; constructing a matching network model driven by feature joint, training the model, and obtaining the physical structure features and deep semantic features corresponding to each scene graph slice based on the trained model; splicing the physical structure features and deep semantic features corresponding to the scene graph slices end to end to obtain feature vectors; obtaining the distances between the features in the feature vectors of the optical scene graph slices and the features in the feature vectors of the synthetic aperture radar scene graph slices to obtain matching point pairs in the scene graphs, and registering the optical scene graph and the synthetic aperture radar based on the matching point pairs. This application can make full use of the shallow structure information and deep semantic information of images to achieve image matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method and device for registering heterogeneous images driven by feature combination. Background Art

[0002] Image registration is to align two or more images captured at different times or angles, and has been widely applied in fields such as image stitching, image fusion, change detection, and land detection. At the same time, with the rapid development of remote sensing technology, the earth observation images generated by various types of sensors such as visible light, infrared, and synthetic aperture radar are becoming increasingly rich. Therefore, the registration of heterogeneous images generated by various types of sensors also has a wide range of applications.

[0003] In the prior art, most traditional image registration methods establish local correspondence relationships between two images based on the extracted physical features, and derive the transformation parameters for image registration based on this correspondence relationship to achieve registration, such as SIFT, SURF, FAST, etc. A series of SIFT variant algorithms have also emerged for heterogeneous image registration. In 2020, Li et al. proposed the RIFT algorithm. This algorithm improves the method of using image intensity information to detect and describe feature points in RIFT to address the problem of large radiation differences in heterogeneous images, uses phase consistency information instead of image intensity information, and at the same time proposes the concept of the maximum index map. With the continuous exploration of the image registration field and the continuous expansion of application scenarios, a series of variant algorithms of RIFT only perform feature description in the image pixel domain and cannot well mine deeper information of the image, having certain limitations in heterogeneous registration.

[0004] At the same time, with the rapid development of deep learning and the emergence of some data sets, deep networks have made major breakthroughs in the field of image matching with their powerful representation and feature extraction capabilities, and have higher matching accuracy compared to traditional methods. In 2017, two data sets that can be used for optical image and SAR image matching appeared successively. Hughes first used this data set and adopted a dual-branch convolutional neural network to learn the similarity between visible light and SAR images, and finally transformed the matching problem into a binary classification problem. Although the dual-branch network is used to extract heterogeneous image features respectively, this similarity measurement method of judging whether to match according to a fixed threshold is absolute. Although the deep semantic features extracted based on the deep network can fully mine the deep semantic information of the image, their interpretability is weak, and the description of physical features such as the structural details and texture information of the image is not sufficient.

[0005] Therefore, it is urgent to improve the defects in the prior art. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the present invention provides a method and apparatus for registering heterogeneous images based on feature joint driving. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] In a first aspect, the present application provides a method for registering heterogeneous images based on feature joint driving, including:

[0008] Obtain paired optical scene maps and synthetic aperture radar scene maps;

[0009] Obtain feature points in the optical scene map and feature points in the synthetic aperture radar scene map, crop the optical scene map centered on the feature points in the optical scene map to obtain an optical scene map slice, and crop the synthetic aperture radar scene map centered on the feature points in the synthetic aperture radar scene map to obtain a synthetic aperture radar scene map slice;

[0010] Construct a matching network model based on feature joint driving, train the model to obtain a trained matching network model based on feature joint driving, obtain the physical structure features and deep semantic features corresponding to the optical scene map slice according to the trained matching network model based on feature joint driving, and obtain the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene map slice; concatenate the physical structure features and deep semantic features corresponding to the optical scene map slice head-to-tail to obtain a feature vector of the optical scene map slice; concatenate the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene map slice head-to-tail to obtain a feature vector of the synthetic aperture radar scene map slice; wherein, the matching network model based on feature joint driving includes a feature extraction module, a feature fusion module and a matching module, the feature extraction module includes a first branch module and a second branch module, the first branch module is a RIFT model, and the second branch module is a HardNet model;

[0011] Obtain the distances between the features in the feature vector of the optical scene map slice and the features in the feature vector of the synthetic aperture radar scene map slice, obtain the matching point pairs in the optical scene map and the synthetic aperture radar scene map according to the nearest neighbor matching strategy, eliminate the wrong matching points according to RANSAC, and register the optical scene map and the synthetic aperture radar with the remaining matching points.

[0012] In a second aspect, the present invention further provides a device for registering heterogeneous images based on feature joint driving, including:

[0013] An acquisition module, configured to acquire an optical scene map and a synthetic aperture radar scene map;

[0014] A cutting module, configured to obtain feature points in an optical scene graph and feature points in a synthetic aperture radar scene graph, cut the optical scene graph centered on the feature points in the optical scene graph to obtain an optical scene graph slice, and cut the synthetic aperture radar scene graph centered on the feature points in the synthetic aperture radar scene graph to obtain a synthetic aperture radar scene graph slice;

[0015] A processing module, configured to construct a matching network model driven by feature combination and train the model to obtain a trained matching network model driven by feature combination, obtain physical structure features and deep semantic features corresponding to the optical scene graph slice according to the trained matching network model driven by feature combination, and obtain physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice; concatenate the physical structure features and deep semantic features corresponding to the optical scene graph slice head-to-tail to obtain a feature vector of the optical scene graph slice; concatenate the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice head-to-tail to obtain a feature vector of the synthetic aperture radar scene graph slice; wherein, the matching network model driven by feature combination includes a feature extraction module, a feature fusion module and a matching module, the feature extraction module includes a first branch module and a second branch module, the first branch module is a RIFT model, and the second branch module is a HardNet model;

[0016] A registration module, configured to obtain the distances between the features in the feature vector of the optical scene graph slice and the features in the feature vector of the synthetic aperture radar scene graph slice, obtain matching point pairs in the optical scene graph and the synthetic aperture radar scene graph according to the nearest neighbor matching strategy, and remove incorrect matching points according to RANSAC, and use the remaining matching points to register the optical scene graph and the synthetic aperture radar.

[0017] Advantages of the present invention:

[0018] A method and device for registering heterogeneous images driven by feature combination provided by the present invention, on the one hand, constructs a registration network model driven by feature combination, and more comprehensively describes image information through the joint drive of physical features and deep features to achieve robust registration of heterogeneous images; on the other hand, constructs a suitable difficult sample sampling strategy, so that the network can fully mine difficult-to-separate samples for network optimization during the training process, so as to have better matching performance in actual tasks.

[0019] The following will further describe the present invention in detail with reference to the accompanying drawings and embodiments. Description of the Drawings

[0020] Figure 1 is a flowchart of a method for registering heterogeneous images driven by feature combination provided by an embodiment of the present invention;

[0021] Figure 2 It is another flowchart of the heterologous image registration method based on feature joint driving provided by the embodiments of the present invention;

[0022] Figure 3 It is a schematic structural diagram of a heterologous image registration device based on feature joint driving provided by the embodiments of the present invention. Specific embodiments

[0023] The present invention will be further described in detail below in conjunction with specific embodiments, but the embodiments of the present invention are not limited thereto.

[0024] In the prior art, most heterologous image registration technologies extract single-domain features to describe image information, such as extracting physical features that describe the texture structure information of the image domain or deep semantic features based on a feature extraction network; however, a single feature still has certain limitations in describing image information in different scenarios: Although physical features are not sensitive to changes in the scene, they are based on simple shallow information and have insufficient ability to mine image information; Deep semantic features can fully mine the deep semantic information of images, but their interpretability is weak, and the description of physical features such as the structural details and texture information of the original image is not sufficient.

[0025] Secondly, most of the existing heterologous image registration methods based on deep learning are aimed at how to design a feature extraction network to extract more comprehensive and accurate feature descriptors, without considering the situation that there are often difficult-to-separate sample data in actual registration tasks.

[0026] In view of this, the present invention provides a heterologous image registration method based on feature joint driving. On the one hand, a registration network model driven by feature joint is constructed to more comprehensively describe image information through the joint driving of physical features and deep features, so as to achieve robust heterologous image registration; on the other hand, a suitable difficult sample sampling strategy is constructed, so that the network can fully mine difficult-to-separate samples for network optimization during the training process, so as to have better matching performance in actual tasks.

[0027] Please refer to Figure 1 , Figure 1 It is a flowchart of a heterologous image registration method based on feature joint driving provided by the embodiments of the present invention. A heterologous image registration method based on feature joint driving provided by the present application includes:

[0028] S101. Obtain the paired optical scene map and synthetic aperture radar scene map;

[0029] S102, acquiring feature points in the optical scene graph and feature points in the synthetic aperture radar scene graph, and cutting the optical scene graph with the feature points in the optical scene graph as the center to obtain optical scene graph slices, and cutting the synthetic aperture radar scene graph with the feature points in the synthetic aperture radar scene graph as the center to obtain synthetic aperture radar scene graph slices;

[0030] S103, constructing a matching network model based on feature joint drive, and training the model to obtain a trained matching network model driven by feature joint drive, obtaining physical structure features and deep semantic features corresponding to the optical scene graph slice according to the trained matching network model driven by feature joint drive, and obtaining physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice; splicing the physical structure features and deep semantic features corresponding to the optical scene graph slice head to tail to obtain a feature vector of the optical scene graph slice; splicing the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice head to tail to obtain a feature vector of the synthetic aperture radar scene graph slice; wherein the matching network model based on feature joint drive includes a feature extraction module, a feature fusion module and a matching module, the feature extraction module includes a first branch module and a second branch module, the first branch module is a RIFT model, and the second branch module is a HardNet model;

[0031] S104, obtaining the distance between each feature in the feature vector of the optical scene graph slice and each feature in the feature vector of the synthetic aperture radar scene graph slice, obtaining matching point pairs in the optical scene graph and the synthetic aperture radar scene graph according to the nearest neighbor matching strategy, and eliminating erroneous matching points according to RANSAC, and using the remaining matching points to align the optical scene graph and the synthetic aperture radar.

[0032] Specifically, the feature-jointly driven heterogeneous image registration method provided in this embodiment can fully utilize the shallow structural information and deep semantic information of the image to achieve more accurate image matching by constructing a matching network model jointly driven by physical features and semantic features.

[0033] In an optional embodiment of the present invention, the RIFT model includes:

[0034] Constructing a maximum index map of an optical scene graph slice using a two-dimensional logarithmic Gabor convolution sequence, and constructing a maximum index map of a synthetic aperture radar scene graph slice using a two-dimensional logarithmic Gabor convolution sequence;

[0035] According to the maximum index map of the optical scene graph slice, a corresponding first distribution histogram is constructed to obtain the physical structure characteristics of the optical scene graph slice; according to the maximum index map of the synthetic aperture radar slice, a corresponding second distribution histogram is constructed to obtain the physical structure characteristics of the synthetic aperture radar slice.

[0036] In an alternative embodiment of the present invention, constructing the maximum index map of the optical scene graph slice using a two-dimensional logarithmic Gabor convolution sequence includes:

[0037] Obtain a two-dimensional logarithmic Gabor filter, and its expression is:

[0038]

[0039] where (ρ, θ) are logarithmic polar coordinates, s is the scale of the filter, o is the orientation of the filter, (ρ s , θ so ) is the center frequency of the filter, σ ρ is the bandwidth of ρ, and σ θ is the bandwidth of θ;

[0040] Perform the inverse Fourier transform on the two-dimensional logarithmic Gabor filter to obtain the spatial domain filter L(x, y, s, o), and its expression is:

[0041] L(x, y, s, o) = L even (x, y, s, o) + iL odd (x, y, s, o);

[0042] where the real part L even (x, y, s, o) and the imaginary part L odd (x, y, s, o) are an even-symmetric filter and an odd-symmetric filter respectively; (x, y) are the position coordinates of each point in the spatial domain filter;

[0043] Convolve the optical scene graph slice with the two-dimensional logarithmic Gabor filter to obtain the first convolution component E so (x, y) and the second convolution component O so (x, y), and their expressions are:

[0044] [E so (x, y), O so (x, y)] = [I o (x, y) * L even (x, y, s, o), I o (x, y) * L odd (x, y, s, o)];

[0045] where I o (x, y) is the optical scene graph slice;

[0046] After passing through the two-dimensional logarithmic Gabor filter with scale s and orientation o, obtain the amplitude A so (x, y) of the convolution graph, and its expression is:

[0047]

[0048] Sum the amplitudes A of the convolution maps corresponding to all scales in the direction o so (x, y) to obtain A o (x, y), and its expression is:

[0049]

[0050] where N s is the total number of scales;

[0051] According to the amplitude of the convolution map and A o (x, y), construct an N o -layer convolution map sequence N o is the number of different directions,

[0052] Based on each pixel point (x j , y j ) in the convolution map sequence, obtain its maximum value A o in N max (x j , y j ) in the directions, and the layer number ω o corresponding to the maximum value in the N max -layer convolution map sequence, and its expression is:

[0053]

[0054] Obtain each pixel point in the optical scene map slice, obtain its corresponding ω max , and use it as the pixel value of each pixel point in the maximum index map to obtain the maximum index map of the optical scene map slice.

[0055] In an optional embodiment of the present invention, constructing the maximum index map of the synthetic aperture radar scene map slice using a two-dimensional logarithmic Gabor convolution sequence includes:

[0056] Obtain a two-dimensional logarithmic Gabor filter, and its expression is:

[0057]

[0058] where (ρ, θ) are logarithmic polar coordinates, s is the scale of the filter, o is the direction of the filter, (ρ s , θ so ) is the center frequency of the filter, σ ρ is the bandwidth of ρ, and σ θ is the bandwidth of θ;

[0059] The two-dimensional logarithmic Gabor filter is subjected to inverse Fourier transform to obtain the spatial domain filter L(x, y, s, o), and its expression is:

[0060] L(x, y, s, o) = L even (x, y, s, o) + iL odd (x, y, s, o);

[0061] Among them, the real part L even (x, y, s, o) and the imaginary part L odd (x, y, s, o) are an even-symmetric filter and an odd-symmetric filter respectively; (x, y) are the position coordinates of each point in this spatial domain filter;

[0062] The synthetic aperture radar scene map slice is convolved with the two-dimensional logarithmic Gabor filter to obtain the first convolution component E so (x, y) and the second convolution component O so (x, y), and its expression is:

[0063] [E so (x, y), O so (x, y)] = [I s (x, y) * L even (x, y, s, o), I s (x, y) * L odd (x, y, s, o)];

[0064] Among them, I s (x, y) is the synthetic aperture radar scene map slice;

[0065] After passing through the two-dimensional logarithmic Gabor filter with a scale of s and a direction of o, the amplitude A so (x, y) of the convolution map is obtained, and its expression is:

[0066]

[0067] The amplitudes A so (x, y) of the convolution maps corresponding to all scales with a direction of o are summed to obtain A o (x, y), and its expression is:

[0068]

[0069] Among them, N s is the total number of scales;

[0070] According to the amplitude of the convolution map and A o (x, y), an N o -layer convolution map sequence N ois the number of different directions,

[0071] Based on each pixel point (x j , y j ) in the convolutional graph sequence, obtain its maximum value A o in N max (x j , y j ) in the directions, and the layer number ω o corresponding to the maximum value in the N max -layer convolutional graph sequence. Its expression is:

[0072]

[0073] Obtain each pixel point in the synthetic aperture radar scene graph slice, obtain its corresponding ω max , and use it as the pixel value of each pixel point in the maximum index graph to obtain the maximum index graph of the synthetic aperture radar scene graph slice.

[0074] In an optional embodiment of the present invention, the HardNet model includes multiple convolutional layers, multiple Batch Norm layers, and multiple ReLU activation functions; each convolutional layer corresponds to one Batch Norm layer and one ReLU activation function.

[0075] In an optional embodiment of the present invention, the convolutional layer includes seven layers. Among them, the convolutional kernel sizes of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer are all 3×3, and the convolutional kernel size of the seventh convolutional layer is 8×8.

[0076] In an optional embodiment of the present invention, the structures of the HardNet model corresponding to the optical scene graph slice and the HardNet model corresponding to the synthetic aperture radar scene graph slice are the same, and the weights are not shared.

[0077] In an optional embodiment of the present invention, a triplet loss function is used to constrain the feature vectors of the optical scene graph slice and the feature vectors of the synthetic aperture radar scene graph slice;

[0078] Among them, the expression of the triplet loss function is:

[0079] L tri = max(d(a, p)-d(a, n)+margin, 0);

[0080] Among them, a is the reference sample, p is the matching positive example sample, n is the non-matching negative example sample, margin is a constant greater than 0, and d(·) is the Euclidean distance between two vectors.

[0081] Specifically, in this embodiment, during the network training phase, both hard positive sample mining and hard negative sample sampling strategies are adopted, achieving the full mining of difficult-to-separate samples in the actual registration task.

[0082] In an alternative embodiment of the present invention, please refer to Figure 2 , Figure 2 which is another flowchart of the heterologous image registration method based on feature joint driving provided by the embodiments of the present invention, and the registration of the optical scene graph and the synthetic aperture radar scene graph is achieved through the following steps.

[0083] (1) Image preprocessing

[0084] Select scene graphs with rich information to construct the dataset SEN1-2, which includes paired registered optical scene graphs and SAR (Synthetic Aperture Radar) scene graphs; among them, both the optical scene graphs and SAR scene graphs in the dataset are images with rich feature points to ensure that the scene graphs used for training and testing contain rich scene information.

[0085] Based on the dataset SEN1-2, obtain the training set and the test set. Select 256×256 scene graphs to divide the training set and the test set, and use the RIFT model to extract the feature points of the optical scene graphs and SAR scene graphs. Crop around each feature point to obtain m 64×64 optical scene graph slices and obtain m 64×64 SAR scene graph slices to form m pairs of matching sample data for network training and testing where m 1 = m 2 = m, m 1 is the number of optical scene graph slices, and m 2 is the number of SAR scene graph slices.

[0086] (2) Construct a feature joint-driven matching network model;

[0087] The feature joint-driven matching network model includes a feature extraction module, a feature fusion module, and a matching module.

[0088] a. Feature extraction module, which includes a two-branch network structure. Among them, the first branch is a RIFT structure feature extraction module that can process 64×64 scene graph slices to obtain the corresponding physical structure features r 0 and r s , where the feature r 0 and r sAll are physical features; the second branch is a semantic feature extraction module, which can adopt the HardNet model and can process a 64×64 scene graph slice to obtain the corresponding semantic feature f o and f s , the feature f o and f s are both deep semantic features.

[0089] Among them, the RIFT structure feature extraction module mainly includes two parts: feature space construction and feature descriptor construction. It uses a two-dimensional logarithmic Gabor convolution sequence to construct the maximum index map of the scene graph slice, including the optical scene graph slice and the synthetic aperture radar scene graph slice; then constructs the corresponding distribution histogram for the maximum index map to obtain the RIFT feature of the scene graph slice, that is, the physical structure feature; for details, please refer to the RIFT models for obtaining the physical features of the optical scene graph slice and the RIFT models for obtaining the physical features of the synthetic aperture radar scene graph slice provided in the above embodiments;

[0090] In actual operation, for the extraction of RIFT structure features, a window with a size of l×l slides from the initial position in the upper left corner of the 64×64 maximum index map with a sliding step of s until the window reaches the lower right corner position of the scene graph slice. A 1×o-dimensional distribution histogram is constructed for each small window, and the final m-dimensional RIFT structure feature of the scene graph slice is obtained by connecting all the histograms and normalizing them. The maximum index maps obtained from the optical image slice and the SAR image slice are respectively used to obtain the m-dimensional RIFT structure feature of the corresponding optical scene graph slice and the m-dimensional RIFT structure feature of the SAR scene graph slice based on the above operations.

[0091] The deep semantic feature extraction module, the second branch is the HardNet network, including multiple convolutional layers, multiple BatchNorm layers and multiple ReLU activation functions, which are set in sequence as follows: the first convolutional layer, the first Batch Norm layer, the first ReLU activation function, the second convolutional layer, the second Batch Norm layer, the second ReLU activation function, the third convolutional layer, the third Batch Norm layer, the third ReLU activation function, the fourth convolutional layer, the fourth Batch Norm layer, the fourth ReLU activation function, the fifth convolutional layer, the fifth Batch Norm layer, the fifth ReLU activation function, the sixth convolutional layer, the sixth Batch Norm layer, the sixth ReLU activation function, the seventh convolutional layer, the seventh Batch Norm layer, the seventh ReLU activation function. Among them, the convolutional kernel sizes of the first six convolutional layers are all 3×3, the padding methods are all equal-sized padding, and the numbers of convolutional kernels are 32, 32, 64, 64, 128 respectively; the size of the seventh convolutional kernel is 8×8, and the padding method is equal-sized padding; it should be noted that the HardNet network structures for extracting optical scene graph slices and SAR scene graph slices are the same, and the weights are not shared.

[0092] The 64×64 optical scene graph slices and SAR scene graph slices are cropped into 32×32 scene graph slices, processed by the deep semantic feature extraction module, and normalized by L2-Norm to obtain a 128-dimensional feature vector that can represent the deep semantic information of the scene graph slices.

[0093] b. Feature fusion module, based on the features r o and r s extracted by the RIFT feature extraction module, and based on the features f o and f s extracted by the deep semantic feature extraction module, and normalize them, and concatenate the obtained 128-dimensional feature vectors head-to-tail to obtain a 256-dimensional feature vector f o ' and f s ', that is, the 256-dimensional feature vector of the optical scene graph slice is f o ', and the 256-dimensional feature vector of the SAR scene graph slice is f s '.

[0094] c. Feature matching module, in the training stage, for the feature f o ' of the optical scene graph slice and the feature f s'Constrained optimization is performed through a triplet loss function. The triplet loss function needs to construct three samples, which form a triplet. The three samples are a reference sample, a positive example sample that matches the reference sample, and a negative example sample that does not match the reference sample. It should be noted that the purpose of optimizing through the triplet loss function is to make the distance between the optical scene graph features and the corresponding matching SAR scene graph features as close as possible, and the distance between the non-matching SAR scene graph features as far as possible. For the specific triplet loss function, please refer to the above embodiments, and the present invention will not elaborate here

[0095] In the test stage, directly calculate the Euclidean distance between the optical scene graph slice feature f o ' and the SAR scene graph slice feature f s ', and obtain the matching optical-SAR image slice pairs according to the nearest neighbor rule to judge the accuracy of the feature jointly driven matching network model.

[0096] (3) Iterative training of the feature jointly driven matching network model;

[0097] a. Set the number of iterations as q, the maximum number of iterations as Q, Q≥500 and q = 0. In this embodiment, Q = 1000;

[0098] b. Randomly initialize the weights of the feature extraction network F o and the feature extraction network F s using normal distribution random points, and initialize the network bias term with a uniform distribution with a constant value of 0.01 to obtain the initialized feature extraction networks F o and F s ;

[0099] c. Train the constructed matching network model through a jointly optimized strategy;

[0100] Randomly select N 1 pairs of optical-SAR scene graph slices with matching labels from the training set and input them into the feature extraction networks F o and F s that do not share weights to obtain the semantic features f 1 of N o optical scene graph slices (x i ) and the semantic features f 1 of N s SAR scene graph slices (x i ), where i = 1, 2, K, N 1 ; Optionally, in this embodiment, N 1 = 64. In this embodiment, the Adam optimizer is used in the training process, and the Adam decay factor is 0.9.

[0101] The obtained semantic feature f o (x i ) and f s (x i ) are input into the matching network to calculate the loss function L o (x i ), i = 1, 2, ..., N} and {f s (x i ), i = 1, 2, ..., N}, and use the gradient descent method to optimize the loss function L tri to update the parameters of the feature extraction networks F tri and F o and F s so that the distance between the hidden layer features of the matched scene graph slices is as small as possible, and an updated deep semantic feature extraction module is obtained.

[0102] Randomly select N 2 optical-SAR scene graph slices with matching labels from the training set and input them into the RIFT structure feature extraction module and the above-mentioned trained deep semantic feature extraction module, and respectively obtain the structure features r 2 of N o (x i ) and semantic features f o (x i ) of optical scene graph slices, and the structure features r 2 of N s (x i ) and semantic features f s (x i ) of SAR scene graph slices, where i = 1, 2, ..., N 2 , optionally, in this embodiment N 2 = 128.

[0103] Furthermore, the structure features and semantic features of the optical scene graph slices and the SAR scene graph slices are fused respectively to obtain the fused features f o '(x i ) and f s '(x i ), where i = 1, 2, ..., N 2 , and input into the matching module to re-optimize the loss function L tri using the gradient descent method to update the parameters of the feature extraction networks F o and F sPerform parameter fine-tuning, where the value of the loss function is calculated based on the fused features; so that the trained matching network model can well achieve the joint drive of physical features and deep semantic features for the matching of optical scene graph slices and SAR scene graph slices. Optionally, in this embodiment, the momentum SGD optimizer is used during the training process, and the momentum value is 0.93. It should be noted that in this embodiment, one of the sub-modules is pre-trained first, and then the pre-trained module is fine-tuned through joint training, which greatly reduces the training difficulty, improves the optimization efficiency, and can also achieve better matching performance.

[0104] d. Sample sampling strategy;

[0105] In the actual cross-source image registration task, there are often difficult-to-distinguish samples, which makes it difficult to optimize the training process of the matching network. In this embodiment, the hard negative sample mining strategy and the hard positive sample mining strategy are used to optimize the loss function, so that the network can have better matching ability in the actual task.

[0106] Hard negative sample mining strategy: For N pairs of optical-SAR scene graph slices with matching labels randomly selected from the training set during the network training process, each time an iteration update is performed, randomly select the optical scene graph slice X o i as the reference sample and the SAR scene graph slice X s i as the positive example sample. Select the sample with the closest distance to the i-th pair of scene slices from the remaining N - 1 pairs of optical-SAR scene graph slice pairs as the negative example sample, that is:

[0107] j min = argmin j=1,2,K,n,j≠i d(X o i , X j ), k min = argmin k=1,2,K,n,k≠i d(X s i , X k );

[0108] If then select as the negative example sample, otherwise select as the negative example sample. Construct a triple with the above three samples to form a triple loss function for optimization.

[0109] Hard positive sample mining strategy: For randomly selected N pairs of optical-SAR scene graph slices with matching labels, perform forward propagation in the network, and construct triples according to the above steps. Calculate each pair of optical-SAR scene graph slices X oi and X s i The distance d(X o i , X s i ), where i = 1, 2, …, N; it should be noted that only the first K pairs of image slices with the farthest distance are selected as hard positive samples for backpropagation, and the remaining N - K pairs of easy-to-separate samples no longer go through the backpropagation of the network. Optionally, in this embodiment, K:N = 2:3.

[0110] (4) Implementing heterologous image registration based on the trained feature joint-driven matching network model

[0111] Input the paired optical scene graph and synthetic aperture radar scene graph into the trained feature joint-driven matching network model to obtain 256-dimensional feature vectors, respectively:

[0112]

[0113]

[0114] Calculate the distances between the features f o ' in the optical scene graph I o ' and the features f s ' in the SAR scene graph I s ', obtain the matching point pairs in the optical scene graph and the SAR scene graph according to the nearest neighbor matching strategy, and eliminate the wrong matching point pairs according to RANSAC. Finally, calculate the transformation model according to the remaining correct matching point pairs to achieve image registration.

[0115] Based on the same inventive concept, please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a heterologous image registration device based on feature joint drive provided by an embodiment of the present invention. The present invention also provides a heterologous image registration device based on feature joint drive. For the heterologous image registration device based on feature joint drive provided in the above embodiments of the present invention, please refer to the above embodiments and will not be elaborated here. It includes:

[0116] An acquisition module 201, configured to acquire an optical scene graph and a synthetic aperture radar scene graph;

[0117] A cropping module 202, configured to acquire the feature points in the optical scene graph and the feature points in the synthetic aperture radar scene graph, and crop the optical scene graph with the feature points in the optical scene graph as the center to obtain optical scene graph slices, and crop the synthetic aperture radar scene graph with the feature points in the synthetic aperture radar scene graph as the center to obtain synthetic aperture radar scene graph slices;

[0118] The processing module 203 is used to construct a matching network model driven by feature combination, train the model to obtain a trained matching network model driven by feature combination, obtain the physical structure features and deep semantic features corresponding to the optical scene graph slice according to the trained matching network model driven by feature combination, and obtain the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice; concatenate the physical structure features and deep semantic features corresponding to the optical scene graph slice head to tail to obtain the feature vector of the optical scene graph slice; concatenate the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice head to tail to obtain the feature vector of the synthetic aperture radar scene graph slice; wherein, the matching network model driven by feature combination includes a feature extraction module, a feature fusion module and a matching module, the feature extraction module includes a first branch module and a second branch module, the first branch module is a RIFT model, and the second branch module is a HardNet model;

[0119] The registration module 204 is used to obtain the distances between the features in the feature vector of the optical scene graph slice and the features in the feature vector of the synthetic aperture radar scene graph slice, obtain the matching point pairs in the optical scene graph and the synthetic aperture radar scene graph according to the nearest neighbor matching strategy, eliminate the wrong matching points according to RANSAC, and use the remaining matching points to register the optical scene graph and the synthetic aperture radar.

[0120] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the article or device including the element. "Connection" or "connected" and other similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "up", "down", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0121] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0122] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A heterogeneous image registration method based on feature joint drive, It is characterized in that include: Obtaining a paired optical scene graph and a synthetic aperture radar scene graph; Acquire feature points in the optical scene graph and feature points in the synthetic aperture radar scene graph, and cut the optical scene graph with the feature points in the optical scene graph as the center to obtain an optical scene graph slice, and cut the synthetic aperture radar scene graph with the feature points in the synthetic aperture radar scene graph as the center to obtain a synthetic aperture radar scene graph slice; Constructing a matching network model based on feature joint drive, and training the model to obtain a trained matching network model driven by feature joint drive, and obtaining the physical structure features and deep semantic features corresponding to the optical scene graph slice according to the trained matching network model driven by feature joint drive, and obtaining the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice; The physical structure features and deep semantic features corresponding to the optical scene graph slice are spliced ​​head to tail to obtain a feature vector of the optical scene graph slice; the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice are spliced ​​head to tail to obtain a feature vector of the synthetic aperture radar scene graph slice; wherein the matching network model based on feature joint drive includes a feature extraction module, a feature fusion module and a matching module, the feature extraction module includes a first branch module and a second branch module, the first branch module is a RIFT model, and the second branch module is a HardNet model; The distance between each feature in the feature vector of the optical scene graph slice and each feature in the feature vector of the synthetic aperture radar scene graph slice is obtained, matching point pairs in the optical scene graph and the synthetic aperture radar scene graph are obtained according to a nearest neighbor matching strategy, and erroneous matching points are eliminated according to RANSAC, and the optical scene graph and the synthetic aperture radar are registered using the remaining matching points.

2. According to claim 1, the heterogeneous image registration method based on feature joint drive, It is characterized in that The RIFT model includes: Constructing a maximum index map of the optical scene graph slice using a two-dimensional logarithmic Gabor convolution sequence, and constructing a maximum index map of the synthetic aperture radar scene graph slice using a two-dimensional logarithmic Gabor convolution sequence; According to the maximum index map of the optical scene graph slice, a corresponding first distribution histogram is constructed to obtain the physical structure characteristics of the optical scene graph slice; according to the maximum index map of the synthetic aperture radar slice, a corresponding second distribution histogram is constructed to obtain the physical structure characteristics of the synthetic aperture radar slice.

3. According to claim 2, the heterogeneous image registration method based on feature joint drive, It is characterized in that The method of constructing the maximum index graph of the optical scene graph slice using the two-dimensional logarithmic Gabor convolution sequence comprises: Get the two-dimensional logarithmic Gabor filter, which is expressed as: where (ρ, θ) are the log-polar coordinates, s is the scale of the filter, o is the orientation of the filter, (ρ s , θ so ) is the center frequency of the filter, σ ρ is the bandwidth of ρ, and σ θ is the bandwidth of θ; The two-dimensional logarithmic Gabor filter is subjected to inverse Fourier transform to obtain the spatial domain filter L(x, y, s, o), which is expressed as: L(x,y,s,o) = L even (x,y,s,o) + iL odd (x,y,s,o); where the real part \(L\) even (x, y, s, o) and the imaginary part \(L\) odd (x, y, s, o) are an even-symmetric filter and an odd-symmetric filter respectively; (x, y) are the position coordinates of each point in this spatial-domain filter; Slice the optical scene graph I o Convolve (x,y) with a two-dimensional log Gabor filter to obtain a first convolution component E so (x,y) and a second convolution component O so (x,y), and its expression is: [E so (x,y),O so (x,y)] = [I o (x,y)*L even (x,y,s,o),I o (x,y)*L odd (x,y,s,o)]; After passing through a two-dimensional logarithmic Gabor filter with scale s and orientation o, the amplitude A of the convolution map is obtained so (x, y), and its expression is: Sum the amplitudes A of the convolution maps corresponding to all scales in the direction o so (x, y) to obtain A o (x, y), and its expression is: Among them, N s is the total number of scales; According to the amplitude of the convolution map and A o (x, y), form N o sequences of convolutional maps with N N o being the number of different directions; Based on each pixel point (x j , y j ) in the convolutional graph sequence, obtain its maximum value A o in N max directions, (x j , y j ), and the layer number ω o corresponding to the maximum value in the N max -layer convolutional graph sequence. Its expression is: Obtain each pixel point in the optical scene graph slice, and obtain its corresponding ω max , and use it as the pixel value of each pixel point in the maximum index graph to obtain the maximum index graph of the optical scene graph slice.

4. The method for registering heterogeneous images driven by feature combination according to claim 2, wherein, the step of constructing the maximum index map of the synthetic aperture radar scene graph slice using the two-dimensional logarithmic Gabor convolution sequence includes: obtaining a two-dimensional logarithmic Gabor filter, the expression of which is: Among them, (ρ,θ) are the log-polar coordinates, s is the scale of the filter, o is the direction of the filter, (ρ s ,θ so ) is the center frequency of the filter, σ ρ is the bandwidth of ρ, σ θ is the bandwidth of θ; performing an inverse Fourier transform on the two-dimensional logarithmic Gabor filter to obtain a spatial domain filter L(x, y, s, o), the expression of which is: L(x,y,s,o) = L even (x,y,s,o) + iL odd (x,y,s,o); Among them, the real part L even (x, y, s, o) and the imaginary part L odd (x, y, s, o) are an even-symmetric filter and an odd-symmetric filter respectively; (x, y) are the position coordinates of each point in this spatial domain filter; Slice the synthetic aperture radar scene map I s Convolve (x, y) with a two-dimensional logarithmic Gabor filter to obtain a first convolution component E so (x, y) and a second convolution component O so (x, y), and its expression is: [E so (x,y),O so (x,y)]=[I s (x,y)*L even (x,y,s,o),I s (x,y)*L odd (x,y,s,o)]; After passing through a two-dimensional logarithmic Gabor filter with a scale of s and a direction of o, the amplitude A of the convolution map is obtained so (x, y), and its expression is: Sum the amplitudes A of the convolution maps corresponding to all scales in the direction o so (x, y) to obtain A o (x, y), and its expression is: Among them, N s is the total number of scales; According to the amplitude value of the convolution graph and A o (x, y), form N o layers of convolution graph sequences N o is the number of different directions; Based on each pixel point (x j , y j ) in the convolutional graph sequence, obtain its maximum value A o in N max directions, (x j , y j ), and the layer number ω o corresponding to the maximum value in the N max -layer convolutional graph sequence. Its expression is: Obtain each pixel point in the synthetic aperture radar scene map slice, and obtain its corresponding ω max , and use it as the pixel value of each pixel point in the maximum index map to obtain the maximum index map of the synthetic aperture radar scene map slice.

5. The method for registering heterogeneous images driven by feature combination according to claim 1, wherein, the HardNet model includes multiple convolutional layers, multiple Batch Norm layers and multiple ReLU activation functions; each convolutional layer corresponds to one Batch Norm layer and one ReLU activation function.

6. The method for registering heterogeneous images driven by feature combination according to claim 5, wherein, the convolutional layer includes seven layers. Among them, the convolutional kernel sizes of the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, the fifth convolutional layer and the sixth convolutional layer are all 3×3, and the convolutional kernel size of the seventh convolutional layer is 8×8.

7. The method for registering heterogeneous images driven by feature combination according to claim 1, wherein, the structures of the HardNet model corresponding to the optical scene graph slice and the HardNet model corresponding to the synthetic aperture radar scene graph slice are the same, but the weights are not shared.

8. The method for registering heterogeneous images driven by feature combination according to claim 1, wherein, it further includes: using a triplet loss function to constrain the feature vectors of the optical scene graph slice and the feature vectors of the synthetic aperture radar scene graph slice; wherein, the expression of the triplet loss function is: L tri = max(d(a,p) - d(a,n) + margin, 0); where a is the reference sample, p is the matching positive example sample, n is the non-matching negative example sample, margin is a constant greater than 0, and d(·) is the Euclidean distance between two vectors.

9. A device for registering heterogeneous images driven by feature combination, wherein, it includes: an acquisition module for acquiring an optical scene graph and a synthetic aperture radar scene graph; a cropping module for obtaining feature points in the optical scene graph and feature points in the synthetic aperture radar scene graph, and cropping the optical scene graph centered on the feature points in the optical scene graph to obtain an optical scene graph slice, and cropping the synthetic aperture radar scene graph centered on the feature points in the synthetic aperture radar scene graph to obtain a synthetic aperture radar scene graph slice; a processing module for constructing a matching network model driven by feature combination and training the model to obtain a trained matching network model driven by feature combination, and obtaining the physical structure features and deep semantic features corresponding to the optical scene graph slice according to the trained matching network model driven by feature combination, and obtaining the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slice. Concatenate the physical structure features and deep semantic features corresponding to the optical scene graph slices from beginning to end to obtain the feature vector of the optical scene graph slices; concatenate the physical structure features and deep semantic features corresponding to the synthetic aperture radar scene graph slices from beginning to end to obtain the feature vector of the synthetic aperture radar scene graph slices; wherein, the matching network model based on feature joint driving includes a feature extraction module, a feature fusion module and a matching module, the feature extraction module includes a first branch module and a second branch module, the first branch module is the RIFT model, and the second branch module is the HardNet model; A registration module, configured to obtain the distances between the features in the feature vector of the optical scene graph slices and the features in the feature vector of the synthetic aperture radar scene graph slices, obtain the matching point pairs in the optical scene graph and the synthetic aperture radar scene graph according to the nearest neighbor matching strategy, remove the incorrect matching points according to RANSAC, and use the remaining matching points to register the optical scene graph and the synthetic aperture radar.

Citation Information

Patent Citations

  • Method for mapping SAR (Synthetic Aperture Radar) image to optical image based on space-frequency characteristic consistency

    CN112883908A

  • Multi-modal image registration method and system based on depth global features

    CN113223068A