Multi-modal image matching method and system, and terminal device and storage medium
Through self-supervised feature extraction and dual-branch network methods, deep feature extraction and repeatability problems in heterologous image matching are solved, and image matching with higher accuracy and robustness are achieved.
Patent Information
- Application Number
- PCT/CN2024/103035
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2024-07-02
- Publication Date
- 2025-07-17
AI Technical Summary
In the prior art, feature-based methods cannot obtain deeper features between heterologous images for matching. Region-based methods are sensitive to grayscale transformation and have a large amount of calculation. Neural network methods ignore repeatable feature point extraction.
The self-supervised feature extraction method is adopted to match the optical image and SAR image through a dual-branch network, and the neural network is used to learn repeatable feature points, and the deep features are extracted in combination with the transformer-CNN framework. The training method of description then detection is used to connect the feature detection and description modules.
It improves the accuracy and robustness of heterologous image matching, has strong anti-interference ability, can extract richer deep features, effectively improving the matching effect.
Smart Images

Figure CN2024103035_17072025_PF_FP_ABST
Abstract
Description
Multimodal image matching method, system, terminal device and storage medium Technical Field
[0001] The present invention relates to the fields of deep learning and image technology, and in particular to a multimodal image matching method, system, terminal device and storage medium. Background Art
[0002] SAR-optical image matching is a key step in multimodal image information fusion. Image matching methods can be divided into feature-based and region-based methods.
[0003] Feature-based methods include classic methods such as SIFT and SAR-SIFT, which are widely used. To address the matching challenges caused by nonlinear radiation distortion and significant image intensity differences between SAR and optical images, Li proposed a phase-consistency-based method, RIFT, which uses phase congruence instead of pixel intensity for feature extraction. Phase congruence is based on frequency domain detection features and is well-suited to nonlinear radiation differences between images. Similarly, the LINFT method proposes that the HOG feature is robust to intensity variations and can achieve good results when applied to optical and SAR image matching tasks.
[0004] Region-based methods use the grayscale information of images to measure the similarity between them. Common similarity metrics include normalized cross correlation (NCC), mutual information (MI), and mean square error (SSD). To address the nonlinear radiometric distortion between heterogeneous images, Ye proposed a new similarity metric, CFOG, based on the structural similarity of images. CFOG is based on the HOG feature and represents the image pixel by pixel by calculating the gradient direction histogram of the pixels. Furthermore, CFOG transforms structural features into frequency space and accelerates matching using the fast Fourier transform.
[0005] To obtain deep features, many researchers have adopted deep neural networks for feature extraction and description. Wu used a CNN network to generate feature maps that correlate between heterogeneous images pixel by pixel, and then used template matching to obtain matching relationships. Learning-based methods eliminate the need for manually designing complex descriptors. Liao et al. designed a deep convolutional network, MatchosNet, which contains dense blocks and cross-stage local networks to generate deep descriptors for matching SAR and optical images. Furthermore, Mihai Dusmanu proposed a method that uses a single convolutional neural network to simultaneously detect features and extract dense feature descriptors. By moving the detection of feature points to a later stage of processing, more stable key points are obtained. These methods all use deep networks to obtain high-level features that correlate between heterogeneous images.
[0006] However, all of the above methods have their own shortcomings. Feature-based methods, such as SIFT and SAR-SIFT, typically rely on image gradient information for feature extraction and description, and are not very effective in heterogeneous matching of SAR and optical images. Methods such as RIFT and LINFT, which adapt SAR and optical images, can combat nonlinear intensity differences to a certain extent, but they cannot capture the deep and complex feature representations shared by heterogeneous images, and instead rely on shallow features (such as corners and edges) for matching. This approach is easily affected by noise.
[0007] SAR images and optical images suffer from nonlinear radiometric distortion and large differences in pixel intensity. Region-based methods, which rely on the grayscale information of the image, are not applicable and are susceptible to image distortion. The computational complexity is high, limiting their application in multimodal image matching.
[0008] Current methods for using neural networks to extract high-level features related to heterogeneous images typically use a single deep convolutional network (CNN) for feature extraction and description. This approach suffers from poor interpretability and struggles to mine global feature information from heterogeneous images. Furthermore, these methods fail to obtain repeatable feature points, increasing the indistinguishability between matching positive samples and similar negative samples.
[0009] Summary of the Invention
[0010] The technical problems addressed by the present invention are that feature-based methods are unable to obtain deeper features between heterogeneous images for matching; region-based methods are sensitive to grayscale transformations, easily affected by image distortion, and require high computational complexity; and typical neural network-based heterogeneous image matching methods use a single CNN network for feature extraction and ignore the extraction of repeatable feature points. To address these technical problems, the present invention provides a multimodal image matching method, system, terminal device, and storage medium.
[0011] In a first aspect, an embodiment of the present invention provides a multimodal image matching method, comprising:
[0012] performing self-supervised feature extraction on the optical image and the SAR image to obtain repeatable feature points between the optical image and the SAR image;
[0013] Segmenting the optical image and the SAR image into a first image block sequence according to the repeatable feature points, and performing feature extraction on the first image block sequence using a dual-branch network to obtain feature description vectors for the optical image and the SAR image, respectively; the dual-branch network includes a first branch network for extracting global features and a second branch network for extracting local features;
[0014] Feature matching is performed on the optical image and the SAR image according to the feature description vector to obtain matching point pairs between the optical image and the SAR image.
[0015] Preferably, before performing self-supervised feature extraction on the optical image and the SAR image, the method further comprises:
[0016] A local normalization filter is used to convert optical images and SAR images into normalized images.
[0017] Preferably, performing self-supervised feature extraction on the optical image and the SAR image to obtain repeatable feature points between the optical image and the SAR image includes:
[0018] Combining the optical image and the SAR image into a SAR-optical image pair;
[0019] Performing self-supervised learning on the SAR-optical image pair to obtain a keypoint score map;
[0020] A loss function is constructed according to the key point score map, and the self-supervised learning is optimized using the loss function to learn repeatable feature points between the optical image and the SAR image.
[0021] Preferably, constructing a loss function according to the key point score map includes:
[0022] Dividing the key point score map into a second image block sequence; the second image block sequence includes a SAR image block sequence and an optical image block sequence;
[0023] Constructing a loss function based on the second image block sequence; the loss function includes a first loss function, a second loss function and a third loss function;
[0024] The first loss function is expressed as follows:
[0025] Among them, I s represents the SAR image, I o represents the optical image, U represents the affine transformation, Score o (i) represents the i-th image block in the optical image block sequence, Score s U(i) represents the i-th image block of the SAR image block sequence, and m represents the number of image blocks in the second image block sequence;
[0026] The second loss function is expressed as follows:
[0027] Where I represents the SAR-optical image pair, Score(i) represents the i-th image block of the second image block sequence;
[0028] The third loss function is expressed by the following formula:
[0029] L3=1-[AP ij Score ij +γ(1-Score ij )]
[0030] Among them, AP ij Score represents the average accuracy of the pixel at position (i, j) of the SAR-optical image pair. ij represents the score corresponding to the pixel at position (i, j) of the SAR-optical image pair, and γ represents a hyperparameter; the average precision represents the similarity between the feature description vectors of the pixels at corresponding positions of the SAR-optical image pair;
[0031] The first loss function, the second loss function and the third loss function are weightedly summed to obtain the loss function.
[0032] Preferably, the dual-branch network further includes a convolutional layer for extracting shallow features;
[0033] The step of segmenting the optical image and the SAR image into a first image block sequence according to the repeatable feature points, and performing feature extraction on the first image block sequence through a dual-branch network to obtain feature description vectors for the optical image and the SAR image, respectively, includes:
[0034] Segmenting the optical image and the SAR image into a first image block sequence with the repeatable feature point as the center;
[0035] Inputting the first image block sequence into the convolutional layer to obtain shallow features of the first image block sequence;
[0036] Inputting the shallow features into the first branch network and the second branch network respectively to obtain global features and local features of the first image block sequence;
[0037] The global features and the local features are concatenated and processed by convolution to obtain a feature description vector of the first image block sequence.
[0038] Preferably, the first branch network includes a Lite-Transformer, a fully connected layer and a normalization layer; the second branch network includes a first CSP module, a connection layer and a second CSP module.
[0039] Preferably, performing feature matching on the optical image and the SAR image according to the feature description vector to obtain matching point pairs between the optical image and the SAR image includes:
[0040] Performing feature matching on the optical image and the SAR image using a nearest neighbor search method to obtain initial matching point pairs between the optical image and the SAR image;
[0041] The initial matching point pairs are purified using random sampling consistency to obtain target matching point pairs between the optical image and the SAR image.
[0042] In a second aspect, an embodiment of the present invention provides a multimodal image matching system, comprising:
[0043] A feature point extraction module is used to perform self-supervised feature extraction on the optical image and the SAR image to obtain repeatable feature points between the optical image and the SAR image;
[0044] a feature description module, configured to segment the optical image and the SAR image into a first image block sequence based on the repeatable feature points, and perform feature extraction on the first image block sequence using a dual-branch network to obtain feature description vectors for the optical image and the SAR image, respectively; the dual-branch network includes a first branch network for extracting global features and a second branch network for extracting local features;
[0045] A feature matching module is used to perform feature matching on the optical image and the SAR image according to the feature description vector to obtain matching point pairs between the optical image and the SAR image.
[0046] In a third aspect, an embodiment of the present invention provides a terminal device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the multimodal image matching method as described above when executing the computer program.
[0047] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the multimodal image matching method as described above.
[0048] Compared with the prior art, the multimodal image matching method, system, terminal device, and storage medium of the embodiments of the present invention have the following advantages: a neural network is used to autonomously learn areas with strong matching ability between SAR images and optical images as feature points, which has advantages over traditional geometric operators in extracting repeatable feature points from heterogeneous images and effectively improves matching accuracy; a dual-branch network composed of a transformer-CNN joint framework can extract richer deep features shared by heterogeneous SAR and optical images, and has stronger anti-interference ability and better robustness than traditional feature extraction methods and methods based on convolutional neural networks; and a two-stage training method of description then detection is used to link the mutually isolated feature detection module and feature description module in traditional feature-based matching methods, so that the matching relationship between images constrains the selection of feature points, thereby obtaining feature points that are more conducive to matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] FIG1 is a schematic diagram of a flow chart of a multimodal image matching method according to an embodiment of the present invention;
[0050] FIG2 is a schematic diagram of a process for performing self-supervised feature extraction on optical images and SAR images according to an embodiment of the present invention;
[0051] FIG3 is a schematic diagram of a process for constructing a loss function based on a key point score graph according to an embodiment of the present invention;
[0052] FIG4 is a schematic diagram of a repeatable feature point extraction result according to an embodiment of the present invention;
[0053] FIG5 is a schematic diagram of a local normalized filtering result according to an embodiment of the present invention;
[0054] FIG6 is a schematic diagram of the structure of a dual-branch network according to an embodiment of the present invention;
[0055] 7 is a schematic diagram of a process for extracting features from a dual-branch network according to an embodiment of the present invention;
[0056] FIG8 is a schematic diagram of the structure of a first branch network according to an embodiment of the present invention;
[0057] 9 is a schematic diagram of the structure of the second branch network embodiment of the present invention;
[0058] 10 is a schematic diagram of a process for performing feature matching on an optical image and a SAR image according to a feature description vector according to an embodiment of the present invention;
[0059] FIG11 is a schematic diagram of feature matching results according to an embodiment of the present invention;
[0060] FIG12 is a schematic structural diagram of a multimodal image matching system according to an embodiment of the present invention;
[0061] FIG13 is a schematic structural diagram of a terminal device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0063] As shown in FIG1 , an embodiment of the present invention provides a multimodal image matching method, comprising the steps of:
[0064] S1, perform self-supervised feature extraction on the optical image and SAR image to obtain repeatable feature points between the optical image and the SAR image;
[0065] In feature-based heterogeneous image matching, feature point repeatability is a key issue. Due to the distortion of SAR and optical images, traditional geometric operators often fail to extract repeatable feature points. Therefore, this embodiment of the present invention treats feature point extraction as a self-supervised task, using a neural network to learn repeatable feature points between SAR and optical images.
[0066] Furthermore, in a specific embodiment, as shown in FIG2 , step S1 includes:
[0067] S101, combining the optical image and the SAR image into a SAR-optical image pair;
[0068] S102, performing self-supervised learning on the SAR-optical image pair to obtain a key point score map;
[0069] Specifically, a fully convolutional network is used as the backbone network to perform self-supervised learning on SAR-optical image pairs to obtain a keypoint score map, in which the local maximum points in the keypoint score map are selected as keypoints.
[0070] S103. Construct a loss function based on the key point score map, and use the loss function to optimize self-supervised learning to learn repeatable feature points between the optical image and the SAR image.
[0071] In order to learn sparse and repeatable feature points and obtain more effective feature points through the matching relationship between SAR images and optical images, this embodiment proposes three new loss functions to optimize self-supervised learning.
[0072] Furthermore, in a specific embodiment, as shown in FIG3 , constructing a loss function according to the key point score graph in step S103 includes:
[0073] S103-A, dividing the key point score map into a second image block sequence;
[0074] Specifically, during the training phase, the input to the fully convolutional network is a pair of SAR-optical image pairs {I s ,I o After preprocessing, the key point score map {Score s U,Score o}. Score s U is Score s The key point score map is obtained by the known transformation U. The goal of this embodiment is to obtain the common salient features of heterogeneous images, so the output Score s U and Score o Consistent, that is, Score o All local maximum points and Score s The measure taken by this embodiment is to maximize the block average cosine similarity between them. Therefore, the key point score map {Score s U,Score o} is divided into a second image block sequence of N×N size. The second image block sequence includes a SAR image block sequence Score s U(i) (i=1...m) and optical image block sequence Score o (i) (i=1...m).
[0075] S103-B. Construct a loss function based on the second image block sequence.
[0076] Specifically, the loss function includes a first loss function, a second loss function, and a third loss function;
[0077] The first loss function is expressed as follows:
[0078] Among them, I s represents the SAR image, I o represents the optical image, U represents the affine transformation, Score o (i) represents the i-th image block in the optical image block sequence, Score s U(i) represents the i-th image block in the SAR image block sequence, and m represents the number of image blocks in the second image block sequence.
[0079] At the same time, in order to enable each sub-block to learn a local maximum, a second loss function is constructed; the second loss function is expressed by the following formula:
[0080] Wherein, I represents the SAR-optical image pair, and Score(i) represents the i-th image block of the second image block sequence.
[0081] The first loss function and the second loss function can learn the common repeatable features between heterogeneous images and obtain evenly distributed feature points, which is conducive to feature matching.
[0082] The third loss function is used to make the neural network learn key points that are easier to match. s ,I o}∈R W×H After preprocessing, the dense feature description vector {V S ,V O}∈R W*H×256 Then, the global metric average precision (AP) is used to characterize the similarity between the feature description vectors of the corresponding pixels of the heterogeneous images, and the average precision AP corresponding to each pixel of the image is obtained. ij (i=1...H, j=1...W). A larger AP indicates a higher degree of similarity of the feature description vectors. Therefore, in order to obtain key points that are easier to match and remove interference points in flat areas or repeated areas, this embodiment constructs a third loss function; the third loss function is expressed by the following formula:
[0083] L3=1-[AP ij Score ij +γ(1-Score ij )]
[0084] Among them, AP ij Score represents the average accuracy of the pixel at position (i, j) of the SAR-optical image pair. ij represents the score corresponding to the pixel at position (i, j) of the SAR-optical image pair, γ∈[0,1] represents the hyperparameter, representing the minimum expected AP ij Numeric value.
[0085] When the corresponding AP ij When it is less than γ, in order to minimize L3, the neural network learns the score of the position ij Should tend to 0. Similarly, when AP ij When Score is greater than γ, ij should approach 1. Therefore, to enable the neural network to learn features that facilitate correct matching between heterogeneous images, the first, second, and third loss functions are weighted and summed to obtain the loss function. Using the loss function to optimize self-supervised learning, repeatable feature points between optical and SAR images are learned, as shown in Figure 4.
[0086] It should be noted that before step S1, the optical image and the SAR image are converted into normalized images using a local normalization filter. Taking into account nonlinear radiation distortion, this embodiment uses a local normalization filter to retain relevant detail information between the two modal images, thereby improving matching performance.
[0087] Specifically, the mathematical definition of the local normalization filter is expressed as follows:
[0088] Among them, I(x,y) represents the original image, I norm (x,y) represents the normalized image, and M(x,y,d) is a local window centered at (x,y) with a size of (2*d+1×2*d+1). The local normalization filter is implemented by taking the difference between the original image and its average filtered result. The purpose of the average filter is to remove detailed structural information from the original image. Therefore, the local normalization filter preserves image detail components and helps extract common structural features between heterogeneous images. The filtering results of the local normalization filter can be seen in Figure 5.
[0089] S2. Segmenting the optical image and the SAR image into a first image block sequence according to the repeatable feature points, and performing feature extraction on the first image block sequence through a dual-branch network to obtain feature description vectors of the optical image and the SAR image, respectively;
[0090] To mine multi-scale, multimodal image features, this embodiment proposes a dual-branch network, as shown in Figure 6. Specifically, the dual-branch network includes a first branch network, the Global Transformer Branch, for extracting global features, and a second branch network, the Detail CNN Branch, for extracting local features. Furthermore, the dual-branch network also includes a convolutional layer for extracting shallow features.
[0091] Furthermore, in a specific embodiment, as shown in FIG7 , step S2 includes:
[0092] S201, dividing the optical image and the SAR image into a first image block sequence with the repeatable feature point as the center;
[0093] Specifically, with the extracted repeatable feature points as the center, this embodiment divides the optical image and the SAR image into a first image block sequence of 64×64 size.
[0094] S202, inputting the first image block sequence into a convolutional layer to obtain shallow features of the first image block sequence;
[0095] Specifically, as shown in Figure 6, the convolution layer consists of two 2D convolution layers ConvBR (convolution kernel is 3×3, stride is 2, padding size is 1) and ResNet Block. The first image block sequence of size 64×64 is taken as input, and the shallow features of the first image block sequence are extracted through the convolution layer. The mathematical expression is expressed as follows:
[0096] ConvBR(I)=bn(relu(conv(I)),RES(I)=(bn(relu(conv(I)))+I
[0097] S203, inputting the shallow features into the first branch network and the second branch network respectively to obtain global features and local features of the first image block sequence;
[0098] Based on the good global context information extraction capability of transformer, the first branch network of this embodiment includes Lite-Transformer, fully connected layer and normalization layer. This embodiment uses the spatial self-attention mechanism of transformer to construct the first branch network, as shown in Figure 8. In the self-attention mechanism, the shallow features of the input are first flattened into a series of feature vectors A∈R for each pixel. h*w×c -, c_ is the dimension of the eigenvector. Then use the transfer matrix {W q ,W k ,W v Generate the corresponding triple (query, key, value), namely (Q, K, V), which is mathematically expressed using the following formula:
[0099] The generated triples are then fed into the Attention layer: the weighted inner product E of K and Q is calculated, and the corresponding weights are obtained through the softmax function. The obtained weights are then weighted summed with V to obtain the final attention feature vector B. E and B are specifically expressed using the following formula:
[0100] Β=attention(E,V)=V·softmax(E)
[0101] Finally, the global features of the first image block sequence are obtained through a series of fully connected layers and layer normalization Considering the balance between performance and computational efficiency, this embodiment adopts LT-Block as the basic unit of the first branch network.
[0102] Furthermore, the second branch network of this embodiment includes a first CSP module, a connection layer and a second CSP module. The second branch network uses CSP-Resnet as the backbone network, as shown in Figure 9. The first CSP module maps the input shallow features into two parts along the channel dimension using a 1×1 convolution layer. One part extracts local information through a conventional ResNet Block, and the other part is directly spliced with the output of the former and input into the Transition layer. The Transition layer contains a batch normalization layer, a ReLu activation layer and a 1×1 convolution layer. The second branch network includes two CSP-ResNet Blocks, and the connection layer uses a 3×3 convolution to expand the channel dimension, and the output is a local feature of size 8×8×128.
[0103] It is understood that the first CSP Block shown in Figure 9 is the first CSP module, and the second CSP Block is the second CSP module. Since the working principle of the second CSP module is the same as that of the first CSP module, it has been described above and will not be repeated here.
[0104] S204 , concatenating the global features and the local features and performing convolution processing to obtain a feature description vector of the first image block sequence.
[0105] Specifically, the global feature Φ global With local features Φ detail The final feature description vector {V S ,V O}∈R 1×256 Among them, each convolutional layer is batch normalized and activated by rectified linear unit (ReLu).
[0106] It should be noted that in order to optimize the dual-branch network, this embodiment designs a loss function hard distance loss, which aims to make the network output feature description vector of the image blocks corresponding to the matching SAR image and the optical image, that is, the positive sample The relative distance between them should be as small as possible, and the negative samples should be increased. For each pair of matching SAR-optical images, sample n-1 negative samples and calculate the relative distance between all positive and negative samples corresponding images Select the negative sample with the smallest relative distance With positive samples Forming triples: Specifically, the loss function used is expressed as follows:
[0107] S3. Perform feature matching on the optical image and the SAR image according to the feature description vector to obtain matching point pairs between the optical image and the SAR image.
[0108] Specifically, this embodiment uses the nearest neighbor search method (NN) and the random sampling consensus method (RANSAC) to perform feature matching and feature matching pair purification.
[0109] Furthermore, in a specific embodiment, as shown in FIG10 , step S3 includes:
[0110] S301, performing feature matching on the optical image and the SAR image using a nearest neighbor search method to obtain initial matching point pairs between the optical image and the SAR image;
[0111] Specifically, the Euclidean distance between feature description vectors is used to search for feature points in the SAR image and their nearest and next-nearest neighbors in the optical image. A pre-set first threshold is used to compare the ratio of the distances between the feature description vectors of the feature points in the SAR image and their nearest and next-nearest neighbors in the optical image. If the ratio is greater than the first threshold, the two feature points do not match; otherwise, they do. Therefore, adjusting the number of matching points by controlling the first threshold helps control the accuracy of key points.
[0112] S302: purify the initial matching point pairs using random sampling consistency to obtain target matching point pairs between the optical image and the SAR image.
[0113] The random sampling consistency method can use a continuous iterative method to find the optimal parameter model in a data set containing "outliers". Specifically, this embodiment randomly samples 4 sample matching pairs from the initial matching point pairs, calculates the transformation matrix H between the images, and records it as model M. Then calculate the projection error between all matching point pairs and model M. If the projection error is less than the second threshold, the inlier set I is added. If the number of elements in the inlier set I is greater than the optimal inlier set I_best, update I_best = I. At the same time, update the number of iterations k, repeat the above process to obtain the optimal transformation matrix and then obtain the purified matching point pairs. The specific matching results can be seen in Figure 11.
[0114] An embodiment of the present invention provides a multimodal image matching method having the following beneficial effects: using a neural network to autonomously learn highly matching regions of SAR images and optical images as feature points, which has advantages over traditional geometric operators in extracting repeatable feature points from heterogeneous images, effectively improving matching accuracy; using a dual-branch network composed of a transformer-CNN joint framework to extract richer deep features shared by heterogeneous SAR and optical images, which has stronger anti-interference capabilities and better robustness than traditional feature extraction methods and methods based on convolutional neural networks; and using a two-stage training method of description then detection to link the mutually isolated feature detection module and feature description module in traditional feature-based matching methods, so that the matching relationship between images constrains the selection of feature points, thereby obtaining feature points that are more conducive to matching.
[0115] Based on the above multimodal image matching method, as shown in FIG12 , an embodiment of the present invention further provides a multimodal image matching system, including:
[0116] Feature point extraction module 1, used to perform self-supervised feature extraction on the optical image and the SAR image to obtain repeatable feature points between the optical image and the SAR image;
[0117] Feature description module 2, configured to segment the optical image and the SAR image into a first sequence of image blocks based on repeatable feature points, and extract features from the first sequence of image blocks using a dual-branch network to obtain feature description vectors for the optical image and the SAR image, respectively; the dual-branch network includes a first branch network for extracting global features and a second branch network for extracting local features;
[0118] The feature matching module 3 is used to perform feature matching on the optical image and the SAR image according to the feature description vector to obtain matching point pairs between the optical image and the SAR image.
[0119] It should be noted that each module in the above-mentioned multimodal image matching system can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules. For the specific definition of a multimodal image matching system, please refer to the definition of a multimodal image matching method above. The two have the same functions and effects and will not be repeated here.
[0120] An embodiment of the present invention further provides a terminal device, comprising:
[0121] processor, memory, and bus;
[0122] The bus is used to connect the processor and the memory;
[0123] The memory is used to store operation instructions;
[0124] The processor is configured to call the operation instruction, and the executable instruction enables the processor to perform operations corresponding to the multimodal image matching method described above in the present invention.
[0125] In an optional embodiment, a terminal device is provided, as shown in FIG13 . The terminal device 5000 shown in FIG13 includes a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, for example, via a bus 5002. Optionally, the terminal device 5000 may further include a transceiver 5004. It should be noted that in actual applications, the number of transceivers 5004 is not limited to one, and the structure of the terminal device 5000 does not constitute a limitation on the embodiments of the present invention.
[0126] Processor 5001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, a transistor logic device, a hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 5001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0127] Bus 5002 may include a path for transmitting information between the aforementioned components. Bus 5002 may be a PCI bus or an EISA bus, for example. Bus 5002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG13 shows only one thick line, but this does not indicate that there is only one bus or only one type of bus.
[0128] The memory 5003 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0129] The memory 5003 is used to store application code for executing the solution of the present invention, and the execution is controlled by the processor 5001. The processor 5001 is used to execute the application code stored in the memory 5003 to implement the content shown in any of the above method embodiments.
[0130] Among them, terminal devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0131] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the multimodal image matching method of the present invention is implemented.
[0132] Yet another embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding contents of the aforementioned method embodiments.
[0133] In addition, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.
[0134] In summary, the embodiments of the present invention provide a multimodal image matching method, system, terminal device, and storage medium. These methods utilize a neural network to autonomously learn highly matching areas of SAR and optical images as feature points. Compared with traditional geometric operators, these methods have advantages in extracting repeatable feature points from heterogeneous images, effectively improving matching accuracy. A dual-branch network composed of a transformer-CNN joint framework can extract richer deep features shared by heterogeneous SAR and optical images. Compared with traditional feature extraction methods and methods based on convolutional neural networks, these networks have stronger anti-interference capabilities and better robustness. A two-stage training method, description then detection, is used to link the isolated feature detection module and feature description module in traditional feature-based matching methods. This allows the matching relationship between images to constrain the selection of feature points, thereby obtaining feature points that are more conducive to matching.
[0135] Each embodiment in this specification is described in a progressive manner, and the same or similar parts of each embodiment can be directly referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. It should be noted that the various technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0136] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and substitutions can be made without departing from the technical principles of the present invention. These improvements and substitutions should also be regarded as the scope of protection of the present invention.
Claims
1. A multimodal image matching method, characterized in that, Including: Performing self-supervised feature extraction on the optical image and the SAR image to obtain repeatable feature points between the optical image and the SAR image; Segmenting the optical image and the SAR image into a first sequence of image patches according to the repeatable feature points, and performing feature extraction on the first sequence of image patches through a dual-branch network to respectively obtain feature description vectors of the optical image and the SAR image; the dual-branch network includes a first branch network for extracting global features and a second branch network for extracting local features; Performing feature matching on the optical image and the SAR image according to the feature description vectors to obtain a pair of matching points between the optical image and the SAR image.
2. The multimodal image matching method according to claim 1, wherein Before performing self-supervised feature extraction on the optical image and the SAR image, including: Converting the optical image and the SAR image into normalized images by using a local normalization filter.
3. The multimodal image matching method according to claim 1, characterized in that, The performing self-supervised feature extraction on the optical image and the SAR image to obtain repeatable feature points between the optical image and the SAR image includes: Combining the optical image and the SAR image into a SAR-optical image pair; Performing self-supervised learning on the SAR-optical image pair to obtain a key point score map; Constructing a loss function according to the key point score map, and optimizing the self-supervised learning by using the loss function to learn the repeatable feature points between the optical image and the SAR image.
4. The multimodal image matching method according to claim 3, characterized in that The constructing a loss function according to the key point score map includes: Dividing the key point score map into a second sequence of image patches; the second sequence of image patches includes a SAR image patch sequence and an optical image patch sequence; Constructing a loss function based on the second sequence of image patches; the loss function includes a first loss function, a second loss function, and a third loss function; The first loss function is expressed by the following formula: Among them, I s represents the SAR image, I o represents the optical image, U represents the affine transformation, Score o (i) represents the i-th image block of the optical image block sequence, Score s U(i) represents the i-th image block of the SAR image block sequence, m represents the number of image blocks in the second image block sequence; The second loss function is represented by the following formula: wherein, I represents the SAR-optical image pair, and Score(i) represents the i-th image patch of the second sequence of image patches; The third loss function is represented by the following formula: L3 = 1 - [AP ij Score ij + γ(1 - Score ij )] Among them, AP ij represents the average precision corresponding to the pixel at the position (i, j) of the SAR-optical image pair, and Score ij represents the score corresponding to the pixel at the position (i, j) of the SAR-optical image pair, and γ represents a hyperparameter; the average precision represents the similarity degree between the feature description vectors of the pixels at the corresponding positions of the SAR-optical image pair; Performing weighted summation on the first loss function, the second loss function, and the third loss function to obtain the loss function.
5. The multimodal image matching method according to claim 1, wherein The dual-branch network further includes a convolutional layer for extracting shallow features; The segmenting the optical image and the SAR image into a first sequence of image patches according to the repeatable feature points, and performing feature extraction on the first sequence of image patches through a dual-branch network to respectively obtain feature description vectors of the optical image and the SAR image includes: Centering on the repeatable feature points, segmenting the optical image and the SAR image into a first sequence of image patches; Inputting the first sequence of image patches into the convolutional layer to obtain shallow features of the first sequence of image patches; Respectively inputting the shallow features into the first branch network and the second branch network to obtain global features and local features of the first sequence of image patches; Concatenating the global features and the local features and performing convolutional processing to obtain a feature description vector of the first sequence of image patches.
6. The multimodal image matching method according to claim 5, wherein The first branch network includes a Lite-Transformer, a fully connected layer, and a normalization layer; the second branch network includes a first CSP module, a connection layer, and a second CSP module.
7. The multimodal image matching method according to claim 1, wherein Performing feature matching on the optical image and the SAR image according to the feature description vector to obtain a matching point pair between the optical image and the SAR image includes: Performing feature matching on the optical image and the SAR image by using the nearest neighbor search method to obtain an initial matching point pair between the optical image and the SAR image; Performing purification on the initial matching point pair by using random sample consensus to obtain a target matching point pair between the optical image and the SAR image.
8. A multimodal image matching system, characterized in that, Including: A feature point extraction module, configured to perform self-supervised feature extraction on an optical image and a SAR image to obtain repeatable feature points between the optical image and the SAR image; A feature description module, configured to segment the optical image and the SAR image into a first sequence of image patches according to the repeatable feature points, and perform feature extraction on the first sequence of image patches through a dual-branch network to respectively obtain feature description vectors of the optical image and the SAR image; the dual-branch network includes a first branch network for extracting global features and a second branch network for extracting local features; A feature matching module, configured to perform feature matching on the optical image and the SAR image according to the feature description vector to obtain a matching point pair between the optical image and the SAR image.
9. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the multi-modal image matching method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the multi-modal image matching method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-source image registration method based on the combination of depth learning and artificially designed features
CN109064502A
Infrared-visible light image registration method and system
CN114092531A
Infrared and visible light image registration method and system based on hierarchical matching
CN114612698A
Multi-modal image matching method and system, terminal equipment and storage medium
CN117541833A
Learning keypoints and matching RGB images to CAD models
WO2020086217A1
Cited By
Land resource dynamic monitoring and early warning method and system based on multi-source remote sensing data fusion
CN121075102A
SAR directed target detection method based on multi-scale context sensing
CN121582554A