A multimodal image matching method of geometric and radiometric invariance

CN122530627BActive Publication Date: 2026-09-22SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611031458.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-22
Estimated Expiration
2046-07-13

AI Technical Summary

Technical Problem

[0011]本发明旨在解决现有多模态图像匹配方法在非线性辐射差异、大范围旋转和显著尺度变化并存时,关键点邻域重叠不足、公共结构表征不稳定、局部感受野受限以及训练特征空间缺少结构约束的问题,提供一种具有几何和辐射不变性的多模态图像匹配方法

Benefits of technology

[0036](1)几何不变性增强:并行利用笛卡尔空间与对数极坐标空间,避免单一采样方式在空间布局保持和旋转、尺度鲁棒性之间的取舍。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530627B_ABST
    Figure CN122530627B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal image matching method with geometric and radiation invariance, and belongs to the technical field of image processing. The method prepares cross-modal positive and negative samples based on known geometric correspondence; detects key points and determines the key point neighborhood, performs Cartesian sampling and logarithmic polar coordinate sampling on the same neighborhood; performs intra-modal, sampling branch and cross-modal feature interaction through an RRSI feature description network; introduces a bidirectional cross-modal structure decoder in the training stage to reconstruct the descriptor under the loss constraint; removes the decoder in the inference stage, and outputs the matching result through the descriptor neighbor search and robust geometric estimation. The method improves the matching stability under nonlinear radiation difference, rotation and scale change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically to a multimodal image matching method with geometric and radiation invariance. Background Technology

[0002] Multimodal image matching is used to determine the corresponding positions between images obtained from different sensors, imaging mechanisms, times, or observation conditions. It is fundamental to tasks such as image registration, image fusion, change detection, target localization, 3D reconstruction, and assisted diagnosis. Typical multimodal images include visible light images, infrared images, synthetic aperture radar images, lidar images, maps, nighttime images, and various medical imaging sequences.

[0003] Different modalities typically exhibit both geometric transformations and nonlinear radiative differences. Geometric transformations include translation, rotation, scale changes, affine transformations, and local deformations; nonlinear radiative differences arise from variations in sensor response, imaging band, illumination, weather, noise, and scattering mechanisms, causing the same ground feature or tissue to exhibit different grayscale, texture, and edge responses in different modalities. Consequently, descriptive methods relying on raw intensity or single gradient statistics struggle to establish stable and consistent features across two modalities.

[0004] Existing feature matching methods typically employ a processing chain of keypoint detection, neighborhood description, and descriptor matching. To achieve rotation invariance, some methods require estimating the principal orientation or traversing multiple candidate angles; to achieve scale invariance, some methods require constructing multi-scale candidates and performing multiple matching operations. This type of processing increases computational cost when the rotation range is large or the scale differences are significant, and angle and scale estimation errors propagate to subsequent descriptions and matching. When using only Cartesian neighborhoods, the effective overlapping area of ​​image patches at different scales decreases rapidly; when using only logarithmic polar coordinate neighborhoods, the original local spatial layout may be lost.

[0005] Deep feature description methods can learn cross-modal common structures through training, but when convolutionally encoding only a single local image patch, the receptive field is limited, making it difficult to utilize complementary information from other key regions in the same image and from another modality. Directly performing large-scale feature interactions may introduce modality-specific responses and background interference into the descriptor. On the other hand, if training relies solely on distance constraints between matched and unmatched samples, deep features may exhibit feature drift or over-smoothing under strong nonlinear radiative differences, failing to explicitly preserve reversible spatial structures.

[0006] Therefore, a technical solution is needed for the practical multimodal image matching process. This solution should preserve the neighborhood spatial structure of key points while absorbing the equivariant properties of log-polar coordinates with rotation and scale changes. It should organize information interaction within a unified feature space, between sampling branches, and between modes, and further constrain descriptors through structural reconstruction during the training phase without increasing the burden on the actual matching inference phase.

[0007] Regarding training data, randomly selected non-corresponding regions that are far apart are usually easy to distinguish and cannot fully force the model to learn precise localization capabilities. Local negative samples that are highly similar to the anchored region but have a small displacement at their center are more valuable for training. Meanwhile, rotation and scale changes can be artificially generated through known geometric transformations, while the nonlinear radiation differences between real sensors are difficult to accurately simulate using only brightness or contrast transformations. Therefore, it is necessary to combine real-registered cross-modal samples, local displacement sampling, and controllable geometric augmentation.

[0008] For keypoint detection, the response of detection directly on the original grayscale image is easily affected by mode inversion, local saturation, speckle noise, and thermal radiation diffusion. By covering different structure sizes with a multi-scale pyramid and detecting corner points on the gradient magnitude map, more boundary locations and structural transitions can be utilized. Although the detection stage cannot guarantee that two images have completely identical point sets, it can provide a sufficient number of relatively stable candidate regions for description and matching.

[0009] For local region sampling, the Cartesian coordinate system directly represents the spatial topology, but the overlap area of ​​two support regions decreases quadratically under large scale differences. The log-polar coordinate system can map angles and logarithmic radii to translation coordinates, making it more compatible with rotation and scale changes, but it alters the original positional relationships. If two types of image patches are simply concatenated at the input, the network may favor the branch with stronger texture, failing to form a clear complementarity. This invention preserves the independent encoding paths of the two types of image patches and performs controlled interactions in the deep feature space.

[0010] In terms of cross-modal feature learning, sharing all convolutional weights may force modalities with significant differences to use the same low-level filter, while not sharing any high-level interactions makes it difficult to align common structures. This invention adopts a pseudo-twin structure with low-level modality-specific and deep-level unified interactions: the convolutional coding parameters of the two modalities are not shared to adapt to their respective imaging statistics; intra-modal double-sampled features and inter-modal fusion features enter the unified space through an attention mechanism. Summary of the Invention

[0011] This invention aims to address the problems of insufficient overlap of key point neighborhoods, unstable representation of common structures, limited local receptive fields, and lack of structural constraints in the training feature space when existing multimodal image matching methods suffer from nonlinear radiation differences, large-scale rotations, and significant scale changes. It provides a multimodal image matching method with geometric and radiation invariance.

[0012] The technical solution adopted in this invention is as follows:

[0013] A geometrically and radiation-invariant multimodal image matching method includes the following steps:

[0014] S1. Preparing Training Data: Acquire a first modal image and a second modal image with a known geometric correspondence. Map sampling points in the first modal image to the second modal image according to the geometric correspondence to obtain corresponding points. Crop anchored image blocks centered on the sampling points, crop cross-modal corresponding image blocks centered on the corresponding points, and crop locally displaced image blocks after offsetting within the local neighborhood of the sampling points or corresponding points. Define image blocks whose center points satisfy the geometric correspondence as corresponding regions, and define image blocks whose center points do not satisfy the geometric correspondence as non-corresponding regions. Perform data augmentation on the obtained image blocks.

[0015] S2. Construct a multimodal image matching network: The multimodal image matching network includes a key point detection module, a dual-head region sampling module, an RRSI feature description network, and a descriptor matching module;

[0016] The key point detection module establishes scale pyramids for the first modal image and the second modal image respectively. For each scale layer image, it calculates the sum of the absolute values ​​of the partial derivatives in two orthogonal directions to obtain a gradient magnitude map. Key points are detected on each gradient magnitude map by a corner detector. Key points are aggregated along the scale dimension, scale information is recorded, and key points are mapped to the original image coordinate system. The neighborhood of each key point is determined according to its position and scale.

[0017] The dual-head region sampling module takes the same key point as the center and the neighborhood of the key point as the sampling support region, and performs Cartesian sampling and log-polar coordinate sampling respectively to obtain Cartesian image blocks and log-polar coordinate image blocks corresponding to the same key point.

[0018] The RRSI feature description network performs convolutional encoding, intra-modal region interaction, inter-sampling branch interaction, and cross-modal interaction on two types of sampled image blocks to obtain RRSI descriptors;

[0019] The descriptor matching module establishes matching relationships based on the obtained RRSI descriptors;

[0020] S3. Training a multimodal image matching network: Construct triplet training samples based on the corresponding and non-corresponding regions. The triplet training samples include anchored regions, cross-modal corresponding regions, and non-corresponding regions. Add two cross-modal structure decoders with opposite directions to the output of the RRSI feature description network. Jointly optimize the network parameters using triplet marginal loss, descriptor similarity regularization loss, and bidirectional cross-modal generation and reconstruction loss. Remove the two cross-modal structure decoders after training.

[0021] S4. Use the trained multimodal image matching network for matching: detect key points in the multimodal image to be matched and determine the neighborhood of the key points. Perform dual-head region sampling on the neighborhood of the same key point and extract RRSI descriptors. Establish candidate matching relationships based on the descriptor distance. Eliminate erroneous matches through spatial consistency screening and robust geometric estimation. Output matching point pairs and / or geometric transformation models.

[0022] Furthermore, in step S1, the geometric correspondence includes at least one of homography transformation, affine transformation, dense displacement field, sensor imaging model, geolocation relationship, or manual annotation; the first sampling point in the first modal image is mapped to the corresponding point in the second modal image according to the geometric correspondence;

[0023] An anchored image block is obtained by cropping to a set size with the first sampling point as the center; a cross-modal corresponding image block is obtained by cropping to the same size with the corresponding point as the center, and the anchored image block and the cross-modal corresponding image block constitute a corresponding region and serve as positive samples; a second sampling point is obtained by randomly shifting along the horizontal and / or vertical directions in the local neighborhood of the first sampling point, and a local displacement image block is obtained by cropping with the second sampling point as the center, and the local displacement image block or image blocks with different numbers constitute a non-corresponding region and serve as negative samples.

[0024] Furthermore, the data augmentation method described in step S1 is as follows: for at least one of the anchored image patch, the cross-modal corresponding image patch, and the local displacement image patch, at a horizontal azimuth angle... to Rotation enhancement is applied within the range, and optionally at the pitch angle. to Perspective enhancement is performed within the range; image patches are scaled within a preset scale range and cropped or resampled around the center; radiometric difference samples are obtained based on real registered multimodal image pairs, and at least one of the following enhancements is performed: brightness, contrast, gamma correction, piecewise linear grayscale transformation, blur, impulse noise, Gaussian noise, speckle noise, shearing, and flipping.

[0025] Furthermore, the dual-head region sampling module uses the same key point as the sampling center and the neighborhood of the key point determined by its position and scale as the sampling support region: Cartesian coordinate sampling is used to obtain Cartesian image blocks that maintain the spatial layout of the neighborhood; logarithmic polar coordinate sampling is used to obtain logarithmic polar coordinate image blocks, such that tangential rotation around the key point corresponds to displacement in the first coordinate axis direction of the logarithmic polar coordinate image block, and radial scale change around the key point corresponds to displacement in the second coordinate axis direction; the Cartesian image blocks and logarithmic polar coordinate image blocks are mutually independent ordered image block pairs corresponding to the same key point.

[0026] Furthermore, the logarithmic polar coordinate sampling is based on , Determine the sampling location of the original image, where, For key point locations, For the maximum sampling radius, To output the coordinates in the image patch, and These represent the width and height of the output image patch, respectively.

[0027] Furthermore, the RRSI feature description network includes two modality coding paths corresponding to the first modality and the second modality respectively, and whose weights are not shared; each modality coding path includes a Cartesian image block convolutional coding branch, a log-polar coordinate image block convolutional coding branch, a self-attention layer acting on each sampling branch, a first cross-attention layer acting on the two sampling branches, and a multilayer perceptron fusion layer; the fused features of the two modality coding paths are cross-modal interacted through the second cross-attention layer to output the first modality RRSI descriptor and the second modality RRSI descriptor;

[0028] The self-attention layer allows different key point regions of the same modality and the same sampling branch to exchange context; the first cross-attention layer allows Cartesian sampling features and log-polar coordinate sampling features to interact with each other; the second cross-attention layer allows the first modality fusion feature and the second modality fusion feature to interact with each other, so as to preserve cross-modal common structural responses and suppress modality-specific responses.

[0029] Furthermore, in step S3, the triplet training samples include anchor descriptors, positive sample descriptors, and negative sample descriptors; wherein, the anchor descriptor is obtained from the RRSI descriptor corresponding to the anchored image block, the positive sample descriptor is obtained from the RRSI descriptor corresponding to the cross-modal corresponding image block, and the negative sample descriptor is obtained from the RRSI descriptor corresponding to the local displacement image block or the image block with different indexes; for each anchor descriptor, the descriptor with the smallest mixing distance is selected from all negative sample descriptors as the most difficult negative sample, and the mixing distance is determined by the angle relationship between the descriptors and the Euclidean geometric relationship of the descriptors.

[0030] Furthermore, in step S3, the first cross-modal structure decoder takes the first modal RRSI descriptor and the source-side multi-scale features of the first modal coding path as input to generate a second modal Cartesian structure image block corresponding to the cross-modal image block; the second cross-modal structure decoder takes the second modal RRSI descriptor and the source-side multi-scale features of the second modal coding path as input to generate a first modal Cartesian structure image block corresponding to the anchored image block; neither of the two cross-modal structure decoders reads the intermediate features of the target modal coding path;

[0031] The total loss used satisfies , ;in, For the marginal loss of the triple, To describe the sub-similarity regularization loss, To generate pixel-level reconstruction loss between structural image patches and their corresponding target modality true Cartesian image patches, To determine the feature-level perceptual consistency loss between the features of the generated structural image patch re-encoded by the target modality convolutional encoder and the features of the real Cartesian image patch before attention processing. , and As weight;

[0032] calculate At that time, a stopping gradient processing is performed on the encoded features of the real Cartesian image patch used as the supervision target, freezing the parameters of the target modal convolutional encoder used to re-encode the generated structured image patch, while retaining the gradient relative to the generated structured image patch, so that the error signal is passed to the cross-modal structure decoder and the RRSI feature description path.

[0033] Further, step S4 includes: normalizing the extracted descriptors; establishing initial correspondences using nearest neighbor search; filtering initial correspondences using at least one of bidirectional nearest neighbor, nearest neighbor distance ratio, and local displacement consistency; mapping keypoints at different scales back to the original image coordinate system; estimating homography, affine transformation, or local geometric model using a random sampling consistency algorithm or a degenerate sample consistency algorithm and deleting correspondences that do not satisfy the reprojection constraint.

[0034] Furthermore, the scale pyramid comprises five scale layers, each scaled relative to the original image by a scaling factor of [missing information]. , , , and Both the Cartesian image patch and the log-polar image patch were resampled to... Pixels; Convolutional coding branch outputs 256-dimensional features; Training rotation enhances coverage. to The training scale enhancement range is 0.75 to 1.5, the original sampling region side length is 96 to 192 pixels and resampled to 128 pixels.

[0035] The beneficial effects of this invention are as follows:

[0036] (1) Enhanced geometric invariance: Parallel use of Cartesian space and logarithmic polar coordinate space avoids the trade-offs between spatial layout preservation and rotation and scale robustness in a single sampling method.

[0037] (2) Multi-layer information interaction: By interacting within a modality, fusing between sampling branches and aligning across modalities, the receptive field limitation of a single key point neighborhood convolution is overcome and modality-specific redundant responses are suppressed.

[0038] (3) Radiation difference robustness: bidirectional cross-modal generation and reconstruction enables descriptors to not only satisfy distance discrimination, but also carry information that can restore the modal structure layout of the other party, thereby anchoring the modality-invariant geometric topology under strong radiation difference.

[0039] (4) Controllable inference overhead: The cross-modal structure decoder is only used during the training phase and is removed after training is completed, without increasing the number of parameters and computational cost in the actual image matching inference phase.

[0040] (5) Wide range of applications: It can be used for visible light-infrared, visible light-SAR, day and night, image-map, remote sensing, medical and other image pairs with obvious imaging differences. Attached Figure Description

[0041] Figure 1 This is a schematic diagram of the overall framework of RRSI multimodal image matching in this invention, as well as the training and inference paths.

[0042] Figure 2 This is a schematic diagram illustrating the combination of spatial sampling and data augmentation in the training data preparation process of this invention.

[0043] Figure 3 This is a schematic diagram of the key point detection process of the integrated scale pyramid, gradient magnitude map and FAST detector of the present invention.

[0044] Figure 4 This is a schematic diagram illustrating the Cartesian sampling and logarithmic polar coordinate sampling performed by the present invention at different sampling scales.

[0045] Figure 5 This is a schematic diagram of the structure of the self-attention module and the cross-attention module of the present invention.

[0046] Figure 6 This is a schematic diagram of the bidirectional cross-modal generation and reconstruction constraints and source-side jump connection during the training phase of this invention. Detailed Implementation

[0047] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0048] The technical solution of this invention aims at co-modeling with geometric invariance and radiation invariance, and uses dual-head region sampling and cross-modal generation and reconstruction constraints during the training phase as core technical means, including the following steps S1 to S4. These steps together form an implementable training and application process. Figure 1 The overall network framework, the auxiliary branches generated and reconstructed during training, and the paths retained during inference are illustrated. Specifically, they include:

[0049] S1, Prepare training data. Obtain first modal images covering the same scene or object. Second mode image And obtain the homography matrix, affine matrix, dense displacement field, sensor imaging model, geographic location relationship or manual annotation correspondence between the two. From the first modal sampling point Obtain the corresponding point of the second mode This forms cross-modal positive samples with a clearly defined center.

[0050] S1-1, Perform spatial sampling. For example... Figure 2 As shown, an anchored image block is cropped with the sampling point as the center; then, a horizontal and / or vertical offset is applied to the center point within the local neighborhood to crop a local displacement image block with the same size as the anchored image block. The anchored image block and the corresponding modal image block form positive samples, and the anchored image block and the local displacement image block form difficult negative samples with similar content but misaligned centers; image blocks with different corresponding indices constitute global negative samples.

[0051] S1-2, Perform data augmentation. Apply viewpoint rotation enhancement, scale resolution enhancement, and radiometric distortion enhancement to the anchored image patch and the corresponding cross-modal image patch. Viewpoint rotation enhancement is performed in... to The azimuth angle can be changed within the range, and it can be used to change the azimuth angle within the range. to Within a set range, the pitch angle is varied to simulate perspective differences; scale enhancement changes the sampling area within a set range and resamples to a fixed size; radiometric enhancement first uses real registered image pairs to provide cross-modal nonlinear radiometric differences, and then superimposes brightness, contrast, gamma curves, piecewise linear grayscale transformation, blur, impulse noise, Gaussian noise, or speckle noise. Geometric enhancement synchronously updates the point correspondence.

[0052] S2, Construct a multimodal image matching network. The network includes a keypoint detection module, a dual-head region sampling module, an RRSI feature description network, and a descriptor matching module; during the training phase, a first cross-modal structure decoder is also connected to the output of the RRSI feature description network. Second transmodal structure decoder .

[0053] S2-1, Construct the key point detection module. For example... Figure 3 As shown, a multi-scale pyramid is constructed for the input image; the sum of the absolute values ​​of the horizontal and vertical partial derivatives is calculated at each scale level to obtain a gradient magnitude map; FAST corner detection and non-maximum suppression are performed on each gradient magnitude map; keypoints are merged along the scale dimension, and the original image coordinates, scale, and response value of the keypoints are recorded. This processing utilizes structural changes rather than absolute intensity to detect local salient locations.

[0054] S2-2, Construct a dual-head region sampling module. For example... Figure 4 As shown, Cartesian sampling is performed independently on the same key point. And logarithmic polar coordinate sampling Obtain the Cartesian image block And logarithmic polar coordinate image patches The former preserves the intuitive spatial layout of edges, corners, and textures; the latter transforms rotations and scale changes around key points into displacements along mutually orthogonal coordinate axes, sampling with denser samples at the center and sparser samples at the periphery. The two types of image patches form ordered pairs rather than being simply stitched together at the input.

[0055] S2-3, Construct the RRSI feature description network. For the first and second modalities, set pseudo-Twin convolutional encoding paths with non-shared weights. Each path encodes a Cartesian image patch and a log-polar image patch, respectively. The initial features of each branch first pass through a self-attention module for information exchange between regions within the same modality, and then through a first cross-attention module for information exchange between double-sampled branches. The two branches are then summed after passing through a multilayer perceptron to form intra-modal fused features. The fused features of the first and second modalities then pass through a second cross-attention module for cross-modal information exchange, outputting the RRSI descriptor. The self-attention and cross-attention structures are as follows: Figure 5 As shown.

[0056] S2-4, Construct a cross-modal structure decoder during training. For example... Figure 6 As shown, Only accept the first mode RRSI descriptor Source-side multi-scale features of the first modal encoder To generate second-modal Cartesian structure image blocks; Only accept the second-mode RRSI descriptor Source-side multi-scale features of the second-mode encoder This generates a first-modal Cartesian structure image patch. Neither decoder accesses intermediate features of the target modality coding path; the decoder can be composed of transposed convolution, instance normalization, ReLU activation, and terminal Tanh layers.

[0057] S3, train the multimodal image matching network. Each training batch starts from... Medium sampling Each distinct point is mapped to ,form For positive samples and For negative samples; select the most difficult negative sample for each anchor point based on the mixing distance, and calculate the triplet marginal loss. And utilize descriptor similarity regularization loss. Constrain the magnitude relationship of the corresponding descriptors.

[0058] S3-1, Calculate the bidirectional cross-modal generation and reconstruction loss. Pixel-level reconstruction loss Constraining the spatial layout differences between the generated structural image patches and the target modality's true Cartesian image patches; feature-level perceptual consistency loss. The generated image patch is re-fed into the corresponding target modality Cartesian convolutional encoder and compared with the encoded features of the real image patch before the attention layer. The real encoded features are used to stop the gradient; during re-encoding, the convolutional encoder parameters are frozen, but the gradients relative to the generated image patch are preserved. The total loss is... .

[0059] S3-2, backpropagate the total loss and update the trainable parameters of the convolutional coding path, attention module, fusion layer, and cross-modal architecture decoder. Remove the [processor / module] when the validation metrics meet the requirements or after a set number of training epochs. and And branches used only for reconstruction loss calculation, retaining the parameters required for keypoint detection, dual-head region sampling, RRSI description, and descriptor matching.

[0060] S4. Multimodal image matching is performed using the trained network. Keypoints are detected on the images to be matched and mapped to the original image coordinate system. Cartesian and log-polar coordinate image patches are generated for each keypoint, and RRSI descriptors are extracted and normalized. Initial correspondences are obtained using nearest neighbor search. The first purification is performed by at least one of bidirectional nearest neighbor, distance ratio, and local displacement consistency. Then, homography, affine, or local geometric models are estimated using RANSAC or DEGENSAC. Correspondences that do not meet the reprojection constraints are deleted, and the matching point pairs, matching confidence, and / or geometric transformation models are output.

[0061] Example:

[0062] Select the first modality image for the same scene Second mode image In this example, the first mode is a visible light image, and the second mode is an infrared image or a SAR image. The image pair has a known homography matrix. For scenarios that cannot be fully described by a single homography, multiple homography, affine models, dense flow, or sensor models can be used in local areas to provide a correspondence.

[0063] exist Randomly selected from Valid sampling points ,according to calculate .by Crop the first modality anchored image patch to the center, The second modality cross-modal corresponding image block is cropped around the center; anchored image blocks with the same index and cross-modal corresponding image blocks constitute positive sample correspondences, and image blocks with different indexes constitute candidate negative samples. Sampling points that fall outside the effective image range or have incomplete neighborhoods after mapping are discarded.

[0064] like Figure 2 As shown, training data preparation includes a spatial sampling module and a data augmentation module. The spatial sampling module includes at least an anchor sampler, a cross-modal correspondence sampler, and a local displacement sampler. The anchor sampler uses... Center-cropping anchored image patches; cross-modal corresponding samplers with Cropping cross-modal corresponding image patches to the center; local displacement sampler in Surrounding choices As the new center, crop image blocks of the same size. The value is not zero and is limited to a range that makes the contents of two image patches highly similar but whose center points should not be determined to be the same. This results in local negative samples that are more difficult to distinguish than random, distant negative samples.

[0065] Figure 2 The relationship between the anchor box and the displacement box is also shown. The first center represents the anchor point, and the second center represents the local displacement point. The lower limit of the offset can be related to the correct matching pixel threshold, and the upper limit of the offset can be related to the non-maximum suppression radius of the keypoint or the size of the local support region. Global negative samples can also be obtained from other locations in the image for the same anchor point, so that the batch simultaneously contains positive samples, locally difficult negative samples, and global negative samples.

[0066] The data augmentation module applies three types of enhancement to the spatially sampled image patches. The first type is rotation or viewpoint enhancement: at the azimuth angle... to Rotate randomly within the range, and if necessary, at the pitch angle. to The first category is perspective aberration caused by simulated terrain undulations or viewing angles within a certain range. The second category is scale resolution enhancement: changing the side length of the sampling area within a preset range, and then restoring the fixed input size by center clipping or interpolation. The third category is radiometric enhancement: superimposing gamma transform, piecewise linear contrast stretching, brightness variation, blurring, impulse noise, Gaussian noise, or speckle noise on the inherent radiometric differences of the corresponding blocks in the real multimodal model.

[0067] Geometric enhancement can be updated synchronously based on artificially generated transformations. Therefore, rotation and scaling supervision can be obtained through self-supervision; cross-modal radiometric correspondences depend on real registered image pairs or known cross-sensor mappings, thus radiometric consistency supervision is provided by real correspondences. Combining these two types of supervision in the same training batch allows descriptors to simultaneously learn geometric equivalence relations and cross-modal common structures.

[0068] In one implementation, the network input size for training image patches is... Pixels. Using 128 pixels as a reference side length, random samples are taken from the original image. Pixel to The region of pixels is then dynamically resampled. The pixel size corresponds to a scale enhancement range of 0.75 to 1.5. The image patch size can also be set to other values ​​based on ground resolution, target size, and video memory capacity.

[0069] When augmenting training images, geometric augmentation updates the correspondences simultaneously to ensure that the correct matching position can still be determined after augmentation. During the training phase, random points can be used instead of detection keypoints to obtain uniform and stable supervision samples; during the inference phase, the output of the keypoint detection module is used.

[0070] like Figure 3 As shown, a five-level scale pyramid is constructed for the input image, and its relative scale can be set to... , , , and For the first Layer Image Calculate the partial derivatives in the horizontal and vertical directions respectively, and obtain the gradient magnitude diagram. :

[0071] ,

[0072] At various scales Candidate points are detected using the FAST corner detector, and non-maximum suppression is performed. The coordinates, scale, and response value of each keypoint are recorded, and its coordinates are converted to the original image coordinate system. Then, cross-scale duplicate points are removed based on the response value and spatial distance. The number of keypoints in a single image can be limited to 5000, but this value can be adjusted according to the image size and computing resources.

[0073] Gradient magnitude maps emphasize local structural transitions while downplaying absolute grayscale. For optical and infrared images, the same building boundary may exhibit opposite brightness relationships; similarly, for optical and SAR images, the texture statistics of the same road or plot edge may differ significantly. By summing the absolute values ​​of partial derivatives, these boundaries can still generate a strong response, thereby increasing the probability of key points recurring across different modes.

[0074] After candidate points at different scales are mapped to the original image, they can be clustered by spatial radius, retaining only the candidate points with the largest response, or multiple scale hypotheses can be retained for subsequent description. When retaining multiple scale hypotheses, each descriptor corresponds one-to-one with its keypoint coordinates and scale, and after matching, they are uniformly mapped to the original coordinate system.

[0075] like Figure 4 As shown, for key points The keypoint neighborhood is determined by the keypoint's location and scale; the dual-head region sampling module uses the keypoint as the center and its neighborhood as the sampling support region, while simultaneously performing Cartesian sampling. And logarithmic polar coordinate sampling , respectively form and Both can be resampled to Pixels, and all correspond to the same key point.

[0076] Cartesian sampling clips a local region centered on keypoints and scales and rotates the local coordinates according to the keypoint scale and optional orientation. Its purpose is to preserve the intuitive positional relationships of edges, corners, and textures within the local neighborhood.

[0077] Logarithmic polar coordinate sampling uses key points as poles, based on the output coordinates. Determine the sampling coordinates (x, y) in the original image:

[0078] ,

[0079] ,

[0080] in, For the maximum sampling radius, and This outputs the width and height of the image patch. The transformation makes the tangential rotation around the keypoints primarily manifest as… Directional displacement causes radial scale changes to manifest primarily as Directional displacement; simultaneously, the sampling density gradually decreases from the center outwards. Thus, each keypoint yields an ordered pair of image patches. , ).

[0081] Cartesian sampling employs an approximately uniform spatial grid. When the sampling range is reduced to 0.5 times its original size, the effective overlap area of ​​two Cartesian image patches at different scales can be reduced to approximately 25%, and continues to decrease as the scale difference further increases. Log-polar sampling allocates a higher sampling density to the vicinity of the keypoint center, allowing more central structure to be preserved at different sampling radii, while transforming rotation and scale changes into translational forms that are easier for convolution and attention modules to handle.

[0082] The two types of sampling are not interchangeable. Using only Cartesian sampling is sensitive to large rotations and scale changes; using only logarithmic polar coordinate sampling weakens the intuitive local positional relationships. This invention encodes the two types of image patches independently first, and then interacts with them through cross-attention, so that one branch provides spatial localization capabilities, and the other branch provides rotation and scale resistance capabilities.

[0083] like Figure 1 and Figure 5 As shown, the set of Cartesian image patches for the first mode. Set of logarithmic polar coordinate image patches Perform convolutional encoding separately to obtain initial features. and The second modality is obtained using a convolutional encoding path with non-weight sharing. and The non-shared weights allow each modal encoder to adapt to its own imaging statistics, while the subsequent unified feature space is responsible for aligning the common structure.

[0084] Within each mode, and Passing through Layer self-attention enables different keypoint regions within the same sampling branch to exchange context; subsequently, both undergo... Layered cross-attention allows the Cartesian spatial structure to complement the logarithmic polar coordinate rotation and scale-related information. The outputs of the two branches are processed separately by a multilayer perceptron and then summed to obtain the intramodal fusion feature. and .

[0085] and Further process Cross-modal cross-attention. This process enables the first modal features to selectively absorb responses from the second modality that correspond to their structure, and performs symmetric operations on the second modality, ultimately obtaining the RRSI descriptor subset. and In one implementation, , and All Each attention module has The size is 256, and the feature dimension is 256. The above values ​​are for illustrative purposes only and do not constitute the sole limitation on the network structure.

[0086] Specifically, the convolutional encoder can employ a VGG-type network. For each The pixel image blocks are progressively convolved and downsampled to output 256-dimensional local features, and the source-side multi-scale features are preserved in several resolution layers. For the same modality, the Cartesian branch and the log-polar coordinate branch can use encoders with the same structure; for different modalities, the encoder parameters are not shared, so as to absorb the intensity, texture and noise statistics of each modality respectively.

[0087] The self-attention module includes layer normalization, multi-head self-attention, residual connections, multilayer perceptrons, and re-residual connections. The input sequence consists of features from multiple keypoint regions in a batch or a single image, enabling a keypoint region to reference the structural responses of other distant regions in the same modality, thus mitigating the problem of insufficient receptive field for a single local image patch.

[0088] Optionally, category labels can be set in the input sequence of the self-attention module or the cross-attention module to aggregate the overall representation of the key point region. Before generating the query, key and value, the input features can be divided, concatenated or linearly mapped. The processing is used to adapt to multi-head attention computation without changing the one-to-one correspondence between the key point region and the RRSI descriptor.

[0089] The first cross-attention module generates queries using Cartesian branch features and keys and values ​​using log-polar coordinate branch features, and can perform reverse symmetric interactions. The interaction results are fused with the residuals of the original branch features. Since the two branches correspond one-to-one with the same key point, this interaction can inject rotation, scale, and other variable information from the log-polar coordinate branch into the Cartesian branch, while simultaneously injecting the explicit spatial layout of the Cartesian branch into the log-polar coordinate branch.

[0090] The second cross-attention module establishes query, key, and value relationships between the first and second modal fusion features. During training, the known corresponding points enable the network to learn cross-modal common structures. During inference, even if the original intensity relationship between the two modalities is non-linear, the attention weights can still select relevant regions based on the deep structural response and suppress modality-specific noise. Ultimately, each keypoint receives a fixed-dimensional RRSI descriptor, maintaining a one-to-one correspondence between keypoints, dual-head image patches, and descriptors.

[0091] In one implementation, the number of self-attention layers within a branch Number of dual-sampling branch cross-attention layers and the number of cross-modal attention layers All are set to 2, and each attention layer is set to... The number of heads and the number of embedded dimensions are 256. The number of layers, heads, and dimensions can also be increased or decreased based on the amount of data and computing resources.

[0092] For the first in the batch A corresponding point, an anchored image patch descriptor As anchoring descriptors, descriptors corresponding to image patches across modalities As a positive sample descriptor; from of Or select with in the local displacement image block descriptor The sample with the smallest distance is used as the hard-negative sample descriptor. A hybrid distance combining cosine and Euclidean geometric terms is used. .make To calculate the cosine similarity between two descriptors, we can use:

[0093] ,

[0094] The marginal loss of the ternary set is:

[0095] ,

[0096] in, Let j*(i) be the marginal parameter, representing the index of the hardest-to-bear sample. The descriptor similarity regularization loss can be expressed as:

[0097]

[0098] like Figure 6 As shown, during training, the first structure decoder by Source-side multi-scale features generated by the first modality convolutional encoder As input, generate second-modal Cartesian structure image patches. Second structure decoder by Source-side multi-scale features As input, generate a first-modal Cartesian structure image patch. Each decoder can be composed of a cascade of transposed convolutions, instance normalization, and nonlinear activation layers, and gradually recovers spatial details through source-side skip connections.

[0099] The pixel-level reconstruction loss is:

[0100] ,

[0101] make and These are two Cartesian convolutional encoders with different modalities. The initial encoded features of the real image patch are: and The feature-level perceptual consistency loss is:

[0102] ,

[0103] Where sg represents the stopping gradient. Calculation Freeze and The parameters are set, but gradients are allowed to propagate with respect to the input generated image patches. Thus, the real encoded features remain stably supervised, and the generated image patches can still transmit error signals back to the decoder and the RRSI description path.

[0104] ,

[0105] ,

[0106] In this example, it is set to , , , Batch Points After training and optimization are complete, delete. and And branches used only for calculating reconstruction loss, storing parameters of the two-modal encoder, attention layer, and fusion layer.

[0107] The set of positive samples can be represented as The negative sample set can be represented as For each ,from Select the one with the smallest mixing distance As the most difficult negative sample; it can also be used symmetrically. The most difficult negative sample is selected from the first modality descriptor as the anchor point. This batch-based mining method avoids using only easily distinguishable, distant negative samples.

[0108] Similarity regularization loss It can be calculated based on the squared difference in the magnitude of the corresponding descriptor. It, along with the triplet marginal loss, respectively constrains the magnitude relationship of the descriptor and the distance boundary between positive and negative samples. The combination of the two helps to mitigate the feature scale drift caused by changes in brightness or radiation.

[0109] like Figure 6 As shown, generative reconstruction does not require a complete translation of one modality into another, but rather projects the common structure in the descriptor into Cartesian structure image patches of the other modality. Therefore, under strong radiation differences... It is not required that the convergence to zero be achieved; its weights... This is used to control the impact of the auxiliary task on the main matching task. Source-side multi-scale skip connections are used to recover spatial details step by step, but it is strictly forbidden to bypass the intermediate features of the target modality encoder and input them into the decoder, so as to avoid the decoder bypassing the descriptor and directly copying the target structure.

[0110] A phased strategy can be adopted during training: first, with and The network is pre-trained and then jointly trained with two structural decoders; alternatively, all three types of losses can be optimized simultaneously from scratch. If a staged strategy is adopted, the learning rate of the convolutional encoder can be reduced after adding the decoders, allowing the generated reconstruction to primarily refine the unified feature space without excessively compromising existing discriminativity. During the validation phase, only the inference path after removing the decoders is run to ensure that the selected model parameters correspond to actual deployment performance.

[0111] like Figure 1 As shown, the first and second modal images to be matched are input into the trained network. Multi-scale gradient keypoint detection is performed separately, generating Cartesian image patches and log-polar image patches for each keypoint, and obtaining descriptors through the RRSI feature description network.

[0112] Perform on descriptor Normalization is performed, and candidate matches are found using FLANN, KD-tree, graph-based nearest neighbor search, or other nearest neighbor search algorithms. Nearest neighbor ratio, bidirectional nearest neighbor, or confidence thresholds can be used to filter candidate relationships. After mapping the coordinates in the scale pyramid back to the original image, RANSAC, DEGENSAC, or other robust estimation methods are used to fit homography, affine, or local geometric models, and matching point pairs that do not satisfy reprojection consistency are deleted.

[0113] For visible-infrared and visible-SAR images, the output matching points can be used for image registration and fusion; for medical multimodal images, the output matching points can be used for spatial alignment between different sequences; for image-map or day / night images, the matching results can be used for localization and change analysis.

[0114] The initial refinement of candidate matches can be performed in three levels. The first level is descriptor-level filtering: retaining bidirectional nearest neighbors, or retaining correspondences where the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than a threshold. The second level is spatial-level filtering: calculating the local displacement, neighborhood ranking, or adjacency relationship of candidate matches in the two images, and deleting correspondences that are significantly inconsistent with neighboring candidates. The third level is model-level filtering: using RANSAC or DEGENSAC to repeatedly sample the minimum point set to fit a geometric model, using candidates with reprojection errors less than a threshold as inliers.

[0115] When the scene mainly consists of a plane or distant view, global homography can be fitted; when the viewpoint changes are small and there is no obvious perspective, an affine model can be fitted; when there are terrain undulations or local non-rigid deformations, block homography, local affine, thin plate splines, or dense deformation fields can be used. The geometric model type can be pre-specified according to the application scenario, or it can be selected from multiple candidate models based on the number of interior points and reprojection error.

[0116] The matching output can include not only coordinate pairs, but also the distance between each descriptor pair, bidirectional consistency flag, local consistency score, robust estimate interior point flag, and final confidence score. The downstream registration module can then perform matching based on the geometric transformation model. Resampling to Coordinate systems are used for image fusion, change detection, target localization, or multimodal joint analysis.

[0117] Other alternative implementation methods:

[0118] The keypoint detector can be replaced with other corner, edge, or learned detectors, as long as they can output keypoints with location and optional scale information.

[0119] The convolutional encoder can be replaced with a residual network, a lightweight convolutional network, or other locally structured encoding networks; the attention structure can be replaced with functionally equivalent self-attention and cross-attention modules.

[0120] The target for generating the reconstruction can include Cartesian structured image patches, edge maps, gradient maps, or other structural representations that preserve spatial topology; pixel loss can be achieved using... Charbonnier or structural similarity loss can be used, and feature loss can be achieved using a frozen modal encoder or structural encoder.

[0121] When the image has local non-rigid deformation, robust geometric estimation can be extended from global homography to block homography, thin plate spline or dense deformation field estimation.

[0122] The method of the present invention can be implemented by software, hardware or a combination of software and hardware, and each module can be deployed on a server, workstation, edge device or sensor processing platform.

Claims

1. A multimodal image matching method with geometric and radiation invariance, characterized in that, The steps include the following: S1. Prepare training data: Obtain a first modal image and a second modal image with a known geometric correspondence, and map the sampling points in the first modal image to the second modal image according to the geometric correspondence to obtain the corresponding points; An anchored image block is cropped with the sampling point as the center, and a cross-modal corresponding image block is cropped with the corresponding point as the center. After offsetting within the local neighborhood of the sampling point or the corresponding point, a locally displaced image block is cropped. Image blocks whose center points satisfy the geometric correspondence are defined as corresponding regions, and image blocks whose center points do not satisfy the geometric correspondence are defined as non-corresponding regions; data augmentation is then performed on the obtained image blocks; S2. Construct a multimodal image matching network: The multimodal image matching network includes a key point detection module, a dual-head region sampling module, an RRSI feature description network, and a descriptor matching module; The key point detection module establishes scale pyramids for the first modal image and the second modal image respectively. For each scale layer image, it calculates the sum of the absolute values ​​of the partial derivatives in two orthogonal directions to obtain a gradient magnitude map. Key points are detected on each gradient magnitude map by a corner detector. Key points are aggregated along the scale dimension, scale information is recorded, and key points are mapped to the original image coordinate system. The neighborhood of each key point is determined according to its position and scale. The dual-head region sampling module takes the same key point as the center and the neighborhood of the key point as the sampling support region, and performs Cartesian sampling and log-polar coordinate sampling respectively to obtain Cartesian image blocks and log-polar coordinate image blocks corresponding to the same key point. The RRSI feature description network performs convolutional encoding, intra-modal region interaction, inter-sampling branch interaction, and cross-modal interaction on two types of sampled image blocks to obtain RRSI descriptors; The descriptor matching module establishes matching relationships based on the obtained RRSI descriptors; S3. Training a multimodal image matching network: Construct triplet training samples based on the corresponding and non-corresponding regions. The triplet training samples include anchored regions, cross-modal corresponding regions, and non-corresponding regions. Add two cross-modal structure decoders with opposite directions to the output of the RRSI feature description network. Jointly optimize the network parameters using triplet marginal loss, descriptor similarity regularization loss, and bidirectional cross-modal generation and reconstruction loss. Remove the two cross-modal structure decoders after training. S4. Use the trained multimodal image matching network for matching: detect key points in the multimodal image to be matched and determine the neighborhood of the key points. Perform dual-head region sampling on the neighborhood of the same key point and extract RRSI descriptors. Establish candidate matching relationships based on the descriptor distance. Eliminate erroneous matches through spatial consistency screening and robust geometric estimation. Output matching point pairs and / or geometric transformation models.

2. The method according to claim 1, characterized in that, In step S1, the geometric correspondence includes at least one of homography transformation, affine transformation, dense displacement field, sensor imaging model, geographic location relationship, or manual annotation; the first sampling point in the first modal image is mapped to the corresponding point in the second modal image according to the geometric correspondence. An anchored image block is obtained by cropping to a set size with the first sampling point as the center; a cross-modal corresponding image block is obtained by cropping to the same size with the corresponding point as the center, and the anchored image block and the cross-modal corresponding image block constitute a corresponding region and serve as positive samples; a second sampling point is obtained by randomly shifting along the horizontal and / or vertical directions in the local neighborhood of the first sampling point, and a local displacement image block is obtained by cropping with the second sampling point as the center, and the local displacement image block or image blocks with different numbers constitute a non-corresponding region and serve as negative samples.

3. The method according to claim 2, characterized in that, The data augmentation method described in step S1 is as follows: for at least one of the anchored image patch, the cross-modal corresponding image patch, and the local displacement image patch, at the horizontal azimuth angle... to Rotation enhancement is applied within the range, and optionally at the pitch angle. to Render perspective enhancement within the area; Scaling image patches within a preset scale range and cropping or resampling around the center; Radiometric difference samples are obtained based on real registered multimodal image pairs, and at least one of the following enhancements is applied: brightness, contrast, gamma correction, piecewise linear grayscale transformation, blurring, impulse noise, Gaussian noise, speckle noise, shearing, and flipping.

4. The method according to claim 1, characterized in that, In step S2, the dual-head region sampling module uses the same key point as the sampling center and the neighborhood of the key point determined by the position and scale of the key point as the sampling support region: Cartesian coordinate sampling is used to obtain Cartesian image blocks that maintain the spatial layout of the neighborhood; logarithmic polar coordinate sampling is used to obtain logarithmic polar coordinate image blocks, such that the tangential rotation around the key point corresponds to the displacement in the first coordinate axis direction of the logarithmic polar coordinate image block, and the radial scale change around the key point corresponds to the displacement in the second coordinate axis direction; the Cartesian image blocks and the logarithmic polar coordinate image blocks are mutually independent ordered image block pairs corresponding to the same key point.

5. The method according to claim 4, characterized in that, The logarithmic polar coordinate sampling is based on , Determine the sampling location of the original image, where, For key point locations, For the maximum sampling radius, To output the coordinates in the image patch, and These represent the width and height of the output image patch, respectively.

6. The method according to claim 1, characterized in that, In step S2, the RRSI feature description network includes two modality coding paths corresponding to the first modality and the second modality respectively and whose weights are not shared; each modality coding path includes a Cartesian image block convolutional coding branch, a log-polar coordinate image block convolutional coding branch, a self-attention layer acting on each sampling branch, a first cross-attention layer acting on the two sampling branches, and a multilayer perceptron fusion layer; the fused features of the two modality coding paths are cross-modal interacted through the second cross-attention layer to output the first modality RRSI descriptor and the second modality RRSI descriptor; The self-attention layer enables different key point regions of the same modality and the same sampling branch to exchange context; the first cross-attention layer enables Cartesian sampling features and log-polar coordinate sampling features to interact with each other. The second cross-attention layer enables the first modality fusion feature and the second modality fusion feature to interact with each other, so as to preserve the cross-modal common structural response and suppress the modality-specific response.

7. The method according to claim 1, characterized in that, In step S3, the triplet training samples include anchor descriptors, positive sample descriptors, and negative sample descriptors; wherein, the anchor descriptor is obtained from the RRSI descriptor corresponding to the anchored image patch, the positive sample descriptor is obtained from the RRSI descriptor corresponding to the cross-modal corresponding image patch, and the negative sample descriptor is obtained from the RRSI descriptor corresponding to the local displacement image patch or the image patch with different indexes; for each anchor descriptor, the descriptor with the smallest mixing distance is selected from all negative sample descriptors as the most difficult negative sample, and the mixing distance is determined by the angle relationship between the descriptors and the Euclidean geometry relationship of the descriptors.

8. The method according to claim 1, characterized in that, In step S3, the first cross-modal structure decoder takes the first modal RRSI descriptor and the source-side multi-scale features of the first modal coding path as input to generate a second modal Cartesian structure image block corresponding to the cross-modal image block. The second cross-modal structure decoder takes the second modal RRSI descriptor and the source-side multi-scale features of the second modal coding path as input to generate a first modal Cartesian structure image block corresponding to the anchored image block; neither of the two cross-modal structure decoders reads the intermediate features of the target modal coding path; The total loss used satisfies , ;in, For the marginal loss of the triple, To describe the sub-similarity regularization loss, To generate pixel-level reconstruction loss between structural image patches and their corresponding target modality true Cartesian image patches, To determine the feature-level perceptual consistency loss between the features of the generated structural image patch re-encoded by the target modality convolutional encoder and the features of the real Cartesian image patch before attention processing. , and As weight; calculate At that time, a stopping gradient processing is performed on the encoded features of the real Cartesian image patch used as the supervision target, freezing the parameters of the target modal convolutional encoder used to re-encode the generated structured image patch, while retaining the gradient relative to the generated structured image patch, so that the error signal is passed to the cross-modal structure decoder and the RRSI feature description path.

9. The method according to claim 1, characterized in that, Step S4 includes: normalizing the extracted descriptors; establishing initial correspondences using nearest neighbor search; filtering initial correspondences using at least one of bidirectional nearest neighbor, nearest neighbor distance ratio, and local displacement consistency; mapping keypoints at different scales back to the original image coordinate system; estimating homography, affine transformation, or local geometric model using a random sampling consistency algorithm or a degenerate sample consistency algorithm, and deleting correspondences that do not satisfy the reprojection constraints.

10. The method according to any one of claims 1 to 9, characterized in that, The scale pyramid comprises five scale layers, each scaled relative to the original image at a different ratio. , , , and Both the Cartesian image patch and the log-polar image patch were resampled to... Pixels; convolutional coding branch outputs 256-dimensional features; Training rotation to enhance coverage to The training scale enhancement range is 0.75 to 1.5, the original sampling region side length is 96 to 192 pixels and resampled to 128 pixels.

Citation Information

Patent Citations

  • Multi-modal remote sensing image matching method based on modal reconstruction and feature disturbance learning

    CN122135051A

  • Neural network-based pose estimation and registration method and device for heterogeneous images, and medium

    US20240169584A1