Multi-modal image matching method for scene matching navigation

By transforming images from the spatial domain to the frequency domain and back to the spatial domain, and combining wavelet response sets and structural tensor analysis, the matching problem caused by radiation and rotation distortion of heterogeneous images in scene matching navigation is solved, and high-precision multimodal image matching is achieved.

CN121392484APending Publication Date: 2026-01-23BEIJING ZHONGKE GUIDANCE & CONTROL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411894482.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing technologies are sensitive to radiation and rotation distortion in heterogeneous image matching, resulting in insufficient matching accuracy, especially in scene matching navigation where it is difficult to achieve all-weather navigation.

Method used

By converting the image from the spatial domain to the frequency domain for filtering, and then converting it back from the frequency domain to the spatial domain, wavelet response sets and maximum moments are obtained. The orientation information of feature points is obtained by combining structural tensor analysis. Feature matching is performed using Log-Gabor filter and SIFT method to filter the final results.

Benefits of technology

It achieves insensitivity to radiation distortion and rotation invariance, improving the accuracy and stability of multimodal image matching, and is suitable for scene matching navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121392484A_ABST
    Figure CN121392484A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal image matching method for scene matching navigation, which relates to the technical field of image matching, and comprises the following steps of: filtering a first modal image and a second modal image to obtain even symmetry and odd symmetry of each pixel point, further obtaining the maximum moment of each pixel point, and further obtaining boundary information; and further obtaining feature points, obtaining feature descriptors of the feature points, and performing feature matching to obtain a final matching result of the first modal image and the second modal image. The method is insensitive to radiation distortion, has rotation invariance, and solves the matching problem caused by non-linear radiation distortion of different-source images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image matching technology, and in particular to a multimodal image matching method for scene matching navigation. Background Technology

[0002] Heterogeneous image registration technology is widely used in remote sensing information fusion, military reconnaissance, and scene matching navigation. Image matching is a crucial step in the image registration process, and its accuracy directly affects the performance of the final image registration. Especially in scene matching navigation tasks, to achieve all-weather navigation, airborne imaging cameras typically use infrared cameras, while the reference ground... Figure One Generally, infrared images are visible light images. However, the imaging characteristics of visible light and infrared images differ significantly, leading to severe nonlinear radiometric distortion between them. Furthermore, the movement of the UAV causes rotational distortion between the infrared and visible light images. Traditional image matching algorithms, such as Scale-Invariant Feature Transform (SIFT) and SpeededUp Robust Features (SURF), use pixel intensity and gradients for feature detection, making them sensitive to radiometric distortion and completely unsuitable for multimodal image matching. Algorithms based on Radiation-variation Insensitive Feature Transform (RIFT) use phase congruency (PC) to construct feature descriptions, exhibiting strong robustness to nonlinear radiometric distortion. However, due to the loss of significant detail information during maximum index mapping, their robustness to rotational distortion is insufficient, particularly placing high demands on system heading constraints when applied to scene matching navigation. Summary of the Invention

[0003] The purpose of this invention is to provide a multimodal image matching method for scene matching navigation that is insensitive to radiation distortion, has rotation invariance, and solves the matching problem caused by nonlinear radiation distortion of heterogeneous images.

[0004] A multimodal image matching method for scene matching navigation, comprising:

[0005] The first modality image is converted from the spatial domain to the frequency domain and then filtered, and then converted from the frequency domain to the spatial domain to obtain the first spatial domain image. The second modality image is converted from the spatial domain to the frequency domain and then filtered, and then converted from the frequency domain to the spatial domain to obtain the second spatial domain image.

[0006] obtaining a first even-symmetric wavelet response set and a first odd-symmetric wavelet response set of each pixel point in the first spatial domain image, and obtaining a second even-symmetric wavelet response set and a second odd-symmetric wavelet response set of each pixel point in the second spatial domain image;

[0007] obtaining a maximum moment of each pixel point in the first spatial domain image based on the first even-symmetric wavelet response set and the first odd-symmetric wavelet response set, and obtaining a maximum moment of each pixel point in the second spatial domain image based on the second even-symmetric wavelet response set and the second odd-symmetric wavelet response set;

[0008] determining a first boundary based on the maximum moment of each pixel point in the first spatial domain image, taking the pixel point of the first modality image located in the first boundary as a first feature point, and obtaining a first feature point set;

[0009] determining a second boundary based on the maximum moment of each pixel point in the second spatial domain image, taking the pixel point of the second modality image located in the second boundary as a second feature point, and obtaining a second feature point set;

[0010] obtaining pixel coordinates and direction information of each first feature point and pixel coordinates and direction information of each second feature point based on structure tensor analysis; specifically:

[0011] obtaining the pixel coordinates of each first feature point;

[0012] obtaining a horizontal gradient image and a vertical gradient image of each neighborhood scale within a set neighborhood scale of each first feature point based on a Sobel gradient operator;

[0013] obtaining a horizontal gradient variance, a vertical gradient variance and a covariance of each first feature point in each scale within a set neighborhood scale based on the horizontal gradient image and the vertical gradient image of each neighborhood scale within the set neighborhood scale; the expression is:

[0014]

[0015] In the formula: is the horizontal gradient variance of the first feature point when the neighborhood scale is s, is the vertical gradient variance of the first feature point when the neighborhood scale is s, G x s y is the covariance of the first feature point when the neighborhood scale is s, s∈S, S is a set neighborhood scale of the first feature point, G xk is the gradient value of the kth pixel point of the horizontal gradient image of the first feature point when the neighborhood scale is s, G ykLet K be the gradient value of the k-th pixel of the vertical gradient image of the first feature point at a neighborhood scale of s, where K is the total number of pixels in the neighborhood scale of s centered on the first feature point.

[0016] Based on the horizontal gradient variance, vertical gradient variance, and covariance of the first feature point at each scale within a defined neighborhood, the fused horizontal gradient variance G of the first feature point is obtained. xx Fusion of vertical gradient variance G yy and fusion covariance G xy The expression is:

[0017]

[0018] In the formula: The weights are defined when the domain scale is s.

[0019] The orientation information of the first feature point is obtained by fusing the horizontal gradient variance, the vertical gradient variance, and the covariance. The expression is as follows:

[0020]

[0021] In the formula: Φ represents the direction information of the first feature point;

[0022] A first feature descriptor for each first feature point is obtained based on the pixel coordinates and orientation information of each first feature point, and a second feature descriptor for each second feature point is obtained based on the pixel coordinates and orientation information of each second feature point.

[0023] Feature matching is performed based on each of the first feature descriptors and each of the second feature descriptors to obtain the initial matching results of the first modal image and the second modal image. The initial matching results are then filtered to obtain the final matching results of the first modal image and the second modal image.

[0024] Optionally, filtering can be performed based on a Log-Gabor filter.

[0025] Optionally, the expression for the Log-Gabor filter is:

[0026]

[0027] In the formula: L(ρ,θ,h,q) represents the filter function, (ρ,θ) is the polar coordinate system of the filter, ρ is the radial coordinate of the filter, θ is the angular coordinate of the filter, h is the neighborhood scale of the filter, q is the direction of the filter, and ρ h Let θ be the center frequency of the filter. hq σ is the center direction of the filter. p σ is the radial bandwidth of the filter. θAn angular bandwidth of the filter.

[0028] Optionally, the first modality image is converted from spatial domain to frequency domain, filtered, and then converted from frequency domain to spatial domain to obtain the first spatial domain image, and the second modality image is converted from spatial domain to frequency domain, filtered, and then converted from frequency domain to spatial domain to obtain the second spatial domain image, comprising:

[0029] performing fast Fourier transform on the first modality image to convert from spatial domain to frequency domain to obtain the first frequency domain image, and performing fast Fourier transform on the second modality image to convert from spatial domain to frequency domain to obtain the second frequency domain image;

[0030] filtering the first frequency domain image to obtain the first filtered image, and filtering the second frequency domain image to obtain the second filtered image;

[0031] performing inverse fast Fourier transform on the first filtered image to convert from frequency domain to spatial domain to obtain the first spatial domain image, and performing inverse fast Fourier transform on the second filtered image to convert from frequency domain to spatial domain to obtain the second spatial domain image.

[0032] Optionally, the first even-symmetric wavelet response set and the first odd-symmetric wavelet response set of each pixel point in the first spatial domain image are obtained, and the second even-symmetric wavelet response set and the second odd-symmetric wavelet response set of each pixel point in the second spatial domain image are obtained, comprising:

[0033] obtaining the first even-symmetric wavelet set and the first odd-symmetric wavelet set of each pixel point in the first spatial domain image, and obtaining the second even-symmetric wavelet set and the second odd-symmetric wavelet set of each pixel point in the second spatial domain image;

[0034] obtaining the first even-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first even-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image, and obtaining the first odd-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first odd-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image;

[0035] obtaining the second even-symmetric wavelet response set of each pixel point in the second spatial domain image based on the second even-symmetric wavelet set and the coordinate value of each pixel point in the second spatial domain image, and obtaining the second odd-symmetric wavelet response set of each pixel point in the second spatial domain image based on the second odd-symmetric wavelet set and the coordinate value of each pixel point in the second spatial domain image.

[0036] Optionally, the first even-symmetric wavelet response set of each pixel in the first spatial domain image is obtained based on the first even-symmetric wavelet set and coordinate values ​​of each pixel; the first odd-symmetric wavelet response set of each pixel in the first spatial domain image is obtained based on the first odd-symmetric wavelet set and coordinate values ​​of each pixel, expressed as:

[0037]

[0038] E so (x,y,s,o) represents the first even-symmetric wavelet response of pixel (x,y) at a neighborhood scale of s and direction of o. even (x,y,s,o) represents the first even-symmetric wavelet of pixel (x,y) at neighborhood scale s and direction o, where o∈O is the number of directions of pixel (x,y). The first even-symmetric wavelet responses of pixel (x,y) at each neighborhood scale and direction constitute the first even-symmetric wavelet set, F. so (x,y,s,o) represents the first odd-symmetric wavelet response of pixel (x,y) at a neighborhood scale of s and direction of o. odd (x,y,s,o) represents the first odd-symmetric wavelet of pixel (x,y) at a neighborhood scale of s and a direction of o. The first odd-symmetric wavelet set is formed by the first odd-symmetric wavelet responses of pixel (x,y) at each neighborhood scale and each direction. I(x,y) represents the coordinates of pixel (x,y).

[0039] Optionally, obtaining the maximum moment of each pixel in the first spatial domain image based on each of the first even-symmetric wavelet responses and each of the first odd-symmetric wavelet responses includes:

[0040] The amplitude and phase components of each pixel in the first spatial domain image are obtained at each neighborhood scale and in each direction based on the first even-symmetric wavelet response and the first odd-symmetric wavelet response of each pixel in the first spatial domain image.

[0041] The expression for the amplitude component is:

[0042] The phase component expression is:

[0043] In the formula: Let A be the phase component of the pixel (x, y) at a neighborhood scale of s and direction of 0. so (x,y,s,o) represents the amplitude components of the pixel (x,y) at a neighborhood scale of s and a direction of o.

[0044] The phase consistency value of each pixel point in the first spatial domain image in each direction within a set neighborhood scale S is obtained based on the amplitude component and the phase component of each pixel point in the first spatial domain image in each neighborhood scale and each direction; the expression is:

[0045]

[0046] In the formula, PC(x, y, S, o) is the phase consistency value of the pixel point (x, y) in the o direction within the set neighborhood scale S, w o (x, y) is the weight of the pixel point (x, y) in the o direction, and ξ is a parameter value, which represents a non-negative value, T is a noise compensation term, which is a phase deviation function, which is the phase component of the pixel point (x, y) in the neighborhood scale s and the o direction minus the phase component of the pixel point (x, y) in the neighborhood scale s and the o direction;

[0047] The maximum moment of each pixel point in the first spatial domain image is obtained based on the PC function value of each pixel point in the first spatial domain image in each direction within the set neighborhood scale S; the expression is:

[0048]

[0049] In the formula, M ψ (x, y) is the maximum moment of the pixel point (x, y), a is a first intermediate variable, b is a second intermediate variable, and c is a third intermediate variable,

[0050] θ o is the angle value of the o direction.

[0051] Optionally, the first feature descriptor of each first feature point is obtained based on the pixel coordinates and the direction information of each first feature point, and the second feature descriptor of each second feature point is obtained based on the pixel coordinates and the direction information of each second feature point, specifically:

[0052] The pixel coordinates and the direction information of each first feature point are taken as the feature information of each first feature point, and the pixel coordinates and the direction information of each second feature point are taken as the feature information of each second feature point.

[0053] The first feature descriptor of each first feature point is obtained based on the feature information of each first feature point by using the SIFT method, and the second feature descriptor of each first feature point is obtained based on the feature information of each second feature point by using the SIFT method.

[0054] Optionally, an initial matching result of the first modality image and the second modality image is obtained based on a nearest neighbor point matching method.

[0055] Optionally, the initial matching result is filtered based on a RANSAC method to obtain a final matching result of the first modality image and the second modality image.

[0056] Effects of the present application are as follows:

[0057] The multi-modality image matching method for scene matching navigation of the present application performs phase detection based on phase consistency, is not sensitive to radiation distortion compared to intensity and gradient detection of traditional algorithms, has rotation invariance, and solves the matching problem caused by non-linear radiation distortion of heterogeneous images.

[0058] The multi-modality image matching method for scene matching navigation of the present application obtains direction information of feature points based on structure tensor analysis, constructs a three-dimensional feature point descriptor with phase information and direction information, has rotation invariance, and solves the heterogeneous image matching problem caused by rotation distortion. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 is a flowchart of the multi-modality image matching method for scene matching navigation of the present application;

[0060] Figure 2 is a matching result schematic diagram of the first modality image after being rotated by 30° with the second modality image;

[0061] Figure 3 is a matching result schematic diagram of the first modality image after being rotated by 60° with the second modality image;

[0062] Figure 4 is a matching result schematic diagram of the first modality image after being rotated by 90° with the second modality image;

[0063] Figure 5 is a matching result schematic diagram of the first modality image after being rotated by 120° with the second modality image;

[0064] Figure 6 is a matching result schematic diagram of the first modality image after being rotated by 150° with the second modality image;

[0065] Figure 7 is a matching result schematic diagram of the first modality image after being rotated by 180° with the second modality image. DETAILED DESCRIPTION

[0066] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings.

[0067] Figure 1is a flow chart of a multi-modal image matching method for scene matching navigation according to the present application. As shown in Figure 1 The present application provides a multi-modal image matching method for scene matching navigation, which comprises:

[0068] S1, filtering the first modal image after converting from spatial domain to frequency domain, and then converting from frequency domain to spatial domain to obtain a first spatial domain image, and filtering the second modal image after converting from spatial domain to frequency domain, and then converting from frequency domain to spatial domain to obtain a second spatial domain image.

[0069] Specifically, S1 comprises:

[0070] S11, performing fast Fourier transform on the first modal image to convert from spatial domain to frequency domain to obtain a first frequency domain image, and performing fast Fourier transform on the second modal image to convert from spatial domain to frequency domain to obtain a second frequency domain image.

[0071] S12, filtering the first frequency domain image to obtain a first filtered image, and filtering the second frequency domain image to obtain a second filtered image.

[0072] Preferably, the filtering is based on a Log-Gabor filter.

[0073] The expression of the Log-Gabor filter is:

[0074]

[0075] In the formula, L(ρ,θ,h,q) represents the filter function, (ρ,θ) is the polar coordinate system of the filter, ρ is the radial coordinate of the filter, θ is the angle coordinate of the filter, h is the neighborhood scale of the filter, q is the direction of the filter, ρ h is the center frequency of the filter, θ hq is the center direction of the filter, σ p is the radial bandwidth of the filter, σ θ is the angle bandwidth of the filter.

[0076] S13, performing fast inverse Fourier transform on the first filtered image to convert from frequency domain to spatial domain to obtain a first spatial domain image, and performing fast inverse Fourier transform on the second filtered image to convert from frequency domain to spatial domain to obtain a second spatial domain image.

[0077] S2, obtaining a first even-symmetry wavelet response set and a first odd-symmetry wavelet response set of each pixel point in the first spatial domain image, and obtaining a second even-symmetry wavelet response set and a second odd-symmetry wavelet response set of each pixel point in the second spatial domain image.

[0078] Further, S2 comprises:

[0079] S21, obtaining a first even-symmetric wavelet set and a first odd-symmetric wavelet set of each pixel point in the first spatial domain image; and obtaining a second even-symmetric wavelet set and a second odd-symmetric wavelet set of each pixel point in the second spatial domain image.

[0080] S22, obtaining a first even-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first even-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image; and obtaining a first odd-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first odd-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image.

[0081] The expression is:

[0082]

[0083] E so (x,y,s,o) is a first even-symmetric wavelet response of a pixel point (x,y) at a neighborhood scale s and a direction o, L even (x,y,s,o) is a first even-symmetric wavelet of a pixel point (x,y) at a neighborhood scale s and a direction o, o∈O is the number of directions of the pixel point (x,y), the first even-symmetric wavelet responses of the pixel point (x,y) at each neighborhood scale and each direction constitute a first even-symmetric wavelet set, F so (x,y,s,o) is a first odd-symmetric wavelet response of a pixel point (x,y) at a neighborhood scale s and a direction o, L odd (x,y,s,o) is a first odd-symmetric wavelet of a pixel point (x,y) at a neighborhood scale s and a direction o, the first odd-symmetric wavelet responses of the pixel point (x,y) at each neighborhood scale and each direction constitute a first odd-symmetric wavelet set, I(x,y) is a coordinate value of the pixel point (x,y).

[0084] S23, obtaining a second even-symmetric wavelet response set of each pixel point in the second spatial domain image based on the second even-symmetric wavelet set and the coordinate value of each pixel point in the second spatial domain image; and obtaining a second odd-symmetric wavelet response set of each pixel point in the second spatial domain image based on the second odd-symmetric wavelet set and the coordinate value of each pixel point in the second spatial domain image. The specific obtaining process is referred to S22.

[0085] S3, obtaining a maximum moment of each pixel point in the first spatial domain image based on each first even-symmetric wavelet response set and each first odd-symmetric wavelet response set; and obtaining a maximum moment of each pixel point in the second spatial domain image based on each second even-symmetric wavelet response set and each second odd-symmetric wavelet response set.

[0086] Preferably, S3 comprises:

[0087] S31, based on the first even-symmetric wavelet response and the first odd-symmetric wavelet response of each pixel in the first spatial domain image at each neighborhood scale and in each direction, obtain the amplitude component and phase component of each pixel in the first spatial domain image at each neighborhood scale and in each direction.

[0088] The expression for the amplitude component is:

[0089] The phase component expression is:

[0090] In the formula: Let A be the phase component of the pixel (x, y) at a neighborhood scale of s and direction of 0. so (x,y,s,o) represents the amplitude components of pixel (x,y) at a neighborhood scale of s and a direction of o.

[0091] S32, based on the amplitude and phase components of each pixel in the first spatial domain image at each neighborhood scale and in each direction, obtain the phase consistency value of each pixel in the first spatial domain image in each direction within a set neighborhood scale S. The expression is:

[0092]

[0093] In the formula: PC(x,y,S,o) is the phase consistency value of pixel (x,y) in the o direction within a set neighborhood scale S, w o (x, y) represents the weight of pixel (x, y) when the direction is 0, and ξ is the parameter value. This indicates that the value should be non-negative, and T is the noise compensation term. For phase deviation function, The phase component of pixel (x, y) at the neighborhood scale of s and direction of o is the same as the phase component of pixel (x, y) at the neighborhood scale of s and direction of o.

[0094] S33, the maximum moment of each pixel in the first spatial domain image is obtained based on the PC function values ​​of each pixel in each direction within a set neighborhood scale S. The expression is:

[0095]

[0096] Where: M ψ (x, y) represents the maximum moment of pixel (x, y), a is the first intermediate variable, b is the second intermediate variable, and c is the third intermediate variable.

[0097] θ o This is the angle value in the o direction.

[0098] S4, determining a first boundary based on the maximum moments of each pixel point in the first spatial domain image, taking the pixel points of the first modality image located within the first boundary as first feature points, and obtaining a first feature point set.

[0099] S5, determining a second boundary based on the maximum moments of each pixel point in the second spatial domain image, taking the pixel points of the second modality image located within the second boundary as second feature points, and obtaining a second feature point set.

[0100] S6, obtaining the pixel coordinates and direction information of each first feature point and the pixel coordinates and direction information of each second feature point based on structural tensor analysis.

[0101] Specifically, S6 includes:

[0102] S61, obtaining the pixel coordinates of each first feature point.

[0103] S62, obtaining the horizontal gradient image and the vertical gradient image of each neighborhood scale within the set neighborhood scale of each first feature point based on the Sobel gradient operator. Specifically, the image within a certain neighborhood scale centered on the first feature point is convolved with the Sobel gradient operator to obtain the horizontal gradient image and the vertical gradient image of the first feature point within the certain neighborhood scale.

[0104] S63, obtaining the horizontal gradient variance, the vertical gradient variance and the covariance of each scale within the set neighborhood scale of the first feature point based on the horizontal gradient image and the vertical gradient image of each neighborhood scale within the set neighborhood scale. The expression is:

[0105]

[0106] In the formula: G is the horizontal gradient variance of the first feature point when the neighborhood scale is s, G is the vertical gradient variance of the first feature point when the neighborhood scale is s, G is the covariance of the first feature point when the neighborhood scale is s, s∈S, S is the set neighborhood scale of the first feature point, G xk G is the gradient value of the kth pixel point of the horizontal gradient image of the first feature point when the neighborhood scale is s, yk G is the gradient value of the kth pixel point of the vertical gradient image of the first feature point when the neighborhood scale is s, K is the total number of pixel points within the neighborhood scale s centered on the first feature point.

[0107] S64, obtaining the fusion horizontal gradient variance G xx , the fusion vertical gradient variance G yy and the fusion covariance G xyThe expression is:

[0108]

[0109] In the formula: The weights are for a domain scale of s.

[0110] S65, the orientation information of the first feature point is obtained based on the fusion of horizontal gradient variance, fusion of vertical gradient variance, and fusion of covariance. The expression is:

[0111]

[0112] In the formula: Φ represents the direction information of the first feature point.

[0113] S7, obtain the first feature descriptor of each first feature point based on the pixel coordinates and orientation information of each first feature point, and obtain the second feature descriptor of each second feature point based on the pixel coordinates and orientation information of each second feature point.

[0114] Specifically, the pixel coordinates and orientation information of each first feature point are used as the feature information of each first feature point, and the pixel coordinates and orientation information of each second feature point are used as the feature information of each second feature point.

[0115] Based on the feature information of each first feature point, the SIFT method is used to obtain the first feature descriptor of each first feature point; based on the feature information of each second feature point, the SIFT method is used to obtain the second feature descriptor of each first feature point.

[0116] After obtaining the position coordinates and orientation information of the feature points, the descriptor is obtained as follows:

[0117] A 36x36 window is created around the feature point, which is divided into 36 sub-blocks of size 6x6. For each sub-block, 6 bin orientation histograms are created, and the 6x6x6 orientation gives 216 bin values. These 216 bin values ​​are represented as feature vectors to form feature point descriptors.

[0118] The orientation histograms of 6 bins cover 360 degrees. Assuming that the orientation of a certain pixel is 40 degrees, it will enter the bin of 0 to 60 degrees. Each sub-block has 6 bins, and there are a total of 6*6 sub-blocks. Therefore, the feature descriptor is 6*6*6.

[0119] S8. Based on each first feature descriptor and each second feature descriptor, feature matching is performed to obtain the initial matching results of the first modal image and the second modal image. The initial matching results are then filtered to obtain the final matching results of the first modal image and the second modal image.

[0120] Preferably, the initial matching result of the first modality image and the second modality image is obtained based on a nearest neighbor point matching method.

[0121] The initial matching result is filtered based on a RANSAC method to obtain the final matching result of the first modality image and the second modality image.

[0122] Specifically, the first modality image is infrared image, and the second modality image is visible light image, the first modality image is matched with the second modality image after being rotated by 30°, 60°, 90°, 120°, 150° and 180° respectively, the 30° matching result is as shown in FIG. 3, the 60° matching result is as shown in FIG. 4, the 90° matching result is as shown in FIG. 5, the 120° matching result is as shown in FIG. 6, the 150° matching result is as shown in FIG. 7, the 180° matching result is as shown in FIG. 8. Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7

[0123] It is found through the above experiment that the infrared image and the visible light image of different modalities can be normally matched at various angles.

[0124] The above-described embodiments are only used to describe the preferred embodiments of the present application, and do not limit the scope of the present application, and various modifications and improvements to the technical solutions of the present application made by those skilled in the art without departing from the design spirit of the present application shall fall within the protection scope of the present application defined by the claims.​​​​​​

Claims

1. A multi-modal image matching method for scene matching navigation, characterized in that, It comprises: filtering the first modality image after converting from spatial domain to frequency domain, and then converting from frequency domain to spatial domain to obtain a first spatial domain image, and filtering the second modality image after converting from spatial domain to frequency domain, and then converting from frequency domain to spatial domain to obtain a second spatial domain image; obtaining a first even-symmetric wavelet response set and a first odd-symmetric wavelet response set of each pixel point in the first spatial domain image, and obtaining a second even-symmetric wavelet response set and a second odd-symmetric wavelet response set of each pixel point in the second spatial domain image; obtaining the maximum moment of each pixel point in the first spatial domain image based on each first even-symmetric wavelet response set and each first odd-symmetric wavelet response set; obtaining the maximum moment of each pixel point in the second spatial domain image based on each second even-symmetric wavelet response set and each second odd-symmetric wavelet response set; determining a first boundary based on the maximum moment of each pixel point in the first spatial domain image, taking the pixel point of the first modality image located in the first boundary as a first feature point to obtain a first feature point set; determining a second boundary based on the maximum moment of each pixel point in the second spatial domain image, taking the pixel point of the second modality image located in the second boundary as a second feature point to obtain a second feature point set; obtaining the pixel coordinates and direction information of each first feature point and the pixel coordinates and direction information of each second feature point based on structure tensor analysis; specifically: obtaining the pixel coordinates of each first feature point; obtaining the horizontal gradient image and the vertical gradient image of each neighborhood scale within a set neighborhood scale of each first feature point based on a Sobel gradient operator; obtaining the horizontal gradient variance, the vertical gradient variance and the covariance of each scale within the set neighborhood scale of the first feature point based on the horizontal gradient image and the vertical gradient image of each neighborhood scale within the set neighborhood scale; the expression is: wherein: is the horizontal gradient variance of the first feature point at the field scale s, is the vertical gradient variance of the first feature point at the field scale s, is the covariance of the first feature point at the field scale s, s e S, S is a set neighborhood scale of the first feature point, G xk is the gradient value of the kth pixel point of the horizontal gradient image of the first feature point at the field scale s, yk is the gradient value of the kth pixel point of the vertical gradient image of the first feature point at the field scale s, k is the total number of pixel points within the field scale s with the first feature point as the center. Based on horizontal gradient variance, vertical gradient variance and covariance of each scale within the set neighborhood scale of the first feature point, a fusion horizontal gradient variance G of the first feature point is obtained xx , a fusion vertical gradient variance G yy and a fusion covariance G xy of the first feature point are obtained; the expression is as follows: wherein: is the weight for a field of size s; obtaining the direction information of the first feature point based on the fusion horizontal gradient variance, the fusion vertical gradient variance and the fusion covariance; the expression is: In the formula, Φ is the direction information of the first feature point; obtaining the first feature descriptor of each first feature point based on the pixel coordinates and direction information of each first feature point, and obtaining the second feature descriptor of each second feature point based on the pixel coordinates and direction information of each second feature point; performing feature matching based on each first feature descriptor and each second feature descriptor to obtain an initial matching result of the first modality image and the second modality image, and performing screening on the initial matching result to obtain a final matching result of the first modality image and the second modality image.

2. The multi-modal image matching method for scene matching navigation according to claim 1, characterized in that, Filtering based on a Log-Gabor filter.

3. The multi-modal image matching method for scene matching navigation according to claim 2, c h a r a c t e r i z e d b y, The expression of the Log-Gabor filter is: where L(p, q, h, q) represents the filter function, (p, q) is the polar coordinate system of the filter, p is the radial coordinate of the filter, q is the angular coordinate of the filter, h is the neighborhood scale of the filter, q is the direction of the filter, p h is the center frequency of the filter, q hq is the center direction of the filter, p p is the radial bandwidth of the filter, q θ is the angular bandwidth of the filter.

4. The multi-modal image matching method for scene matching navigation according to claim 1, characterized in that, The filtering of the first modality image after converting from spatial domain to frequency domain, and then converting from frequency domain to spatial domain to obtain a first spatial domain image, and filtering the second modality image after converting from spatial domain to frequency domain, and then converting from frequency domain to spatial domain to obtain a second spatial domain image, comprises: performing a fast Fourier transform on the first modality image to transform from a spatial domain to a frequency domain to obtain a first frequency domain image, and performing a fast Fourier transform on the second modality image to transform from the spatial domain to the frequency domain to obtain a second frequency domain image; filtering the first frequency domain image to obtain a first filtered image, and filtering the second frequency domain image to obtain a second filtered image; performing an inverse fast Fourier transform on the first filtered image to transform from the frequency domain to the spatial domain to obtain a first spatial domain image, and performing an inverse fast Fourier transform on the second filtered image to transform from the frequency domain to the spatial domain to obtain a second spatial domain image.

5. The multi-modal image matching method for scene matching navigation according to claim 1, wherein, The method for obtaining the first even-symmetric wavelet response set and the first odd-symmetric wavelet response set of each pixel point in the first spatial domain image, and obtaining the second even-symmetric wavelet response set and the second odd-symmetric wavelet response set of each pixel point in the second spatial domain image comprises: obtaining the first even-symmetric wavelet set and the first odd-symmetric wavelet set of each pixel point in the first spatial domain image, and obtaining the second even-symmetric wavelet set and the second odd-symmetric wavelet set of each pixel point in the second spatial domain image; obtaining the first even-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first even-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image, and obtaining the first odd-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first odd-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image; obtaining the second even-symmetric wavelet response set of each pixel point in the second spatial domain image based on the second even-symmetric wavelet set and the coordinate value of each pixel point in the second spatial domain image, and obtaining the second odd-symmetric wavelet response set of each pixel point in the second spatial domain image based on the second odd-symmetric wavelet set and the coordinate value of each pixel point in the second spatial domain image.

6. The multi-modal image matching method for scene matching navigation according to claim 5, c h a r a c t e r i z e d b y, obtaining the first even-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first even-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image, and obtaining the first odd-symmetric wavelet response set of each pixel point in the first spatial domain image based on the first odd-symmetric wavelet set and the coordinate value of each pixel point in the first spatial domain image, and the expression is: E so (x,y,s,o) is the first even-symmetric wavelet response of pixel point (x, y) at neighborhood scale s and direction o, L even (x,y,s,o) is the first even-symmetric wavelet of pixel point (x, y) at neighborhood scale s and direction o, o∈O is the number of directions of pixel point (x, y), the first even-symmetric wavelet responses of pixel point (x, y) at each neighborhood scale and each direction constitute a first even-symmetric wavelet set, F so (x,y,s,o) is the first odd-symmetric wavelet response of pixel point (x, y) at neighborhood scale s and direction o, L odd (x,y,s,o) is the first odd-symmetric wavelet of pixel point (x, y) at neighborhood scale s and direction o, the first odd-symmetric wavelet responses of pixel point (x, y) at each neighborhood scale and each direction constitute a first odd-symmetric wavelet set, I(x,y) is the coordinate value of pixel point (x, y).

7. The multi-modal image matching method for scene matching navigation according to claim 6, c h a r a c t e r i z e d b y, The method for obtaining the maximum moment of each pixel point in the first spatial domain image based on the first even-symmetric wavelet response and the first odd-symmetric wavelet response comprises: obtaining the amplitude component and the phase component of each pixel point in the first spatial domain image in each neighborhood scale and each direction based on the first even-symmetric wavelet response and the first odd-symmetric wavelet response of each pixel point in the first spatial domain image in each neighborhood scale and each direction; The amplitude component expression is: The phase component expression is: wherein: A(x,y,s,o) is the amplitude component of the pixel point (x,y) at the neighborhood scale s and the direction o, so (x,y,s,o) is the amplitude component of the pixel point (x,y) at the neighborhood scale s and the direction o; obtaining the phase consistency value of each pixel point in the first spatial domain image in each direction within a set neighborhood scale S based on the amplitude component and the phase component of each pixel point in the first spatial domain image in each direction within the set neighborhood scale S, and the expression is: wherein: PC(x, y, S, o) is the phase consistency value of the pixel point (x, y) in the direction o in the set neighborhood scale S, w o (x, y) is the weight of the pixel point (x, y) in the direction o, and ξ is a parameter value, represents taking a non-negative value, T is a noise compensation term, is a phase deviation function, is the phase component of the pixel point (x, y) in the neighborhood scale s and the direction o minus the phase component of the pixel point (x, y) in the neighborhood scale s and the direction o. obtaining the maximum moment of each pixel point in the first spatial domain image based on the PC function value of each pixel point in the first spatial domain image in each direction within the set neighborhood scale S, and the expression is: In the formula, M ψ (x, y) is the maximum moment of the pixel point (x, y), a is a first intermediate variable, b is a second intermediate variable, and c is a third intermediate variable, θ o is an angle value of the o direction.

8. The multi-modal image matching method for scene matching navigation according to claim 1, wherein, The pixel coordinates and the direction information of each first feature point are used to obtain a first feature descriptor of each first feature point, and the pixel coordinates and the direction information of each second feature point are used to obtain a second feature descriptor of each second feature point, specifically as follows: The pixel coordinates and the direction information of each first feature point are used as the feature information of each first feature point, and the pixel coordinates and the direction information of each second feature point are used as the feature information of each second feature point. The SIFT method is used to obtain the first feature descriptor of each first feature point based on the feature information of each first feature point, and the SIFT method is used to obtain the second feature descriptor of each first feature point based on the feature information of each second feature point.

9. The multi-modal image matching method for scene matching navigation according to claim 1, wherein, An initial matching result of the first modality image and the second modality image is obtained based on a nearest neighbor point matching method.

10. The multi-modal image matching method for scene matching navigation according to claim 1, characterized in that, The initial matching result is filtered based on a RANSAC method to obtain a final matching result of the first modality image and the second modality image.