A three-dimensional phase orientation invariant feature transform multi-modal image matching method and system

By employing a three-dimensional phase orientation invariant feature transformation method, combined with differential phase feature extraction and a three-dimensional phase consistency orientation histogram, the geometric and radiometric differences in multimodal remote sensing image matching are solved, achieving high-precision image alignment and collaborative processing.

CN120182640BActive Publication Date: 2025-11-18WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510210448.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-11-18
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Multimodal remote sensing image matching faces challenges such as geometric and radiometric differences between modes, missing texture information, noise interference, and nonlinear radiometric distortion, which increase the difficulty of matching. Existing methods are not effective in dealing with these problems.

Method used

A three-dimensional phase orientation invariant feature transformation method is adopted. By extracting differential phase features and constructing a three-dimensional phase consistency orientation histogram, a multi-directional differential map and a multi-scale three-dimensional phase consistency optimal index map are constructed to enhance the matching stability and accuracy.

Benefits of technology

It effectively reduces spectral differences between multimodal images, improves feature point extraction capabilities, and enhances matching accuracy and robustness. It is particularly suitable for situations with large modal differences, severe noise interference, or missing texture information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182640B_ABST
    Figure CN120182640B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional phase orientation invariant feature transformation multi-modal image matching method and system, and belongs to the field of image processing. Firstly, a multi-directional difference image is constructed through a difference filter, then a three-dimensional phase tensor feature is calculated by using a three-dimensional phase consistency model, and a phase difference image is calculated. Then, a maximum direction value of the phase difference image is calculated to obtain a significant difference phase feature. Finally, a feature point is obtained in a difference phase moment space. Then, optimal feature indexing is performed and the feature point position is located under the three-dimensional phase feature of multi-scale and multi-direction, a direction gradient histogram of a local region of the feature point is calculated and counted, and a three-dimensional phase orientation feature descriptor is constructed. Finally, a Euclidean distance is used as a matching measure, a same-named point is obtained by using a ratio of a nearest neighbor distance and a second nearest neighbor distance of the three-dimensional phase orientation feature, mis-matching points are removed by using a random sample consensus algorithm, and multi-modal image matching is completed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of image processing, and particularly relates to a multi-modal image matching method and system of three-dimensional phase orientation invariant feature transformation. BACKGROUND

[0002] Remote sensing image matching plays a crucial role in the field of computer vision and photogrammetry, as it serves as the foundation for key technologies such as image fusion, change detection, positioning and navigation, and three-dimensional reconstruction, providing strong support for various remote sensing applications. With the increasing demand for remote sensing applications, a single data source has been unable to meet the processing application needs of complex geographic scenes, and integrating images from different data sources has gradually become a new research hotspot. In recent years, image acquisition technology has developed rapidly, and new sensors such as visible light cameras, infrared cameras, and synthetic aperture radar (SAR) have emerged, making the acquisition of multi-modal remote sensing images increasingly diverse, covering various surface information and environmental changes. However, with the increase in data modalities, the significant differences between different modalities further increase the difficulty of matching. Therefore, how to effectively eliminate the differences between modalities and improve the precision and robustness of image alignment has become the key to the collaborative processing of multi-modal images.

[0003] The core task of multi-modal image matching is to overcome the matching problems caused by differences in sensor types, shooting angles, lighting conditions, etc. between different modalities, in order to achieve high-precision image alignment. The specific challenges of multi-modal image matching include: 1. Geometric differences, imaging angles, sensor characteristics, and changes in ground objects may lead to differences in image geometry and spatial position; 2. Spectral differences, optical and radar images of different modalities have significant differences in spectral response, and the same ground object may exhibit different characteristics under different modalities; 3. Noise and nonlinear radiation distortion can cause loss of image details and color distortion, increasing interference in the feature extraction process. In view of these difficulties, experts and scholars have proposed a series of solutions. For example, the scale-invariant feature transform (SIFT) algorithm proposed by Lowe et al. effectively alleviates the geometric difference problem by extracting scale and rotation invariant features. However, this method still has limitations in dealing with radiation differences, lighting differences, and nonlinear radiation distortion. Therefore, researchers have begun to explore phase features in the frequency domain of images, and through methods such as Log-Gabor filter optimized matching (LGHD), extended phase correlation algorithm (LGEPC), and phase consistency direction histogram (HOPC), they have successfully improved geometric and radiation differences. However, in the face of complex situations such as lack of texture information, noise interference, or nonlinear radiation distortion, these methods may still fail to match. Deep learning methods can effectively deal with nonlinear matching problems due to their ability to learn high-level features, but due to the lack of multi-modal data samples, their matching generalization has certain limitations. SUMMARY

[0004] To solve the problems in the background art, the present application proposes a multi-modal image matching method based on three-dimensional phase orientation invariant feature transformation, aiming to effectively deal with geometric and radiometric differences. This method combines three-dimensional phase difference maps and three-dimensional phase consistency orientation histograms. First, the present application innovatively proposes a differential phase feature extraction technique, which constructs multi-directional difference maps through a differential filter and calculates three-dimensional phase features by combining a three-dimensional phase consistency model. The three-dimensional phase difference map can effectively reduce the spectral differences between multi-modal images, thereby significantly improving the feature point extraction capability. Second, the present application designs a three-dimensional phase consistency orientation histogram, which further enhances the matching stability by constructing a multi-scale, multi-directional three-dimensional phase consistency optimal index map. In particular, when dealing with cases where the modal difference is too large, such as depth maps or electronic maps, the noise interference is severe, and the texture information is missing, the three-dimensional phase consistency orientation histogram can effectively capture the local and global geometric characteristics of the image, ensure the invariance of the features, and improve the matching accuracy.

[0005] Through the innovative method, the multi-modal image matching method based on three-dimensional phase orientation invariant feature transformation effectively overcomes the geometric and radiometric differences in multi-modal image matching, and solves the challenges brought by texture detail loss, noise interference, and nonlinear radiometric distortion, providing strong technical support for the accurate alignment and collaborative processing of multi-modal images.

[0006] The present application proposes a multi-modal image matching method based on three-dimensional phase orientation invariant feature transformation to solve the matching problem of multi-modal images.

[0007] The technical solution adopted by the present application is: a multi-modal image matching method based on three-dimensional phase orientation invariant feature transformation, comprising the following steps:

[0008] Step 1, initialize the calculation parameters of multi-modal image matching, and perform nonlinear diffusion on the multi-modal images;

[0009] Step 2, calculate the multi-channel differential gradient feature map using the nonlinear diffusion result;

[0010] Step 3, perform image domain calculation on the multi-channel differential gradient feature map, and estimate the image phase feature using a three-dimensional phase consistency model, thereby generating a three-dimensional phase feature;

[0011] Step 4, perform Gaussian differential filtering using the three-dimensional phase feature, and count the maximum feature values in each channel to generate a saliency differential phase feature map, and extract feature points in the saliency differential phase feature map;

[0012] Step 5, decompose the three-dimensional phase amplitude feature in the three-dimensional phase feature, and count the feature intensity of each channel, record the optimal feature index, and obtain the optimal feature index map;

[0013] Step 6, count the direction gradient histogram features of the feature point neighborhood in the optimal feature index map, and then construct a three-dimensional phase directional feature descriptor vector feature;

[0014] Step 7, similarity measurement of nearest neighbor and second nearest neighbor distance of three-dimensional phase directional feature descriptor vector, establish initial correspondence relationship through bidirectional matching strategy, identify correct homonymy point pair through gross error elimination method, and complete multi-modal image matching.

[0015] Further, the number of image scale space layers, the number of feature points, the feature point descriptor neighborhood window and the gross error elimination threshold parameters are initialized in step 1.

[0016] Further, the multi-channel difference gradient feature map in step 2 is defined as:

[0017] G r,c,o =|cos o*Gx r,c,o +sin o*Gy r,c,o |

[0018] Wherein G r,c,o represents a multi-channel difference gradient feature map; r is the number of image rows, c is the number of image columns, o is the direction, No is the number of directions, Gx r,c,o is the x-direction gradient feature, and Gy r,c,o is the y-direction gradient feature.

[0019] Further, the mathematical expression of the three-dimensional phase consistency model in step 3 is defined as:

[0020]

[0021] In the formula, PC ρ,Φ (l) represents the gradient phase feature result under the three-dimensional phase consistency model, ω Φ (l) is a frequency expansion weight factor, l represents a polar coordinate, A ρ,Φ (l) is the filter convolution response amplitude of the three-dimensional Log-Gabor filter under the scale p and the three-dimensional space direction Φ at the position l, A ρ,Φ (l)=|EO ρ,Φ |,EO ρ,Φ is the complex response of the three-dimensional Log-Gabor filter under the scale p and the three-dimensional space direction Φ, ΔΨ ρ,Φ (l) represents the phase deviation under l, scale p and direction Φ, which is used to measure the significance of local phase, T is a noise threshold, and ε is a minimum factor to prevent division by zero;

[0022]

[0023] where, and are the Fourier transform and inverse Fourier transform, respectively, filter ρ,o is a frequency domain filter at scale p and three-dimensional spatial direction o, used to extract the frequency components of the image at a specific scale and direction.

[0024] Further, the specific implementation in step 4 is as follows.

[0025] First, the three-dimensional phase orientation feature is summed up, and the specific mathematical expression is as follows:

[0026]

[0027] where, is a three-dimensional tensor representing the phase feature value at position (x, y), direction j, and channel k, j is the direction index representing the directionality of the feature, k is the channel index, usually representing different feature channels or filter responses, and K represents the number of channels; then a three-dimensional Gaussian filter is used to smooth the three-dimensional phase feature to generate the smoothed feature map PC3D_m;

[0028] The difference before and after Gaussian smoothing is used to generate Dog_m1, which is used to capture the significant changes in the feature;

[0029] Dog_m1(x, y, j) = PC3D_m1(x, y, j) - PC_m1(x, y, j)

[0030] In the phase difference map, the maximum value of each pixel in all directions is selected as the final feature response PC_m1, and it is normalized to obtain the locally most significant directional feature;

[0031] PC_m1(x, y) = maxDog_m1(x, y, j)

[0032]

[0033] where, S Dog represents the significant difference phase feature map, min(·) and max(·) represent the minimum and maximum calculation functions, respectively; feature point extraction is performed in the significant difference phase feature map, including FSAT, Harris detection.

[0034] Further, the specific implementation of step 5 includes:

[0035] Firstly, the calculation of the three-dimensional phase feature convolution sequence, the amplitude of the two-dimensional maximum direction phase feature map under each scale ρ is added, and the specific formula is as follows:

[0036]

[0037] In the formula, CS(x, y, j) represents the sequence accumulation value of the phase feature tensor at the position of pixel (x, y) and the jth layer channel; EO i,j (x, y, z) represents the value of the ith input feature map at the jth layer, with z channels, and ns is the number of feature maps calculated for each j channel; max(·) represents taking the maximum value;

[0038] After completing the amplitude addition of the two-dimensional maximum direction phase feature result under each scale ρ, the amplitude maximum value under different directions is counted, that is, the optimal feature index BoIM of the three-dimensional phase feature is completed, and the specific mathematical expression is as follows:

[0039] MI CS (x, y) = argmax j CS(x, y, j)

[0040] MI CS (x, y) represents the optimal feature index result of the three-dimensional phase feature.

[0041] Further, the specific implementation mode of step 6 is as follows:

[0042] In the optimal feature index BoIM, the method of direction gradient histogram descriptor is used for feature vector calculation; for each feature point, a region with a size of K×K is determined, and the region is divided into a K×K grid, in each grid, the pixel histogram of the BoIM map is calculated, and according to the direction of the pixel value, it is distributed to N directions; each small area generates an N-dimensional feature vector, which represents the histogram information of N directions of the area, therefore, the descriptor vector of each feature point is K×K×N-dimensional, which contains the significant information related to the feature point.

[0043] Further, in step 7, the Euclidean distance is used for similarity measurement, and a bidirectional matching strategy is used to ensure one-to-one correspondence of the matching points, and a random sample consensus algorithm is used to eliminate false matches, so as to realize stable matching of multi-modal images.

[0044] Further, step 8 is also included, which evaluates the matching effect of multi-modal images by using correct homonym points, for each image pair, the root mean square error RMSE of homonym points and the number of homonym point pairs are used for quantitative test, wherein the unit of RMSE is pixel.

[0045] The present invention also provides a multimodal image matching system with three-dimensional phase orientation invariant feature transform, comprising:

[0046] The system includes a processor and a memory. The memory stores program instructions, and the processor calls the stored instructions in the memory to execute a multimodal image matching method based on a three-dimensional phase-oriented invariant feature transformation, as described in the above technical solution.

[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0048] This invention proposes a multimodal image matching method based on three-dimensional phase-oriented invariant feature transformation. By introducing an innovative differential phase feature extraction technique and utilizing multi-directional differential filters and three-dimensional phase feature calculation, the spectral differences between multimodal images are effectively reduced, significantly improving the feature point extraction capability. Furthermore, the designed three-dimensional phase-consistent oriented histogram enhances feature representation while maintaining geometric consistency, making it particularly suitable for multimodal images with weak texture and lack of detail. Results show that the proposed method can achieve stable multimodal image matching with good performance. Attached Figure Description

[0049] Figure 1 : Flowchart of the method of this invention;

[0050] Figure 2 : Schematic diagram of feature point extraction from saliency difference phase feature map;

[0051] Figure 3 : Schematic diagram of optimal feature index graph calculation;

[0052] Figure 4 : Multimodal image dataset, where (a) is multi-temporal images, (b) is infrared images and optical images, (c) is electronic navigation maps and optical images, and (d) is SAR images and optical images;

[0053] Figure 5 : Multimodal image matching results, where (a) is the matching result of multi-temporal images, (b) is the matching result of infrared images and optical images, (c) is the matching result of electronic navigation maps and optical images, and (d) is the matching result of SAR images and optical images. Detailed Implementation

[0054] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0055] Please see Figure 1The flow chart shows a three-dimensional phase orientation invariant feature transformation multi-modal matching method provided by the application, which comprises the following steps:

[0056] Step 1: initializing the calculation parameters of multi-modal image matching, and performing nonlinear diffusion on the multi-modal images;

[0057] Preferably, the parameters of the number of image scale space graph layers, the number of feature points, the feature point descriptor neighborhood window and the coarse error elimination threshold in step 1 are initialized. According to a large amount of experimental experience, the number of scale space graph layers is 4, the number of feature points is set to 3500, the feature point descriptor neighborhood window is set to 96 pixels, and the coarse error elimination threshold is set to 3 pixels.

[0058] Step 2: creating an image scale space by using the nonlinear diffusion result to perform Gaussian filtering, calculating a multi-channel image gradient graph, and calculating a multi-channel differential gradient feature graph.

[0059] Preferably, in step 2, the nonlinear diffusion result image generates a set of gradually smoothed image gradient graphs. The Gaussian filter and its derivative help the image to suppress details and extract features through smoothing operation and multi-scale representation, and generate smooth images of different scales. The window size W of the Gaussian filter is determined by the standard deviation σ, and is defined as:

[0060] W = 2 * |2 * σ| + 1 (1)

[0061] Based on the image filtered by the Gaussian filter, a multi-channel Sobel gradient feature is calculated. The gradient filter template H in the x direction is defined as:

[0062]

[0063] The filter template in the y direction is set to the transpose of the filter template in the x direction. After calculating the differential gradient graphs of the image in the x direction and the y direction by using the Sobel filter templates, the gradient direction o is divided into a plurality of discrete directions based on the x direction (usually defined as the horizontal direction, i.e. 0 degree direction). The number of directions No determines the granularity of direction division, and 360 degrees is evenly divided into No directions:

[0064]

[0065] Then the multi-channel gradient feature is defined as:

[0066] G r,c,o = |cos o * Gx r,c,o + sin o * Gy r,c,o | (4)

[0067] Where G r,c,orepresents a multi-channel gradient feature map; r is the number of image rows, c is the number of image columns, o is the direction, No is the number of directions, Gx r,c,o is the x-direction gradient feature, Gy r,c,o is the y-direction gradient feature.

[0068] Step 3: image domain calculation is performed on the multi-channel differential gradient feature map by using Fourier transform, and image phase features are estimated by using a three-dimensional phase consistency model, so as to generate three-dimensional phase features.

[0069] As preferred, the three-dimensional phase consistency model first completes multi-scale filtering, combines the amplitude and phase information of the filtering response, calculates the three-dimensional phase features of the image, and enhances the feature saliency through noise suppression and directionality weight adjustment. In order to better calculate the three-dimensional phase features, it is assumed that the number of rows r and the number of columns c of the image point are known, and a plane direction o is determined, and the corresponding gradient value G r,c,o can also be determined by the scale p and the three-dimensional space direction F (F is composed of an elevation angle q and an azimuth angle d on a unit radius sphere).

[0070] Among them, the key to calculating the three-dimensional phase features is to introduce three-dimensional Log-Gabor filtering, which is similar to two-dimensional Log-Gabor filtering using a phase consistency model for feature detection. The present application extends the two-dimensional phase consistency model to three dimensions to construct a three-dimensional phase consistency model. In order to better load the multi-channel image gradient map as input into the three-dimensional phase consistency model, the present application calculates the multi-direction phase consistency feature map of the phase consistency model and the convolution response value of the three-dimensional Log-Gabor filter, respectively. The mathematical expression of the three-dimensional phase consistency model is defined as:

[0071]

[0072] In the formula, PC ρ,Φ (l) represents the gradient phase feature result under the three-dimensional phase consistency model, ω Φ (l) is a frequency expansion weight factor, l represents the polar coordinate, A ρ,Φ (l) is the filtering convolution response amplitude A ρ,Φ (l) = |EO ρ,Φ |, EO ρ,Φ is the complex response of the three-dimensional Log-Gabor filter at scale p and direction F, and its amplitude A ρ,Φ (l) = |EO ρ,Φ | represents the local signal intensity; ΔΨ ρ,Φ(l) represents the phase deviation at scale p and direction F, used to measure the saliency of local phase, T is a noise threshold, and e is a minimum factor to prevent division by zero.

[0073]

[0074] wherein, and are the Fourier transform and inverse Fourier transform, respectively, filter ρ,o is a frequency domain filter (such as a Log-Gabor filter) at scale p and direction o, used to extract the frequency components of the image at a specific scale and direction.

[0075] Step 4: Perform Gaussian difference filtering on the three-dimensional phase feature, and count the maximum feature value in each channel to generate a saliency difference phase feature map. Perform FAST feature detector on the saliency difference phase feature map to extract feature points.

[0076] As a preferred, after processing the multi-channel gradient feature using the three-dimensional phase consistency model, in order to extract more salient features and for description, first obtain the three-dimensional phase tensor feature, and then perform three-dimensional Gaussian convolution on it. Then, subtract the feature tensor after Gaussian convolution from the feature tensor before convolution to obtain a three-dimensional Gaussian difference phase map. According to the number of division directions No, the dimension of the three-dimensional Gaussian difference phase map is r*c*No. Finally, take the maximum value of the three-dimensional Gaussian difference phase map along the No direction, and the final saliency difference phase feature map is obtained.

[0077] First, perform three-dimensional phase directional feature summation, and the specific mathematical expression is as follows:

[0078]

[0079] wherein, is a three-dimensional tensor representing the phase feature value at position (x, y), direction j, and channel k, j is the direction index representing the directionality of the feature (for example, j can be 1 to No, where No is the number of directions), k is the channel index, which generally represents different feature channels or filter responses, and K represents the number of channels (in the embodiment of the present application, the range of k is 1 to 6). In order to reduce noise interference, a three-dimensional Gaussian filter is used to smooth the three-dimensional phase feature to generate a smoothed feature map PC3D_m.

[0080] Use the difference before and after Gaussian smoothing to generate Dog_m1, which is used to capture the saliency change of the feature.

[0081] Dog_m1(x, y, j) = PC3D_m1(x, y, j) - PC_m1(x, y, j) (8)

[0082] The maximum value of each pixel point in all directions is selected as the final feature response PC_m1 in the phase difference map, and normalized to obtain the local most significant direction feature.

[0083]

[0084] where S Dog represents the minimum and maximum value calculation functions. Feature point extraction is performed in the saliency difference phase feature map, and the FSAT detector is used for feature point extraction, as shown in Figure 2 .

[0085] Step 5: decompose the three-dimensional phase amplitude feature in the three-dimensional phase feature, and count the feature intensity of each channel, record the optimal feature index, and obtain the optimal feature index map.

[0086] As a preferred, when calculating the three-dimensional phase consistency model, a set of three-dimensional orthogonal bases and can be obtained for each position l(ρ,Φ) in the image, and the spectral feature EQ(e ρ,Φ ,q ρ,Φ ) is obtained by combining the multi-directional difference gradient map. In the scale ρ, the three-dimensional phase feature tensor is obtained in the three-dimensional space direction Φ, with a dimension of r*c*No, and then the maximum value along the o direction is counted, i.e. the two-dimensional maximum direction phase feature is obtained.

[0087] First, the calculation of the three-dimensional phase feature convolution sequence, the amplitude of the two-dimensional maximum direction phase feature map in each scale ρ is added, and the specific formula is as follows:

[0088]

[0089] In the formula, CS(x,y,j) represents the sequence accumulation value of the phase feature tensor at the position of pixel (x,y) and the jth layer channel; EO i,j (x,y,z) represents the value of the ith input feature map in the jth layer, with z channels; EO i,j (x,y,z) is a complex response value evolved from EO ρ,Φ , which increases the z channel dimension, represents multi-channel feature information, and its role in the formula is to take the maximum value of z channels and accumulate to generate the sequence accumulation value of the phase feature tensor. ns is the number of feature maps calculated for each j channel; max(·) represents the maximum value.

[0090] After the amplitude addition of the two-dimensional maximum direction phase feature results under each scale p, the maximum amplitude in different directions is calculated, that is, the best-optimal index map (BoIM) of the three-dimensional phase feature is completed, and the specific mathematical expression is as follows:

[0091] MI CS (x,y)=argmax j CS(x,y,j) (12)

[0092] MI CS (x,y) represents the best-optimal index map of the three-dimensional phase feature, and the specific process is as shown in Figure 3

[0093] Step 6: The direction gradient histogram features of the neighborhood of the feature points in the best-optimal index map are counted, and a three-dimensional phase directional feature descriptor vector is constructed.

[0094] As a preferred, in the BoIM, the direction gradient histogram descriptor method is used for feature vector calculation. For each feature point, a region with a size of KxK is determined, and the region is divided into a KxK grid (in the present application, the grid is set to 8x8 according to the existing experience). In each grid, the pixel histogram of the BoIM map is calculated, and the pixel value is assigned to 6 directions (from 0 degree direction to 360 degree direction, if 6 directions are set, 1 corresponds to 60 degrees, 2 corresponds to 120 degrees, and so on, and 6 corresponds to 360 degrees) according to the direction. A 6-dimensional feature vector is generated for each small region, representing the histogram information of the 6 directions of the region. Therefore, the descriptor vector of each feature point is 384-dimensional (8x8x6), which contains significant information related to the feature point.

[0095] Step 7: The Euclidean distance is used as a matching measure to measure the similarity of the nearest neighbor and the second nearest neighbor distance of the three-dimensional phase directional feature descriptor vector, and the initial corresponding relationship is established through a bidirectional matching strategy. Then, the correct homonym pairs are obtained through the elimination of gross errors, and the multi-modal image matching is completed, and the result is as shown in Figure 5

[0096] Step 8: The matching effect of the multi-modal images is evaluated by using the correct homonym points. The performance of the algorithm is tested by using 4 groups of multi-modal images, and the data set is as shown in Figure 4 ​​For each image pair, the Root-Mean-Square Error (RMSE) of the corresponding points and the number of the corresponding point pairs are used for quantitative test, wherein the unit of the RMSE is pixel. The multi-modal image matching method proposed in the application is named as DoIFT algorithm, and is compared with several optimal image matching methods (LGHD, PSO-SIFT and RIFT). The comparison results are shown in Table 1.

[0097] Table 1 Comparison of several image matching methods

[0098]

[0099] As shown in Table 1, in the multi-modal remote sensing image data, the DoIFT algorithm can obtain more corresponding point pairs than the LGHD, PSO-SIFT and RIFT algorithms. The DoIFT algorithm proposed in the application can achieve a relatively optimal result. The RMSE of the DoIFT algorithm is better than the LGHD, PSO-SIFT and RIFT methods. The value of the RMSE of the DoIFT algorithm proposed in the application is less than 2 pixels, which further proves that the DoIFT algorithm not only greatly increases the number of the corresponding point pairs, but also maintains good matching accuracy. Meanwhile, through a large number of experiments, it is found that when the multi-modal image extraction is difficult, the size of the descriptor field window can be increased; on the contrary, when the multi-modal image texture is rich, the value of the parameter can be appropriately reduced. The number of the image scale space image layers is set to be between 3 and 6.

[0100] On the other hand, the embodiment of the application further provides a multi-modal image matching system based on three-dimensional phase orientation invariant feature transformation, comprising:

[0101] A processor and a memory, the memory is used for storing program instructions, and the processor is used for calling the storage instructions in the memory to execute the multi-modal image matching method based on three-dimensional phase orientation invariant feature transformation as described in the above technical solution.

[0102] It should be understood that the parts not described in detail in the specification are all prior art.

[0103] It should be understood that the above description of the preferred embodiments is more detailed, and therefore should not be considered as a limitation on the scope of patent protection of the application. Those skilled in the art can make substitutions or modifications without departing from the scope of protection of the claims of the application, and all fall within the scope of protection of the application. The scope of protection of the application should be subject to the appended claims.

Claims

1. A multimodal image matching method using three-dimensional phase orientation invariant feature transform, characterized in that, Includes the following steps: Step 1: Initialize the calculation parameters for multimodal image matching and perform nonlinear diffusion on the multimodal image; Step 2: Calculate the multi-channel differential gradient feature map using the nonlinear diffusion results; Step 3: Perform image domain calculation on the multi-channel differential gradient feature map, and use the three-dimensional phase consistency model to estimate the image phase features, thereby generating three-dimensional phase features; The mathematical expression for the three-dimensional phase consistency model in step 3 is defined as follows: In the formula, This represents the gradient phase feature results under the three-dimensional phase consistency model. For the frequency spread weighting factor, Represents polar coordinates. In scale and three-dimensional spatial direction The three-dimensional Log-Gabor filter at position The amplitude of the filtered convolution response at the point, , It is the complex response of a three-dimensional Log-Gabor filter at scale ρ and three-dimensional spatial direction Φ. Indicates in ,scale and direction Φ The phase deviation below is used to measure the significance of local phase. It is a noise threshold. It is the smallest factor that prevents division by zero; in, and These are the Fourier transform and the inverse Fourier transform, respectively. This represents a multi-channel differential gradient feature map. In scale and three-dimensional spatial direction The frequency domain filter is used to extract the frequency components of the image at a specific scale and direction. Step 4: Perform Gaussian difference filtering using three-dimensional phase features, and count the maximum eigenvalue in each channel to generate a significant differential phase feature map. Extract feature points from the significant differential phase feature map. The specific implementation method in step 4 is as follows: First, the three-dimensional phase orientation features are summed, and the specific mathematical expression is as follows: in, It is a three-dimensional tensor representing the position ( x , y ),direction j and channels k The phase eigenvalues ​​below, j It is a direction index, indicating the directionality of the feature. k It is a channel index, which typically represents different characteristic channels or filter responses. K The number of channels is indicated; then, a three-dimensional Gaussian filter is used to smooth the three-dimensional phase features, generating a smoothed feature map. ; Generate using the difference before and after Gaussian smoothing It is used to capture significant changes in features; In the phase difference map, the maximum value of each pixel in all directions is selected as the final feature response. Then, normalize the values ​​to obtain the most significant directional features in the local area; in, This represents a significant differential phase feature map. and These represent the functions for calculating the minimum and maximum values, respectively. Step 5: Decompose the three-dimensional phase amplitude features in the three-dimensional phase features, count the feature intensity of each channel, record the optimal feature index, and obtain the optimal feature index map. Step 6: Statistically analyze the directional gradient histogram features of the neighborhood of feature points in the optimal feature index map, and then construct the three-dimensional phase orientation feature descriptor vector features. Step 7: Measure the similarity of the three-dimensional phase orientation feature descriptor vectors by the nearest neighbor and second nearest neighbor distances, establish the initial correspondence through a bidirectional matching strategy, and then identify the correct corresponding point pairs through the gross error elimination method to complete the multimodal image matching.

2. The multimodal image matching method based on three-dimensional phase-oriented invariant feature transform according to claim 1, characterized in that: Step 1 initializes the number of layers in the image scale space, the number of feature points, the feature point descriptor neighborhood window, and the gross error removal threshold parameter.

3. The multimodal image matching method based on three-dimensional phase orientation invariant feature transform according to claim 1, characterized in that: In step 2, the multi-channel differential gradient feature map is defined as follows: in Represents a multi-channel differential gradient feature map; For the number of image rows, For the number of image columns, As direction, For direction number, for x Oriented gradient features for y Oriented gradient characteristics.

4. The multimodal image matching method with three-dimensional phase-oriented invariant feature transform according to claim 1, characterized in that: In step 4, feature points are extracted from the saliency differential phase feature map, including FSAT and Harris detection.

5. The multimodal image matching method based on three-dimensional phase-oriented invariant feature transform according to claim 1, characterized in that: The specific implementation of step 5 includes: First, the calculation of the three-dimensional phase feature convolution sequence involves various scales. The amplitudes of the two-dimensional maximum direction phase feature maps are summed, as shown in the following formula: In the formula, Indicated in pixels Position and number j The sequence accumulation value of the phase feature tensor on the layer channel; Indicates the first i The input feature map at the th ... j The value of the layer has z One channel, ns For each j The number of feature maps computed per channel; This indicates taking the maximum value; Complete each scale After summing the amplitudes of the two-dimensional maximum direction phase feature results, the maximum amplitude values ​​in different directions are used to complete the optimal feature index BoIM for the three-dimensional phase feature. The specific mathematical expression is as follows: The optimal feature index result representing the three-dimensional phase features.

6. The multimodal image matching method based on three-dimensional phase-oriented invariant feature transform according to claim 1, characterized in that: The specific implementation method of step 6 is as follows; In the Optimal Feature Index BoIM, the directional gradient histogram descriptor method is used to calculate feature vectors; for each feature point, a descriptor of size is determined. K × K The area, and divided into K × K The BoIM image is divided into grids. Within each grid, a pixel histogram is calculated, and the pixel values ​​are assigned to N directions based on their orientation. Each small region generates an N-dimensional feature vector representing the histogram information of the region in N directions. Therefore, the descriptor vector for each feature point is... K × K It is ×N-dimensional and contains significant information related to feature points.

7. The multimodal image matching method based on three-dimensional phase-oriented invariant feature transform according to claim 1, characterized in that: In step 7, Euclidean distance is used for similarity measurement, and a bidirectional matching strategy is used to ensure one-to-one correspondence of matching points. A random sampling consensus algorithm is used to eliminate erroneous matches, thereby achieving stable matching of multimodal images.

8. The multimodal image matching method based on three-dimensional phase-oriented invariant feature transform according to claim 1, characterized in that: It also includes step 8, which evaluates the matching effect of multimodal images using correct corresponding points. For each image pair, quantitative verification is performed using the root mean square error (RMSE) of corresponding points and the number of matching corresponding point pairs, where the unit of RMSE is pixels.

9. A multimodal image matching system with three-dimensional phase orientation invariant feature transform, characterized in that, include: The processor and memory are provided, wherein the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the multimodal image matching method of three-dimensional phase orientation invariant feature transformation as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Dual-polarization SAR image system

    CN112558066A

  • Multi-modal remote sensing image matching method based on weighted phase directional description

    CN115861792A