A Multimodal Remote Sensing Image Matching Method Based on Weighted Phase Orientation Description
Through the methods of aggregate feature extraction and weighted phase orientation description, the nonlinear distortion and geometric difference of multimodal remote sensing images are solved, more robust image matching is achieved, and the application effect of multimodal remote sensing images in fields such as natural disaster assessment is improved.
Patent Information
- Application Number
- CN202211397518.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-09
AI Technical Summary
Due to different imaging mechanisms, multimodal remote sensing images have significant nonlinear distortion and geometric differences, which lead to difficulty in extracting feature points and robust description, affecting the matching effect, especially in the fields of natural disaster assessment, disaster relief search, change detection, image stitching and three-dimensional reconstruction.
The method of aggregate feature extraction and weighted phase orientation description is adopted to solve the maximum and minimum moment characteristics of the image through the phase consistency model, and the spot and corner characteristics are extracted in combination with the Heisen matrix and the Fast detector. The weighted phase orientation feature map is constructed using the weighted bandwidth function, and regularized polar coordinate descriptors are generated, and matched by Euclidean distance.
It effectively overcomes the problems of phase extreme value mutation and direction reversal, improves the matching stability and accuracy of multimodal remote sensing images, increases the descriptor expression ability, and achieves more robust image matching.
Smart Images

Figure CN115861792B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image processing methods, and particularly relates to a multi-modal remote sensing image matching method for aggregating feature extraction and weighted phase orientation description. Background Art
[0002] Multi-modal remote sensing images generally refer to two or more images taken at different times under different sensors or different imaging conditions (such as: point cloud depth maps, infrared images, navigation maps, SAR images, night light images, etc.). Due to the different imaging mechanisms of such images, they have significant non-linear distortions and geometric differences, and often cannot achieve the extraction and robust description of feature points, resulting in difficult and ineffective multi-modal remote sensing image matching. However, multi-modal remote sensing images play an important role in fields such as natural disaster assessment, disaster relief search, change detection, image stitching, aerial triangulation, and 3D reconstruction. Therefore, it is very necessary to conduct research on this.
[0003] With the continuous development of computer vision and image processing technologies, remote sensing image matching methods can be roughly divided into three categories, namely region-based methods, feature-based methods, and deep learning-based methods. Region-based methods mainly rely on features such as image intensity and mutual information to calculate the similarity of images, and have the advantage of high matching accuracy, but there are problems such as large computational complexity and poor scale and rotation invariance of images, which restrict their application scope. Feature-based methods mainly focused on image gradient features in the early stage, such as many methods like SIFT and SIFT-like. Such methods can achieve satisfactory results under translation, scale, and rotation differences. However, in multi-modal images, due to the sensitivity of image gradient features to the non-linear distortion and poor contrast of multi-modal remote sensing images, they are not suitable for multi-modal matching. Later, scholars applied the phase features of images to multi-modal remote sensing image matching and achieved breakthroughs, such as algorithms like phase consistency orientation histogram and rotation invariant feature transform. However, they all restrict the scale and rotation invariance of the algorithms to a certain extent. Although some people have tried to improve them, problems such as phase direction reversal and phase extreme value mutation in multi-modal remote sensing images still exist. Deep learning-based methods, although they have shown great development potential and achieved remarkable results in image matching. However, due to the difficulty of collecting sample data of multi-modal remote sensing images and the complexity of remote sensing image scene applications, deep learning matching methods still require a certain amount of development time.
[0004] In summary, the ability to extract effective feature points and robust descriptors is the key to the success of multimodal image matching. In feature extraction, most scholars focus on finding corner or blob features between images, while there is less research on the comprehensive extraction of image corners and blobs. In feature description, the phase features directly extracted using the phase consistency model have significant problems such as significant direction reversal and sudden change of phase extreme values. Phase essentially represents the shape component of an image, and the above problems change the image shape component to a certain extent, thus destroying the structural integrity of the image, resulting in the inability of the phase orientation feature to correctly characterize the direction change between images and increasing the description difficulty of the feature descriptor. Based on this, the present invention proposes a multimodal remote sensing image matching method of aggregated feature extraction and weighted phase orientation description to achieve their robust matching. Summary of the Invention
[0005] The present invention proposes a multimodal remote sensing image matching method of weighted phase orientation description to solve the matching problem of multimodal remote sensing images.
[0006] The technical solution adopted by the present invention is: a multimodal remote sensing image matching method of aggregated feature extraction and weighted phase orientation description, including the following steps:
[0007] Step 1, initialize the calculation parameters for multimodal remote sensing image matching, and perform non-linear diffusion on the multimodal remote sensing images;
[0008] Step 2, respectively solve the maximum moment feature and the minimum moment feature in the image scale space through the phase consistency model, and output the solution results;
[0009] Step 3, use the Hessian matrix to extract blob features from the maximum moment, and use the Fast detector to extract corner features from the minimum moment. Complete the extraction of the corresponding layers in sequence and summarize them into a feature point set, and output the feature point set;
[0010] Step 4, use the aggregated feature optimization model to filter the feature points extracted in Step 3, and determine the final feature point set by setting the feature point significance detection threshold;
[0011] Step 5, use the weighted bandwidth function to calculate the phase shape component of the multimodal remote sensing image, and obtain the weight coefficient of the image in the phase feature;
[0012] Step 6, use the weight coefficient to construct a weighted phase orientation feature model, and calculate the weighted phase orientation feature map of the multimodal remote sensing image;
[0013] Step 7, according to the obtained weighted phase orientation feature map, calculate the regularized logarithmic polar coordinate descriptor of each feature point, and output the descriptor vector set of the feature points;
[0014] Step 8: Match according to the descriptor vector set of feature points, eliminate gross errors, and complete the matching of multimodal remote sensing images.
[0015] Further, in Step 1, initialize the number of layers of the non-linear image scale space, the range of the non-maximum threshold window of feature points, the filtering threshold of feature points, and the initial neighborhood range size parameter of the descriptor.
[0016] Further, the formulas for the maximum moment and minimum moment features in Step 2 are shown in (1) and (2):
[0017]
[0018]
[0019] In formula (1), PC(x, y) represents the result of the phase consistency measure of the image; w O (x, y) represents the weighting function; A SO (x, y) represents the amplitude component; S represents the scale; O represents the convolution direction; ξ represents a minimum value; represents taking zero when the enclosed quantity is negative; ΔΦ SO (x, y) represents the phase deviation function; T represents the noise compensation term; in formula (2), M max represents the maximum moment of the phase consistency of the image; M min represents the minimum moment of the phase consistency of the image; PC(θo) represents the mapping of the image in the o direction; A, B, and C are intermediate quantities for phase moment calculation; θ is the angle of direction o.
[0020] Further, the specific implementation method in Step 4 is as follows;
[0021]
[0022] In formula (3), S points represents the set of spots and corners after mask optimization, f Blob represents the spot extraction function; f Corner represents the corner extraction function, M max represents the maximum moment of the phase consistency of the image; M min represents the minimum moment of the phase consistency of the image; represents the mask function, R represents the mask radius; N w represents the neighborhood window, where the mask function is defined as
[0023]
[0024] In formula (4), Mask(x, y) represents the solution result of the mask function; I xRepresents the pixel value of the original image in the x direction; I y Represents the pixel value of the original image in the y direction; N w Represents the neighborhood window;
[0025] Construct a saliency score extraction equation, and its mathematical expression is shown in Equation (5):
[0026]
[0027] In Equation (5), f nms (·) represents the non-maximum suppression function obtained by the feature points; AF represents the final key points; S score Represents the set of points filtered by the saliency score; k represents the filtering threshold of the feature points, and this value is not a fixed value and is set and adjusted according to the strength of the texture of different remote sensing images; Represents its phase intensity value, f nms (·) represents the feature points retained after non-maximum suppression, and S points Represents the points removed from the edge.
[0028] Furthermore, the specific implementation of step 5 includes,
[0029] First, perform Log-Gabor convolution on the multi-modal remote sensing image to extract the phase information of the image, and decompose the image into two parts, and its formula is shown in (6):
[0030]
[0031] In Equation (6), s and o represent the scale and direction of the Log-Gabor filter; Represents the even-symmetric filter of Log-Gabor filtering; Represents the odd-symmetric filter; the symbol i represents the imaginary unit of the complex number; I(x, y) represents the image; EO(x, y) represents the response result of the image on the real part filter; OO(x, y) represents the response result of the image on the imaginary part filter; Represents the convolution operation;
[0032] Secondly, a weighted bandwidth function is designed according to the response results of the real part and the imaginary part of the Log-Gabor filter, and its mathematical expression is shown in (7):
[0033]
[0034] In Equation (7), wc′ represents the maximum weighted coefficient of the image; wc″ represents the minimum weighted coefficient of the image; exp represents the exponential calculation; Cutoff represents the fractional scale of the frequency distribution; width irepresents the frequency score in the current image; g represents the sharpness of the transformation in the phase consistency model; ξ represents a minimum value to prevent a zero value.
[0035] Furthermore, in step 6, for the solution of the weighted phase orientation feature, its mathematical expression is as shown in Equation (8):
[0036]
[0037] In Equation (8), W represents the weighted phase orientation feature; OO i represents the odd-symmetric convolution result on the i-th direction layer; EO i represents the even-symmetric convolution result on the i-th direction layer; θ represents the rotation angle.
[0038] Furthermore, the specific implementation method of step 7 is as follows;
[0039] On the weighted phase orientation feature map, the descriptor neighborhood range is divided into 4 layers and equally divided into 12 parts, that is, divided into 4 concentric circles. The four concentric circles are regularly equally divided, and finally the entire feature neighborhood range is divided into a 48-sub-region grid of epipolar coordinates; where the horizontal direction in each grid represents the polar angle of the position of the circular neighborhood pixel points. After calculating the direction histogram of each feature point, each dimension is divided every 45°, and the 0-360° direction is divided into 8 dimensions; therefore, the adjacent points of each sub-region grid have a gradient position orientation histogram of 8 dimensions, and finally a 384-dimensional regularized log-polar coordinate descriptor is generated;
[0040] The mathematical expression of the regularized log-polar coordinate descriptor is as shown in (9):
[0041] RGLOH = [D1, D2, …, D N T (9)
[0042] In Equation (9), RGLOH represents the descriptor set of all feature points; D i represents the 384-dimensional characteristic vector value of the i-th feature point, that is, the 384-dimensional regularized log-polar coordinate descriptor; T represents the matrix transpose character; N represents the number of feature points.
[0043] Furthermore, in step 8, the Euclidean distance is used to measure the descriptors of the feature points, so as to complete the nearest neighbor matching, and the gross errors are removed by the fast sample consensus algorithm.
[0044] Furthermore, it also includes step 9, using the correct homologous points to evaluate the matching effect of the multi-modal remote sensing images. In step 9, the root mean square error of the solved homologous points and the number of homologous point pairs are used to quantitatively test the matching accuracy.
[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0046] The multi-modal remote sensing image matching method proposed by the present invention is divided into three parts: feature point extraction with feature aggregation, descriptor construction, matching, and gross error rejection. First, the phase consistency model is used to establish the maximum moment and minimum moment features of the image and extract the spot and corner features respectively, and then the key points are extracted by optimizing the feature points through the designed aggregation feature equation. Secondly, the weighted phase orientation features are constructed by the designed weighted bandwidth function, and the regularized polar coordinate descriptors are generated through these features. Finally, the Euclidean distance is used for nearest neighbor image matching, and the wrong points are removed by means of the fast sample consensus algorithm to complete the image matching. The weighted phase orientation features constructed by the proposed weighted bandwidth function can better overcome problems such as sudden change of phase extreme value and direction reversal, and increase the expression ability of the descriptor. The results show that the method proposed by the present invention can better achieve the matching of multi-modal remote sensing images and is more stable than the traditional method. Brief Description of the Drawings
[0047] Figure 1 : Flow chart of the method of the present invention;
[0048] Figure 2 : Schematic diagram of the optimization of aggregated feature points;
[0049] Figure 3 : Schematic diagram of the logarithmic polar coordinate descriptor of the weighted phase orientation feature;
[0050] Figure 4 : Multi-modal remote sensing image dataset, where (a) is the point cloud depth map and the optical image, (b) is the infrared image and the optical image, (c) is the navigation map and the optical image, (d) is the SAR image and the optical image, (e) is the night light image and the optical image;
[0051] Figure 5 : Multi-modal remote sensing image matching result. Detailed Embodiment
[0052] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0053] Please refer to Figure 1 the flow chart. A multi-modal remote sensing image matching method with weighted phase orientation description provided by the present invention includes the following steps:
[0054] Step 1: Initialize the calculation parameters for multi-modal remote sensing image matching, and perform non-linear diffusion on the multi-modal remote sensing image.
[0055] Preferably, for the non-linear image scale space in step 1, during the construction of the scale space, the initialization is completed mainly by the number of layers of the non-linear image scale space, the size of the non-maximum suppression threshold window of the feature points, the filtering threshold of the feature points, and the size of the initial neighborhood range of the descriptor. According to a large number of experimental experiences, the number of layers of the non-linear image scale space, the size of the non-maximum suppression threshold window of the feature points, the filtering threshold of the feature points, and the size of the initial neighborhood range of the descriptor are set to 3, 3, 0.85, and 38 respectively.
[0056] Step 2: Calculate the maximum moment feature and the minimum moment feature of the image scale space respectively through the phase congruency model, and output the calculation results. The phase congruency model is a model that can effectively detect the local energy features of an image. This model is convenient for extracting the edge features and corner features of an image. The maximum moment and the minimum moment are used to better describe the saliency features of an image. According to the moment analysis algorithm, the axis corresponding to the minimum moment is called the principal axis, and this principal axis usually represents the direction information of the feature. The axis of the maximum moment is perpendicular to the principal axis and reflects the saliency of the feature. Extracting feature points of multi-modal remote sensing images is more conducive through the maximum moment and minimum moment features.
[0057] Therefore, the present invention calculates the maximum moment and minimum moment features of the image through the phase congruency model. The phase congruency measurement is mainly calculated through local Fourier transform. It can well extract the frequency domain feature information of the image. The present invention uses the phase congruency model to extract the maximum moment and minimum moment features of multi-modal remote sensing images, and its formulas are shown in (1) and (2):
[0058]
[0059]
[0060] In formula (1), PC(x, y) represents the phase congruency measurement result of the image; w O (x, y) represents the weighting function; A SO (x, y) represents the amplitude component; S represents the scale; O represents the convolution direction; ξ represents a minimum value; represents taking zero when the enclosed quantity is negative; ΔΦ SO (x, y) represents the phase deviation function; T represents the noise compensation term. In formula (2), M max represents the maximum moment of the phase congruency of the image; M min represents the minimum moment of the phase congruency of the image; PC(θo) represents the mapping of the image in the o direction; A, B, and C are intermediate quantities for phase moment calculation; θ is the angle of the direction o.
[0061] Step 3: Use the Hessian matrix to extract the blob features of the maximum moment, and use the Fast detector to extract the corner features of the minimum moment. Complete the extraction of the corresponding layers in sequence and summarize them into a feature point set, and output this feature point set. Feature point extraction is an important step in multimodal image matching. Common feature points can be mainly divided into two categories: blobs and corners. Blobs usually refer to areas with color and grayscale differences from their surroundings, and they have the advantages of strong noise resistance and good stability. Corners usually refer to the corner areas of objects in the image or the intersection parts between lines, and the significance of corners is relatively high.
[0062] Step 4: Use the designed aggregated feature optimization model to filter the feature points extracted in Step 3, and determine the final feature point set by setting the feature point significance detection threshold.
[0063] Traditional detectors only extract one type of key points, which is not beneficial for image matching. Inspired by this, the present invention designs an aggregated feature strategy to optimize the feature points, thereby improving the richness of the feature points. The aggregated feature optimization strategy has three steps: removal of boundary region points, non-maximum suppression, and significance score filtering. The mathematical expression for the removal of boundary region points is shown in Equation (3):
[0064]
[0065] In Equation (3), S points represents the set of blobs and corners after mask optimization. f Blob represents the blob extraction function; f Corner represents the corner extraction function; represents the mask function, R represents the mask radius; N w represents the neighborhood window. Among them, the mask function is defined as,
[0066]
[0067] In Equation (4), Mask(x, y) represents the solution result of the mask function; I x represents the pixel value of the original image in the x direction; I y represents the pixel value of the original image in the y direction; N w represents the neighborhood window.
[0068] Non-maximum suppression is a commonly used method, and the present invention will not elaborate on it.
[0069] Finally, calculate the significance scores of the feature points after the first two operations, and its mathematical expression is shown in Equation (5):
[0070]
[0071] In Equation (5), fnms (·) represents the non-maximum suppression function for feature points; AF represents the final key points; S score represents the set of points filtered by the saliency score; k represents the filtering threshold for feature points, which is a non-fixed value and is set and adjusted according to the strength of the texture of different remote sensing images; represents its phase intensity value, f nms (·) represents the feature points retained after non-maximum suppression, and S points represents the points removed from the edge, and the result is as Figure 2 shown.
[0072] Step 5: Calculate the change of the phase shape component of the multimodal remote sensing image using the designed weighted bandwidth function to obtain the noise weight coefficient of the image in the phase feature.
[0073] First, perform convolution processing on the input image using a Log-Gabor filter to extract local phase information. The Log-Gabor filter is decomposed into two parts in the spatial domain, which can better resist the noise and gray level differences of the image. Its formula is as shown in (6):
[0074]
[0075] In formula (6), s and o represent the scale and direction of the Log-Gabor filter; represents the even-symmetric filter of the Log-Gabor filter; represents the odd-symmetric filter; the symbol i represents the imaginary unit of a complex number; I(x, y) represents the image; EO(x, y) represents the response result of the image on the real part filter; OO(x, y) represents the response result of the image on the imaginary part filter; represents the convolution operation;
[0076] Secondly, a weighted bandwidth function is designed according to the response results of the real part and the imaginary part of the Log-Gabor filter, and its mathematical expression is as shown in (7):
[0077]
[0078] In formula (7), wc′ represents the maximum weighted coefficient of the image; wc″ represents the maximum weighted coefficient of the image; exp represents the exponential calculation; Cutoff represents the fractional scale of the frequency distribution; width i represents the frequency score in the current image; g represents the sharpness of the conversion in the phase consistency model; ξ represents a minimum value to prevent the value from being zero.
[0079] Step 6: Construct a weighted phase orientation feature model and calculate the weighted phase orientation feature map of the multimodal remote sensing image.
[0080] Due to the influence of phase energy extreme value mutation and direction reversal in the phase orientation features of multi-modal images, the descriptors of feature points are significantly unstable. To better solve such problems, the present invention introduces the result of the weighted bandwidth function into the odd-symmetric function and even-symmetric function of Log-Gabor. This operation can effectively overcome the negative effects brought by phase extreme value mutation and direction reversal, thereby enhancing the robustness of the feature point descriptors.
[0081] Substitute the weighted bandwidth function into the calculation of the phase orientation feature, and its mathematical expression is as shown in (8):
[0082]
[0083] In formula (8), W represents the weighted phase orientation feature; OO i represents the odd-symmetric convolution result on the i-th direction layer; EO i represents the even-symmetric convolution result on the i-th direction layer; θ represents the rotation angle. Secondly, migrate the orientation feature to between [0, 360] to generate the final weighted phase orientation feature.
[0084] Step 7: According to the obtained weighted phase orientation feature map, calculate the regularized log-polar coordinate descriptor of each feature point, and output the descriptor vector set of the feature points, and the result is as Figure 3 shown.
[0085] Considering the stability and robustness of the descriptor, the present invention divides the descriptor neighborhood range into 4 layers and performs 12 equal divisions on the weighted phase orientation feature map. That is, it is divided into 4 concentric circles, and the four concentric circles are regularly equally divided. Finally, the entire characteristic neighborhood range is divided into a pair of polar number coordinate grids of (12×4) = 48 sub-region grids. This grid division method makes up for the instability of the descriptor caused by a smaller area, making the description of the feature point neighborhood range more detailed and accurate.
[0086] The areas of each polar coordinate sub-region are approximately the same. Among them, the horizontal direction in each grid represents the polar angle of the position of the circular neighborhood pixel points. After calculating the direction histogram of each feature point, divide one dimension every 45°, and divide the 0-360° direction into 8 dimensions. Therefore, the adjacent points of each sub-region grid have an 8-dimensional gradient position orientation histogram, and finally a 384-dimensional regularized log-polar coordinate descriptor is generated.
[0087] The mathematical expression of the regularized log-polar coordinate descriptor is as shown in (9) below:
[0088] RGLOH = [D1, D2, …, D N T , (9)
[0089] In Equation (9), RGLOH represents the descriptor set of all feature points; D i represents the 384-dimensional characteristic vector value of the i-th feature point, that is, the 384-dimensional regularized log-polar descriptor; T represents the matrix transpose character; N represents the number of feature points.
[0090] Step 8: After completing the descriptor calculation, the Euclidean distance is used to measure the descriptors of the feature points, thereby completing the nearest neighbor matching. The gross errors are removed through the Random Sample Consensus (RANSAC) algorithm to complete the matching of multimodal remote sensing images, and the results are as Figure 5 shown.
[0091] Step 9: Use the correct homologous points to evaluate the matching effect of multimodal remote sensing images. The present invention uses 5 groups of multimodal remote sensing images to test the performance of the algorithm, and the data set is shown in Figure 4 . For each image pair, the Root-Mean-Square Error (RMSE) of the homologous points and the number of matching homologous point pairs are used for quantitative inspection, where the unit of RMSE is pixels. The multimodal remote sensing image matching method proposed in the present invention is named the HORP algorithm, and it is compared with several optimal image matching methods (LGHD, PSO-SIFT, and RIFT). The comparison results are shown in Table 1.
[0092] Table 1 Comparison of several image matching methods
[0093]
[0094] As can be seen from Table 1, in the multimodal remote sensing image data, the HORP algorithm can obtain more homologous point pairs compared with the LGHD, PSO-SIFT, and RIFT algorithms. The relatively optimal result can be achieved through the HORP algorithm proposed in the present invention. Among them, the RMSE of the HORP algorithm is better than that of the LGHD, PSO-SIFT, and RIFT methods. The RMSE values of the HORP algorithm proposed in the present invention are all less than 2 pixels, which further proves that the HORP algorithm not only greatly increases the number of matching homologous points but also maintains good matching accuracy. At the same time, through a large number of experiments, it is found that when the extraction difficulty of multimodal remote sensing images is relatively large, the size of the descriptor domain window can be increased; conversely, when the multimodal remote sensing images are rich in texture, the parameter values can be appropriately reduced. Among them, the number of layers parameter of the image scale space is set between 2 and 6.
[0095] It should be understood that the parts not elaborated in detail in this specification all belong to the prior art.
[0096] It should be understood that the above description of the preferred embodiment is relatively detailed, and it should not be considered as a limitation on the protection scope of the present invention patent. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the scope protected by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of protection claimed by the present invention shall be subject to the appended claims.
Claims
1. A multimodal remote sensing image matching method based on weighted phase orientation description, characterized in that, It includes the following steps: Step 1, initialize the calculation parameters for multimodal remote sensing image matching, and perform non-linear diffusion on the multimodal remote sensing image; Step 2, respectively calculate the maximum moment feature and the minimum moment feature of the image scale space through the phase consistency model, and output the calculation results; Step 3, use the Hessian matrix to extract the speckle features of the maximum moment, and use the Fast detector to extract the corner features of the minimum moment. Sequentially complete the extraction of the corresponding layers and summarize them into a feature point set, and output the feature point set; Step 4, use the aggregated feature optimization model to filter the feature points extracted in Step 3, and determine the final feature point set by setting the feature point significance detection threshold; Step 5, use the weighted bandwidth function to calculate the phase shape component of the multimodal remote sensing image, and obtain the weight coefficient of the image in the phase feature; Step 6, use the weight coefficient to construct a weighted phase orientation feature model, and calculate the weighted phase orientation feature map of the multimodal remote sensing image; Step 7, according to the obtained weighted phase orientation feature map, calculate the regularized log-polar coordinate descriptor of each feature point, and output the descriptor vector set of the feature points; Step 8, perform matching according to the descriptor vector set of the feature points, and eliminate outliers to complete the matching of the multimodal remote sensing image.
2. The multimodal remote sensing image matching method described by weighted phase orientation according to claim 1, wherein: In Step 1, initialize the number of layers of the non-linear image scale space, the size of the non-maximum threshold window range of the feature points, the filtering threshold of the feature points, and the initial neighborhood range size parameter of the descriptor.
3. A multimodal remote sensing image matching method described by weighted phase orientation according to claim 1, characterized in that: The formulas for the maximum moment and the minimum moment features in Step 2 are shown in (1) and (2): In Equation (1), PC(x, y) represents the result of the phase consistency measure of the image; w O (x, y) represents the weighting function; A SO (x, y) represents the amplitude component; S represents the scale; O represents the convolution direction; ξ represents a minimum value; represents taking zero when the enclosed quantity is negative; ΔΦ SO (x, y) represents the phase deviation function; T represents the noise compensation term; in Equation (2), M max represents the maximum moment of the phase consistency of the image; M min represents the minimum moment of the phase consistency of the image; PC(θo) represents the mapping of the image in the o direction; A, B, and C are intermediate quantities for phase moment calculation; θ is the angle of direction o.
4. A multimodal remote sensing image matching method described by weighted phase orientation according to claim 1, characterized in that: The specific implementation method in Step 4 is as follows; In formula (3), S points represents the set of spots and corners after mask optimization, f Blob represents the spot extraction function; f Corner represents the corner extraction function, M max represents the maximum moment of the phase consistency of the image; M min represents the minimum moment of the phase consistency of the image; represents the mask function, R represents the mask radius; N w represents the neighborhood window, where the mask function is defined as, In Equation (4), Mask(x, y) represents the solution result of the mask function; I x represents the pixel value of the original image in the x direction; I y represents the pixel value of the original image in the y direction; N w represents the neighborhood window; Construct a significance score extraction equation, and its mathematical expression is shown in Equation (5): In formula (5), f nms (·) represents the non-maximum suppression function obtained from the feature points; AF represents the final key points; S score represents the set of points filtered by the significance score; k represents the filtering threshold of the feature points, which is a non-fixed value and is set and adjusted according to the strength of different remote sensing image textures; represents its phase intensity value, f nms (·) represents the feature points retained after non-maximum suppression, and S points represents the points removed from the edge.
5. A multimodal remote sensing image matching method described by weighted phase orientation according to claim 1, characterized in that: The specific implementation of Step 5 includes First, perform Log-Gabor convolution on the multimodal remote sensing image to extract the phase information of the image, and decompose the image into two parts, and its formula is shown in (6): In Equation (6), s and o represent the scale and orientation of the Log-Gabor filter; represents the even-symmetric filter for Log-Gabor filtering; represents the odd-symmetric filter; the symbol i represents the imaginary unit of a complex number; I(x, y) represents the image; EO(x, y) represents the response result of the image on the real part filter; OO(x, y) represents the response result of the image on the imaginary part filter; represents the convolution operation; Secondly, design a weighted bandwidth function according to the response results of the real part and the imaginary part of the Log-Gabor filter, and its mathematical expression is shown in (7): In Equation (7), wc′ represents the maximum weighting coefficient of the image; wc″ represents the minimum weighting coefficient of the image; exp represents exponential calculation; Cutoff represents the fractional scale of the frequency distribution; width i represents the frequency score in the current image; g represents the sharpness of the transformation in the control phase consistency model; ξ represents a minimum value to prevent a zero value.
6. A multi-modal remote sensing image matching method described by weighted phase orientation according to claim 5, characterized in that: In Step 6, for the calculation of the weighted phase orientation feature, its mathematical expression is shown in Equation (8): In Equation (8), W represents the weighted phase orientation feature; OO i represents the odd-symmetric convolution result on the i-th directional layer; EO i represents the even-symmetric convolution result on the i-th directional layer; θ represents the rotation angle.
7. A multimodal remote sensing image matching method based on weighted phase orientation description according to claim 1, characterized in that: The specific implementation method of Step 7 is as follows; On the weighted phase orientation feature map, divide the descriptor neighborhood range into 4 layers and divide it into 12 equal parts, that is, divide it into 4 concentric circles, and perform regularized equal division on the four concentric circles. Finally, divide the entire characteristic neighborhood range into a 48-sub-region grid of polar number coordinates; where the horizontal direction in each grid represents the polar angle of the position of the circular neighborhood pixel points. After calculating the direction histogram of each feature point, divide one dimension every 45°, and divide the 0-360° direction into 8 dimensions; therefore, the adjacent points of each sub-region grid have an 8-dimensional gradient position orientation histogram, and finally a 384-dimensional regularized log-polar coordinate descriptor is generated; The mathematical expression of the regularized log-polar coordinate descriptor is shown in (9): RGLOH = [D1, D2, …, D N T (9) In Equation (9), RGLOH represents the descriptor set of all feature points; D i represents the 384-dimensional characteristic vector value of the i-th feature point, that is, the 384-dimensional regularized log-polar descriptor; T represents the matrix transpose character; N represents the number of feature points.
8. A multimodal remote sensing image matching method described by weighted phase orientation according to claim 1, characterized in that: In Step 8, use the Euclidean distance to measure the descriptors of the feature points, so as to complete the nearest neighbor matching, and eliminate outliers through the fast sample consensus algorithm.
9. The multimodal remote sensing image matching method described by weighted phase orientation according to claim 1, wherein: It also includes step 9 of evaluating the matching effect of multimodal remote sensing images using correct homologous points. In step 9, the root mean square error of the solved homologous points and the number of homologous point pairs are used to quantitatively test the matching accuracy.
Citation Information
Patent Citations
Self-adaptive weak texture remote sensing image registration method for descriptor neighborhood
CN112288784A
Multi-modal remote sensing image feature extraction method based on neural network
CN113313002A