Multimodal Remote Sensing Image Template Matching Method and System Based on Multidirectional Tensor Features
Through the multimodal remote sensing image template matching method based on multi-directional tensor features, the problems of insufficient accuracy and low efficiency in multimodal remote sensing image matching are solved, efficient and robust feature point matching is achieved, and the matching accuracy and feature point distribution uniformity of multimodal images are improved.
Patent Information
- Application Number
- CN202410114935.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-01-26
AI Technical Summary
The traditional multimodal remote sensing image matching method has insufficient accuracy, low computational efficiency, uneven distribution of feature points, and difficult to achieve the matching accuracy and efficiency of single-modal images.
The multimodal remote sensing image template matching method based on multidirectional tensor features is adopted, and uniform distribution and efficient matching of feature points are achieved through image normalized grid division, extraction of maximum significant points in the enclosing box, tensor feature calculation, coherent enhancement diffusion and fast Fourier transform similarity measurement.
The accuracy and efficiency of multimodal remote sensing image matching is improved, the distribution uniformity of feature points is better, the matching time is greatly shortened, the accuracy is significantly improved, and the number and distribution uniformity of identified points of the same name are better than traditional methods.
Smart Images

Figure CN117853762B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image processing methods, and particularly relates to a multi-modal remote sensing image template matching method and system based on multi-directional tensor features. Background Art
[0002] Multi-modal remote sensing image (MRSI) matching is a process of identifying corresponding points of images from different sensors and different modalities. Multi-modal image matching can provide technical support for map calibration, precise positioning, feature extraction, target recognition, surface change monitoring, 3D reconstruction, and stereo vision, etc.
[0003] Although traditional feature matching has achieved good accuracy (such as in multi-temporal, infrared image and other modality types), and rich corresponding points have been identified in data types with large modality differences, its point position accuracy is difficult to reach within two-pixel error. Compared with the small data difference between single-modal images (or called homologous images), due to the different imaging mechanisms between multi-modal images, problems such as noise difference, texture structure difference, contrast difference, and resolution difference between images are further amplified, thus limiting the accuracy of the feature matching method itself and making it difficult to achieve the accuracy obtained by single-modal image matching.
[0004] To address the above problems, it is necessary to break through the limitations of feature matching and consider a template matching method with higher matching accuracy advantages. As is well known, there are three main limitations of template matching. First, the flexibility of the template matching method is poor (especially in scale and rotation differences, and preliminary prior correspondence is required to perform better); second, due to its pixel-by-pixel calculation logic, the template matching method has extremely low running efficiency, which greatly restricts the application value of template matching. Third, the uniformity of the distribution of feature points is also an important issue affecting the matching result. Therefore, how to overcome the noise difference, contrast difference, and NRD difference of MRSI while solving the dependence problem of template matching on the initial transformation relationship, improving the low calculation efficiency, and increasing the uniformity of the distribution of feature points will become the core of multi-modal template matching research.
[0005] Therefore, to solve the problems of low calculation efficiency, limited accuracy, and uneven distribution of feature points in template matching. The present invention proposes a lightweight multi-directional tensor feature description model to achieve fast and robust template matching. Summary of the Invention
[0006] The present invention proposes a multi-modal remote sensing image template matching method based on multi-directional tensor features to solve the problems of low calculation efficiency, limited accuracy, and uneven distribution of feature points in template matching.
[0007] The technical solution adopted by the present invention is as follows: The multi-modal remote sensing image template matching method based on multi-directional tensor features mainly includes feature point extraction, descriptor construction, and similarity measurement, and specifically includes the following steps:
[0008] Step 1, Image normalization grid division and grid point extraction: Calculate the number of grids in the image by dividing the number of input feature points by the boundary of the image. Then, calculate the bounding box boundaries of each grid point according to the number of grids in the vertical and horizontal directions of the image;
[0009] Step 2, Extraction of the maximum significant point within the bounding box: Establish a square bounding box for each feature point, calculate the corner feature intensity map of the image within this bounding box, and take the coordinates of the maximum pixel intensity value within the bounding box as the key feature point of this feature point;
[0010] Step 3, Tensor feature calculation: Use tensor features to represent the size and direction information of image edges;
[0011] Step 4, Coherence-enhanced diffusion model optimization: After completing the calculation of the image tensor features, introduce the coherence-enhanced diffusion function to enhance the structural feature information;
[0012] Step 5, Generation of multi-directional tensor features: Generate a multi-directional tensor feature map set using the logarithmic polar coordinate parameterization expression as the feature descriptor;
[0013] Step 6, Similarity measurement for realizing fast template matching: Obtain the feature descriptors of all key feature points on the reference image and the image to be matched, and use a fast Fourier transform model to perform similarity measurement operations until all feature point matches are completed.
[0014] Further, the specific implementation method of Step 1 is as follows;
[0015] Calculate the boundaries of the bounding box of each grid point according to the number of grids in the vertical and horizontal directions of the image. Its mathematical expression is as shown in (1):
[0016]
[0017] In formula (1), Box B represents the boundary length of the bounding box; represents rounding up; max represents the maximum value function; d1 represents the length of the image; d2 represents the width of the image.
[0018] Further, in Step 2, first construct a Box B ×Box B square bounding box for each feature point. Box BDenote the boundary length of the bounding box; then calculate the corner feature intensity map of the image within the bounding box, and take the coordinates of the maximum pixel intensity value within the bounding box as the final feature point of the feature point, and its calculation formula is as shown in Equation (2):
[0019]
[0020] In Equation (2), P′ i represents the maximum significant point within the bounding box of the i-th grid point, represents the set of Harris feature points extracted within the i-th bounding box; f Harris (·) represents the Harris feature extraction function; ↑ represents the set of feature points sorted from top to bottom in
[0021] Furthermore, in step 3, the specific implementation method for completing the tensor feature calculation is as follows;
[0022] First, perform a second-order gradient calculation on the image. Use the Sobel template [-1, 0, 1; -2, 0, 2; -1, 0, 1] to obtain the second-order gradient amplitudes of the image in the horizontal and vertical directions respectively, and its calculation equation is as shown in Equation (3):
[0023]
[0024] In Equation (3), is the first-order gradient result of the image; L(x, y) is the gray value of the image; σ is the Gaussian standard deviation; is the template of the Sobel operator in the horizontal direction; is the template of the Sobel operator in the vertical direction; perform a second-order gradient calculation on the first-order gradient, and the equation is as shown in Equation (4):
[0025]
[0026] In Equation (4), L(x, y) xx and L(x, y) yy respectively represent the sum of the squares of the second-order gradients in the x direction and the y direction; L(x, y) xy represents the trace of the second-order gradient of the image, represents the convolution operator;
[0027] Therefore, the structure tensor representation is as shown in Equation (5):
[0028]
[0029] In Equation (5), G σ is the Gaussian kernel function with a standard deviation of σ; L(x, y) yxand L(x,y) xy The results are the same. At the same time, after the image tensor calculation, the feature vectors parallel and orthogonal to the tensor feature T(x,y) are calculated respectively, which are recorded as: V p (x,y) and V o (x,y).
[0030] Furthermore, the mathematical definition of the coherent enhancement diffusion function is shown in formula (6):
[0031]
[0032]
[0033] In formula (6), D represents the tensor matrix after coherent enhancement diffusion, V p (x,y) and V o (x, y) represents the eigenvectors parallel and orthogonal to the tensor feature, α represents the small difference value, λ1 represents the eigenvector in the direction parallel to the image gradient; λ2 represents the eigenvector in the direction orthogonal to the image gradient, k and m represent constants, and e represents an exponent; C m represents a constant to correct for the bias in the original Perona-Malik diffusion function.
[0034] Furthermore, the specific implementation of step 5 is as follows;
[0035] First, the enhanced diffusion tensor matrix calculated in the diffusion model optimization step is decomposed in the x-direction and the y-direction, and the decomposed features are recorded as CoT x and CoT y , and then the feature image is directional rotated, the rotation range is [0~π], the rotation interval angle is π / o, and after the rotation is completed, a fast Fourier transform is performed to further filter the detail noise. The final mathematical expression is shown in formula (8):
[0036]
[0037] In formula (8), F o (x, y) represents the enhanced diffusion tensor feature of the oth layer; o represents the number of layers of the directional tensor feature calculation; cos(.) and sin(.) are trigonometric function expressions; FFT represents the Fourier transform function; ||·|| represents the absolute value operator. At this point, the construction of the multi-directional tensor feature set is completed;
[0038] Finally, the coherence-enhanced diffusion tensor feature F calculated in formula (8) is o The feature combination of (x, y) forms a multi-channel feature map, and its mathematical formula is defined as shown in formula (9):
[0039] MoTF(x,y)={Fi (x,y)}, i = [1, … o] (9)
[0040] In Equation (9), MoTF(x, y) represents the finally calculated multi-directional tensor feature map.
[0041] Furthermore, the specific implementation method of the similarity measurement operation in Step 6 is as follows;
[0042] For the multi-directional diffusion tensor feature maps of the image to be registered and the reference image obtained from Step 5, use the fast Fourier transform to convert the images from the spatial domain to the frequency domain respectively. Then, multiply the Fourier transform results of the image to be registered with the Fourier transform results of the reference image in terms of spectral values. Next, perform the inverse Fourier transform on the result of the spectral value multiplication to convert it back to the spatial domain. Finally, find the position of the maximum matching response value in the result of the inverse Fourier transform.
[0043] Furthermore, Step 6 also includes using the fast sample consensus method to eliminate gross errors, complete the multi-modal remote sensing image matching, and perform accuracy verification.
[0044] The present invention also provides a multi-modal remote sensing image template matching system based on multi-directional tensor features, including the following modules:
[0045] Image normalization grid division and grid point extraction module, which is used to calculate the number of grids in the image by dividing the number of input feature points by the boundary of the image. Then, according to the number of grids in the vertical and horizontal directions of the image, calculate the bounding box boundaries of each grid point.
[0046] Maximum significant point extraction module within the bounding box, which is used to establish a square bounding box for each feature point, calculate the corner feature intensity map of the image within this bounding box, and take the coordinates of the maximum pixel intensity value within the bounding box as the key feature point of this feature point.
[0047] Tensor feature calculation module, which is used to represent the size and direction information of the image edge using tensor features.
[0048] Coherence-enhanced diffusion model optimization module, which is used to introduce a coherence-enhanced diffusion function after completing the calculation of the image tensor features to enhance the structural feature information.
[0049] Multi-directional tensor feature generation module, which is used to generate a multi-directional tensor feature map set using the logarithmic polar coordinate parameterization expression as a feature descriptor.
[0050] Similarity measurement module for fast template matching, which is used to obtain the feature descriptors of all key feature points on the reference image and the image to be matched, and adopt a fast Fourier transform model to perform similarity measurement operations until all feature point matches are completed. Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0051] The multi-modal remote sensing image template matching method and system based on multi-directional tensor features proposed by the present invention aims to solve problems such as low computational efficiency, limited accuracy, and uneven distribution of feature points in template matching. This refined template matching method is based on the Joint Multi-orientation Tension Feature (MOTF for short) and includes the following three main steps: feature point extraction, descriptor construction, and similarity measurement. First, in the feature point extraction stage, the grid box maximum significant point extraction method is used to ensure the uniform distribution of feature points. Then, in the descriptor construction stage, the generation of the descriptor depends on the introduced multi-directional tensor feature model to form a multi-directional tensor feature map. Finally, for similarity measurement, the fast multi-dimensional Fourier transform model is used to ensure the high efficiency of template matching. The advantage of this innovative method is that it overcomes the problems of low computational efficiency, limited accuracy, and uneven distribution of feature points in traditional template matching. The results show that the method proposed by the present invention can better achieve the matching of multi-modal remote sensing images and is more robust than traditional methods. Description of the Drawings
[0052] Figure 1 : Flowchart of the method of the present invention;
[0053] Figure 2 : Schematic diagram of the extraction of the grid box maximum significant point;
[0054] Figure 3 : Multi-modal remote sensing image dataset. From left to right, each column is respectively multi-temporal optical image, infrared image and optical image, night light image and optical image, SAR image and optical image, navigation map and optical image, optical image and depth image;
[0055] Figure 4 : Matching results of five algorithms in multi-temporal image types.
[0056] Figure 5 : Matching results of five algorithms in infrared image and visible light image types.
[0057] Figure 6 : Matching results of five algorithms in night light image and visible light image types.
[0058] Figure 7 : Matching results of five algorithms in SAR image and visible light image types.
[0059] Figure 8 : Matching results of five algorithms in electronic navigation map and visible light image types.
[0060] Figure 9: Matching results of five algorithms in point cloud depth maps and visible light image types. Detailed implementation method
[0061] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0062] Please refer to Figure 1 the flowchart. The present invention provides a multi-modal remote sensing image matching method based on multi-directional tensor features, including the following steps:
[0063] Step 1: Perform image normalization grid division and grid point extraction. Calculate the number of grids in the image by dividing the number of input feature points by the boundary of the image. Then, calculate the bounding box boundaries of each grid point according to the number of grids in the vertical and horizontal directions of the image;
[0064] First, calculate the number of grids in the image by dividing the number of input feature points by the boundary of the image. Then, calculate the boundaries of the bounding box of each grid point according to the number of grids in the vertical and horizontal directions of the image. The mathematical expression is shown in Formula 1:
[0065]
[0066] In Formula (1), Box B represents the boundary length of the bounding box; represents rounding up; max represents the maximum value function; d1 represents the length of the image; d2 represents the width of the image.
[0067] Step 2: Perform extraction of the maximum significant point within the bounding box, and establish a square bounding box for each feature point. Calculate the Harris feature intensity map of the image within this bounding box, and use the coordinates of the maximum pixel intensity value within the bounding box as the final feature point of this feature point;
[0068] First, construct a square bounding box of Box B ×Box B for each feature point according to Formula (1). Then, calculate the Harris feature intensity map of the image within this bounding box, and use the coordinates of the maximum pixel intensity value within the bounding box as the final feature point of this feature point. The calculation formula is shown in Formula (2):
[0069]
[0070] In Formula (2), Pi i represents the maximum significant point within the bounding box of the i-th grid point. denotes the set of Harris feature points extracted within the \(i\)-th bounding box; \(f\) Harris \(f(\cdot)\) represents the Harris feature extraction function; \(\uparrow\) represents the set of feature points sorted from top to bottom in
[0071] Step 3. To achieve correct matching of the feature points obtained in Step 2, it is necessary to construct feature descriptors for the reference image and the image to be matched. The feature descriptors in the present invention are mainly calculated through multi-directional tensor features. To complete this process, first, tensor feature calculation is performed. Tensors are used to represent the magnitude and direction information of image edges. The tensor model helps to extract the structural features of the image, especially the edge magnitudes that change due to contrast changes while the direction remains unchanged;
[0072] First, second-order gradient calculation of the image is performed. The Sobel template \([-1, 0, 1; -2, 0, 2; -1, 0, 1]\) is used to obtain the second-order gradient amplitudes of the image in the horizontal and vertical directions respectively. Its calculation equation is shown in Equation (3):
[0073]
[0074] In Equation (3), is the first-order gradient result of the image; \(L(x, y)\) is the gray value of the image; \(\sigma\) is the Gaussian standard deviation; is the template of the Sobel operator in the horizontal direction; is the template of the Sobel operator in the vertical direction. Second-order gradient calculation is performed on the first-order gradient, and the equation is shown in Equation (4):
[0075]
[0076] In Equation (4), \(L(x, y)\) xx and \(L(x, y)\) yy respectively represent the sum of squares of the second-order gradients in the \(x\) and \(y\) directions; \(L(x, y)\) xy represents the trace of the second-order gradient of the image. represents the convolution operator.
[0077] Tensors are used to represent the magnitude and direction information of image edges. Although the magnitude of the edge changes due to contrast changes, the direction remains unchanged. Therefore, using the tensor model can help to extract the structural features of the image. The classical structural tensor representation is shown in Equation (5):
[0078]
[0079] In Equation (5), \(G\) σ is the Gaussian kernel function with a standard deviation of \(\sigma\); \(L(x, y)\) yxSame as L(x,y) xy The results are the same. At the same time, after performing image tensor calculations, the eigenvectors parallel and orthogonal to the feature tensor T(x,y) are calculated respectively, denoted as: V p (x,y) and V o (x,y).
[0080] Step 4, after completing the calculation of the image tensor features, it is found that the multi-modal noise still exists, which will affect the registration robustness of the multi-modal images. Based on this, a coherence-enhanced diffusion function is introduced. This function can improve the multiplicative speckle noise interference in multi-modal images, reduce the noise on the uniform regions of the images, retain weak edges, and enhance the structural feature information.
[0081] Coherence-enhanced diffusion calculations are performed on the tensor eigenvalues, and their mathematical definitions are shown in (6) and (7):
[0082]
[0083]
[0084] In equations (6) and (7), D represents the tensor matrix after coherence-enhanced diffusion, α represents a small difference value (α = 0.05) even if there is no preferred direction and k acts as the value of (λ1 - λ2) 2 threshold, λ1 represents the eigenvector in the direction parallel to the image gradient; λ2 represents the eigenvector in the direction orthogonal to the image gradient; T represents the transpose symbol. In addition, e represents the exponent; C m represents a constant to correct the deviation in the original Perona-Malik diffusion function.
[0085] Step 5, first, the enhanced diffusion tensor matrix calculated in the diffusion model optimization step is decomposed in the x direction and the y direction, and the decomposed features are denoted as CoT x and CoT y . Then, the feature image is rotated directionally, the rotation range is [0 to π], and the rotation interval angle is (π / o). After the rotation is completed, a fast Fourier transform is performed to further filter out the detailed noise, and the final mathematical expression is shown in equation (8):
[0086]
[0087] In equation (8), x and y respectively represent the horizontal and vertical coordinate axes of the image, F o (x,y) represents the enhanced diffusion tensor feature of the o-th layer; o represents the number of layers for calculating the directional tensor features (set to 6 in this embodiment); cos(.) and sin(.) are trigonometric function expressions; FFT represents the Fourier transform function; ||·|| represents the absolute value operator. Thus, the construction of the multi-directional tensor feature set is completed.
[0088] Finally, the coherence-enhanced diffusion tensor feature calculated in formula (8), i.e., F o (x,y), is combined with its features to form a multi-channel feature map. Its mathematical formula is defined as shown in formula (9):
[0089] MoTF(x,y) = {F i (x,y)}, i = [1,…o] (9)
[0090] In formula (9), MoTF(x,y) represents the finally calculated multi-directional tensor feature map.
[0091] Step 6: Implement the similarity measurement for fast template matching. A fast Fourier transform model is adopted to complete the similarity measurement of template matching. After completing step 5, the similarity measurement operation needs to be performed on each feature point extracted in step 2 until all feature point matching is completed. At the same time, the fast sample consensus method is used to eliminate outliers for all matched feature point pairs to complete the correct matching of multi-modal remote sensing images.
[0092] Specifically, for the multi-directional diffusion tensor feature maps of the to-be-registered image and the reference image obtained in step 5, the fast Fourier transform is respectively used to convert the images from the spatial domain to the frequency domain. Then, the Fourier transform results of the to-be-registered image are multiplied by the spectral values of the Fourier transform results of the reference image. Next, the inverse Fourier transform is performed on the result of the spectral value multiplication to convert it back to the spatial domain. Finally, the position of the maximum response value of the match is found in the result of the inverse Fourier transform.
[0093] The present invention also evaluates the matching effect of multi-modal remote sensing images. The present invention uses 6 groups of multi-modal remote sensing image test methods to test the performance, and the data set is shown in Figure 3 . For each image pair, four indicators are used to evaluate the performance of the multi-modal remote sensing image matching algorithm. The indicators are: (1) the ratio of correctly matched points (RCM), (i.e., the ratio of the number of correctly matched points to all matching numbers); (2) the matching running time (PT); (3) the root mean square error of correct matching (RMSE); and the uniformity of feature point distribution (standard deviation of box-counting, SDBC). The unit of PT is seconds, the unit of RCM is %, the unit of RMSE is pixels, and SDBC represents the standard deviation. And it is compared with several optimal image matching methods (MI, DLSS, HOPC, and CFOG), and the comparison results are shown in Table 1.
[0094] Table 1 Comparison of several image matching methods
[0095]
[0096] As can be seen from Table 1, in multi-modal remote sensing image data, the algorithm proposed by the present invention can quickly achieve template matching of multi-modal remote sensing images, greatly saving time costs. While ensuring that the matching accuracy is sub-pixel, the matching time consumption is improved by about 49 times compared with the MI algorithm, about 3 times compared with the FSURF algorithm, about 1.4 times compared with the DLSS algorithm, about 3 times compared with the FHOG algorithm, and about 6.5 times compared with the HOPC algorithm. The algorithm proposed by the present invention can achieve robust matching of multi-modal remote sensing images, and can obtain richer matching homologous points. In terms of RCM, it is improved by at least 1 percentage point and at most 20.1 percentage points, with remarkable effects. The distribution uniformity of the homologous points identified by the algorithm proposed by the present invention is better. The average SDBC calculated from the homologous points identified by the method proposed by the present invention is improved by 2.3 times compared with the FAST method, 2.6 times compared with the KAZE method, 1.5 times compared with the SURF method, and 2 times compared with the Harris method. As Figures 4 to 9 can be seen, among the 5 comparison methods, as the modal difference of the MI algorithm increases, the number of matching homologous points gradually decreases; the DLSS and HOPC algorithms have poor matching results in weak texture areas (see Figure 7 ); the CFOG algorithm has better results than the other three methods, but the distribution uniformity of the extracted homologous points is inferior to the method proposed by the present invention. Generally speaking, the algorithm proposed by the present invention can achieve the best results.
[0097] The embodiment of the present invention also provides a multi-modal remote sensing image template matching system based on multi-directional tensor features, including the following modules:
[0098] An image normalization grid division and grid point extraction module, which is used to calculate the number of grids in the image by dividing the number of input feature points by the boundary of the image, and then calculate the bounding box boundary of each grid point according to the number of grids in the vertical and horizontal directions of the image;
[0099] A maximum significant point extraction module within the bounding box, which is used to establish a square bounding box for each feature point, calculate the corner feature intensity map of the image within the bounding box, and use the coordinates of the maximum pixel intensity value within the bounding box as the key feature point of the feature point;
[0100] A tensor feature calculation module, which is used to represent the size and direction information of the image edge by using tensor features;
[0101] A coherent enhanced diffusion model optimization module, which is used to introduce a coherent enhanced diffusion function after the image tensor feature calculation to enhance the structural feature information;
[0102] A multi-directional tensor feature generation module, which is used to generate a multi-directional tensor feature atlas by using the logarithmic polar coordinate parameterization expression method as a feature descriptor;
[0103] A similarity measurement module for fast template matching, which is used to obtain the feature descriptors of all key feature points on the reference image and the image to be matched, and adopts a fast Fourier transform model to perform similarity measurement operations until all feature point matching is completed.
[0104] The specific implementation methods of each module correspond to the respective steps, and are not described in this invention.
[0105] It should be understood that the parts not elaborated in detail in this specification all belong to the prior art.
[0106] It should be understood that the above description of the preferred embodiment is relatively detailed, and it should not be considered as a limitation to the protection scope of the invention patent of this invention. Under the inspiration of this invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the protection scope defined by the claims of this invention, and all fall within the protection scope of this invention. The scope of protection claimed in this invention shall be subject to the appended claims.
Claims
1. A multimodal remote sensing image template matching method based on multi-directional tensor features, characterized in that, It includes the following steps: Step 1, Image normalization grid division and grid point extraction: Calculate the number of grids in the image by dividing the number of input feature points by the boundary of the image. Then, calculate the bounding box boundaries of each grid point according to the number of grids in the vertical and horizontal directions of the image. Step 2, Extraction of the largest significant point within the bounding box: Establish a square bounding box for each feature point, calculate the corner feature intensity map of the image within this bounding box, and use the coordinates of the maximum pixel intensity value within the bounding box as the key feature point of this feature point. In step 2, first, a Box is constructed for each feature point according to formula (2). B ×Box B of the square bounding box, where Box B represents the boundary length of the bounding box; Then, calculate the corner feature intensity map of the image within this bounding box, and use the coordinates of the maximum pixel intensity value within the bounding box as the final feature point of this feature point. Its calculation formula is as shown in Equation 2: In formula (2), P′ i represents the maximum significant point within the bounding box of the i-th grid point, and represents the set of Harris feature points extracted within the i-th bounding box; f Harris (·) represents the Harris feature extraction function; ↑ represents the set of feature points sorted from top to bottom; Step 3, Tensor feature calculation: Use tensor features to represent the size and direction information of image edges. Step 4, Coherent enhanced diffusion model optimization: After completing the calculation of image tensor features, introduce a coherent enhanced diffusion function to enhance the structural feature information. Step 5, Generation of multi-directional tensor features: Use the logarithmic polar coordinate parameterization expression to generate a multi-directional tensor feature map set, which is used as a feature descriptor. The specific implementation method of Step 5 is as follows: First, decompose the enhanced diffusion tensor matrix calculated in the diffusion model optimization step in the x and y directions, and denote the decomposed features as CoT x and CoT y , then perform an orientational rotation on the feature image, with the rotation range in [0~π] and the rotation interval angle of π / o. After the rotation is completed, perform a fast Fourier transform to further filter out the detailed noise. The final mathematical expression is shown in Equation (8): In Equation (8), F o (x, y) represents the enhanced diffusion tensor feature of the o-th layer; o represents the number of layers for calculating the orientation tensor feature; cos(.) and sin(.) are trigonometric function expressions; FFT represents the Fourier transform function; ||·|| represents the absolute value operator. Thus, the construction of the multi-directional tensor feature set is completed; Finally, the coherence-enhanced diffusion tensor feature calculated in formula (8), i.e., F o (x, y), is combined to form a multi-channel feature map, and its mathematical formula is defined as shown in formula (9): MoTF(x,y) = {F i (x,y)}, i = [1,…o] (9) In Equation (9), MoTF(x, y) represents the finally calculated multi-directional tensor feature map. Step 6, Similarity measurement for realizing fast template matching: Obtain the feature descriptors of all key feature points on the reference image and the image to be matched, and use a fast Fourier transform model to perform similarity measurement operations until all feature point matches are completed.
2. The multi-modal remote sensing image template matching method based on multi-directional tensor features according to claim 1, wherein: The specific implementation method of Step 1 is as follows; Calculate the boundaries of the bounding boxes of each grid point according to the number of grids in the vertical and horizontal directions of the image. Its mathematical expression is as shown in (1): In formula (1), Box B represents the boundary length of the bounding box; represents rounding up; max represents the maximum value function; d1 represents the length of the image; d2 represents the width of the image.
3. The multimodal remote sensing image template matching method based on multi-directional tensor features according to claim 1, characterized in that: In Step 3, the specific implementation method for completing tensor feature calculation is as follows; First, perform a second-order gradient calculation on the image. Use the Sobel template [-1, 0, 1; -2, 0, 2; -1, 0, 1] to obtain the second-order gradient amplitudes of the image in the horizontal and vertical directions respectively. Its calculation equation is as shown in Equation (3): In Equation (3), is the first-order gradient result of the image; L(x, y) is the gray value of the image; σ is the Gaussian standard deviation; is the template of the Sobel operator in the horizontal direction; is the template of the Sobel operator in the vertical direction; The second-order gradient calculation is performed on the first-order gradient, and the equation is shown in Equation (4): In formula (4), L(x, y) xx and L(x, y) yy respectively represent the sum of squares of the second-order gradients in the x and y directions; L(x, y) xy represents the trace of the second-order gradient of the image, represents the convolution operator; Therefore, the structure tensor representation is as shown in Equation (5): In formula (5), G σ is a Gaussian kernel function with a standard deviation of σ; L(x, y) yx is the same as L(x, y) xy As a result, after performing image tensor calculations, the eigenvectors parallel and orthogonal to the tensor feature T(x, y) are calculated respectively, and are denoted as: V p (x, y) and V o (x, y).
4. The multi-modal remote sensing image template matching method based on multi-directional tensor features according to claim 1, characterized in that: The mathematical definition of the coherent enhanced diffusion function is as shown in Equation (6): In Equation (6), D represents the tensor matrix after coherence-enhanced diffusion, V p (x, y) and V o (x, y) represent the eigenvectors parallel and orthogonal to the tensor feature respectively, α represents a small difference value, λ1 represents the eigenvector in the direction parallel to the image gradient; λ2 represents the eigenvector in the direction orthogonal to the image gradient, k and m represent constants, e represents the exponent; C m represents a constant to correct the deviation in the original Perona-Malik diffusion function.
5. The multimodal remote sensing image template matching method based on multi-directional tensor features according to claim 1, characterized in that: The specific implementation method of the similarity measurement operation in Step 6 is as follows; For the multi-directional diffusion tensor feature maps of the image to be registered and the reference image obtained in Step 5, respectively use the fast Fourier transform to convert the image from the spatial domain to the frequency domain. Then, multiply the Fourier transform results of the image to be registered with the Fourier transform results of the reference image in terms of spectral values. Next, perform an inverse Fourier transform on the result of the spectral value multiplication to convert it back to the spatial domain; finally, find the position of the maximum matching response value in the result of the inverse Fourier transform.
6. The multi-modal remote sensing image template matching method based on multi-directional tensor features according to claim 1, wherein: Step 6 also includes using the fast sample consensus method to eliminate gross errors, complete the multi-modal remote sensing image matching, and perform accuracy verification.
7. A multimodal remote sensing image template matching system based on multi-directional tensor features, characterized in that It includes the following modules: The image normalization grid division and grid point extraction module is used to calculate the number of grids in the image by dividing the number of input feature points by the boundary of the image, and then calculate the bounding box boundaries of each grid point according to the number of grids in the vertical and horizontal directions of the image; The maximum significant point extraction module within the bounding box is used to create a square bounding box for each feature point, calculate the corner feature intensity map of the image within this bounding box, and use the coordinates of the maximum pixel intensity value within the bounding box as the key feature point of this feature point; the specific implementation method is as follows: First, construct a Box for each feature point according to formula (2). B × Box B of the square bounding box, where Box B represents the side length of the bounding box; Then calculate the corner feature intensity map of the image within this bounding box, and use the coordinates of the maximum pixel intensity value within the bounding box as the final feature point of this feature point, and its calculation formula is as shown in Equation 2: In formula (2), P′ i represents the maximum significant point within the bounding box of the i-th grid point, and represents the set of Harris feature points extracted within the i-th bounding box; f Harris (·) represents the Harris feature extraction function; ↑ represents the set of feature points sorted from top to bottom; The tensor feature calculation module is used to represent the size and direction information of the image edge using tensor features; The coherent enhanced diffusion model optimization module is used to introduce a coherent enhanced diffusion function after the image tensor feature calculation to enhance the structural feature information; The multi-directional tensor feature generation module is used to generate a multi-directional tensor feature atlas using the logarithmic polar coordinate parameterization expression as a feature descriptor; the specific implementation method is as follows: First, decompose the enhanced diffusion tensor matrix calculated in the diffusion model optimization step in the x-direction and y-direction, and denote the decomposed features as CoT x and CoT y , then perform an orientation rotation on the feature image, with the rotation range in [0~π] and the rotation interval angle of π / o. After the rotation is completed, perform a fast Fourier transform to further filter out the detailed noise. The final mathematical expression is shown in Equation (8): In Equation (8), F o (x, y) represents the enhanced diffusion tensor feature of the o-th layer; o represents the number of layers for calculating the orientation tensor feature; cos(.) and sin(.) are trigonometric function expressions; FFT represents the Fourier transform function; ||·|| represents the absolute value operator. Thus, the construction of the multi-directional tensor feature set is completed; Finally, the coherence-enhanced diffusion tensor feature calculated in formula (8), i.e., F o (x, y) feature combination forms a multi-channel feature map, and its mathematical formula is defined as shown in formula (9): MoTF(x,y) = {F i (x,y)}, i = [1,…o] (9) In Equation (9), MoTF(x, y) represents the finally calculated multi-directional tensor feature map; The similarity measurement module for fast template matching is used to obtain the feature descriptors of all key feature points on the reference image and the image to be matched, and perform similarity measurement operations using a fast Fourier transform model until all feature point matches are completed.
Citation Information
Patent Citations
End-to-end template matching method for multi-modal image
CN113723447A
Fast and robust multimodal remote sensing image matching method and system
WO2019042232A1