Aerial Image Matching Method Based on the Consistency of Scene Salient Regions

Through aerial image matching method based on the consistency of the scene significance region, deep learning and graph cutting method extract the significant target area, combined with template matching and nearest neighbor method, the accuracy and speed problems of aerial image matching are solved, fast and accurate feature matching is achieved, and a variety of high-level computer vision tasks are supported.

CN114882258BActive Publication Date: 2025-07-29ANHUI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210505619.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2025-07-29
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

The existing aerial image matching methods have low accuracy and long time-consuming, making it difficult to meet the real-time and efficient requirements of computer vision systems.

Method used

Aerial image matching method based on the consistency of the scene significance region is adopted. By calculating the feature matching results between the salient target area in the query aerial image and the matching area of the reference aerial image, and combining the feature matching results of the non-striking area, the salient target area is extracted using deep learning and graph cutting method, and the matching accuracy and efficiency are improved using template matching and nearest neighbor methods.

Benefits of technology

It significantly improves the accuracy and speed of aerial image matching, can perform superiorly on international aerial image datasets than traditional methods, achieves fast and accurate feature matching, and supports high-level computer vision tasks such as three-dimensional reconstruction, image retrieval, visual positioning, object tracking, virtual reality and augmented reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882258B_ABST
    Figure CN114882258B_ABST
Patent Text Reader

Abstract

The present invention discloses an aerial image matching method based on the consistency of scene saliency regions, which includes calculating feature points and feature descriptors in a query aerial image and a reference aerial image, calculating the salient target regions in the query aerial image, finding corresponding matching regions in the reference aerial image for the salient regions of the query aerial image, calculating the feature matching results between the salient regions in the query aerial image and the corresponding matching regions, calculating the feature matching results between the regions outside the salient regions in the query aerial image and the non-salient regions in the reference aerial image, and combining the two matching results to obtain the feature matching results between the two aerial images. The present invention utilizes the salient target regions in the query aerial image to transform the problem of high-resolution aerial image matching into the feature matching problems of salient regions and non-salient regions, which not only avoids the time overhead caused by feature matching in non-overlapping regions but also improves the accuracy of image matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing, computer vision, and artificial intelligence technologies, and particularly relates to an aerial image matching method based on the consistency of scene saliency regions. Background Art

[0002] 3D reconstruction and perceptual computing based on aerial images are one of the important ways to help humans accurately and effectively understand and perceive the physical world. Recovering the 3D model of a scene from images through the structure-from-motion method in the field of multi-view geometry is an important research direction in the fields of computer vision and computer graphics. The aerial image matching method can provide effective 2D feature point information for high-level computer vision systems such as image-based 3D reconstruction and scene perception. Currently, aerial image matching has been widely applied in computer vision and image processing systems, such as image-based 3D reconstruction (Patent: 202010012402.9), image retrieval (Patent: 200910081069), etc.

[0003] Aerial image matching has always been a research issue that has received extensive attention. Currently, the theoretical research on aerial image matching classifies the existing aerial image matching methods into two categories: (1) Aerial image matching based on global features; (2) Aerial image matching methods based on local features.

[0004] The aerial image matching method based on local features realizes image matching by calculating the similarity between local feature descriptors on the query aerial image and the reference aerial image. For example, Lowe et al. proposed "Distinctive image features from scale-invariant keypoints" in 2000. This method calculates the similarity between descriptors through a brute-force matching method, which is very time-consuming and cannot meet the requirements of modern computer vision and image processing systems for the time performance of algorithms. Jiang et al. proposed "Robust Feature Matching Using SpatialClustering With Heavy Outliers" in 2020. This method eliminates incorrect feature matches through spatial clustering. This method is only a method for eliminating incorrect local feature matches and can be used as a post-processing step.

[0005] Currently, relevant domestic patents in the field of aerial image matching include: An image matching method and an image matching device (Patent No.: 201310685205.3). This method takes several seconds to process two images, making it difficult to meet the time efficiency requirements of some real-time processing application systems.

[0006] Problems existing in the field of aerial image matching: (1) The accuracy of existing aerial image matching methods is low; (2) Existing aerial image matching methods are very time-consuming and difficult to meet the real-time requirements of algorithms for computer vision systems. Summary of the Invention

[0007] Object of the Invention: The object of the present invention is to solve the deficiencies existing in the existing aerial image matching technology, and provide an aerial image matching method based on the consistency of scene salient regions. By the consistency of salient regions in aerial images, the aerial image matching problem is transformed into: the feature matching between salient regions and the feature matching between non-salient regions, so as to quickly calculate the local feature matching results between two aerial images, making some high-level computer vision tasks based on aerial image matching possible.

[0008] Technical Solution: An aerial image matching method based on the consistency of scene salient regions of the present invention includes the following steps:

[0009] S1. Calculate the local feature points and corresponding feature descriptors in the query aerial image I l , denoted as {k(x i , y i ), d(k(x i , y i ))}; i ∈ [1, M], where M represents the number of feature points detected from the query image I l ;

[0010] Calculate the local feature points and corresponding feature descriptors in the reference aerial image I r , denoted as {k(x j , y j ), d(k(x j , y j ))}, j ∈ [1, N], where N represents the number of local feature points detected from the reference image I r ;

[0011] S2. Extract the salient object regions in the query aerial image I l : First, calculate the saliency map of the query aerial image I l , and then use the graph cut method to extract the salient object regions from the saliency map, which is Q represents the number of salient object regions; k represents the kth salient object region extracted; (c x , c y ) represents the center coordinates of the kth salient object region; (w, h) represents the width and length of the kth salient object region;

[0012] S3, from the reference aerial image I r The query aerial image I l The significant target region Region l Find a matching region r ;

[0013] S4. Calculate and query the aerial image I l The salient target region Region l and the corresponding matching region r The feature matching results between Matches1;

[0014] S5. Calculate and query aerial image I l The area outside the salient area and the reference aerial image I r Matches2, the feature matching results between non-salient regions;

[0015] S6: Merge (take the union) the matching results of step S4 and step S5 to obtain the query aerial image I l and reference aerial image I r Matches are the feature matching results between them.

[0016] Furthermore, the specific method of calculating local feature points and corresponding feature descriptors using a local feature detection method based on deep learning in step S1 is:

[0017] Step S1.1, pre-training of local feature points: Create corresponding 3D objects and crop pictures of these objects from a single perspective to obtain 2D images. All known local feature points in these 2D images are used to train the deep neural network model.

[0018] Step S1.2, feature point self-labeling: Use ImageNet as the training and test datasets; use synthetic scenes to train a basic feature point detection network model, and use the basic feature point detection network model to extract feature points on the ImageNet dataset, which is feature point self-labeling;

[0019] Step S1.3, joint training: Perform geometric transformation on the image used in the previous step to obtain several image pairs, input the corresponding image pairs into the basic feature point detection network, extract feature points and feature descriptors, and perform joint training to obtain local features based on deep learning, so as to detect local feature points and calculate feature descriptors.

[0020] Furthermore, the detailed steps of obtaining the saliency map in step S2 are:

[0021] S2.1. Calculate the coarse-grained saliency map using a fully connected convolutional neural network (FCNN) and use it as foreground information;

[0022] S2.2. Use the superpixels in the boundary region of the original query aerial image as background seed samples, and calculate the coarse-grained saliency map in the reference aerial image using a propagation method based on non-linear regression;

[0023] S2.3. Use graph Laplacian regularization non-linear regression to combine and optimize the coarse-grained foreground and background saliency maps, so as to detect the precise saliency map.

[0024] Furthermore, the detailed process of step S3 is as follows:

[0025] The method for finding a matching region in the reference aerial image for each saliency region in the query image is: use a fast and accurate template matching method, and the specific calculation process is as follows:

[0026] S3.1. Calculate the histogram of gradient directions of the saliency region l in the query aerial image I and denote it as

[0027] S3.2. Starting from the upper left corner of the reference aerial image I r , search in the X-axis direction with a step size of width w and in the Y-axis direction with a step size of height h to find the candidate matching regions;

[0028] S3.3. Calculate the histogram of gradient directions of the candidate matching regions obtained in step S3.2;

[0029] S3.4. Calculate the sum of absolute differences (SAD) between the histogram of gradient directions of the saliency region and the corresponding histogram of gradient directions of the candidate matching regions;

[0030] S3.5. Take the region with the smallest sum of absolute errors in the reference aerial image I r as the final matching region

[0031]

[0032] Among them, represents the midpoint coordinates, width, and length of the f-th matching region; P represents the number of matching regions found in the reference aerial image I r .

[0033] Further, the specific process of step S4 for calculating the feature matching results between each significant target region and the corresponding matching region is as follows:

[0034] First, use the nearest neighbor method to calculate the initial feature matching, and then eliminate the incorrect feature matching to obtain the final matching result Matches1, which is

[0035]

[0036] where represents the feature points in the non-significant region of the query aerial image I l and represents the feature points in the non-significant region of the reference aerial image I r .

[0037] The above matching process can avoid error accumulation and the feature point matching of non-overlapping region images, greatly improving the time efficiency of the overall algorithm.

[0038] Further, the method for calculating the feature matching outside the significant region in step S5 is as follows:

[0039] First, use the nearest neighbor method to calculate the initial feature matching, and then eliminate the incorrect feature matching to obtain the final matching result Matches2, which is:

[0040]

[0041] where represents the feature points in the non-significant region of the query aerial image I l and represents the feature points in the non-significant region of the reference aerial image I r .

[0042] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0043] (1) The present invention makes full use of the significant target regions in the query aerial image, converts the high-resolution aerial image matching problem into the feature matching problems of the significant region and the non-significant region, which not only avoids the time overhead caused by the feature matching of non-overlapping regions, but also improves the accuracy of image matching.

[0044] (2) The present invention has achieved results significantly superior to traditional methods on the largest international aerial image dataset (Image Matching Benchmark).

[0045] (3) This facial mask can quickly find accurate matching points for the feature points in the query aerial image from the reference aerial image, and the query results can be applied to high-level computer vision systems such as image-based 3D reconstruction, image retrieval, visual positioning, object tracking, virtual reality, and augmented reality.

[0046] (4) To further improve the time efficiency, the present invention adopts a parallel computing method, where each thread is responsible for the feature matching task within a salient region, and the real-time performance of the algorithm is well guaranteed. Brief Description of the Drawings

[0047] Figure 1 is the overall matching flowchart of the present invention;

[0048] Figure 2 is the schematic flowchart of an embodiment of the present invention

[0049] Figure 3 is the original query aerial image of an embodiment of the present invention;

[0050] Figure 4 is the original reference aerial image of an embodiment of the present invention;

[0051] Figure 5 is the final feature matching result of an embodiment of the present invention;

[0052] Figure 6 is the schematic diagram for comparing the effects of the present invention and the prior art solutions. Detailed Description of the Embodiments

[0053] The technical solutions of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0054] As Figure 1 shown, the present invention discloses an aerial image matching method based on the consistency of scene salient regions. The present invention can quickly find corresponding matching points for the feature points in the query aerial image from the reference aerial image. The feature matching results calculated by the present invention can be used as a series of high-level computer vision applications, including: image-based 3D reconstruction, object detection, object tracking, simultaneous localization and mapping, structure from motion, image retrieval, map navigation, augmented reality, virtual reality, and the metaverse.

[0055] Embodiment 1: As Figure 2 shown, the specific matching steps of this embodiment include:

[0056] Step a

[0057] As Figure 3 and Figure 4 shown, input a pair of original aerial images, namely the query aerial image Il with reference aerial image I r . Among them, the aerial image can be collected by a consumer-grade drone device or a high-end drone device.

[0058] Step b

[0059] Detect the feature points in the query aerial image I l and the reference aerial image I r respectively, and calculate the corresponding feature descriptors.

[0060] Denote the feature points in the query aerial image I l as: The feature descriptors are: where represents the feature descriptor of the feature point , M represents the number of feature points in the aerial image I l .

[0061] Denote the feature points in the reference aerial image I r as: The feature descriptors are: where represents the feature descriptor of the feature point , N represents the number of feature points in the aerial image I r .

[0062] Step c

[0063] First, calculate the saliency map in the aerial image I l , then use the graph cut method to extract the salient object region from the saliency map, and denote the extracted salient object region as:

[0064]

[0065] where k represents the number of the extracted salient object regions; (c x , c y ) represents the center coordinates of the k-th salient object region; (w, h) represents the width and length of the k-th salient object region.

[0066] Adopt the template matching method to find a corresponding matching region in the reference aerial image I r for each salient region in Region l , and denote the calculated matching region as:

[0067]

[0068] where ​Denote the midpoint coordinates, width, and length of the f-th matching region; P represents the number of matching regions found in the reference aerial image I r from which

[0069] It should be noted here that: P in Equation (2) is not necessarily equal to Q in Equation (1); P in Equation (2) is equal to Q in Equation (1) if and only if there is a matching region in the reference aerial image I l for each significant region in Region r from which

[0070] Step d

[0071] Calculate the feature matching results in the aerial images I l and I r First, calculate the feature matching between the significant regions and the corresponding matching regions Adopt the K-nearest neighbor method based on local perception hashing to find two candidate matching points for the feature points which are and

[0072] At this point, the following ratio can be calculated:

[0073]

[0074] where represents the feature descriptor of the feature point ; represents the feature descriptor of the feature point ; represents the feature descriptor of the feature point . If then it is considered that and are a pair of correct matching points, otherwise they are incorrect matching points

[0075] According to Equation (3), the corresponding feature matching results can be calculated for all significant regions and matching regions, denoted as:

[0076]

[0077] Similarly, according to Equation (3), the corresponding matching points can also be found for the feature points in the non-significant regions, denoted as:

[0078]

[0079] where represents the feature points in the non-significant regions in the aerial image I l from which Denote the aerial image I r Feature points within the non-significant region in Central Africa

[0080] By combining Matches1 in Equation (4) and Matches2 in Equation (5), the aerial image I can be obtained l and the aerial image I r between the feature matching results

[0081]

[0082] Among them, M is the number of feature points in the aerial image I l ; N represents the number of feature points in the aerial image I r in the number of feature points

[0083] The final matching result of this implementation is as Figure 5 shown, where each connection line represents a pair of correct matching points, and the two ends of the connection line represent the feature points in the query aerial image and the feature points in the reference aerial image respectively

[0084] It can be seen from the above embodiments that the present invention first calculates the feature matching result between the significant regions in the query aerial image and the reference aerial image, and then calculates the feature matching result in the non-significant region, avoiding the errors caused by calculating the feature matching in the significant region and the non-significant region; at the same time, it reduces the number of matching times and reduces the time overhead of the algorithm (for images with 4k resolution

[0085] Such as Figure 6 shown Figure 6 (a) of and Figure 6 (b) of are the matching effect diagrams of the same image by the existing technical solution and the inventive technical solution respectively; the average time consumption of the present invention is 0.021 seconds, and the average matching accuracy is 99.67%). The present invention has achieved significantly superior results to traditional methods on the existing largest aerial image dataset (Image Matching Benchmark), and there are many incorrect matching points in the matching results of the existing technical solution

Claims

1. An aerial image matching method based on the consistency of scene saliency regions, characterized in that: The following steps are involved: S1. Calculate the local feature points and corresponding feature descriptors in the query aerial image I l and denote them as {k(x i , y i ), d(k(x i , y i ))}; i ∈ [1, M], where M represents the number of feature points detected from the query aerial image I l ; Calculate the local feature points and corresponding feature descriptors in the reference aerial image I r and denote them as {k(x j , y j ), d(k(x j , y j ))}, where j ∈ [1, N] and N represents the number of local feature points detected from the reference aerial image I r ; S2. Extract the significant target region in the query aerial image I l : First, calculate the saliency map of the query aerial image I l . Then, use the graph cut method on the saliency map to extract the significant target region, which is Q represents the number of significant target regions; k represents the k-th significant target region extracted; (c x , c y ) represents the center coordinates of the k-th significant target region; (w, h) represents the width and length of the k-th significant target region S3. From the reference aerial image I r find the significant target region Region l in the query aerial image I l and find a matching region Region r ; The detailed process is: S3.

1. Calculate the significant region in the query aerial image I l and denote the histogram of oriented gradients of the significant region as ​ S3.

2. Starting from the upper left corner of the reference aerial image I r search in the X-axis direction with a step size of width w and in the Y-axis direction with a step size of width h to find the candidate matching regions in the reference aerial image; S3.

3. Calculate the gradient direction histogram of the candidate matching area obtained in step S3.2; S3.

4. Calculate the significant region Calculate the sum of absolute differences SAD between the histogram of oriented gradients of S3.

5. Select the area with the smallest sum of absolute errors in the reference aerial image I r as the final matching area Among them, represents the midpoint coordinates, width, and length of the f-th matching region; P represents the number of matching regions found from the reference aerial image I r ; S4. Calculate the query aerial image I l The significant target region Region l And the corresponding matching region Region r The feature matching result Matches1 between them; S5. Calculate the query aerial image I l The feature matching result Matches2 between the area outside the significant area in r and the non-significant area in the reference aerial image I; S6. Merge the matching results of step S4 and step S5 to obtain the feature matching results Matches between the query aerial image I l and the reference aerial image I r ​ 2. The aerial image matching method based on the consistency of scene salient regions according to claim 1, wherein: The specific method of calculating local feature points and corresponding feature descriptors using the local feature detection method based on deep learning in step S1 is: Step S1.1, pre-training of local feature points: Create corresponding 3D objects and crop pictures of these objects from a single viewing angle to obtain 2D images. All known local feature points in these 2D images are used for network training. Step S1.2, feature point self-labeling: Use ImageNet as the training and test datasets; use synthetic scenes to train a basic feature point detection network model, and use the basic feature point detection network model to extract feature points on the ImageNet dataset, which is feature point self-labeling; Step S1.3, joint training: Perform geometric transformation on the image used in the previous step, and input the corresponding image pairs into the basic feature point detection network to extract feature points and descriptors. Joint training is performed to obtain local features based on deep learning, thereby detecting local feature points and calculating feature descriptors.

3. The aerial image matching method based on the consistency of scene salient regions according to claim 1, wherein: The detailed steps for obtaining the saliency map in step S2 are: S2.

1. Use a fully connected convolutional neural network (FCNN) to calculate a coarse-grained saliency map and use it as foreground information. S2.

2. Using the superpixels in the boundary area of the original query aerial image as background seed samples, a saliency map in the reference aerial image is calculated using a propagation method based on nonlinear regression. S2.

3. Using graph Laplace regularized nonlinear regression, the coarse-grained foreground and background saliency maps are combined and optimized to detect accurate saliency maps.

4. The aerial image matching method based on the consistency of scene saliency regions according to claim 1, wherein: The specific process of calculating the feature matching result between each salient target region and the corresponding matching region in step S4 is as follows: First, use the nearest neighbor method to calculate the initial feature matching, and then eliminate the wrong feature matching to obtain the final matching result Matches1, which is Among them, represents the feature points in the non-significant region of the query aerial image I l , and represents the feature points in the non-significant region of the reference aerial image I r .

5. The aerial image matching method based on the consistency of scene saliency regions according to claim 1, wherein: The feature matching method outside the salient area is calculated in step S5 as follows: First, use the nearest neighbor method to calculate the initial feature matching, and then eliminate the erroneous feature matching to obtain the final matching result Matches2, which is: Among them, represents the feature points in the non-significant area of the query aerial image I l , and represents the feature points in the non-significant area of the reference aerial image I r .

Citation Information

Patent Citations

  • Image matching method and image matching device

    CN103617625A

  • Quick robust feature tracking method for large-scale three-dimensional reconstruction

    CN111209965A

  • High-resolution remote sensing image feature matching method

    CN103456022A

  • Hierarchical image matching method

    CN114332510A