An unsupervised unmanned aerial vehicle visual positioning method based on scale constraint and feature consistency

By constructing a scale-receptive field coupling mapping model and a two-way matching reliability assessment, the problems of positioning accuracy and adaptability of UAV visual geolocation in large-scale satellite imagery were solved, enabling UAVs to achieve autonomous positioning and navigation in complex environments.

CN122115567APending Publication Date: 2026-05-29BEIJING UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF TECH
Filing Date
2026-02-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing UAV visual geolocation methods suffer from reduced positioning accuracy and insufficient adaptability in large-scale satellite imagery, especially in environments with limited GNSS signals. Furthermore, their reliance on complex preprocessing and manual annotation leads to high computational complexity.

Method used

By constructing a scale-receptive field coupling mapping model, determining the validity of features, and performing bidirectional matching reliability calculation and consistency constraint evaluation, unsupervised localization of UAV aerial images in large-scale satellite images is achieved, avoiding scale mismatch and mismatch.

Benefits of technology

It improves the accuracy and stability of UAV positioning without the need for manual annotation and complex preprocessing, making it suitable for autonomous navigation in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115567A_ABST
    Figure CN122115567A_ABST
Patent Text Reader

Abstract

The application discloses a kind of non-supervised unmanned aerial vehicle visual positioning methods based on scale constraint and feature consistency, it is related to unmanned aerial vehicle visual positioning technical field.First, the imaging parameters of aerial image and satellite image are used to establish unified ground physical scale mapping relationship, construct scale-receptive field coupling model, according to the effective geographical coverage area of aerial image Effective feature validity determination threshold is set, so as to obtain the effective feature that meets geographical scale constraint;Second, based on the effective aerial image features and effective satellite image features screened, the bidirectional matching relationship of aerial image pointing to satellite candidate area and satellite candidate area pointing to aerial image is respectively constructed, the first matching confidence and the second matching confidence are calculated, and the consistency constraint evaluation is carried out on the bidirectional matching result.The application can effectively improve the matching stability and positioning reliability of cross-view visual positioning, and has strong engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and remote sensing image processing technology, and in particular to an unsupervised UAV visual localization method based on scale constraints and feature consistency. Background Technology

[0002] Visual geolocation for unmanned aerial vehicles (UAVs) estimates their current location by matching aerial images taken by the UAV with geotagged satellite reference images. This technology has broad application prospects in urban navigation, military reconnaissance, and agricultural monitoring. Traditional UAV positioning primarily relies on the Global Navigation Satellite System (GNSS), but its accuracy significantly decreases or even fails in environments with limited GNSS signals, such as urban areas, dense forests, and indoors. Visual geolocation, as an absolute position estimation method, offers advantages such as not accumulating errors over time and adaptability to complex environments.

[0003] Existing visual geolocation methods are mainly divided into two categories: feature point matching and feature block matching. The former achieves localization by extracting and matching local feature points, while the latter maps the image to a feature space for similarity comparison. However, these methods often perform poorly when faced with large-scale differences, viewpoint changes, and appearance distortions between aerial images and satellite images. In existing technologies, to achieve matching and localization of UAV aerial images in large-scale satellite images, satellite images are typically pre-segmented or cropped to construct a candidate image library, and then region-by-region matching is performed based on this candidate image library. This type of method divides the original large-scale satellite image into multiple local regions during the preprocessing stage, resulting in the overall spatial structure information of the original image being dispersed across different regions. Furthermore, when the number of candidate regions is large, it significantly increases the complexity of feature storage and matching computation. In addition, most visual geolocation methods use supervised learning for model training, relying on manually labeled precise location information or matching relationships as supervision signals during training. This leads to a strong dependence of the model on specific data distributions and labeling formats, limiting the model's adaptability when the imaging viewpoint, scale range, or scene type changes in practical applications. Therefore, there is an urgent need in this field for a visual geolocation method that can directly locate the position of aerial images in large-scale satellite images without supervision and without complex preprocessing. Summary of the Invention

[0004] To overcome the shortcomings of the prior art, this invention provides an unsupervised UAV visual localization method based on scale constraints and feature consistency.

[0005] The technical solution adopted in this invention is: an unsupervised UAV visual localization method based on scale constraints and feature consistency, the method comprising the following steps:

[0006] S1. Scale-Constrained Feature Validity Marking: By acquiring the imaging parameters of UAV aerial images and satellite images, a scale mapping relationship is constructed between the aerial and satellite images. Based on this scale mapping relationship, the aerial and satellite images are aligned in a unified ground physical scale domain. Furthermore, based on the receptive field parameters corresponding to each feature layer, a scale-receptive field coupling mapping model is constructed to map the receptive field size of the feature layer to the geospatial coordinate system, thereby obtaining the coverage range of each feature layer in geospatial space. The aerial image is input into a multi-level feature extraction network to obtain spatial feature responses. A feature validity judgment threshold is set according to the relationship between the feature coverage range and the effective geographical coverage area of ​​the aerial image. When extracting satellite image features, a feature validity judgment operation is performed. When the coverage range of a single pixel in the satellite image features is... The coverage area at a uniform ground physical scale is no larger than the effective area threshold. At that time, then the The feature generates a corresponding valid feature label and determines the feature as a valid feature.

[0007] S2. Visual Localization Based on Consistent Matching of Effective Features: First, based on the effective aerial image features and effective satellite image features selected in step S1, the aerial image features are used as query features, and the features of each candidate region in the satellite image are used as response features to calculate the first matching confidence score between the aerial image and the satellite candidate region. Then, the satellite candidate region features are used as query features, and the aerial image features are used as response features to calculate the second matching confidence score between the satellite candidate region and the aerial image. The matching confidence score is calculated based on a feature similarity function and normalized to form a probability distribution to reflect the correspondence between different matching directions. After obtaining the first and second matching confidence scores, a joint matching confidence score is constructed, and a bidirectional consistency constraint evaluation is performed on the results of the two matching directions. When the matching confidence scores of a candidate region in both matching directions simultaneously meet the preset joint confidence score threshold, and the difference in confidence scores between the two directions is within an allowable range, the candidate region is determined to be a valid candidate region, and the region with the highest confidence score among all valid candidate regions is selected as the final localization region.

[0008] Compared to previous UAV visual positioning methods, this invention constructs a scale-receptive field coupling mapping model to determine feature validity at a unified ground physical scale, ensuring that features participating in the matching calculation have clear geographical coverage significance and avoiding interference from scale-mismatched features in the positioning results. It achieves adaptive feature selection at the feature level through an effective feature labeling mechanism; and it suppresses mismatched regions caused by unidirectional high similarity through bidirectional matching confidence calculation and consistency constraint evaluation mechanisms. Furthermore, it completes the positioning of UAV aerial images in large-scale satellite images without the need for manual data annotation, demonstrating strong applicability and stability.

[0009] The beneficial effects of this invention are:

[0010] This invention constructs a scale-receptive field coupling mapping model to determine feature validity at a unified ground physical scale, ensuring that features participating in the matching calculation have clear geographical coverage significance and avoiding interference from scale-mismatched features in the positioning results. Through a feature effective labeling mechanism, it achieves adaptive filtering at the feature level. Furthermore, through a bidirectional matching confidence calculation and consistency constraint evaluation mechanism, it suppresses mismatched regions caused by unidirectional high similarity. This invention achieves reliable positioning of UAV aerial images in large-scale satellite imagery without introducing additional manual annotations or external prior information, demonstrating strong versatility and practical value. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings with creative effort.

[0012] Figure 1 This is a schematic diagram of the overall process according to an embodiment of the present invention;

[0013] Figure 2 This is a schematic diagram of the visual matching results in an embodiment of the present invention.

[0014] Figure 3 This is a schematic diagram of the visual matching results in an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art are within the scope of protection of the present invention.

[0016] According to an embodiment of the present invention, a UAV visual localization method based on scale-constrained feature validity determination and bidirectional consistency evaluation can be applied in practical applications as follows: Figure 1 As shown, the structure of the present invention is deployed, including:

[0017] A scale-adaptive deep convolutional feature extraction method is used. A single drone aerial image and a large-scale satellite image are input, and their features are extracted by a scale-adaptive feature extraction network. By calculating the receptive field size of each layer and comparing it with the size of the aerial image, the most suitable feature extraction layer is adaptively selected to obtain the feature maps of the aerial image and the satellite image for matching.

[0018] A feature consistency-based matching module receives feature maps from aerial images and satellite images for matching, and then outputs high-confidence candidate regions.

[0019] To facilitate understanding of the above technical solutions of the present invention, the following detailed description of the above technical solutions of the present invention will be provided through actual deployment and application.

[0020] Firstly, to eliminate the large scale difference between UAV aerial images and satellite images, the aerial images need to be appropriately processed to ensure that the two types of images are at a uniform scale. To ensure the accuracy of the image adjustment ratio, a scaling factor based on the ratio between the UAV aerial images and satellite images can be set. This invention determines the Ground Physical Scale (GSD) based on the flight altitude H (m), camera pixel size d (µm), and focal length parameter f (mm) of the UAV aerial images. The formulaic expression of GSD is as follows:

[0021] GSD=

[0022] The uniform ground physical scale of the aerial images was obtained using the GSD calculation formula. Unified ground physical scale with satellite imagery Let F be the scaling factor between aerial and satellite images, and its formulaic expression is:

[0023]

[0024] Then the drone aerial images are scaled using F. To standardize the ground physical scale for aerial images. To standardize the ground physical scale of satellite imagery and ensure consistent proportions with satellite imagery, this improves matching accuracy between two heterogeneous image sources. Definition To adjust the size of the obtained drone aerial image, The initial size of the aerial image can be expressed as follows:

[0025]

[0026] Scale-Adaptive Deep Convolution Feature Extraction: Feature extraction methods based on Convolutional Neural Networks (CNNs) extract features through each layer of the network. Since remote sensing images contain a very large number of pixels, and most feature extraction processes involve convolution operations, the computational and time costs of remote sensing image processing increase with increasing network depth. The hierarchical structure of the receptive field is a core characteristic of CNNs, representing the size of the region mapped from the pixels on the feature map output by each layer to the original image. After preprocessing, aerial images have relatively low resolution, so shallow features are directly used for feature extraction. When the size of our aerial image is much smaller than the receptive field of the feature map of the satellite image obtained through a single CNN layer, directly matching these two types of images would process many meaningless regions, resulting in significant computational waste. Therefore, to address this issue, we designed a scale-adaptive deep convolution feature extraction method. Its core principle is that the receptive field of the selected target layer should match the size of the aerial image as closely as possible. This allows the system to adaptively select a suitable target layer based on the size of the UAV image and extract feature maps from it for subsequent processing. Its formulaic expression is:

[0027]

[0028] in The receptive field size of the k-th layer of effective satellite image features. The size of the receptive field of the (k-1)th layer. Let K be the kernel size of the k-th layer. Let be the step size of the i-th layer. Features are extracted sequentially from shallow to deep layers using the model. For each layer, the size is determined by its corresponding receptive field size. Determine if the minimum size of the aerial image is met. The coverage requirement. When Meanwhile, the model continues to extract deeper features. Once the receptive field of a certain layer exceeds the size of the aerial image, that is... If the feature map is too deep, then extraction stops, and the previous layer is selected as the final feature layer. This strategy can effectively avoid the problem of oversampling caused by excessively deep feature maps, while ensuring that the extracted features have sufficient context awareness to cover the aerial image area.

[0029] Candidate Region Localization: In order to find the best matching region in satellite imagery, this patent quantitatively evaluates the matching confidence level. The formalized expression is as follows:

[0030]

[0031] Where R represents all candidate regions in the satellite image, r represents the feature representation of a specific location in all R, t represents the feature representation of a specific location in the aerial image, and Quality(r,t) is the matching confidence of location r in the image with aerial image t.

[0032] In order to obtain This patent needs to define Quality(r,t), that is, how to measure the confidence of the match between (r,t) and the other two. Let's first assume... , Let denot aerial images and satellite images be the feature representations, respectively. Then, the similarity between the two types of images at a certain feature location is defined as ρ( , ), where ρ(⋅,⋅) is a function used to measure feature similarity (such as cosine similarity). To improve the reliability of matching confidence, the similarity score can be converted into a probability distribution using the softmax function, which is formalized as:

[0033]

[0034] This formula represents the confidence level of a match between the features of a region in an aerial image and every region in the entire satellite image. Similarly, the confidence level of a match between the features of a region in a satellite image and every region in the entire aerial image can be derived, and its formulaic expression is:

[0035]

[0036] Interactive matching mechanisms exist in information retrieval and sentence matching tasks. These mechanisms consider not only the similarity between the query and candidate options but also the candidate options' responsiveness to the query. This bidirectional interactive matching approach is more stable and achieves higher confidence levels compared to ordinary matching (which uses only one direction). Therefore, drawing inspiration from this interactive matching mechanism, its formulaic expression is as follows:

[0037]

[0038] when and When similar, a suitable similarity metric ρ(⋅,⋅) will yield a high value; otherwise, a low value should be obtained. and When a match actually occurs, ρ( , The highest score should be 1, which, ideally, is 1 after activation using the softmax function. The value is also 1. Finally, the confidence scores of the first match in the two matching directions are obtained. Second matching confidence and joint confidence level The evaluation of bidirectional consistency constraints is formalized as follows:

[0039]

[0040]

[0041] definition For the joint matching confidence threshold, The confidence difference threshold for bidirectional matching is used. When the matching confidence of a candidate region in both matching directions simultaneously meets the preset consistency constraint threshold, the candidate region is determined to be a valid candidate region with high matching confidence. Finally, the candidate region with the highest confidence is selected as the final location region.

[0042] In summary, by utilizing the aforementioned technical solutions, this invention proposes an unsupervised UAV visual localization method based on scale constraints and feature consistency. This framework constructs a novel solution that can directly locate aerial images within large-scale satellite imagery without complex preprocessing. Specifically, in the first stage, this invention designs a candidate region localization method based on scale-adaptive depth convolution feature extraction. By introducing scale-constrained feature validity labeling and a visual localization strategy based on effective feature consistency matching, it effectively overcomes the scale difference problem between aerial and satellite images, significantly improving the accuracy and efficiency of candidate region selection. The proposed method not only achieves high localization accuracy but also maintains good robustness, providing a practical technical solution for autonomous localization and navigation of UAVs in complex environments, filling a research gap in the field of cross-scale visual geolocation.

[0043] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An unsupervised UAV visual localization method based on scale constraints and feature consistency, characterized in that, The method includes the following steps: S1. Scale-Constrained Feature Validity Labeling: First, by combining aligned aerial and satellite images at a unified ground physical scale with information from each feature layer extracted by a pre-trained feature extraction network, a scale-receptive field coupling mapping model is constructed to obtain the coverage range of each feature layer at a unified ground physical scale. Then, based on the relationship between the feature coverage range and the effective geographic coverage area of ​​the aerial image, a feature validity threshold is set. Features of both the aerial and satellite images are extracted, and a feature validity determination operation is performed when extracting satellite image features. When the coverage range of a single pixel in the satellite image features... The coverage area at a uniform ground physical scale is no larger than the effective area threshold. At that time, then the The feature generates a corresponding valid feature label and determines the feature as a valid feature; S2. Visual localization based on effective feature consistency matching: First, based on the effective aerial image features and effective satellite image features selected in step S1, the first matching confidence of the aerial image features pointing to the satellite candidate region and the second matching confidence of the satellite candidate region pointing to the aerial image features are calculated respectively to characterize the correspondence between the aerial image features and the satellite candidate region under different matching directions. Then, the consistency of the two is jointly evaluated to determine the matching stability of the candidate region in the bidirectional matching space. Only when the result of the bidirectional matching of the candidate region simultaneously satisfies the bidirectional consistency constraint condition is it determined as a valid candidate region, and the candidate positioning region in the large-scale satellite image of the UAV aerial image is determined accordingly.

2. The unsupervised UAV visual localization method based on scale constraints and feature consistency according to claim 1, characterized in that, The specific process of step S1 is as follows: S11. The unified ground physical scale (GSD) mentioned in step S1 is determined based on the flight altitude H, camera pixel size d, and focal length parameter f of the UAV aerial image, and its formulaic expression is as follows: GSD= ; The uniform ground physical scale of the aerial images was obtained using the GSD calculation formula. Unified ground physical scale with satellite imagery Let F be the scaling factor between aerial and satellite images, and its formulaic expression is: ; Then, based on the calculated scaling factor F between the aerial and satellite images, the original width and height were... Aerial images are adjusted to obtain aerial images of uniform scale. Its formal expression is as follows: ; definition To determine the coverage area of ​​a single pixel in the k-th layer of effective satellite image features, a scale-receptive field coupled geographic scale mapping relationship is constructed based on the principle that the receptive field of each layer of the convolutional network expands progressively with each layer. Its formal expression is as follows: ; in Let be the size of the coverage area of ​​a single pixel in the corresponding feature of the (k-1)th layer. Let K be the kernel size of the k-th layer. Let i be the step size of the i-th layer. Represents all step sizes from layer 1 to layer k-1. The product of consecutive products; S12. As described in step S12, when the receptive field of a certain layer of features... When the size first exceeds the size of the drone aerial image, the features of the layer above that layer are automatically selected as the final feature representation for candidate region matching, i.e., based on its corresponding... Whether the size meets the size requirement of the area covered by the aerial image. The demand, when When that happens, the previous layer is selected as the final effective feature layer.

3. The unsupervised UAV visual localization method based on scale constraints and feature consistency according to claim 1, characterized in that, The specific process of step S2 is as follows: S21. As described in step S21, match the confidence level. Defined as: ; Where R represents all candidate regions in the satellite image, r represents the feature representation of a specific location in all R, t represents the feature representation of a specific location in the aerial image, and Quality(r,t) is the matching confidence of location r in the image with aerial image t. Subsequently, using aerial image features as query features and satellite candidate region features as response features, a first-direction matching relationship is established between the aerial image and the satellite candidate region. Simultaneously, using satellite candidate region features as query features and aerial image features as response features, a second-direction matching relationship is established between the satellite candidate region and the aerial image. The formula for calculating the first matching confidence of the aerial image features in the global satellite candidate region set is defined as follows: ;in , These are the feature representations of aerial images and satellite images, respectively. The similarity at a certain feature location, By combining the matching confidence in both matching directions, we obtain the bidirectional feature matching relationship, which can be formally expressed as: ; S22. As described in step S21, the first matching confidence scores obtained in the two matching directions in step S21 are... Second Match Confidence and joint confidence level The evaluation of bidirectional consistency constraints can be formalized as follows: ; ;definition For the joint matching confidence threshold, The confidence difference threshold for bidirectional matching is used. When the matching confidence of a candidate region in both matching directions simultaneously meets the preset consistency constraint threshold, the candidate region is determined to be a valid candidate region with high matching confidence. Finally, the candidate region with the highest confidence is selected as the final location region.