Unmanned system scene matching and positioning method in area denial scene

By constructing a semantic topological relationship graph and a graph convolutional network for the coastline scene, the positioning accuracy and reliability issues of unmanned systems in complex scenarios were solved, and the suppression and adaptive adjustment of multi-source interference were achieved, thereby improving positioning accuracy and reliability.

CN120833504AActive Publication Date: 2025-10-24NORTHWESTERN POLYTECHNICAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511325375.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-24
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

In complex coastal and near-ground scenarios, the scene matching and localization method of unmanned systems faces problems such as feature space aliasing caused by multi-source interference, spatiotemporal changes in ground feature features, and cross-domain feature heterogeneity, resulting in insufficient positioning accuracy and reliability.

Method used

We employ a feature consistency representation method that embeds causal semantic knowledge. By constructing a semantic topology graph of the coastline scene and using graph convolutional networks for feature propagation and refinement, we combine inertial navigation data for position calculation and accuracy evaluation, thereby achieving suppression and adaptive adjustment of multi-source interference.

Benefits of technology

It improves the positioning accuracy and reliability of unmanned systems in complex scenarios, effectively suppresses the influence of interference factors such as atmospheric turbulence and sea surface ripples, maintains stable feature matching performance, and improves positioning accuracy and reliability through multi-source information fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833504A_ABST
    Figure CN120833504A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned system scene matching and positioning method in a region denial scene, and relates to the technical field of image matching and positioning. The method comprises the following steps: firstly, processing an unmanned aerial vehicle image and a satellite base map, identifying semantic entities in the map, calculating a relation strength weight, and constructing a causal semantic map; a key point detection mask is generated, a key point set in the unmanned aerial vehicle image and the satellite base map is obtained, and the incidence relation between the key points and the semantic entity is established; calculating semantic association strength between the key points and obtaining enhanced features, screening key point matching pairs, and establishing an initial matching set; screening the initial matching set according to a weighted fusion result to obtain a refined matching set; optimizing the node feature matrix and evaluating the refined matching set to obtain a final matching set; and calculating a deviation vector between the actual position and the estimated position of the unmanned aerial vehicle by using the final matching set to obtain fused position estimation. According to the invention, the scene matching and positioning precision in the area denial scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image matching positioning, in particular to a scene matching positioning method for an unmanned system in a regional denial scenario. BACKGROUND

[0002] High-altitude unmanned systems have built a regular patrol mechanism in many coastal areas due to their wide-area monitoring, remote sensing, and integrated task capabilities, and are widely used in marine monitoring, ecological protection, disaster response, and security operation and maintenance fields, showing important technical support value. However, during the execution of long-distance and long-time coastal cruising tasks, unmanned systems are easily disturbed by complex electromagnetic environments, typical means including satellite navigation signal regional denial, navigation deception, and communication interference, with an influence range of up to hundreds of kilometers. At the same time, due to the error accumulation of airborne inertial navigation systems (for example, the hourly error of pure inertial navigation of some high-altitude unmanned systems can reach 1 nautical mile), in environments where satellite signals are denied or severely disturbed, the positioning accuracy is difficult to meet the needs of continuous operation.

[0003] Scene matching navigation technology has become a feasible solution for achieving high-precision autonomous positioning in a satellite navigation-free environment due to its strong autonomy, good environmental adaptability, high precision, and lower hardware cost. This technology acquires real-time ground images through imaging equipment carried by unmanned systems, matches features with pre-stored satellite base maps, and combines pose information to achieve accurate positioning. The matching accuracy between real-time images and satellite base maps directly determines the accuracy and system stability of the positioning results.

[0004] Currently commonly used scene matching methods mainly include:

[0005] Feature point-based matching methods (such as SIFT, SURF, ORB, etc.), which extract key points in images and generate feature descriptors for matching; region-based matching methods (such as cross-correlation matching, mutual information matching, etc.), which rely on calculating the similarity between regions of two images to achieve matching; and deep learning-based methods (such as SuperGlue, LoFTR, etc.), which use neural networks to extract high-level features and establish matching relationships.

[0006] However, in complex coastal and near-earth scenarios, existing methods still face several technical challenges:

[0007] First, multi-source interference easily leads to feature space aliasing. Images taken from high altitudes are easily disturbed by atmospheric turbulence, sea surface ripple changes, intertidal rock exposure, and other multi-source environmental disturbances, leading to unstable image features. Existing methods rely heavily on data-driven shallow statistical correlations and fail to fully incorporate semantic and topological constraints, limiting their ability to resist interference and effectively extract features. Moreover, they lack effective mechanisms for fusing prior information with image features, making it difficult to construct a robust feature representation.

[0008] Secondly, the spatiotemporal variations in terrain features introduce uncertainty into matching. Coastal topography is subject to dynamic changes due to tides, seasons, and human activity. However, satellite imagery updates are limited, preventing real-time synchronization with on-site conditions. This results in significant spatiotemporal heterogeneity in matching features. Existing algorithms often have fixed structures, making it difficult to adaptively adjust network structures or task coordination strategies based on differences in the spatial distribution of matching images and changes in the temporal dimension.

[0009] Third, cross-domain feature heterogeneity affects the model's generalization capabilities. Imagery captured by unmanned systems differs from satellite imagery in terms of viewing angle, sensor type, and imaging conditions, leading to inconsistent feature distributions. Mainstream supervised learning methods struggle to fully adapt to the demands of multi-view, multimodal data training, lacking hierarchical gradient propagation mechanisms and the ability to incrementally learn domain-invariant features from unlabeled data.

[0010] Therefore, in order to meet the application needs in coastal and complex geographical environments under navigation signal denial conditions, it is urgent to develop a scene matching positioning method with strong robustness and high adaptability to improve the positioning accuracy and reliability of unmanned systems in real complex scenarios. Summary of the Invention

[0011] In response to the problems existing in the prior art, the present invention proposes a scene matching positioning method for an unmanned system in an area denial scenario, comprising the following steps:

[0012] Step 1: Real-time drone images collected by high-altitude unmanned systems Perform geometric correction to obtain the corrected drone image ; For the corrected drone image And pre-stored satellite basemaps Perform equalization processing to obtain the processed drone image and satellite basemaps ; According to the characteristics of the coastline, the processed drone image and satellite basemaps Divided into sea areas , coastline area and land areas Three sub-regions, and assign weights to different regions ;

[0013] Step 2: Identify drone images and satellite basemaps The semantic entities in the causal semantic graph are constructed by calculating the spatial relationship between the semantic entities and calculating the relationship strength weights between the semantic entities. ;

[0014] Step 3: generating a key point detection mask according to the causal semantic graph , the pixel value at a position in the mask representing the priority of the position as a key point; then using the key point detection mask to perform key point detection on the unmanned aerial vehicle image and the satellite base map based on the self-supervised deep learning architecture, to obtain the respective key point sets and of the unmanned aerial vehicle image and the satellite base map , and extracting the depth feature descriptors and of the key points, and finally establishing the association between the key points and the semantic entities ;

[0015] Step 4: calculating the semantic association strength between key points based on the causal semantic graph , constructing a feature enhancement network based on self-attention based on the semantic association strength, and fusing the input features of the feature enhancement network with external environment and time factors to obtain enhanced features and ; using the enhanced features and to construct an optimal transport problem with causal semantic constraints, and obtaining an optimal matching matrix by solving the optimal transport matching problem, and using the optimal matching matrix to filter the key point matching pairs between the unmanned aerial vehicle image and the satellite base map to establish an initial matching set ;

[0016] Step 5: for the initial matching set , performing local neighborhood consistency verification and global structure verification based on causal semantic relationships, and weighting and fusing the local consistency scores obtained by local neighborhood consistency verification, the global consistency scores obtained by global structure verification, and the initial matching scores of the matching pairs, filtering low-confidence matches according to the fusion results to obtain a refined matching set ;

[0017] Step 6: combining the causal semantic graph , constructing an enhanced adjacency matrix for the matching set ; then using the adjacency matrix , applying a multi-layer graph convolution network to optimize the node feature matrix; using the optimized node feature matrix to re-evaluate the key point matching pairs in the matching set , eliminating matches that do not conform to global semantic consistency, to obtain a final matching set ;

[0018] Step 7: According to the final matching set Calculate the deviation vector between the actual position of the UAV and the estimated position ; fuse the deviation vector with the inertial navigation data to obtain the fused position estimate.

[0019] Further, in step 1, the area weight , wherein is the shoreline area weight, is the land area weight, is the sea area weight, and .

[0020] Further, in step 2, the causal semantic graph , wherein is a node set with semantic entities as nodes, is an edge set composed of relationships between two semantic entities with explicit spatial association in the node set , and is an edge weight set composed of strength weights of each relationship in the spatial relationship set .

[0021] Further, in step 2, according to the formula:

[0022]

[0023] Calculate the relationship strength weight between semantic entities and ; wherein the function is a stability function for evaluating the stability of the relationship between semantic entities, the function is a persistence function for evaluating the degree of change of the relationship in the time scale, the function is a recognition function for evaluating the visual recognition and uniqueness of the relationship, the function is a relationship weight function, taking:

[0024]

[0025] wherein is the shoreline area indicator variable, is the land area indicator variable, is the sea area indicator variable.

[0026] Further, in step 3, according to the formula:

[0027]

[0028] Generate a key point detection mask where denotes the guidance mask value at position denotes the priority of the position as a keypoint, is an indicator function that returns 1 when the condition inside the parentheses is true, and 0 otherwise, denotes a semantic entity denotes the spatial region covered by the satellite footprints denotes the importance weight of a semantic entity

[0029] Further, in step 4, the semantic association strength between a keypoint

[0030]

[0031] is calculated based on the causal semantic graph where is an indicator function that returns 1 when the condition inside the parentheses is true, and 0 otherwise; denotes a mapping relationship between a keypoint and a semantic entity denotes a mapping relationship between a keypoint and a semantic entity is the relationship strength weight between a semantic entity and a semantic entity

[0032] Further, in step 4, the optimal transport problem with causal semantic constraints is constructed using the enhanced features and

[0033]

[0034]

[0035]

[0036] where denotes the feature vector of the corresponding keypoint in the enhanced feature denotes the feature vector of the corresponding keypoint in the enhanced feature ​​​​​​​​​​​​​preset coefficient; key points cost matrix between key points optimal matching matrix matching probability of key points in UAV image and key points in satellite base map constraint condition ensure that the sum of matching probabilities of key points meets a preset distribution and .

[0037] Further, in step 5, for key point matching pairs in the initial matching set , the fusion result is calculated according to the formula:

[0038]

[0039] wherein , and are weight coefficients, satisfying , is the initial matching score of the key point matching pair , is the local consistency score of the key point matching pair , is the global structural consistency score of the key point matching pair .

[0040] Further, in step 6, the formula of the node feature matrix is:

[0041]

[0042] wherein and respectively represent the node feature matrix of the i-th layer and the j-th layer, is the degree matrix of the enhanced adjacency matrix , is the learnable weight matrix of the i-th layer, is a nonlinear activation function; the enhanced adjacency matrix is obtained according to the formula:

[0043]

[0044] wherein is the original adjacency matrix constructed based on the matching set ,​​​​​​ is a unit matrix of the same size as matrix is a unit matrix of the same size as matrix is a weight coefficient of the causal relationship, is a causal relationship adjacency matrix.

[0045] Further, in step 6, according to the formula:

[0046]

[0047] get the final matching set , wherein is the L2 norm of the feature vector corresponding to the key point in the node feature matrix , is the L2 norm of the feature vector corresponding to the key point in the node feature matrix , is the L2 norm of the feature vector corresponding to the key point in the node feature matrix is the graph convolution refinement threshold.

[0048] Advantages:

[0049] The present application proposes a scene matching positioning method for unmanned systems in a regional denial scenario, which has the following advantages:

[0050] 1、The present application adopts a feature consistency representation method of causal semantic knowledge embedding, constructs a semantic topological relationship graph of the coastline scene, and performs feature propagation and refinement through a graph convolution network, thereby being able to solve the feature space aliasing and distribution drift problem caused by multi-source interference, realize the stability and consistency of feature expression, and effectively suppress the influence of atmospheric turbulence, sea surface ripple and other interference factors on the matching accuracy.

[0051] 2、The present application realizes accurate identification and matching of coastline features through a key point extraction and matching strategy enhanced by causal semantics. By explicitly modeling the semantic relationship between key points, stable feature matching performance can be maintained even in the case of environmental interference or changes in viewing angle, greatly improving the matching accuracy and reliability.

[0052] 3、The present application adopts a multi-source information fusion position solution and accuracy evaluation method, optimally fuses the scene matching result with inertial navigation data, and introduces a reliability evaluation mechanism, so that the system can evaluate the positioning accuracy in real time and perform adaptive adjustment, improving the accuracy and reliability of positioning.

[0053] Additional aspects and advantages of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood by the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0054] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description, taken in conjunction with the following drawings of which:

[0055] Figure 1 Flow chart of the scene matching positioning method of the regional denial scene unmanned system. DETAILED DESCRIPTION

[0056] Embodiments of the present application are described below in detail, which are exemplary and intended to explain the present application, and cannot be understood as a limitation of the present application.

[0057] The present embodiment is aimed at the aerial unmanned aerial vehicle scene matching positioning demand in the coastal regional denial scene, and proposes a corresponding positioning method, including the following steps:

[0058] Step 1: Multi-source heterogeneous data preprocessing and adaptive enhancement:

[0059] The aerial unmanned aerial vehicle collects images of the coastline region in real time , based on the unmanned aerial vehicle attitude parameters and the camera internal parameters , geometric correction is performed on the unmanned aerial vehicle images to eliminate perspective distortion, and the corrected unmanned aerial vehicle images are obtained; multi-scale adaptive histogram equalization processing is performed on the corrected unmanned aerial vehicle images and the pre-stored satellite base map to improve the image contrast, and the processed unmanned aerial vehicle images and satellite base map are obtained; and according to the characteristics of the coastline, the processed unmanned aerial vehicle images and satellite base map are divided into sea area , coastline area and land area three sub-regions, and different region weights are given.

[0060] In the present embodiment, the input of step 1 is the unmanned aerial vehicle images collected by the aerial unmanned aerial vehicle in real time and the pre-stored satellite base map , the unmanned aerial vehicle attitude parameters and the camera internal parameters ;

[0061] The processing process is:

[0062] Step 1.1: Based on the unmanned aerial vehicle attitude parameters and the camera internal parameters , geometric correction is performed to eliminate perspective distortion:

[0063]

[0064] In this step, the geometric correction function The perspective transformation matrix is calculated by using the attitude parameters (roll angle, pitch angle and yaw angle) of the unmanned aerial vehicle and the camera intrinsic matrix, and the oblique image is corrected to an image with an approximate vertical viewing angle, so that it is more consistent with the observation angle of the satellite base map and has the same size, thereby laying a foundation for subsequent matching. The geometric correction function .

[0065] Step 1.2: Multi-scale adaptive histogram equalization is used to improve the contrast of the image:

[0066] ,

[0067] In this step, Multi-scale adaptive histogram equalization algorithm is used, which performs histogram equalization at multiple scales and then adaptively fuses the results of each scale. The parameter controls the degree of local contrast enhancement, and appropriate increase can improve the visibility of details; controls the global brightness adjustment, which is used to compensate for different lighting conditions; is a detail enhancement factor used to enhance texture features. This processing can effectively deal with the imaging differences caused by weather conditions and changes in lighting in the coastline scene. The multi-scale adaptive histogram equalization algorithm is prior art in the field.

[0068] Step 1.3: According to the characteristics of the coastline, the images processed in steps 1.1 and 1.2 are divided into three sub-regions: sea area , coastline area and land area , and different weights are assigned:

[0069]

[0070] Among them, is the weight of the coastline area, is the weight of the land area, is the weight of the sea area, , which reflects the importance of the coastline area as the main matching target. In this embodiment, is 0.8, is 0.6, is 0.3. This region division is achieved with image segmentation techniques, which identify the boundaries of different regions by analyzing color, texture, and spatial continuity features of the image. The weight assignment reflects the importance of each region in the matching process, with particular emphasis on the shoreline region, as shoreline morphology is generally highly time-stable and visually recognizable.

[0071] The final output of this step is the pre-processed UAV image , satellite base map , and region weight vector .

[0072] Step 2: Constructing the coastal scene causal semantic graph:

[0073] Identify semantic entities in the pre-processed UAV image and satellite base map processed in step 1, and establish spatial relationships between semantic entities while calculating the strength weights of the relationships between semantic entities to obtain a complete causal semantic graph:

[0074]

[0075] where is a node set with semantic entities as nodes, is an edge set composed of relationships between two semantic entities with explicit spatial association in the node set , and is an edge weight set composed of the strength weights of each relationship in the spatial relationship set .

[0076] In this embodiment, the input of step 2 is the pre-processed UAV image , satellite base map , and region weight vector obtained in step 1;

[0077] The processing process is as follows:

[0078] Step 2.1: Identify the main semantic entities in the coastal scene:

[0079]

[0080] In this step, the function uses a fully convolutional neural network to perform pixel-level semantic segmentation on the input UAV image , satellite base map , and identify various key features. The function first extracts multi-scale features, then captures global information through a context aggregation module, and finally generates a high-precision segmentation map through a decoder to obtain the node set including semantic entities such as shoreline edge, reef, building, etc. These semantic entities are the basic nodes for building the causal semantic graph.

[0081] Step 2.2: Establish spatial relationship between semantic entities:

[0082]

[0083] Function For analyzing the geometric and spatial configuration relationship between semantic entities, when there is a clear spatial association between semantic entities and , an edge is established. Spatial relationships include: adjacent relationship (such as the wharf adjacent to the shoreline), containment relationship (such as the reef group contained by the bay), parallel relationship (such as the road parallel to the shoreline) and intersection relationship (such as the bridge intersecting the river). These relationships reflect the spatial layout of objects in the scene, providing spatial constraints for subsequent matching.

[0084] Step 2.3: Calculate the relationship strength weight based on regional characteristics and prior knowledge:

[0085]

[0086] This formula calculates the strength weight of the relationship between semantic entities, considering multiple factors:

[0087] The function is a stability function, which in this embodiment takes , where is the stability coefficient of semantic entity , is the stability coefficient of semantic entity . In this embodiment, the stability coefficients of fixed ground objects such as buildings, roads, bridges, etc. are set to 0.9, the stability coefficients of semi-fixed ground objects such as shorelines, large reefs, etc. are set to 0.7, the stability coefficients of changing ground objects such as vegetation, sand, etc. are set to 0.5, and the stability coefficients of dynamic ground objects such as ships, water ripples, etc. are set to 0.2. The function is used to evaluate the stability of the relationship between semantic entities, for example, the relationship between the shoreline and the reef will fluctuate with the tide, so the stability is low; while the relationship between the shoreline and the lighthouse is relatively stable.

[0088] The function is a persistence function, which in this embodiment takes , where is the time variation rate of semantic entity , is the time variation rate of semantic entity In this embodiment, the time change rate of semantic entities such as permanent structures, such as concrete buildings and rocky coasts, is 0.05, that is, an annual change rate of 5%. The time change rate of semantic entities such as long-term stable structures, such as asphalt roads and large docks, is 0.15, that is, an annual change rate of 15%. The time change rate of semantic entities such as medium-term changing features, such as vegetation cover and sandy coasts, is 0.35, that is, an annual change rate of 35%. The time change rate of semantic entities such as short-term changing features, such as seasonal vegetation and intertidal zones, is 0.60, that is, an annual change rate of 60%. The time change rate of semantic entities such as high-frequency changing features, such as water surface conditions and temporary facilities, is 0.85, that is, an annual change rate of 85%. The function is used to evaluate the degree to which the relationship varies over time scales, taking into account changes in time scales, such as seasonal changes that have a greater impact on vegetation but little impact on buildings.

[0089] Function is the recognition function, in this embodiment, ,in Semantic Entity The recognition Semantic Entity In this embodiment, the recognition degree of semantic entities such as unique landmarks such as lighthouses and buildings with special shapes is 0.45, the recognition degree of semantic entities such as linear features such as roads, rivers, and coastlines is 0.35, the recognition degree of semantic entities such as planar features such as building complexes and forests is 0.25, the recognition degree of semantic entities such as texture features such as farmland and grassland is 0.15, and the recognition degree of semantic entities such as uniform areas such as water surfaces and open spaces is 0.05. The function is used to evaluate the visual recognition and uniqueness of the relationship. For example, the relationship between unique landmarks (such as a bay with a unique shape and a lighthouse) is highly recognizable, and the constraints are , that is, the total recognition does not exceed 0.9.

[0090] The values ​​of the above three functions are set according to the stability, persistence and recognition between specific semantic entities.

[0091] The function is a relationship weight function, which can adjust the relationship strength weight according to the regional weight vector determined in step 1, so as to give priority to the semantic entity relationship of the shoreline area. In this embodiment, ,in is the shoreline area indicator variable, is the land area indicator variable, is the indicator variable for sea area, The value is 0, 0.5 or 1. When there are semantic entities in a region and When the region indicator variable is 1, when there are only semantic entities in a region and When there is one of the semantic entities in a region, the region indicator variable is 0.5. and , the area indicator variable is 0.

[0092] The above four factors are multiplied together to obtain the comprehensive relationship strength weight. The higher the value, the more important the relationship is, and the greater the weight in subsequent feature matching.

[0093] Step 2.4: Construct a complete causal semantic graph:

[0094]

[0095] This step will obtain the node set (semantic entity), edge set (spatial relationships) and edge weight sets (relationship strength) into a complete causal semantic graph As the output of this step, the graph is stored as a data structure, either an adjacency matrix or an adjacency list, containing node information, edge connections, and edge weights. This semantic graph serves as the core knowledge representation guiding feature extraction and matching in subsequent steps.

[0096] The final output of this step is the causal semantic graph .

[0097] Step 3: Causal semantics-guided keypoint detection and feature extraction:

[0098] According to the causal semantic graph , generate a key point detection mask of the same size as the image in step 1, the pixel value of a certain position in the mask represents the priority of the position as a key point; then use the key point detection mask to detect the UAV image separately and satellite basemaps Perform key point detection based on self-supervised deep learning architecture to obtain drone images and satellite basemaps The corresponding key point sets and , and extract the deep feature descriptor of the key points and , and finally establish the association relationship between key points and semantic entities .

[0099] In this embodiment, the input of step 3 is the pre-processed drone image. , satellite basemap and causal semantic graphs ;

[0100] The processing process is:

[0101] Step 3.1: Based on the causal semantic graph , generate semantically guided keypoint detection masks:

[0102]

[0103] function The causal semantic graph is converted into a guidance mask for key point detection. The specific function expression in this embodiment is:

[0104]

[0105] in are the pixel coordinates in the image, Indicates location The guide mask value at , indicating the priority of this position as a key point, is an indicator function, which is 1 when the condition in the brackets is met, otherwise it is 0. Representing semantic entities On satellite basemap The spatial area covered by Semantic Entity In this embodiment, the importance weight of semantic entities such as coastline is 0.9, the importance weight of semantic entities such as permanent buildings is 0.8, the importance weight of semantic entities such as large landmarks is 0.7, the importance weight of semantic entities such as road intersections is 0.6, the importance weight of semantic entities such as general land features is 0.4, and the importance weight of semantic entities such as change areas is 0.2. This function analyzes the causal semantic graph The importance and stability of each semantic entity in the causal semantic graph are converted into a mask of the same size as the image in step 1. , where the pixel value represents the priority of the position as a key point for key point detection. The mask constructed by this function highlights the following areas:

[0106] The shoreline edge area was given the highest priority because the shoreline is a stable feature that distinguishes land from sea;

[0107] Permanently built areas (e.g., lighthouses, piers) have the next highest priority; these artificial structures are generally stable and easily recognizable.

[0108] Landmark areas with distinct features (such as mountains and special terrain) have medium priority;

[0109] Waterway-bounded areas (e.g., estuaries, bay entrances) are given a lower but meaningful priority;

[0110] This mask-guided mechanism makes the keypoint detection process more focused on semantically meaningful and stable regions, reducing sensitivity to volatile areas such as water ripples.

[0111] Step 3.2: Keypoint detection using self-supervised deep learning architecture:

[0112]

[0113]

[0114] Function is a keypoint detector using a deep learning architecture, which receives preprocessed images and semantic-guided keypoint detection masks as input, and outputs a set of key points The working mechanism of the keypoint detector is as follows:

[0115] First, a multi-scale feature pyramid is constructed through a fully convolutional network (FCN) to capture feature information at different scales;

[0116] Then, a keypoint heat map of the image is generated through the multi-scale feature pyramid, representing the likelihood of each pixel position as a key point;

[0117] Next, the semantic-guided keypoint detection mask is multiplied with the heat map to enhance the response value of semantically important regions;

[0118] Finally, local optimal points are selected through non-maximum suppression (NMS), and a threshold is used to filter out the final set of key points;

[0119] This deep learning method has stronger feature expression ability compared to traditional keypoint detectors, and the introduction of semantic guidance makes the detection more focused on valuable regions.

[0120] Step 3.3: Extracting basic feature descriptors of key points:

[0121]

[0122]

[0123] Function is used to extract feature descriptors for detected key points. It is implemented based on a deep neural network, using a shared encoder architecture to ensure computational efficiency of feature extraction, and through local feature aggregation, a 256-dimensional compact feature descriptor is extracted within the receptive field around the key point, and the compact feature descriptor is processed through L2 normalization to facilitate subsequent similarity calculation; and through metric learning training, the discriminative ability of the descriptor is optimized, making the descriptors of the same object closer and the descriptors of different objects farther apart. Through this function The extracted deep feature descriptors have partial invariance and are robust to illumination changes, view angle changes and scale changes. These feature descriptors capture the visual characteristics such as texture, shape and structure of the region around the key point, and are the basis for subsequent feature matching.

[0124] Step 3.4: Establish the association between the key points and the semantic entities:

[0125]

[0126] This step establishes the mapping relationship between the key points and the semantic entities , that is, to determine which semantic entity each key point belongs to. The function outputs the semantic entity covered by the spatial region, and when the key point falls within the region, the association is established. This association relationship is the basis for subsequent utilization of semantic information to enhance feature representation, enabling the system to adjust the feature matching strategy according to the semantic relationship.

[0127] The final output of this step is the key point set , , the basic feature descriptor , and the association relationship between the key points and the semantic entities .

[0128] Step 4: Causally semantic attention enhanced feature representation and optimal matching:

[0129] Based on the causal semantic graph , the semantic association strength between key points is calculated; based on the semantic association strength, a self-attention-based feature enhancement network is constructed, and the input features of the feature enhancement network are fused with external environment and time factors to obtain enhanced features and ; using the enhanced features and , an optimal transport problem with causal semantic constraints is constructed, and by solving the optimal transport matching problem, an optimal matching matrix is obtained, which is used to filter the key point matching pairs between the unmanned aerial vehicle image and the satellite base map , and establish an initial matching set .

[0130] In this embodiment, the input of step 4 is the key point set , , the basic feature descriptor , ​, the association relationship between the key points and the semantic entities and the causal semantic graph ;

[0131] The processing procedure is as follows:

[0132] Step 4.1: calculating the key point set based on the causal semantic graph the semantic association strength between a key point in the key point set and a key point in the key point set

[0133]

[0134] wherein is a key point in the key point set , and is a key point in the key point set , the semantic association strength between and is calculated by the formula wherein is an indicator function, returning 1 when the condition is met, and 0 otherwise. denotes that the key point has a mapping relationship with the semantic entity , denotes that the key point has a mapping relationship with the semantic entity , is the relationship strength weight of the semantic entity and the semantic entity .

[0135] The specific logic of the formula is as follows:

[0136] First, it is judged whether the key point belongs to the semantic entity , and whether the key point belongs to the semantic entity , i.e., whether there is a mapping relationship; if both belong, the relationship strength weight between the semantic entity and is calculated; finally, the above calculation is performed on all possible semantic entity pairs and summed. In this way, when two key points respectively belong to semantic entities with close semantic association (such as one belonging to a lighthouse and the other belonging to a nearby wharf), their semantic association strength is higher; and the semantic association strength of a key point pair belonging to irrelevant semantic entities (such as one belonging to a far-off building and the other belonging to a near-sea wave) is lower. This provides prior information at the semantic level for subsequent feature matching.

[0137] Step 4.2: Build a self-attention based feature enhancement network:

[0138]

[0139]

[0140] in Representing an image The 0th layer features, , Representing an image The basic feature descriptor of Representing an image No. Layer features, Representing an image No. Layer features, functions represents the multi-head attention mechanism, Represents the strength of semantic association; the above formula describes the processing process of the multi-layer self-attention network:

[0141] Initialize the 0th layer features Basic feature descriptor ; For the Layer, updates the features through residual connection, that is, the features of the current layer are equal to the features of the previous layer plus the output of the multi-head attention mechanism; function Represents a multi-head attention mechanism, receiving the previous layer features as queries ,key Sum , and the strength of semantic association as additional input.

[0142] The specific calculation formula of the multi-head attention mechanism is:

[0143]

[0144]

[0145] Here, each attention head The attention weights are calculated independently and then aggregated; 、 、 is a learnable projection matrix that projects the input feature Z into the query, key, and value space; Calculation query and key The similarity matrix between Normalize to prevent the gradient disappearance problem; Incorporate semantic association strength into attention calculation, where is a weight coefficient of semantic information; The function converts similarity into a probability distribution; finally multiplied by the value matrix to get the weighted aggregated feature representation.

[0146] The above self-attention-based feature enhancement network has layers, and the final feature output is the corresponding feature of the UAV image and the corresponding feature of the satellite base map . .

[0147] The innovation of this attention mechanism is to introduce semantic association information , so that the feature enhancement process not only considers the similarity of the features themselves, but also considers the relevance at the semantic level, thereby generating a more robust feature representation.

[0148] Step 4.3: Fuse and with external environment and time factors to further enhance feature representation:

[0149]

[0150]

[0151]

[0152]

[0153] where is a dynamic adjustment coefficient, calculated by a multi-layer perceptron according to image quality , platform motion parameters and environmental weather conditions , the dynamic adjustment coefficient value is between 0 and 1, reflecting the degree of dependence on prior knowledge under current conditions, image quality , platform motion parameters and environmental weather conditions are set according to the actual scene; is a prior mask based on time and geographic information, generated by a convolutional neural network based on the time of day , season , tidal state and geographic location ; and are obtained by existing technologies in the art; represents element-level multiplication, i.e. element-wise multiplication; finally get the enhanced feature corresponding to the UAV image Enhanced features corresponding to satellite base map .

[0154] By this step, the feature representation is flexibly adjusted according to external conditions. For example, in low light conditions, the dependence on prior knowledge can be increased; in the period when the tide changes obviously, the influence of the tide on the coastline shape is compensated through the mask.

[0155] Step 4.4: Utilize enhanced features and An optimal transport problem with causal semantic constraints is constructed and solved:

[0156]

[0157]

[0158]

[0159] The above formula models the feature matching problem as a constrained optimal transport problem; wherein represents the enhanced feature corresponding to the key point , represents the enhanced feature corresponding to the key point , is a preset coefficient; is the cost matrix between the key point and the key point , which consists of two parts: the first part is the cosine distance of feature similarity, and the smaller the value is, the more similar the features are; the second part is the complement of semantic association strength, whose influence is controlled by the coefficient ; is the optimal matching matrix, represents the matching probability of the key point in the unmanned aerial vehicle image and the key point in the satellite base map; the constraint condition ensures that the sum of the matching probabilities of the key points meets the preset distribution and , and in this embodiment, it is assumed that and are uniform distributions.

[0160] The advantage of this optimal transport problem model is that it can consider all potential matching pairs globally, rather than simply filtering based on a threshold, thereby improving the overall consistency and accuracy of the matching. The introduction of semantic constraints makes the point pairs with strong semantic association more likely to be matched.

[0161] Since the direct solution of the optimal transport problem has high computational complexity, Sinkhorn algorithm is used to solve its regularized version to reduce the computational complexity. The flow of the algorithm is:

[0162] Initialize the matching matrix , where is the cost matrix, is the temperature parameter, which controls the softness of the matching; and the row normalization and column normalization are alternately performed by iteration:

[0163]

[0164] where and are the scaling vectors that make the row and column of satisfy the constraint condition, and finally the optimal matching matrix is obtained. Sinkhorn algorithm usually converges within dozens of iterations, greatly reducing the computational complexity and enabling the system to process a large number of key point matches in real time. The smaller the parameter is, the closer the matching is to the hard assignment (one-to-one matching); the larger the parameter is, the softer the matching is (one-to-many matching).

[0165] After obtaining the optimal matching matrix , the key point matching pairs with high confidence are screened by the threshold to construct the initial matching set :

[0166]

[0167] where is the element value in the optimal matching matrix . The threshold controls the sparsity and confidence of the matching. A higher threshold value obtains more reliable but fewer matches, and a lower threshold value obtains more but possibly contains false matches.

[0168] The final output of this step is the enhanced feature and and the initial matching set .

[0169] Step 5: False matching screening based on causal semantic constraints:

[0170] For the initial matching set , respectively, local neighborhood consistency verification and global structure verification based on causal semantic relationship, and the local consistency score obtained by local neighborhood consistency verification, the global consistency score obtained by global structure verification and the initial matching score of the matching pair are weighted and fused, low confidence matches are filtered according to the fusion result, and a refined matching set is obtained .

[0171] In this embodiment, the input of step 5 is the initial matching set , the key point set , and the causal semantic graph ;

[0172] The processing process is as follows:

[0173] For a key point matching pair , local neighborhood consistency verification is performed according to the following formula:

[0174]

[0175]

[0176]

[0177] Wherein denotes the neighborhood set of key point , containing all key points with a distance less than threshold from , denotes the neighborhood set of key point , is a distance function; is the local consistency score of the key point matching pair , that is, how many matching pairs exist in the initial matching set within the respective neighborhood; wherein the numerator calculates the number of matches between the two neighborhoods, and the denominator takes the smaller value of the size of the two neighborhoods, and the result is normalized to the interval [0, 1].

[0178] This local consistency test is based on the assumption that correct matches should exhibit clustering effects in space, that is, points in the neighborhood of one point should be matched to points in the neighborhood of another point. A high consistency score indicates that the spatial structure around the matching points remains consistent, and a low score may be a false match.

[0179] Similarly, for a key point matching pair , global structure verification is performed according to the following formula:

[0180]

[0181] wherein denotes a keypoint the semantic entity to which the keypoint belongs the relationship strength weight of the semantic entity to which the keypoint belongs, denotes a keypoint the semantic entity to which the keypoint belongs the relationship strength weight of the semantic entity to which the keypoint belongs; denotes a keypoint matching pair the confidence of the keypoint matching pair in the optimal matching matrix . This formula calculates the global structure consistency score by weighted average of the reliability of the current matching pair and other matching pairs with semantic association, and the weight is determined by the semantic relationship strength. This verification mechanism captures a wider range of structural relationships, not limited to spatial proximity, and can detect mismatching that violates semantic constraints but may satisfy local consistency.

[0182] After obtaining the local consistency score and the global structure consistency score, combine the initial matching score to obtain the fusion result of the keypoint matching pair :

[0183]

[0184] wherein is the initial matching score of the keypoint matching pair , taking its value in the optimal matching matrix , and , and are weight coefficients, satisfying . Then filter low-confidence matches based on the threshold to obtain the refined matching set :

[0185]

[0186] This multi-constraint screening mechanism can effectively remove mismatching and improve the accuracy of matching.

[0187] The final output of this step is the refined matching set .

[0188] Step 6: Optimize the matching result using multi-level causal graph convolution:

[0189] First, combine the causal semantic graph to construct an enhanced adjacency matrix for the matching set ​ ; then the adjacency matrix is applied to optimize the node feature matrix using a multi-layer graph convolution network; the keypoint matching pairs in the matching set are re-evaluated using the optimized node feature matrix to remove the matches that do not conform to the global semantic consistency, and a final matching set is obtained .

[0190] In this embodiment, the input of step 6 is the refined matching set and the causal semantic graph .

[0191] The processing procedure is as follows:

[0192] Step 6.1: An enhanced adjacency matrix is constructed for subsequent graph convolution operations:

[0193]

[0194] wherein is an original adjacency matrix constructed based on the matching set , reflecting the direct connection relationship between the points, is an identity matrix of the same size as the matrix , is a weight coefficient of the causal relationship, controlling the influence degree of the causal knowledge, is a causal relationship adjacency matrix, capturing the indirect causal influence between the points, and the specific formula is:

[0195]

[0196] wherein denotes the adjacency relationship between the semantic entity and the semantic entity in the causal relationship adjacency matrix, and all possible paths from the semantic entity to the semantic entity are considered in the calculation. The edge weights on each path are multiplied, and then all the paths are summed. This design can capture the indirect relationship between distant nodes, so that information can propagate along the causal chain of the semantic graph. Wherein denotes is a certain path, denotes the semantic entity and the semantic entity are two nodes on the path , is the relationship strength weight of the semantic entity and the semantic entity .

[0197] Step 6.2: Utilize the adjacency matrix , and apply a multi-layer graph convolution network to optimize the node feature matrix, where the formula of the node feature matrix is:

[0198]

[0199] where and denote the node feature matrix of the th layer and the th layer, respectively, is the degree matrix of , is the learnable weight matrix of the th layer, controlling the feature transformation, is a nonlinear activation function, such as ReLU or LeakyReLU.

[0200] This graph convolution operation allows each node to aggregate information from its neighbors, which are defined by the augmented adjacency matrix . By stacking multiple layers, nodes can acquire more extensive contextual information, generating more globalized feature representations. The normalization term ensures that features do not diverge or disappear as the number of layers increases.

[0201] Step 6.3: Further refine the matching results based on the optimized node feature matrix:

[0202]

[0203] The function re-evaluates the confidence of matching pairs in the matching set , eliminates matches that do not conform to global semantic consistency, and obtains the final matching set . In this embodiment, the specific expression is:

[0204]

[0205] where is the L2 norm of the feature vector corresponding to the key point in the node feature matrix , is the L2 norm of the feature vector corresponding to the key point in the node feature matrix , is the graph convolution refinement threshold, used to control the strictness of matching screening, and in this embodiment, the value is ​​​The advantage of this step is that the graph convolution network can capture high-order relationships and global structures between matching points, going beyond simple pairwise matching constraints, thereby further improving the accuracy and robustness of matching.

[0206] The final output of this step is the final matching set .

[0207] Step 7: Perform position calculation and precision evaluation based on multi-source information fusion:

[0208] According to the final matching set , calculate the deviation vector between the actual position and the estimated position of the UAV ; fuse the deviation vector with the inertial navigation data to obtain the fused position estimate and calculate the positioning precision evaluation index, realize the UAV position correction and navigation update.

[0209] In this embodiment, the input of step 7 is the final matching set and the initial position estimate provided by the inertial navigation system ;

[0210] The processing process is as follows:

[0211] Step 7.1: According to the final matching set , calculate the deviation vector between the actual position and the estimated position of the UAV :

[0212]

[0213] Where is the coordinate difference set of the key point matching pairs in the final matching set , i.e. the difference between the geographic coordinates of the key points in the satellite base map and the projection coordinates of the matching key points in the UAV image, is the coordinate transformation matrix; based on the common view geometry principle, the corresponding relationship of the matching key points is inversely deduced to the camera position through coordinate transformation.

[0214] Step 7.2: According to the formula:

[0215]

[0216] The optimal fusion of inertial navigation data and scene matching results is realized, where is the initial position estimate provided by the inertial navigation system, is the Kalman gain matrix, which is dynamically adjusted according to the uncertainty of the inertial navigation system and the matching result, is the final position estimate after fusion. The Kalman gain in it The calculation takes into account the reliability of the two information sources: when the inertial navigation system is just reset or the matching result is very reliable, a larger value, and is more inclined to adopt the matching result; on the contrary, when the inertial navigation system is relatively accurate in the short term or the matching result is not very reliable, a smaller value, and is more conservative in adjusting the position.

[0217] The final output of this step is the fused final position estimate The fused final position estimate is finally fed back to the unmanned aerial vehicle navigation system to adjust the navigation parameters such as heading, speed, etc., to realize navigation state updating and ensure that the unmanned aerial vehicle can safely cruise according to the predetermined task trajectory.

[0218] Experimental comparison:

[0219] Using the coastline scene matching dataset, the performance of the method of the present application is compared with that of various existing methods, and the results are as follows:

[0220] Table 1: Performance comparison

[0221]

[0222] The anti-interference ability is comprehensively evaluated by the matching stability under the interference conditions of strong light reflection, sea fog, tidal changes, etc. It can be seen from the experimental results that the method of the present application is superior to the existing methods in terms of matching accuracy, matching time consumption, false alarm rate and positioning error, and especially outstanding in anti-interference ability, which can effectively cope with the challenges of complex environment in the coastline area.

[0223] Although the embodiments of the present application have been shown and described above, it should be understood that the above-mentioned embodiments are exemplary and should not be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above-mentioned embodiments within the scope of the present application without departing from the principles and purposes of the present application.

Claims

1. A scene matching positioning method for an unmanned system in an area denial scenario, characterized by: The method comprises the following steps: Step 1: Real-time collection of UAV images by high-altitude unmanned system Geometric correction is performed to obtain corrected UAV images ; corrected unmanned aerial vehicle image and a pre-stored satellite base map to obtain a processed unmanned aerial vehicle image and the satellite base map ; according to the coastline characteristics, the processed unmanned aerial vehicle image and the satellite base map are divided into three sub-regions of a sea area , a coastline area and a land area , and different regions are given different weights ; Step 2: identifying semantic entities in the UAV image and the satellite base map and establishing spatial relationships between the semantic entities, while calculating the relationship strength weight between the semantic entities, to construct a causal semantic graph ; Step 3: generating a keypoint detection mask according to the causal semantic graph , generating a keypoint detection mask, a pixel value at a position in the mask representing a priority of the position as a keypoint; Then, the key point detection mask is used to detect the UAV images and satellite basemaps Perform key point detection based on self-supervised deep learning architecture to obtain drone images and satellite basemaps The corresponding key point sets and , and extract the deep feature descriptor of the key points and , and finally establish the association relationship between key points and semantic entities ; Step 4: Causal Semantic Graph Based Computing the semantic correlation strength between key points; Based on the semantic correlation strength, a feature enhancement network based on self-attention is constructed, and the input features of the feature enhancement network are fused with external environment and time factors to obtain enhanced features and ; using enhanced features and , an optimal transport problem with causal semantic constraints is constructed, and an optimal matching matrix is obtained by solving the optimal transport matching problem , the optimal matching matrix is used to screen the key point matching pairs between the unmanned aerial vehicle image and the satellite base map , and an initial matching set is established ; Step 5: for the initial matching set , respectively, carry out local neighborhood consistency verification and global structure verification based on causal semantic relationship, and weight and fuse the local consistency score obtained by local neighborhood consistency verification, the global consistency score obtained by global structure verification, and the initial matching score of the matching pair, filter low-confidence matches according to the fusion result, and obtain a refined matching set ; Step 6: Combine the causal semantic graph , and the matching set , construct an enhanced adjacency matrix ; then use the adjacency matrix , apply a multi-layer graph convolution network to optimize the node feature matrix; use the optimized node feature matrix to re-evaluate the key point matching pairs in the matching set , eliminate the matches that do not conform to the global semantic consistency, and obtain the final matching set ; Step 7: According to the final matching set Calculate the deviation vector between the actual position of the UAV and the estimated position ; fuse the deviation vector with the inertial navigation data to obtain the fused position estimate.

2. The method of claim 1, wherein: In step 1, the zone weight wherein is a shore zone weight, is a land zone weight, is a sea zone weight, and .

3. The method of claim 1, wherein: In step 2, the causal semantic graph wherein is a set of nodes with semantic entities as nodes, is a set of nodes with semantic entities as nodes, is a set of edges consisting of relations between two semantic entities with explicit spatial association in is a set of edges consisting of relations between two semantic entities with explicit spatial association in is a set of edge weights consisting of strength weights of each relation in 4. The method of claim 2, wherein: In step 2, the formula is Computing relationship strength weights between semantic entities and wherein the function is a stability function for assessing the stability of the relationship between the semantic entities, the function is a persistence function for assessing the degree of change in the relationship over a time scale, the function is a recognizability function for assessing the visual recognizability and uniqueness of the relationship, the function is a relationship weight function taking wherein is a variable indicating a shore line area, is a variable indicating a land area, is a variable indicating a sea area.

5. The method of claim 1, wherein: In step 3, the formula is Generate keypoint detection mask ,in Indicates location The guide mask value at , indicating the priority of this position as a key point, Is an indicator function, which is 1 when the condition in the brackets is met, otherwise it is 0. Representing semantic entities On satellite basemap The spatial area covered by Semantic entity The importance weight of .

6. The method of claim 1, wherein: In step 4, the formula is Computing a set of key points based on a causal semantic graph a certain key point in the set of key points a certain key point in the set of key points a certain key point in the set of key points a certain key point in the set of key points wherein is an indicator function that returns 1 when the condition inside holds, and 0 otherwise; denotes that there is a mapping relationship between the key point and the semantic entity denotes that there is a mapping relationship between the key point and the semantic entity denotes that there is a mapping relationship between the key point and the semantic entity denotes that there is a mapping relationship between the key point is the relationship strength weight of the semantic entity and the semantic entity .

7. The method of claim 6, wherein: In step 4, the enhanced features are utilized and The optimal transportation problem constructed with causal semantic constraints is: wherein denotes the enhanced feature corresponding key point in the feature vector, denotes the enhanced feature corresponding key point in the feature vector, is a preset coefficient; is a cost matrix between the key point and the key point , is an optimal matching matrix, denotes a matching probability of the key point in the unmanned aerial vehicle image and the key point in the satellite base map; constraint condition ensures that a total sum of the matching probabilities of the key points satisfies a preset distribution and .

8. The method of claim 6, wherein: In step 5, for the keypoint match pairs in the initial match set , the following formula is used to calculate the matching score Computing the fusion result wherein , and are weight coefficients satisfying , is an initial matching score for the keypoint match pair , is a local consistency score for the keypoint match pair , is a global structure consistency score for the keypoint match pair .

9. The method of claim 6, wherein: In step 6, the formula of the node feature matrix is in and In turn, they represent Layer and The node feature matrix of the layer, is the enhanced adjacency matrix The degree matrix of It is The learnable weight matrix of the layer, is a nonlinear activation function; the enhanced adjacency matrix According to the formula is obtained, wherein is based on the matching set is the original adjacency matrix constructed, is the identity matrix of the same size as the matrix is the identity matrix of the same size as the matrix is the weight coefficient of the causal relationship, is the causal relationship adjacency matrix.

10. The method of claim 9, wherein: In step 6, the formula is obtaining a final matching set wherein is a key point is an L2 norm of a corresponding feature vector in a node feature matrix is an L2 norm of a corresponding feature vector in a node feature matrix is a key point is an L2 norm of a corresponding feature vector in a node feature matrix is an L2 norm of a corresponding feature vector in a node feature matrix is a graph convolution refinement threshold.

Citation Information

Patent Citations

  • Rapid scene matching and positioning method for unmanned aerial vehicle in large-range scene

    CN119251455A

  • Unmanned aerial vehicle visual positioning method based on adaptive area routing mixed attention

    CN120298502A

  • Semantic segmentation-based unmanned aerial vehicle image georeferencing method, and related device

    WO2024221946A1