Foundation pit surveying unmanned aerial vehicle pose correction method based on visual feature matching
By employing multi-level feature processing and error adjustment mechanisms, the problem of mis-association in visual feature matching during foundation pit surveying was solved, achieving high accuracy and stability in UAV pose estimation and reducing safety risks in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2026-03-24
AI Technical Summary
In the surveying of large-scale foundation pit projects, existing technologies are prone to misassociations due to the visual feature matching algorithm, which can lead to deviations in UAV pose estimation, resulting in flight path errors and abnormal feedback from the control system, posing safety risks, especially in complex environments.
Through multi-level feature processing and error adjustment mechanisms, including texture information entropy screening, structural consistency analysis, and error-sensitive region label modeling, a closed-loop process is constructed to dynamically adjust the feature point removal threshold and solution parameters, adaptively responding to texture interference and error fluctuations.
It significantly improves the accuracy and stability of UAV attitude estimation, ensures high-precision control of flight path and image registration results, has strong robustness and high adaptability, avoids mismatch interference, and reduces equipment damage and safety risks.
Smart Images

Figure CN121074433B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of surveying and mapping engineering technology, and specifically to a method for UAV pose correction in foundation pit surveying based on visual feature matching. Background Technology
[0002] "Visual Feature Matching-Based UAV Attitude Correction for Foundation Pit Surveying" refers to the process of acquiring on-site images using visual sensors (such as cameras) onboard a UAV during foundation pit surveying. By extracting key visual feature points (such as edges, corners, and textures) from these images, the system matches these features with features in existing reference images or digital models. This identifies deviations in the UAV's position (position coordinates) and attitude (flight angle, direction) caused by factors such as GPS drift, sensor errors, or wind interference during actual flight. Based on the matching results, the system estimates the UAV's current actual attitude error in real time and dynamically corrects the navigation system to ensure its flight path remains consistent with the predetermined surveying trajectory. This improves the data accuracy and spatial consistency of surveying tasks such as 3D modeling and contour extraction of the foundation pit area.
[0003] Existing technologies suffer from the following shortcomings: In the mapping scenarios of large-scale foundation pit engineering projects, there are often a large number of structurally repetitive components, such as precast blocks on retaining walls, evenly spaced support piles, and regular textures of steel mesh. These engineering structures form highly similar texture distributions in visual images, leading to problems such as feature redundancy, strong morphological symmetry, and low discriminability in local image areas. In the process of UAV pose correction based on visual feature matching, the system usually relies on key points in the image for spatial registration and pose estimation. However, when multiple structural regions have similar geometric distributions and texture features, the matching algorithm is prone to misassociation, that is, misidentifying similar features at different locations as the same target point, resulting in mis-locking of feature points. This type of mismatch behavior will directly interfere with the accuracy of UAV attitude calculation, causing pose estimation to shift or even reverse, further leading to errors in flight path judgment and abnormal feedback from the control system. If the error continues to accumulate, the UAV will experience global flight trajectory drift during the mapping process, resulting in spatial misalignment between the collected image data and the real scene, ultimately leading to severe distortion of the 3D modeling, orthophoto stitching, and other output data. More seriously, in confined spaces such as narrow construction areas and the edge of deep foundation pits, incorrect attitude estimation may cause flight control module failure and abnormal path planning, resulting in uncontrolled flight or collision accidents of the drone, posing a high risk of equipment damage and on-site safety.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a pose correction method for UAVs used in foundation pit mapping based on visual feature matching. Through multi-level feature processing and error adjustment mechanisms, a closed-loop process from feature extraction to pose optimization is constructed. Specifically, error-sensitive region labels and confidence interpolation modeling are introduced to effectively suppress mismatch interference. The system can dynamically adjust the rejection threshold and convergence parameters, adaptively responding to texture interference and error fluctuations in complex scenarios, significantly improving the accuracy and stability of UAV attitude estimation. It possesses advantages such as strong robustness and high adaptability, thus solving the problems mentioned in the background technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for UAV pose correction in foundation pit mapping based on visual feature matching, comprising the following steps:
[0007] S1, Obtain the current image frame, extract regions with high texture complexity as salient regions based on the texture information entropy distribution in the image, remove feature points in regions with low texture variability, and generate the first candidate feature region set;
[0008] S2, based on the first candidate feature region set, combined with the region boundary contour, directional gradient distribution and local neighborhood density, structural consistency evaluation is performed, and image feature points with high structural integrity are selected to generate the second candidate feature region set;
[0009] S3. Based on the second candidate feature region set, perform image registration, construct an image registration error distribution map, calculate the matching residual value of image feature points in each region, identify the region in the matching error set, and generate an error-sensitive region label map.
[0010] S4. Based on the error-sensitive region label map, a dynamic confidence backoff process is executed. Image feature points in the region with high matching residual values are locally excluded. A confidence interpolation weight map is constructed based on the distribution of image feature points in the neighboring region with low matching residual values. The local confidence adjustment matrix is then output.
[0011] S5, input the local confidence adjustment matrix into the pose calculation process, combine the spatial distribution density of image feature points with the matching residual strength to generate a local confidence weight distribution map, and simultaneously calculate the matching residuals of all image feature points to obtain the overall image matching error energy map;
[0012] S6, based on the spatial coupling relationship between the overall image matching error energy map and the local confidence weight distribution map, constructs a multi-scale residual adjustment strategy to dynamically update the image feature point removal threshold and pose calculation convergence parameters.
[0013] Preferably, step S1 includes:
[0014] The current image frame is acquired, and the image is locally segmented using a window sliding method. A gray-level co-occurrence matrix is constructed, and the texture information entropy value is calculated to generate a texture information entropy heatmap.
[0015] Based on the texture information entropy heatmap, a saliency threshold is set to extract salient regions with high texture complexity and construct a salient region mask map.
[0016] Image feature points are extracted based on the salient region mask map constraint, and a local feature response value filtering mechanism is introduced to remove feature points with low response intensity, unstable orientation, and non-repeatable scale.
[0017] Spatial clustering analysis is performed on the selected image feature points. Density suppression is performed based on the spatial density distribution of the clustered regions to remove redundant points and generate a set of first candidate feature regions.
[0018] Preferably, step S2 includes:
[0019] For each region in the first candidate feature region set, the boundary contour is extracted, the boundary continuity, closure and edge intensity changes are calculated, and regions with incomplete structures are eliminated.
[0020] Perform directional gradient distribution analysis on the retained regions, construct directional gradient histograms, and filter out regions with concentrated directions and poor directional responsiveness;
[0021] Based on the changes in the number, distribution direction, and response value of feature points around each feature point according to the local neighborhood density statistics, regions with poor structural consistency are eliminated to generate a second set of candidate feature regions.
[0022] Preferably, step S3 includes:
[0023] Image feature matching is performed using all high-quality image feature points in the second candidate feature region set. A two-way verification mechanism and a distance ratio filtering strategy are used to ensure the uniqueness of matching point pairs in the descriptor space.
[0024] Perform geometric verification on all matching point pairs and remove abnormal matching points that do not meet spatial consistency requirements;
[0025] Calculate the registration residual value for each matching point pair, construct a registration error distribution map, and record the residual mean, residual variance, and local outlier density for each region;
[0026] Error clustering analysis is performed on the registration error distribution map. The sliding window and density estimation techniques are used to calculate the proportion and clustering trend of high residual points in each window area, identify error concentration areas, and generate error-sensitive area label maps.
[0027] Preferably, step S4 includes:
[0028] Identify high-matching residual regions based on the error-sensitive region label map, filter out low-confidence feature points within the regions, and locally exclude them;
[0029] Low-error feature points are extracted from the neighboring regions of the error-sensitive region, a feature point set is constructed, and the confidence level is evaluated based on the distribution density, orientation consistency, and residual stability.
[0030] Based on the confidence of neighboring feature points, weighted distance modeling or Gaussian kernel interpolation is performed to generate a confidence interpolation weight map.
[0031] A local confidence adjustment matrix is generated based on the confidence interpolation weight map, providing input parameters for weighted error control in subsequent pose calculations.
[0032] Preferably, step S5 includes:
[0033] Receive and load the local confidence adjustment matrix, match its position with the set of feature points in the current image, and assign each feature point a confidence adjustment weight corresponding to its position.
[0034] By combining the spatial distribution density and geometric location of feature points, the number of feature points, directional gradient distribution and response value intensity in each region are statistically analyzed, and local confidence fusion modeling is performed to generate a local confidence weight distribution map.
[0035] For all feature point matching pairs, the residual values are calculated and weighted and accumulated according to the confidence of each feature point to form a residual field indexed by pixel coordinates;
[0036] Based on the residual field, regional aggregation and statistical modeling are performed to generate an overall image matching error energy map.
[0037] Preferably, step S6 includes:
[0038] Spatial coupling analysis is performed on the overall matching error energy map and the local confidence weight distribution map of the image to calculate the correspondence between the residual intensity and confidence level in each image region and identify different error regions;
[0039] A multi-scale residual adjustment framework is constructed, which dynamically sets the feature point removal threshold in the small-scale region, adjusts the regional weights in the medium-scale region, and evaluates the overall map error control in the large-scale region.
[0040] Based on the multi-scale adjustment results, the convergence parameters in the pose calculation process are updated in real time to ensure error control and calculation stability.
[0041] The adjusted feature point removal threshold and pose calculation parameters are applied to the pose estimation process of the current image frame, and the processing results are fed back to the error history record to update the adjustment strategy for the next frame.
[0042] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0043] This invention begins with feature extraction from the image source and constructs a closed-loop feedback mechanism from feature selection to error adjustment and pose optimization through multi-level processing including texture entropy filtering, structural consistency analysis, and residual aggregation identification. In particular, the introduction of error-sensitive region labels and confidence interpolation modeling effectively avoids interference from mismatched points on the estimation results. Furthermore, the system dynamically adjusts the feature point removal threshold and solution convergence parameters throughout the estimation process, enabling it to adaptively cope with texture interference and error distribution fluctuations in different scenarios. Ultimately, this achieves high-precision control and stability assurance of UAV flight path and image registration results in complex engineering environments, demonstrating significant advantages such as strong robustness, high adaptability, and strong engineering feasibility. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0045] Figure 1 This is a flowchart of the UAV pose correction method for foundation pit mapping based on visual feature matching according to the present invention. Detailed Implementation
[0046] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0047] This invention provides, for example Figure 1 The method for UAV pose correction for foundation pit mapping based on visual feature matching, as shown, includes the following steps:
[0048] S1, acquire the current image frame, extract regions with high texture complexity as salient regions based on the texture information entropy distribution in the image, remove feature points in regions with low texture variability, and generate the first candidate feature region set.
[0049] To address issues such as image feature redundancy and feature matching ambiguity caused by structurally repetitive components, and to improve the pose estimation accuracy of UAVs in foundation pit mapping, this process first performs texture information entropy analysis on the acquired current image frame, constructs an image salient region extraction process, and then filters key regions with high texture complexity in the image to form a preliminary set of candidate feature regions that can be used for spatial registration. The specific steps of this process are as follows:
[0050] The process involves acquiring real-time image frames captured by the drone during flight, dividing the entire image into raster sections using a sliding window mechanism, and calculating the texture entropy value of each sub-region at a fixed scale (e.g., 16×16 pixels). Specifically, the texture entropy calculation uses a gray-level co-occurrence matrix to construct a basic gray-level probability distribution, and the Shannon entropy function is used to evaluate the texture complexity of each local region, resulting in an image texture entropy heatmap. In this heatmap, regions with more complex texture structures have higher entropy values, reflecting richer texture details and stronger feature uniqueness and recognition value. This step aims to quantify the texture performance of different regions in the image, thereby providing data support for subsequent salient region extraction.
[0051] The specific steps for constructing a basic gray-level probability distribution using a gray-level co-occurrence matrix and evaluating the texture complexity of local image regions using the Shannon entropy function to form a texture information entropy heatmap for the entire image are as follows:
[0052] The input image is converted to grayscale and then divided into local areas using a sliding window approach, with each small window considered as an independent image region.
[0053] Within each window, the relative spatial relationships between gray levels are statistically analyzed to construct a gray-level co-occurrence matrix for that region. This matrix reflects the joint occurrence frequency of different gray-level pairs at specific directions and distances, thereby capturing the texture structure pattern of that region.
[0054] Normalize the count values in the co-occurrence matrix to obtain the gray-level probability distribution of the local region, that is, the probability of different gray-level pairs appearing simultaneously in the region.
[0055] Based on this set of probability data, the Shannon entropy function is used to evaluate the texture information of the region. The information entropy value reflects the complexity of the gray-level distribution: the more diverse the gray-level relationships and the more balanced the distribution, the higher the entropy value, indicating that the texture is more complex; conversely, it indicates that the texture is simple or highly repetitive.
[0056] The entropy value corresponding to each window region is mapped back to its corresponding position in the original image, forming a texture information entropy heatmap covering the entire image. The brightness value of each pixel or small block in this heatmap represents the texture complexity of that region, which is used for subsequent salient region extraction and feature selection operations. The entire process does not depend on specific image content, has good versatility and texture structure perception capabilities, and is especially suitable for detecting problems such as repetitive structures and texture degradation.
[0057] Based on the constructed texture entropy heatmap, a saliency threshold parameter is set, and image regions with entropy values higher than the set threshold are identified as salient regions. To ensure the spatial continuity and feature density of the extracted regions, region connectivity analysis is further performed. Neighboring regions that meet the entropy requirements are merged, and fragmented regions with small areas, strong isolation, or severe edge interference are removed, thereby generating a saliency region mask map with continuous morphology and rich texture. This mask map is used for regional constraint in the subsequent feature extraction process, ensuring that feature points are extracted only within recognized high texture complexity regions. This process effectively eliminates invalid regions with simple textures and low entropy caused by structural repetition in the image, which is beneficial to improving the discriminative ability of feature points.
[0058] After constructing the salient region mask, a feature point extraction algorithm (such as SIFT or ORB) is invoked. Under the constraint of the salient mask, key point extraction is performed only in high information entropy regions. Simultaneously, to prevent feature quality instability caused by local interference or illumination variations within high-texture regions, a local feature response value filtering mechanism is introduced. This mechanism calculates the response intensity, directional stability, and scale repeatability parameters for each extracted feature point. Weak response points below a threshold are removed, further optimizing the stability and discriminability of the feature point set. By limiting the extraction range and introducing an intensity filtering mechanism, the discriminative power and spatial uniqueness of feature points are effectively enhanced, reducing the probability of structural repetition interfering with subsequent matching algorithms.
[0059] After constructing the salient region mask, this mask is used as the region restriction input for the feature extraction process, ensuring that feature point detection only occurs in image segments identified as high information entropy regions. Specifically, the original image and the salient region mask are first aligned pixel-by-pixel, with each pixel in the mask corresponding one-to-one with the original image. Then, when executing feature point extraction algorithms (such as SIFT or ORB), before detecting each candidate keypoint, it is checked whether the pixel location belongs to a salient region in the mask. If the location is marked as a non-salient region (e.g., mask value is zero), the pixel is skipped to avoid extracting invalid feature points in regions with simple textures or repetitive structures. For high information entropy regions that pass the screening, keypoint detection, orientation assignment, and feature description steps are performed normally.
[0060] This constraint mechanism focuses the feature extraction process only on regions with complex textures and unique structures, reducing feature point redundancy and the risk of mismatches, and improving the efficiency and accuracy of subsequent feature matching and pose estimation. This region-limited extraction method, combining saliency masking and feature algorithms, is an important means to enhance image understanding and resistance to structural interference.
[0061] The local feature response value screening mechanism refers to the process of evaluating the quality of each feature point after initial feature point extraction. Based on indicators such as response intensity, directional stability, and scale adaptability within a local image region, feature points are further screened and optimized. Its core function is to eliminate low-quality feature points with weak responses, blurred edges, or susceptibility to noise from the detected candidate keypoints, retaining only those with strong responses generated in regions with drastic grayscale changes, clear orientation, and high scale repeatability. This mechanism can significantly improve the stability and matching robustness of feature points under image transformations (such as scaling, rotation, and brightness changes), avoiding misleading features in areas with repetitive structures or degraded textures, thereby improving the accuracy of feature point matching and providing more reliable basic data for subsequent pose estimation and image registration. In short, this mechanism is a crucial step in "quality control" of feature points, ensuring that each used feature point possesses strong discriminative power and stability.
[0062] The high-quality feature points obtained through the above steps are spatially clustered according to their salient regions. The redundancy and uniformity of the overall feature points are then evaluated based on the spatial density distribution after clustering. If densely distributed feature points with similar texture structures are found in local regions, point density suppression is applied based on a spatial density threshold. Key points with strong representativeness and high response values are retained, while redundant points are eliminated to prevent excessive concentration of feature points in homogeneous regions from causing subsequent matching ambiguities. Finally, a set of first-candidate feature regions with reasonable structural distribution, high texture complexity, and stable point quality is output, laying a solid foundation for subsequent contextual structure consistency analysis and image registration error modeling.
[0063] The main purpose of spatial clustering analysis is to evaluate and optimize the distribution of feature points in image space, avoiding situations where feature points are too dense in some areas and sparse in others, thereby improving the representativeness and matching balance of feature points throughout the image. Specifically, while high-quality feature points extracted from salient regions may possess good texture response, strong local textures or concentrated repetitive structures can lead to a large cluster of similar feature points in some areas, forming "dense redundancy." This not only wastes computational resources but also increases the risk of matching ambiguity. Spatial clustering analysis can identify the clustering trends and spatial distribution patterns of feature points in an image, and thereby evaluate the point density within each cluster region. In terms of implementation, density-based clustering algorithms (such as DBSCAN) or grid-based methods are typically used to divide the image into multiple spatial units, counting the number of feature points and their distribution characteristics within each unit. If feature points in a certain cluster region are too concentrated, the feature points with the highest response intensity and best location representativeness are retained based on a density threshold, while redundant points are removed, thus maintaining the spatial uniformity, comprehensive coverage, and recognition efficiency of the overall feature point set. This spatial clustering mechanism not only optimizes the quality structure of feature points, but also provides more stable and less ambiguous input for subsequent image registration and pose estimation, improving robustness and accuracy in complex scenes.
[0064] This step aims to improve the quality and discriminative power of feature points from the source, providing a reliable, stable, and representative feature foundation for subsequent image feature matching and UAV pose estimation. In complex foundation pit engineering environments, the presence of numerous structurally repetitive components (such as retaining piles, retaining walls, and steel mesh) easily leads to regions with highly similar textures and redundant information in the image. Feature points extracted from these regions often lack uniqueness and are easily misassociated during feature matching, causing matching errors and directly affecting the accuracy of UAV attitude calculation. To avoid this problem, this step introduces texture information entropy as a criterion, performing texture complexity analysis on the image and calculating local information entropy values region by region using a sliding window approach to identify significant regions with rich texture variations and prominent structural details. Subsequently, regions with low information entropy and monotonous textures are removed to avoid extracting feature points from these regions. This mechanism enables the final first candidate feature region set to have higher spatial resolution and texture recognition capabilities, significantly reducing feature ambiguity and matching interference caused by repetitive structures. Overall, this step plays a fundamental role in "selecting the best and eliminating the worst," laying a high-quality data input foundation for the entire visual positioning and pose correction process. It is a key preliminary step to ensure the accuracy of UAV positioning and the reliability of surveying data.
[0065] S2, based on the first candidate feature region set, combined with the region boundary contour, directional gradient distribution and local neighborhood density, structural consistency evaluation is performed, and image feature points with high structural integrity are selected to generate the second candidate feature region set;
[0066] To further improve the matching accuracy of image feature points against a background of structural repetition and avoid mismatches caused by incomplete local feature information or blurred directional features, a structural consistency evaluation process is executed after the extraction of the first candidate feature region set. This process identifies and selects image feature points with high structural integrity and clear geometric features, thereby constructing a second candidate feature region set with greater discriminative power and spatial stability. The structural consistency evaluation process includes the following steps:
[0067] For each region in the first candidate feature region set, boundary contours are extracted, and gradient-based edge detection algorithms (such as the Canny operator) are used to identify the contour lines of the region. By calculating the continuity, closure, and edge intensity changes of each region's boundary, it is determined whether the region possesses complete geometric boundary features. If the region's contour has problems such as breaks, blurred boundaries, or insignificant edge changes, it is considered to have weak structural expressive power and lack stable spatial reference significance, and is therefore discarded. This step aims to eliminate structurally incomplete regions caused by uneven lighting, reflection interference, or image blurring.
[0068] When extracting the boundary contour of each region in the first candidate feature region set, the specific implementation steps are as follows:
[0069] The original image is converted to grayscale to make subsequent edge detection processing more efficient and accurate.
[0070] For each candidate feature region, a local sub-image region is cropped within its corresponding image coordinate range and used as input data for edge detection.
[0071] Image gradient calculation methods are applied within this sub-image region. Typically, the Sobel operator or similar methods are used to extract the gradient images in the horizontal and vertical directions to obtain the gradient magnitude and direction information of each pixel.
[0072] The gradient information is further processed using the Canny edge detection algorithm: first, the region image is denoised by Gaussian filtering, then significant edge points are identified by the double threshold edge determination method, and the edge lines are refined by non-maximum suppression technology, making the extracted contour lines more accurate and continuous.
[0073] The edge lines extracted from each candidate region undergo closure and connectivity analysis to determine whether the contour possesses complete and clear structural features. The entire boundary contour extraction process aims to accurately reconstruct the geometric morphology of the candidate regions, providing a fundamental geometric basis for subsequent structural integrity assessment and feature point selection. This approach not only eliminates regions with blurred boundaries or incomplete structures but also improves the stability and reliability of feature regions in spatial registration.
[0074] An directional gradient distribution analysis is performed on the retained candidate regions. By constructing a gradient direction histogram within the region, the gradient change trend of feature points in multiple directions is evaluated to determine whether the region has a significant main directional structure or multi-directional texture interlacing features. If the gradient direction within the region is highly concentrated or exhibits uniformity, it indicates that the texture directionality of the region is too strong, and matching failure is easily caused by angle changes; while regions with multi-directional interlacing and a clear main direction indicate that they have good directional responsiveness and anti-rotation ability, making them suitable as stable feature recognition regions. This process can further filter out the "directional convergence" problem exhibited in repetitive structures, thereby enhancing the spatial uniqueness of feature points.
[0075] The structural integrity of feature points is evaluated based on the local neighborhood density distribution of candidate regions. Specifically, for each candidate feature point, the number, distribution direction, and response value changes of neighboring feature points within a certain radius are statistically analyzed. If the feature points in this neighborhood are evenly distributed, have diverse directions, and exhibit small differences in response values, it indicates that the region possesses good structural consistency and feature stability. Conversely, if the feature point density is too high, the direction is concentrated, or the response fluctuates significantly, it indicates that the region may have texture overlap or local interference, resulting in poor matching stability, and should be filtered out. Neighborhood density analysis effectively improves the ability to discriminate local structural integrity and feature representation capabilities.
[0076] The high structural integrity regions selected through the three evaluation steps described above are used as the second candidate feature region set, providing the input basis for subsequent image registration and error analysis steps. When outputting this set, a structural integrity score is assigned to each region to serve as a confidence reference parameter for dynamic adjustment in subsequent processes.
[0077] Through the above structural consistency evaluation process, not only is the geometric stability and feature representation ability of each region comprehensively evaluated from multiple dimensions such as spatial contour, directional characteristics and neighborhood distribution, but the feature recognition misjudgment problem caused by structural repetition regions is also effectively avoided, significantly improving the accuracy of image feature point matching and the robustness of pose estimation.
[0078] This step aims to conduct a deeper geometric structure analysis and stability evaluation of the initially selected first candidate feature region set, thereby further improving the spatial discriminative power and matching reliability of image feature points. Although the first candidate feature regions have already eliminated regions with overly simple or repetitive textures through texture information entropy filtering, in complex engineering scenarios, there may still be some regions with blurred structural boundaries, unclear orientations, or uneven feature distributions. Feature points in these regions may still lead to erroneous associations in subsequent matching processes due to incomplete structures or poor orientation consistency. Therefore, this step introduces three structural indicators—region boundary contour, directional gradient distribution, and local neighborhood density—to conduct a multi-level evaluation of the image structure of the candidate regions, selecting regions with clear geometric features, stable directional responses, and reasonable feature point distributions as the second candidate feature region set.
[0079] Among these, boundary contour analysis identifies whether a region possesses a closed and continuous geometric structure, avoiding feature instability caused by blurred edges; directional gradient distribution evaluation detects texture directionality, filtering out regions with overly unidirectional orientations or those easily affected by viewing angles; and local neighborhood density analysis, starting from the spatial distribution of feature points, determines whether there is feature overlap or dense interference in local areas, thereby improving the distribution balance and anti-interference ability of feature points. Through the fusion of these three evaluation results, the system can effectively eliminate feature regions with incomplete structures, blurred orientations, or dense repetitions, retaining only those image regions with clear structural expression and strong spatial independence as the final key input. This process not only optimizes the spatial structure quality of feature points but also provides a more stable and reliable feature foundation for image registration and pose estimation. It is an indispensable key link in the entire visual correction process, significantly enhancing the system's matching robustness and recognition accuracy in complex environments.
[0080] S3. Perform image registration based on the second candidate feature region set, construct an image registration error distribution map, calculate the feature point matching residual value in each region, identify the region in the matching error set, and generate an error-sensitive region label map.
[0081] To identify the distribution of registration errors in an image caused by unstable feature matching or structural repetition interference, image registration is performed based on a pre-selected set of second candidate feature regions. During this process, a registration error distribution map is constructed. The matching accuracy of each region is analyzed through matching residual calculation, thereby generating an error-sensitive region label map reflecting the spatial distribution characteristics of the errors. This provides refined residual reference information for subsequent confidence adjustment and pose correction. The processing flow includes the following steps:
[0082] Using all high-quality image feature points in the second candidate feature region set, an image-to-image feature matching process is performed. Specifically, feature points in the current frame image are matched one-to-one with feature points in the reference frame image or digital modeling map. The matching process employs a two-way verification mechanism and a distance ratio filtering strategy to ensure that the similarity of matched point pairs in the descriptor space is sufficiently high and that they are unique. To further improve the geometric consistency of the matching, after the matched point pairs are established, a geometric verification is performed on the point pair set through the calculation of the fundamental matrix or homography matrix to eliminate abnormal matching pairs that are outside the spatial consistency range, ensuring the accuracy and physical rationality of subsequent residual evaluation.
[0083] The purpose of employing a two-way verification mechanism and a distance ratio screening strategy is to improve the accuracy and uniqueness of image feature matching, effectively avoid mismatch problems caused by insufficient similarity of feature descriptions or structural repetition, and ensure that the matching point pairs finally used for registration and pose calculation have high spatial consistency and descriptor reliability.
[0084] In practice, initial matching is performed first. For each feature point extracted from the current frame image, a descriptor metric (such as Euclidean distance or Hamming distance) is used to find its best matching point in the reference frame image or digitally modeled map. Next, during bidirectional verification, the established matching point pairs are checked in reverse. That is, starting from the feature points in the reference frame or map, the best matching point in the current frame is searched again. Only when both directions consider each other as the best match is the point pair retained as a valid match. This mechanism eliminates the risk of asymmetry and false matches in the matching results. Simultaneously, a distance ratio filtering strategy is used. For the best and second-best matches of each feature point, the distance ratio between them is calculated. If the distance ratio between the best and second-best matches is less than a set threshold (e.g., 0.7), it indicates that the match has strong uniqueness and discriminative power, and the point pair is retained. Otherwise, it is considered a low-confidence match pair that may be ambiguous and is discarded. The synergistic application of these two screening mechanisms can significantly improve the reliability and geometric consistency of feature point matching pairs, providing a high-quality input foundation for subsequent image registration, error modeling, and pose estimation. This is especially suitable for complex mapping environments with repetitive structures and texture similarities.
[0085] After initial registration, residual calculation is performed on all geometrically verified matching point pairs. Specifically, for each matching point pair, the Euclidean distance between the transformed projected position of the feature point in the current frame and the true position in the target frame is calculated, and this distance is used as the matching residual value for that point. Subsequently, all residual values are mapped back to the spatial distribution of the original image according to their corresponding spatial coordinates, constructing a registration error distribution map. This map uses image coordinates as a reference and records the average residual, residual variance, and local outlier density of each region to quantify the stability and reliability of each region in the registration process.
[0086] Error clustering analysis was performed on the registration error distribution map. A sliding window and density estimation technique was used to divide the image into blocks, calculating the residual concentration and the proportion of high-error points within each local window region. Regions with high residual concentration and average residual values far exceeding a set threshold were marked as potential high-incidence areas of mismatch. To improve the accuracy of error identification, an inter-region comparison mechanism was introduced, comparing local residual values with the mean residual values of their surrounding regions to identify error abrupt change areas and residual abnormal boundaries, thereby accurately defining the range of image sub-regions with concentrated errors.
[0087] The purpose of error clustering analysis on registration error distribution maps is to accurately identify local regions in images where mismatches are highly concentrated or frequently occur, thereby achieving quantitative perception and explicit labeling of registration risks at specific spatial locations. Traditional registration error evaluation often relies on global average residual values, which cannot effectively reveal local error fluctuations, especially in complex environments such as structural repetition or occlusion, where errors may be highly localized. Therefore, a more refined spatial analysis mechanism is needed. To this end, sliding window and density estimation techniques are used to segment the registration error distribution map.
[0088] The specific steps are as follows:
[0089] The entire image is scanned by sliding a window of a fixed size (e.g., 32×32 pixels) to gradually cover the entire image. Within each sliding window area, the registration residual values of all valid matching points in that area are counted, and the proportion of high residual points (i.e., residual values higher than a set threshold) is further calculated. At the same time, the residual mean, standard deviation, and residual change trend of that area are evaluated.
[0090] Density estimation methods are used to determine whether high-error points exhibit a clustering pattern. For example, kernel density estimation or local peak detection techniques can be employed to quantify the spatial concentration of high residual points. If the proportion of high-error points within a certain window is significantly higher than that of the surrounding area or forms a density peak, then that window is marked as an error clustering region. This process enables spatial identification and labeling of error anomalies, effectively supporting subsequent confidence backoff and pose error adjustment mechanisms, and significantly improving the responsiveness and robustness to local anomalies in complex texture backgrounds.
[0091] Based on the above identification results, an error-sensitive region label map is constructed. This label map assigns an error label to each image region based on image pixel coordinates; the label value represents the matching residual level or error risk level of that region. The label map retains the location, boundary shape, and confidence index of high-error regions, providing precise spatial guidance for subsequent confidence backoff, feature point removal, and weight adjustment.
[0092] By using this error-sensitive region label map, the transformation from feature point registration error to explicit expression of image spatial error is realized, so that subsequent processing no longer depends on global average error, but has the ability to accurately perceive local error changes.
[0093] The purpose of this step is to systematically identify abnormal regions in the image where matching errors are concentrated after image registration using high-quality feature points. This allows for the construction of an error-sensitive region label map that expresses the spatial distribution of errors, providing refined data support for error control and dynamic adjustment in subsequent pose estimation. Although the second candidate feature region set has already screened feature points for structural integrity, in actual registration, factors such as texture similarity, occlusion, lighting changes, or viewpoint deviations may still lead to inaccurate matching of some feature points. These mismatched regions often exhibit locality and clustering; if not identified and isolated, they can easily interfere with the overall results in subsequent pose calculations, even causing serious problems such as estimation drift and orientation reversal. Therefore, the key task of this step is to perform spatial modeling and cluster analysis on the residual information from the image registration process.
[0094] Specifically, the process begins by precisely matching feature points between the current frame and the reference frame, and then using geometric consistency verification to remove some erroneous point pairs. Next, the registration residual value, representing the positional deviation of the transformed points, is calculated for all remaining matching point pairs. Subsequently, spatial analysis methods such as sliding windowing and density analysis are used to perform localized block processing on the image, statistically analyzing the density and concentration of high residual points within each window to identify regions with concentrated errors. Finally, these high-error regions are explicitly labeled on the image, generating an error-sensitive region label map. This map serves as the input for subsequent critical steps such as feature point removal, confidence backoff, and weight adjustment. This label map is a spatial representation of mismatch sensitivity, enabling pose estimation to move beyond a global mean strategy and achieve an adjustment mechanism based on dynamic feedback of local errors.
[0095] In summary, this step not only improves the ability to identify local error risks, but also introduces a quantifiable, traceable, and adjustable error control mechanism for the overall visual navigation system, significantly enhancing the positioning accuracy and robustness of UAVs in complex foundation pit scenarios, demonstrating high engineering practical value and technological innovation.
[0096] S4. Based on the error-sensitive region label map, execute the dynamic confidence backoff process to locally exclude image feature points in the high error region, and construct a confidence interpolation weight map based on the feature point distribution in the neighboring low error region, and output the local confidence adjustment matrix.
[0097] To further enhance the robustness of the pose estimation process to mismatch perturbations, the system constructs a dynamic confidence backoff mechanism based on the error-sensitive region label map generated in the previous stage. This mechanism performs feature point removal and confidence adjustment operations on local regions in the image where mismatches frequently occur. Through this mechanism, based on the identification of error-prone regions, the influence weight of untrusted feature points on the estimation model can be locally reduced, while the contribution of trustworthy regions is enhanced through interpolation inference, thereby improving the stability and accuracy of the entire pose estimation process. The processing flow includes the following steps:
[0098] Based on the error-sensitive region label map, regions marked with high matching residuals in the image are identified, i.e., error-sensitive regions. In this stage, pixel-level or region-level encapsulation analysis is performed on each labeled region to determine the region boundary range and error level. For these regions, all feature points located within the error-sensitive regions are selected from the feature point set in the current image and marked as "low-confidence points." To avoid these error points causing weight interference in subsequent pose estimation, a local exclusion operation is performed on them, temporarily excluding them from subsequent pose estimation calculations, thus isolating potential mismatches at the source.
[0099] To compensate for the feature sparsity problem that may result from excluding feature points in high-error regions, low-error feature points are extracted from the neighboring regions of the error-sensitive area to construct a feature point set that can be used for local information compensation. Based on the spatial distribution density, directional consistency, and residual stability of these feature points, their confidence levels are evaluated, and spatial interpolation is performed accordingly. Specifically, methods such as weighted distance functions or Gaussian kernel functions are used to continuously model the spatial weights of high-confidence feature points in the neighboring regions, thereby forming a confidence interpolation weight map covering the entire error-sensitive region.
[0100] The core objective of using methods such as weighted distance functions or Gaussian kernel functions to model the spatial weights of high-confidence feature points in neighboring regions is to achieve continuous and smooth interpolation expansion by leveraging the confidence information of neighboring regions when reliable feature points are lacking in error-sensitive areas. This fills the confidence gap in mismatched areas and improves the continuity and robustness of overall pose estimation. Feature points in error-sensitive regions are excluded due to issues such as unstable matching or excessive residuals, directly causing a "gap" in the feature contribution of this region. Without confidence compensation for this region, local estimation gaps will inevitably occur, affecting the accuracy and stability of global registration. Therefore, by introducing weighted distance functions or Gaussian kernel functions, the influence range of the confidence of neighboring high-confidence feature points can be extended into the sensitive region with continuous weights based on their location and distance. Points closer to each other have a greater impact on the interpolation result, while the impact of points farther away gradually weakens, thus ensuring that the interpolation has physical meaning and spatial smoothness. The final generated confidence interpolation weight map not only restores the confidence continuity in the sensitive area, but also provides a spatial weight basis for subsequent feature point weighting, error adjustment and attitude fusion. It is a key technical means to achieve a balance between local suppression of error interference and global stability control of the system.
[0101] Based on the generated confidence interpolation weight map and the image coordinate system, each pixel or sub-region is assigned a confidence value derived from its neighboring points, thus generating a complete local confidence adjustment matrix. This matrix not only describes the confidence weight distribution of each region in the current image but also serves as an important parameter input for weighted error control in subsequent pose calculations. At this stage, the residuals of each region can be weighted and fused according to the adjustment matrix, so that the error contribution and spatial confidence form a dynamic coupling relationship, achieving reliable enhancement and unreliable suppression of error information.
[0102] To ensure the stability and adaptability of the confidence backoff mechanism, a dynamic update mechanism based on residual trend feedback was designed. This mechanism monitors the changing trend of image registration error energy in real time and dynamically adjusts the threshold for judging error-sensitive regions and the weighting factors of the interpolation strategy. Based on the historical error distribution of each frame or multiple frame sequences, this mechanism determines whether error fluctuations are stabilizing and appropriately shrinks or expands the response range of the error region, ensuring that the confidence adjustment strategy has sufficient adaptability and fault tolerance. Finally, the output local confidence adjustment matrix is used as a weight adjustment parameter input to the pose estimation model to suppress mismatch interference and enhance the accuracy of pose solving.
[0103] The main function of this step is to construct a dynamic response mechanism to adjust the weight contribution of feature points in sensitive areas with concentrated errors during image registration in real time during pose estimation. This improves the ability to suppress local mismatch interference and maintain the stability of global estimation. In complex tasks such as UAV pit mapping, images often contain numerous repetitive structures, occlusions, and lighting variations, which can easily lead to feature point mismatches. Although the preceding processing has identified these high-risk areas through error-sensitive region labeling, continuing to allow feature points from these areas to participate in pose estimation will introduce significant error noise. Therefore, it is necessary to locally exclude and adjust the confidence level of these high-error areas.
[0104] This step first performs a dynamic confidence backoff operation, selectively removing feature points located within the error-sensitive region from the estimation model to avoid direct interference from these potential mismatches. To prevent data gaps and discontinuities in the estimation results caused by feature removal, the system further models the influence of high-confidence feature points within the spatial range using methods such as weighted distance functions or Gaussian kernel functions, based on the distribution of low-error feature points around the error-sensitive region, generating a confidence interpolation weight map covering the error region. This map effectively propagates information from neighboring confident regions into the error region, achieving confidence compensation for missing regions. The resulting local confidence adjustment matrix can be used as a weighting factor input in the pose estimation model. During error minimization, it automatically reduces the influence of low-confidence regions and increases the dominance of high-confidence regions, thereby achieving targeted suppression of registration errors and stable optimization of pose estimation results.
[0105] S5. Input the local confidence adjustment matrix into the pose calculation process, combine the spatial distribution density of image feature points with the matching residual strength to generate a local confidence weight distribution map, and simultaneously calculate the matching residuals of all image feature points to obtain the overall image matching error energy map.
[0106] To differentiate the impact of feature points in different regions of the image on error and enhance the adaptability of pose estimation to local mismatch perturbations, the local confidence adjustment matrix generated in the previous step is input into the pose calculation process. This matrix is combined with the spatial distribution density of feature points and the matching residual strength to construct a local confidence weight distribution map. Based on this, the matching residual values of all feature points in the image are simultaneously calculated to form an overall image matching error energy map. This process specifically includes the following steps:
[0107] The system receives and loads the local confidence adjustment matrix output from the previous stage. This matrix contains the confidence weight value for each image region or pixel unit, reflecting the confidence level of each location in the current image participating in pose estimation. This matrix is then matched against the established set of all feature points in the current image, assigning each feature point a confidence adjustment weight corresponding to its location. This ensures that the error contribution of feature points in the pose estimation model is no longer treated equally, but rather differentiated based on their error risk level in the image, achieving error-driven feature point weighting.
[0108] For all feature points with assigned confidence levels, local confidence fusion modeling is performed by combining their distribution density and geometric location in the image space. Specifically, in the image coordinate system, the number of feature points, directional gradient distribution, and response value intensity are statistically analyzed within each sub-region according to a region division method. These indicators are then used to locally smooth the region confidence values, forming a locally confident weight distribution map with good continuity and clear boundaries. This map not only retains the error label information in the original adjustment matrix but also rationalizes the weights of sparsely populated local feature regions through a spatial density compensation strategy, enhancing the completeness and granularity of the image confidence representation.
[0109] After constructing the local confidence weight distribution map, error analysis and residual accumulation processing are performed on all feature point matching pairs. For each feature point pair, after a matching relationship is formed between the current frame and the reference frame, the single-point matching residual is calculated based on the Euclidean distance between its transformed projection displacement and the target position. The matching residuals of all valid feature points are recorded and weighted with their respective confidence weights to form a residual field indexed by pixel coordinates. This residual field, through region aggregation and statistical modeling, generates an overall matching error energy map covering the entire image. The pixel value of each region in the map represents the cumulative energy level of the feature point residual at that location, used to express the spatial intensity distribution characteristics of the overall image registration error.
[0110] The core function of this step is to transition image feature point error control from the "individual feature point level" to the "spatial region structure level." It utilizes a local confidence adjustment matrix to dynamically adjust the weights of feature points with differentiation, thereby constructing a robust optimization mechanism that responds to error changes and possesses spatial coupling relationships during pose estimation. In complex foundation pit mapping scenarios, due to the presence of numerous repetitive structural components and textured regions, mismatches are often unavoidable in the matching process between image feature points. These mismatches can interfere with the pose estimation system in the form of local clusters, causing pose calculation shifts, jumps, or even system non-convergence. Traditional approaches typically use a uniform weight model or a simple elimination mechanism to control the impact of error points, making it difficult to flexibly adjust according to spatial variations in error risk.
[0111] This step inputs the local confidence adjustment matrix into the pose calculation process, so that the error contribution weight of each feature point is no longer a static constant, but dynamically determined by the confidence value and mismatch risk of its region. Simultaneously, by combining the spatial distribution density of image feature points with the matching residual intensity, it can identify whether feature points are concentrated in a certain local area, and whether there are anomalous point clusters with excessively high density but abnormal residuals. This information is incorporated into the weighted modeling, further strengthening the feature influence of high-confidence regions and weakening the interference ability of low-confidence regions. By constructing a local confidence weight distribution map, a spatialized expression model of the contribution of feature points is obtained, setting differentiated influence parameters for each sub-region of the image.
[0112] Furthermore, the matching residuals of all image feature points are calculated simultaneously and then weighted and fused with the confidence scores to form an overall image matching error energy map. This energy map comprehensively presents the error intensity distribution of the entire image under the current registration state, which is not only used for the identification of mismatched regions, but also provides data support for the adaptive control of subsequent multi-scale residual adjustment and pose calculation convergence strategies.
[0113] S6, based on the spatial coupling relationship between the overall image matching error energy map and the local confidence weight distribution map, constructs a multi-scale residual adjustment strategy, dynamically updates the image feature point removal threshold and pose calculation convergence parameters, and realizes adaptive adjustment of UAV attitude estimation error under structural repetition interference;
[0114] To address issues such as uneven residual distribution, strong error clustering, and decreased estimation stability during UAV image matching under interference from repetitive structural components, a multi-scale residual adjustment strategy is constructed based on the spatial coupling relationship between the overall image matching error energy map and the local confidence weight distribution map generated in the previous steps. This strategy dynamically updates the image feature point removal threshold and pose calculation convergence parameters accordingly. This strategy not only achieves a closed-loop adjustment process from local error identification to global estimation optimization but also possesses adaptive capabilities, automatically adjusting parameters based on real-time changes in error states in different image scenarios, thereby improving the accuracy and robustness of pose calculation in complex environments. The specific implementation steps are as follows:
[0115] Spatial coupling analysis is performed on the overall matching error energy map and the local confidence weight distribution map of the image. Specifically, the two are registered at the pixel level in the image coordinate system, and joint statistical modeling is performed. Based on the coupling model, the correspondence between residual intensity and confidence level in each image region is calculated, thereby identifying three typical spatial states: "high residual-low confidence" region (potential mismatch core region), "high residual-high confidence" region (possible structural interference region), and "low residual-high confidence" region (credible reference region). This provides a regional classification basis for subsequent residual adjustment.
[0116] A multi-scale residual adjustment framework is constructed, defining hierarchical feature point removal threshold strategies for regions with different spatial scales and error levels in the image. Within a small-scale window (e.g., 32×32 pixels), the removal threshold is dynamically set based on the changes in local variance and confidence gradient of the residuals to accurately remove abnormally high residual feature points. In a medium-scale region (e.g., 128×128 pixels), the weight of the overall region or its labeling as a "weak feature region" is determined by combining regional statistical information and historical error trends. At the large-scale global image level, the global removal ratio and threshold benchmark are dynamically evaluated based on the error energy change trend between the current frame and previous frames to ensure an overall balance between feature point distribution quality and error control capability across the entire image.
[0117] Based on the multi-scale adjustment results, the convergence parameters in the pose calculation process are updated in real time, including but not limited to key values such as the maximum number of iterations in the optimization algorithm, the error convergence tolerance threshold, and the residual reweighting coefficient. For regions with high error density and unstable matching states, the convergence criteria are appropriately tightened, and the number of iterations is forcibly increased to ensure the stability of the solution results; for image scenes dominated by high-confidence regions, the iteration tolerance can be appropriately relaxed to improve the solution efficiency. This parameter adjustment process is dynamically adjusted based on the actual error state of each frame of the image, and has the ability to protect the stability of the estimation in real time.
[0118] The adjusted feature point removal threshold and pose calculation parameters are simultaneously applied to the pose estimation process of the current image frame, and the processing result of this frame is fed back to the error history record to update the prior adjustment strategy in the next frame, thus realizing the continuous evolution and optimization of the estimation strategy. Through this multi-scale residual adjustment mechanism, the system can not only effectively suppress the influence of mismatch in the context of repeated structural interference and improve the robustness of estimation, but also has frame-by-frame self-learning and adaptive capabilities, providing technical support for stable mapping of UAVs in complex environments.
[0119] This step aims to construct a pose estimation optimization mechanism with spatial awareness and dynamic adjustment capabilities. Utilizing the spatial coupling relationship between the overall image matching error energy map and the local confidence weight distribution map, it identifies potential mismatch interference regions in the image in real time and formulates refined, multi-scale residual adjustment strategies accordingly. This enables adaptive control of UAV attitude estimation errors in complex visual scenarios such as repetitive structures. In actual UAV mapping tasks, especially when facing repetitive pattern areas such as equally spaced components, regular steel mesh, and prefabricated structures frequently encountered in foundation pit engineering, conventional pose estimation algorithms struggle to distinguish spatial differences between similar feature points. This can easily lead to problems such as feature point mismatch, attitude calculation drift, and even abnormal flight control paths, severely impacting the accuracy and stability of 3D modeling and orthophoto stitching.
[0120] This step introduces an overall image matching error energy map to comprehensively quantify the error distribution of the current image frame during the registration process. It then performs spatial fusion analysis with the local confidence weight distribution map to identify key interference areas such as "high error - low confidence" regions and "error abrupt change zones." Simultaneously, it implements hierarchical control of error trends at different spatial scales, from setting local point removal thresholds at a small scale, to controlling regional error accumulation at a medium scale, and finally updating the global estimation strategy at a large scale, achieving dynamic error perception and differentiated response across the entire image. This process dynamically updates feature point removal thresholds, avoiding accidental or missed deletions of abnormal points due to improper static threshold settings. Furthermore, it adjusts the convergence parameters in the pose calculation process based on the error status, ensuring sufficient robustness against interference without sacrificing computational efficiency.
[0121] Ultimately, this strategy enables adaptive capabilities during each frame of image processing, continuously learning from error feedback and gradually forming an understanding of the current mapping environment and a self-adjusting estimation model. This frame-by-frame feedback-adjustment-optimization closed-loop mechanism significantly enhances the system's attitude estimation accuracy and stability in dynamic, structurally complex, and texture-overlapping mapping environments, ensuring a high degree of consistency between UAV trajectory control and image registration. It is a key component in achieving robust operation and high-precision estimation within the entire pose correction method.
[0122] The aforementioned visual feature matching-based UAV pose correction method for foundation pit mapping effectively suppresses image feature matching errors and significantly improves pose estimation accuracy under conditions of structural repetition interference. Starting with feature extraction from the image source, this method employs multi-level processing, including texture entropy filtering, structural consistency analysis, and residual aggregation identification, to construct a closed-loop feedback mechanism from feature selection to error adjustment and then to pose optimization. In particular, the introduction of error-sensitive region labels and confidence interpolation modeling effectively avoids interference from mismatched points on the estimation results. Furthermore, the system dynamically adjusts the feature point removal threshold and solution convergence parameters throughout the estimation process, enabling it to adaptively cope with texture interference and error distribution fluctuations in different scenarios. Ultimately, this method achieves high-precision control and stability assurance of UAV flight path and image registration results in complex engineering environments, demonstrating significant advantages such as strong robustness, high adaptability, and strong engineering feasibility.
[0123] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0124] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
[0125] It should be noted that, in this document, the use of relational terms such as "first" and "second" is merely for distinguishing one entity or operation from another, and does not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0126] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0127] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0128] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0129] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0130] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0131] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0132] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A method for UAV pose correction in foundation pit surveying based on visual feature matching, characterized in that, Includes the following steps: S1, Obtain the current image frame, extract regions with high texture complexity as salient regions based on the texture information entropy distribution in the image, remove feature points in regions with low texture variability, and generate the first candidate feature region set; S2, based on the first candidate feature region set, combined with the region boundary contour, directional gradient distribution and local neighborhood density, structural consistency evaluation is performed, and image feature points with high structural integrity are selected to generate the second candidate feature region set; S3. Based on the second candidate feature region set, perform image registration, construct an image registration error distribution map, calculate the matching residual value of image feature points in each region, identify the region in the matching error set, and generate an error-sensitive region label map. S4. Based on the error-sensitive region label map, a dynamic confidence backoff process is executed. Image feature points in the region with high matching residual values are locally excluded. A confidence interpolation weight map is constructed based on the distribution of image feature points in the neighboring region with low matching residual values. The local confidence adjustment matrix is then output. Step S4 includes: Identify high-matching residual regions based on the error-sensitive region label map, filter out low-confidence feature points within the regions, and locally exclude them; Low-error feature points are extracted from the neighboring regions of the error-sensitive region, a feature point set is constructed, and the confidence level is evaluated based on the distribution density, orientation consistency, and residual stability. Based on the confidence of neighboring feature points, weighted distance modeling or Gaussian kernel interpolation is performed to generate a confidence interpolation weight map. A local confidence adjustment matrix is generated based on the confidence interpolation weight map, providing input parameters for weighted error control in subsequent pose calculations; S5, input the local confidence adjustment matrix into the pose calculation process, combine the spatial distribution density of image feature points with the matching residual strength to generate a local confidence weight distribution map, and simultaneously calculate the matching residuals of all image feature points to obtain the overall image matching error energy map; S6. Based on the spatial coupling relationship between the overall image matching error energy map and the local confidence weight distribution map, a multi-scale residual adjustment strategy is constructed to dynamically update the image feature point removal threshold and pose calculation convergence parameters. Step S6 includes: Spatial coupling analysis is performed on the overall matching error energy map and the local confidence weight distribution map of the image to calculate the correspondence between the residual intensity and confidence level in each image region and identify different error regions; A multi-scale residual adjustment framework is constructed, which dynamically sets the feature point removal threshold in the small-scale region, adjusts the regional weights in the medium-scale region, and evaluates the overall map error control in the large-scale region. Based on the multi-scale adjustment results, the convergence parameters in the pose calculation process are updated in real time to ensure error control and calculation stability. The adjusted feature point removal threshold and pose calculation parameters are applied to the pose estimation process of the current image frame, and the processing results are fed back to the error history record to update the adjustment strategy for the next frame.
2. The method for UAV pose correction for foundation pit mapping based on visual feature matching according to claim 1, characterized in that, Step S1 includes: The current image frame is acquired, and the image is locally segmented using a window sliding method. A gray-level co-occurrence matrix is constructed, and the texture information entropy value is calculated to generate a texture information entropy heatmap. Based on the texture information entropy heatmap, a saliency threshold is set to extract salient regions with high texture complexity and construct a salient region mask map. Image feature points are extracted based on the salient region mask map constraint, and a local feature response value filtering mechanism is introduced to remove feature points with low response intensity, unstable orientation, and non-repeatable scale. Spatial clustering analysis is performed on the selected image feature points. Density suppression is performed based on the spatial density distribution of the clustered regions to remove redundant points and generate a set of first candidate feature regions.
3. The method for UAV pose correction for foundation pit mapping based on visual feature matching according to claim 1, characterized in that, Step S2 includes: For each region in the first candidate feature region set, the boundary contour is extracted, the boundary continuity, closure and edge intensity changes are calculated, and regions with incomplete structures are eliminated. Perform directional gradient distribution analysis on the retained regions, construct directional gradient histograms, and filter out regions with concentrated directions and poor directional responsiveness; Based on the changes in the number, distribution direction, and response value of feature points around each feature point according to the local neighborhood density statistics, regions with poor structural consistency are eliminated to generate a second set of candidate feature regions.
4. The method for UAV pose correction for foundation pit mapping based on visual feature matching according to claim 1, characterized in that, Step S3 includes: Image feature matching is performed using all high-quality image feature points in the second candidate feature region set. A two-way verification mechanism and a distance ratio filtering strategy are used to ensure the uniqueness of matching point pairs in the descriptor space. Perform geometric verification on all matching point pairs and remove abnormal matching points that do not meet spatial consistency requirements; Calculate the registration residual value for each matching point pair, construct a registration error distribution map, and record the residual mean, residual variance, and local outlier density for each region; Error clustering analysis is performed on the registration error distribution map. The sliding window and density estimation techniques are used to calculate the proportion and clustering trend of high residual points in each window area, identify error concentration areas, and generate error-sensitive area label maps.
5. The method for UAV pose correction for foundation pit mapping based on visual feature matching according to claim 1, characterized in that, Step S5 includes: Receive and load the local confidence adjustment matrix, match its position with the set of feature points in the current image, and assign each feature point a confidence adjustment weight corresponding to its position. By combining the spatial distribution density and geometric location of feature points, the number of feature points, directional gradient distribution and response value intensity in each region are statistically analyzed, and local confidence fusion modeling is performed to generate a local confidence weight distribution map. For all feature point matching pairs, the residual values are calculated and weighted and accumulated according to the confidence of each feature point to form a residual field indexed by pixel coordinates; Based on the residual field, regional aggregation and statistical modeling are performed to generate an overall image matching error energy map.
Citation Information
Patent Citations
Dynamic environment self-adaptive intelligent navigation method and system
CN120063287A
Port internal and external fleet positioning system and method
CN120282264A