A narrow field of view unmanned aerial vehicle image accurate positioning method and system based on external attitude prior

CN122835330APending Publication Date: 2026-09-29HENAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610970208.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

但当GPS观测质量较低(如EXIF GPS存在重复、量化噪声)时,其约束效果有限

Benefits of technology

在本实施例所采用的参考评价点条件下,引入外部完整相机姿态软先验后,检查点水平定位RMSE由约49.5m降低至约0.5m,表明该完整姿态框架能够显著提高长焦窄视场影像定位稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122835330A_ABST
    Figure CN122835330A_ABST
Patent Text Reader

Abstract

A narrow field of view unmanned aerial vehicle image precise positioning method and system based on external attitude prior. The method obtains EXIF GPS camera station coordinates and external SfM complete camera attitude (including camera center coordinates and rotation matrix) of a narrow field of view image sequence; the EXIF GPS is converted from WGS-84 to a local three-dimensional rectangular coordinate frame, and is weighted according to repeated GPS and track quality; the external SfM camera center and rotation attitude are unified to the local coordinate frame through quality weighted seven-parameter similarity transformation; a bundle adjustment model containing re-projection error, GPS weak constraint, camera center soft prior and SO (3) rotation soft prior is constructed, and camera parameters and three-dimensional points are optimized; then robust multi-view forward intersection is carried out, and ground point coordinates and reliability are output. The method is suitable for long-focus narrow field of view image positioning under the condition of no measured ground control point and no high-precision IMU / RTK.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of UAV photogrammetry and computer vision technology, specifically a method and system for precise localization of narrow field-of-view UAV images based on external attitude priors. Background Technology

[0002] With the increasing application of drones in fields such as power line inspection, ancient building protection, and geological disaster monitoring, drones equipped with telephoto lenses (narrow field of view) can acquire high-resolution, detailed images from a distance. However, narrow field-of-view images are characterized by a small field of view (typically less than 15°), limited overlap between adjacent images, a weak datum-to-height ratio, and susceptibility to camera attitude drift. In practical operations, consumer-grade or lightweight drones often lack high-precision inertial measurement units (IMUs), real-time dynamic differential positioning (RTK) or post-processing dynamic differential positioning (PPK) tracks and ground control points, and can only read low-precision GPS coordinates from exchangeable image files (EXIF). Existing aerial triangulation algorithms struggle to accurately locate narrow field-of-view images with only image-corresponding points and EXIF ​​GPS data, often resulting in positioning errors on the order of tens of meters.

[0003] In existing technologies, GPS-assisted bundle adjustment introduces camera station coordinates as weighted observations into the error equation, which can constrain the model's translation, scale, and absolute position to some extent. However, its constraint effect is limited when the quality of GPS observations is low (e.g., EXIF ​​GPS contains duplication and quantization noise). Furthermore, while attitude continuity and relative attitude priors can improve local camera relationships, they cannot fundamentally eliminate overall attitude drift. Existing research shows that, without ground control points, the absolute accuracy of UAVSfM results is highly dependent on the initial positioning quality and external attitude constraints. Therefore, how to stably recover the complete camera attitude of narrow field-of-view images without high-precision hardware support is crucial for achieving accurate positioning. Summary of the Invention

[0004] This invention aims to provide a method and system for precise positioning of narrow field-of-view UAV images based on external attitude priors. Under conditions where there are no measured ground control points involved in adjustment, no high-precision IMU / RTK observations, and only low-precision EXIF ​​GPS, this invention addresses the problem of unstable positioning caused by small intersection angles, limited overlap, and easy attitude drift in narrow field-of-view long-focal-length UAV images. It utilizes the complete camera attitude of external SfM to construct a soft-constrained attitude framework, thereby achieving target geolocation based on pixel-selective points.

[0005] To address the above technical problems, the specific solution adopted in this invention is as follows: a method for precise localization of narrow field-of-view UAV images based on external attitude priors. This involves acquiring the EXIF ​​GPS station coordinates and the complete external SfM camera attitude of a narrow field-of-view image sequence, where the complete camera attitude includes the camera center coordinates and rotation matrix of each image; transforming the external SfM complete camera attitude into the EXIF ​​GPS coordinate system using a three-dimensional seven-parameter similarity transformation as the complete attitude prior; using the reprojection error of corresponding points in the images as the primary constraint, the EXIF ​​GPS station coordinates as the weak constraint, and the complete attitude prior as the initial value or a weighted prior for bundle adjustment optimization to calculate the optimized camera parameters; and performing multi-view forward intersection on the same image point in multiple images based on the optimized camera parameters to obtain the ground point coordinates.

[0006] Preferably, it includes the following steps: S1: Obtain the narrow field-of-view image sequence captured by the drone and the corresponding EXIF ​​GPS camera station coordinates; S2: Perform feature matching on the image sequence to construct the trajectory of corresponding points and the image connection map; S3: Obtain the complete camera pose of the external SfM, which includes the camera center coordinates and rotation matrix of each image; S4: Through three-dimensional seven-parameter similarity transformation, the attitude of the external SfM complete camera is transformed to the EXIF ​​GPS coordinate system to obtain the transformed complete attitude prior; S5: Using the reprojection error of the corresponding points in the image as the main constraint, the EXIF ​​GPS camera station coordinates as the weak constraint, and the transformed complete attitude prior as the initial value or weighted prior constraint, perform bundle adjustment optimization to solve the optimized camera parameters for each image. S6: Based on the optimized camera parameters, perform multi-view forward intersection on the same image point in multiple images to obtain the coordinates of the ground point.

[0007] Preferably, the feature matching in step S2 includes: firstly, using the SIFT algorithm to perform basic matching, and then using RANSAC to remove outliers; for weakly connected image pairs with fewer than a threshold number of matched inliers, further supplementary matching is performed using the detectorless matching method LoFTR based on Transformer.

[0008] Preferably, the external SfM complete camera pose mentioned in step S3 is derived from the results of processing the same set of images using mature 3D reconstruction software.

[0009] Preferably, the three-dimensional seven-parameter similarity transformation in step S4 is solved by weighted least squares to obtain the scale factor, rotation matrix and translation vector, where the weights are set according to the quality of EXIF ​​GPS.

[0010] Preferably, the bundle adjustment in step S5 also incorporates GPS track smoothing: interpolating the repeated coordinate segments in the EXIF ​​GPS coordinates and smoothing the track to reduce the negative impact of repeated coordinates and quantization noise on weak constraints.

[0011] Preferably, the bundle adjustment in step S5 also incorporates SO(3) relative rotation continuity constraints between adjacent images to suppress abrupt attitude changes.

[0012] Preferably, the narrow field-of-view image refers to UAV image acquired by a telephoto lens with a field of view of less than 15°.

[0013] Preferably, in step S6, during multi-view forward intersection, abnormal observations are identified based on the intersection residuals of each ray, and rays with large deviations are eliminated or downweighted, and a positioning reliability index is output. The reliability index includes the number of participating images, the average intersection residual, the maximum intersection angle, and the bundle adjustment coverage.

[0014] A narrow field-of-view UAV imagery-based precise positioning system based on external attitude priors includes: The image matching module is used to match features in narrow field-of-view image sequences and construct trajectories of corresponding points and image connection maps. The GPS preprocessing module is used to perform repeating segment interpolation and smoothing on EXIF ​​GPS camera station coordinates. The external attitude acquisition and reduction module is used to acquire the complete attitude of the external SfM camera and transform it to the EXIF ​​GPS coordinate system through a three-dimensional seven-parameter similarity transformation to obtain the complete attitude prior. The bundle adjustment module is used to jointly optimize camera parameters and object 3D points with reprojection error as the main constraint, EXIF ​​GPS camera station coordinates as the weak constraint, and complete attitude prior as the initial value or weighted prior. The multi-view forward intersection module is used to perform spatial ray intersection on the same image point in multiple images based on optimized camera parameters, and solve for the coordinates of the ground point. The credibility evaluation module is used to evaluate the credibility of the positioning results based on the number of participating images, intersection residuals, intersection angles, and bundle adjustment coverage.

[0015] Beneficial effects Under the reference evaluation point conditions used in this embodiment, after introducing an external complete camera attitude soft prior, the horizontal positioning RMSE of the check point decreased from about 49.5m to about 0.5m, indicating that the complete attitude framework can significantly improve the positioning stability of telephoto narrow field-of-view images.

[0016] Ablation experiments show that, under the conditions of repeated EXIF ​​GPS and weak intersection in this embodiment, simply increasing the number of matching points, GPS smoothing, or relative attitude constraints is insufficient to fundamentally eliminate overall attitude drift; a complete attitude framework that includes both the camera center and rotational attitude can more effectively constrain the camera network.

[0017] This invention employs a SIFT+LoFTR combined matching strategy, which increases the number of corresponding points in weak and repetitive texture regions several times over, significantly improving the integrity of the image connectivity map and providing more sufficient observation data for subsequent bundle adjustment.

[0018] The complete attitude prior of the external SfM of this invention can be calculated by mature 3D reconstruction software from the same set of images, without the need for expensive equipment such as airborne IMU and RTK, and is suitable for consumer drones. Attached Figure Description

[0019] Figure 1 This is a distribution map of the image centers in the experimental area of ​​this invention, showing the planar distribution of the centers of 68 images in the local area of ​​Fleurac.

[0020] Figure 2 This is a sample image of a multi-view image, showing the different perspectives of the same checkpoint (CP01) in multiple images.

[0021] Figure 3 This is a comparison chart of the original EXIF ​​GPS track and the smoothed GPS track.

[0022] Figure 4 It is a GPS step size variation map of adjacent images, showing the phenomenon of step size approaching zero (repeated GPS) and step size variation being large.

[0023] Figure 5 It is an image matching connectivity heatmap, which reflects the matching strength between different image pairs.

[0024] Figure 6 This is a comparison chart of the number of matching interior points between SIFT and LoFTR.

[0025] Figure 7 It is a baseline and weight distribution map of the MIPMAP relative attitude prior image.

[0026] Figure 8 This is a residual distribution map of the similarity reduction between the MIPMAP camera center and EXIF ​​GPS.

[0027] Figure 9 This is a comparison chart of RMSE for horizontal positioning using different positioning schemes.

[0028] Figure 10 This is a comparison chart of SIFT, LoFTR, and SIFT+LoFTR matched ablation.

[0029] Figure 11 This is a comparison chart of the checkpoint errors between the baseline scheme and the complete attitude prior scheme.

[0030] Figure 12 This is a comparison chart of sensitivity to prior weights of complete pose and multi-look consistency.

[0031] Figure 13 This is a diagnostic diagram of the posture difference between EXIF ​​BA and MIPMAP.

[0032] Figure 14 This is the overall architecture diagram of the visualization positioning system.

[0033] Figure 15 This is an example diagram of the system interface and image click positioning.

[0034] Figure 16 This is a comparison chart of RMSE for different multi-view positioning strategies.

[0035] Figure 17 It is a geometric diagram of the intersection of multiple views.

[0036] Figure 18 This is a schematic diagram of the narrow field-of-view error amplification mechanism.

[0037] Figure 19 This is a comparison chart of attitude and camera center disturbance sensitivity.

[0038] Figure 20 This is a comparison chart of pixel selection error perturbation sensitivity. Detailed Implementation

[0039] The invention will now be described in further detail with reference to the accompanying drawings and specific experimental data. This embodiment uses 68 long-range UAV images acquired in the Fleurac region of France as experimental data. The image field of view is less than 15° (i.e., the narrow field of view image described in this invention), the image size is 1840×1228 pixels, there are no ground control points or high-precision IMU attitude data, and only low-precision GPS coordinates can be read from EXIF. For example... Figure 1 As shown, the image center distribution area is small, and the baselines of adjacent images are short, which is a typical condition of narrow field of view and weak intersection geometry. The viewing angle changes and local occlusion of the same inspection point in different images are shown below. Figure 2 As shown, this requires the use of multi-view forward intersection in subsequent processing to reduce the point selection error of a single image. Tables 1 and 2 show the image parameters and EXIF ​​GPS quality diagnosis: 42 images have duplicate GPS coordinates (21 sets of duplicate segments), 29 adjacent step lengths are close to zero, the original median step length is 70.309m, and the smoothed median step length is 44.135m.

[0040] Table 1 Experimental images and camera parameters

[0041] Table 2 EXIF ​​GPS Quality Diagnostic Statistics

[0042] S1: Obtain the EXIF ​​GPS station coordinates of the narrow field-of-view image sequence and perform preprocessing.

[0043] GPS coordinates are extracted from the EXIF ​​information of JPG images captured by the UAV. To address the issues of duplicate coordinates and quantization noise in the original EXIF ​​GPS data, interpolation is performed on duplicate coordinate segments and the track is smoothed (i.e., GPS track smoothing). Figure 3 Comparing the tracks before and after smoothing, it was found that the original tracks had local repetitions and discontinuities (some image station positions overlapped or changed abnormally). The track continuity was improved after interpolation smoothing. Figure 4 The results show the variation in GPS step size between adjacent imagery. Some step sizes approaching zero indicate repeated GPS issues, while abrupt changes in step size indicate the discreteness of EXIF ​​GPS. After smoothing, the average correction of the GPS prior relative to the original coordinates is 14.737m, and the median correction is 20.508m. However, it should be emphasized that these results are only for weak GPS priors and cannot be equated with high-precision RTK / PPK trajectories.

[0044] S2: Perform feature matching on the image sequence to construct the trajectory of corresponding points and the image connection map.

[0045] This step employs a matching strategy combining SIFT and LoFTR. First, candidate matching pairs are generated based on image sequence proximity and GPS distance. SIFT is then used for basic matching, and RANSAC is used to eliminate outliers based on epipolar geometric constraints. The epipolar constraint formula is as follows: For weakly connected image pairs with fewer than a threshold number of matched inliers, a detector-free matching method based on Transformer, LoFTR, is used for supplementary matching. LoFTR utilizes self-attention and cross-attention mechanisms to enhance the global contextual representation of features. The attention formula is as follows: .

[0046] Figure 5 This is a heatmap of image matching connectivity; the wider the connecting edge, the more corresponding points there are. Figure 5 It can be seen that most image pairs have effective connections, but some image pairs have weak connections, which reflects the limited overlap area between adjacent images in narrow field-of-view images. Figure 6Comparing the number of matched inliers between SIFT and LoFTR in some weakly connected image pairs, it is evident that LoFTR can significantly increase the number of inliers. Table 3 lists detailed data for some image pairs. For example, the number of SIFT / RANSAC inliers for image pair 4-5 increased from 1313 to 7016, and for image pair 5-6 from 1504 to 7062, demonstrating that LoFTR can effectively compensate for the insufficient matching of traditional SIFT in weakly textured or repetitive textured regions.

[0047] Table 3. LoFTR Supplementary Matching Intrapoints (Partial)

[0048] Matching ablation experiments (Table 4) further demonstrate that SIFT+LoFTR increases the total number of points within the RANSAC from 164,181 to 517,168, and the number of corresponding point trajectories from 17,898 to 114,083. However, it should be noted that while increasing the number of matching points primarily improves image connectivity, whether it can improve absolute positioning accuracy still needs to be determined in conjunction with subsequent BA results—under a purely weak GPS BA framework, the checkpoint RMSE remains approximately 52m, indicating that increasing the number of matching points does not solve the absolute attitude drift problem.

[0049] Table 4 Ablation Experiment Results

[0050] S3: Obtain the complete camera pose from the external SfM.

[0051] This embodiment obtains the complete camera pose results obtained by pre-processing the same set of 68 images using MIPMAP 3D reconstruction software (a mature commercial 3D reconstruction software, belonging to the "mature 3D reconstruction software" described in this invention). These results include the camera center coordinates (3D coordinates) and rotation matrix (3×3 matrix) for each image. It is important to emphasize that this external SfM complete pose prior is different from the measured ground control points. Its function is to provide a stable camera pose framework for narrow field-of-view images to analyze the impact of the complete pose framework on positioning accuracy. It should be noted that the external SfM complete pose prior is different from the measured ground control points. Its function is to provide a relative pose framework for the camera network within the same image sequence. The checkpoints in this embodiment are only used for positioning accuracy evaluation, do not participate in bundle adjustment, do not participate in seven-parameter reduction, and are not used as control points.

[0052] S4: Transform the attitude of the external SfM complete camera to the EXIF ​​GPS coordinate system through a three-dimensional seven-parameter similarity transformation.

[0053] The external SfM attitude is usually located in an independent model coordinate system and needs to be transformed to a GPS coordinate frame through coordinate transformation. This step uses a three-dimensional seven-parameter similarity transformation, and the transformation formula is as follows:

[0054] Where s is the scale factor, RT is the rotation matrix, and t is the translation vector. The residual term has a weight wi set according to the EXIFGPS quality. By solving the seven parameters using weighted least squares, the complete camera pose of the MIPMAP can be unified into the local coordinate frame used in this embodiment. Figure 8 The distribution of the similarity reduction residuals between the MIPMAP camera center and the EXIF ​​GPS is shown. The reduction residuals of most camera centers are small, indicating that the seven-parameter similarity transformation can align the external SfM attitude frame with the GPS coordinate frame well. However, some images still have large residuals, which are related to the repetition and quantization noise of EXIF ​​GPS. Figure 7 The image-baseline and weight distribution in the MIPMAP relative pose prior are shown. This relative pose prior can be used as an auxiliary constraint, but it cannot fix the global coordinate frame when used alone.

[0055] S5: The bundle adjustment optimization is performed using the reprojection error of the image corresponding points as the main constraint, the EXIF ​​GPS camera station coordinates as the weak constraint, and the transformed complete attitude prior as the initial value or weighted prior constraint.

[0056] Bundle adjustment is a core method in photogrammetry for estimating camera parameters and 3D point coordinates. Its basic model is based on the collinearity equation:

[0057] The objective function for bundle adjustment in this step includes three constraints: (1) Master constraint for reprojection error: Minimize the reprojection error of all image point observations, with the objective function being... , where K is the camera intrinsic parameter matrix, Ri is the camera rotation matrix, Ci is the camera center, and ρ is the robust kernel function.

[0058] (2) EXIF ​​GPS station weak constraints: The GPS station coordinates preprocessed in step S1 are added as weak observations to the adjustment model, and the augmented objective function is: Because EXIFGPS has limited accuracy and suffers from duplicate coordinates, this embodiment uses it as a weak constraint rather than a high-precision POS ground truth value.

[0059] (3) Complete attitude prior constraints: The complete attitude prior transformed in step S4 is used as the initial value or weighted prior constraints and added to the adjustment. The complete attitude prior provides both the camera center and rotation matrix, providing a stable initial attitude framework for the entire camera network.

[0060] As an optional enhancement, this step can also incorporate SO(3) relative rotation continuity constraints between adjacent images to suppress abrupt attitude changes, with the residual formula being: .

[0061] To verify the effects of different constraints, four ablation experiments were designed in this embodiment, and the results are shown in Table 5. Figure 9 As shown: Table 5 Ablation Experiment Results of Different Positioning Schemes

[0062] Option A (weak constraints only, EXIF ​​GPS): Using only image-related points and weak priors from EXIF ​​GPS stations, the checkpoint level RMSE is 49.544m.

[0063] Option B (GPS Smoothing + Attitude Continuity BA): Based on Option A, track smoothing and adjacent attitude continuity constraints are added. The RMSE at the checkpoint is 52.890m, with no significant improvement.

[0064] Option C (MIPMAP Relative Attitude Prior): Further introduces the relative rotation or relative translation direction between image pairs. The checkpoint RMSE is 51.932m, which is still in the tens of meters range.

[0065] Option D (MIPMAP Complete Attitude Prior): Using the complete camera center and rotation matrix from the external SfM results as the attitude framework and recalculating with GPS seven parameters, the checkpoint RMSE is reduced to 0.492m, and the positioning accuracy is improved by about 99%.

[0066] Figure 11 Further comparisons were made between the baseline scheme (Scheme A) and the complete attitude prior scheme (Scheme D) at two checkpoints (CP01 and CP02). The latter's errors decreased from 38.884m and 58.286m to 0.572m and 0.394m, respectively. Table 6 lists the details of the positioning errors and estimated latitude and longitude of the two checkpoints. Figure 9 A direct comparison of the positioning errors of the four schemes shows that the complete camera attitude frame has a decisive influence on the positioning stability of telephoto narrow field-of-view images.

[0067] Table 6. Details of Checkpoint Positioning Errors

[0068] The weight sensitivity experiment of the complete attitude prior (Table 7) shows that, when the complete attitude framework is relatively stable, the checkpoint RMSE is stable between 0.489 and 0.492 m under different prior weights (rotation σ from 0.02 to 0.32, translation σ from 3.000 to 40.000), the joint intersection RMSE is stable at 0.705 m, and the leave-one-out method mean displacement is approximately 0.078 m, indicating that this method is not sensitive to the selection of attitude prior weights. However, it is worth noting that if the camera center after MIPMAP reduction is completely fixed and not included in the optimization (i.e., the "fixed center" scheme), the checkpoint RMSE increases to 69.728 m, and the mean intersection residual increases to 4.331 m. This indicates that the external SfM results cannot be simply locked and used, but should be used as the complete attitude prior and initial framework, and then fine-tuned by combining image reprojection error and GPS seven-parameter reduction.

[0069] Table 7. Sensitivity of complete pose prior weights and results of multi-look consistency.

[0070] To explain the significant differences between the weak GPS BA (Scheme A) and the complete attitude prior scheme (Scheme D), this embodiment performs attitude difference diagnosis. Figure 13 The results of the differential diagnosis between EXIF ​​BA and MIPMAP attitude are presented. The direct camera center horizontal RMSE is 25.259m, the local center RMSE after seven-parameter similarity transformation is 25.642m, and the median rotation difference is 24.995°. Significant differences in attitude exist, with some images exhibiting large camera center and rotation differences. Table 8 lists representative images with significant differences, such as image "8130480.jpg" with a local center difference of 78.695m and an attitude difference of 22.776°; and image "8130448.jpg" with a center difference of 42.474m and an attitude difference of 29.865°. These data indicate that the error of pure weak GPS BA is not a simple overall translational deviation, but includes significant differences in camera center and rotational attitude. For telephoto and narrow field-of-view images, these attitude differences are amplified during image point backprojection and multi-view intersection, ultimately leading to ground positioning errors on the order of tens of meters.

[0071] Table 8. Images with significant pose differences between EXIF ​​weakly constrained BA and MIPMAP. Figure 18 The mechanism of narrow field-of-view error amplification is illustrated: due to the small intersection angle of light rays, even a small image point error or attitude error can lead to a large positional deviation in the object space. Disturbance sensitivity experiment (Table 9); Figure 19This was quantitatively verified: the RMSE of the checkpoint was 0.492m without disturbance, increasing to 1.338m after adding a 0.2° random attitude disturbance, 12.138m at 0.5°, 22.765m at 1°, and 28.110m at 2°. This indicates that even small attitude errors in telephoto narrow field-of-view images can be amplified into significant ground positioning errors. The influence of camera center disturbance is related to factors such as the disturbance direction and the spatial location of the checkpoint, and its performance is less monotonic than that of attitude disturbance. Meanwhile, pixel selection error disturbance experiments (Table 10) were also conducted. Figure 20 The results show that the RMSE of the multi-view intersection checkpoint is 0.708m without disturbance, 0.712m with a 0.5px disturbance, 0.735m with a 1px disturbance, 1.056m with a 5px disturbance, and 1.629m with a 10px disturbance. This indicates that when the number of multi-view observations is sufficient, the system has a certain robustness to random point selection errors within 1px, but the error increases significantly beyond 5px. Therefore, the image scaling point selection function provided by this invention is of practical significance for maintaining positioning accuracy.

[0072] Table 9. Influence of Attitude and Camera Center Disturbances on Positioning Accuracy

[0073] Table 10. Impact of Pixel Selection Error Perturbation on Multi-Look Positioning Accuracy

[0074] S6: Based on the optimized camera parameters, perform multi-view forward intersection on the same image point in multiple images to obtain the coordinates of the ground point and output the positioning reliability index.

[0075] Given that the camera's internal and external orientation elements are known, image points on the image can be back-projected into spatial rays. Let the image point coordinates be... Then its back projection process is as follows: image point homogeneous coordinates Camera coordinate system direction vector The direction of the ray after transformation to the object coordinate system is The spatial ray parametric equation is: .

[0076] When the same ground feature is observed in multiple images, its object location can be determined by multiple spatial rays. The objective function of multi-view least squares intersection is: In practice, abnormal observations are identified based on the intersection residuals of each ray, and rays with large deviations are discarded or downweighted, thereby improving the stability of multi-view positioning results. Figure 17This is a geometric diagram of multi-view forward intersection, showing the basic process of back-projecting corresponding image points in multiple images into spatial rays and solving for the object point coordinates using the least squares method. When the camera attitude is accurate and the intersection conditions of multiple rays are good, the intersection point position is relatively stable; conversely, attitude errors and pixel selection errors will cause the rays to fail to intersect stably.

[0077] Multi-view intersection results (Table 11; Figure 16 , Figure 16 The joint intersection error is significantly lower with the complete attitude prior scheme. The results show that under the pure EXIF ​​GPS weakly constrained BA scheme (Scheme A), the joint intersection RMSE is 46.934m; after introducing the external SfM complete attitude prior (Scheme D), the joint intersection RMSE drops to 0.705m, a reduction of approximately 98%. This demonstrates that a stable complete camera attitude framework is a key foundation for achieving pixel-level geolocation.

[0078] Table 11. Details of intersection errors for multiple viewpoints of the same name This step also outputs positioning reliability indicators, including: number of participating images (the number of images participating in multi-view intersection; more images indicate more redundant observations), average intersection residual (the average distance from a spatial point to multiple rays; smaller residuals indicate better ray consistency), maximum intersection angle (the maximum angle between participating rays; a small intersection angle will amplify depth errors), and bundle adjustment coverage (the coverage of the image with corresponding points in the BA; low coverage indicates weaker attitude reliability). See Table 12 for the specific definitions and reliability judgment criteria of these indicators. The system comprehensively judges the reliability of the positioning results based on these indicators and outputs a reliability level prompt to the user.

[0079] Table 12 Location Reliability Evaluation Indicators As an application vehicle of this invention, a visual positioning system was developed. This system includes: an image matching module (for performing feature matching and connectivity graph construction in step S2), a GPS preprocessing module (for performing repeated segment interpolation and smoothing in step S1), an external attitude acquisition and reduction module (for performing external attitude acquisition and seven-parameter transformation in steps S3 and S4), a bundle adjustment module (for performing joint optimization in step S5), a multi-view forward intersection module (for performing spatial ray intersection solution in step S6), and a reliability evaluation module (for calculating and outputting reliability indicators in step S6). The overall system architecture is as follows: Figure 14 As shown, the system comprises a data layer, a model layer, a computational service layer, an interactive display layer, and an evaluation and diagnostic layer. An example of the system interface and image click-and-locate functionality is shown below. Figure 15As shown, it supports image zoom point selection (improving the accuracy of manual pixel point selection), single-point positioning, and multi-view fusion positioning.

[0080] In summary, this invention, under conditions of no ground control points and only low-precision EXIF ​​GPS, reduces the horizontal positioning RMSE of checkpoints in narrow field-of-view UAV imagery from 49.544m to 0.492m (an improvement of approximately 99% in accuracy) and the multi-view forward intersection RMSE from 46.934m to 0.705m (a reduction of approximately 98%) by introducing an external SfM complete attitude prior (after seven-parameter similarity transformation as the initial value and prior for BA), achieving sub-meter-level positioning. This method does not require expensive equipment such as airborne IMUs and RTKs, making it suitable for fine-grained observation scenarios using consumer-grade UAVs.

Claims

1. A method for precise localization of narrow field-of-view UAV images based on external attitude priors, characterized in that, include: Obtain narrow field-of-view UAV image sequences and the corresponding EXIF ​​GPS camera station coordinates for each image; The EXIF ​​GPS station coordinates are converted from WGS-84 ellipsoidal coordinates to spatial rectangular coordinates, and a local three-dimensional rectangular coordinate framework based on the survey area is established. Quality diagnosis is performed on the EXIF ​​GPS station coordinates to identify duplicate station coordinate segments and abnormal step sizes, and quality weights are assigned to each station coordinate based on the diagnosis results. Feature matching and outlier removal are performed on the image sequence to construct the trajectory of corresponding points and an image connectivity map. The complete external SfM camera pose for the same image sequence is obtained, including the camera center coordinates and rotational pose of each image. A three-dimensional seven-parameter similarity transformation is solved based on the quality weights to transform the external SfM camera center and rotational pose to the local three-dimensional rectangular coordinate framework, obtaining a complete pose soft prior. A framework is constructed including image corresponding point reprojection errors and EXIF... The bundle adjustment objective function, based on weak constraints of GPS camera stations, soft priors of camera center, and soft priors of SO(3) rotation, jointly optimizes camera parameters and object-side 3D points while maintaining optimizable camera center and rotation attitude. Based on the optimized camera parameters, robust multi-view forward intersection is performed on the same target image point in multiple images to obtain the ground coordinates of the target point and output the positioning reliability.

2. The method for precise localization of narrow field-of-view UAV images based on external attitude prior, as described in claim 1, is characterized in that: The local three-dimensional rectangular coordinate frame is an ENU coordinate frame. The EXIF ​​GPS station coordinates are first converted from longitude, latitude, and elevation to ECEF spatial rectangular coordinates, and then converted to ENU coordinates with the reference point of the survey area as the origin. The quality weight is determined based on the distance between adjacent stations, the length of consecutive repeated coordinates, and the degree of track step anomaly. The GPS observations corresponding to repeated coordinate segments and abnormal step lengths are assigned lower weights, including the following steps: S1: Obtain the narrow field-of-view image sequence captured by the drone and the corresponding EXIF ​​GPS camera station coordinates; S2: Perform feature matching on the image sequence to construct the trajectory of corresponding points and the image connection map; S3: Obtain the complete camera pose of the external SfM, which includes the camera center coordinates and rotation matrix of each image; S4: Through three-dimensional seven-parameter similarity transformation, the attitude of the external SfM complete camera is transformed to the EXIF ​​GPS coordinate system to obtain the transformed complete attitude prior; S5: Using the reprojection error of the corresponding points in the image as the main constraint, the EXIF ​​GPS camera station coordinates as the weak constraint, and the transformed complete attitude prior as the initial value or weighted prior constraint, perform bundle adjustment optimization to solve the optimized camera parameters for each image. S6: Based on the optimized camera parameters, perform multi-view forward intersection on the same image point in multiple images to obtain the coordinates of the ground point.

3. The method for precise localization of narrow field-of-view UAV images based on external attitude prior, as described in claim 2, is characterized in that: The feature matching in step S2 includes: first, using the SIFT algorithm for basic matching and using RANSAC to remove outliers; for weakly connected image pairs with fewer than a threshold number of matched inliers, a detectorless matching method LoFTR based on Transformer is further used for supplementary matching.

4. The method according to claim 2, characterized in that, The complete camera pose of the external SfM is obtained from the same image sequence through external 3D reconstruction or SfM processing, and includes at least the image number corresponding to the original image, camera center coordinates, rotation matrix or quaternion, camera interior orientation parameters, and the coordinate frame information of the SfM model to which it belongs.

5. The method according to claim 2, characterized in that, The three-dimensional seven-parameter similarity transformation includes a scale factor, a three-dimensional rotation, and a translation vector. The seven parameters are solved by weighted least squares from the center of the external SfM camera and the corresponding EXIF ​​GPS local coordinates. The three-dimensional rotation is used to simultaneously perform a frame transformation on the rotational attitude of the external SfM camera, so that the camera center and the rotational attitude together form a complete attitude soft prior.

6. The method for precise localization of narrow field-of-view UAV images based on external attitude prior, as described in claim 2, is characterized in that: The bundle adjustment described in step S5 also incorporates GPS track smoothing: interpolating the repeated coordinate segments in the EXIF ​​GPS coordinates and smoothing the track to reduce the negative impact of repeated coordinates and quantization noise on weak constraints.

7. The method for precise localization of narrow field-of-view UAV images based on external attitude prior, as described in claim 2, is characterized in that: The bundle adjustment described in step S5 also incorporates SO(3) relative rotation continuity constraints between adjacent images to suppress abrupt attitude changes.

8. The method for precise localization of narrow field-of-view UAV images based on external attitude prior, as described in claim 2, is characterized in that: The narrow field-of-view image refers to UAV images acquired by a telephoto lens with a field of view of less than 15°.

9. The method for precise localization of narrow field-of-view UAV images based on external attitude prior, as described in claim 2, is characterized in that: In step S6, during multi-view forward intersection, abnormal observations are identified based on the intersection residuals of each ray. Rays with large deviations are removed or downweighted, and a positioning reliability index is output. The reliability index includes the number of participating images, the average intersection residual, the maximum intersection angle, and the bundle adjustment coverage.

10. A narrow field-of-view UAV imagery-based precise positioning system based on external attitude priors, characterized in that: include: The image matching module is used to match features in narrow field-of-view image sequences and construct trajectories of corresponding points and image connection maps. The GPS preprocessing module is used to perform repeating segment interpolation and smoothing on EXIF ​​GPS camera station coordinates. The external attitude acquisition and reduction module is used to acquire the complete attitude of the external SfM camera, convert the EXIF ​​GPS to a local three-dimensional rectangular coordinate frame, and unify the center and rotation attitude of the external SfM camera to this local coordinate frame through a mass-weighted seven-parameter similarity transformation to obtain the complete attitude soft prior. The bundle adjustment module is used to jointly optimize camera parameters and object 3D points with reprojection error as the main constraint, EXIF ​​GPS camera station coordinates as the weak constraint, and complete attitude prior as the initial value or weighted prior. The multi-view forward intersection module is used to perform spatial ray intersection on the same image point in multiple images based on optimized camera parameters, and solve for the coordinates of the ground point. The credibility evaluation module is used to evaluate the credibility of the positioning results based on the number of participating images, intersection residuals, intersection angles, and bundle adjustment coverage.