Real-time seamless splicing method for remote sensing images of unmanned aerial vehicle

By combining the LightGlue deep learning model and intelligent pathfinding algorithm, subpixel-level precision feature matching and efficient registration of UAV remote sensing images are achieved, solving the problem of high computational complexity in large-scale image processing using traditional methods and meeting the needs of high-timeliness applications.

CN120931481AActive Publication Date: 2025-11-11SOUTH CHINA NORMAL UNIV

Patent Information

Application Number
CN202510927674.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-11
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Traditional UAV remote sensing image stitching methods have high computational complexity, making it difficult to meet the requirements of real-time or near-real-time stitching. In particular, they consume a lot of computational resources when processing large-scale, high-resolution images. Furthermore, traditional feature extraction algorithms are inefficient and cannot meet the needs of high-timeliness visualization analysis.

Method used

The LightGlue deep learning model is used for subpixel-level feature matching. Combined with the intelligent pathfinding stitching algorithm and the multi-stage registration framework of GNSS/IMU positioning data, the translation parameters are optimized by the least squares method to achieve efficient image registration and seamless stitching.

Benefits of technology

It enables real-time seamless stitching of large-scale images, improves processing efficiency, meets the needs of high-timeliness application scenarios such as disaster emergency response and wide-area dynamic monitoring, and provides an efficient and reliable technical solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931481A_ABST
    Figure CN120931481A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image data processing, and provides an unmanned aerial vehicle remote sensing image real-time seamless splicing method which comprises the following steps: preprocessing an unmanned aerial vehicle remote sensing image to obtain a geometric correction image; performing feature matching processing on the geometric correction image through a LightGlue deep learning model to obtain sub-pixel-level matching point pairs; performing image registration on the unmanned aerial vehicle remote sensing image according to the sub-pixel-level matching point pair to obtain a geometric registration image; performing splicing line search on the geometric registration image through an intelligent path-finding splicing line algorithm to obtain an optimal splicing path; and performing seamless splicing on the geometric registration image according to the optimal splicing path to obtain a seamless spliced image. According to the method, the overall splicing efficiency is improved, the unmanned aerial vehicle image can be processed in real time or near real time, and a more efficient splicing technology is provided for large-scale remote sensing application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for real-time seamless stitching of UAV remote sensing images. Background Technology

[0002] Rapid image stitching technology for UAV remote sensing is a key technology in computer vision and remote sensing applications, widely used in disaster emergency response, wide-area scene monitoring, agricultural census, and urban dynamic observation. Traditional image stitching methods are typically based on 3D reconstruction and orthorectification processes. First, 3D point clouds and camera poses of the scene are reconstructed using Structure of Motion (SfM) or Multi-View Stereo (MVS) techniques. Then, orthorectified images are generated based on the 3D model, and finally, stitching is performed. For feature matching, traditional methods mainly rely on classic feature extraction algorithms, such as Scale Invariant Feature Transform (SIFT) and Speed-Up Robust Feature Transform (SURF), which demonstrate good stability and robustness. In recent years, deep learning technology has made significant progress in computer vision, especially in feature extraction and matching tasks. Models based on convolutional neural networks (CNNs) and attention mechanisms (such as SuperPoint, SuperGlue, LoFTR, and LightGlue) have shown powerful performance, leveraging the advantages of GPU parallel computing.

[0003] However, existing technologies have significant shortcomings: while traditional 3D reconstruction and orthorectification processes can guarantee geometric accuracy, their computational complexity is extremely high, especially when processing large-scale, high-resolution images (such as 5000×5000 pixels and above). The 3D reconstruction and dense matching processes consume a large amount of computational resources, resulting in low processing efficiency and making it difficult to meet the needs of real-time or near-real-time stitching. On the other hand, traditional feature extraction algorithms have low computational efficiency on large-scale images. As the number of images or resolution increases, the time complexity of feature extraction and matching increases non-linearly, becoming a bottleneck in the entire stitching process and failing to meet the visualization analysis needs with high real-time requirements. Summary of the Invention

[0004] This invention provides a method for real-time seamless stitching of UAV remote sensing images to overcome the shortcomings of existing technologies.

[0005] This invention provides a method for real-time seamless stitching of UAV remote sensing images, including: S1: Preprocess the UAV remote sensing images to obtain geometrically corrected images; S2: The geometrically corrected image is processed by feature matching using the LightGlue deep learning model to obtain sub-pixel level matching point pairs; S3: Perform image registration on the UAV remote sensing image based on the sub-pixel level matching point pairs to obtain a geometrically registered image; S4: The optimal stitching path is obtained by searching the stitching lines of the geometric registration image through an intelligent path-finding stitching line algorithm; S5: Seamlessly stitch the geometrically registered images according to the optimal stitching path to obtain a seamless stitched image.

[0006] According to the present invention, a method for real-time seamless stitching of UAV remote sensing images, step S1 further includes: S11: Distortion removal processing is performed on the original UAV remote sensing images using camera calibration parameters to obtain distortion-removed images; S12: Perform illumination correction processing on the distortion-reduced image using an illumination equalization algorithm to obtain an illumination-corrected image; S13: Perform attitude correction processing on the illumination correction image based on the UAV pose data to obtain a geometric correction image.

[0007] According to the present invention, a method for real-time seamless stitching of UAV remote sensing images, step S13 further includes: S131: Construct a rotation correction matrix based on the UAV pose data, eliminate the tilt component of the illumination correction image, and obtain a rotation correction image; S132: Optimize translation parameters by combining the GNSS positioning information in the UAV pose data to geographically align the rotation-corrected image and obtain a geometrically corrected image.

[0008] According to the present invention, a method for real-time seamless stitching of UAV remote sensing images, step S2 further includes: S21: Construct a training sample library; S22: Pre-train the LightGlue model using the MegaDepth dataset to obtain a pre-trained model; S23: Fine-tune the pre-trained model according to the training sample library to obtain the fine-tuned LightGlue model; S24: The geometrically corrected image is subjected to feature detection and matching through the fine-tuned LightGlue model to obtain the sub-pixel level matching point pair.

[0009] According to the present invention, a method for real-time seamless stitching of UAV remote sensing images, step S21 further includes: S211: Perform sparse 3D reconstruction on the geometrically corrected image to obtain multi-view matching point pairs between multi-view images; S212: Perform geometric consistency verification on the multi-view matching point pairs to obtain reliable feature points; S213: Filter the reliable feature points by the reprojection error threshold to obtain the training sample library.

[0010] According to the present invention, a method for real-time seamless stitching of UAV remote sensing images, step S3 further includes: S31: Establish a reference coordinate system, and perform coarse registration on the geometrically corrected image through the reference coordinate system to obtain a coarsely registered image; S32: Solve the preset objective function using the least squares method to optimize the sub-pixel level matching points and obtain the optimal translation parameters; S33: Perform fine registration on the coarse registration image according to the optimal translation parameters to obtain the geometric registration image.

[0011] According to the real-time seamless stitching method for UAV remote sensing images provided by the present invention, the expression of the objective function in step S32 is: ; ; in, This refers to the group index value of the sub-pixel level matching points. The total number of groups of sub-pixel level matching points. The x-axis term represents the optimal translation parameters. The ordinate term represents the optimal translation parameters. For reference image number The x-coordinates of sub-pixel level matching points. For the target image, the first The x-coordinates of sub-pixel level matching points. For reference image number The ordinate of the sub-pixel level matching points. For the target image, the first The ordinate of the sub-pixel level matching points.

[0012] According to the present invention, a method for real-time seamless stitching of UAV remote sensing images, step S4 further includes: S41: Measure the disparity of sub-pixel level matching points by Euclidean distance calculation to obtain disparity data; S42: Filter the disparity data according to the dynamic disparity threshold to obtain a candidate point set; S43: Use the A* pathfinding algorithm to perform path search on the candidate point set to obtain the optimal splicing path.

[0013] According to the real-time seamless stitching method for UAV remote sensing images provided by the present invention, the expression for the cost function of path search for the candidate point set in step S43 is as follows: ; in, These are the coordinate points to be selected. Candidate coordinate points The total cost function, As the first weighting coefficient, This is the second weighting coefficient. The third weighting coefficient, Candidate coordinate points The path smoothness cost, Candidate coordinate points The cost of local disparity consistency Candidate coordinate points The cost of texture continuity.

[0014] According to the present invention, a method for real-time seamless stitching of UAV remote sensing images, step S5 further includes: S51: Smooth the optimal splicing path to obtain a smooth splicing line; S52: Segment the geometrically registered image according to the smooth stitching line to obtain the stitching area; S53: The stitching area is edge-blended using feathering technology to obtain a seamless stitched image.

[0015] This invention provides a real-time seamless stitching method for UAV remote sensing images. It employs the LightGlue deep learning model for feature matching, utilizing the model's adaptive inference mechanism and lightweight attention structure. This method achieves sub-pixel-level precision feature detection and matching in large-scale image processing, significantly improving processing efficiency compared to the traditional SIFT algorithm for large-pixel images. Furthermore, it intelligently allocates computational resources by dynamically evaluating matching confidence, effectively addressing the bottleneck of non-linearly increasing computational complexity in high-resolution image processing using traditional methods. Secondly, the introduction of an intelligent pathfinding stitching algorithm improves the A* pathfinding algorithm to construct a comprehensive cost function. This algorithm simultaneously considers three key factors: path smoothness, disparity consistency, and texture continuity. It uses Euclidean distance calculation for disparity measurement and dynamic disparity threshold filtering, effectively avoiding high-disparity areas such as building edges and moving objects, thus increasing the SSIM value of the stitching seam and significantly improving the problems of mismatches and obvious stitching artifacts in complex scenes, which are common with traditional methods. Finally, the method combines coarse registration based on GNSS / IMU positioning data with least squares optimization. The proposed multi-stage registration framework, through precise solution of optimal translation parameters, controls registration errors in complex urban scenes to the sub-pixel level, avoiding the computational burden of high-precision 3D reconstruction required by traditional methods, and achieving an efficient registration process from coarse to fine. It also employs cubic uniform B-spline interpolation to ensure the continuity of stitching lines, combined with edge blending using feathering technology, and achieves a natural transition effect through pixel-value weighted calculation using a Gaussian weighting function, effectively eliminating the obvious seam problems caused by traditional direct stitching methods. Furthermore, by constructing a high-precision sample library based on the SfM tool, combined with geometric consistency verification and reprojection error filtering, this invention provides reliable supervisory data for the fine-tuning training of the LightGlue model, enabling the model to perform excellently in scenarios with unique perspective changes and scale differences in UAV remote sensing imagery, improving matching accuracy compared to general pre-trained models in specific application scenarios. The geometric correction processing using affine transformation and rotation correction matrices effectively eliminates geometric inconsistencies caused by camera pose in oblique photography, providing a standardized data foundation for subsequent processing and avoiding the computational overhead of complex orthorectification required by traditional methods.

[0016] Overall, this invention achieves real-time processing capability for single-pair image processing time by organically combining the parallel computing advantages of deep learning algorithms with intelligent path optimization. This meets the urgent needs of high-timeliness application scenarios such as disaster emergency response and wide-area dynamic monitoring, and provides an efficient and reliable technical solution for rapid visualization and analysis of large-scale remote sensing images. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a method for real-time seamless stitching of UAV remote sensing images provided by the present invention; Figure 2 This is a schematic diagram of the sub-pixel level matching point pair acquisition method provided by the present invention; Figure 3 This is a schematic diagram of the geometric registration image acquisition method provided by the present invention; Figure 4 This is a schematic diagram of the optimal splicing path acquisition method provided by the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, embodiments of this invention, and should not be construed as limiting the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention. In the description of this invention, it should be understood that the terminology used is for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0020] The embodiments of the present invention are described below with reference to the figures.

[0021] like Figure 1 As shown, the present invention provides a method for real-time seamless stitching of UAV remote sensing images, comprising: S1: Preprocess the UAV remote sensing imagery to obtain geometrically corrected images.

[0022] Step S1 further includes: S11: Distortion removal processing is performed on the original UAV remote sensing images using camera calibration parameters to obtain distorted images.

[0023] Furthermore, the camera calibration parameters include a camera intrinsic parameter matrix and a distortion coefficient vector. The intrinsic parameter matrix K is a 3×3 matrix containing focal length parameters and optical center coordinates, and the distortion coefficient vector D contains radial distortion coefficients and tangential distortion coefficients. Specifically, the distortion correction process first reads the pixel coordinates of the original UAV remote sensing image and converts them into normalized image coordinates. Then, it calculates the distance from the normalized coordinates to the optical center, and applies the radial distortion correction formula to calculate the corrected normalized coordinates (x', y'). Subsequently, tangential distortion correction is applied. Finally, the corrected normalized coordinates are converted back to pixel coordinates. This invention uses a bilinear interpolation algorithm to calculate the pixel value corresponding to the new pixel coordinates, obtaining the corrected pixel value. In step S11 of this invention, the distortion correction process traverses each pixel of the original image, generating a distorted image through the aforementioned mathematical transformations and interpolation calculations. Each pixel in the distorted image has undergone distortion correction, eliminating the influence of lens distortion on the geometric accuracy of the image.

[0024] S12: Perform illumination correction processing on the distortion-free image using an illumination equalization algorithm to obtain an illumination-corrected image.

[0025] Furthermore, the illumination equalization algorithm employs an adaptive contrast-limited histogram equalization algorithm. This invention uses this algorithm to divide the distortion-reduced image into multiple non-overlapping rectangular regions. Specifically, firstly, a histogram is calculated for each region to statistically analyze the distribution of pixels at gray levels 0-255 within the region, forming a 256-dimensional histogram vector. Next, a contrast limit threshold is set, and the cumulative distribution function for each region is calculated. For histograms exceeding the contrast limit threshold, the excess pixels are evenly distributed to other histograms, resulting in a redistributed histogram. Then, a new cumulative distribution function is calculated based on the corrected histogram, establishing a gray-level mapping relationship. For each pixel, the pixel value is transformed according to the gray-level mapping relationship of its region. For pixels located at region boundaries, bilinear interpolation is used to calculate the final pixel value. The interpolation weight is determined based on the distance from the pixel to the center of each region. By performing the above illumination equalization processing on all pixels of the distortion-reduced image, an illumination-corrected image is obtained, exhibiting uniform illumination distribution and enhanced contrast.

[0026] S13: Perform attitude correction processing on the illumination correction image based on the UAV pose data to obtain a geometric correction image.

[0027] Step S13 further includes: S131: Construct a rotation correction matrix based on the UAV pose data, eliminate the tilt component of the illumination correction image, and obtain a rotation correction image.

[0028] Furthermore, in step S1131, the angle value is first converted from degrees to radians, and then rotation matrices around the X, Y, and Z axes are constructed respectively. Subsequently, a coincidence matrix is ​​calculated through matrix multiplication. Then, the rotation matrix is ​​expanded into homogeneous coordinate form for transformation, and finally, the rotated pixel coordinates are obtained through perspective division. Through the above rotation transformation and interpolation calculation, a rotation-corrected image is generated, eliminating the image tilt component caused by changes in UAV attitude.

[0029] S132: Optimize translation parameters by combining the GNSS positioning information in the UAV pose data to geographically align the rotation-corrected image and obtain a geometrically corrected image.

[0030] Furthermore, the GNSS positioning information includes the latitude and longitude coordinates and altitude of the UAV. In data processing, the present invention first converts the latitude and longitude coordinates into planar coordinates in the UTM coordinate system, then optimizes the translation parameters, obtains the optimal translation parameters, performs translation transformation on the rotation-corrected image, and finally calculates the translated pixel values ​​through a bilinear interpolation algorithm to generate a geometrically corrected image. The obtained image simultaneously eliminates attitude tilt and position deviation, achieving accurate geographic alignment.

[0031] S2: The geometrically corrected image is processed by feature matching using the LightGlue deep learning model to obtain sub-pixel level matching point pairs.

[0032] Step S2 further includes: S21: Construct a training sample library.

[0033] In step S21, this invention constructs a refined training sample library of UAV images, aiming to extract high-precision matching point pairs between multi-view images as supervised ground truth. During the reconstruction process, reliable feature points are first screened based on geometric consistency verification, and then outliers with large reprojection errors are filtered out. Finally, a sample library containing image pairs, matching point coordinates, and camera poses is constructed. At the same time, data augmentation methods such as illumination transformation and affine perturbation are introduced to improve sample diversity and ensure the domain adaptability and robustness of model training.

[0034] Step S21 further includes: S211: Perform sparse 3D reconstruction on the geometrically corrected image to obtain multi-view matching point pairs between multi-view images.

[0035] Furthermore, the sparse 3D reconstruction of this invention employs the Structure for Motion Restoration (SfM) algorithm. First, feature point detection is performed on the geometrically corrected image obtained in step S1. Corner and edge points are detected using the SIFT algorithm. The SIFT algorithm constructs a Gaussian pyramid containing multiple scale layers. Each layer of the image is generated by convolving the original image with a Gaussian kernel. The algorithm calculates the difference images of adjacent scale layers. Subsequently, candidate feature points are found through local extremum detection. The detection condition is that the response value of a pixel is greater than or less than all pixel values ​​in its neighborhood. Then, the feature points are precisely located. Subpixel-level coordinates are calculated using Taylor expansion. The precise position is obtained by finding the derivative to be zero. Then, the principal direction of the feature point is calculated. The gradient magnitude and direction are calculated in the neighborhood of the feature point, and a direction histogram is constructed. The histogram contains 36 bins, each covering a 10-degree angle range. After generating a 128-dimensional SIFT descriptor, the neighborhood of the feature point is divided into 4×4 sub-regions. Gradient histograms of 8 directions are calculated for each sub-region, forming a 128-dimensional feature vector. For N geometrically corrected images, M SIFT feature points are extracted from each image, forming an N×M feature point set. Next, feature point matching is performed, calculating the Euclidean distance between feature descriptors, applying a ratio test to filter mismatches, and then using the RANSAC algorithm to estimate camera parameters. Eight randomly selected matching point pairs are used to calculate the fundamental matrix, and the distances from the remaining matching points to the epipolar line are calculated; points with distances less than a threshold are considered inliers. This process is repeated, and the fundamental matrix with the most inliers is selected as the optimal solution. Finally, the Levenberg-Marquardt algorithm is used to optimize the camera parameters and 3D point coordinates, resulting in multi-view matching point pairs. Each matching point pair contains 2D coordinates and corresponding 3D coordinates from different viewpoints.

[0036] S212: Perform geometric consistency verification on the multi-view matching point pairs to obtain reliable feature points.

[0037] Furthermore, the geometric consistency verification is based on the epipolar geometric constraint principle. First, the epipolar constraint error of the fundamental matrix is ​​calculated. After calculating the epipolar constraint error, an epipolar constraint threshold pixel is set. Matching point pairs whose epipolar constraint error is less than the epipolar constraint threshold pixel pass the epipolar constraint test. Combining the above geometric consistency verification conditions, matching point pairs that meet the above conditions are marked as reliable feature points.

[0038] S213: Filter the reliable feature points by the reprojection error threshold to obtain the training sample library.

[0039] Furthermore, the reprojection error is verified based on triangulation. First, triangulation is used to calculate the coordinates of three-dimensional points. A homogeneous linear equation system is constructed using linear triangulation. Then, the reprojection error is calculated, and a threshold pixel for the reprojection error is set. Matching point pairs whose reprojection errors from all cameras are less than the threshold pixel pass the triangulation verification. Finally, the data structure in the constructed training sample library contains information such as image pair ID, feature point coordinates, feature descriptors, camera parameters, and matching labels for each sample.

[0040] S22: Pre-train the LightGlue model using the MegaDepth dataset to obtain a pre-trained model.

[0041] Furthermore, in step S22, the present invention first loads the MegaDepth dataset, which contains one million images of 196 scenes, each image labeled with depth information and camera pose. The LightGlue model adopts an encoder-decoder architecture. The encoder includes a feature extraction network and a position encoding module, and the decoder includes an attention mechanism and a matching head. The feature extraction network is based on the ResNet-50 backbone network and outputs a 256-dimensional feature map. The attention mechanism adopts multi-head self-attention and cross-attention. The matching head adopts bidirectional nearest neighbor matching. After calculating the similarity matrix, the matching probability is calculated using the softmax function. The total training batch size is 16, the total number of training rounds is 100, and each training round contains 1000 batches. Finally, the pre-trained model converges on the MegaDepth dataset.

[0042] S23: Fine-tune the pre-trained model according to the training sample library to obtain the fine-tuned LightGlue model.

[0043] Furthermore, the fine-tuning process of this invention employs a transfer learning strategy. First, the weight parameters of the pre-trained model are loaded, the first three convolutional layers of the feature extraction network are frozen, and only the parameters of subsequent layers and the matching head are fine-tuned. During training, the training sample library is divided into training, validation, and test sets in an 8:1:1 ratio. The training set is used for model parameter updates, the validation set for hyperparameter tuning and early stopping strategies, and the test set for final performance evaluation. The data loader uses a random sampling strategy, with each batch containing 8 pairs of images, and 500 feature points are randomly selected from each pair for matching. During training, data augmentation includes geometric augmentation and photometric augmentation. Geometric augmentation includes random rotation, scaling, translation, and perspective transformation, while photometric augmentation includes brightness adjustment, contrast adjustment, saturation adjustment, and hue adjustment. The training process employs an early stopping strategy, stopping training when the validation set loss does not decrease for 10 consecutive rounds. The model saving strategy uses the best model saving, selecting the optimal model weights based on the matching accuracy on the validation set. The fine-tuning training consisted of 100 rounds, with each round taking approximately 30 minutes. The fine-tuned model achieved a matching accuracy of over 92% on the UAV imagery test set.

[0044] S24: The geometrically corrected image is subjected to feature detection and matching through the fine-tuned LightGlue model to obtain the sub-pixel level matching point pair.

[0045] Specifically, in step S24, the geometrically corrected image is first input into the fine-tuned LightGlue model. After the image input, a multi-scale feature map is generated through the ResNet-50 backbone network. Subsequently, the multi-scale features are fused from top to bottom through a Feature Pyramid Network (FPN) structure. Then, the position encoding module generates a position encoding vector for each feature point, using sine and cosine encoding. The attention calculation stage uses a multi-head attention mechanism and self-attention for calculation, and then establishes the association between the features of the two images through cross-attention. The detection matching head adopts a bidirectional soft matching strategy. After calculating the similarity matrix, the optimal transfer matrix is ​​solved using the Sinkhorn algorithm. Subsequently, the gradient descent method is used for sub-pixel refinement. After multiple iterations, sub-pixel level matching point pairs are obtained. Each matching point pair contains the precise coordinates and matching confidence scores in the two images.

[0046] S3: Perform image registration on the UAV remote sensing image based on the sub-pixel level matching point pairs to obtain a geometrically registered image.

[0047] Step S3 further includes: S31: Establish a reference coordinate system, and perform coarse registration on the geometrically corrected image through the reference coordinate system to obtain a coarsely registered image.

[0048] Furthermore, in step S31, the present invention first reads the EXIF ​​metadata of each image to extract GNSS coordinates, then reads IMU data to obtain attitude angles, corresponding to pitch, roll, and yaw angles respectively. The reference coordinate system adopts the UTM projected coordinate system, converting geographic coordinates into planar coordinates. After establishing the reference coordinate system, the position and orientation of each image in the reference coordinate system are calculated, and then translation, rotation, and scaling are performed using an affine transformation matrix for coarse registration. Specifically, firstly, an initial translation vector is calculated based on the GNSS coordinate difference, and a rotation matrix is ​​calculated based on the IMU attitude angles. Then, the target image is transformed to the reference image coordinate system. The transformed image coordinates are the product of the rotation matrix and the original coordinates, plus the sum of the translation vectors. The coarsely registered images achieve preliminary spatial alignment.

[0049] S32: Solve the preset objective function using the least squares method to optimize the sub-pixel level matching points and obtain the optimal translation parameters.

[0050] In this invention, the objective function for image registration is defined as the sum of squared Euclidean distances between all matching point pairs. During data processing, the system first iterates through all matching point pairs, calculates the coordinate difference for each pair, and then averages all differences to obtain the optimal translation parameters. It is important to note that this invention employs double-precision floating-point arithmetic in the calculation process to ensure sub-pixel accuracy. The advantage of the least squares method lies in the existence and uniqueness of its analytical solution, avoiding the computational complexity of iterative optimization. Ultimately, the translation parameters obtained through the least squares method minimize the geometric errors of all matching point pairs, achieving sub-pixel registration accuracy.

[0051] The expression for the objective function in step S32 is: ; ; in, This refers to the group index value of the sub-pixel level matching points. The total number of groups of sub-pixel level matching points. The x-axis term represents the optimal translation parameters. The ordinate term represents the optimal translation parameters. For reference image number The x-coordinates of sub-pixel level matching points. For the target image, the first The x-coordinates of sub-pixel level matching points. For reference image number The ordinate of the sub-pixel level matching points. For the target image, the first The ordinate of the sub-pixel level matching points.

[0052] S33: Perform fine registration on the coarse registration image according to the optimal translation parameters to obtain the geometric registration image.

[0053] Furthermore, in step S33, based on the optimal translation parameters obtained in step S32, the present invention performs a geometric transformation on the coarsely registered image to perform fine registration. Specifically, firstly, a coordinate transformation is performed, applying a translation transformation to the coordinates of each pixel in the target image. Since the translation parameters are usually non-integer values, the transformed coordinates may not correspond to integer pixel positions, therefore pixel interpolation is required. The pixel interpolation uses a bilinear interpolation algorithm, which uses a weighted average of the gray values ​​of the four nearest neighbor pixels around the transformed coordinates. Through bilinear interpolation, the system generates a geometrically registered image, and the obtained image achieves sub-pixel alignment with the reference image in spatial position.

[0054] S4: The stitching line search is performed on the geometric registration image using an intelligent pathfinding stitching line algorithm to obtain the optimal stitching path.

[0055] Step S4 further includes: S41: Disparity data is obtained by calculating the disparity of sub-pixel level matching points using Euclidean distance.

[0056] In step S411, this invention quantifies the spatial distance between two points on a two-dimensional plane using Euclidean distance calculation. In the disparity metric of sub-pixel-level matching points, the coordinate difference between the reference image and the target image is calculated for each pair of matching points. Specifically, a matching point pair data structure is first established, then all matching point pairs are traversed one by one, and the coordinate difference is calculated for each pair to obtain the Euclidean distance of each pair of sub-pixel-level matching points, i.e., disparity data. The disparity data reflects the degree of spatial offset of the matching point pair; a smaller disparity value indicates better spatial consistency of the matching point pair, while a larger disparity value indicates the presence of geometric deformation or matching error.

[0057] S42: Filter the disparity data according to the dynamic disparity threshold to obtain a candidate point set.

[0058] Furthermore, the dynamic disparity threshold of this invention dynamically adjusts the disparity screening criteria based on scene complexity and image features. The calculation of the dynamic disparity threshold is based on the statistical characteristics of disparity data. The formula for calculating the dynamic disparity threshold is: ; in, To calculate the obtained dynamic disparity threshold, The mean of the disparity values. The standard deviation of the disparity values. It's an adjustment coefficient, set according to the scene type. For urban building cluster scenes... The value is set to 1.5 for farmland and plains scenarios. Set the value to 2.0.

[0059] During the filtering process, all matching point pairs need to be traversed, and the disparity value of each pair is compared with a dynamic disparity threshold. Matching point pairs whose disparity value is less than the dynamic disparity threshold are retained and added to the candidate point set. The data structure of the candidate point set contains the coordinate information and disparity value information of each candidate point. The filtered candidate point set removes outliers in high disparity regions and retains matching point pairs with good geometric consistency, providing a reliable set of nodes for subsequent stitching line search.

[0060] S43: Use the A* pathfinding algorithm to perform path search on the candidate point set to obtain the optimal splicing path.

[0061] In the stitching line search, this invention uses the A* pathfinding algorithm to search for the optimal path from the left edge to the right edge of the image within a candidate point set. Algorithm initialization includes creating an open list and a closed list. The open list stores nodes to be evaluated, and the closed list stores evaluated nodes. During algorithm execution, when a target node is added to the closed list, the optimal path is reconstructed by backtracking the parent node pointer. The final optimal stitching path is output as a sequence of node coordinates, including the coordinates of all intermediate nodes from the starting point to the ending point.

[0062] The expression for the cost function of path search on the candidate point set in step S43 is as follows: ; in, These are the coordinate points to be selected. Candidate coordinate points The total cost function, As the first weighting coefficient, This is the second weighting coefficient. The third weighting coefficient, Candidate coordinate points The path smoothness cost, Candidate coordinate points The cost of local disparity consistency Candidate coordinate points The cost of texture continuity.

[0063] Furthermore, the path smoothness is the change in the turning angle of adjacent points, the local disparity consistency represents the disparity variance within the sliding window, the texture continuity is the gradient consistency along the path direction, and multiple weight coefficients are selected empirically.

[0064] S5: Seamlessly stitch the geometrically registered images according to the optimal stitching path to obtain a seamless stitched image.

[0065] Step S5 further includes: S51: Smooth the optimal splicing path to obtain a smooth splicing line.

[0066] Furthermore, this invention employs a B-spline curve representation method, generating a smooth curve with C² continuity through control points and basis functions. In the smoothing process, a cubic uniform B-spline interpolation algorithm is used to fit the optimal splicing path obtained in step S43. Specifically, first, the node vector array is calculated; then, the corresponding basis function value is calculated for each parameter value; next, the basis function values ​​are weighted and summed with the control point coordinates to obtain the coordinates of the corresponding points on the curve; then, the parameters are discretized according to a preset sampling density, with each sampling point corresponding to a coordinate point on the curve, forming a coordinate sequence of the smooth splicing line. The smooth splicing line has continuous first and second derivatives, eliminating sharp turns and discontinuities in the original path.

[0067] S52: Segment the geometric registration image according to the smooth stitching line to obtain the stitching area.

[0068] Further, in step S52, the present invention divides the geometrically registered image into different stitching regions based on the smooth stitching line. The smooth stitching line serves as the segmentation boundary, dividing the overlapping region into two sub-regions: the retained region and the discarded region. The segmentation algorithm employs a scan-line filling method, scanning image pixels line by line and determining the positional relationship of each pixel relative to the stitching line. Specifically, firstly, a mathematical representation of the stitching line is established, converting the coordinate sequence of the smooth stitching line into a piecewise linear function. Then, for each scan line, the coordinates of its intersection with the stitching line are calculated. After traversing each scan line, pixels with a horizontal coordinate less than the intersection coordinate are marked as reference image regions, and pixels with a horizontal coordinate greater than the intersection coordinate are marked as target image regions. The final segmentation result generates a binary mask image, where pixels with a mask value of 1 belong to the reference image region, and pixels with a mask value of 0 belong to the target image region. The data structure of the stitching region includes the coordinate information, grayscale information, and region identification information of each pixel. The segmentation operation precisely divides the original overlapping region according to the stitching line, providing clear region boundaries for subsequent feathering processing.

[0069] S53: The stitching area is edge-blended using feathering technology to obtain a seamless stitched image.

[0070] Furthermore, the feathering process employs a Gaussian weighted fusion algorithm, establishing symmetrical feather bands on both sides of the stitching line. The fusion weight of each pixel within the feather band is calculated using a Gaussian function, and the distance calculation uses the point-to-line distance formula. For a line segment formed by two adjacent points on the stitching line, the shortest distance from each pixel in the overlapping area to that line segment is calculated. Then, a weighted fusion operation is performed on each pixel within the feather band. The fusion process is performed pixel by pixel, with weight calculation and pixel value fusion performed on the three color channels (RGB) respectively. Finally, the feathering process generates a seamless stitched image, eliminating grayscale abrupt changes and color differences at the stitching line.

[0071] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0072] This invention provides a real-time seamless stitching method for UAV remote sensing images. It proposes a rapid UAV image stitching technology for real-time visualization needs. Through deep learning feature matching and intelligent path optimization algorithms, it achieves efficient synthesis from local images to wide-area scenes. Compared to traditional surveying-grade stitching schemes relying on 3D reconstruction, this invention employs an innovative 2D direct stitching strategy, ensuring visual consistency and achieving a speed improvement of over 80% in processing large-scale images (5000×5000 pixels and above). The core of this invention lies in achieving sub-pixel-level high-precision feature matching based on deep learning, combined with intelligent pathfinding algorithms to effectively avoid large parallax areas such as buildings, ensuring seamless stitching. This invention is particularly suitable for high-timeliness applications such as disaster emergency response and wide-area dynamic monitoring, providing reliable technical support for rapidly acquiring panoramic visualization data. It has significant application value and promising prospects in fields such as smart city construction and agricultural resource surveys.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for real-time seamless stitching of UAV remote sensing images, characterized in that, include: S1: Preprocess the UAV remote sensing images to obtain geometrically corrected images; S2: The geometrically corrected image is processed by feature matching using the LightGlue deep learning model to obtain sub-pixel level matching point pairs; S3: Perform image registration on the UAV remote sensing image based on the sub-pixel level matching point pairs to obtain a geometrically registered image; S4: The optimal stitching path is obtained by searching the stitching lines of the geometric registration image through an intelligent path-finding stitching line algorithm; S5: Seamlessly stitch the geometrically registered images according to the optimal stitching path to obtain a seamless stitched image.

2. The method for real-time seamless stitching of UAV remote sensing images according to claim 1, characterized in that, Step S1 further includes: S11: Distortion removal processing is performed on the original UAV remote sensing images using camera calibration parameters to obtain distortion-removed images; S12: Perform illumination correction processing on the distortion-reduced image using an illumination equalization algorithm to obtain an illumination-corrected image; S13: Perform attitude correction processing on the illumination correction image based on the UAV pose data to obtain a geometric correction image.

3. The method for real-time seamless stitching of UAV remote sensing images according to claim 2, characterized in that, Step S13 further includes: S131: Construct a rotation correction matrix based on the UAV pose data, eliminate the tilt component of the illumination correction image, and obtain a rotation correction image; S132: Optimize translation parameters by combining the GNSS positioning information in the UAV pose data to geographically align the rotation-corrected image and obtain a geometrically corrected image.

4. The method for real-time seamless stitching of UAV remote sensing images according to claim 1, characterized in that, Step S2 further includes: S21: Construct a training sample library; S22: Pre-train the LightGlue model using the MegaDepth dataset to obtain a pre-trained model; S23: Fine-tune the pre-trained model according to the training sample library to obtain the fine-tuned LightGlue model; S24: The geometrically corrected image is subjected to feature detection and matching through the fine-tuned LightGlue model to obtain the sub-pixel level matching point pair.

5. A method for real-time seamless stitching of UAV remote sensing images according to claim 4, characterized in that, Step S21 further includes: S211: Perform sparse 3D reconstruction on the geometrically corrected image to obtain multi-view matching point pairs between multi-view images; S212: Perform geometric consistency verification on the multi-view matching point pairs to obtain reliable feature points; S213: Filter the reliable feature points by the reprojection error threshold to obtain the training sample library.

6. The method for real-time seamless stitching of UAV remote sensing images according to claim 1, characterized in that, Step S3 further includes: S31: Establish a reference coordinate system, and perform coarse registration on the geometrically corrected image through the reference coordinate system to obtain a coarsely registered image; S32: Solve the preset objective function using the least squares method to optimize the sub-pixel level matching points and obtain the optimal translation parameters; S33: Perform fine registration on the coarse registration image according to the optimal translation parameters to obtain the geometric registration image.

7. A method for real-time seamless stitching of UAV remote sensing images according to claim 6, characterized in that, The expression for the objective function in step S32 is: ; ; in, This refers to the group index value of the sub-pixel level matching points. The total number of groups of sub-pixel level matching points. The x-axis term represents the optimal translation parameters. The ordinate term represents the optimal translation parameters. For reference image number The x-coordinate of the sub-pixel level matching points. For the target image, the first The x-coordinate of the sub-pixel level matching points. For reference image number The ordinate of the sub-pixel level matching points. For the target image, the first The ordinate of the sub-pixel level matching points.

8. A method for real-time seamless stitching of UAV remote sensing images according to claim 1, characterized in that, Step S4 further includes: S41: Measure the disparity of sub-pixel level matching points by Euclidean distance calculation to obtain disparity data; S42: Filter the disparity data according to the dynamic disparity threshold to obtain a candidate point set; S43: Use the A* pathfinding algorithm to perform path search on the candidate point set to obtain the optimal splicing path.

9. A method for real-time seamless stitching of UAV remote sensing images according to claim 8, characterized in that, The expression for the cost function of path search on the candidate point set in step S43 is: ; in, These are the coordinate points to be selected. Candidate coordinate points The total cost function, As the first weighting coefficient, This is the second weighting coefficient. The third weighting coefficient, Candidate coordinate points The path smoothness cost, Candidate coordinate points The cost of local disparity consistency Candidate coordinate points The cost of texture continuity.

10. A method for real-time seamless stitching of UAV remote sensing images according to claim 1, characterized in that, Step S5 further includes: S51: Smooth the optimal splicing path to obtain a smooth splicing line; S52: Segment the geometrically registered image according to the smooth stitching line to obtain the stitching area; S53: The stitching area is edge-blended using feathering technology to obtain a seamless stitched image.

Citation Information

Patent Citations

  • Large-area complex-terrain-region unmanned plane sequence image rapid seamless splicing method

    CN104156968A

  • Unmanned aerial vehicle remote image mosaicing method based on projection-similarity transformation

    CN106447601A

  • Aerial photography map generation system and method based on quadrotor

    CN106485655A

  • Low-altitude remote sensing image splicing method and device

    CN112750075A

  • Multi-view machine vision image splicing method and system and storage medium

    CN113793266A

Cited By

  • Lightweight remote sensing image geometric fine correction method and system based on B / S architecture

    CN121660943A