A depth-estimation-free pure-rotation optimization unmanned aerial vehicle panoramic image stitching method

CN122089565BActive Publication Date: 2026-08-07江苏省地质测绘大队
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
江苏省地质测绘大队
Filing Date
2026-04-24
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本发明的目的是提供一种免深度估计的纯旋转优化无人机全景影像拼接方法,能够解决现有技术中依赖深度信息、影像对齐精度低、视觉连续性差等问题,同时提高无人机全景影像拼接的效率和鲁棒性

Benefits of technology

本发明,在无需对场景进行深度测量或三维重建,仅利用影像间的纯旋转信息,即可实现高精度全景拼接,显著降低计算复杂度和对采集环境的依赖,通过旋转矩阵和方向统一处理,保证多张影像在方向层面的一致性,实现重叠区域的视觉连续,避免因平移误差或三维重建误差导致的拼接失真,充分利用无人机在同一空间位置悬停时的旋转采集特性,提高影像采集效率,保证全景覆盖完整,结合矫正拍摄方向和球面坐标映射,实现影像像素级融合,使最终全景影像视觉效果平滑连续,适用于无人机航拍、自然资源管理、耕地保护、虚拟现实、地理信息获取等多种应用场景,无需依赖环境特征的深度信息,对不同场景均适用,并能有效应对重叠区域特征复杂或稀疏的情况,提高全景影像拼接的稳定性和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089565B_ABST
    Figure CN122089565B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of panoramic image stitching, and particularly relates to a depth-estimation-free pure-rotation-optimized unmanned aerial vehicle panoramic image stitching method. The present application does not need to measure the depth of a scene or reconstruct the scene in three dimensions, but only uses the pure-rotation information between images to realize high-precision panoramic stitching, significantly reduces the calculation complexity and dependence on the collection environment, ensures the consistency of multiple images in the direction layer through rotation matrix and direction unified processing, realizes visual continuity in the overlapping area, avoids stitching distortion caused by translation error or three-dimensional reconstruction error, fully utilizes the rotation collection characteristics of the unmanned aerial vehicle when hovering at the same spatial position, improves the image collection efficiency, guarantees the completeness of panoramic coverage, combines the rectification of the shooting direction and the mapping of the spherical coordinates, realizes the image pixel-level fusion, makes the visual effect of the final panoramic image smooth and continuous, and is suitable for various application scenarios such as unmanned aerial vehicle aerial photography, natural resource management, farmland protection, virtual reality, geographic information acquisition and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of panoramic image stitching technology, specifically relating to a pure rotation optimization method for UAV panoramic image stitching without depth estimation. Background Technology

[0002] With the development of drone aerial photography and intelligent image processing technologies, panoramic imagery is increasingly being used in fields such as natural resource surveys, urban planning, environmental monitoring, digital exhibition halls, and virtual reality. Drones can flexibly hover and move in the air, acquiring large-scale image information through multi-angle and multi-directional shooting, providing a rich data source for panoramic image stitching. However, several technical problems still need to be solved in the actual panoramic image generation process.

[0003] Traditional panoramic image stitching methods typically rely on scene depth information or 3D spatial point coordinates, achieving image alignment through 3D reconstruction or stereo matching. These methods are prone to matching errors in scenes with sparse textures, repetitive patterns, or complex scenes, leading to ghosting, misalignment, or breaks in the panoramic image. Furthermore, depth estimation and 3D reconstruction are computationally intensive and inefficient, failing to meet the demands of rapid panoramic image generation from UAVs. When a UAV hovers in a fixed spatial location, the shooting directions of multiple images differ. Improper handling of rotational changes can easily cause directional misalignment during image stitching, especially in overlapping fields of view, potentially leading to visual discontinuities or localized blurring. Summary of the Invention

[0004] The purpose of this invention is to provide a pure rotation-optimized UAV panoramic image stitching method that does not require depth estimation. This method can solve the problems of reliance on depth information, low image alignment accuracy, and poor visual continuity in the prior art, while improving the efficiency and robustness of UAV panoramic image stitching.

[0005] The specific technical solution adopted by this invention is as follows: A depth-estimation-free, purely rotation-optimized UAV panoramic image stitching method includes: Acquire multiple raw images by rotating the drone around the optical center of the camera while it is hovering in the same spatial position, and obtain the corresponding camera intrinsic parameters; Feature extraction and feature matching are performed on multiple original images. For original image pairs with overlapping fields of view, corresponding image points located in the overlapping field of view are obtained. The rotation matrix of the difference in the shooting direction of the images is obtained based on the directional changes of the corresponding image points in their respective image imaging planes. Based on the rotation matrix corresponding to each original image pair, the shooting direction changes of multiple original images are correlated and unified as a whole to obtain the unified shooting direction result of multiple original images at the direction level. Based on the unified shooting direction result, obtain the initial shooting direction corresponding to each original image in the unified global direction coordinate system; Based on the camera's internal parameters, the shooting direction of each original image is corrected to obtain the corrected shooting direction; Based on the corrected shooting direction, each original image is mapped to a unified spherical coordinate system and the images are fused to generate a drone spherical panoramic image.

[0006] In a preferred embodiment, multiple raw images are acquired by the UAV while hovering at the same spatial position and rotating angularly around the optical center of the camera, and the corresponding camera intrinsic parameters are obtained, including: Obtain a preset spatial location and control the drone to hover within that location; The drone camera rotates continuously around the camera's optical center at a predetermined angle to acquire multiple raw images covering the panoramic view. During each image acquisition, camera intrinsic parameters are automatically recorded, including focal length, principal point coordinates, and distortion coefficients.

[0007] In a preferred embodiment, feature extraction and feature matching are performed on multiple original images. For original image pairs with partially overlapping fields of view, corresponding image points located within the overlapping field of view are obtained. A rotation matrix representing the difference in image shooting direction is obtained based on the directional changes of these corresponding image points in their respective image imaging planes, including: Feature extraction is performed on multiple original images, and feature matching is performed between different original images; For original image pairs whose fields of view partially overlap, select the corresponding image points located within the common field of view from the matching results; Based on the difference in shooting direction corresponding to the same image points in their respective image imaging planes, the rotation matrix of the original image with respect to the difference in shooting direction is obtained.

[0008] In a preferred embodiment, feature extraction is performed on multiple original images separately, and feature matching is performed between different original images, including: Based on each original image and its corresponding camera intrinsic parameters, local feature points are extracted from each original image, and corresponding feature description information is generated for each local feature point. Obtain directional reference information of the drone camera's orientation at the moment of shooting, and obtain relative shooting direction information between different original images based on the directional reference information; Based on the relative shooting direction information, original images with overlapping fields of view are selected as image pairs to be matched. Based on the feature description information of each image pair to be matched, the similarity between feature points is obtained, and an initial feature matching set is constructed. Based on the initial feature matching set, and combined with the relative shooting direction information of the image pairs to be matched, the initial feature matches that do not conform to the preset direction constraints are eliminated, and the feature matching results that meet the preset direction constraints are retained.

[0009] In a preferred embodiment, based on the difference in shooting direction corresponding to the same image points in their respective image imaging planes, a rotation matrix of the original image with respect to the difference in shooting direction is obtained, including: Based on the pixel position of each corresponding image point in the corresponding original image imaging plane, and combined with the camera intrinsic parameters of the corresponding original image, the position of each corresponding image point in the imaging plane is converted into the shooting direction with the camera optical center as the starting point. For the same set of image points with the same name, obtain the corresponding shooting direction in the original image pair, and obtain the difference in shooting direction of the same image points in different original images based on the change in direction between the two. Based on the differences in shooting directions corresponding to multiple sets of image points with the same name, rotation parameters are obtained to make the overall differences in shooting directions consistent. Based on the rotation parameters, a rotation matrix is ​​generated to represent the difference in shooting direction between the original image pairs.

[0010] In a preferred embodiment, based on the rotation matrix corresponding to each original image pair, the changes in the shooting orientation of multiple original images are comprehensively correlated and unified to obtain a unified shooting orientation result for multiple original images at the orientation level, including: For each original image, a consistency check is performed on the corresponding rotation matrix. Rotation matrices that do not match the shooting direction changes of most original images are removed, and only rotation matrices with stable shooting direction changes are retained as valid rotation matrices. Among the multiple original images involved in the effective rotation matrix, one original image is selected as the orientation reference image, and the shooting direction of the orientation reference image is used as the reference direction. Based on the effective rotation matrix, according to the shooting order between the original images, the shooting direction change is gradually transferred from the direction reference image to the remaining original images to obtain the shooting direction parameters of each original image relative to the reference direction. The shooting direction of each original image obtained through direction transfer is uniformly processed so that the shooting directions of multiple original images form corresponding shooting direction parameters under the same direction reference, thereby obtaining a unified shooting direction result of multiple original images at the direction level.

[0011] In a preferred embodiment, based on the unified shooting direction result, the initial shooting direction corresponding to each original image in a unified global orientation coordinate system is obtained, including: Based on the unified shooting direction result, a global orientation coordinate system for common orientation reference of multiple original images is obtained; Obtain the camera coordinate system corresponding to each original image during image acquisition, and obtain the orientation transformation result of the camera coordinate system relative to the global orientation coordinate system based on the unified shooting direction result; Based on the orientation transformation result, the shooting orientation of the original image in the camera coordinate system is transformed to the global orientation coordinate system, and the orientation representation result of the original image under a unified orientation reference is obtained. The orientation representation results are parameterized to extract the orientation parameters of the original image relative to the global orientation coordinate system. The orientation parameter is used as the initial shooting orientation of the corresponding original image in a unified global orientation coordinate system.

[0012] In a preferred embodiment, the initial shooting direction of each original image is corrected based on camera intrinsic parameters to obtain a corrected shooting direction, including: Based on the camera intrinsic parameters, the pixel coordinates of each corresponding image point in each original image are converted into a direction vector originating from the camera optical center. The directional deviation is obtained based on the directional vectors of the corresponding image points between the current image and adjacent images. Obtain a preset deviation threshold and determine whether the directional deviation is lower than the preset deviation threshold; If the directional deviation is lower than the preset deviation threshold, the initial shooting direction of the original image will be used as the corrected shooting direction. If the directional deviation is not lower than the preset deviation threshold, the initial shooting direction of the current image is updated by combining the camera intrinsic parameters to obtain the updated shooting direction. The directional deviation is re-judged based on the updated shooting direction until the directional deviation of each original image is lower than the preset threshold, and the corrected shooting direction of each original image is obtained.

[0013] In a preferred embodiment, each original image is mapped to a unified spherical coordinate system based on the corrected shooting direction, and the images are fused to generate a UAV spherical panoramic image, including: Based on the corrected shooting direction of each original image, the coordinates of each pixel in the camera coordinate system are converted into a direction vector in the spherical coordinate system. Align the orientation vectors of multiple original images in a unified spherical coordinate system to obtain the orientation vector alignment result. Based on the orientation vector alignment results, the pixel information of each original image is fused to generate a drone spherical panoramic image.

[0014] And, a depth estimation-free, purely rotation-optimized UAV panoramic image stitching terminal, comprising: One or more processors; A storage device on which one or more programs are stored; When one or more programs are executed by one or more processors, the one or more processors implement a pure rotation-optimized UAV panoramic image stitching method that eliminates depth estimation.

[0015] The technical effects achieved by this invention are as follows: This invention achieves high-precision panoramic stitching without requiring depth measurement or 3D reconstruction of the scene, utilizing only rotational information between images. This significantly reduces computational complexity and dependence on the acquisition environment. Through rotation matrix and orientation unification processing, it ensures consistency in orientation across multiple images, achieving visual continuity in overlapping areas and avoiding stitching distortion caused by translation or 3D reconstruction errors. It fully leverages the rotational acquisition characteristics of drones hovering in the same spatial location, improving image acquisition efficiency and ensuring complete panoramic coverage. Combined with corrected shooting direction and spherical coordinate mapping, it achieves pixel-level image fusion, resulting in a smooth and continuous visual effect for the final panoramic image. It is suitable for various applications such as drone aerial photography, natural resource management, farmland protection, virtual reality, and geographic information acquisition. It does not rely on depth information from environmental features, is applicable to different scenarios, and effectively handles situations with complex or sparse features in overlapping areas, improving the stability and reliability of panoramic image stitching. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method provided by the present invention. Detailed Implementation

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0018] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0019] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments.

[0020] Furthermore, the present invention will be described in detail with reference to the schematic diagrams. When describing the embodiments of the present invention in detail, the schematic diagrams are merely examples for ease of explanation and should not limit the scope of protection of the present invention.

[0021] Please see the appendix Figure 1As shown, a method for stitching panoramic UAV images with pure rotation optimization and no depth estimation is provided, including: S1. Acquire multiple raw images taken by the UAV while hovering in the same spatial position and rotating around the optical center of the camera, and obtain the corresponding camera intrinsic parameters; S2. Perform feature extraction and feature matching on multiple original images. For original image pairs with overlapping fields of view, obtain the corresponding image points located in the overlapping field of view area. Obtain the rotation matrix of the difference in image shooting direction based on the direction change of the corresponding image points in their respective image imaging planes. S3. Based on the rotation matrix corresponding to each original image pair, perform overall correlation and unified processing on the shooting direction changes of multiple original images to obtain the unified shooting direction result of multiple original images at the direction level. S4. Based on the unified shooting direction result, obtain the initial shooting direction corresponding to each original image in the unified global direction coordinate system; S5. Based on the camera's intrinsic parameters, the shooting direction of each original image is corrected to obtain the corrected shooting direction; S6. Based on the corrected shooting direction, map each original image to a unified spherical coordinate system and perform image fusion to generate a drone spherical panoramic image.

[0022] As described in steps S1 to S6 above, the drone hovers at a preset spatial position. By controlling the camera to rotate around the optical center at a predetermined angle, multiple original images covering the panoramic field of view are continuously acquired. Simultaneously, camera intrinsic parameters, including focal length, principal point coordinates, and distortion coefficients, are automatically recorded. Feature extraction and feature matching are performed on the acquired original images to identify corresponding image points located within overlapping field-of-view areas. By analyzing the directional changes of these corresponding image points in their respective image imaging planes, a rotation matrix reflecting the differences in image shooting directions is obtained. This rotation matrix only reflects the directional differences between images, without involving translation components or requiring scene depth information. Based on the rotation matrix of each image pair, multiple original images are... The shooting directions are correlated and unified as a whole. During this process, through direction transfer and consistency judgment, a unified shooting direction reference is formed for each image at the direction level, achieving directional continuity and rotational consistency between images. Based on the direction unification result, the shooting direction of each original image in the camera coordinate system is transformed to a unified global direction coordinate system to obtain the initial shooting direction of each image. Combined with camera intrinsic parameters, the initial shooting direction of each original image is iteratively corrected. By converting the pixel coordinates of corresponding image points in the image into direction vectors, the direction deviation is obtained, and the initial direction is updated. This iteration is repeated until the deviation is lower than a preset threshold, resulting in an accurate... The corrected shooting direction is used to map each original image to a direction vector in a unified spherical coordinate system. Alignment is then performed within the spherical orientation space, and pixel information from each image is fused to generate a UAV spherical panoramic image. Through alignment with pure rotation constraints, the stitched panoramic image maintains visual continuity in overlapping areas. It does not rely on depth information or 3D point coordinates, nor does it require depth measurement or 3D reconstruction of the scene. High-precision panoramic stitching can be achieved using only pure rotation information between images, significantly reducing computational complexity and dependence on the acquisition environment. Through rotation matrix and orientation unification processing, consistency in orientation among multiple images is ensured. This technology achieves visual continuity in overlapping areas, avoiding stitching distortion caused by translation or 3D reconstruction errors. It fully utilizes the rotational acquisition characteristics of drones hovering in the same spatial position to improve image acquisition efficiency and ensure complete panoramic coverage. By combining corrected shooting direction and spherical coordinate mapping, it achieves pixel-level image fusion, resulting in a smooth and continuous visual effect in the final panoramic image. It is suitable for various application scenarios such as drone aerial photography, natural resource management, farmland protection, virtual reality, and geographic information acquisition. It does not rely on depth information of environmental features, is applicable to different scenarios, and can effectively cope with complex or sparse features in overlapping areas, improving the stability and reliability of panoramic image stitching.

[0023] In a preferred embodiment, multiple raw images are acquired by the UAV while hovering at the same spatial position and rotating angularly around the optical center of the camera, and the corresponding camera intrinsic parameters are obtained, including: S101. Obtain the preset spatial position and control the drone to hover at the preset spatial position; S102. Make the drone camera rotate continuously around the camera optical center at a predetermined angle to acquire multiple original images covering the panoramic view. S103. During each image acquisition, automatically record the camera intrinsic parameters, including focal length, principal point coordinates, and distortion coefficient.

[0024] As described in steps S101 to S103 above, the drone is first controlled to hover in a preset spatial position to ensure that the positions of all acquired images are consistent. During the hovering process, the spatial position remains unchanged, thereby ensuring that the parallax between each image is mainly caused by the change in the camera's shooting direction, rather than the translation error caused by position movement. In the hovering state, the drone camera is controlled to rotate continuously around the optical center at a predetermined angle to acquire multiple original images covering the panoramic field of view. Since only the shooting direction is changed while the drone position remains fixed, the difference in perspective between each image comes only from the change in shooting direction. At the same time as image acquisition, the camera intrinsic parameters are automatically recorded, including focal length, principal point coordinates, and distortion coefficients. The intrinsic parameter information can be used to convert the image pixel coordinates into a direction vector with the camera's optical center as the starting point. Since the drone position is fixed, there is only a difference in perspective between each image caused by the change in shooting direction. This condition makes the geometric relationship between images completely described by rotation, without relying on scene depth or 3D reconstruction information.

[0025] In a preferred embodiment, feature extraction and feature matching are performed on multiple original images. For original image pairs with partially overlapping fields of view, corresponding image points located within the overlapping field of view are obtained. A rotation matrix representing the difference in image shooting direction is obtained based on the directional changes of the corresponding image points in their respective image imaging planes, including: S201. Extract features from multiple original images and perform feature matching between different original images; S202. For original image pairs whose shooting fields of view partially overlap, select the same image points located within the common shooting field of view from the matching results. S203. Based on the difference in shooting direction corresponding to the same image points in their respective image imaging planes, obtain the rotation matrix of the original image with respect to the difference in shooting direction.

[0026] As described in steps S201 to S203 above, local feature points are extracted from each original image, and descriptive information corresponding to each feature point is generated. Feature matching is performed between different original images to identify corresponding feature points in the same scene, i.e., homonymous image points. For original image pairs with overlapping fields of view (i.e., the two images cover the same area in space), homonymous image points located in the common field of view are selected from the feature matching results. Based on the directional changes of homonymous image points in their respective image imaging planes, a rotation matrix describing the difference in shooting direction between the original image pairs is obtained. The obtained rotation matrix only reflects the directional change, does not contain translation components, and does not depend on scene depth information, thus achieving pure rotation description, avoiding complex depth estimation or 3D reconstruction calculations, and reducing algorithm complexity.

[0027] In a preferred embodiment, feature extraction is performed on multiple original images respectively, and feature matching is performed between different original images, including: S2011. Based on each original image and the corresponding camera intrinsic parameter information, extract local feature points in each original image and generate corresponding feature description information for each local feature point. S2012. Obtain the direction reference information of the drone camera's orientation at the time of shooting, and obtain the relative shooting direction information between different original images based on the direction reference information; S2013. Based on the relative shooting direction information, select original images with overlapping field of view as image pairs to be matched. S2014. Obtain the similarity between feature points based on the feature description information of each image pair to be matched, and construct an initial feature matching set; S2015. Based on the initial feature matching set, and combined with the relative shooting direction information of the image pairs to be matched, the initial feature matches that do not conform to the preset direction constraints are eliminated, and the feature matching results that meet the preset direction constraints are retained.

[0028] As described in steps S2011 to S2015 above, for each original image, local feature points are extracted using camera intrinsic parameters (including focal length, principal point coordinates, distortion coefficients, etc.). Each local feature point generates corresponding feature description information for matching between different images. The feature description information can use local image feature descriptors (such as SIFT, ORB, SURF, etc.) to characterize the local texture and gradient information of the feature points, thereby achieving robust feature recognition. Using the UAV's orientation reference information at the time of shooting (such as IMU attitude, rotation angle, or GPS heading data), the relative shooting directions between different original images are obtained. Based on this relative orientation information, image pairs with overlapping fields of view are selected to ensure that feature matching is only performed on images with the same spatial coverage, thereby improving matching accuracy. To reduce unnecessary computation, for each image pair to be matched, the similarity between feature points is obtained using feature description information (the similarity between two feature points can be quantified by calculating the Euclidean distance or Hamming distance between feature description vectors; the smaller the distance, the more similar the features). Based on the similarity, the feature point pair with the highest similarity is used as the initial feature matching set. Combining the relative shooting direction information of the image pair to be matched, the initial feature matching results are filtered based on preset direction constraints (setting the maximum allowable shooting direction deviation range between images). Feature matching points that are inconsistent with the relative direction information or exceed the allowable deviation are eliminated, and matching results that meet the direction constraints are retained. Feature matching is only performed on image pairs with overlapping fields of view, avoiding meaningless matching of non-overlapping areas, reducing computational load, and improving system processing efficiency.

[0029] In a preferred embodiment, based on the difference in shooting direction corresponding to the same image points in their respective image imaging planes, a rotation matrix of the original image with respect to the difference in shooting direction is obtained, including: S2031. Based on the pixel position of each corresponding image point in the imaging plane of the corresponding original image, and combined with the camera intrinsic parameters of the corresponding original image, convert the position of each corresponding image point in the imaging plane into the shooting direction with the camera optical center as the starting point. S2032. For the same set of image points with the same name, obtain the corresponding shooting direction in the original image pair, and obtain the difference in shooting direction of the same image points in different original images based on the change in direction between the two. S2033. Based on the differences in shooting directions corresponding to multiple sets of image points with the same name, obtain rotation parameters that make the overall differences in shooting directions consistent. S2034. Based on the rotation parameters, generate a rotation matrix representing the difference in shooting direction between the original image pairs.

[0030] As described in steps S2031 to S2034 above, for each corresponding image point, its pixel coordinates in the original image imaging plane and the corresponding camera intrinsic parameters (focal length, principal point coordinates, distortion coefficients, etc.) are combined to convert the pixel coordinates into a direction vector originating from the camera optical center. For the same group of corresponding image points, the direction vector corresponding to it in the original image pair is obtained. By the angle or rotation relationship between the vectors, the direction difference of the corresponding image point in the two images is obtained, reflecting the change in shooting direction between the images, without relying on the actual depth or position change of the scene. Based on the direction difference of multiple groups of corresponding image points, by minimizing the overall error of the direction difference, the rotation parameters that make all direction differences consistent are obtained. In this process, no translation information between images is introduced, nor is scene depth information relied upon, ensuring the universality of the method in cases where the UAV is stationary or only rotates for shooting. Based on the obtained rotation parameters, a rotation matrix between the original image pairs is constructed to represent the difference in shooting direction between the image pairs.

[0031] In a preferred embodiment, based on the rotation matrix corresponding to each original image pair, the changes in the shooting direction of multiple original images are comprehensively correlated and unified to obtain a unified shooting direction result for multiple original images at the direction level, including: S301. Perform a consistency check on the rotation matrix corresponding to each original image pair, remove rotation matrices that do not match the shooting direction changes of most original images, and retain only rotation matrices with stable shooting direction changes as valid rotation matrices. S302. Among the multiple original images involved in the effective rotation matrix, select one original image as the direction reference image, and use the shooting direction of the direction reference image as the reference direction. S303. Based on the effective rotation matrix, according to the shooting order between the original images, the shooting direction change is gradually transferred from the direction reference image to the remaining original images to obtain the shooting direction parameters of each original image relative to the reference direction. S304. The shooting direction of each original image obtained through direction transfer is uniformly processed so that the shooting directions of multiple original images form corresponding shooting direction parameters under the same direction reference, thereby obtaining a unified shooting direction result of multiple original images at the direction level.

[0032] As described in steps S301 to S304 above, statistical analysis is performed on the rotation matrix corresponding to each original image pair. Rotation matrices that do not conform to the rotation trend of most original images are identified, and abnormal or unstable rotation matrices are eliminated. Only rotation matrices that conform to the overall shooting direction trend are retained as valid rotation matrices, ensuring the stability of the rotation relationship and reducing the impact of individual matching errors or abnormal images on the global orientation uniformity. In the set of original images involved in the valid rotation matrix, one image is selected as the orientation reference image, and its shooting direction is used as the global orientation benchmark. The orientation reference image is usually selected to cover the center of the panorama or the most stable image to ensure that the reference direction is representative. According to the acquisition order of the original images, the shooting direction of the orientation reference image is gradually passed to the adjacent images. The specific transmission method is through the image pairs. The effective rotation matrix is ​​used to multiply the reference image direction by the rotation matrix to obtain the shooting direction of the adjacent image relative to the reference direction. Then, the rotation matrix of the next image is used to continue the transfer until the shooting direction of all images is calculated. This step-by-step transfer method ensures that the shooting direction of each image is expressed under the same directional reference, while maintaining the consistency of the rotation relationship between images. The directional parameters of each image obtained through the transfer are uniformly processed, such as by optimizing the global minimization of rotation error or averaging the rotation parameters, so that the shooting direction of each image forms corresponding directional parameters under the same reference. The result after uniform processing is the unified shooting direction result at the directional level. The unified directional result of multiple original images can be directly used for iterative correction and spherical panoramic image fusion to improve stitching accuracy, without relying on scene depth information.

[0033] In a preferred embodiment, based on the unified shooting direction result, the initial shooting direction corresponding to each original image in a unified global orientation coordinate system is obtained, including: S401. Based on the unified shooting direction result, obtain the global orientation coordinate system of multiple original images for common orientation reference; S402. Obtain the camera coordinate system corresponding to each original image during image acquisition, and obtain the direction transformation result of the camera coordinate system relative to the global direction coordinate system based on the unified shooting direction result. S403. Based on the orientation transformation result, the shooting orientation of the original image in the camera coordinate system is transformed to the global orientation coordinate system to obtain the orientation representation result of the original image under a unified orientation reference. S404. Perform parameterization on the orientation representation results to extract the orientation parameters of the original image relative to the global orientation coordinate system. S405. Use the orientation parameter as the initial shooting orientation of the corresponding original image in a unified global orientation coordinate system.

[0034] As described in steps S401 to S405 above, based on the unified shooting direction result, a global orientation coordinate system shared by multiple original images is defined. The global orientation coordinate system serves as a unified reference, used to map the shooting directions of different images to the same representation standard. For each original image, the camera coordinate system information at the time of image acquisition is obtained. Based on the unified shooting direction result, the rotation relationship between the camera coordinate system and the global orientation coordinate system is obtained, i.e., the orientation transformation matrix, which realizes the mapping basis from the local camera coordinate system to the global orientation coordinate system. Using the orientation transformation result, the shooting direction of each image in the camera coordinate system is mapped to the global orientation coordinate system. After the transformation, the orientation information of each image is represented as a orientation vector or rotation parameter relative to the unified orientation reference, realizing the comparability and composability of different image orientations. The transformed orientation information is parameterized (e.g., extracting the rotation matrix or Euler angles, quaternion-represented orientation parameters), and the parameterized orientation result is used as the initial shooting direction of each original image in the global orientation coordinate system.

[0035] In a preferred embodiment, the initial shooting direction of each original image is corrected based on camera intrinsic parameters to obtain a corrected shooting direction, including: S501. Based on the camera intrinsic parameters, convert the pixel coordinates of each corresponding image point in each original image into a direction vector originating from the camera optical center. S502. Obtain the direction deviation based on the direction vector of the same image point between the current image and the adjacent images; S503. Obtain a preset deviation threshold and determine whether the directional deviation is lower than the preset deviation threshold. If the directional deviation is lower than the preset deviation threshold, the initial shooting direction of the original image will be used as the corrected shooting direction. If the directional deviation is not lower than the preset deviation threshold, the initial shooting direction of the current image is updated by combining the camera intrinsic parameters to obtain the updated shooting direction. The directional deviation is re-judged based on the updated shooting direction until the directional deviation of each original image is lower than the preset threshold, and the corrected shooting direction of each original image is obtained.

[0036] As described in steps S501 to S503 above, by combining the camera intrinsic parameters (including focal length, principal point coordinates, and distortion coefficients) of each original image, the pixel coordinates of each corresponding image point in the image are converted into a direction vector originating from the camera optical center. In this way, the two-dimensional position on the pixel plane is mapped to a three-dimensional direction space, allowing the acquisition of rotation differences to be based on direction vectors rather than pixel coordinates. This avoids the influence of scene depth on rotation calculations. The direction vectors of corresponding image points in each image are compared with those in its neighboring images to obtain the direction deviation, which is used to characterize the rotation differences between images. The direction deviation reflects the rotation error between the image's shooting direction and that of its neighboring images and is the core quantity for correcting the shooting direction. A preset orientation deviation threshold is used to determine whether the orientation accuracy of the current image meets the requirements. If the orientation deviation is lower than the threshold, the initial shooting orientation of the image is directly used as the corrected shooting orientation, indicating that the orientation of the image has met the accuracy requirements. If the orientation deviation is not lower than the threshold, the initial shooting orientation is updated according to the orientation deviation and camera intrinsic parameters to obtain a new shooting orientation. The updated orientation continues to be used for deviation calculation and threshold determination, and this process is repeated iteratively until the orientation deviation is lower than the preset threshold to obtain the final corrected shooting orientation of each image. The method is entirely based on orientation vectors and rotation relationships, without relying on scene depth information or external measurement data, and achieves orientation correction without depth estimation, which is suitable for complex or texture-sparse scenes.

[0037] In a preferred embodiment, based on the corrected shooting direction, each original image is mapped to a unified spherical coordinate system and fused to generate a UAV spherical panoramic image, including: S601. Based on the corrected shooting direction of each original image, convert the coordinates of each pixel in the camera coordinate system into a direction vector in the spherical coordinate system. S602. Align the direction vectors of multiple original images in a unified spherical coordinate system to obtain the direction vector alignment result. S603. Based on the direction vector alignment result, the pixel information of each original image is fused to generate a UAV spherical panoramic image.

[0038] As described in steps S601 to S603 above, the corrected shooting direction of each original image is used to convert the two-dimensional coordinates of each pixel in the image into a three-dimensional direction vector originating from the camera's optical center, and then mapped to a unified spherical coordinate system. During the mapping process, precise projection is performed using camera intrinsic parameters (focal length, principal point coordinates, distortion coefficients), ensuring that each pixel has corresponding direction information in the spherical coordinate system without relying on scene depth or three-dimensional point coordinates. The direction vectors mapped from multiple images are aligned within the unified spherical coordinate system, ensuring that the direction vectors in the overlapping field of view remain continuous and consistent. The alignment process is based on the corrected shooting direction and the rotation between the images. By correcting directional deviations through a rotation matrix, global consistency in orientation between images is achieved. This method ensures the spatial orientation continuity of pixels in overlapping areas without requiring scene depth information or 3D spatial coordinates, guaranteeing stitching accuracy and visual consistency. Based on the orientation vector alignment results, pixel information from multiple original images is fused to generate a drone spherical panoramic image. During the fusion process, the continuity of the corrected shooting direction and orientation vector is fully utilized to ensure smooth transition of pixels in overlapping areas, reducing ghosting or breaks, and improving the accuracy and visual continuity of the panoramic image. The entire process does not rely on depth estimation, achieving pure rotation-optimized panoramic image generation.

[0039] And, a depth estimation-free, purely rotation-optimized UAV panoramic image stitching terminal, comprising: One or more processors; A storage device on which one or more programs are stored; When one or more programs are executed by one or more processors, the one or more processors implement a pure rotation-optimized UAV panoramic image stitching method that eliminates depth estimation.

[0040] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.

Claims

1. A method for stitching panoramic images from unmanned aerial vehicles (UAVs) with pure rotation optimization and no depth estimation required, characterized in that, include: Acquire multiple raw images by rotating the drone around the optical center of the camera while it is hovering in the same spatial position, and obtain the corresponding camera intrinsic parameters; Feature extraction and feature matching are performed on multiple original images. For original image pairs with overlapping fields of view, corresponding image points located in the overlapping field of view are obtained. The rotation matrix of the difference in the shooting direction of the images is obtained based on the directional changes of the corresponding image points in their respective image imaging planes. Based on the rotation matrix corresponding to each original image pair, the shooting direction changes of multiple original images are correlated and unified as a whole to obtain the unified shooting direction result of multiple original images at the direction level. Based on the unified shooting direction result, obtain the initial shooting direction corresponding to each original image in the unified global direction coordinate system; Based on the camera's internal parameters, the shooting direction of each original image is corrected to obtain the corrected shooting direction; Based on the corrected shooting direction, each original image is mapped to a unified spherical coordinate system and the images are fused to generate a drone spherical panoramic image. Based on the rotation matrix corresponding to each original image pair, the changes in the shooting direction of multiple original images are correlated and unified as a whole, resulting in a unified shooting direction result for multiple original images at the direction level, including: For each original image, a consistency check is performed on the corresponding rotation matrix. Rotation matrices that do not match the shooting direction changes of most original images are removed, and only rotation matrices with stable shooting direction changes are retained as valid rotation matrices. Among the multiple original images involved in the effective rotation matrix, one original image is selected as the orientation reference image, and the shooting direction of the orientation reference image is used as the reference direction. Based on the effective rotation matrix, according to the shooting order between the original images, the shooting direction change is gradually transferred from the direction reference image to the remaining original images to obtain the shooting direction parameters of each original image relative to the reference direction. The shooting direction of each original image obtained through direction transfer is uniformly processed so that the shooting directions of multiple original images form corresponding shooting direction parameters under the same direction reference, thereby obtaining a unified shooting direction result of multiple original images at the direction level.

2. The method for stitching panoramic UAV images with pure rotation optimization and no depth estimation as described in claim 1, characterized in that, Acquire multiple raw images taken by the drone while hovering at the same spatial position and rotating angularly around the optical center of the camera, and obtain the corresponding camera intrinsic parameters, including: Obtain a preset spatial location and control the drone to hover within that location; The drone camera rotates continuously around the camera's optical center at a predetermined angle to acquire multiple raw images covering the panoramic view. During each image acquisition, camera intrinsic parameters are automatically recorded, including focal length, principal point coordinates, and distortion coefficients.

3. The method for stitching panoramic UAV images with pure rotation optimization and no depth estimation as described in claim 1, characterized in that, Feature extraction and feature matching are performed on multiple original images. For original image pairs with overlapping fields of view, corresponding image points located within the overlapping field of view are obtained. A rotation matrix representing the difference in image shooting direction is obtained based on the directional changes of these corresponding image points in their respective image imaging planes, including: Feature extraction is performed on multiple original images, and feature matching is performed between different original images; For original image pairs whose fields of view partially overlap, select the corresponding image points located within the common field of view from the matching results; Based on the difference in shooting direction corresponding to the same image points in their respective image imaging planes, the rotation matrix of the original image with respect to the difference in shooting direction is obtained.

4. The method for stitching panoramic UAV images with pure rotation optimization and no depth estimation as described in claim 3, characterized in that, Feature extraction is performed on multiple original images, and feature matching is performed between different original images, including: Based on each original image and its corresponding camera intrinsic parameters, local feature points are extracted from each original image, and corresponding feature description information is generated for each local feature point. Obtain directional reference information of the drone camera's orientation at the moment of shooting, and obtain relative shooting direction information between different original images based on the directional reference information; Based on the relative shooting direction information, original images with overlapping fields of view are selected as image pairs to be matched. Based on the feature description information of each image pair to be matched, the similarity between feature points is obtained, and an initial feature matching set is constructed. Based on the initial feature matching set, and combined with the relative shooting direction information of the image pairs to be matched, the initial feature matches that do not conform to the preset direction constraints are eliminated, and the feature matching results that meet the preset direction constraints are retained.

5. The method for stitching panoramic UAV images with pure rotation optimization and no depth estimation as described in claim 3, characterized in that, Based on the difference in shooting direction corresponding to the same image points in their respective image imaging planes, the rotation matrix of the original image with respect to the difference in shooting direction is obtained, including: Based on the pixel position of each corresponding image point in the corresponding original image imaging plane, and combined with the camera intrinsic parameters of the corresponding original image, the position of each corresponding image point in the imaging plane is converted into the shooting direction with the camera optical center as the starting point. For the same set of image points with the same name, obtain the corresponding shooting direction in the original image pair, and obtain the difference in shooting direction of the same image points in different original images based on the change in direction between the two. Based on the differences in shooting directions corresponding to multiple sets of image points with the same name, rotation parameters are obtained to make the overall differences in shooting directions consistent. Based on the rotation parameters, a rotation matrix is ​​generated to represent the difference in shooting direction between the original image pairs.

6. The method for stitching panoramic UAV images with pure rotation optimization and no depth estimation as described in claim 1, characterized in that, Based on the unified shooting direction results, the initial shooting direction of each original image in the unified global orientation coordinate system is obtained, including: Based on the unified shooting direction result, a global orientation coordinate system for common orientation reference of multiple original images is obtained; Obtain the camera coordinate system corresponding to each original image during image acquisition, and obtain the orientation transformation result of the camera coordinate system relative to the global orientation coordinate system based on the unified shooting direction result; Based on the orientation transformation result, the shooting orientation of the original image in the camera coordinate system is transformed to the global orientation coordinate system, and the orientation representation result of the original image under a unified orientation reference is obtained. The orientation representation results are parameterized to extract the orientation parameters of the original image relative to the global orientation coordinate system. The orientation parameter is used as the initial shooting orientation of the corresponding original image in a unified global orientation coordinate system.

7. The method for stitching panoramic images of UAVs with pure rotation optimization and no depth estimation as described in claim 1, characterized in that, Based on the camera's intrinsic parameters, the initial shooting direction of each original image is corrected to obtain the corrected shooting direction, including: Based on the camera intrinsic parameters, the pixel coordinates of each corresponding image point in each original image are converted into a direction vector originating from the camera optical center. The directional deviation is obtained based on the directional vectors of the corresponding image points between the current image and adjacent images. Obtain a preset deviation threshold and determine whether the directional deviation is lower than the preset deviation threshold; If the directional deviation is lower than the preset deviation threshold, the initial shooting direction of the original image will be used as the corrected shooting direction. If the directional deviation is not lower than the preset deviation threshold, the initial shooting direction of the current image is updated by combining the camera intrinsic parameters to obtain the updated shooting direction. The directional deviation is re-judged based on the updated shooting direction until the directional deviation of each original image is lower than the preset threshold, and the corrected shooting direction of each original image is obtained.

8. The method for stitching panoramic images of UAVs with pure rotation optimization and no depth estimation as described in claim 1, characterized in that, Based on the corrected shooting direction, each original image is mapped to a unified spherical coordinate system and then fused to generate a drone spherical panoramic image, including: Based on the corrected shooting direction of each original image, the coordinates of each pixel in the camera coordinate system are converted into a direction vector in the spherical coordinate system. Align the orientation vectors of multiple original images in a unified spherical coordinate system to obtain the orientation vector alignment result. Based on the orientation vector alignment results, the pixel information of each original image is fused to generate a drone spherical panoramic image.

9. A pure rotation-optimized UAV panoramic image stitching terminal without depth estimation, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When one or more programs are executed by one or more processors, the one or more processors implement the depth estimation-free pure rotation optimization UAV panoramic image stitching method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Human body three-dimensional reconstruction method and system based on camera

    CN121392136A

  • Vehicle-mounted panoramic image generation method, computer device, computer storage medium, computer program product and mobile platform

    WO2025246691A1