Unmanned aerial vehicle video stitching method, device and system based on Gaussian radiation field
Through the Gaussian radiation field-based method, 3DGS is used to perform three-dimensional reconstruction and rendering of drone images, solving the problems of low accuracy and weak parallax resistance in high parallax conditions by traditional stitching methods, and achieving high-quality and high-resolution drone video stitching.
Patent Information
- Application Number
- CN202510229474.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
The traditional drone image video stitching method has low accuracy when processing images with large parallaxes, resulting in poor stitching effect and weak parallax resistance.
Using a Gaussian radiation field-based method, the image sequence collected by the drone is three-dimensionally reconstructed through 3DGS to generate a panoramic three-dimensional point set, and the image is rendered using this point set. By optimizing the viewing angle to minimize the loss between images, high-resolution drone video is generated.
The quality of drone video stitching is improved, parallax resistance is enhanced, splicing efficiency is improved, and high-resolution wide-field drone video is generated.
Smart Images

Figure CN120182085A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of aerial image processing, and more specifically, relates to a method, device, and system for unmanned aerial vehicle (UAV) video stitching based on a Gaussian radiation field. Background Art
[0002] When a UAV collects aerial video information, it needs to rise to a relatively high altitude to obtain a large-field-of-view image. However, due to the limitations of the performance of the onboard camera, the resolution of the obtained large-field-of-view image is relatively low. When flying at a low altitude, although the UAV can capture high-resolution images, the field of view is limited, and the amount of information obtained is limited. To solve this problem, multiple UAV images can be stitched together so that a large field of view can be presented in one image, thereby achieving the acquisition of high-resolution aerial images.
[0003] Traditional UAV image video stitching methods generally achieve coordinate transformation by matching the feature points of two UAV images. However, for two large-baseline UAV images with a large parallax, the accuracy of existing feature point matching methods is relatively low, resulting in a low stitching accuracy of the images. In addition, due to the large baseline between adjacent UAVs, there will be an obvious parallax between UAV images, which will cause artifacts in the overlapping area of the two images after stitching, resulting in poor visual effects. Therefore, traditional UAV image stitching methods have disadvantages such as highly relying on the quality of feature point matching and weak anti-parallax ability.
[0004] Currently, many technologies for image stitching using deep learning have been proposed at home and abroad. These methods extract features by using convolutional neural networks or autoencoders, and then align and fuse different images through algorithms (such as transformation networks, image registration). Deep learning can also optimize the geometric transformation and color transition between images, reduce seams and distortions in traditional stitching, and make the stitching effect more natural and delicate. However, these methods still cannot handle UAV images with a large parallax well, and the quality of the finally stitched UAV video still needs to be further improved. Summary of the Invention
[0005] Aiming at the defects and improvement requirements of the prior art, the present invention provides a method, device, and system for UAV video stitching based on a Gaussian radiation field, aiming to improve the quality of the stitched UAV video.
[0006] To achieve the above object, according to one aspect of the present invention, a method for UAV video stitching based on a Gaussian radiation field is provided. The UAV video is collected by an UAV cluster across scales; the UAV cluster includes a main UAV and n 2Sub - drones; the height of each sub - drone from the ground is the same, and there is an overlap between the fields of view of adjacent sub - drones; the height of the main drone from the ground is greater than that of the sub - drones, and the field of view of the main drone exactly covers the fields of view of all sub - drones; n is a positive integer greater than 1;
[0007] The method for stitching drone videos includes:
[0008] S1: Use 3DGS to perform three - dimensional reconstruction on the image sequences collected by each sub - drone respectively, obtain the three - dimensional point sets under the fields of view of each sub - drone, and stitch them into a panoramic three - dimensional point set R;
[0009] S2: Select the first image from the image sequence collected by the main drone as the reference image Q, and randomly initialize the viewing angle C;
[0010] S3: Render an image I from the current viewing angle C using the panoramic three - dimensional point set R, as the rendered image I corresponding to the current reference image Q;
[0011] S4: Calculate the loss L between the reference image Q and the rendered image I, and optimize the viewing angle C with the goal of minimizing the loss L, and use the rendered image corresponding to the optimized viewing angle C as the rendering result of the current reference image Q;
[0012] S5: Determine whether all the images collected by the main drone have been rendered. If so, go to S6; otherwise, select the next image of the current reference image Q from the image sequence collected by the main drone as the new reference image Q, and go to S3;
[0013] S6: Use the sequence of rendering results corresponding to the image sequence collected by the main drone as the stitched drone video.
[0014] Furthermore, in S1, stitching the three - dimensional point sets under the fields of view of each drone into a panoramic three - dimensional point set R includes:
[0015] S11: Divide all the three - dimensional point sets into multiple groups, each group contains two three - dimensional point sets with an overlapping area, or contains only one three - dimensional point set, and the number of groups containing only one three - dimensional point set is not greater than 1;
[0016] S12: For each group containing two three - dimensional point sets, perform the following steps:
[0017] S121: Render images image1 and image2 from the viewing angles C1 and C2 respectively using the three - dimensional point sets R1 and R2 in the current group; there is an overlapping area between images image1 and image2;
[0018] S122: Extract the feature points in images image1 and image2 respectively, and obtain the matching relationship between the feature points of the two images;
[0019] S123: Calculate the geometric transformation matrix between the three-dimensional point sets R1 and R2 according to the feature matching relationship between images image1 and image2, and map all the three-dimensional points in the three-dimensional point set R2 to the coordinate system of the three-dimensional point set R1 according to the geometric transformation matrix to obtain the three-dimensional point set R2';
[0020] S124: Concatenate the three-dimensional point set R1 and the three-dimensional point set R2' into a new three-dimensional point set; when concatenating, for the points that are located in the overlapping area and match each other in the three-dimensional point sets R1 and R2, a weighted fusion method is used for concatenation;
[0021] S13: If the number of three-dimensional point sets is greater than 1 after concatenation, go to S11; otherwise, use the three-dimensional point set containing all the three-dimensional points under the fields of view of the sub-unmanned aerial vehicles as the panoramic three-dimensional point set R.
[0022] Further, in S12, if there are multiple groups to be concatenated, the concatenation of multiple groups is executed in parallel.
[0023] Further, in S124, for the points that are located in the overlapping area and match each other in the three-dimensional point sets R1 and R2, when performing weighted fusion, the weights of the two points are the same.
[0024] Further, S122 further includes: after obtaining the matching relationship between the feature points of the two images, removing the matching relationships between the incorrect feature points.
[0025] Further, the expression of the loss L between the reference image Q and the rendered image I is as follows:
[0026] L=(1 - λ)L com +λL ma
[0027] where, L com represents the comparison loss between the reference image Q and the rendered image I, and L ma represents the registration loss between the reference image Q and the rendered image I; λ represents the balance coefficient, and λ ∈ [0, 1].
[0028] According to another aspect of the present invention, there is provided a computer program product, including a computer program; when the computer program is executed by a processor, it implements the above-mentioned method for stitching unmanned aerial vehicle videos based on Gaussian radiation fields provided by the present invention.
[0029] According to another aspect of the present invention, there is provided a computer-readable storage medium, including a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the above-mentioned method for stitching drone videos based on the Gaussian radiation field provided by the present invention.
[0030] According to another aspect of the present invention, there is provided a device for stitching drone videos based on the Gaussian radiation field, including:
[0031] A computer-readable storage medium for storing a computer program;
[0032] And a processor for reading the computer program stored in the computer-readable storage medium to implement the above-mentioned method for stitching drone videos based on the Gaussian radiation field provided by the present invention.
[0033] According to another aspect of the present invention, there is provided a drone video acquisition system, including: a cross-scale drone cluster and the above-mentioned device for stitching drone videos based on the Gaussian radiation field provided by the present invention;
[0034] The drone cluster includes a main drone and n 2 sub-drones; the height of each sub-drone from the ground is the same, and there is an overlap between the fields of view of adjacent sub-drones; the height of the main drone from the ground is greater than that of the sub-drones, and the field of view of the main drone exactly covers the fields of view of all sub-drones; n is a positive integer greater than 1.
[0035] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0036] (1) The present invention first uses 3DGS to perform three-dimensional reconstruction on the detailed image sequences collected by each sub-drone to obtain the corresponding three-dimensional point sets of each sub-drone, and can utilize the advantages of 3DGS in high-quality three-dimensional reconstruction and view-independent real-time rendering; then, the three-dimensional point sets corresponding to each sub-drone are stitched into a panoramic three-dimensional point set, and the panoramic three-dimensional point set can be rendered to obtain high-resolution images from any perspective; on this basis, for each image in the panoramic image sequence collected by the main drone, the present invention uses the panoramic three-dimensional point set to render the corresponding image, and by constructing the loss between the two and iteratively optimizing the rendering perspective, the problem of lacking pose information of the required perspective can be effectively solved, and finally a drone image sequence with a wide field of view and high resolution, that is, a drone video, is generated. Compared with the traditional method for stitching drone videos, the quality of the finally stitched drone video is effectively improved.
[0037] (2) The present invention gradually stitches small 3D point sets with overlapping regions into a large 3D point set in a grouped stitching manner until a panoramic 3D point set containing all the 3D point sets from the perspectives of the sub-unmanned aerial vehicles (UAVs) is obtained, thereby effectively improving the stitching efficiency of the 3D point sets and effectively alleviating the problems of long time consumption and high computational resource consumption introduced by 3D reconstruction. In a further preferred embodiment, the stitching of the 3D point sets in multiple groups is executed in parallel, which can further improve the stitching efficiency of the UAV videos.
[0038] (3) When calculating the loss between the reference image Q in the image sequence collected by the master UAV in the calculation and the rendered image I rendered from the panoramic 3D point set at a specific perspective, the present invention takes into account both the comparison loss and the registration loss between the two. The former can achieve pixel-level image difference analysis and effectively capture the subtle changes in the camera poses, while the latter can directly measure the geometric position difference between images. The combination of the two can more accurately measure the difference between images and provide a more accurate basis for the iterative optimization of the pose (perspective). Description of the Drawings
[0039] Figure 1 is a schematic plan view of the arrangement of the cross-scale UAV cluster provided by an embodiment of the present invention;
[0040] Figure 2 is a schematic three-dimensional view of the arrangement of the cross-scale UAV cluster provided by an embodiment of the present invention;
[0041] Figure 3 is a flowchart of the UAV video stitching method based on the Gaussian radiation field provided by an embodiment of the present invention. Detailed Embodiments
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0043] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0044] To solve the problems of the existing UAV video stitching method, which highly depends on the feature point matching quality and has weak anti-parallax ability, considering the advantages of 3DGS (3D Gaussian Splatting) in high-quality 3D reconstruction and view-independent real-time rendering, etc., the present invention proposes to use multiple groups of UAV images for 3D reconstruction, and then render the required large field of view and high-resolution UAV images and videos. However, for the existing large-scale 3D reconstruction methods, there are problems of long time consumption and large consumption of computing resources. And due to the lack of pose information of the required viewpoints, how to iteratively optimize the required large field of view UAV images and videos is also one of the problems to be solved. Therefore, even the neural radiance field, which has better performance and flexibility in processing complex 3D scene rendering, cannot be directly applied to UAV image and video stitching.
[0045] To solve the above problems, the present invention further proposes that the master UAV and slave UAVs in a cross-scale master-slave UAV cluster collect panoramic images and detail images from multiple scales. First, 3DGS is used to perform 3D reconstruction on the image sequences collected by each slave UAV to obtain the 3D point sets in the Gaussian fields corresponding to each slave UAV. Then, the transformation matrix between two Gaussian fields is calculated using two images with a certain overlapping area in adjacent two Gaussian fields, and the Gaussian fields of the slave UAVs are fused in turn. Then, in a self-supervised manner, using the loss between the rendered image and the master UAV image, the stable high-resolution image sequence with the same field of view as the master UAV is iteratively optimized, so as to make the multi-scale UAV video stitching more accurate and the stitching efficiency higher.
[0046] The following are embodiments.
[0047] Embodiment 1:
[0048] A UAV video stitching method based on Gaussian radiance field.
[0049] In this embodiment, the UAV video is collected by a cross-scale UAV cluster; as Figure 1 and Figure 2As shown in the figure, the UAV cluster includes a main UAV and 4 sub-UAVs; the height of each sub-UAV from the ground is the same, and there is an overlap between the fields of view of adjacent sub-UAVs, so there is also a certain overlap between the images captured by adjacent sub-UAVs, ensuring that the three-dimensional scenes of the sub-UAVs can be stitched; the height of the main UAV from the ground is greater than that of the sub-UAVs, and the field of view of the main UAV exactly covers the fields of view of all sub-UAVs. Specifically, in this embodiment, the 4 sub-UAVs responsible for capturing detailed scenes are arranged at the four corners of a square, with the same height from the ground, all being 55m, the distance between adjacent UAVs being 50m, and the overlapping area between the images captured by adjacent UAVs being 5% - 10%. In order to make the panoramic image exactly contain all the scene contents of the 4 detailed images, the main UAV responsible for capturing the panoramic image is located at the center of the 4 sub-UAVs and has a height of 100m from the ground. The target scenes include parks, stadiums, parking lots, etc., and the image resolution is 1920*1080.
[0050] Through the above cross-scale UAV cluster, in this embodiment, 4 detailed image sequences captured by 4 sub-UAVs and one panoramic image sequence captured by the main UAV can be collected.
[0051] It should be noted that the cross-scale UAV cluster described here is only an optional UAV cluster of the present invention and should not be understood as the only limitation of the present invention. In practical applications, the number of sub-UAVs can be flexibly adjusted, but the number needs to be n 2 (n is a positive integer greater than 1). At the same time, the height of each sub-UAV from the ground is the same, and there is an overlap between the fields of view of adjacent sub-UAVs; the height of the main UAV from the ground is greater than that of the sub-UAVs, and the field of view of the main UAV exactly covers the fields of view of all sub-UAVs.
[0052] For the UAV video collected by the above cross-scale UAV cluster, as Figure 3 shown, the UAV video stitching method proposed in this embodiment includes:
[0053] S1: Use 3DGS to perform three-dimensional reconstruction on the image sequences collected by each sub-UAV respectively to obtain the three-dimensional point sets under the fields of view of each sub-UAV, and stitch them into a panoramic three-dimensional point set R;
[0054] S2: Select the first image from the image sequence collected by the main UAV as the reference image Q, and randomly initialize the viewing angle C;
[0055] S3: Use the panoramic three-dimensional point set R to render the image I from the current viewing angle C as the rendered image I corresponding to the current reference image Q;
[0056] S4: Calculate the loss L between the reference image Q and the rendered image I, and optimize the viewing angle C with the goal of minimizing the loss L. Use the rendered image corresponding to the optimized viewing angle C as the rendering result of the current reference image Q.
[0057] S5: Determine whether all the images collected by the master UAV have been rendered. If so, go to S6; otherwise, select the next image of the current reference image Q from the image sequence collected by the master UAV as the new reference image Q, and go to S3.
[0058] S6: Use the sequence of rendering results corresponding to the image sequence collected by the master UAV as the stitched UAV video.
[0059] In this embodiment, first, 3DGS is used to perform three-dimensional reconstruction on the detailed image sequences collected by each sub-UAV to obtain the three-dimensional point sets corresponding to each sub-UAV, taking advantage of the advantages of 3DGS in high-quality three-dimensional reconstruction and view-independent real-time rendering. Then, the three-dimensional point sets corresponding to each sub-UAV are stitched into a panoramic three-dimensional point set, which can be rendered to obtain high-resolution images at any viewing angle. On this basis, for each image in the panoramic image sequence collected by the master UAV, the panoramic three-dimensional point set is used to render the corresponding image. By constructing the loss between the two and iteratively optimizing the rendering viewing angle, the problem of lacking pose information of the required viewing angle can be effectively solved, and finally a UAV image sequence with a large field of view and high resolution, that is, a UAV video, is generated. Compared with traditional UAV video stitching methods, the quality of the finally stitched UAV video is effectively improved.
[0060] To solve the problems of long time consumption and high computing resource consumption introduced by three-dimensional reconstruction, in this embodiment, when stitching the three-dimensional point sets in the fields of view of each sub-UAV, a grouped stitching method is adopted. Specifically, in S1, stitching the three-dimensional point sets in the fields of view of each UAV into a panoramic three-dimensional point set R includes:
[0061] S11: Divide all the three-dimensional point sets into multiple groups, each group containing two three-dimensional point sets with overlapping regions, or only containing one three-dimensional point set, and the number of groups containing only one three-dimensional point set is no more than 1.
[0062] S12: For each group containing two three-dimensional point sets, perform the following steps:
[0063] S121: Through perspective projection, use the three-dimensional point sets R1 and R2 in the current group to render images image1 and image2 from viewing angles C1 and C2 respectively; there is an overlapping region between images image1 and image2.
[0064] S122: extracting feature points from images image1 and image2 respectively, and obtaining a matching relationship between the feature points of the two images;
[0065] Optionally, in this embodiment, S122 uses the SIFT algorithm to extract feature points in the images image1 and image2, and the result of the feature point extraction includes the coordinates and descriptors of the feature points; the extracted single feature point can be expressed as kp=(x, y, d), where x and y represent the horizontal coordinate and the vertical coordinate of the feature point respectively, and d represents the feature point descriptor;
[0066] Based on the feature point extraction results, the Euclidean distance between the feature points is calculated according to the feature point descriptor, and the matching relationship between the feature points in image1 and image2 can be obtained by using matching algorithms such as brute force matching.
[0067] In order to ensure the accuracy of the acquired feature point matching relationship, as a preferred implementation, in this embodiment, after obtaining the matching relationship between the feature points of the two images, S122 will remove the wrong matching relationship between the feature points; in practical applications, the RANSAC algorithm can be used to identify the wrong matching relationship;
[0068] S123: Calculate the geometric transformation matrix between the three-dimensional point sets R1 and R2 according to the feature matching relationship between the images image1 and image2, and map the three-dimensional points in the three-dimensional point set R2 to the coordinate system of the three-dimensional point set R1 according to the geometric transformation matrix to obtain the three-dimensional point set R2';
[0069] Specifically, according to the feature matching relationship between image1 and image2, the geometric transformation matrix H of image2 in the image1 coordinate system is calculated:
[0070]
[0071] Among them, H 11 Indicates the scaling and rotation in the x direction, H 12 represents the shear and rotation in the x direction, H 13 represents the translation in the x direction, H 21 represents shear and rotation in the y direction, H 22 Indicates the scaling and rotation in the y direction, H 23 represents the translation in the y direction, H 31 represents the perspective transformation in the x direction, H 32 represents the perspective transformation in the y direction, H 33 Represents the overall scaling factor, usually set to 1.
[0072] According to the geometric transformation matrix H between image1 and image2, taking R1 as the reference object, mapping the three-dimensional point set in R2 to the coordinate system of R1, the three-dimensional point set R1' can be obtained;
[0073] S124: Concatenate the three-dimensional point set R1 and the three-dimensional point set R2' into a new three-dimensional point set; when concatenating, for the points located in the overlapping area and matching each other in the three-dimensional point sets R1 and R2, a weighted fusion method is used for concatenation, and when concatenating, the weights of the two points are the same;
[0074] S13: If the number of three-dimensional point sets is greater than 1 after concatenation, go to S11; otherwise, use the three-dimensional point set containing all the three-dimensional point sets under the fields of view of the sub-unmanned aerial vehicles as the panoramic three-dimensional point set R.
[0075] In this embodiment, the UAV cluster includes 4 UAVs, and 4 three-dimensional point sets are obtained accordingly. When concatenating, the three-dimensional point sets will first be divided into 2 groups, each group containing the three-dimensional point sets under the fields of view of two sub-UAVs. After concatenating the three-dimensional point sets in these 2 groups respectively, 2 new three-dimensional point sets will be obtained. Then, these two new three-dimensional point sets will be divided into a new group, and after concatenating the 2 three-dimensional point sets in this new group, the panoramic three-dimensional point set can be obtained.
[0076] The above process realizes the concatenation of three-dimensional point sets in a grouped manner, which can effectively improve the concatenation efficiency. To further improve the concatenation efficiency of UAV videos, in S12 of this embodiment, if there are multiple groups that need to be concatenated, the concatenation of multiple groups is executed in parallel. In addition, in S1, the operation of three-dimensional reconstruction of the image sequences collected by each sub-UAV using 3DGS can also be executed in parallel.
[0077] When using Figure 1 the cross-scale UAV cluster shown to collect image sequences, there will be a problem of lacking the pose information of the required viewpoints. In this case, in order to effectively iterate the required large-field-of-view UAV image video, based on obtaining the panoramic three-dimensional point set R, this embodiment proposes an iterative optimization rendering method based on comparison loss and registration loss. Among them, the registration loss evaluates the pose matching degree by directly measuring the geometric position difference between images, which is highly consistent with the core principle of pose calculation; intuitively, when the spatial distribution of the matching feature points in two images is closer, the corresponding camera poses are more consistent; the comparison loss uses pixel-level image difference analysis and can effectively capture the subtle changes between camera poses; this method provides a more refined evaluation dimension for pose estimation by accurately quantifying the pixel differences between images, thereby achieving sub-pixel-level prediction accuracy of camera poses.
[0078] Accordingly, when calculating the loss L between the reference image Q and the corresponding rendered image I, the present invention adopts a combined loss function based on the comparison loss L com and the registration loss L ma .
[0079] The comparison loss L com uses the mean squared error (MSE) to compare the rendered image I and the reference image Q, and the calculation formula is as follows:
[0080] L com = MSE(I, Q)
[0081] The introduction of the comparison loss can achieve iterative optimization through pixel-level alignment;
[0082] The registration loss L ma uses the Euclidean distance between the matched feature points as a metric to quantify the difference between two poses. Optionally, in this embodiment, a pre-trained model LoFTR (a detector-free local feature matching method) is used to identify the corresponding feature point pairs between the rendered image I and the reference image Q to achieve iterative optimization by reducing the Euclidean distance between the feature point pairs:
[0083]
[0084] where k represents the number of feature point pairs, and f1 j and represent the feature points located in the rendered image I and the reference image Q in the jth matched feature point pair respectively.
[0085] Finally, the expression of the combined loss L is as follows:
[0086] L = (1 - λ)L com + λL ma
[0087] where L com represents the comparison loss between the reference image Q and the rendered image I, and L ma represents the registration loss between the reference image Q and the rendered image I; λ represents the balance coefficient, and λ ∈ [0, 1].[[]END]]
[0088] Through the iterative optimization of the combined loss L, this embodiment can obtain a high-resolution image with the same field of view as the reference image Q. At this time, the viewing angle C is the optimized viewing angle C, and this viewing angle C will be used as the input viewing angle for rendering the next image in the panoramic image sequence collected by the main drone to obtain the corresponding rendering result. Repeating this process can generate a wide-field-of-view, high-resolution drone image sequence and finally form a wide-field-of-view, high-resolution drone video.
[0089] Embodiment 2:
[0090] A computer program product includes a computer program; when the computer program is executed by a processor, it implements the method for stitching drone videos based on a Gaussian radiation field provided in the above-mentioned Embodiment 1.
[0091] Embodiment 3:
[0092] A computer-readable storage medium includes a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the method for stitching drone videos based on a Gaussian radiation field provided in the above-mentioned Embodiment 1.
[0093] Embodiment 4:
[0094] A device for stitching drone videos based on a Gaussian radiation field includes:
[0095] A computer-readable storage medium for storing a computer program;
[0096] And a processor for reading the computer program stored in the computer-readable storage medium to implement the method for stitching drone videos based on a Gaussian radiation field provided in the above-mentioned Embodiment 1.
[0097] Embodiment 5:
[0098] A drone video acquisition system includes: a cross-scale drone cluster and the device for stitching drone videos based on a Gaussian radiation field provided in the above-mentioned Embodiment 4;
[0099] The drone cluster includes a main drone and 4 sub-drones; the height of each sub-drone from the ground is the same, and there is an overlap between the fields of view of adjacent sub-drones, so that there is also a certain overlap between the images captured by adjacent sub-drones, ensuring that the three-dimensional scenes of the sub-drones can be stitched; the height of the main drone from the ground is greater than that of the sub-drones, and the field of view of the main drone exactly covers the fields of view of all sub-drones. Specifically, in this embodiment, the 4 sub-drones responsible for capturing detailed scenes are arranged at the 4 corners of a square, with the same height from the ground, all being 55m. In order to ensure that the panoramic image exactly contains all the scene contents of the 4 detailed images, the main drone responsible for capturing the panoramic image is located at the center of the 4 sub-drones and has a height from the ground of 100m.
[0100] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A UAV video stitching method based on Gaussian radiation field, characterized in that: The drone video is collected by a cross-scale drone cluster; the drone cluster includes a master drone and n 2 sub-drones; the heights of the sub-drones above the ground are consistent, and there is overlap between the fields of view of adjacent sub-drones; the height of the main drone above the ground is greater than the heights of the sub-drones above the ground, and the field of view of the main drone just covers the fields of view of all sub-drones; n is a positive integer greater than 1; The drone video stitching method comprises: S1: Use 3DGS to perform 3D reconstruction on the image sequences collected by each sub-UAV, obtain the 3D point set in the field of view of each sub-UAV, and splice them into a panoramic 3D point set R; S2: Select the first image from the image sequence collected by the main drone as the reference image Q, and randomly initialize the view angle C; S3: using the panoramic 3D point set R to render an image I from the current viewing angle C as the rendered image I corresponding to the current reference image Q; S4: Calculate the loss L between the reference image Q and the rendered image I, and optimize the viewing angle C with the goal of minimizing the loss L, and use the rendered image corresponding to the optimized viewing angle C as the rendering result of the current reference image Q; S5: Determine whether all images collected by the master drone have been rendered. If so, proceed to S6; otherwise, select the next image of the current reference image Q from the image sequence collected by the master drone as the new reference image Q, and proceed to S3; S6: The rendering result sequence corresponding to the image sequence collected by the main drone is used as the stitched drone video.
2. The UAV video stitching method based on Gaussian radiation field according to claim 1, characterized in that: In S1, the 3D point sets under the field of view of each drone are stitched into a panoramic 3D point set R, including: S11: Divide all three-dimensional point sets into a plurality of groups, each group contains two three-dimensional point sets with overlapping areas, or contains only one three-dimensional point set, and the number of groups containing only one three-dimensional point set is not greater than 1; S12: For each group containing two three-dimensional point sets, perform the following steps: S121: using the 3D point sets R1 and R2 in the current group to render images image1 and image2 from the perspectives C1 and C2 respectively; there is an overlapping area between the images image1 and image2; S122: extracting feature points from images image1 and image2 respectively, and obtaining a matching relationship between the feature points of the two images; S123: Calculate the geometric transformation matrix between the three-dimensional point sets R1 and R2 according to the feature matching relationship between the images image1 and image2, and map the three-dimensional points in the three-dimensional point set R2 to the coordinate system of the three-dimensional point set R1 according to the geometric transformation matrix to obtain the three-dimensional point set R2'; S124: splicing the three-dimensional point set R1 and the three-dimensional point set R2' into a new three-dimensional point set; when splicing, for the points in the three-dimensional point sets R1 and R2 that are located in the overlapping area and match each other, a weighted fusion method is used for splicing; S13: If the number of three-dimensional point sets after stitching is greater than 1, proceed to S11; otherwise, the three-dimensional point set containing the three-dimensional point sets in the field of view of all sub-drones is used as the panoramic three-dimensional point set R.
3. The UAV video stitching method based on Gaussian radiation field as claimed in claim 2, characterized in that: In S12, if there are multiple groups that need to be spliced, the splicing of the multiple groups is performed in parallel.
4. The UAV video stitching method based on Gaussian radiation field as claimed in claim 2, characterized in that: In S124, for the points in the three-dimensional point sets R1 and R2 that are located in the overlapping area and match each other, when weighted fusion is performed, the weights of the two points are the same.
5. The UAV video stitching method based on Gaussian radiation field as claimed in claim 2, characterized in that: S122 also includes: after obtaining the matching relationship between the feature points of the two images, eliminating the matching relationship between the wrong feature points.
6. The UAV video stitching method based on Gaussian radiation field according to any one of claims 1 to 5, characterized in that: The loss L between the reference image Q and the rendered image I is expressed as follows: L=(1-λ)L com +λL ma Among them, L com represents the comparison loss between the reference image Q and the rendered image I, L ma represents the registration loss between the reference image Q and the rendered image I; λ represents the balance coefficient, λ∈[0,1].
7. A computer program product, characterized in that It includes a computer program; when the computer program is executed by a processor, it implements the drone video stitching method based on Gaussian radiation field as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that: It includes a stored computer program; when the computer program is executed by a processor, it controls the device where the computer-readable storage medium is located to execute the drone video stitching method based on Gaussian radiation field according to any one of claims 1 to 6.
9. A drone video splicing device based on Gaussian radiation field, characterized in that: include: A computer-readable storage medium for storing a computer program; And a processor, used to read the computer program stored in the computer-readable storage medium to implement the drone video stitching method based on Gaussian radiation field as described in any one of claims 1 to 6.
10. A drone video acquisition system, characterized in that: include: A cross-scale drone cluster and a drone video stitching device based on Gaussian radiation field as described in claim 9; The drone cluster includes a main drone and n 2 sub-UAVs; the heights above the ground of each sub-UAV are consistent, and there is overlap between the fields of view of adjacent sub-UAVs; the height above the ground of the main UAV is greater than the heights above the ground of the sub-UAVs, and the field of view of the main UAV just covers the fields of view of all sub-UAVs; n is a positive integer greater than 1.