Aerial sequence image incremental reconstruction method based on three-dimensional Gaussian splash representation
Patent Information
- Application Number
- CN202610798514.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-08-21
AI Technical Summary
[0011]1、目的:本发明的目的在于提供一种基于三维高斯泼溅表示的航拍序列图像增量式重建与新视角合成方法,以解决现有航拍遥感三维重建与新视角合成方法在连续图像输入条件下高斯初始化不稳定、历史区域容易在渐进式优化中退化、大范围场景重建计算和存储开销高的问题,从而实现面向无人机航拍序列图像的高效三维重建和高保真新视角图像合成
[0053]本发明围绕“多视图深度估计引导的三维高斯初始化、基于连续学习的渐进式高斯优化、基于BEV网格的地图动态加载”三个核心步骤形成完整技术方案,具有如下优点:首先,通过多视图深度估计为三维高斯原语提供可靠的几何先验,相比随机初始化或单目深度初始化,能够更好适应航拍图像成像距离远、帧间视差小的特点,减少高斯位置偏差、漂浮伪影和几何错位,提高后续三维重建和新视角合成的几何基础质量;其次,通过基于连续学习的渐进式高斯优化,使系统能够在航拍图像按序输入时持续吸收新观测信息,并利用历史视角回放、深度一致性约束和逐高斯学习率调度保持已重建区域的稳定性,从而缓解历史区域退化和灾难性遗忘问题,提升连续建图过程中新旧区域的整体渲染一致性;通过基于BEV网格的地图动态加载,将大范围三维高斯地图划分为空间网格并进行CPU-GPU动态调度,使GPU显存占用主要取决于当前活动区域而非完整场景规模,从而支持大范围航拍遥感场景的可扩展建图、快速渲染和未采集视角下的高质量新视角合成。
Smart Images

Figure CN122618104A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical fields of remote sensing image 3D reconstruction, UAV photogrammetry, neural rendering, novel perspective synthesis, and 3D scene representation. Specifically, it relates to an incremental reconstruction and novel perspective synthesis method for aerial image sequences based on 3D Gaussian splash representation. This method is applicable to progressive 3D mapping of UAV aerial image sequences, large-scale remote sensing scene representation, and arbitrary perspective image synthesis. It can progressively construct a renderable 3D Gaussian scene representation under the condition of continuous input of aerial images in chronological order, and generate high-quality aerial remote sensing images from unacquired perspectives based on the constructed 3D Gaussian scene representation. Background Technology
[0002] Unmanned aerial vehicle (UAV) remote sensing has been widely applied in fields such as urban 3D modeling, disaster emergency assessment, agricultural and forestry monitoring, infrastructure inspection, geographic mapping, and digital twins. UAV platforms are characterized by flexible deployment, low cost, high resolution, and the ability to plan flight paths according to mission requirements, enabling them to continuously acquire high-resolution aerial images of large-scale terrain scenes. For these continuously acquired aerial images, the key challenge in remote sensing 3D reconstruction and new perspective synthesis lies in how to quickly, stably, and accurately reconstruct the 3D structure of large-scale scenes and support subsequent new perspective image synthesis and scene visualization.
[0003] Traditional 3D reconstruction methods typically rely on motion-based structure reconstruction and multi-view stereo reconstruction techniques. These methods first estimate camera pose and generate sparse point clouds through image feature matching, then utilize multi-view geometry to reconstruct dense depth or mesh models. While traditional methods offer good geometric interpretation in small to medium-scale scenes, they still have several shortcomings in large-scale aerial photography scenarios. First, aerial images are usually acquired sequentially by drones along their flight paths, resulting in a large number of images and a wide scene range. Performing a one-time global reconstruction on all images would incur high computational costs and memory consumption. Second, aerial images are mostly from a top-down or near-top-down perspective, with small inter-frame parallax and a long imaging distance from the ground to the camera, leading to significant uncertainty in depth estimation. Third, remote sensing scenes often contain various terrain features such as buildings, roads, farmland, mountains, and vegetation, with complex texture distributions and large scale spans. Geometric errors, holes, or noise are prone to occur in areas with weak textures, repetitive textures, and occlusions.
[0004] In practical remote sensing applications, the ultimate goal of 3D reconstruction is often not limited to obtaining discrete point clouds or mesh models; acquiring a 3D scene representation capable of interactive rendering is currently a major requirement. For UAV aerial photography missions, due to limitations such as flight altitude, flight path planning, battery life, airspace restrictions, and safe distances, the actual acquired images usually only cover a limited range of perspectives. If high-quality images of unacquired perspectives can be synthesized based on the acquired aerial photography sequence, functions such as virtual aerial photography, observation from any perspective, continuous viewpoint roaming, scene supplementary perspective display, and verification of key areas can be achieved without re-flying and acquiring new images.
[0005] In recent years, 3D Gaussian splashing technology has demonstrated high efficiency and rendering quality in 3D scene representation and novel perspective synthesis. The 3D Gaussian splashing method uses a set of 3D Gaussian primitives to represent the scene. Each Gaussian primitive includes attributes such as spatial position, scale, rotation, opacity, and color, and is rendered efficiently through differentiable rasterization. Compared to implicit volume rendering methods based on neural radiation fields, 3D Gaussian splashing offers advantages such as fast training speed, high rendering efficiency, and strong explicit editability, making it suitable as a fundamental representation for large-scale remote sensing scene reconstruction and novel perspective synthesis.
[0006] However, most existing 3D Gaussian splash reconstruction methods are geared towards offline scenarios, typically assuming that all training images have been acquired before reconstruction and performing global optimization on the complete dataset. This approach is unsuitable for the actual acquisition process of UAV aerial images that are continuously input in chronological order. In applications such as disaster monitoring, emergency mapping, and inspection modeling, the system often needs to gradually update the scene representation as new images arrive during UAV flight, and cannot wait until all data acquisition is complete before processing them uniformly. Furthermore, directly applying existing incremental 3D Gaussian methods to aerial remote sensing scenarios also faces the following technical difficulties.
[0007] First, aerial images are characterized by large-distance imaging and small parallax. Using single-frame depth estimation or random depth initialization can easily lead to inaccurate Gaussian primitive positions, resulting in floating artifacts, geometric noise, and unstable structures. Since 3D Gaussian splashing is highly sensitive to the primitive center position, incorrect initial positions will be amplified in subsequent optimizations, affecting the overall reconstruction quality and the quality of new perspective synthesis.
[0008] Second, aerial image sequences typically have limited perspective changes. If only the current input image is used to optimize newly added Gaussian primitives, the model is prone to overfitting to the current perspective. Furthermore, during the continuous optimization of new images, the Gaussian primitives of the reconstructed regions may be perturbed by subsequent gradients, causing geometric degradation of historical regions and a decline in rendering quality—a catastrophic forgetting problem in incremental learning. Therefore, incremental aerial reconstruction needs to simultaneously consider the adaptability to new regions, the stability of old regions, and rendering consistency under uncaptured perspectives.
[0009] Third, large-scale remote sensing scenes require a large number of 3D Gaussian primitives to represent surface details. As drones continue to fly and maps continue to expand, the number of Gaussian primitives may continue to grow, making it difficult for a single GPU memory block to store a complete global map for a long time. If the Gaussian map is not managed properly, memory usage will increase linearly with the scene size, limiting the application of methods in reconstruction and new perspective rendering in large-scale scenes at the kilometer level.
[0010] To address the aforementioned problems, this invention proposes an incremental reconstruction and novel perspective synthesis method for aerial image sequences based on 3D Gaussian splash representation. This method, tailored to the sequential input characteristics of UAV aerial images, incorporates Gaussian initialization guided by multi-view dense depth estimation, progressive optimization based on continuous learning, and BEV grid map management, thereby achieving efficient, high-fidelity, and scalable 3D reconstruction and novel perspective synthesis of large-scale remote sensing scenes from aerial image sequences. Summary of the Invention
[0011] 1. Purpose: The purpose of this invention is to provide an incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation, in order to solve the problems of unstable Gaussian initialization under continuous image input conditions, easy degradation of historical areas during progressive optimization, and high computational and storage overhead for large-scale scene reconstruction in existing aerial remote sensing three-dimensional reconstruction and new perspective synthesis methods. This will enable efficient three-dimensional reconstruction and high-fidelity new perspective image synthesis for UAV aerial image sequences.
[0012] 2. Technical Solution: The present invention is achieved through the following technical solution: an incremental reconstruction and new perspective synthesis method for aerial image sequence based on three-dimensional Gaussian splash representation includes the following three core steps: three-dimensional Gaussian initialization guided by multi-view depth estimation, progressive Gaussian optimization based on continuous learning, and dynamic map loading based on BEV grid.
[0013] Step 1: 3D Gaussian Initialization Guided by Multi-View Depth Estimation
[0014] The system acquires a sequence of drone aerial images and their camera parameters, including camera intrinsics and camera pose, input sequentially over time. For each currently input aerial image, it determines whether to use it as a keyframe based on its spatial distance from selected keyframes, viewpoint changes, and the coverage of that viewpoint by the current 3D Gaussian map. For the current image selected as a keyframe, it selects neighboring views from historical keyframes that have a co-view relationship with the current image and whose baselines meet the requirements, forming the multi-view depth estimation input. Let the current keyframe be... Camera internal parameters are The camera pose is The selected set of neighboring reference views is Multi-view dense depth estimation is performed based on the current keyframe and its neighboring views to obtain the depth map corresponding to the current keyframe. and uncertainty diagram In multi-view geometry, depth, focal length, baseline, and parallax satisfy the following relationship:
[0015]
[0016] Among them, z Indicates the depth corresponding to a pixel. Indicates the camera's focal length. This indicates the baseline length between the current view and the reference view. This represents parallax. Since aerial images typically have long imaging distances and small inter-frame parallax, selecting neighboring views with sufficient baselines and maintaining co-view relationships can improve the stability and scale consistency of depth estimation.
[0017] For depth-reliable pixels By back-projecting the depth map onto the world coordinate system, three-dimensional points are obtained. Its calculation form is:
[0018]
[0019] in, Represents pixels homogeneous coordinates Represents pixels The corresponding estimated depth, This represents the transformation from the camera coordinate system to the world coordinate system. Only if the uncertainty diagram satisfies... Only then, the 3D points obtained by backprojecting the pixel are used to initialize the 3D Gaussian primitive, where This represents the uncertainty threshold.
[0020] This invention uses three-dimensional Gaussian primitives to represent scenes. A three-dimensional Gaussian map is denoted as:
[0021]
[0022] in, Indicates the first A three-dimensional Gaussian primitive, Indicates the location of the center of Gauss. Represents the covariance matrix or the anisotropic shape determined by scale and rotation parameters. Indicates opacity. This represents the color or spherical harmonic color coefficient. For a 3D point obtained from reliable depth backprojection, it is used as the center position of a new Gaussian primitive, and its scale, rotation, opacity, and color attributes are initialized according to the neighborhood point spacing, image color, and initial scale rules. During subsequent keyframe inputs, the system first renders the current 3D Gaussian map from the perspective of that keyframe to obtain a rendered color map. Rendering depth map and rendering opacity map For pixels Its rendered color and opacity can be represented as:
[0023]
[0024]
[0025] in, Indicates projection onto pixels Gaussian set, Indicates the first The contribution of each Gaussian to the opacity of that pixel. Indicates its color, This represents the transmittance accumulated from the preceding Gaussian sequence. Based on the current map rendering results and multi-view depth estimation results, new observation areas are determined, which are then used to initialize the Gaussian sequence for these new areas.
[0026] Step 2: Progressive Gaussian Optimization Based on Continuous Learning
[0027] After initializing the newly added 3D Gaussian primitives, the optimization process of the 3D Gaussian map under the aerial image sequence is modeled as a continuous learning process. For the current input image, the system uses the 3D Gaussian map to render a synthetic image and depth map from the current viewpoint, and optimizes the Gaussian parameters by jointly constraining the depth through real images, multi-view depth estimation, and historical view playback.
[0028] Let the current 3D Gaussian map be The current input image is Its rendered image is Rendering depth is The depth of the multi-view estimation is The image reconstruction loss from the current perspective includes color consistency loss and structural similarity loss, which can be expressed as:
[0029]
[0030] in, This represents the pixel-level absolute error between the real image and the rendered image. Represents structural similarity loss. where represents the weighting coefficients. To enhance the geometric stability of the 3D Gaussian map, a depth consistency loss is introduced between the rendered depth and the estimated depth:
[0031]
[0032] Therefore, the reconstruction loss from the current perspective can be expressed as:
[0033]
[0034] in, and These represent the weights of the color reconstruction loss and the depth consistency loss, respectively. To avoid destroying historical regions when continuously inputting new aerial images, this invention sets up a historical view playback buffer. This buffer stores processed keyframes or observation frames. When optimizing the current image, a subset of historical viewpoints is selected from the buffer for training. Unlike random replay, this invention determines the replay weights based on the reconstruction loss of historical viewpoints, allowing viewpoints with higher historical losses to receive higher sampling probabilities. For historical viewpoints... Its sampling weight can be expressed as:
[0035]
[0036] in, This represents the set of historical perspectives sampled from the historical buffer. Therefore, the total loss of the asymptotic Gaussian optimization can be expressed as:
[0037]
[0038] in, This approach balances adaptation to the current input image with preservation of historical regions. Through this continuous learning optimization objective, the 3D Gaussian map can absorb new observation information while avoiding significant geometric degradation and rendering quality decline in historical regions.
[0039] Step 3: Dynamic Map Loading and New Perspective Synthesis Based on BEV Grid
[0040] As drone aerial photography sequences are continuously input, the size of the 3D Gaussian map continues to grow. If the complete global 3D Gaussian map is always stored in GPU memory, memory usage will increase rapidly as the scene expands. Considering that aerial photography scenes primarily unfold horizontally along the ground surface and cameras mostly observe from a top-down or tilted top-down perspective, this invention employs a dynamic map loading method based on BEV (Bird's-Eye View, BEV) grids to manage the global 3D Gaussian map.
[0041] Let the global 3D Gaussian map be... A two-dimensional BEV mesh is set on the horizontal plane of the world coordinate system. For any Gaussian primitive... According to its central position horizontal coordinates Assign it to the corresponding BEV mesh cell:
[0042]
[0043] in, Gaussian primitives The corresponding BEV grid cell index, and This represents the reference boundary of the global map on the horizontal plane. This indicates the size of the mesh cell. Each BEV mesh cell stores the 3D Gaussian primitive index and attribute data within its spatial range. When processing current real aerial images or virtual viewpoints to be synthesized, the coverage area of the viewpoint on the BEV plane is calculated based on the camera pose, camera intrinsics, and view frustum, and the set of active meshes to be loaded is determined.
[0044]
[0045] in, Indicates the camera pose and camera internal reference A defined three-dimensional view frustum, This indicates that the 3D view frustum is projected onto the BEV plane. Indicates the first Each BEV grid cell. The system stores the complete 3D Gaussian map in CPU memory, only storing the active grid set. The corresponding local Gaussian primitives are loaded into the GPU to obtain the currently active Gaussian set:
[0046]
[0047] in, Represents BEV mesh element The system stores a set of Gaussian primitives. For mesh cells no longer within the current active area, their Gaussian data is unloaded from the GPU and written back to the CPU; for mesh cells newly entering the field of view or virtual view coverage, they are loaded from the CPU to the GPU. In this way, GPU memory usage is primarily determined by the size of the currently active region, rather than by the overall scene size.
[0048] During new perspective compositing, the user specifies the virtual camera pose. and camera internal reference The system according to The active mesh set is calculated and the corresponding Gaussian primitives are loaded. Then, a virtual viewpoint image is generated through 3D Gaussian splash rasterization. For pixels... Its composite color is:
[0049]
[0050] in, This represents the projection onto pixels in a virtual viewpoint. Gaussian primitive set Indicates transmittance. Indicates contribution to opacity. This indicates the color. If the user specifies a continuous virtual aerial photography path, the system updates the active BEV mesh set frame by frame along the path and reuses the Gaussian data already loaded in adjacent viewpoints, thereby achieving continuous new viewpoint rendering in a large-scale scene.
[0051] Through the above three core steps, this invention can progressively construct, optimize and manage a renderable 3D Gaussian scene representation during the continuous input of UAV aerial images. While reducing the computation and memory overhead of large-scale scenes, it can achieve high-quality incremental 3D reconstruction and synthesis of aerial remote sensing images from unacquired perspectives.
[0052] 3. Advantages and effects:
[0053] This invention forms a complete technical solution around three core steps: "3D Gaussian initialization guided by multi-view depth estimation, progressive Gaussian optimization based on continuous learning, and dynamic map loading based on BEV grids." It has the following advantages: First, multi-view depth estimation provides reliable geometric priors for 3D Gaussian primitives. Compared to random initialization or monocular depth initialization, it better adapts to the characteristics of long imaging distances and small inter-frame parallax in aerial images, reducing Gaussian positional bias, floating artifacts, and geometric misalignment, thus improving the geometric foundation quality of subsequent 3D reconstruction and new perspective synthesis. Second, through progressive Gaussian optimization based on continuous learning, the system can dynamically load maps in aerial... As images are input sequentially, new observation information is continuously absorbed. Historical view playback, depth consistency constraints, and Gaussian learning rate scheduling are used to maintain the stability of reconstructed areas, thereby mitigating the problems of historical area degradation and catastrophic forgetting, and improving the overall rendering consistency of old and new areas during continuous mapping. Through dynamic map loading based on BEV grids, large-scale 3D Gaussian maps are divided into spatial grids and CPU-GPU dynamic scheduling is performed, so that GPU memory usage depends mainly on the current active area rather than the scale of the entire scene. This supports scalable mapping, fast rendering, and high-quality synthesis of new perspectives from unacquired viewpoints for large-scale aerial remote sensing scenes. Attached Figure Description
[0054] Figure 1 This is an incremental reconstruction algorithm framework based on 3D Gaussian splash representation.
[0055] Figure 2 This is a schematic diagram of 3D Gaussian initialization based on multi-view depth guidance.
[0056] Figure 3 This is a schematic diagram of progressive 3D Gaussian optimization based on continuous learning.
[0057] Figure 4 This is a composite image showing the new perspective effect of the present invention in multiple aerial photography scenarios. Detailed Implementation
[0058] The specific embodiments of the present invention will now be described with reference to the accompanying drawings. It should be understood that the following embodiments are only used to further illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the parameters, module implementation methods, or data processing flows in each step without departing from the core idea of the present invention should be included within the scope of protection of the present invention.
[0059] This invention provides an incremental reconstruction and new perspective synthesis method for aerial image sequences based on 3D Gaussian splash representation. The method uses a sequence of aerial images acquired sequentially by a UAV as input, and combines camera intrinsic parameters, camera extrinsic parameters, or camera pose information obtained through photogrammetry to progressively construct a renderable 3D Gaussian scene map. The method mainly includes three core processes: 3D Gaussian initialization guided by multi-view depth estimation, progressive Gaussian optimization based on continuous learning, and dynamic map loading based on BEV grids. Through these processes, this invention can continuously expand the 3D scene with continuous input of aerial images and supports the synthesis of new perspectives from previously unacquired viewpoints.
[0060] Step 1: Initialize the 3D Gaussian based on multi-view depth estimation.
[0061] As attached Figure 1 and attached Figure 2 As shown, the first step is to obtain the drone aerial image sequence and its corresponding camera parameters. Let the input aerial image sequence be: ,in, Indicates the first Frame-by-frame aerial images, each frame corresponding to camera intrinsic parameters and camera pose First, keyframes are selected from the input image sequence. An image is designated as a keyframe if the spatial distance, viewpoint change, or proportion of a newly added observation area between the current image and the nearest keyframe meets preset conditions. For the current keyframe... From the historical keyframe set, select neighboring views that have a co-view relationship and whose baselines meet the requirements to form a reference view set. In aerial photography, the camera is far from the ground, and the parallax between adjacent frames is small, making depth estimation prone to instability. In multi-view geometry, depth... ,focal length baseline and parallax satisfy:
[0062]
[0063] This embodiment improves the reliability of depth estimation under low parallax conditions in aerial photography by selecting neighboring views with sufficient baseline and maintaining common viewing relationships.
[0064] Subsequently, based on the current keyframe and reference view collection Perform multi-view dense depth estimation to obtain the depth map of the current keyframe. and uncertainty diagram The depth map provides the geometric initialization location for the 3D Gaussian primitives, while the uncertainty map filters out regions with unreliable depth. For pixels... Its homogeneous coordinates are: When the pixel satisfies At that time, its depth was considered reliable, and it was back-projected onto the world coordinate system based on the camera intrinsic parameters and camera pose:
[0065]
[0066] in, This represents the three-dimensional points obtained by back projection. This represents the uncertainty threshold.
[0067] When inputting subsequent keyframes, instead of re-initializing the entire scene, the current 3D Gaussian map is first rendered from the perspective of that keyframe, resulting in a rendered color map, a rendered depth map, and a rendered opacity map. For pixels... If the current map does not cover the area at that pixel sufficiently, or if the rendered depth differs significantly from the multi-view estimated depth, then it is identified as a newly added observation area.
[0068]
[0069] in, Indicates rendering opacity. Indicates rendering depth. Opacity threshold This is the depth difference threshold. For conditions satisfying... and The pixels are then used to perform backprojection and generate new 3D Gaussian primitives.
[0070] Step 2: Progressively optimize Gaussian by replaying continuously learned data.
[0071] As attached Figure 3 As shown, after initializing the newly added 3D Gaussian primitives, the 3D Gaussian map is progressively optimized. This embodiment models the 3D Gaussian map optimization process as a continuous learning process, enabling the model to maintain historical regional stability while absorbing new observation information.
[0072] Let the current 3D Gaussian map be The current input image is The image rendered from the 3D Gaussian map is Rendering depth is The depth of the multi-view estimation is The image reconstruction loss from the current perspective includes color error and structural similarity error:
[0073]
[0074] in, These are the weighting coefficients. This represents the structural similarity loss.
[0075] Based on the rendered depth map and the multi-view estimated depth map, calculate the depth consistency loss:
[0076]
[0077] The reconstruction loss from the current perspective is:
[0078]
[0079] in, and These are the weights for image reconstruction loss and depth consistency loss, respectively.
[0080] Based on the set historical view replay buffer This buffer is used to store processed keyframes or observation frames. When optimizing the current input image, a subset of historical viewpoints is selected from the historical buffer to participate in training. (Regarding historical viewpoints...) The replay weight is determined based on its historical reconstruction losses:
[0081]
[0082] in, Representing a historical perspective The reconstruction loss. Historical perspectives with larger losses typically correspond to regions that are difficult to preserve or prone to degradation, and therefore have a higher playback probability. The preservation loss for historical perspective playback is expressed as:
[0083]
[0084] in, This represents the set of historical perspectives sampled from the historical buffer. Finally, the total loss of the asymptotic Gaussian optimization is:
[0085]
[0086] in, This indicates the weight of historical losses.
[0087] For each newly added image sequence, the Gaussian distribution in the scene is iteratively optimized following the above process. The viewpoints of historical frames are selected from the playback viewpoint set according to the reconstruction loss to avoid local overfitting. Based on this process, the image sequence is processed progressively.
[0088] Step 3: Dynamic Map Loading and New Perspective Synthesis for Building the BEV Mesh
[0089] As attached Figure 1As shown, with the continuous input of drone aerial photography sequences, the size of the 3D Gaussian map continues to grow. This embodiment adopts a dynamic map loading mechanism based on the BEV grid. Aerial remote sensing scenes are usually unfolded horizontally along the ground surface, and the camera is mostly observing from above or at an angle. Based on this characteristic, a BEV grid is established on the horizontal plane of the world coordinate system, and the 3D Gaussian primitives are organized according to their horizontal positions. Let any Gaussian primitive... The central location is: Based on its horizontal coordinates Assign it to the corresponding BEV mesh cell:
[0090]
[0091] Gaussian primitives The corresponding BEV grid cell index, and The reference boundary representing the horizontal extent of the global map. This indicates the size of the mesh cell. Each BEV mesh cell stores the 3D Gaussian primitive indexes and attribute data within its range. The complete 3D Gaussian map is stored in CPU memory, and only the active local Gaussian required for current optimization or rendering is loaded to the GPU. For the current real aerial viewpoint or the virtual viewpoint to be synthesized, the BEV mesh set covered by the current viewpoint is calculated based on the camera pose T and camera intrinsic parameters K, and the corresponding 3D Gaussian primitives are loaded from the CPU global map to the GPU to form the current active Gaussian set:
[0092]
[0093] in, This represents the set of active BEV meshes covered by the current viewpoint. Represents grid cells The set of Gaussian primitives stored in the middle.
[0094] For grid cells that are no longer within the current active range, their updated Gaussian data is written back from the GPU to the CPU and the video memory is released; for grid cells that newly enter the field of view or the virtual view coverage area, they are loaded from the CPU to the GPU.
[0095] When performing new perspective compositing, the user specifies the virtual camera pose. and camera internal reference The system according to The active BEV mesh set is calculated and the corresponding Gaussian primitives are loaded. Then, a virtual viewpoint image is generated using 3D Gaussian sputtering rasterization. For pixels in the virtual viewpoint... Its composite color is:
[0096]
[0097] This represents the projection onto pixels in a virtual viewpoint. Gaussian primitive set Indicates transmittance. Indicates the Gaussian opacity contribution. This represents Gaussian color.
[0098] When a user specifies a continuous virtual aerial photography path, the system updates the virtual camera pose sequentially along the path and reuses the BEV mesh data already loaded in adjacent viewpoints. Since there is usually a large overlap between continuous viewpoints, the system only needs to dynamically load newly visible meshes and unload meshes far away from the current view frustum to achieve continuous new viewpoint rendering in a large-scale scene.
[0099] In summary, this embodiment improves the reliability of Gaussian geometry initialization in low parallax aerial photography scenarios by using multi-view depth estimation-guided 3D Gaussian initialization; it alleviates the historical area degradation problem in the continuous mapping process by using progressive Gaussian optimization based on continuous learning; and it achieves scalable management and efficient new perspective synthesis of 3D Gaussian maps in large-scale scenarios by using dynamic map loading based on BEV grids.
[0100] Experimental Results: To verify the reconstruction and new perspective synthesis effects of the method of this invention in aerial remote sensing scenes, several typical UAV aerial photography scenes were selected for experiments, including different types of areas such as factories, villages, farmland, parking lots, towns, and mountains. In the experiments, aerial images collected by UAVs in sequence were input into the method of this invention. Multi-view depth estimation guided 3D Gaussian initialization, and progressive Gaussian optimization was performed during continuous image input to finally construct a 3D Gaussian scene representation that can be used for rendering. Reconstruction was performed based on sequential input images of various scenes, and one image was selected at 20 equal intervals as the ground truth for new perspective synthesis for evaluation. The evaluation metrics used were Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). Higher PSNR and SSIM indicate more realistic synthesis effects, while lower LPIPS indicates better synthesized image quality. The results of the quantitative experiments are shown in the table below.
[0101] method PSNR SSIM LPIPS MonoGS 18.57 0.476 0.621 VINGS-Mono 21.48 0.681 0.376 This invention 22.93 0.720 0.317
[0102] Experimental results show that, compared with MonoGS, the new perspective synthesis algorithm of VINGS-Mono significantly improves the quality of new perspective synthesized images in remote sensing scenes, verifying the effectiveness of the design of each part of the present invention.
[0103] To qualitatively evaluate the quality of the new perspective synthesis, attached Figure 4 The invention demonstrates novel perspective rendering results in multiple aerial photography scenarios. (See attached image.) Figure 4 As can be seen, this invention can generate clear, continuous, and relatively detailed synthetic images under different land cover types and scene scales, demonstrating good expressive ability for contours, textures, and geometry. Simultaneously, it maintains good geometric and appearance consistency even under uncaptured viewpoints, reducing floating artifacts, local blurring, and structural misalignment. Experimental results show that this invention can not only achieve incremental 3D reconstruction from aerial image sequences but also generate high-quality new perspective images based on the constructed 3D Gaussian map, making it suitable for applications such as virtual aerial photography and digital twin display in large-scale remote sensing scenes.
Claims
1. A method for incremental reconstruction and new perspective synthesis of aerial image sequences based on three-dimensional Gaussian splash representation, characterized in that, Includes the following steps: Step 1: 3D Gaussian Initialization Guided by Multi-View Depth Estimation The system acquires a sequence of drone aerial images input in chronological order and their corresponding camera parameters, including camera intrinsics and camera pose. Based on the spatial distance between the current input aerial image and selected keyframes, the viewpoint change, and the coverage of the current 3D Gaussian map over the current viewpoint, it determines whether to identify the current input aerial image as a keyframe. For the current image identified as a keyframe, it selects neighboring views from historical keyframes that have a co-view relationship with the current image and whose baselines meet the requirements, thus forming the multi-view depth estimation input. Perform multi-view dense depth estimation based on the current keyframe and its neighboring views to obtain the depth map and uncertainty map corresponding to the current keyframe; filter depth-reliable pixels according to the depth map and uncertainty map, and back-project the depth-reliable pixels to the world coordinate system and initialize them as three-dimensional Gaussian primitives; Step 2: Progressive Gaussian Optimization Based on Continuous Learning After initializing the newly added 3D Gaussian primitives, the optimization process of the 3D Gaussian map under the aerial image sequence is modeled as a continuous learning process. The current 3D Gaussian map is rendered using the current input image to obtain the rendered image and rendered depth map from the current viewpoint. The current viewpoint reconstruction loss is constructed based on the real aerial image, the rendered image, the multi-view estimated depth map, and the rendered depth map. At the same time, a historical viewpoint replay buffer is set up, and historical views are selected from the historical viewpoint replay buffer to participate in training. The corresponding replay weights are determined based on the reconstruction loss of the historical views to construct the history preservation loss. The parameters of the 3D Gaussian primitives in the 3D Gaussian map are progressively optimized by combining the current viewpoint reconstruction loss and the history preservation loss. Step 3: Dynamic Map Loading and New Perspective Synthesis Based on BEV Grid A BEV mesh is established on the horizontal plane of the world coordinate system. 3D Gaussian primitives from the global 3D Gaussian map are assigned to corresponding BEV mesh cells according to their center horizontal coordinates. The complete 3D Gaussian map is stored in CPU memory. The currently active BEV mesh set is determined based on the camera intrinsics, camera pose, and frustum of the current real aerial viewpoint or the virtual viewpoint to be synthesized. Only the 3D Gaussian primitives corresponding to the currently active BEV mesh set are loaded into GPU memory for rendering or optimization. For BEV mesh cells no longer within the current active range, their updated 3D Gaussian primitive data is written back from GPU memory to CPU memory and the memory is released. For BEV mesh cells newly entering the current active range, their corresponding 3D Gaussian primitives are loaded from CPU memory into GPU memory. When synthesizing a new viewpoint, the corresponding active BEV mesh set is loaded according to the user-specified virtual camera pose and camera intrinsics, and a virtual viewpoint image is generated through 3D Gaussian splash rasterization.
2. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: In step one, when selecting a neighboring view from historical keyframes, the neighboring view simultaneously satisfies the following conditions: it shares a common viewing area with the current keyframe and the camera baseline length between it and the current keyframe is greater than a preset baseline threshold. The depth, focal length, baseline, and parallax satisfy the following relationship: in, Indicates the depth corresponding to a pixel. Indicates the camera's focal length. This indicates the baseline length between the current view and the reference view. Indicates parallax.
3. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: In step one, after performing multi-view dense depth estimation based on the current keyframe and its neighboring views, the depth map corresponding to the current keyframe is obtained. and uncertainty diagram For the pixels in the current keyframe When its uncertainty satisfies When this happens, the pixel is determined to be a depth-reliable pixel, where, The uncertainty threshold is used; based on the pixel coordinates, depth value, camera intrinsic parameters, and camera pose of the reliable depth pixel, it is back-projected to the world coordinate system to obtain a 3D point: in, This represents a 3D point in the world coordinate system obtained by back projection. This indicates the camera pose of the current keyframe. Indicates camera intrinsic parameters. Represents pixels homogeneous coordinates Represents pixels The corresponding depth value; the three-dimensional point As the central location of the newly added three-dimensional Gaussian primitive, the newly added three-dimensional Gaussian primitive is represented as: in, , This indicates the center position of the newly added three-dimensional Gaussian primitive. This represents the covariance matrix determined by the scaling and rotation parameters. Indicates opacity. The color or spherical harmonic color coefficient is represented; the scale parameter is initialized based on the neighborhood point spacing of the back-projected 3D point; the rotation parameter is initialized using unit rotation or based on the local neighborhood geometric direction; the opacity is initialized using a preset initial value; and the color or spherical harmonic color coefficient is based on the pixels in the current keyframe. The color values are initialized, thus completing the initialization of the 3D Gaussian primitives based on the multi-view depth estimation results.
4. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: In step one, during subsequent keyframe input, the current 3D Gaussian map is first rendered from the current keyframe perspective to obtain a rendered color map, a rendered depth map, and a rendered opacity map. Based on the rendered depth map, rendered opacity map, and the depth map obtained from multi-view estimation, the newly added observation region in the current keyframe is determined, and 3D Gaussian primitive initialization is performed only within the newly added observation region. The newly added observation region is masked by a new region mask. It is confirmed that the newly added region mask satisfies the following relationship: in, Indicates pixel position, This indicates the current 3D Gaussian map in pixels. The rendering opacity at that location This indicates the current 3D Gaussian map in pixels. Rendering depth at that location This represents the depth value obtained from multi-view estimation. Indicates the opacity threshold. Indicates the depth difference threshold. Represents a logical OR operation. For simultaneously satisfying the masking requirements of the new area and uncertainty constraints For pixels that do not satisfy the newly added region mask or uncertainty constraints, perform back projection and generate new 3D Gaussian primitives; for pixels that do not satisfy the newly added region mask or uncertainty constraints, do not generate new 3D Gaussian primitives.
5. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: In step two, the image reconstruction loss at the current viewpoint includes color consistency loss and structural similarity loss, and the image reconstruction loss is expressed as: in, This indicates the currently input aerial image. This represents the image rendered from the current 3D Gaussian map. Indicates pixel-level absolute error. Represents structural similarity loss. This represents the weight of the structural similarity loss. The depth consistency loss is calculated based on the rendered depth map and the multi-view estimated depth map. The depth consistency loss is expressed as: in, This represents the depth map obtained from multi-view estimation. This represents the depth map rendered from the current 3D Gaussian map; the reconstruction loss from the current viewpoint is represented as: in, and These represent the weights of the image reconstruction loss and the depth consistency loss, respectively.
6. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: In step two, the historical view replay buffer is used to store processed keyframes or observation frames. When optimizing the current input aerial image, some historical views are selected from the historical view replay buffer to participate in the training together, and their replay weights are determined according to the historical reconstruction loss corresponding to each historical view, so that historical views with larger historical reconstruction losses have higher replay probabilities. Historical perspective Replay weight Represented as: in, Representing a historical perspective The reconstruction losses, This represents the historical view replay buffer or the set of candidate historical views selected from the historical view replay buffer. Representing a historical perspective The reconstruction losses. The historical retention loss based on historical perspective playback is expressed as: in, This represents the set of historical perspectives sampled from the historical perspective playback buffer. Representing a historical perspective Replay weight, Representing a historical perspective The corresponding reconstruction loss. The total loss of the asymptotic Gaussian optimization is expressed as:
7. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: In step three, the BEV mesh is set on the horizontal plane of the world coordinate system. For any three-dimensional Gaussian primitive... According to its central position horizontal coordinates It is then assigned to the corresponding BEV mesh cell.
8. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: In step three, as the drone aerial image sequence is continuously input, the current active BEV mesh set is dynamically updated as the camera view changes. For a new BEV mesh cell entering the current active BEV mesh set, its corresponding 3D Gaussian primitive is loaded from CPU memory to GPU memory. For a BEV mesh cell leaving the current active BEV mesh set, its updated 3D Gaussian primitive is written back to CPU memory and released from GPU memory.
9. The incremental reconstruction and new perspective synthesis method for aerial image sequences based on three-dimensional Gaussian splash representation according to claim 1, characterized in that: When performing new perspective synthesis, the user specifies the virtual camera pose and camera intrinsic parameters. The system determines the active BEV mesh set corresponding to the virtual perspective based on the virtual camera pose and camera intrinsic parameters, loads the 3D Gaussian primitive corresponding to the active BEV mesh set, and generates the virtual perspective image through 3D Gaussian splash rasterization.