True orthophoto generation method based on three-dimensional Gaussian model, storage medium and equipment

Through an improved true projection image generation method based on the three-dimensional Gaussian model algorithm, the problems of low image quality, slow processing efficiency and loss of detailed features in the prior art are solved, and real-time acquisition of high-quality true projection images in large-scale scenarios is achieved.

CN120163933APending Publication Date: 2025-06-17WUHAN TIANYUANSHI TECH
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510200418.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-17

Smart Images

  • Figure CN120163933A_ABST
    Figure CN120163933A_ABST
Patent Text Reader

Abstract

The invention relates to a true orthophoto generation method based on a three-dimensional Gaussian model, a storage medium and equipment, and aims to solve the problems of low image quality, low generation speed, detail loss and the like in the prior art. The method comprises the following steps: firstly, acquiring image data through an unmanned aerial vehicle or aerial photography, and extracting a camera attitude and sparse point cloud by using a motion recovery structure (SfM) technology; secondly, constructing a three-dimensional Gaussian model based on the sparse point cloud, and performing iterative training through top view orthographic projection in combination with a gradient descent algorithm; in the training process, a densification strategy, a point deleting strategy, a blocking strategy and an image pyramid strategy are innovatively introduced, so that the detail expressive force and the overall quality of the image are remarkably improved. Specifically, according to the densification strategy, fine detail reconstruction is achieved by dynamically increasing Gaussian ball density, and according to the point deletion strategy, rendering efficiency is improved and computing resource allocation is optimized by eliminating redundant Gaussian balls. The image pyramid strategy generates a multi-level visual effect through multi-scale training, and the blocking strategy improves the reconstruction precision through local optimization. Finally, on the basis of the trained three-dimensional Gaussian model, real-time generation of a high-quality true orthophoto in a large-scale scene can be realized. The method has remarkable advantages in the aspects of efficiency, precision and practicability, provides important technical support for the fields of geographic information systems, urban planning, disaster monitoring and the like, and has wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of surveying and mapping remote sensing, and particularly relates to a true orthophoto image generation method, a storage medium and a device based on a three-dimensional Gaussian model algorithm. Background Art

[0002] A digital true orthophoto map (TDOM) is a remotely sensed image product that has been precisely geometrically corrected and has both the geometric accuracy of a map and the texture characteristics of an image. This image can accurately reflect the topographic and geomorphic features of the earth's surface and support precise spatial measurement and analysis, and has important application values in fields such as urban planning, land resource management, environmental monitoring and emergency response.

[0003] Currently, the methods for generating digital orthophoto images can be mainly divided into two categories: The first category is the differential correction and stitching method based on images, which geometrically corrects and radiometrically corrects the original images, and then realizes the stitching and fusion of multiple images; The second category is the orthographic projection method based on three-dimensional scene modeling, which first constructs a three-dimensional model of the scene and then generates an orthophoto image through top-down orthographic projection. These two methods each have their own technical characteristics and application scenarios.

[0004] The process of generating digital orthophoto images inevitably involves the image stitching link. Currently, the method based on image stitching has formed a relatively complete technical system, and significant progress has been made especially in the optimal seam line selection algorithm. However, there are still several technical bottlenecks in this kind of method: First, it is difficult to completely eliminate the geometric misalignment problem between images; Second, the color transition in the stitching area is often not natural enough; Finally, the stitching effect of linear features such as building edges and roads is still not ideal.

[0005] The orthophoto image generation method based on three-dimensional reconstruction adopts a completely different technical route, which obtains a true orthophoto image by top-down orthographic projection of the reconstructed three-dimensional scene. Although this method avoids the complex image stitching process, its output quality highly depends on the accuracy of three-dimensional reconstruction. Traditional three-dimensional reconstruction techniques usually use the Structure from Motion (SfM) algorithm to obtain the camera pose and generate a sparse point cloud, and then combine it with the Multi-View Stereo (MVS) algorithm for dense reconstruction, and finally construct a fine three-dimensional model. However, this kind of method still faces many challenges in terms of real-time performance, computational resource consumption and adaptability to complex scenes.

[0006] In recent years, Neural Radiance Fields (NeRF) has emerged as a revolutionary 3D reconstruction technology, breaking through the limitations of traditional methods. NeRF does not rely on explicit camera models and geometric constraints. Instead, it directly learns the 3D representation of a scene from image features through a neural network, effectively eliminating the errors caused by photographic tilt angles. However, NeRF still has obvious deficiencies: firstly, its training process consumes a large amount of computing resources and time; secondly, there are efficiency bottlenecks in dealing with large-scale scenes; finally, the reconstruction effect under complex lighting conditions still needs to be improved.

[0007] In summary, how to achieve the real-time acquisition of high-quality orthophotos in large-scale scenes remains a key scientific problem that urgently needs to be solved. Summary of the Invention

[0008] The present invention aims to solve the technical problems existing in the existing orthophoto generation methods, such as low image quality, slow processing efficiency, and loss of detail features, and proposes an improved orthophoto generation method based on the 3D Gaussian Splatting (3DGS) model. The method specifically includes the following steps:

[0009] S1. Data Acquisition and Preprocessing

[0010] Obtain UAV or aerial photography images as input data, and divide the dataset into a training set, a validation set, and a test set according to a preset ratio to ensure the rationality of data distribution and the reliability of model training.

[0011] S2. Sparse Point Cloud Generation and Gaussian Point Cloud Initialization

[0012] Use the Structure from Motion (SfM) algorithm to convert the input images into a sparse point cloud. Specifically, SfM reconstructs the camera motion trajectory and estimates the 3D spatial positions of scene feature points by analyzing the feature point matching relationships between multi-view images. The present invention directly calls the SfM functional module in the COLMAP library, which provides a complete SfM and Multi-View Stereo (MVS) processing pipeline, enabling efficient and accurate sparse point cloud reconstruction.

[0013] After obtaining the sparse point cloud, the 3DGS algorithm initializes each point as a 3D Gaussian ellipsoid. Each Gaussian point contains the following key attribute parameters:

[0014] ● Position Parameters: Define the center position of the 3D Gaussian ellipsoid, determined by the mean vector of the Gaussian distribution. Its initial value is directly inherited from the spatial coordinates of the sparse point cloud, providing an accurate initial spatial positioning for subsequent optimization.

[0015] ● Covariance Matrix: Controls the spatial geometric characteristics of the three-dimensional Gaussian distribution. It defines the scale and orientation of Gaussian points in space through a 3×3 symmetric positive definite matrix. Specifically, its eigenvalues determine the stretching and shrinking degrees of Gaussian points in the three principal axis directions, and the eigenvectors determine the principal axis directions of the Gaussian distribution, thus achieving precise control over anisotropic geometric forms.

[0016] ● Opacity: Characterizes the visibility of Gaussian points, with a value range of

[0017] [0, 1]. This parameter not only controls the visible and invisible states of individual Gaussian points but also realizes the transparency effect and level-of-detail performance of complex scenes through the superposition of the opacities of multiple Gaussian points. It is one of the key factors affecting rendering quality.

[0018] ● Spherical Harmonics: Used to describe the view-dependent color characteristics of Gaussian points. Low-order spherical harmonic coefficients (usually of the third order) are used to encode the color changes of Gaussian points under different viewpoints, effectively controlling the storage overhead while ensuring rendering quality. This function supports color expression under complex lighting conditions,

[0019] and can accurately reproduce the lighting effects and material properties of the scene.

[0020] In the initialization stage, the covariance matrix, opacity, and spherical harmonics are all given preset default values, and these parameters will be gradually adjusted in the subsequent optimization process.

[0021] S3. Three-Dimensional to Two-Dimensional Projection Transformation

[0022] To achieve multi-view image rendering, through the view transformation matrix W i and the Jacobian matrix J i , the 3D Gaussian point cloud is projected from the world coordinate system to the 2D image plane. This projection process ensures an accurate mapping from the three-dimensional scene to the two-dimensional image.

[0023] However, the original 3DGS rendering only supports perspective projection, while our goal is to generate true orthophotos. Therefore, we need to make corresponding improvements to the rendering module of 3DGS to support orthographic projection, and this change is an important innovation of this invention.

[0024] In orthographic projection, the camera viewport is uniquely determined by the near plane and far plane n, f, and the left and right boundaries l, r and top and bottom boundaries t, b within the near plane. The corresponding orthographic projection matrix can be expressed as:

[0025]

[0026] In orthographic projection, since there is no perspective relationship where objects appear larger when closer and smaller when farther away, all light rays are considered parallel, and the camera is assumed to be located infinitely far behind the projection plane. Therefore, the transformation from the camera coordinate system to the ray coordinate system can be achieved as follows:

[0027] t = φ(t) = Pt,

[0028] where P = I3 is the identity matrix, so no first-order Jacobian approximation is required. The transformation from the world coordinate system to the camera coordinate system can be expressed as where W is the transformation matrix from world coordinates to camera coordinates, and d is the translation vector. Therefore, the transformation from the world coordinate system to the ray coordinate system can be expressed as:

[0029]

[0030] According to the properties of the Gaussian distribution, the 3D Gaussian G k After being projected onto the 2D screen space, its covariance matrix becomes:

[0031] ∑′ k = W∑ k W T ,

[0032] Under this transformation, it can be equivalently considered that the corresponding Jacobian matrix is the third-order identity matrix I3, and no further first-order Jacobian approximation is required. In addition, in orthographic projection, the focal length of the camera can be regarded as the ratio of the image width and height to the viewport size, that is:

[0033]

[0034] Therefore, we can render orthographic views that meet different resolution requirements by adjusting the focal length of the orthographic camera. However, when the number of Gaussian basis elements in the viewport is too large, it may cause memory overflow and the rendering cannot be completed. For high-precision and large-scale Gaussian reconstruction models, it is not feasible to render the entire scene picture at once. For this reason, we adopt a sliding scan strategy: use a small viewport to scan the entire scene in segments, generate corresponding small-sized pictures, and finally stitch these pictures together in sequence to form a complete orthographic view. The specific process is as Figure 6 shown. This method not only effectively solves the memory limitation problem but also ensures the integrity and accuracy of large-scale scene reconstruction.

[0035] S4. Rendering Process

[0036] To calculate the color of each pixel, 3DGS adopts a classic method based on Neural Point-Based Rendering. Specifically, a ray r is projected from the camera center, and its interaction with 3D Gaussian distributions is calculated along this ray to aggregate color and density information. When rendering at a new viewpoint, the 3D Gaussians are projected onto the 2D plane, and the final color value is calculated along the given ray r. The color calculation formula for this ray is as follows:

[0037]

[0038] where N represents the number of sampling points on ray r, c n represents the color and opacity of the nth Gaussian, and

[0039]

[0040] Here, α n represents the opacity of the nth Gaussian, represents the projection of the nth Gaussian basis element onto the 2D screen space. The product term is used to simulate the attenuation of transparency accumulated from front to back, so that the influence of the front-layer Gaussian on the back-layer Gaussian conforms to the physical rendering law.

[0041] To achieve efficient overall rendering and fast sorting, we adopt a tile-based rasterizer for processing Gaussian ellipsoids. This method not only supports approximate α-blending (including anisotropic ellipsoids), but also breaks through the limit on the number of ellipsoids that can receive gradients. By pre-sorting the projection points of the entire image, we avoid the additional computational overhead caused by pixel-by-pixel sorting in traditional α-blending schemes. Specifically, we divide the screen into tiles of 16×16 pixels, and each tile serves as an independent processing unit. In this framework, only the Gaussian ellipsoids that intersect with the view frustum or the tile are retained, especially those with 99% confidence intervals overlapping with the view frustum. To further optimize the computational efficiency, we introduce a guard band to exclude some ellipsoids at extreme positions, such as those with means close to the near plane but far from the view frustum, to avoid instability in the calculation of projection covariance.

[0042] During the rasterization process, each Gaussian ellipsoid is instantiated according to the number of tiles it covers and assigned a key containing the view-space depth and the tile ID. Subsequently, we use GPU radix sort to efficiently sort these instances without additional pixel-level sorting. This sorting method may introduce approximate errors in some cases, but when the ellipsoid size is close to the pixel size, the error is negligible. In addition, this method can significantly improve the rendering and training performance and does not produce obvious artifacts in the convergence scenario.

[0043] After sorting, we generate a rasterization list for each tile, recording the first and last Gaussian ellipsoids of the sorted depth. In the rasterization stage, each tile is processed by a threadblock, which collaboratively loads the relevant Gaussian ellipsoids into the shared memory and traverses the list from front to back, accumulating the color and α value for each pixel in turn. When the pixel reaches the target saturation, the corresponding thread stops computing. In addition, we regularly check the thread status and terminate the calculation of the tile in advance when all pixels within the tile reach full saturation (α = 1), thus further improving the computational efficiency.

[0044] S5. Block Training and Detail Optimization

[0045] During the training process, we construct a loss function based on the difference between the rendered image and the ground truth image and use the gradient descent algorithm for iterative optimization. Among them, the densification strategy realizes the accurate reconstruction of complex structures by increasing the Gaussian sphere density in the high-frequency detail area; while the point deletion strategy significantly improves the computational efficiency and optimizes the reconstruction quality by removing redundant or less contributing Gaussian spheres. In addition, to support the Gaussian reconstruction of large-scale scenes and generate multi-level of detail (LOD) results, we further introduce a block strategy and an image pyramid method (as Figure 2 shown). The block strategy divides the scene into local regions for independent optimization, reducing the computational complexity; the image pyramid method realizes efficient reconstruction from global to local through multi-scale analysis. These strategies will be elaborated in detail below.

[0046] S5-1. Image Pyramid Strategy

[0047] An image pyramid is a hierarchical structure formed by arranging the same image at different resolutions, with its resolution increasing gradually from the top layer to the bottom layer, that is, gradually refining from low resolution to high resolution. During the Gaussian training process, since the goal of the model is to fit the original image, adopting the image pyramid training strategy can generate Gaussian models corresponding to each resolution level. This multi-scale training method enables the model to fully learn the feature expressions of the image at different resolutions, thereby generating rendering results with multi-level details (LOD) when outputting. Through this hierarchical mechanism, the finally output image can not only maintain a consistent visual effect globally, but also present different precisions in local details, thus flexibly adapting to the requirements of various application scenarios from long-distance viewing to close inspection. In the present invention, the number of layers of the image pyramid is dynamically determined according to the scene scale, and usually 4 to 10 layers are selected to balance detail retention and computational efficiency.

[0048] S5-2. Blocking strategy

[0049] In the implementation of the blocking strategy, taking the first layer of the image pyramid as an example, we first calculate the axis-aligned bounding box (AABB) of the scene based on the sparse SfM point cloud, with its length and width being L and H respectively. Assuming the size of the AABB of the training scene is l0×l0, the entire scene is thus divided into regular grid blocks. For each grid block, we filter out the corresponding points contained therein and select the cameras with visibility values exceeding the preset threshold γ. The visibility value is defined as the ratio of the projected area of the grid block on the image plane to the total area of the image. To further improve the reconstruction accuracy and avoid unnecessary point cloud splitting or movement, we also include the corresponding points related to the current grid block in each image in the calculation scope, even if these points are outside the current grid block. This design ensures the integrity of local reconstruction and global consistency, while optimizing the computational efficiency.

[0050] To further improve the training efficiency, we adaptively adjust the total number of iterations according to the number of pictures in each single-block training, so as to flexibly adapt to the learning needs of different data scales and avoid the problems of waste of training resources or insufficient training. At the same time, aiming at the characteristics of different levels, we dynamically set the splitting interval and reset the opacity to optimize the internal state of the model, making the training process more efficient and accurate, and ensuring that the model can maintain excellent performance under different numbers of pictures.

[0051] S5-3. Densification strategy

[0052] In the densification stage, we adopt an adaptive splitting strategy based on the gradient threshold. As Figure 3 shown, in the traditional 3DGS method, both cases where the two gradients are close to 0 will trigger the densification operation. However, the gradient being close to 0 is only a sufficient condition for densification rather than a necessary condition. For example, Figure 3The gradient situation shown on the left actually does not require densification. To solve this problem, we use the absolute value of the gradient as the judgment basis and combine it with the number of pixels occupied by the Gaussian projection on the image for comprehensive evaluation. As Figure 4 shown, this method effectively avoids unnecessary splitting of overly small Gaussian ellipsoids.

[0053] This gradient-based adaptive splitting strategy can not only ensure that the number of Gaussian ellipsoids grows exponentially during the splitting process and reaches the preset target number at the end of the iteration, thus accurately capturing scene details, especially performing well in the reconstruction of small objects and ensuring their accurate reconstruction in subsequent steps. In addition, this strategy also significantly improves the stability of the splitting process, effectively avoiding the waste of computing resources and the extension of training time caused by over-splitting.

[0054] S5-4. Point Deletion Strategy

[0055] In terms of the point deletion strategy, we mainly remove Gaussian ellipsoids with relatively low opacity and little contribution to the final rendering effect in the point cloud file. Specifically, during the iteration process, when the iteration times reach 10000 and 16000, the system will automatically remove Gaussian ellipsoids with opacity close to 0 because these ellipsoids have little contribution to the rendering result. As Figure 5 shown, this strategy significantly reduces the total number of Gaussian spheres. This point deletion mechanism not only improves the rendering efficiency but also greatly reduces the memory occupancy, especially being particularly effective when dealing with large-scale scenes.

[0056] By adaptively controlling the splitting speed and combining the point deletion strategy, we achieved peak control of the number of Gaussian ellipsoids twice during the training process. This method not only optimizes the accuracy of detail reconstruction but also effectively suppresses the problem of floating objects caused by over-splitting. This refined control mechanism ensures the rendering quality while significantly optimizing the utilization rate of computing resources, thereby overall improving the rendering performance.

[0057] S6. Generating True Orthophotos

[0058] After the training is completed, we perform a top-down orthographic projection on the model to generate a true orthophoto, thereby obtaining the final reconstruction result.

[0059] The present invention provides a computer-readable storage medium for generating true orthophotos based on three-dimensional Gaussian model technology. The storage medium stores a series of computer programs for implementing the generation of true orthophotos based on 3DGS technology. The programs include a data preprocessing module, a core algorithm module, and a post-processing module: the data preprocessing module is used to process the input image data and generate Gaussian point clouds; the core algorithm module is responsible for executing the 3DGS algorithm to construct a 3DGS model; the post-processing module generates true orthophotos based on the generated 3DGS model. These programs are optimized and designed to run efficiently on a processor, thus ensuring high efficiency and high precision in the true orthophoto generation process.

[0060] In addition, the present invention also provides an electronic device, including a processor and a memory, which are efficiently connected through a communication bus. The memory pre-stores specially designed computer programs, which contain a series of computer-readable instructions for executing a method for generating true orthophotos based on three-dimensional Gaussian technology. The processor is specially configured to efficiently call and execute these instructions, quickly start the program running, and thus achieve high-efficiency and high-quality true orthophoto generation.

[0061] Beneficial effects brought by the technical solution provided by the present invention

[0062] The present invention proposes an innovative method for generating large-scale digital orthophotos based on 3DGS technology. This method can generate high-quality digital orthophotos in real time while avoiding the cumbersome image stitching process in traditional methods and effectively eliminating errors introduced by the photographic tilt angle. To further improve the reconstruction efficiency and accuracy of 3DGS, the present invention adopts a resolution pyramid and a block training strategy to optimize 3DGS. In addition, in the densification stage of 3DGS, by introducing a gradient threshold to control the splitting process, the utilization rate of computing resources is significantly improved, and redundant initial point clouds are deleted, which not only improves the reconstruction accuracy but also effectively suppresses the generation of floating objects, thus achieving higher-quality scene reconstruction. Brief Description of the Drawings

[0063] Figure 1 is a flowchart of a method for generating true orthophotos based on a three-dimensional Gaussian model according to an embodiment of the present invention.

[0064] Figure 2 is the image block and image pyramid strategy according to an embodiment of the present invention

[0065] Figure 3 is the case where the gradient is 0 according to an embodiment of the present invention

[0066] Figure 4 is the densification strategy according to an embodiment of the present invention.

[0067] Figure 5It is the dot deletion strategy of the embodiments of the present invention.

[0068] Figure 6 It is the DOM generation pipeline based on 3DGS of the embodiments of the present invention.

[0069] Figure 7 It is a block diagram of a storage redundant electronic device in an exemplary embodiment of Embodiment 1 of the present invention.

[0070] The embodiments disclosed in the present invention are intended to provide clear guidance for those skilled in the relevant art so that they can implement or apply the present invention smoothly. These embodiments are not fixed but highly flexible. Based on their professional knowledge and practical experience, those skilled in the art can make various modifications and adjustments to these embodiments to meet different application scenarios and specific requirements.

[0071] In the implementation process of the present invention, we emphasize the understanding and application of the basic principles. These basic principles form the solid foundation of the present invention and support its overall architecture. Without departing from the core spirit and overall scope of the present invention, those skilled in the art can flexibly apply them in other embodiments according to the basic principles described herein, so as to develop more practical solutions with application value.

[0072] Therefore, the scope of the present invention should not be limited to the specific embodiments listed herein. On the contrary, its scope should cover the broadest scope consistent with the principles and innovative features disclosed herein. This broad inclusiveness not only reflects the openness of the present invention but also provides sufficient space for future technological development and innovation, enabling the present invention to continuously play its value in the evolving technological environment and make positive contributions to the progress of the relevant fields.

Claims

1. A true orthogonal image generation method based on a three-dimensional Gaussian model, the specific implementation of which includes the following steps: S1. Generate data set: Generally, drone images or aerial images are used, and the data set is divided into training set, validation set and test set. S2. Initialization: The acquired image is converted into a sparse point cloud through the SfM technology, and the three-dimensional Gaussian model is initialized based on the sparse point cloud. S3, Projection: In order to render true projection images from different viewpoints, the perspective transformation matrix W is used i and the Jacobian matrix J i , project the 3D Gaussian point cloud from the three-dimensional space to the two-dimensional image plane to achieve orthogonal projection. S4, Rendering: Rendering is performed using a rendering formula in combination with a tile-based rasterizer. S5, Training: The loss function is constructed by combining the 3D Gaussian model and the rendering formula, and the gradient descent method is used to optimize the model. In order to achieve Gaussian reconstruction of large scenes and generate LOD results, the block strategy and image pyramid strategy are introduced. In order to effectively control the number of Gaussian balls, the densification strategy and point deletion strategy are adopted to reduce redundant information while ensuring that the details are fully reconstructed. S6, Reasoning: Use the trained model to perform top-view orthographic projection to generate high-quality true orthographic images in real time.

2. According to the method for generating true orthogonal images based on a three-dimensional Gaussian model as described in claim 1, its innovation lies in: introducing blocking and image pyramid strategies to support efficient reconstruction and LOD generation of large scenes; adopting densification and point deletion strategies to dynamically control the number of Gaussian balls to ensure the accuracy of detail reconstruction and optimize the allocation of computing resources; in the three-dimensional Gaussian model, the traditional perspective projection is modified into a top-view orthographic projection to generate high-quality true orthogonal images in real time.

Citation Information

Cited By

  • Method, device and equipment for generating orthoimage based on point cloud and storage medium

    CN120746921A

  • A method, device and equipment for generating orthographic images based on point clouds, and a storage medium

    CN120746921B

  • Virtual camera simulation data generation method and system based on Gaussian point cloud model

    CN121304946A

  • Three-dimensional Gaussian splash reconstruction method for underwater scene

    CN121639946A

  • A three-dimensional gaussian splatting reconstruction method for underwater scenes

    CN121639946B