Method for generating multi-view data set based on fusion of oblique photography live-action three-dimensional model and Gaussian splash three-dimensional model
By combining UAV oblique photography and 3D Gaussian splashing technology, multi-view datasets are automatically generated, solving the problems of high-cost data acquisition and insufficient image reconstruction in deep learning model training, and achieving low-cost, high-quality image generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies rely on high-cost, manually captured datasets for training deep learning models, and traditional oblique photogrammetry models lack reconstruction capabilities, making it difficult to generate high-quality multi-view images. 3D Gaussian splashing technology is not stable enough with sparse input, making it difficult to generate a large number of multi-view 2D images.
By combining drone oblique photography and 3D Gaussian splashing technology, multi-view datasets are automatically generated, viewpoints are automatically generated in the 3D model space using virtual camera path rules, and high-quality images are rendered using platform tools.
It significantly reduces data acquisition costs, generates images of any number and perspective, covers dangerous environments, and produces high-quality images that meet the training needs of AI models, solving the entire process problem from few images to many images and then to high-quality images.
Smart Images

Figure CN121837564A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technical improvement of multi-view image datasets, specifically a method for generating multi-view datasets based on the fusion of oblique photogrammetry real-world 3D and Gaussian splash 3D models. Background Technology
[0002] Currently, the training of deep learning models (especially in fields such as autonomous driving and remote sensing monitoring) heavily relies on large-scale, multi-view, and multi-distance image datasets. Traditional dataset construction methods mainly rely on manual on-site photography, which has inherent bottlenecks such as long collection cycles, high economic costs, and limited view coverage. In addition, on-site photography is almost impossible in certain dangerous environments or hard-to-reach areas.
[0003] In the field of 3D reconstruction technology, oblique photogrammetry can capture images with multiple overlapping views through a specific flight path, making it an important means of generating realistic 3D models. However, traditional oblique photogrammetry models often suffer from insufficient ability to reconstruct small objects, limiting their ability to generate high-quality 2D images from the models.
[0004] In recent years, 3D Gaussian Splatting (3DGS) technology has become an emerging hot topic in 3D reconstruction due to its excellent rendering speed and quality. However, its stability when processing sparse input images, and how to effectively synthesize large batches of multi-view 2D images from the generated Gaussian model, remain unsolved problems. Summary of the Invention
[0005] Based on the aforementioned problems in the existing technology, the present invention provides a method for generating a multi-view dataset by fusing oblique photogrammetry real-world 3D and Gaussian splash 3D models. This method combines the path planning advantages of oblique photogrammetry with the efficient reconstruction capabilities of 3D Gaussian splash and can automatically generate multi-view image datasets.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for generating a multi-view dataset based on the fusion of oblique photogrammetry real-world 3D and Gaussian splash 3D models, comprising: (1) Real-scene 3D model generation module Data acquisition: Raw images were acquired through actual UAV oblique photography flight and 360° surround path; Core reconstruction technologies: UAV oblique photogrammetry modeling technology and 3D Gaussian splash modeling technology are used; Model fusion: By fusing the Gaussian splash model with the tilt model, latitude and longitude can be accurately located according to the scene; (2) Virtual viewpoint automatic planning module Path and rule definition: Customize or use the default virtual camera path rules; Viewpoint coordinates and parameter generation: Based on the above rules, this module automatically generates a series of virtual camera spatial coordinates (X, Y, Z), orientation (yaw, pitch) and focal length in the 3D model space, forming a complete virtual acquisition path; (3) Multi-view image rendering and post-processing module High-quality image rendering: Using the developed platform tools, the initial 2D image is generated from each planned virtual viewpoint based on the designed shooting position.
[0007] Preferably, the virtual camera path rule can be a surround path, that is, a 360° surround around the model at multiple horizontal levels.
[0008] Preferably, the virtual camera path rule can be a radial advance / retreat path, that is, gradually approaching or moving away from the model body along a fixed angle.
[0009] Preferably, the virtual camera path rule can be an oblique photography simulation path: simulating the five-lens shooting angles of a real drone, including vertical, front, back, left, and right.
[0010] Compared with the prior art, the beneficial effects of the present invention are: 1. Significantly reduce data acquisition costs: Transform the acquisition of AI training data from expensive physical world photography to low-cost, automated digital content generation.
[0011] 2. Unlimited and controllable sample generation: It can generate any number of images from any angle and distance, covering extreme angles and dangerous environments that are difficult to capture by humans.
[0012] 3. High-quality generated data: Based on advanced 3D Gaussian splashing technology, the generated 3D models and derived images have high geometric consistency and visual fidelity. Through post-processing such as tilted 3D fusion, the image quality meets the basic requirements for training professional AI models.
[0013] 4. Technological Integration and Innovation: It creatively combines oblique photography path, 3D Gaussian splash reconstruction, virtual viewpoint planning and image enhancement technology to solve the whole process problem from "few images" to "many images" and then to "high-quality images". Attached Figure Description
[0014] Figure 1 This is a flight image taken along the inclined flight path of the present invention; Figure 2 This is a 360° full-range photograph of the present invention; Figure 3 This is a tilted, real-world 3D model diagram of the present invention; Figure 4 This is a diagram of the Gaussian splashing model of the present invention; Figure 5 This is a fusion diagram of the tilt and Gaussian splash model of the present invention; Figure 6 The image showing the shooting locations generated by the algorithm of this invention; Figure 7 This is a perspective image of the present invention; Figure 8 This is a perspective image from the present invention. Figure 9 This is a perspective image from the present invention. Figure 10 Image 4 is a perspective view of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] The method of the present invention for generating a multi-view dataset based on the fusion of oblique photogrammetry real-world 3D and Gaussian splash 3D models is characterized by comprising: (1) Real-scene 3D model generation module Data acquisition: Raw images were acquired through actual UAV oblique photography flight and 360° surround path; Core reconstruction technologies: UAV oblique photogrammetry modeling technology and 3D Gaussian splash modeling technology are used; Model fusion: By fusing the Gaussian splash model with the tilt model, latitude and longitude can be accurately located according to the scene; (2) Virtual viewpoint automatic planning module Path and rule definition: Customize or use the default virtual camera path rules; The virtual camera path rule can be a surround path, that is, a 360° surround around the model at multiple horizontal levels.
[0017] The virtual camera path rule can be a radial advance / retreat path, that is, gradually approaching or moving away from the model body along a fixed angle.
[0018] The virtual camera path rule can be an oblique photography simulation path: simulating the five-lens shooting angles of a real drone, including vertical, front, back, left, and right.
[0019] Viewpoint coordinates and parameter generation: Based on the above rules, this module automatically generates a series of virtual camera spatial coordinates (X, Y, Z), orientation (yaw, pitch) and focal length in the 3D model space, forming a complete virtual acquisition path; (3) Multi-view image rendering and post-processing module High-quality image rendering: Using the developed platform tools, the initial 2D image is generated from each planned virtual viewpoint based on the designed shooting position.
[0020] In practice, the original images are collected through actual drone oblique photography flights and 360° surround paths, specifically as follows: Figure 1 and 2 As shown; The original 2D image is reconstructed using a 3D Gaussian splashing technique with a voxel alignment strategy to generate a realistic 3D model, such as... Figure 3 , 4 As shown in Figure 5; In the space surrounding the real-world 3D model, an algorithm is used to automatically plan a series of virtual camera viewpoints based on the latitude, longitude, and altitude of the target center point. Figure 6 As shown; Starting from the virtual camera viewpoint, an initial two-dimensional image is generated, and the final synthesized image dataset is output (the following are some rendered images), such as... Figure 7-10 .
[0021] In practice, the coordinates of the center point are known to be: P0 = (x0, y0, z0). Maximum radius: Rmax>0; Number of layers: L∈N+; Number of points per layer: N∈N+N∈N+ Radius decrease factor: γ∈(0,1) Requirements: Generate L×NL×N viewpoints, distributed on a hemisphere, from the outside to the inside and from the bottom to the top.
[0022] Establish the expression of the calculus function
[0023] 1. Continuous parameterization
[0024] Define the continuous parameter domain: D={(u,v)∈R2∣0≤u<2π, 0≤v≤π2}D={(u,v)∈R2∣0≤u<2π,0≤v≤2π} Where: uu: azimuth (longitude); vv: elevation (latitude) 2. Radius-decreasing function Define a radius function r(v)∈C1[0,π / 2], satisfying: r(0) = Rmax (maximum bottom radius) r(π / 2)=γRmaxr(π / 2)=γRmax (minimum top radius) drdv<0dvdr<0 (strictly decreasing) Using the exponential decay model: r(v) = Rmax exp(-λ vπ / 2)r(v)=Rmax exp(-λ π / 2v) Where λ = -lnγ > 0.
[0025] Alternatively, linear decay can be used: r(v) = Rmax [1-(1-γ) [vπ / 2]r(v)=Rmax[1-(1-γ)] π / 2v] A more general expression: r(v) = Rmax [1-(1-γ) [(vπ / 2)α],α>0r(v)=Rmax [1-(1-γ) [(π / 2v)α],α>0 3. Viewpoint position field Define the mapping from the parameter domain to three-dimensional space: X:D→R3X:D→R3X(u,v)=P0+r(v) [sinvcosusinvsinucosv]X(u,v)=P0+r(v) sinvcosusinvsinucosv This is the rectangular coordinate representation of spherical coordinates.
[0026] 4. Discrete Sampling Scheme
[0027] For a given number of layers LL and the number of points per layer NN, perform discrete sampling in the parameter domain: Pitch angle sampling (latitude direction): Let vk be the pitch angle of the k-th layer, k = 0, 1, ..., L-1 Uniform sampling: vk=π2 kL-1,k=0,1,…,L-1vk=2π L-1k,k=0,1,…,L-1 Alternatively, use equal-area sampling (which is better): vk=arccos(1-kL-1 (1-cosvmax))vk=arccos(1-L-1k (1-cosvmax)) Where vmax≤π / 2 is the maximum pitch angle.
[0028] Azimuth sampling (longitude direction): Let ujuj be the azimuth angle of the j-th point, j=0,1,…,N-1 Uniform sampling: uj=2πN j+δkuj=N2π j+δk Where δk is the phase shift of the k-th layer, to avoid symmetry, it can be taken as: δk=πN kL-1δk=Nπ L-1k 5. Formula for calculating discrete viewpoints The k-th viewpoint at layer j: Pk,j=P0+r(vk) [sinvkcosujsinvksinujcosvk]Pk,j=P0+r(vk) sinvkcosujsinvksinujcosvk in: vk = vmin + (vmax - vmin) kL-1vk=vmin+(vmax-vmin) L-1kuj=2πN j+δ kL-1,δ=πN (optional) uj=N2π j+δ L-1k,δ=Nπ(optional)r(vk=Rmax [1-(1-γ) (vk-vminvmax-vmin)α]r(vk)=Rmax [1-(1-γ) (vmax-vminvk-vmin)α] 6. Camera Orientation Calculation The line of sight should be directed towards the center point: dk,j=P0-Pk,jdk,j=P0-Pk,j Normalization: d^k,j=dk,j dk,j d^k,j= dk,j dk,j Calculate Euler angles: Yaw angle (clockwise from north): ψk,j=arctan2(d^k,j,yd^k,j,x)ψk,j=arctan2(d^k,j,xd^k,j,y) Pitch (positive downwards): k,j=arcsin(-d^k,j,z) k,j=arcsin(-d^k,j,z) 7. Differential geometric quantities Tangent vector: Tu= X u=r(v)sinv [-sinucosu0]Tu= u X=r(v)sinv -sinucosu0Tv= X v=r′(v) [sinvcosusinvsinucosv]+r(v) [cosvcosucosvsinu-sinv]Tv= v X=r′(v) sinvcosusinvsinucosv+r(v) cosvcosucosvsinu-sinv Normal vector: N=Tu×Tv Tu×Tv N= Tu×Tv Tu×Tv Curvature: Gaussian curvature: K=1r2(v)K=r2(v)1 (When r′(v)=0, it is a sphere) 8. Area Infinite Elements and Viewpoint Density Area micro-element: dA= Tu×Tv du dv=r(v)sinv[r′(v)]2+[r(v)]2 du dvdA= Tu×Tv dudv=r(v)sinv[r′(v)]2+[r(v)]2dudv For discrete sampling, the area covered by each viewpoint is approximately: ΔAk,j≈1L×N DdA=1LN 0π / 2 02πr(v)sinv[r′(v)]2+[r(v)]2 du dvΔAk,j≈L×N1 DdA=LN1 0π / 2 02πr(v)sinv[r′(v)]2+[r(v)]2dudv 9. Optimize the objective function To achieve a uniform viewpoint distribution, the objective function can be minimized: J=∑k=0L-2∑j=0N-1( Pk+1,j-Pk,j 2+ Pk,(j+1)mod N-Pk,j 2) J=k=0∑L-2j=0∑N-1( Pk+1,j-Pk,j 2+ Pk,(j+1)modN-Pk,j 2) 10. Integral representation in the limiting case As L, N→∞, the viewpoint set approaches a continuous distribution, which can be represented by an integral: Total viewpoint "density": Φ= Dρ(u,v) dAΦ= Dρ(u,v)dA Where ρ(u,v) is the viewpoint surface density function.
[0029] If a uniform distribution is required, then ρ(u,v) = constant.
[0030] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for generating a multi-view dataset by fusing oblique photogrammetry real-world 3D and Gaussian splash 3D models, characterized in that, include: (1) Real-scene 3D model generation module Data acquisition: Raw images were acquired through actual UAV oblique photography flight and 360° surround path; Core reconstruction technologies: UAV oblique photogrammetry modeling technology and 3D Gaussian splash modeling technology are used; Model fusion: By fusing the Gaussian splash model with the tilt model, latitude and longitude can be accurately located according to the scene; (2) Virtual viewpoint automatic planning module Path and rule definition: Customize or use the default virtual camera path rules; Viewpoint coordinates and parameter generation: Based on the above rules, this module automatically generates a series of virtual camera spatial coordinates (X, Y, Z), orientation (yaw, pitch) and focal length in the 3D model space, forming a complete virtual acquisition path; (3) Multi-view image rendering and post-processing module High-quality image rendering: Using the developed platform tools, the initial 2D image is generated from each planned virtual viewpoint based on the designed shooting position.
2. The method for generating a multi-view dataset based on the fusion of oblique photogrammetry real-world 3D and Gaussian splash 3D models as described in claim 1, characterized in that, The virtual camera path rule can be a surround path, that is, a 360° surround around the model at multiple horizontal levels.
3. The method for generating a multi-view dataset based on the fusion of oblique photogrammetry real-world 3D and Gaussian splash 3D models as described in claim 1, characterized in that, The virtual camera path rule can be a radial advance / retreat path, that is, gradually approaching or moving away from the model body along a fixed angle.
4. The method for generating a multi-view dataset based on the fusion of oblique photogrammetry real-world 3D and Gaussian splash 3D models as described in claim 1, characterized in that, The virtual camera path rule can be an oblique photography simulation path: simulating the five-lens shooting angles of a real drone, including vertical, front, back, left, and right.