Rendering optimization method and system for updating cluster building in super-large scale city
By quantifying the three-dimensional Gaussian body properties and adopting a progressive resolution training strategy, the problems of long training time, slow rendering speed and high memory usage when rendering hyperscale urban renewal cluster buildings are solved, and efficient rendering and training are achieved.
Patent Information
- Application Number
- CN202510051602.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
When rendering hyperscale urban renewal cluster buildings, the prior art has problems such as long training time, slow rendering speed and high memory usage.
The three-dimensional Gaussian body properties are expressed in quantified manner, and the rendering process of three-dimensional Gaussian point clouds is optimized through a progressive resolution training strategy from rough to fine.
It significantly reduces memory storage requirements, improves rendering efficiency and training speed, and can achieve real-time rendering in full resolution scenarios.
Smart Images

Figure CN119991903A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to a rendering optimization method and system for super-large-scale urban renewal cluster buildings. Background Art
[0002] There are currently two main methods for rendering through deep learning, one is neural radiation field and the other is three-dimensional Gaussian splattering, but they each have their own problems.
[0003] Problem 1: Neural radiation fields have problems such as long training time, slow rendering speed, and low effect.
[0004] Question 2: Real-time rendering and accelerated training can be achieved through fast and differentiable 3D Gaussian splashing, but this requires a lot of memory resources because each scene requires millions of Gaussian point clouds. Especially for ultra-large-scale urban renewal cluster building scenes, if grid rendering is used, the number of triangles exceeds 5 billion and the number of triangle vertices exceeds 10 billion. To construct a Gaussian point cloud, at least 5 billion are required, which is an ultra-large-scale virtual scene. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a rendering optimization method and system for super-large-scale urban renewal cluster buildings.
[0006] To achieve the above object, the present invention adopts the following technical solution:
[0007] A rendering optimization method for super-large-scale urban renewal cluster buildings, comprising:
[0008] Obtain images of buildings in super-large-scale urban renewal clusters;
[0009] Based on the images of super-large-scale urban renewal cluster buildings, image rendering is performed through three-dimensional Gaussian splashing. Among them, a quantitative method is used to represent the properties of the three-dimensional Gaussian volume, and a progressive resolution training strategy from coarse to fine is adopted.
[0010] As a preferred method, the three-dimensional Gaussian volume attributes are represented in a quantitative manner as follows:
[0011] A set of properties of a 3D Gaussian body include:
[0012] Position vector:
[0013] Scaling factor:
[0014] Rotate a quaternion vector:
[0015] Opacity
[0016] Spherical harmonics coefficients where d = 3f 2 , f represents the harmonic degree;
[0017] Quantization uses 16-bit int type integers to represent 64-bit float type floating point numbers;
[0018] For attribute collections Any attribute in , converts an l-dimensional quantized latent vector consisting of integers Input to the multi-layer perceptron decoder In, we get k-dimensional attributes;
[0019] During the training process, a straight line estimator is used to approximate the continuous Take the integer and pass it directly to the gradient during the reverse transfer process, that is, in, is the desired k-dimensional attribute, yes Continuous approximation of ;
[0020] Only for and SH for quantification.
[0021] As a preferred method, a progressive resolution training strategy from coarse to fine is adopted as follows:
[0022] The super-large-scale urban renewal cluster building images are downsampled to generate resolution images with different levels of details, and 6 coarse resolution images are obtained, with resolutions of 256*256, 128*128, 64*64, 32*32, 16*16, and 8*8 respectively;
[0023] From coarse resolution to full-size resolution, training is done batch by batch; from 8*8 to 512*512, there are 7 batches in total, of which the first 6 batches are trained for 10,000 iterations each, and the last batch is trained for 40,000 iterations, for a total of 100,000 trainings.
[0024] Preferably, during the training phase, the three-dimensional Gaussian point cloud with opacity less than the threshold φ is removed every 100 iterations, that is, Among them, φ=0.5.
[0025] Preferably, during the training phase, every 100 iterations, the segmentation position gradient is greater than the threshold A three-dimensional Gaussian point cloud, where
[0026] Preferably, in step S1, optical remote sensing satellites, manned aircraft, or drones are used to take oblique photographs of densely populated urban areas to take pictures or videos, so as to obtain images of ultra-large-scale urban renewal cluster buildings; wherein, densely populated urban areas include: dense buildings, dense vehicles and pedestrians, and complex traffic.
[0027] The present invention also provides a rendering optimization system for super-large-scale urban renewal cluster buildings, comprising:
[0028] An acquisition module is used to obtain images of super-large-scale urban renewal cluster buildings;
[0029] The rendering optimization module is used to render images based on super-large-scale urban renewal cluster building images through three-dimensional Gaussian splashing; in which a quantized method is used to represent the properties of the three-dimensional Gaussian volume, and a progressive resolution training strategy from coarse to fine is adopted.
[0030] Preferably, during the training phase, the three-dimensional Gaussian point cloud with opacity less than the threshold φ is removed every 100 iterations, that is, Among them, φ=0.5.
[0031] Preferably, during the training phase, every 100 iterations, the segmentation position gradient is greater than the threshold A three-dimensional Gaussian point cloud, where
[0032] Preferably, optical remote sensing satellites, manned aircraft, and drones are used to take oblique photographs of densely populated urban areas to take pictures or videos, so as to obtain images of ultra-large-scale urban renewal cluster buildings; wherein, densely populated urban areas include: dense buildings, dense vehicles and pedestrians, and complex traffic.
[0033] In order to solve the problems of neural radiation field and 3D Gaussian splashing at the same time, the present invention uses quantization properties to significantly reduce memory storage requirements, and a progressive resolution training strategy from coarse to fine to optimize 3D Gaussian point clouds faster and more stably.
[0034] The present invention uses fewer three-dimensional Gaussian volumes and adopts a quantitative method to represent properties, thereby improving the training time and rendering speed of real-time rendering in full-resolution scenes. Specifically, compared with neural radiation fields, the present invention can improve rendering efficiency by about 50 times and save about 100 times of training time; compared with three-dimensional Gaussian splashing, the present invention can save more than 98% of memory resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0036] Figure 1 This is a flow chart of a rendering optimization method for super-large-scale urban renewal cluster buildings according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] Embodiment 1:
[0040] 3D Gaussian splatting is a method of representing a 3D scene with a 3D Gaussian point cloud and rendering it quickly in a rasterization pipeline. A 3D Gaussian point cloud is also a point cloud, initialized with a sparse point cloud. Unlike ordinary point clouds, each point in a 3D Gaussian point cloud is represented by a 3D Gaussian volume. A 3D Gaussian volume is an elliptical 3D spatial structure composed of multiple attributes, such as position, scale, rotation, color, opacity, etc. When rendering, the anisotropic volume of the 3D Gaussian volume is projected onto a 2D plane, that is, the 3D Gaussian volume is affine mapped onto a 2D Gaussian plane, and then a differentiable tile-based rasterization renderer is used to blend different 3D Gaussian volumes together, and finally the attributes are further optimized using the gradient descent method. In short, 3D Gaussian splatting is a method of rasterization rendering using a 3D Gaussian point cloud.
[0041] Three-dimensional Gaussian expression: in is the position vector of a point in three-dimensional space, represents the three-dimensional covariance matrix, and the symbol T represents the transpose operator.
[0042] For a three-dimensional covariance matrix, its definition is Where x, y, z are points in space The coordinates of , σ represents variance, and Cov represents covariance. The initialization method of the three-dimensional covariance matrix is: Where S represents a three-dimensional scaling vector and R represents a four-dimensional rotation vector.
[0043] Will Affine mapping to two dimensions gives Where W represents the projection transformation matrix from world coordinates to camera coordinates, which mainly performs translation and rotation, both of which are affine transformations. However, the projection transformation is non-affine, so the Jacobian matrix J is used instead of the projection, which serves as an affine approximation of the projection transformation.
[0044] The final color C of the pixel is calculated by the pixel contribution covered by N 2D Gaussian regions. k , Among them C k represents the color of the k-th pixel, k represents the screen pixel index, and N represents that there are N 2D Gaussian regions covering the k-th pixel. represents the opacity of the i-th 2D Gaussian region affecting the color of the k-th pixel, Denotes the color of the i-th 2D Gaussian that influences the color of the k-th pixel.
[0045] The loss functions of 3D Gaussian splashing include: total loss function, reconstruction loss function and structural similarity loss function; the total loss function is: Among them, the coefficient of the loss function is λ = 0.2; the reconstruction loss function By comparing the color C of the kth pixel k The true color of the kth pixel The L1 loss of ; the structural similarity loss function is:
[0046]
[0047] Among them, I1 and I2 are two input images, and are the average values of I1 and I2 respectively, and are the standard deviations of I1 and I2, is the covariance of I1 and I2, C1 and C2 are constants used to stabilize the calculation. Structural similarity loss is used to measure the structural similarity between two images. The closer the value is to 1, the more similar the two images are, and the closer the value is to 0, the less similar the two images are.
[0048] like Figure 1 As shown, an embodiment of the present invention provides a rendering optimization method for a super-large-scale urban renewal cluster building, comprising:
[0049] Step S1, obtaining a super-large-scale urban renewal cluster building image;
[0050] Step S2: performing image rendering by three-dimensional Gaussian splashing based on the super-large-scale urban renewal cluster building image; wherein, a quantized method is used to represent the three-dimensional Gaussian volume attributes, and a progressive resolution training strategy from coarse to fine is adopted.
[0051] As an implementation method of an embodiment of the present invention, in step S1, optical remote sensing satellites, manned aircraft, and drones are used to take oblique photography pictures or videos of densely populated urban areas (dense buildings, dense vehicles and pedestrians, and complex traffic) to obtain images of ultra-large-scale urban renewal cluster buildings.
[0052] As an implementation method of the embodiment of the present invention, the three-dimensional Gaussian volume attributes are represented by a quantitative method as follows:
[0053] A set of properties of a 3D Gaussian body Specifically include:
[0054] Position vector:
[0055] Scaling factor:
[0056] Rotate a quaternion vector:
[0057] Opacity
[0058] Spherical harmonics coefficients Where d = 3f 2 , f represents the harmonic degree. For the spherical harmonic function with f = 5, its color coefficient accounts for more than 80% of the dimension of the entire attribute vector; among them, the number of spherical harmonic function coefficients is d = 3f 2 =3×5 2 =75, the number of other parameters is position 3+scaling 3+rotation 4+opacity 1, a total of 75+11=86 dimensions, so 75 / 86=87%.
[0059] The quantitative representation method is as follows:
[0060] The core idea of quantization is to use 16-bit int type integers to represent 64-bit float type floating point numbers;
[0061] For attribute collections Any attribute in , converts an l-dimensional quantized latent vector consisting of integers Input to the multi-layer perceptron decoder In order to obtain k-dimensional attributes;
[0062] because is not differentiable. During the training process, a straight-through estimator (STE) is used to approximate the continuous Round it up and then pass it directly to the gradient during the backward pass, i.e. in is the desired k-dimensional attribute, yes Continuous approximation of ;
[0063] Property Collection Some attributes in the initialization are very sensitive, and quantization will lead to a significant performance degradation, so these parameters are not quantized (including and ), only for and SH for quantification.
[0064] Quantified benefits:
[0065] 1. Converting 64-bit floating point numbers into 16-bit integers can significantly reduce the memory space occupied by data, and can reduce the overall memory usage by about 70%. For each dimension of floating point numbers that occupy 8 bytes, only 2-byte integers are needed. There are 86 dimensions for all attributes, of which 80 dimensions are quantized. Therefore, the memory usage before processing is 8*86=688 bytes, and the memory usage after processing is 2*80+8*6=208 bytes, 1-208 / 688=69.76%.
[0066] 2. By using quantized integer operations instead of floating-point operations, the computing efficiency can be greatly improved, and the overall computing efficiency is increased by about 10 times.
[0067] 3. Quantifying opacity can remove artifacts when rendering new views to a certain extent, thereby improving rendering quality.
[0068] 4. In addition, f=2 can be set to reduce the dimension of the attributes in the three-dimensional Gaussian body. If all parameters are quantized, the memory usage can be greatly reduced.
[0069] As an implementation method of an embodiment of the present invention, in the prior art, the training process of the three-dimensional Gaussian point cloud is to calculate the loss function on the entire input image, which will result in low training efficiency and will produce floating artifacts. The three-dimensional Gaussian point cloud is initialized with a sparse point cloud, and some attributes of the three-dimensional Gaussian body are also roughly estimated, so the initialization of the three-dimensional Gaussian splash is not accurate. Therefore, in the early stages of training, forcibly adapting the three-dimensional Gaussian point cloud to the fine features of the entire scene will result in floating artifacts, and these artifacts cannot be eliminated during the optimization process. In order to solve the problem of floating artifacts in the prior art, the present invention adopts a progressive resolution training strategy from coarse to fine, as follows:
[0070] Downsample the original input image to generate resolution images with different levels of detail. For example, if the resolution of the original input image is 512*512, six coarse resolution images are obtained through downsampling, with resolutions of 256*256, 128*128, 64*64, 32*32, 16*16, and 8*8.
[0071] From coarse resolution to full-size resolution, training is done batch by batch. From 8*8 to 512*512, there are 7 batches in total. The first 6 batches are trained for 10,000 iterations each, and the last batch is trained for 40,000 iterations, for a total of 100,000 trainings.
[0072] Benefits of progressive resolution training:
[0073] 1. Training time: Progressive resolution training can use fewer 3D Gaussian volumes in low-resolution scenes, speed up training, and significantly reduce the training time of low-resolution scenes. Under the premise of a certain total number of training times, the overall training time can be greatly reduced.
[0074] 2. Convergence speed: Progressive resolution training can make the three-dimensional Gaussian point cloud converge to the minimum loss value more easily, achieve faster rendering and back propagation speed, provide better initialization values for the densification process of the three-dimensional Gaussian body, reduce the amount of calculation for each training iteration, and speed up the convergence of training.
[0075] 3. Reconstruction quality: Progressive resolution training can better adapt to input data of different resolutions, reduce blur and noise during training, remove artifacts caused by rasterization rendering of poor 3D Gaussian bodies, reconstruct fine features of the scene, and improve scene reconstruction quality and image rendering effects.
[0076] 4. Stability: Progressive resolution training can provide a more stable optimization process by gradually increasing the resolution of the scene.
[0077] As an implementation method of an embodiment of the present invention, three-dimensional Gaussian densification is the process of copying and splitting a three-dimensional Gaussian body. A large three-dimensional Gaussian body will be divided and a small three-dimensional Gaussian body will be copied to better adapt to the geometric shape. Too many densifications will lead to an explosive growth of three-dimensional Gaussian bodies. In general, the more three-dimensional Gaussian bodies there are, the better the effect, and the more accurate the scene details that can be represented, but redundant three-dimensional Gaussian bodies will be added, thereby increasing the number of training, rendering time and memory usage. The present invention adopts four methods to reduce the three-dimensional Gaussian densification process:
[0078] 1. The current scheme is to perform densification every 100 iterations, and set it to perform densification every 500 iterations.
[0079] 2. Increase the threshold The value of , increases from 0.5 to 1.
[0080] 3. Increase the density interval and adopt a strategy combining importance sampling, random sampling and uniform sampling, that is, distribute more three-dimensional Gaussian bodies in the foreground bounding box and maintain a basically uniform distribution, so that the density interval is more moderate, and use random distribution in other areas; among them, the number of three-dimensional Gaussian bodies N1 in the foreground bounding box and the number of three-dimensional Gaussian ellipsoids N2 in the background are converted according to their areas S1 and S2, that is, the proportional area of the three-dimensional Gaussian ellipsoids in the background is N1*S2 / S1, and the desired result is N1>N1*S2 / S1. If N1≤N1*S2 / S1, the three-dimensional Gaussian body in the foreground bounding box can be further split from 1 to 2, and so on, until N1>N1*S2 / S1 is satisfied.
[0081] 4. During the training phase, every 100 iterations, the segmentation position gradient is greater than the threshold A three-dimensional Gaussian point cloud, where During the training phase, every 100 iterations, the opacity values less than the threshold are removed. The three-dimensional Gaussian point cloud is in
[0082] In the existing 3D Gaussian splashing method, 64-bit floating points are used to represent attribute parameters. For ultra-large-scale virtual scenes, 5 billion (5 billion means that there are 5 billion 3D Gaussian bodies in the 3D Gaussian point cloud) 3D Gaussian point clouds, each 3D Gaussian body f=5, then the memory space occupied is 5 billion*8*86 / 1024 / 1024 / 1024=3203.75GB. Under the joint action of the three optimization strategies of "reducing the number", "controlling densification" and "quantifying attributes", 5 billion 3D Gaussian bodies can be reduced to 1 billion, and f=2 reduces the attribute dimension and reduces the memory usage of the attribute dimension. Specifically, 1 billion*(23*2) / 1024 / 1024 / 1024=42.84GB, reducing the memory usage ratio by 98.66%, equivalent to a reduction of 74.78 times. If f = 2, SH has 12 parameters and the other 11 parameters, so a total of 23 parameters, of which 6 parameters of position and scale are not quantized, and the remaining 17 parameters are quantized. After quantization, 2-byte int can be used to replace the original 8-byte float, so it is 17*2+6*8=82; if further optimization is performed, all parameters are quantized, and the total parameters are 23*2=46. The progressive resolution strategy can greatly speed up the training speed and improve the reconstruction quality.
[0083] Embodiment 2:
[0084] The embodiment of the present invention further provides a rendering optimization system for super-large-scale urban renewal cluster buildings, comprising:
[0085] An acquisition module is used to obtain images of super-large-scale urban renewal cluster buildings;
[0086] The rendering optimization module is used to render images based on super-large-scale urban renewal cluster building images through three-dimensional Gaussian splashing; in which a quantized method is used to represent the properties of the three-dimensional Gaussian volume, and a progressive resolution training strategy from coarse to fine is adopted.
[0087] As an implementation method of the present invention, in the training stage, every 100 iterations, the three-dimensional Gaussian point cloud with opacity less than the threshold φ is removed, that is, Among them, φ=0.5.
[0088] As an implementation method of the present invention, in the training stage, every 100 iterations, the segmentation position gradient greater than the threshold A three-dimensional Gaussian point cloud, where
[0089] As an implementation method of an embodiment of the present invention, optical remote sensing satellites, manned aircraft, and drones are used to take oblique photos or videos of densely populated urban areas to obtain images of ultra-large-scale urban renewal cluster buildings; wherein, densely populated urban areas include: dense buildings, dense vehicles and pedestrians, and complex traffic.
[0090] Embodiment 3:
[0091] An embodiment of the present invention also provides a rendering optimization system for super-large-scale urban renewal cluster buildings, including: a memory and a processor, wherein the memory stores a computer program run by the processor, and when the computer program is run by the processor, it executes a rendering optimization method for super-large-scale urban renewal cluster buildings.
[0092] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.
Claims
1. A rendering optimization method for super-large-scale urban renewal cluster buildings, characterized in that: include: Obtain images of buildings in super-large urban renewal clusters; Based on the images of super-large-scale urban renewal cluster buildings, image rendering is performed through three-dimensional Gaussian splashing. Among them, a quantitative method is used to represent the properties of the three-dimensional Gaussian volume, and a progressive resolution training strategy from coarse to fine is adopted.
2. The rendering optimization method for super-large-scale urban renewal cluster buildings according to claim 1, characterized in that: The quantitative method used to represent the properties of the three-dimensional Gaussian volume is as follows: A set of properties of a 3D Gaussian body include: Position vector: Scaling factor: Rotate a quaternion vector: Opacity Spherical harmonics coefficients where d = 3f 2 , f represents the harmonic degree; Quantization uses 16-bit int type integers to represent 64-bit float type floating point numbers; For attribute collections Any attribute in , converts an l-dimensional quantized latent vector consisting of integers Input to the multi-layer perceptron decoder In, we get k-dimensional attributes; During the training process, a straight line estimator is used to approximate the continuous Take the integer and pass it directly to the gradient during the reverse transfer process, that is, in, is the desired k-dimensional attribute, yes Continuous approximation of ; Only for o and SH were quantified.
3. The rendering optimization method for super-large-scale urban renewal cluster buildings according to claim 2, characterized in that: The specific training strategy of using progressive resolution from coarse to fine is as follows: The super-large-scale urban renewal cluster building images are downsampled to generate resolution images with different levels of details, and 6 coarse resolution images are obtained, with resolutions of 256*256, 128*128, 64*64, 32*32, 16*16, and 8*8 respectively; From coarse resolution to full-size resolution, training is done batch by batch; from 8*8 to 512*512, there are 7 batches in total, of which the first 6 batches are trained for 10,000 iterations each, and the last batch is trained for 40,000 iterations, for a total of 100,000 trainings.
4. The rendering optimization method for super-large-scale urban renewal cluster buildings according to claim 3 is characterized in that: During the training phase, every 100 iterations, the 3D Gaussian point cloud with opacity less than the threshold φ is removed, that is, Among them, φ=0.
5.
5. The rendering optimization method for super-large-scale urban renewal cluster buildings according to claim 4, characterized in that: During the training phase, every 100 iterations, the segmentation position gradient greater than the threshold A three-dimensional Gaussian point cloud, where 6. The rendering optimization method for super-large-scale urban renewal cluster buildings according to claim 5, characterized in that: In step S1, optical remote sensing satellites, manned aircraft, and drones are used to take pictures or videos of densely populated urban areas to obtain images of large-scale urban renewal cluster buildings; wherein, densely populated urban areas include: dense buildings, dense vehicles and pedestrians, and complex traffic.
7. A rendering optimization system for super-large-scale urban renewal cluster buildings, characterized in that: include: An acquisition module is used to obtain images of super-large-scale urban renewal cluster buildings; The rendering optimization module is used to render images based on super-large-scale urban renewal cluster building images through three-dimensional Gaussian splashing; in which a quantized method is used to represent the properties of the three-dimensional Gaussian volume, and a progressive resolution training strategy from coarse to fine is adopted.
8. The rendering optimization system for super-large-scale urban renewal cluster buildings according to claim 5, characterized in that: During the training phase, every 100 iterations, the 3D Gaussian point cloud with opacity less than the threshold φ is removed, that is, Among them, φ=0.
5.
9. The rendering optimization system for super-large-scale urban renewal cluster buildings according to claim 3, characterized in that: During the training phase, every 100 iterations, the segmentation position gradient greater than the threshold A three-dimensional Gaussian point cloud, where 10. The rendering optimization system for super-large-scale urban renewal cluster buildings according to claim 9, characterized in that: By using optical remote sensing satellites, manned aircraft, and drones to take oblique photos or videos of densely populated urban areas, we can obtain images of ultra-large-scale urban renewal cluster buildings; among them, densely populated urban areas include: dense buildings, dense vehicles and pedestrians, and complex traffic.
Citation Information
Patent Citations
Real-time rendering method and device based on multi-level Gaussian sputtering
CN118096972A
Gaussian rendering and reconstruction method based on prior guidance of symbol distance radiation field
CN118314268A
Novel view angle synthesis method based on Gaussian splash and fusing learnable basis function
CN118505541A
Hierarchical progressive coding framework method and system for volume video
CN118890487A
Three-dimensional cloud reconstruction method and device
CN119091043A