Rendering optimization method and system for large-scale urban renewal cluster buildings
By quantifying the properties of 3D Gaussian volumes and implementing a progressive resolution training strategy, we solved the problems of slow neural radiation field rendering and high memory consumption of 3D Gaussian splattering, and achieved efficient rendering of ultra-large-scale urban renewal cluster buildings.
Patent Information
- Application Number
- CN202510051602.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Existing neural radiation field rendering methods take a long time to train and are slow, while the three-dimensional Gaussian splashing method requires a lot of memory resources in ultra-large-scale urban renewal cluster building scenarios, resulting in low rendering efficiency.
The 3D Gaussian volume attributes are represented in a quantitative manner, and a progressive resolution training strategy from coarse to fine is adopted. The 3D Gaussian point cloud is optimized by quantitative attributes and progressive resolution training strategy, which reduces memory requirements and improves rendering speed.
It significantly reduces memory storage requirements, improves rendering efficiency and training speed, improves rendering efficiency by about 50 times compared to neural radiation field, saves about 100 times of training time, and saves more than 98% of memory resources compared to 3D Gaussian splashing.
Smart Images

Figure CN119991903B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to a rendering optimization method and system for super-large-scale urban renewal cluster buildings. Background Art
[0002] There are currently two main methods for rendering through deep learning, one is neural radiance field and the other is 3D Gaussian splattering, but they each have their own problems.
[0003] Problem 1: Neural radiation fields have problems such as long training time, slow rendering speed, and low effect.
[0004] Problem 2: Fast, differentiable 3D Gaussian splatting enables real-time rendering and accelerated training, but this consumes significant memory resources because each scene requires millions of Gaussian point clouds. This is especially true for ultra-large-scale urban renewal cluster building scenes, where mesh rendering can produce over 5 billion triangles and 10 billion triangle vertices. Constructing a Gaussian point cloud also requires at least 5 billion, making this a very large-scale virtual scene. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a rendering optimization method and system for super-large-scale urban renewal cluster buildings.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A rendering optimization method for large-scale urban renewal cluster buildings, comprising:
[0008] Acquire images of large-scale urban renewal cluster buildings;
[0009] Based on the images of super-large-scale urban renewal cluster buildings, image rendering is performed through three-dimensional Gaussian splashing. The three-dimensional Gaussian volume properties are represented in a quantitative manner, and a progressive resolution training strategy from coarse to fine is adopted.
[0010] As a preference, the quantitative method is used to represent the properties of the three-dimensional Gaussian volume as follows:
[0011] A set of properties of a 3D Gaussian body include:
[0012] Position vector:
[0013] Scaling factor:
[0014] Rotate a quaternion vector:
[0015] Opacity
[0016] Spherical harmonic coefficients where d = 3f 2 , f represents the harmonic degree;
[0017] Quantization uses 16-bit int type integers to represent 64-bit float type floating point numbers;
[0018] For attribute collections Any attribute in , a l-dimensional quantized latent vector consisting of integers Input to the multilayer perceptron decoder In , we get k-dimensional attributes;
[0019] During the training process, a straight line estimator is used to approximate the continuous Take the integer and pass it directly to the gradient during the reverse transfer process, that is, in, is the desired k-dimensional attribute, yes Continuous approximation of ;
[0020] Only for and SH were quantified.
[0021] As a preference, a progressive resolution training strategy from coarse to fine is adopted as follows:
[0022] The ultra-large-scale urban renewal cluster building image is downsampled to generate resolution images with different levels of detail. Six coarse resolution images are obtained, with resolutions of 256*256, 128*128, 64*64, 32*32, 16*16, and 8*8 respectively.
[0023] Training is performed batch by batch from coarse resolution to full-size resolution; from 8*8 to 512*512, a total of 7 batches, of which the first 6 batches are trained for 10,000 iterations each, and the last batch is trained for 40,000 iterations, for a total of 100,000 training iterations.
[0024] As a preference, during the training phase, every 100 iterations, the 3D Gaussian point cloud with opacity less than the threshold φ is removed, i.e. Among them, φ=0.5.
[0025] As a preference, during the training phase, every 100 iterations, the segmentation position gradient is greater than the threshold Three-dimensional Gaussian point cloud, where
[0026] Preferably, in step S1, oblique photography is performed on densely populated urban areas by optical remote sensing satellites, manned aircraft, or drones to take pictures or videos to obtain images of large-scale urban renewal cluster buildings; wherein, densely populated urban areas include: dense buildings, dense vehicles and pedestrian flows, and complex traffic.
[0027] The present invention also provides a rendering optimization system for ultra-large-scale urban renewal cluster buildings, comprising:
[0028] Acquisition module, used to obtain images of super-large-scale urban renewal cluster buildings;
[0029] A rendering optimization module is used to render images of large-scale urban renewal cluster buildings using 3D Gaussian splatting. The module uses a quantitative representation of 3D Gaussian volume properties and a coarse-to-fine progressive resolution training strategy.
[0030] As a preference, during the training phase, every 100 iterations, the 3D Gaussian point cloud with opacity less than the threshold φ is removed, i.e. Among them, φ=0.5.
[0031] As a preference, during the training phase, every 100 iterations, the segmentation position gradient is greater than the threshold Three-dimensional Gaussian point cloud, where
[0032] Preferably, optical remote sensing satellites, manned aircraft, and drones are used to take oblique photographs of densely populated urban areas to capture pictures or videos, thereby obtaining images of large-scale urban renewal cluster buildings; wherein, densely populated urban areas include: dense buildings, dense traffic of vehicles and people, and complex traffic.
[0033] In order to simultaneously solve the problems of neural radiation field and 3D Gaussian splashing, the present invention uses quantization properties to significantly reduce memory storage requirements, and a progressive resolution training strategy from coarse to fine to optimize 3D Gaussian point clouds faster and more stably.
[0034] This method uses fewer 3D Gaussian volumes and employs a quantitative method to represent properties, thereby improving training time and rendering speed for real-time rendering at full resolution. Specifically, compared to neural radiation fields, this method can improve rendering efficiency by approximately 50 times and save approximately 100 times the training time; compared to 3D Gaussian splatting, this method can save over 98% of memory resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0036] Figure 1 This is a flow chart of a rendering optimization method for super-large-scale urban renewal cluster buildings according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0038] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] Example 1:
[0040] 3D Gaussian splatting is a method for representing a 3D scene using a 3D Gaussian point cloud and quickly rendering it in a rasterization pipeline. A 3D Gaussian point cloud is also a point cloud, initialized with a sparse point cloud. Unlike a regular point cloud, each point in a 3D Gaussian point cloud is represented by a 3D Gaussian volume. A 3D Gaussian volume is an elliptical 3D spatial structure composed of multiple attributes, such as position, scale, rotation, color, and opacity. During rendering, the anisotropic volume of the 3D Gaussian volume is projected onto a 2D plane, i.e., the 3D Gaussian volume is affine mapped onto the 2D Gaussian plane. A differentiable tile-based rasterizer is then used to blend the different 3D Gaussian volumes together, and finally, gradient descent is used to further optimize the attributes. In short, 3D Gaussian splatting is a method for rasterization rendering using a 3D Gaussian point cloud.
[0041] Three-dimensional Gaussian expression: in is the position vector of a point in three-dimensional space, represents the three-dimensional covariance matrix, and the symbol T represents the transpose operator.
[0042] For a three-dimensional covariance matrix, its definition is Where x, y, and z are points in space The coordinates of , σ represents variance, and Cov represents covariance. The initialization method of the three-dimensional covariance matrix is: Where S represents the three-dimensional scaling vector and R represents the four-dimensional rotation vector.
[0043] Will Affine mapping to two dimensions gives Where W represents the projection transformation matrix from world coordinates to camera coordinates, which mainly performs translation and rotation, both of which are affine transformations. However, the projection transformation is non-affine, so the Jacobian matrix J is used instead of the projection, which acts as an affine approximation of the projection transformation.
[0044] The final color C of the pixel is calculated by the pixel contribution covered by N 2D Gaussian regions k , Among them C k Indicates the color of the k-th pixel, k represents the screen pixel index, and N represents N 2D Gaussian regions covering the k-th pixel. represents the opacity of the i-th 2D Gaussian region affecting the color of the k-th pixel, Represents the color of the i-th 2D Gaussian that influences the color of the k-th pixel.
[0045] The loss function of 3D Gaussian splashing includes: total loss function, reconstruction loss function and structural similarity loss function; among them, the total loss function is: Among them, the coefficient of the loss function is λ = 0.2; the reconstruction loss function By comparing the color C of the kth pixel k The true color of the k-th pixel The L1 loss of ; the structural similarity loss function is:
[0046]
[0047] Among them, I1 and I2 are two input images, and are the average values of I1 and I2 respectively, and are the standard deviations of I1 and I2, is the covariance of I1 and I2, C1 and C2 are constants used to stabilize the calculation. Structural similarity loss is used to measure the structural similarity between two images. The closer the value is to 1, the more similar the two images are, and the closer the value is to 0, the less similar the two images are.
[0048] like Figure 1 As shown, an embodiment of the present invention provides a rendering optimization method for a super-large-scale urban renewal cluster building, comprising:
[0049] Step S1, obtaining an image of a super-large-scale urban renewal cluster building;
[0050] Step S2: Render the image using three-dimensional Gaussian splashing based on the image of the super-large-scale urban renewal cluster buildings; wherein, a quantized method is used to represent the properties of the three-dimensional Gaussian volume, and a progressive resolution training strategy from coarse to fine is adopted.
[0051] As an implementation method of an embodiment of the present invention, in step S1, optical remote sensing satellites, manned aircraft, and drones are used to take oblique photography pictures or videos of densely populated urban areas (dense buildings, dense vehicles and pedestrians, and complex traffic) to obtain images of ultra-large-scale urban renewal cluster buildings.
[0052] As an implementation method of an embodiment of the present invention, the quantitative method is used to represent the properties of a three-dimensional Gaussian volume as follows:
[0053] A set of properties of a 3D Gaussian body Specifically include:
[0054] Position vector:
[0055] Scaling factor:
[0056] Rotate a quaternion vector:
[0057] Opacity
[0058] Spherical harmonic coefficients Where d = 3f 2 , f represents the harmonic degree. For the spherical harmonic function with f = 5, its color coefficient accounts for more than 80% of the dimension of the entire attribute vector; among them, the number of spherical harmonic function coefficients is d = 3f 2 =3×5 2 =75, the number of other parameters is position 3 + scale 3 + rotation 4 + opacity 1, a total of 75+11=86 dimensions, so 75 / 86=87%.
[0059] The quantitative representation method is as follows:
[0060] The core idea of quantization is to use 16-bit int type integers to represent 64-bit float type floating point numbers;
[0061] For attribute collections Any attribute in , a l-dimensional quantized latent vector consisting of integers Input to the multilayer perceptron decoder In order to obtain k-dimensional attributes;
[0062] because Is not differentiable. During the training process, a straight-through estimator (STE) is used to approximate the continuous Round it up and then pass it directly to the gradient during the backward pass, i.e. in is the desired k-dimensional attribute, yes Continuous approximation of ;
[0063] Property Collection Some attributes in the initialization are very sensitive, and quantization will lead to a significant performance degradation, so these parameters are not quantized (including and ), only for and SH were quantified.
[0064] Quantified benefits:
[0065] 1. Converting 64-bit floating-point numbers to 16-bit integers can significantly reduce the memory space occupied by the data, reducing the overall memory usage by about 70%. For each dimension of a floating-point number that used to occupy 8 bytes, only 2-byte integers are needed. All attributes have a total of 86 dimensions, 80 of which are quantized. Therefore, the memory usage before processing is 8*86=688 bytes, and after processing is 2*80+8*6=208 bytes, which is 1-208 / 688=69.76%.
[0066] 2. By using quantized integer operations instead of floating-point operations, the computing efficiency can be greatly improved, and the overall computing efficiency is increased by about 10 times.
[0067] 3. Quantifying opacity can remove artifacts when rendering new views to a certain extent, thereby improving rendering quality.
[0068] 4. In addition, f = 2 can be set to reduce the dimension of the attributes in the three-dimensional Gaussian. If all parameters are quantized, the memory usage can be greatly reduced.
[0069] As an implementation method of an embodiment of the present invention, in the prior art, the training process of the three-dimensional Gaussian point cloud is to calculate the loss function on the entire input image, which will result in low training efficiency and will produce floating artifacts. The three-dimensional Gaussian point cloud is initialized with a sparse point cloud, and some attributes of the three-dimensional Gaussian body are also roughly estimated, so the initialization of the three-dimensional Gaussian splash is not accurate. Therefore, in the early stage of training, forcibly forcing the three-dimensional Gaussian point cloud to adapt to the fine features of the entire scene will lead to floating artifacts, and these artifacts cannot be eliminated during the optimization process. In order to solve the floating artifact problem in the prior art, the present invention adopts a progressive resolution training strategy from coarse to fine, as follows:
[0070] Downsampling the original input image produces images with different levels of detail. For example, if the original input image has a resolution of 512*512, six coarse-resolution images are obtained through downsampling, with resolutions of 256*256, 128*128, 64*64, 32*32, 16*16, and 8*8, respectively.
[0071] Training is done batch by batch from coarse resolution to full resolution, from 8*8 to 512*512, for a total of 7 batches. The first 6 batches are trained for 10,000 iterations each, and the last batch is trained for 40,000 iterations, for a total of 100,000 training iterations.
[0072] Benefits of progressive resolution training:
[0073] 1. Training time: Progressive resolution training can use fewer 3D Gaussian volumes in low-resolution scenes, speeding up training and significantly reducing the training time for low-resolution scenes. Given a certain total number of training times, the overall training time can be significantly reduced.
[0074] 2. Convergence speed: Progressive resolution training can make it easier for 3D Gaussian point clouds to converge to the minimum loss value, achieve faster rendering and backpropagation speed, provide better initialization values for the densification process of 3D Gaussian volumes, reduce the amount of computation per training iteration, and accelerate the convergence of training.
[0075] 3. Reconstruction quality: Progressive resolution training can better adapt to input data of different resolutions, reduce blur and noise during training, remove artifacts caused by rasterization rendering of poor 3D Gaussian volumes, reconstruct the fine features of the scene, and improve the scene reconstruction quality and image rendering effect.
[0076] 4. Stability: Progressive resolution training can provide a more stable optimization process by gradually increasing the resolution of the scene.
[0077] As an implementation method of an embodiment of the present invention, three-dimensional Gaussian densification is the process of copying and splitting a three-dimensional Gaussian body. A large three-dimensional Gaussian body will be divided and a small three-dimensional Gaussian body will be copied to better adapt to the geometric shape. Too many times of densification will lead to an explosive growth of three-dimensional Gaussian bodies. Generally speaking, the more three-dimensional Gaussian bodies there are, the better the effect, and the more accurate the scene details that can be represented, but redundant three-dimensional Gaussian bodies will be added, thereby increasing the number of training, rendering time and memory usage. The present invention adopts four methods to reduce the three-dimensional Gaussian densification process:
[0078] 1. The current scheme is to perform densification every 100 iterations, and the setting is to perform densification every 500 iterations.
[0079] 2. Increase the threshold , increasing the value from 0.5 to 1.
[0080] 3. Increase the density interval and adopt a strategy that combines importance sampling, random sampling, and uniform sampling. That is, distribute more three-dimensional Gaussian bodies within the foreground bounding box and maintain a basically uniform distribution to make the density interval more moderate, while using random distribution in other areas. Among them, the number of three-dimensional Gaussian bodies N1 within the foreground bounding box and the number of three-dimensional Gaussian ellipsoids N2 in the background are converted according to their areas S1 and S2. That is, the proportional area of the three-dimensional Gaussian ellipsoids in the background is N1*S2 / S1. The desired result is N1>N1*S2 / S1. If N1≤N1*S2 / S1, the three-dimensional Gaussian body in the foreground bounding box can be further split from 1 to 2, and so on, until N1>N1*S2 / S1 is satisfied.
[0081] 4. During the training phase, every 100 iterations, the segmentation position gradient is greater than the threshold Three-dimensional Gaussian point cloud, where During the training phase, every 100 iterations, the opacity less than the threshold is removed. The three-dimensional Gaussian point cloud of in
[0082] Existing 3D Gaussian splatting methods use 64-bit floating-point representations for attribute parameters. For ultra-large-scale virtual scenes, with 5 billion 3D Gaussian point clouds (5 billion refers to the 5 billion 3D Gaussian volumes in the 3D Gaussian point cloud), and f = 5 for each volume, the memory usage is 5 billion * 8 * 86 / 1024 / 1024 / 1024 = 3203.75 GB. By combining the optimization strategies of "reducing the number," "controlling densification," and "quantizing the attributes," the 5 billion 3D Gaussian volumes can be reduced to 1 billion. Setting f = 2 reduces the attribute dimension and the memory usage of the attribute dimension. Specifically, 1 billion * (23 * 2) / 1024 / 1024 / 1024 = 42.84 GB, reducing memory usage by 98.66%, a 74.78-fold reduction. If f = 2, SH has 12 parameters and the other 11 parameters, for a total of 23 parameters. Six parameters, including position and scale, are not quantized, while the remaining 17 parameters are quantized. After quantization, 2-byte ints can be used instead of 8-byte floats, resulting in a total of 17*2+6*8=82 parameters. If further optimization is performed and all parameters are quantized, the total number of parameters becomes 23*2=46. Using a progressive resolution strategy can significantly accelerate training and improve reconstruction quality.
[0083] Example 2:
[0084] An embodiment of the present invention further provides a rendering optimization system for ultra-large-scale urban renewal cluster buildings, comprising:
[0085] Acquisition module, used to obtain images of super-large-scale urban renewal cluster buildings;
[0086] A rendering optimization module is used to render images of large-scale urban renewal cluster buildings using 3D Gaussian splatting. The module uses a quantitative representation of 3D Gaussian volume properties and a coarse-to-fine progressive resolution training strategy.
[0087] As an implementation method of the present invention, during the training phase, every 100 iterations, the three-dimensional Gaussian point cloud with opacity less than a threshold value φ is removed, that is, Among them, φ=0.5.
[0088] As an implementation method of the embodiment of the present invention, in the training stage, every 100 iterations, the segmentation position gradient is greater than the threshold Three-dimensional Gaussian point cloud, where
[0089] As an implementation method of an embodiment of the present invention, optical remote sensing satellites, manned aircraft, and drones are used to take oblique photographs of densely populated urban areas to capture pictures or videos, thereby obtaining images of ultra-large-scale urban renewal cluster buildings; wherein, densely populated urban areas include: dense buildings, dense vehicles and pedestrians, and complex traffic.
[0090] Example 3:
[0091] An embodiment of the present invention also provides a rendering optimization system for super-large-scale urban renewal cluster buildings, comprising: a memory and a processor, wherein the memory stores a computer program run by the processor, and when the computer program is run by the processor, it executes a rendering optimization method for super-large-scale urban renewal cluster buildings.
[0092] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A rendering optimization method for large-scale urban renewal cluster buildings, characterized by: include: Acquire images of large-scale urban renewal cluster buildings; Based on images of large-scale urban renewal cluster buildings, the image is rendered using 3D Gaussian splatting. A quantized representation of 3D Gaussian volume properties is used, while a coarse-to-fine resolution training strategy is employed. The quantitative method used to represent the properties of the three-dimensional Gaussian volume is as follows: A set of properties of a 3D Gaussian body include: Position vector: Scaling factor: Rotate a quaternion vector: Opacity Spherical harmonic coefficients where d = 3f 2 , f represents the harmonic degree; Quantization uses 16-bit int type integers to represent 64-bit float type floating point numbers; For attribute collections Any attribute in , a l-dimensional quantized latent vector consisting of integers Input to the multilayer perceptron decoder In , we get k-dimensional attributes; During the training process, a straight line estimator is used to approximate the continuous Take the integer and pass it directly to the gradient during the reverse transfer process, that is, in, is the desired k-dimensional attribute, yes Continuous approximation of ; Only for o and SH were quantified; The specific training strategy of using progressive resolution from coarse to fine is as follows: The ultra-large-scale urban renewal cluster building image is downsampled to generate resolution images with different levels of detail. Six coarse resolution images are obtained, with resolutions of 256*256, 128*128, 64*64, 32*32, 16*16, and 8*8 respectively. Training is done batch by batch from coarse resolution to full resolution; from 8*8 to 512*512, a total of 7 batches, of which the first 6 batches are trained for 10,000 iterations each, and the last batch is trained for 40,000 iterations, for a total of 100,000 training iterations; During the training phase, every 100 iterations, the 3D Gaussian point cloud with opacity less than the threshold φ is removed, i.e. Among them, φ = 0.5; in the training stage, every 100 iterations, the segmentation position gradient is greater than the threshold Three-dimensional Gaussian point cloud, where In step S1, optical remote sensing satellites, manned aircraft, and drones are used to take oblique photographs of densely populated urban areas to capture images or videos, thereby obtaining images of large-scale urban renewal cluster buildings. Urban dense areas include: dense buildings, dense vehicles and people flow, and complex traffic.
2. A rendering optimization system for super-large-scale urban renewal cluster buildings that implements the rendering optimization method for super-large-scale urban renewal cluster buildings according to claim 1, characterized in that: include: Acquisition module, used to obtain images of super-large-scale urban renewal cluster buildings; A rendering optimization module, which uses 3D Gaussian splatting to render images of large-scale urban renewal cluster buildings. This module uses a quantitative representation of 3D Gaussian volume properties and employs a coarse-to-fine resolution training strategy. During the training phase, every 100 iterations, the 3D Gaussian point cloud with opacity less than the threshold φ is removed, i.e. Among them, φ = 0.5; in the training stage, every 100 iterations, the segmentation position gradient is greater than the threshold Three-dimensional Gaussian point cloud, where By using optical remote sensing satellites, manned aircraft, and drones to take oblique photos or videos of densely populated urban areas, we can obtain images of large-scale urban renewal cluster buildings; among them, densely populated urban areas include: dense buildings, dense vehicles and pedestrians, and complex traffic.
Citation Information
Patent Citations
Novel view angle synthesis method based on Gaussian splash and fusing learnable basis function
CN118505541A