Accelerated rendering method for three-dimensional Gaussian splash and pipeline architecture
By dividing the imaging planes in the 3DGS algorithm and sorting and rendering in batches, the time-consuming problem of Gaussian ellipsoid sorting on mobile devices is solved, and the effect of low power consumption and real-time rendering is achieved.
Patent Information
- Application Number
- CN202510148506.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
The 3DGS algorithm needs to sort millions of Gaussian ellipsoids during rendering, resulting in excessive computing time and resource consumption, making it difficult to realize real-time rendering on mobile devices such as VR glasses and mobile phones.
Reduce unnecessary sorting and rendering calculations by dividing the two-dimensional imaging plane into smaller planes, find out the n-quantile distance between the two-dimensional Gaussian ellipse and the imaging plane in each plane, and divide it into n batches for sorting and alpha mixed rendering.
It greatly reduces the sorting and rendering calculation of Gaussian ellipsoids, improves rendering speed, and achieves low-power and real-time rendering effects, suitable for mobile devices.
Smart Images

Figure CN120070699A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional scene reconstruction method and its pipeline architecture, belonging to the fields of graphics rendering and algorithm hardware accelerators. Background Art
[0002] 3D Gaussian Splatting (3DGS) is an advanced three-dimensional scene reconstruction and rendering technology. It represents points in a scene by using a three-dimensional Gaussian function as an ellipsoid (as shown on the right side of [FIGURE], which is a simple scene composed of 3 Gaussian ellipsoids). Each three-dimensional Gaussian ellipsoid is described by spatial position, covariance matrix (describing the shape of the Gaussian ellipsoid), color, and opacity parameters. During rendering, these Gaussian function ellipsoids are projected onto a 2D image plane for rendering, forming elliptical "splatting". Multiple overlapping splatting are blended through alpha blending (when different ellipsoids overlap on the two-dimensional plane, the alpha blending algorithm is used to determine the final color value of the overlapping part) to obtain the final pixel color (such as the process from right to left in [FIGURE]), thereby achieving high-quality and high-efficiency three-dimensional scene reconstruction and new view synthesis, mainly divided into four steps: Figure 1 as shown Figure 1 On the right side of [FIGURE], there is a simple scene composed of 3 Gaussian ellipsoids. Each three-dimensional Gaussian ellipsoid is described by spatial position, covariance matrix (describing the shape of the Gaussian ellipsoid), color, and opacity parameters. During rendering, these Gaussian function ellipsoids are projected onto a 2D image plane for rendering, forming elliptical "splatting". Multiple overlapping splatting are blended through alpha blending (when different ellipsoids overlap on the two-dimensional plane, the alpha blending algorithm is used to determine the final color value of the overlapping part) to obtain the final pixel color (such as the process from right to left in [FIGURE]), thereby achieving high-quality and high-efficiency three-dimensional scene reconstruction and new view synthesis, mainly divided into four steps: Figure 1 From right to left), thus realizing high-quality and high-efficiency three-dimensional scene reconstruction and new view synthesis, mainly divided into four steps:
[0003] First step: According to the position and orientation parameters of the given two-dimensional imaging plane, convert the three-dimensional Gaussian ellipsoids in the model into two-dimensional Gaussian ellipses;
[0004] Second step: Divide the two-dimensional imaging plane (such as 1920×1080 corresponding to 1080p pixels) into smaller planes (tiles, such as 16×16). Each tile will select the two-dimensional Gaussian ellipses that intersect with it. At this time, dividing the tiles is for subsequent parallel computing;
[0005] Third step: Sort the two-dimensional Gaussian ellipses of each tile according to their distances from the imaging plane from near to far;
[0006] Fourth step: Perform parallel alpha blending rendering on the tiles.
[0007] The above 3DGS algorithm comes from the paper "3D Gaussian Splatting for Real-Time Radiance Field Rendering", published in the ACM Transactions on Graphics journal in July 2023. The project website of this algorithm is: https: / / repo-sam.inria.fr / fungraph / 3d-gaussian-splatting / , and the article website of this algorithm is: https: / / repo-sam.inria.fr / fungraph / 3d-gaussian-splatting / 3d_gaussian_splatting_low.pdf 。
[0008] There are a large number of Gaussian ellipsoids in the scene represented by the 3DGS algorithm. During the rendering process, since the alpha blending step among them has a sequential order for Gaussian ellipsoids, all the ellipsoids need to be sorted according to their distances from the two-dimensional plane before alpha blending can be performed. Specifically, the weight of a Gaussian ellipsoid closer to the two-dimensional plane is larger than that of a farther one, and the value of the weight of the closer one determines the value of the weight of the farther one. There is a correlation between the two and they are not independent. Therefore, when rendering with 3DGS, millions of Gaussian ellipsoids need to be sorted, and this step (i.e., the third step mentioned above) will consume a large amount of computing time and resources. Currently, there are some 3DGS algorithms based on software acceleration that can achieve real-time rendering of 1080p images (at a speed of greater than or equal to 30 frames per second) on computers at the PC and server levels, using desktop-class and above graphics cards. However, for mobile devices such as VR glasses and mobile phones, the graphics cards they carry cannot provide sufficient computing power, making it impossible for the 3DGS algorithm to achieve real-time rendering on these mobile devices. Summary of the Invention
[0009] The technical problem to be solved by the present invention is that when rendering with the 3DGS algorithm, millions of Gaussian ellipsoids need to be sorted, which will consume a large amount of computing time and resources. For mobile devices such as VR glasses and mobile phones, the chips they carry cannot provide sufficient computing power for real-time rendering of the 3DGS algorithm. At the same time, due to the limitations of the circuit designs of existing computing units (such as CPUs and GPUs), it is very difficult to further improve the rendering speed of the 3DGS algorithm without significantly increasing indicators such as the power consumption, area, and memory of the computing units.
[0010] To solve the above technical problems, one aspect of the present invention discloses an accelerated rendering method for three-dimensional Gaussian splash, which is characterized by including the following steps:
[0011] Step 1: Convert the target three-dimensional Gaussian ellipsoid into a two-dimensional Gaussian ellipse;
[0012] Step 2: Divide the two-dimensional imaging plane into smaller planes. For each plane, obtain all the two-dimensional Gaussian ellipses intersecting with it, and find the n-th quantiles of the distances between each two-dimensional Gaussian ellipse in each plane and the two-dimensional imaging plane;
[0013] Step 3: Sort all the two-dimensional Gaussian ellipses in each plane into n batches in ascending order based on their distances from the two-dimensional imaging plane according to the n-th quantiles. After the current batch of two-dimensional Gaussian ellipses in the current plane is sorted, immediately enter Step 4;
[0014] Step 4: Perform alpha blending rendering on the current batch of two-dimensional Gaussian ellipses in the plane;
[0015] Step 5: If the opacity of the current plane reaches the requirement, skip all the remaining two-dimensional Gaussian ellipses in the current plane;
[0016] If the opacity of the current plane does not reach the requirement, take the next batch of two-dimensional Gaussian ellipses of the current batch of two-dimensional Gaussian ellipses as the current batch of two-dimensional Gaussian ellipses, and return to Step 4 until the opacity of the current plane reaches the requirement.
[0017] Another aspect of the present invention discloses an accelerated rendering pipeline architecture for three-dimensional Gaussian splashes, which is used to implement the above-mentioned accelerated rendering method, and is characterized by including:
[0018] A segmentation module for obtaining the n quantiles in Step 2. The segmentation module divides a data group with a quantity of k into multiple data subgroups, performs a full permutation on all the data in each data subgroup to obtain the n quantiles of each data subgroup, then performs a full permutation on all the n quantiles of all the data subgroups, and takes the median as the estimated value of the n quantiles of the original data group with a quantity of k;
[0019] A rendering module that integrates the sorting in Step 3 and the alpha blending rendering in Step 4.
[0020] Preferably, the segmentation module includes a hardware sorting circuit for performing a full permutation on all the data in each data subgroup.
[0021] Preferably, the segmentation module includes a median module for performing a full permutation on all the n quantiles of all the data subgroups and taking the median.
[0022] Preferably, the median module is composed of n multi-level sorting queues. Among them, the j-th multi-level sorting queue is used for sorting the j-th quantile.
[0023] Preferably, the multi-level sorting queue is composed of several sorting units. Each sorting unit contains a register and a selector, where:
[0024] The register in the i-th level sorting unit stores the current i-th largest value;
[0025] The selector is used to update the value in the register of the current sorting unit. According to the selection signal generated by the internal circuit of the sorting unit, the selector can select among the following three values:
[0026] The first kind: the value of the register in the i-th level itself;
[0027] The second kind: the newly input value;
[0028] Third, the value of the (i - 1)-th level register;
[0029] When the n quantiles of a set of data enter the median module, the j-th quantile value enters the j-th multi-level sorting queue for sorting; for the j-th multi-level sorting queue, when a new value is input, each sorting unit contained therein compares the value stored in its own register with the newly input value and generates a corresponding selection signal, and the selection signal will select the corresponding value to complete the update of the sorting unit register.
[0030] Preferably, the operation of the rendering module is divided into Process 1 and Process 2, where:
[0031] In Process 1, the current batch of two-dimensional Gaussian ellipses is read, and the color components and related parameters of the current batch of two-dimensional Gaussian ellipses on different pixels in the current plane are calculated respectively, and then the arrangement of the current batch of two-dimensional Gaussian ellipses is completed;
[0032] In Process 2, after the current batch of two-dimensional Gaussian ellipses is read and arranged, the alpha renderer immediately performs alpha blending rendering on the current batch of two-dimensional Gaussian ellipses.
[0033] Preferably, in Process 1, the arrangement of the two-dimensional Gaussian ellipses is completed through a multi-level sorting array.
[0034] Preferably, in Process 1, if the number of the current batch of two-dimensional Gaussian ellipses exceeds the depth of the multi-level sorting array, they are temporarily stored in the overflow cache.
[0035] Preferably, after the rendering module completes the alpha blending rendering of the current batch of two-dimensional Gaussian ellipses, the two-dimensional Gaussian ellipses in the overflow cache are used as the input of the multi-level sorting array to complete Process 1 and Process 2. After the overflow cache is emptied, the next batch of two-dimensional Gaussian ellipses is used as the input of the multi-level sorting array to complete Process 1 and Process 2.
[0036] The present invention is a dedicated acceleration operation pipeline architecture designed for the 3DGS algorithm. This architecture mainly uses a unique hardware circuit design to optimize and accelerate the "chunking", "sorting" and "alpha blending" parts in the 3DGS algorithm, so that the rendering speed of 3DGS can be greatly improved, its computing efficiency can be increased, and thus the effects of low power consumption and real-time rendering can be achieved. The present invention combines algorithm design with hardware design. The algorithm part greatly reduces the operation of sorting Gaussian ellipses, and the hardware part efficiently implements the calculation process of the algorithm, which can achieve the effect of reducing power consumption.
[0037] The architecture disclosed by the present invention can be implemented based on FPGA or programmable GPU, or can be used as a soft / hard IP core in chip design to fabricate an ASIC chip. The main advantages are low power consumption and real-time rendering. It is mainly targeted at edge devices with requirements for low power consumption and small area, and can also be used for server or desktop devices to achieve the goal of reducing power consumption and improving the operation efficiency of the 3DGS algorithm. Description of the Drawings
[0038] Figure 1 It is a flowchart for 3DGS rendering;
[0039] Figure 2 It schematically shows the processing flow of the segmentation module;
[0040] Figure 3 It schematically shows the hardware structure of the median module;
[0041] Figure 4 It schematically shows the hardware structure and process of the rendering module. Detailed Embodiments
[0042] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0043] In the original 3DGS rendering step, usually in the third step, several million Gaussian ellipses need to be sorted, which is the main bottleneck in the calculation of the entire algorithm. The technical solution disclosed in the embodiments of the present invention mainly optimizes the third step. At the same time, in order to cooperate with the optimization, the second and fourth steps are also improved. The main improvements are as follows:
[0044] For the second step, after selecting the two-dimensional Gaussian ellipses that intersect with each tile, then find the n-th quantile of the distances between all the two-dimensional Gaussian ellipses in each tile and the imaging plane.
[0045] For the third step, each tile is sorted according to the n-th quantile of the distances between all the two-dimensional Gaussian ellipses obtained in the second step and the imaging plane. All the two-dimensional Gaussian ellipses are divided into n batches from near to far and enter the sorting unit. After each batch of two-dimensional Gaussian ellipses is sorted, the fourth-step alpha blending rendering is immediately performed. For example, when rendering the first tile, first read 1 / n of the two-dimensional Gaussian ellipses closest to the imaging plane into the sorting unit for sorting, and immediately perform alpha blending rendering on these Gaussian ellipses after sorting; if the opacity of the current tile reaches the requirement, the remaining (n - 1) / n two-dimensional Gaussian ellipses can be skipped. If the opacity of the current tile has not reached the requirement, continue to read 1 / n to 2 / n of the remaining two-dimensional Gaussian ellipses and repeat the previous steps. Such a rendering process greatly reduces the sorting and rendering calculations of unnecessary Gaussian ellipses, thus improving the rendering speed.
[0046] To support the above new rendering process, the embodiment of the present invention also designs a corresponding dedicated hardware architecture.
[0047] For the above second step, the embodiment of the present invention designs a "segmentation module". This segmentation module can estimate the n-th quantile of a set of data with less calculation. To find the n-th quantile of a set of data with a quantity of k, usually, the k data are fully arranged and then the corresponding position data is taken. However, this algorithm requires a large amount of data sorting. As Figure 2 shown, the segmentation module disclosed in the embodiment of the present invention first divides the k data into multiple batch groups, each batch group has b data, so there are k / b batch groups in total. Each batch group performs a full arrangement using a "hardware sorting circuit" and finds its n-th quantile. Among them, the "hardware sorting circuit" can be implemented using a common hardware sorting network. Then, the quantiles of the same n-th quantile of different batch groups are fully arranged in the "median module" and the median is taken as the estimated value of the n-th quantile of the original data group with a quantity of k. Among them, the hardware structure of the "median module" is implemented as Figure 3 shown. The median module consists of n multi-stage sorting queues. Among them, the j-th multi-stage sorting queue is used to sort the j-th quantile. Each multi-stage sorting queue is composed of several sorting units. As Figure 3 shown in the lower right corner of , the sorting unit contains a register, and the register in the i-th stage sorting unit stores the current i-th largest value. The sorting unit also contains a selector for updating the value in the register of the current sorting unit. The selector can select among (1) the value of the i-th stage register itself, (2) the newly input value, and (3) the value of the (i - 1)-th stage register according to the selection signal generated by the internal circuit of the sorting unit.
[0048] When the n n - quantiles of a set of data enter the median module, the j - th quantile value enters the j - th multi - level sorting queue for sorting. For the j - th multi - level sorting queue, when a new value is input, each sorting unit it contains will compare the value stored in its own register with the newly input value and generate a corresponding selection signal. The selection signal will select the corresponding value to complete the update of the sorting unit register. For example, if the value stored in the register of the i - th level sorting unit is smaller than the input value, the selection signal will select the current value and keep it unchanged; if the value of the i - th level is larger than the input value and the value of the i - 1 - th level is smaller than the input value, the selection signal will select the input value, and the value stored in the register of this level unit will be updated to the input value; if the value of the i - th level is larger than the input value and the value of the i - 1 - th level is also larger than the input value, the selection signal will select the value of the i - 1 - th level, and the value stored in the register of this level unit will be updated to the previous level, that is, the value of the i - 1 - th level.
[0049] When all the data is input, the multi - level sorting queue can complete the ascending sorting of all the input data in the next cycle. This algorithm converts the original full permutation of k data into the full permutation of k / b groups of b data plus the full permutation of n groups of k / b data. By converting a large - scale data sorting into multiple small - scale data permutations, the present invention enables the sorting process to be implemented by a hardware circuit, thus achieving the purpose of acceleration.
[0050] For the third and fourth steps, the embodiment of the present invention designs a dedicated "rendering module". This rendering module integrates the sorting function in the third step and the alpha - blending rendering function in the fourth step. The work of the rendering module is divided into two processes. When it renders a tile, as Figure 4 shown in Process 1, according to the n - quantile value calculated in the second step, first read 1 / n two - dimensional Gaussian ellipses from all the two - dimensional Gaussian ellipses near the rendering imaging plane, calculate the color components and related parameters of this two - dimensional Gaussian ellipse on different pixels in this tile respectively, and then enter the "multi - level sorting array" similar to the "median module" to complete the arrangement of the two - dimensional Gaussian ellipses. If the number of two - dimensional Gaussian ellipses in a certain pixel exceeds the depth of the multi - level sorting array, they will be temporarily stored in the "overflow cache". After the current 1 / n two - dimensional Gaussian ellipses are read and arranged, the alpha renderer will immediately perform alpha - blending rendering on the data in the multi - level sorting array, as Figure 4As shown in Process 2. After rendering, the color value of the pixel will be obtained. If there is no "overflow" situation, after the current batch of two-dimensional Gaussian ellipses is rendered, check whether the opacity value of each pixel in the tile meets the requirements. If there is an "overflow" situation, it means that the number of Gaussian ellipses in a batch is too large and exceeds the number that the multi-level sorting array can accommodate. At this time, the overflow Gaussian ellipses are temporarily stored in the "overflow cache". When the Gaussian ellipses in the multi-level sorting array are completed with alpha rendering, the overflow Gaussian ellipses are taken out from the overflow cache and re-sorted and alpha-rendered through the multi-level sorting array. At this time, when the alpha rendering of the overflow Gaussian ellipses is performed, further calculations need to be carried out on the basis of the pixel values that have been rendered before, that is, the obtained pixel color value is used as the initial value and input into the "multi-level sorting array". That is, if there is an overflow of pixels in Process 1, continue to perform the operations of Process 1 and 2 until the "overflow cache" is emptied. At this time, the two-dimensional Gaussian ellipses stored in the "overflow cache" will be used as the input; after the "overflow cache" is emptied, then input the subsequent 1 / n Gaussian ellipsoids from the rendering imaging plane, and repeat the above steps until the current tile is completed with rendering.
Claims
1. A three-dimensional Gaussian splash accelerated rendering method, characterized in that: The following steps are involved: Step 1, converting the target three-dimensional Gaussian ellipsoid into a two-dimensional Gaussian ellipse; Step 2, divide the two-dimensional imaging plane into smaller planes, obtain all two-dimensional Gaussian ellipses intersecting each plane, and find the nth quantile of the distance between each two-dimensional Gaussian ellipse and the two-dimensional imaging plane in each plane; Step 3: sort all the two-dimensional Gaussian ellipses in each plane into n batches according to the n quantiles and based on the distance from the two-dimensional imaging plane, from near to far, and proceed to step 4 immediately after the current batch of two-dimensional Gaussian ellipses in the current plane is sorted; Step 4, perform alpha blending rendering on the current batch of two-dimensional Gaussian ellipses in the plane; Step 5: If the opacity of the current plane reaches the requirement, all remaining two-dimensional Gaussian ellipses in the current plane are skipped; If the opacity of the current plane does not meet the requirement, the next batch of two-dimensional Gaussian ellipses of the current batch of two-dimensional Gaussian ellipses is used as the current batch of two-dimensional Gaussian ellipses, and the process returns to step 4 until the opacity of the current plane meets the requirement.
2. A 3D Gaussian splash accelerated rendering pipeline architecture, used to implement the accelerated rendering method of claim 1, characterized in that: include: A segmentation module for obtaining the n-quantile in step 2, wherein the segmentation module divides a data group of k into a plurality of data subgroups, performs a full permutation on all data in each data subgroup to obtain the n-quantile of each data subgroup, and then performs a full permutation on all n-quantiles of all data subgroups, and takes the median thereof as an estimated value of the n-quantile of the original data group of k; A rendering module that integrates the sorting in step 3 and the alpha blending rendering in step 4.
3. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 2, characterized in that: The segmentation module includes a hardware sorting circuit for fully arranging all data in each data subgroup.
4. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 2, characterized in that: The segmentation module includes a median module for performing a full permutation of all n quantiles of all data subgroups and taking the median thereof.
5. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 4, characterized in that: The median module is composed of n multi-level sorting queues, wherein the j-th multi-level sorting queue is used to sort the j-th quantile.
6. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 5, characterized in that: The multi-level sorting queue is composed of a number of sorting units, each of which includes a register and a selector, wherein: The register in the i-th level sorting unit stores the current i-th largest value; The selector is used to update the value of the register in the current sorting unit. The selector can select from the following three values according to the selection signal generated by the internal circuit of the sorting unit: The first type is the value of the register at level i itself; The second type is new input value; The third type, the value of the register at level i-1; When the n quantiles of a set of data enter the median module, the j-th quantile value enters the j-th multi-level sorting queue for sorting; for the j-th multi-level sorting queue, when a new value is input, each sorting unit it contains will compare the value stored in its own register with the new input value and generate a corresponding selection signal. The selection signal will select the corresponding value to complete the update of the sorting unit register.
7. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 2, characterized in that: The work of the rendering module is divided into process one and process two, wherein: In process one, a current batch of two-dimensional Gaussian ellipses are read, color components and related parameters of the current batch of two-dimensional Gaussian ellipses at different pixels in the current plane are calculated respectively, and then the current batch of two-dimensional Gaussian ellipses are arranged; In process 2, after the current batch of two-dimensional Gaussian ellipses are read and arranged, the alpha renderer immediately performs alpha blending rendering on the current batch of two-dimensional Gaussian ellipses.
8. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 7, characterized in that: In the process 1, the arrangement of the two-dimensional Gaussian ellipses is completed through a multi-level sorting array.
9. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 8, characterized in that: In the process 1, if the number of the current batch of two-dimensional Gaussian ellipses exceeds the depth of the multi-level sorting array, they are temporarily stored in the overflow cache.
10. The accelerated rendering pipeline architecture of 3D Gaussian splashing according to claim 9, characterized in that: If the rendering module completes the alpha blending rendering of the current batch of two-dimensional Gaussian ellipses, the two-dimensional Gaussian ellipses in the overflow cache are used as the input of the multi-level sorting array to complete the process one and the process two. Until the overflow cache is cleared, the next batch of two-dimensional Gaussian ellipses are used as the input of the multi-level sorting array to complete the process one and the process two.