3D Gaussian point cloud multi-level rendering optimization method and device and electronic equipment
By obtaining multiple images of a three-dimensional scene to restore the three-dimensional structure information, dividing multi-level subspaces, and using Gaussian point representation formulas and LOD technology for rendering, the problem of low rendering efficiency caused by the huge number of Gaussian points in large scenes is solved, and efficient multi-level rendering is achieved.
Patent Information
- Application Number
- CN202510063704.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-27
AI Technical Summary
When the existing 3D Gaussian Splatting technology renders large scenes with rich details, the huge number of Gaussian points leads to uneven density distribution and low rendering efficiency.
By obtaining multiple images of a three-dimensional scene, the three-dimensional structure information is restored, the multi-level subspace is divided, and the Gaussian point representation formula is used to calculate the representation sequence of the spatial points of each layer, and converted it into a three-dimensional Gaussian distribution. Then, layer by layer Gaussian point initialization and texture sampling are performed, and hierarchical rendering is combined with LOD technology to optimize 3D Gaussian point cloud multi-level rendering.
The rendering efficiency of 3D Gaussian point cloud multi-level rendering in large scenes is improved, and the authenticity and efficiency of rendered images are improved through hierarchical division and texture sampling.
Smart Images

Figure CN120047589A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a 3D Gaussian point cloud multi-level rendering optimization method, device, and electronic device. Background Art
[0002] In recent years, the 3D Gaussian Splatting (abbreviated as 3DGS) technology has achieved revolutionary breakthroughs in the field of three-dimensional reconstruction, solving the problems of the traditional neural radiance field (NeRF) rendering technology, such as non-intuitive implicit expression, low controllability, low training efficiency, and difficulty in high-resolution real-time rendering.
[0003] Although the existing 3DGS technology has made great progress in real-time rendering, when applied to large scenes with rich details, such as large scenes with tens of millions to billions of 3D Gaussian points in urban streets or campus scenes, the large number of Gaussian points makes the density distribution of Gaussian points uneven and the rendering efficiency low. Summary of the Invention
[0004] The present invention provides a 3D Gaussian point cloud multi-level rendering optimization method, device, and electronic device, which can improve the rendering efficiency of 3D Gaussian point cloud multi-level rendering in large scenes.
[0005] To achieve the above object, a 3D Gaussian point cloud multi-level rendering optimization method provided by the present invention includes:
[0006] Obtain multiple images in a three-dimensional scene, and restore the three-dimensional structure information of the corresponding three-dimensional scene according to the multiple images;
[0007] Obtain the bounding box of the three-dimensional scene according to the three-dimensional structure information, calculate the maximum number of division layers of the bounding box according to the octree hierarchical division formula, and recursively divide the sub-space according to the positions of the spatial points to obtain multiple divided sub-spaces;
[0008] Calculate the representation sequence of the spatial points in each divided sub-space by using the Gaussian point representation formula, convert the center coordinates of the representation sequence in each divided sub-space into a three-dimensional Gaussian distribution, and perform layer-by-layer Gaussian point initialization to obtain a layer-by-layer Gaussian distribution;
[0009] Perform texture sampling on the original real image according to the maximum number of division layers to obtain multi-level real image information;
[0010] Project the layer-by-layer Gaussian distribution onto a two-dimensional plane by using layer-by-layer training, and obtain a rendered image through volume rendering. Calculate the loss by using the multi-level real image information and the rendered image, and perform 3D Gaussian point cloud multi-level rendering optimization based on the loss by using the backpropagation method. After the optimization is completed, a multi-level 3D Gaussian model is obtained;
[0011] Hierarchical rendering of a multi-level 3D Gaussian model using LOD technology.
[0012] Optionally, converting the center coordinates representing the sequence in each divided subspace into a three-dimensional Gaussian distribution includes:
[0013] Query the total number of spatial points in the representation sequence in each divided subspace, and calculate the center coordinates based on the total number of spatial points;
[0014] Convert the center coordinates of each divided subspace into a three-dimensional Gaussian distribution according to a preset Gaussian distribution conversion formula.
[0015] Optionally, converting the center coordinates of each divided subspace into a three-dimensional Gaussian distribution according to a preset Gaussian distribution conversion formula includes:
[0016] Calculate the center coordinate p according to the following formula:
[0017]
[0018] where N is the total number of spatial points, x i is the abscissa of the i-th spatial point, y i is the ordinate of the i-th spatial point, and z i is the coordinate value of the i-th spatial point on the Z-axis;
[0019] Convert the center coordinates of each divided subspace into a three-dimensional Gaussian distribution G(p) using the following formula:
[0020]
[0021] where ∑ is the covariance matrix, μ is the position of the spatial point in the three-dimensional Gaussian distribution, and (·) T is the transpose operation.
[0022] Optionally, texture sampling of the original real image according to the maximum number of division levels includes:
[0023] Obtain the size information and original resolution of the original real image;
[0024] Divide the resolution of each layer of the real image according to the size information, original resolution, and maximum number of division levels, and perform texture sampling according to the hierarchical index of the maximum number of division levels to obtain multi-level real image information.
[0025] Optionally, the texture sampling according to the hierarchical index of the maximum number of division levels further includes: The hierarchical index starts from layer 0, and the original real image is downsampled by average filtering, where each layer of the real image is obtained by downsampling the previous layer of the real image until the real image at the maximum number of division levels is obtained.
[0026] Optionally, obtaining the rendered image through volume rendering includes:
[0027] Projecting the spatial points of each divided subspace onto a two-dimensional plane and performing volume rendering using the following rendering formula:
[0028]
[0029] where I(x) is the rendered image, N is the total number of spatial points, α i is the transparency of the i-th sampling point, C i is the color value of the i-th sampling point, is the cumulative transparency of all spatial points before the i-th sampling point.
[0030] Optionally, hierarchically rendering the multi-level 3D Gaussian model using the LOD technology includes:
[0031] Screening the LOD levels using the hierarchical screening formula and calculating the smoothing factor using the linear interpolation method;
[0032] Performing smoothing processing on different LOD levels according to the smoothing factor, calculating the mixed Gaussian points to be rendered, and performing hierarchical rendering based on the mixed Gaussian points to be rendered.
[0033] To solve the above problems, the present invention also provides a 3D Gaussian point cloud multi-level rendering optimization device, and the device includes:
[0034] A three-dimensional structure information acquisition module, configured to acquire multiple images in a three-dimensional scene and restore the three-dimensional structure information number of the corresponding three-dimensional scene according to the multiple images;
[0035] A subspace division module, configured to obtain the bounding box of the three-dimensional scene according to the three-dimensional structure information, calculate the maximum division layer number of the bounding box according to the octree hierarchical division formula, recursively divide the subspace according to the positions of the spatial points to obtain multiple divided subspaces; calculate the representation sequence of the spatial points in each divided subspace using the Gaussian point representation formula, and convert the central coordinates of the representation sequence in each divided subspace into a three-dimensional Gaussian distribution, perform layer-by-layer Gaussian point initialization to obtain layer-by-layer Gaussian distributions; perform texture sampling on the original real image according to the maximum division layer number to obtain multi-level real image information;
[0036] A multi-level rendering module, configured to project the layer-by-layer Gaussian distribution onto a two-dimensional plane using layer-by-layer training and obtain a rendered image through volume rendering, calculate the loss using the multi-level real image information and the rendered image, and perform 3D Gaussian point cloud multi-level rendering optimization based on the loss using the backpropagation method. After the optimization is completed, a multi-level 3D Gaussian model is obtained; hierarchically render the multi-level 3D Gaussian model using the LOD technology.
[0037] To solve the above problems, the present invention further provides an electronic device, which includes:
[0038] At least one processor; and,
[0039] A memory communicatively connected to the at least one processor; wherein,
[0040] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned 3D Gaussian point cloud multi-level rendering optimization method.
[0041] To solve the above problems, the present invention further provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned 3D Gaussian point cloud multi-level rendering optimization method.
[0042] The present invention can accurately extract three-dimensional geometric information by recovering the three-dimensional structure information of the corresponding three-dimensional scene and the camera parameters corresponding to each image from multiple images. In addition, by recursively dividing the subspaces according to the positions of the spatial points, multiple levels of divided subspaces can be obtained, which can be divided into different levels of detail and provide a basis for hierarchical division for subsequent LOD technology for multi-level rendering. In addition, by converting the central coordinates representing the sequence in each level of divided subspace into a three-dimensional Gaussian distribution and performing layer-by-layer Gaussian point initialization to obtain a layer-by-layer Gaussian distribution, multi-level point cloud representation can be completed through layer-by-layer Gaussian point initialization. And by performing texture sampling on the original real images according to the maximum number of divided levels, texture information at different levels of detail can be obtained, improving the authenticity of the rendered images. Finally, by using the LOD technology to perform hierarchical rendering on the multi-level 3D Gaussian model, the rendering efficiency can be improved through hierarchical rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic flow chart of the 3D Gaussian point cloud multi-level rendering optimization method provided by an embodiment of the present invention;
[0044] Figure 2 It is a functional module diagram of a 3D Gaussian point cloud multi-level rendering optimization device provided by an embodiment of the present invention;
[0045] Figure 3 It is a schematic structural diagram of an electronic device for implementing the above-mentioned 3D Gaussian point cloud multi-level rendering optimization method provided by an embodiment of the present invention.
[0046] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Implementation Manner
[0047] It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0048] An embodiment of the present application provides a 3D Gaussian point cloud multi-level rendering optimization method. The execution subject of the 3D Gaussian point cloud multi-level rendering optimization method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the 3D Gaussian point cloud multi-level rendering optimization method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0049] Referring to Figure 1 As shown, it is a schematic flowchart of a 3D Gaussian point cloud multi-level rendering optimization method provided by an embodiment of the present invention. In this embodiment, the 3D Gaussian point cloud multi-level rendering optimization method includes:
[0050] S1. Obtain multiple images under a three-dimensional scene, and restore the three-dimensional structure information of the corresponding three-dimensional scene according to the multiple images.
[0051] In an embodiment of the present invention, multiple images under a three-dimensional scene refer to multiple images taken from different perspectives in three-dimensional space.
[0052] In an embodiment of the present invention, restoring the three-dimensional structure information of the corresponding three-dimensional scene according to multiple images is implemented through the colmap algorithm. Among them, the colmap algorithm refers to an open-source multi-view Figure 3 dimensional reconstruction system, which is mainly used to restore the three-dimensional structure from multiple two-dimensional images.
[0053] S2. Obtain the bounding box of the three-dimensional scene according to the three-dimensional structure information, calculate the maximum number of division layers of the bounding box according to the octree hierarchical division formula, and recursively divide the sub-space according to the positions of the spatial points to obtain multiple divided sub-spaces.
[0054] In an embodiment of the present invention, the bounding box of the three-dimensional scene refers to the geometric framework that encloses the entire image.
[0055] In the embodiments of the present invention, the octree hierarchical division formula is a formula for calculating the maximum number of division layers of the bounding box. The maximum number of division layers L of the bounding box can be calculated using the following formula:
[0056] L = [log 2 B] + 1
[0057] where B is the coordinate range of the bounding box, and the bounding box is divided into multiple subspaces of division layers according to the maximum number of division layers.
[0058] In the embodiments of the present invention, the coordinate range of the bounding box is realized by querying the minimum and maximum values of the coordinates of all spatial points in the x-axis, y-axis, and z-axis in the space, and can be represented by min_x, min_y, min_z and max_x, max_y, max_z.
[0059] As an embodiment of the present invention, recursively dividing the subspaces according to the positions of the spatial points includes: taking the maximum number of division layers as the number of nodes of the octree; dividing the three-dimensional scene into multiple spaces equal to the number of nodes according to the bounding box; querying the bounding box of each space and the number and positions of the spatial points belonging to each space, and combining the bounding box of each space and the number and positions of the spatial points belonging to each space to obtain the divided subspaces. Among them, each node in the octree represents a divided subspace, and each node includes the bounding box of the current divided subspace and the number and positions of the spatial points in the divided subspace.
[0060] S3. Using the Gaussian point representation formula to calculate the representation sequence of the spatial points in each layer of the divided subspace, and converting the central coordinates of the representation sequence in each layer of the divided subspace into a three-dimensional Gaussian distribution, performing layer-by-layer Gaussian point initialization to obtain a layer-by-layer Gaussian distribution.
[0061] In the embodiments of the present invention, using the Gaussian point representation formula to calculate the representation sequence of the spatial points in each layer of the divided subspace can be calculated using the following formula:
[0062] M = P / 2 l
[0063] where M is the representation sequence, l is the number of layers of the divided subspace, and P is the position coordinate of the spatial point.
[0064] Furthermore, the representation sequence can be expressed as {P / 2 0 , P / 2 1 , …, P / 2 L-1}.
[0065] As an embodiment of the present invention, converting the central coordinates of the representation sequence in each layer of the divided subspace into a three-dimensional Gaussian distribution includes:
[0066] Query the total number of spatial points in the representation sequence in each layer of the divided subspace, and calculate the center coordinates based on the total number of spatial points;
[0067] Convert the center coordinates of each layer of the divided subspace into a three-dimensional Gaussian distribution according to the preset Gaussian distribution conversion formula.
[0068] Further, converting the center coordinates of each layer of the divided subspace into a three-dimensional Gaussian distribution according to the preset Gaussian distribution conversion formula includes:
[0069] Calculate the center coordinate p according to the following formula:
[0070]
[0071] where N is the total number of spatial points, x i is the abscissa of the i-th spatial point, y i is the ordinate of the i-th spatial point, and z i is the coordinate value of the i-th spatial point on the Z-axis;
[0072] Use the following formula to convert the center coordinates of each layer of the divided subspace into a three-dimensional Gaussian distribution G(p):
[0073]
[0074] where ∑ is the covariance matrix, μ is the position of the spatial point in the three-dimensional Gaussian distribution, and (·) T is the transpose operation.
[0075] In the embodiment of the present invention, converting the center coordinates of the representation sequence in each layer of the divided subspace into a three-dimensional Gaussian distribution, in addition to performing layer-by-layer Gaussian point initialization, further includes: assigning the attribute information color value c and opacity α to each spatial point p, where the color value c - fits the appearance related to the viewing angle through spherical harmonic functions, and uses a linear combination of a set of orthogonal bases to fit the light field; the opacity α - is used for rendering Splatting, and when it is projected onto the image plane, the diffusion traces are superimposed through the opacity.
[0076] S4. Perform texture sampling on the original real image according to the maximum number of division layers to obtain multi-level real image information.
[0077] In the embodiment of the present invention, texture sampling refers to a technique that can generate texture images with different resolutions.
[0078] As an embodiment of the present invention, performing texture sampling on the original real image according to the maximum number of division layers includes:
[0079] Obtain the size information and the original resolution of the original real image;
[0080] Divide the resolution of each layer of the real image according to the size information, the original resolution, and the maximum number of division layers, and perform texture sampling according to the hierarchical index of the maximum number of division layers to obtain multi-level real image information.
[0081] In the embodiment of the present invention, the size information is the size information of the original real image.
[0082] Further, performing texture sampling according to the hierarchical index of the maximum number of division layers further includes: the hierarchical index starts from layer 0, and the original real image is downsampled through average filtering, where each layer of the real image is obtained by downsampling the previous layer of the real image until the real image under the maximum number of division layers is obtained.
[0083] Exemplarily, according to the maximum number of division layers, texture sampling is performed on the original real image. The following implementation steps can be adopted:
[0084] Step 1: Extract an original real image I 0 , whose size is W×H, where W and H are the width and height of the image respectively;
[0085] Step 2: Define the maximum level of texture sampling as L levels for one-to-one correspondence during subsequent training;
[0086] Step 3: Extract the resolution of the original real image as 2 n ×2 n , where n is the resolution exponent, then the image resolution of the I-th layer is where I is the hierarchical index, starting from 0. Each layer of image I i is obtained by downsampling the previous layer of image I i-1 , and the downsampling process usually uses average filtering to reduce high-frequency information.
[0087] S5. Project the layer-by-layer Gaussian distribution onto the two-dimensional plane by progressive hierarchical training, and obtain the rendered image through volume rendering. Calculate the loss using the multi-level real image information and the rendered image, and perform 3D Gaussian point cloud multi-level rendering optimization based on the loss using the backpropagation method. After the optimization is completed, a multi-level 3D Gaussian model is obtained.
[0088] As an embodiment of the present invention, obtaining the rendered image through volume rendering includes:
[0089] Project the spatial points of each divided subspace onto the two-dimensional plane, and perform volume rendering using the following rendering formula:
[0090]
[0091] where I(x) is the rendered image, N is the total number of spatial points, and α iis the transparency of the i-th sampling point, C i is the color value of the i-th sampling point, is the cumulative transparency of all spatial points before the i-th sampling point.
[0092] As an embodiment of the present invention, using the LOD technology to perform hierarchical rendering on a multi-level 3D Gaussian model, including:
[0093] Using a hierarchical screening formula to screen LOD levels and using a linear interpolation method to calculate a smoothing factor;
[0094] Smoothing different LOD levels according to the smoothing factor, calculating the mixed Gaussian points to be rendered, and performing hierarchical rendering according to the mixed Gaussian points to be rendered.
[0095] When performing hierarchical rendering according to the mixed Gaussian points to be rendered in the embodiment of the present invention, it further includes calculating the distance d from the viewpoint v to each Gaussian point in real time, and the calculation formula is as follows:
[0096]
[0097] where, (x i , y i , z i ) is the position of Gaussian point i, and (x c , y c , z c ) is the position of the camera.
[0098] The hierarchical screening formula in the embodiment of the present invention can adopt the following formula:
[0099] d i ≤d<d i+1
[0100] where, d i (d 0 <d 1 <…<d L-1 ) is the maximum effective distance of each layer.
[0101] Further, using the linear interpolation method to calculate the smoothing factor β, it can be calculated by the following formula:
[0102]
[0103] where, d is the distance from the viewpoint v to each Gaussian point, d 0 is the distance from the viewpoint v to each Gaussian point in the 0th layer, d 1 is the distance from the viewpoint v to each Gaussian point in the 1st layer.
[0104] Further, calculating the mixed Gaussian points U to be rendered, it can be calculated by the following formula:
[0105] U = (1 - β)p Lod0 + p Lod1
[0106] Wherein, p Lod0 is the Gaussian point in LOD level 0, and p Lod1 is the Gaussian point in LOD level 1.
[0107] S6. And hierarchically render the multi-level 3D Gaussian model by using the LOD technology.
[0108] According to the present invention, the three-dimensional structure information of the corresponding three-dimensional scene and the camera parameters corresponding to each image can be restored from multiple images, and the accurate extraction of three-dimensional geometric information can be realized. In addition, the subspaces are recursively divided according to the positions of the spatial points to obtain multiple divided subspaces, different levels of detail can be divided, and a hierarchical division basis is provided for the subsequent hierarchical rendering by the LOD technology. In addition, the central coordinates representing the sequence in each divided subspace are transformed into a three-dimensional Gaussian distribution, and the Gaussian points are initialized layer by layer to obtain a Gaussian distribution layer by layer, and the point cloud representation of multiple levels can be completed through the Gaussian point initialization layer by layer. According to the maximum number of divided layers, texture sampling is performed on the original real images, and texture information with different levels of detail can be obtained, improving the authenticity of the rendered images. Finally, hierarchically rendering the multi-level 3D Gaussian model by using the LOD technology can improve the rendering efficiency through hierarchical rendering.
[0109] As Figure 2 shown, it is a functional module diagram of a 3D Gaussian point cloud multi-level rendering optimization device provided by an embodiment of the present invention.
[0110] The 3D Gaussian point cloud multi-level rendering optimization device 100 according to the present invention can be installed in an electronic device. According to the implemented functions, the 3D Gaussian point cloud multi-level rendering optimization device 100 can include a three-dimensional structure information acquisition module 101, a subspace division module 102, and a multi-level rendering module 103.
[0111] The module according to the present invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0112] In this embodiment, the functions of each module / unit are as follows:
[0113] The three-dimensional structure information acquisition module 101 is used to acquire multiple images under a three-dimensional scene and restore the three-dimensional structure information of the corresponding three-dimensional scene according to the multiple images.
[0114] In the embodiments of the present invention, multiple images in a three-dimensional scene refer to multiple images taken from different perspectives in a three-dimensional space.
[0115] In the embodiments of the present invention, the three-dimensional structure information of the corresponding three-dimensional scene is restored from multiple images through the COLMAP algorithm. Among them, the COLMAP algorithm refers to an open-source multi-view Figure 3 3D reconstruction system, mainly used to restore the 3D structure from multiple 2D images.
[0116] The subspace division module 102 is configured to obtain the bounding box of the three-dimensional scene according to the three-dimensional structure information, calculate the maximum number of division layers of the bounding box according to the octree hierarchical division formula, and recursively divide the subspace according to the positions of the spatial points to obtain multiple divided subspaces; calculate the representation sequence of the spatial points in each divided subspace by using the Gaussian point representation formula, and convert the central coordinates of the representation sequence in each divided subspace into a three-dimensional Gaussian distribution, and perform layer-by-layer Gaussian point initialization to obtain layer-by-layer Gaussian distributions; perform texture sampling on the original real images according to the maximum number of division layers to obtain multi-level real image information.
[0117] In the embodiments of the present invention, the bounding box of the three-dimensional scene refers to the geometric framework that encloses the entire image.
[0118] In the embodiments of the present invention, the octree hierarchical division formula is a formula used to calculate the maximum number of division layers of the bounding box. The maximum number of division layers L of the bounding box can be calculated by the following formula:
[0119] L = [log 2 B] + 1
[0120] where B is the coordinate range of the bounding box, and the bounding box is divided into multiple divided subspaces according to the maximum number of division layers.
[0121] In the embodiments of the present invention, the coordinate range of the bounding box is realized by querying the minimum and maximum values of the coordinates of all spatial points in the space on the x-axis, y-axis, and z-axis, and can be represented by min_x, min_y, min_z and max_x, max_y, max_z.
[0122] As an embodiment of the present invention, recursively dividing the subspace according to the positions of the spatial points includes: taking the maximum number of division layers as the number of nodes of the octree; dividing the three-dimensional scene into multiple spaces with the same number as the number of nodes according to the bounding box; querying the bounding box of each space and the number and positions of the spatial points belonging to each space, and combining the bounding box of each space and the number and positions of the spatial points belonging to each space to obtain the divided subspaces. Among them, each node in the octree represents a divided subspace, and each node includes the bounding box of the current divided subspace and the number and positions of the spatial points in the divided subspace.
[0123] In the embodiments of the present invention, to calculate the representation sequence of spatial points in each layer of the divided subspace using the Gaussian point representation formula, the following formula can be used for calculation:
[0124] M = P / 2 l
[0125] where M is the representation sequence, l is the number of layers of the divided subspace, and P is the position coordinate of the spatial point.
[0126] Furthermore, the representation sequence can be expressed as {P / 2 0 , P / 2 1 , …, P / 2 L-1}.
[0127] As an embodiment of the present invention, converting the central coordinates of the representation sequence in each layer of the divided subspace into a three-dimensional Gaussian distribution includes:
[0128] Querying the total number of spatial points in the representation sequence in each layer of the divided subspace, and calculating the central coordinates based on the total number of spatial points;
[0129] Converting the central coordinates of each layer of the divided subspace into a three-dimensional Gaussian distribution according to a preset Gaussian distribution conversion formula.
[0130] Furthermore, converting the central coordinates of each layer of the divided subspace into a three-dimensional Gaussian distribution according to a preset Gaussian distribution conversion formula includes:
[0131] Calculating the central coordinate p according to the following formula:
[0132]
[0133] where N is the total number of spatial points, x i is the abscissa of the i-th spatial point, y i is the ordinate of the i-th spatial point, and z i is the coordinate value of the i-th spatial point on the Z-axis;
[0134] Converting the central coordinates of each layer of the divided subspace into a three-dimensional Gaussian distribution G(p) using the following formula:
[0135]
[0136] where ∑ is the covariance matrix, μ is the position of the spatial point in the three-dimensional Gaussian distribution, and (·) T is the transpose operation.
[0137] In the embodiment of the present invention, the central coordinates representing the sequence in each sub - space layer are transformed into a three - dimensional Gaussian distribution. In addition to performing layer - by - layer Gaussian point initialization, it further includes: endowing the attribute information of each spatial point p with a color value c and an opacity α. Among them, the color value c - - fits the appearance related to the viewing angle through spherical harmonic functions, and uses a linear combination of a set of orthogonal bases to fit the light field; the opacity α - - is used for rendering Splatting. When it is projected onto the image plane, the diffusion traces are superimposed together through the opacity.
[0138] In the embodiment of the present invention, texture sampling refers to a technology that can generate texture images with different resolutions.
[0139] As an embodiment of the present invention, texture sampling is performed on the original real image according to the maximum number of division layers, including:
[0140] Obtain the size information and the original resolution of the original real image;
[0141] Divide the resolution of each layer of the real image according to the size information, the original resolution, and the maximum number of division layers, and perform texture sampling according to the hierarchical index of the maximum number of division layers to obtain multi - level real image information.
[0142] In the embodiment of the present invention, the size information is the size information of the original real image.
[0143] Furthermore, performing texture sampling according to the hierarchical index of the maximum number of division layers further includes: the hierarchical index starts from layer 0, and the original real image is downsampled through average filtering. Among them, each layer of the real image is obtained by downsampling the previous layer of the real image until the real image under the maximum number of division layers is obtained.
[0144] Exemplarily, when performing texture sampling on the original real image according to the maximum number of division layers, the following implementation steps can be adopted:
[0145] Step 1: Extract an original real image I 0 , whose size is W×H, where W and H are the width and height of the image respectively;
[0146] Step 2: Define the maximum level of texture sampling as L layers for one - to - one correspondence during subsequent training;
[0147] Step 3: Extract the resolution of the original real image as 2 n ×2 n , where n is the resolution exponent. Then the image resolution of the i - th layer is where i is the hierarchical index, starting from 0. Each layer of image I i is obtained by downsampling the previous layer of image I i-1Obtained by downsampling. The process of downsampling usually uses average filtering to reduce high-frequency information.
[0148] The multi-level rendering module 103 is used to project the layer-by-layer Gaussian distribution onto a two-dimensional plane by means of hierarchical training, obtain a rendered image through volume rendering, calculate the loss using the multi-level real image information and the rendered image, and perform multi-level rendering optimization of the 3D Gaussian point cloud based on the loss using the backpropagation method. After the optimization is completed, a multi-level 3D Gaussian model is obtained; the multi-level 3D Gaussian model is hierarchically rendered using the LOD technology.
[0149] As an embodiment of the present invention, obtaining a rendered image through volume rendering includes:
[0150] Projecting the spatial points in each divided subspace onto a two-dimensional plane, and performing volume rendering using the following rendering formula:
[0151]
[0152] where I(x) is the rendered image, N is the total number of spatial points, α i is the transparency of the i-th sampling point, C i is the color value of the i-th sampling point, is the cumulative transparency of all spatial points before the i-th sampling point.
[0153] As an embodiment of the present invention, hierarchically rendering the multi-level 3D Gaussian model using the LOD technology includes:
[0154] Selecting the LOD level using the level selection formula and calculating the smoothing factor using the linear interpolation method;
[0155] Performing smoothing processing on different LOD levels according to the smoothing factor, calculating the mixed Gaussian points to be rendered, and performing hierarchical rendering according to the mixed Gaussian points to be rendered.
[0156] When hierarchically rendering according to the mixed Gaussian points to be rendered in the embodiment of the present invention, it further includes calculating in real time the distance d from the viewing point v to each Gaussian point, and the calculation formula is as follows:
[0157]
[0158] where (x i , y i , z i ) is the position of the Gaussian point i, and (x c , y c , z c ) is the position of the camera.
[0159] The level selection formula in the embodiment of the present invention can adopt the following formula:
[0160] d i ≤d < d i+1
[0161] where d i (d 0 < d 1 <… < d L-1 ) is the maximum effective distance of each layer.
[0162] Furthermore, the smoothing factor β is calculated by using the linear interpolation method, and can be calculated by the following formula:
[0163]
[0164] where d is the distance from the view point v to each Gaussian point, d 0 is the distance from the view point v to each Gaussian point of the 0th layer, d 1 is the distance from the view point v to each Gaussian point of the 1st layer.
[0165] Furthermore, to calculate the Gaussian point U to be rendered, the following formula can be used:
[0166] U = (1 - β)p Lod0 + p Lod1
[0167] where p Lod0 is the Gaussian point in the LOD level 0, p Lod1 is the Gaussian point in the LOD level 1.
[0168] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing a multi-level rendering optimization method of 3D Gaussian point cloud provided by an embodiment of the present invention.
[0169] The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as a program for a multi-level rendering optimization method of 3D Gaussian point cloud.
[0170] Among them, in some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. By running or executing programs or modules stored in the memory 11 (such as executing a program for a 3D Gaussian point cloud multi-level rendering optimization method, etc.), and by calling the data stored in the memory 11, it performs various functions of the electronic device and processes data.
[0171] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of a 3D Gaussian point cloud multi-level rendering optimization method program, etc., but also to temporarily store data that has been output or will be output.
[0172] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to achieve connection communication between the memory 11 and at least one processor 10, etc.
[0173] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.
[0174] Figure 3 Only the electronic device with components is shown, and those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0175] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0176] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0177] The 3D Gaussian point cloud multi-level rendering optimization method program stored in the memory 11 of the electronic device is a combination of multiple instructions, and when running in the processor 10, it can implement:
[0178] Obtain multiple images in a three-dimensional scene, and restore the three-dimensional structure information of the corresponding three-dimensional scene according to the multiple images;
[0179] Obtain the bounding box of the three-dimensional scene according to the three-dimensional structure information, calculate the maximum number of division layers of the bounding box according to the octree hierarchical division formula, and recursively divide the subspaces according to the positions of the spatial points to obtain multi-level divided subspaces;
[0180] Calculate the representation sequence of spatial points in each layer of the divided subspace using the Gaussian point representation formula, convert the central coordinates of the representation sequence in each layer of the divided subspace into a three-dimensional Gaussian distribution, perform layer-by-layer Gaussian point initialization to obtain a layer-by-layer Gaussian distribution;
[0181] Perform texture sampling on the original real image according to the maximum number of divided layers to obtain multi-level real image information;
[0182] Use layer-by-layer training to project the layer-by-layer Gaussian distribution onto a two-dimensional plane, obtain a rendered image through volume rendering, calculate the loss using the multi-level real image information and the rendered image, and perform multi-level rendering optimization of the 3D Gaussian point cloud based on the loss using the backpropagation method. After the optimization is completed, obtain a multi-level 3D Gaussian model;
[0183] Use the LOD technology to perform hierarchical rendering on the multi-level 3D Gaussian model.
[0184] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to the description of the relevant steps in the corresponding embodiments of the accompanying drawings, which will not be elaborated here.
[0185] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0186] The present invention also provides a computer-readable storage medium, and the readable storage medium stores a computer program, which when executed by the processor of the electronic device, can implement:
[0187] Obtain multiple images in a three-dimensional scene, and restore the three-dimensional structure information of the corresponding three-dimensional scene according to the multiple images;
[0188] Obtain the bounding box of the three-dimensional scene according to the three-dimensional structure information, calculate the maximum number of divided layers of the bounding box according to the octree hierarchical division formula, and recursively divide the subspace according to the positions of the spatial points to obtain multiple divided subspaces;
[0189] Calculate the representation sequence of spatial points in each layer of the divided subspace using the Gaussian point representation formula, convert the central coordinates of the representation sequence in each layer of the divided subspace into a three-dimensional Gaussian distribution, perform layer-by-layer Gaussian point initialization to obtain a layer-by-layer Gaussian distribution;
[0190] Perform texture sampling on the original real image according to the maximum number of division levels to obtain multi-level real image information;
[0191] Use hierarchical training to project the layer-by-layer Gaussian distribution onto a two-dimensional plane, and obtain a rendered image through volume rendering. Calculate the loss using the multi-level real image information and the rendered image, and perform multi-level rendering optimization of the 3D Gaussian point cloud based on the loss using the backpropagation method. After the optimization is completed, a multi-level 3D Gaussian model is obtained;
[0192] Use the LOD technology to perform hierarchical rendering on the multi-level 3D Gaussian model.
[0193] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0194] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0195] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of hardware plus software functional modules.
[0196] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0197] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.
[0198] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0199] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0200] In addition, it is obvious that the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as first and second are used to represent names and do not indicate any specific order.
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A 3D Gaussian point cloud multi-level rendering optimization method, characterized in that: The method comprises: Acquire multiple images of a three-dimensional scene, and restore three-dimensional structure information of the corresponding three-dimensional scene according to the multiple images; Obtain the bounding box of the 3D scene according to the 3D structure information, calculate the maximum number of division levels of the bounding box according to the octree hierarchical division formula, and recursively divide the subspace according to the position of the spatial point to obtain a multi-layer division subspace; The Gaussian point representation formula is used to calculate the representation sequence of the spatial points in each layer of the divided subspace, and the central coordinates of the representation sequence in each layer of the divided subspace are converted into a three-dimensional Gaussian distribution, and Gaussian point initialization is performed layer by layer to obtain a layer by layer Gaussian distribution; Perform texture sampling on the original real image according to the maximum number of division layers to obtain multi-level real image information; The Gaussian distribution layer by layer is projected onto a two-dimensional plane through layer-by-layer training, and a rendered image is obtained through volume rendering. The multi-layer real image information and the rendered image are used to calculate the loss. The back propagation method is used based on the loss to optimize the multi-layer rendering of the 3D Gaussian point cloud. After the optimization is completed, a multi-layer 3D Gaussian model is obtained. LOD technology is used to perform hierarchical rendering on multi-level 3D Gaussian models.
2. The 3D Gaussian point cloud multi-level rendering optimization method according to claim 1, characterized in that: The conversion of the central coordinates of the sequence represented in each layer of the divided subspace into a three-dimensional Gaussian distribution includes: Query the total number of spatial points in the sequence represented in each layer of divided subspace, and calculate the center coordinates based on the total number of spatial points; According to the preset Gaussian distribution conversion formula, the center coordinates of each layer of divided subspace are converted into a three-dimensional Gaussian distribution.
3. The 3D Gaussian point cloud multi-level rendering optimization method according to claim 1 or 2, characterized in that: The method of converting the center coordinates of each layer of divided subspace into a three-dimensional Gaussian distribution according to a preset Gaussian distribution conversion formula includes: The center coordinate p is calculated according to the following formula: Among them, N is the total number of space points, x i is the horizontal coordinate of the i-th spatial point, y i is the ordinate of the i-th spatial point, z i is the coordinate value of the i-th spatial point on the Z axis; The following formula is used to transform the center coordinates of each layer of divided subspace into a three-dimensional Gaussian distribution G(p): Where Σ is the covariance matrix, μ is the position of the spatial point in the three-dimensional Gaussian distribution, (·) T is the transpose operation.
4. The 3D Gaussian point cloud multi-level rendering optimization method according to claim 1, characterized in that: The method of sampling the texture of the original real image according to the maximum number of division layers includes: Get the size information and original resolution of the original real image; The resolution of each layer of the real image is divided according to the size information, the original resolution and the maximum number of division layers, and texture sampling is performed according to the level index of the maximum number of division layers to obtain multi-level real image information.
5. The 3D Gaussian point cloud multi-level rendering optimization method according to claim 4, characterized in that: The texture sampling according to the level index of the maximum number of division layers includes: the level index starts from level 0, and the original real image is downsampled by average filtering, wherein each layer of the real image is obtained by downsampling the real image of the previous layer until the real image with the maximum number of division layers is obtained.
6. The 3D Gaussian point cloud multi-level rendering optimization method according to claim 1, characterized in that: The step of obtaining a rendered image through volume rendering includes: Project the spatial points of each layer of subspace division onto a two-dimensional plane, and perform volume rendering using the following rendering formula: Among them, I(x) is the rendered image, N is the total number of spatial points, α i is the transparency of the i-th sampling point, C i is the color value of the i-th sampling point, is the cumulative transparency of all spatial points before the i-th sampling point.
7. The 3D Gaussian point cloud multi-level rendering optimization method according to claim 1, characterized in that: The method of using LOD technology to perform hierarchical rendering on a multi-level 3D Gaussian model includes: Use the level screening formula to screen the LOD level, and use the linear interpolation method to calculate the smoothing factor; Different LOD levels are smoothed according to the smoothing factor, and the mixed Gaussian points to be rendered are calculated, and hierarchical rendering is performed according to the mixed Gaussian points to be rendered.
8. A 3D Gaussian point cloud multi-level rendering optimization device, characterized in that: The device is used to implement the 3D Gaussian point cloud multi-level rendering optimization method according to any one of claims 1 to 7, and the device includes: A three-dimensional structure information acquisition module is used to acquire multiple images in a three-dimensional scene and restore the three-dimensional structure information of the corresponding three-dimensional scene based on the multiple images; The subspace partitioning module is used to obtain the bounding box of the three-dimensional scene according to the three-dimensional structure information, calculate the maximum number of division layers of the bounding box according to the octree hierarchical partitioning formula, and recursively divide the subspace according to the position of the spatial point to obtain a multi-layer divided subspace; use the Gaussian point representation formula to calculate the representation sequence of the spatial points in each layer of the divided subspace, and convert the central coordinates of the representation sequence in each layer of the divided subspace into a three-dimensional Gaussian distribution, perform layer-by-layer Gaussian point initialization, and obtain a layer-by-layer Gaussian distribution; perform texture sampling on the original real image according to the maximum number of division layers to obtain multi-level real image information; The multi-level rendering module is used to project the Gaussian distribution layer by layer onto a two-dimensional plane through layer-by-layer training, and obtain the rendered image through volume rendering. The multi-level real image information and the rendered image are used for loss calculation, and the back propagation method is used based on the loss to optimize the multi-level rendering of the 3D Gaussian point cloud. After the optimization is completed, a multi-level 3D Gaussian model is obtained; the LOD technology is used to perform hierarchical rendering of the multi-level 3D Gaussian model.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the 3D Gaussian point cloud multi-level rendering optimization method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the 3D Gaussian point cloud multi-level rendering optimization method as described in any one of claims 1 to 7 is implemented.