3D Gaussian field generation method based on three-plane representation

By combining a three-plane representation and a renderer with a diffusion model, the encoding and compression problem of Gaussian scattering data structures is solved, achieving efficient 3D Gaussian field generation and rendering, which is suitable for 3D model design and game scene design.

CN121708192APending Publication Date: 2026-03-20THE CHINESE UNIVERSITY OF HONG KONG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411279835.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Traditional 3D representation methods such as voxels, implicit fields, meshes, and point clouds have limitations in efficiency and flexibility when processing complex 3D data. In particular, the massive data structure of Gaussian scattering (GS) is difficult to generate and encode and compress directly, and multi-view image generation methods cannot guarantee 3D consistency.

Method used

The three-plane representation is used to encode the geometric and Gaussian properties of 3D objects onto three mutually perpendicular planes, construct a three-plane mesh, and encode and decode it through a three-plane renderer and variational autoencoder (VAE). Combined with a diffusion model, it generates 3D Gaussian field point clouds and multi-view images.

Benefits of technology

It achieves efficient 3D Gaussian field generation and rendering, and the generated objects have high-quality rendering results and consistency in multiple viewpoints. It is suitable for fast and editable real-time rendering, compatible with existing rendering software, and can be applied to 3D model design and game scene design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708192A_ABST
    Figure CN121708192A_ABST
Patent Text Reader

Abstract

The invention relates to a three-plane (Triplane)-based 3D Gaussian field generation method, which comprises the following steps of: firstly, inputting and standardizing 3D object data to adapt to three-plane representation; secondly, using a three-plane encoder to encode geometric and Gaussian attributes of the 3D object to three vertical planes, and constructing a three-plane grid; then, geometric and Gaussian field attribute branches of a three-plane renderer are trained, and decoding from a grid to a 3D grid and attributes is achieved; then, the geometrical shape and the attribute of the 3D object are decoded through a three-plane decoder; further, training VAE model compression and reconstructing three-plane representation; compressing the representation into a hidden code by using a VAE encoder, and training a diffusion model in a hidden code space; generating a hidden code of the new 3D object through an inverse diffusion process; and finally, decoding and rendering the implicit code into a 3D Gaussian field point cloud by using a VAE decoder and a three-plane renderer, and generating a high-quality 3D object geometry and rendering result from multiple perspectives. The method ensures the reconstruction quality and the high efficiency of multi-view rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to 3D object generation, and in particular to a method for generating 3D Gaussian fields based on triplane representation. Background Technology

[0002] 3D object generation is an important topic in computer vision and graphics, widely used in games, animation, augmented reality / virtual reality (AR / VR), and other fields. With the development of differentiable rendering techniques, such as mesh-based and neural radiative field (NeRF) rendering, significant progress has been made in 3D content generation. However, traditional 3D representation methods such as voxels, implicit fields, meshes, and point clouds have limitations in efficiency and flexibility when processing complex 3D data.

[0003] Gaussian Splatting (GS), an emerging differentiable rendering technique, uses point clouds (splats) with Gaussian distribution shapes to describe 3D content and achieves real-time rendering through special rasterization techniques. GS has attracted attention due to its high rendering efficiency and good editability, but its complex data structure and multi-channel characteristics make directly generating GS data challenging.

[0004] Specifically, traditional object generation methods primarily focus on generating shape geometry based on different data structures, including voxels, implicit fields, meshes, and point clouds, as shown in studies such as [24,51,38]. However, when it comes to vision-oriented tasks, as shown in studies such as [32,2], merging textures becomes crucial, with these studies generating corresponding colors based on shapes.

[0005] Recently, numerous studies have emerged aiming to combine geometry generation and rendering, which can be broadly categorized into two types. The first type involves multi-view image-to-3D methods, which generate multi-view colored 3D images, then reconstruct 3D shapes and project the images onto textures, such as [31,34,15,19,8,36]. While these methods require only 2D supervision, maintaining 3D consistency across multi-view images is challenging, potentially leading to degraded quality of the generated 3D geometry. The second type primarily operates in 3D space and can be classified based on different 3D representations. For example, [13,12] is able to generate high-quality 3D textured meshes using differentiable rendering. [27,46] generates point clouds and their corresponding colors. [3,22] generates NeRF volumes via latent diffusion.

[0006] Due to its editing flexibility and rendering efficiency, GS is becoming the most popular 3D content format. However, its massive data structure has become a bottleneck for generation. Therefore, some studies have utilized multi-view image-based methods, such as [8,36], where an image generator outputs images of the desired views to recursively enhance GS rendering. Few studies have directly explored GS generation in 3D space. Most relevant to this work is

[52] , which uses a Transformer structure to generate point clouds from 2D images as input and constructs a three-plane encoding mapping from the image to GS attributes.

[0007] It should be noted that the information disclosed in the background section above is only for understanding the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0008] The main objective of this invention is to solve the problems existing in the above-mentioned background technology and provide a method for generating 3D Gaussian fields based on triplane representation.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] In a first aspect, the present invention provides a method for generating 3D Gaussian fields based on a triplane, comprising:

[0011] Input 3D object data and standardize it to meet the requirements of three-plane representation;

[0012] A three-plane encoder is used to encode the geometric and Gaussian properties of a 3D object onto three mutually perpendicular planes to form a three-plane mesh.

[0013] 3D object information captured from different perspectives is stored on three vertical planes to construct a complete three-plane mesh;

[0014] Train the geometry rendering branch and the Gaussian field property rendering branch of the triplane renderer to implement decoding from triplane mesh to 3D mesh and Gaussian property;

[0015] Use a three-plane decoder to decode the geometry and Gaussian properties of 3D objects from a three-plane mesh;

[0016] Train the encoder and decoder of the variational autoencoder (VAE) to compress the three-plane representation into an implicit code and reconstruct the three-plane representation from the implicit code;

[0017] The three-plane representation is compressed into a hidden code using a trained VAE encoder;

[0018] A diffusion model, including a conditional diffusion module and a noise injection mechanism, is built and trained in the implicit code space to learn the data generation process.

[0019] The hidden code of a new 3D object is generated by using a trained diffusion model through a reverse diffusion process.

[0020] The VAE decoder is used to decode the generated implicit code back to a three-plane representation;

[0021] The decoded three-plane representation is rendered into a 3D Gaussian field point cloud using a three-plane renderer, and then multi-view images are generated.

[0022] In the second aspect, a 3D Gaussian field generation method based on a triplane is used in the model training phase, including:

[0023] Input 3D object data and standardize it to meet the requirements of three-plane representation;

[0024] A three-plane encoder is used to encode the geometric and Gaussian properties of a 3D object onto three mutually perpendicular planes to form a three-plane mesh.

[0025] 3D object information captured from different perspectives is stored on three vertical planes to construct a complete three-plane mesh;

[0026] Train the two branches of the triplane renderer: a geometry rendering branch and a Gaussian field property rendering branch, so that it can decode from triplane meshes to 3D meshes and Gaussian properties;

[0027] Training is performed on a three-plane mesh using a three-plane decoder to accurately recover the geometry and Gaussian properties of 3D objects;

[0028] Train a VAE model, including an encoder and a decoder, to compress a three-plane representation into an implicit code and reconstruct the three-plane representation from the implicit code;

[0029] The VAE encoder is trained to compress the three-plane representation into a hidden code, which prepares data for training the diffusion model.

[0030] A diffusion model is constructed in the implicit code space, including a conditional diffusion module and a noise injection mechanism, so that the model can learn the data generation process;

[0031] The training diffusion model generates the implicit code of new 3D objects through a reverse diffusion process;

[0032] The decoder trained on the VAE decodes the implicit code back to a three-plane representation and renders it using a three-plane renderer to generate 3D Gaussian field point clouds and multi-view images.

[0033] In a third aspect, the present invention provides a 3D Gaussian field generation method based on a triplane for the inference stage of a model, comprising:

[0034] Inverse diffusion process: Using the trained diffusion model, the implicit code of a new 3D object is generated through the inverse diffusion process;

[0035] Implicit code decoding: The VAE decoder is used to decode the implicit code obtained during the inverse diffusion process back into a three-plane representation to reconstruct the geometric and Gaussian properties of the 3D object;

[0036] Three-plane rendering: The three-plane renderer is used to render the decoded three-plane representation to generate a 3D Gaussian field point cloud;

[0037] Multi-view image generation: Based on the rendered 3D Gaussian field point cloud, generate multi-view images of the object to achieve a comprehensive visual presentation;

[0038] Output: The generated 3D Gaussian field point cloud or multi-view image will be used as the final output.

[0039] In a fourth aspect, a computer-readable storage medium stores a computer program that, when executed by a processor, implements the aforementioned method for generating 3D Gaussian fields based on a triplane.

[0040] In a fifth aspect, a computer program product includes a computer program that, when executed by a processor, implements the aforementioned triplane-based 3D Gaussian field generation method.

[0041] The present invention has the following beneficial effects:

[0042] This invention proposes an end-to-end 3D Gaussian field generation method. Through a triplane representation, Gaussian scattering is represented as a continuous field resembling an image. This representation effectively encodes geometric and texture information, can be smoothly converted back to a Gaussian point cloud, and rendered into an image using the TriRenderer. The TriRenderer is fully differentiable, thus the rendering loss can supervise the encoding of texture and geometry. Furthermore, the triplane representation can be compressed using a variational autoencoder (VAE) and subsequently used in latent diffusion to generate 3D objects. Experimental results demonstrate that the proposed GS representation achieves satisfactory reconstruction quality, and the generative framework can generate high-quality 3D object geometry and rendering results from multiple perspectives.

[0043] 3D Gaussian fields, as a point cloud-based representation of 3D objects, can be used for fast and editable real-time rendering. However, due to their discretized distribution, uneven density, and large number of data channels, they are difficult to directly encode, compress, and apply to existing generative models. This invention proposes a three-plane-based 3D Gaussian field representation method, compressing the 3D Gaussian field into three mutually perpendicular continuous planar spaces, and designing a corresponding renderer for encoding and decoding. This invention can be directly applied to currently popular generative frameworks such as diffusion models to generate 3D Gaussian fields. The generated objects are compatible with existing Gaussian field rendering software, enabling fast rendering and editing operations, and can be applied to tasks such as 3D model design and game scene design.

[0044] Other beneficial effects of the embodiments of the present invention will be further described below. Attached Figure Description

[0045] Figure 1 This is a schematic diagram illustrating the three planes in an embodiment of the present invention.

[0046] Figure 2 This is a schematic diagram of a three-plane renderer according to an embodiment of the present invention.

[0047] Figure 3 This is a schematic diagram of the generation framework based on the three-plane scheme in an embodiment of the present invention.

[0048] Figure 4 The generated sample is shown using Gaussian scattering rendering.

[0049] Figure 5 The diagram illustrates a three-plane VAE and a two-stage potential diffusion.

[0050] Figure 6 The channel visualization of the three planes of the sample is shown.

[0051] Figure 7 The Gaussian scattering (GS) fitting comparison is shown.

[0052] Figure 8 The comparison of the generated samples is shown.

[0053] Figure 9 This shows the image quality degradation in a VAE; left: the actual situation. Right: VAE reconstruction.

[0054] Figure 10 This indicates that GS attribute generation failed to be performed directly in voxel space.

[0055] Figure 11 The results show the noise generated by direct diffusion across the three planes.

[0056] Figure 12 An example of a generated sample is shown. Detailed Implementation

[0057] The embodiments of the present invention will be described in detail below. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.

[0058] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of the present invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0059] 3D object generation has always been an important topic in computer vision and graphics, and has been widely used in games, animation production and AR / VR. With the advancement of differentiable rendering techniques on various representation methods (such as meshes and neural radiation fields (NeRF)), works like [22,12] can directly generate high-quality textured objects. Recently, Gaussian scattering (GS)

[17] has become a new trend in differentiable rendering due to its high rendering efficiency and good editability. However, due to the complexity of GS data structure, few works directly address the challenge of generating GS data. In this invention, a new method called TriGSGen is proposed, which represents GS as a three-plane representation and introduces a corresponding renderer TriRenderer. Subsequently, latent diffusion is applied on the three-plane representation to generate high-quality GS object data.

[0060] Gaussian scattering (GS) is a newly developed rendering technique that uses "splats" of multichannel point clouds to describe 3D content and renders it into an image using differentiable rasterization. Compared to neural radiation fields, GS offers fast rendering speed and good editing flexibility. However, the sparsity, multichannel nature, and non-uniform density distribution of 3D GS pose significant challenges to direct generation. Therefore, many studies [8,36] have employed the generation of multi-view images to reconstruct 3D content. However, this approach, lacking 3D priors, struggles to fundamentally guarantee 3D consistency between the generated multi-view images. To address this issue, it is worth considering whether the sparse GS can be converted into a continuous field, and then generation can be performed on this field.

[0061] This invention proposes representing GS content as a triplane, which has proven advantageous for representing 3D geometry or even NeRF, as shown in [45,35]. In this invention, the triplane representation is used to encode the geometric information and other channels of the GS. Each 3D object can be represented by a triplane. By training on a batch of 3D objects, a shared TriRenderer can be obtained, which is capable of decoding any triplane into a GS and then rendering it to the screen. The TriRenderer is fully differentiable, allowing the use of rendering loss to supervise texture and geometric information.

[0062] A generative framework is constructed on the proposed triplane representation using a stable diffusion model. First, a variational autoencoder (VAE) is designed to further compress the triplane into the latent space. Two independent decoders are used to separate the decoding of geometric and other GS attributes. The triplane latent is extended to an extended multichannel image, which is then generated using latent diffusion.

[0063] The main concepts and contributions of this invention are as follows: 1) A three-plane representation for GS is proposed, which can reconstruct GS point clouds with satisfactory rendering performance. 2) A fully differentiable TriRenderer is designed to supervise the training of GS to three planes. 3) A three-plane-based GS generation framework is developed, which combines a specially designed VAE and a latent diffusion module. 4) Experiments show that the generation achieves competitive performance in both 3D geometry and multi-view rendering quality.

[0064] See Figure 1 This invention provides a method for generating 3D Gaussian fields based on a triplane, comprising the following steps:

[0065] S1. Data Preprocessing: Input 3D object data and standardize it to meet the requirements of three-plane representation;

[0066] S2. Triplane Representation: The geometric and Gaussian properties of a 3D object are encoded onto three mutually perpendicular planes using a triplane encoder to form a triplane mesh;

[0067] S3. Three-plane mesh construction: Store 3D object information captured from different perspectives on three vertical planes to construct a complete three-plane mesh;

[0068] S4. Triplane Renderer Training: Train the two branches of the triplane renderer: the geometry rendering branch and the Gaussian field property rendering branch, to implement the decoding from triplane mesh to 3D mesh and Gaussian property;

[0069] S5. Geometry and Gaussian Attribute Decoding: Uses a three-plane decoder to decode the geometry and Gaussian attributes of 3D objects from a three-plane mesh;

[0070] S6. Variational Autoencoder (VAE) Encoding and Reconstruction: Train the encoder and decoder of the VAE to compress the three-plane representation into an implicit code, and reconstruct the three-plane representation from the implicit code;

[0071] S7. Implicit Code Generation: The trained VAE encoder is used to compress the three-plane representation into an implicit code to prepare data for the next step of the diffusion model.

[0072] S8. Diffusion Model Training: Construct and train a diffusion model in the implicit code space, including a conditional diffusion module and a noise injection mechanism, to learn the data generation process;

[0073] S9. Implementation of the reverse diffusion process: Using the trained diffusion model, the implicit code of a new 3D object is generated through the reverse diffusion process;

[0074] S10. Implicit code decoding to three-plane representation: The VAE decoder is used to decode the generated implicit code back to a three-plane representation;

[0075] S11. Final Rendering: The decoded three-plane representation is rendered into a 3D Gaussian field point cloud using a three-plane renderer, and further multi-view images are generated.

[0076] In some embodiments, in step S2, the three-plane representation specifically includes: using the three planes formed by the x, y, and z axes in 3D space, obtaining the features of any 3D point through orthogonal projection, and constructing a continuous three-plane GS field; implementing a tensor of size 3×H×H×C to store feature information on the three planes, where H represents the resolution of each axis and C represents the number of feature channels.

[0077] In some embodiments, step S3 specifically includes: performing bilinear interpolation query on any continuous coordinate point in the three-plane grid to obtain the features of the point.

[0078] In some embodiments, step S4, training the two branches of the three-plane renderer specifically includes:

[0079] S4.1 Geometric Branch Training Sub-step: Obtain grid features using a grid coordinate query system with a preset resolution; process the grid features through a first feature decoder based on a fully connected layer to obtain the filling state and TSDF value of each point; apply a differentiable Marching Cube algorithm to generate a mesh of the 3D object based on the TSDF value; perform sampling operations on the generated mesh surface to generate a point cloud for subsequent processing.

[0080] S4.2 Gaussian Field Attribute Branch Training Sub-step: Using the point cloud obtained in step S4.1, query the three planes to obtain the corresponding grid features; obtain the Gaussian field attribute of each point based on the grid features through the second feature decoder based on the fully connected layer; construct the Gaussian field point cloud based on the obtained Gaussian field attribute; set multiple virtual camera positions, and use the three-plane renderer to perform multi-view rendering on the Gaussian field point cloud to generate multi-view images.

[0081] In some embodiments, step S4, the training of the three-plane renderer further includes: S4.3 Three-plane parameter initialization: initializing the three-plane parameters corresponding to each object in the dataset, with the initial value set to zero; S4.4 Renderer decoder weight initialization: randomly initializing the decoder weights in the renderer shared by all objects; S4.5 Parameter training and update: during the training process, truncated signed distance function (TSDF) data and multi-view image data are used simultaneously as supervision signals to train and update the three-plane parameters and decoder weights; wherein, ground truth TSDF data is used to ensure the accuracy of 3D geometric information; multi-view image data is used to ensure the visual effect and consistency of Gaussian field properties.

[0082] In some embodiments, the processing method using Gaussian scattering in steps S4 and S5 includes:

[0083] In step S4, a set of Gaussian spheres with positional mean, covariance, and three-axis scales are used to describe 3D objects, where each point is equipped with an opacity scalar and color information to capture the viewpoint dependence of the scene; spherical harmonics are used to represent the color of each Gaussian point to ensure that the color representation can change according to the viewing angle, achieving a realistic visual effect under viewpoint changes; during the rendering process, the Gaussian points are projected into two-dimensional space, and the covariance matrix is ​​adjusted to adapt to the viewing transformation, ensuring the accuracy and efficiency of the rendering process; the covariance matrix is ​​represented by Cholesky decomposition to ensure its positive semidefiniteness throughout the training process.

[0084] In step S5, a decoder is used to decode the geometry and Gaussian properties of the 3D object from the three-plane mesh, including the opacity scalar and color information represented by spherical harmonics. The position and properties of each Gaussian point are integrated into a unified data structure for effective management during rendering and decoding. The decoded Gaussian point cloud is rendered into an image using a rasterization-based splatting renderer that supports differentiable operations to display the decoding results and can be integrated into a further training and optimization framework.

[0085] In some embodiments, in step S4, the training of the tri-plane renderer specifically includes: encoding 3D geometric information and other GS attributes separately using independent channels to improve training convergence; establishing the TriRenderer structure, which consists of a GS decoder and a GS renderer, wherein the GS decoder includes a geometry decoding branch and a GS attribute decoding branch, used to recover point clouds from the tri-plane and obtain GS attributes; representing the geometry using the truncated signed distance function (TSDF), extracting triangular meshes from the TSDF volume and sampling point clouds using the Marching Cube algorithm; and implementing gradient backpropagation and optimizing the geometry using a fully differentiable Marching Cube algorithm and an interpolation-based point sampler.

[0086] In a preferred embodiment, during the training of the three-plane renderer, L1 loss and cross-entropy regression of TSDF values ​​and binary occupancy are used to calculate geometric loss, and L1 pixel loss and SSIM loss are used to calculate rendering loss; when the occupancy accuracy reaches a certain threshold, backpropagation of rendering loss is enabled to optimize TSDF values.

[0087] In some embodiments, step S6, the variational autoencoder (VAE) training specifically includes: VAE structure design: adopting a UNet-like VAE structure to compress the three planes into the latent space, wherein the three planes are regarded as uniform data encoding features from three orthogonal view directions; data reshaping: reshaping the three-plane data with a batch size of B into a new batch format, from B×3×H×H×C to 3B×H×H×C, to adapt to VAE training; encoder and decoder decoupling: using completely independent encoders and decoders for geometry and other Gaussian properties (GS properties) to avoid blurring of rendering results that may be caused by mixed encoding.

[0088] In a preferred embodiment, the variational autoencoder (VAE) training uses binary cross-entropy loss (BCE) and L1 loss to regress and predict binary occupied voxel grids and TSDF voxel grids, uses L1 pixel loss and SSIM loss to calculate rendering loss, and introduces KL divergence loss to constrain the latent space, but does not use L1 loss between the input and reconstructed three planes.

[0089] The variational autoencoder (VAE) training employs a phased training approach, first training the geometric encoder... The encoder and decoder are then trained using rendering loss;

[0090] In some embodiments, a reconstruction method based on the truncated signed distance function (TSDF) is used for mesh extraction and point sampling, rather than relying on real data.

[0091] In some embodiments, in step S8, based on the decoupled geometric and Gaussian attribute latent codes, a two-stage inverse diffusion process is used to generate geometric latent codes and Gaussian attribute latent codes sequentially. The first stage generates the geometric latent code: during the first stage training, the input samples are mixed with Gaussian noise using a diffusion model to predict and regress the noise values. The second stage generates the Gaussian attribute latent code: during the second stage training, the geometric latent code generated in the first stage is used as a condition to perform a diffusion process on the ground-value Gaussian attribute latent code, and the UNet in the diffusion model is used to predict and regress the noise values ​​under this condition. The three-plane latent representation is unfolded into an image shape and used as the input to the diffusion model, where the class labels are mapped to integers as diffusion conditions. In step S9, the inverse diffusion process is performed using the model trained in both stages to predict and progressively remove noise to reverse the forward diffusion process and generate the latent code for the desired 3D object. In step S10, the generated latent code is converted back to its original shape and decoded using a trained TriRenderer to render multi-view images.

[0092] In a preferred embodiment, in steps S8 and S9, mean squared error loss is used to supervise noise prediction. This loss measures the difference between the noise predicted by the model and the actual noise, and the model is trained by minimizing this difference.

[0093] This invention also provides a method for generating 3D Gaussian fields based on a triplane for the inference stage of a model, comprising the following steps:

[0094] R1. Reverse diffusion process: Using the trained diffusion model, the hidden code of a new 3D object is generated through the reverse diffusion process;

[0095] R2. Implicit Code Decoding: Using the VAE decoder, the implicit code obtained during the inverse diffusion process is decoded back into a three-plane representation to reconstruct the geometric and Gaussian properties of the 3D object;

[0096] R3. Triplane Rendering: Uses a triplane renderer to render the decoded triplane representation and generate a 3D Gaussian field point cloud;

[0097] R4. Multi-view image generation: Generate multi-view images of objects based on the rendered 3D Gaussian field point cloud to achieve a comprehensive visual presentation;

[0098] R5. Output: The generated 3D Gaussian field point cloud or multi-view image will be used as the final output.

[0099] This invention also provides a method for generating 3D Gaussian fields based on a triplane for use in the model training phase, comprising the following steps:

[0100] T1. Data Preprocessing: Input 3D object data and standardize it to meet the requirements of three-plane representation;

[0101] T2. Triplane Representation: The geometric and Gaussian properties of a 3D object are encoded onto three mutually perpendicular planes using a triplane encoder to form a triplane mesh;

[0102] T3. Three-plane mesh construction: Store 3D object information captured from different perspectives on three vertical planes to construct a complete three-plane mesh;

[0103] T4. Triplane Renderer Training: Train the two branches of the triplane renderer: the geometry rendering branch and the Gaussian field property rendering branch, so that it can decode from the triplane mesh to the 3D mesh and Gaussian property;

[0104] T5. Geometry and Gaussian Property Decoding Training: Training to decode a three-plane mesh using a three-plane decoder to accurately recover the geometry and Gaussian properties of a 3D object;

[0105] T6. Variational Autoencoder (VAE) Training: Training a VAE model, including an encoder and a decoder, to compress a three-plane representation into an implicit code and to reconstruct the three-plane representation from the implicit code;

[0106] T7. Implicit Code Generation Training: Train the VAE encoder to compress the three-plane representation into an implicit code, preparing data for the training of the diffusion model;

[0107] T8. Diffusion Model Training: Construct a diffusion model in the implicit code space, including a conditional diffusion module and a noise injection mechanism, so that the model can learn the data generation process;

[0108] T9. Inverse Diffusion Inference: The training diffusion model generates the implicit code of new 3D objects through the inverse diffusion process;

[0109] T10. Implicit Code Decoding and Rendering Training: Train the VAE decoder to decode the implicit code back to a three-plane representation and render it using a three-plane renderer to generate 3D Gaussian field point clouds and multi-view images.

[0110] This invention proposes an innovative 3D Gaussian field generation method based on triplanes. This method significantly improves the generation and rendering efficiency of 3D objects by effectively transforming and compressing complex Gaussian scattering (GS) data structures into continuous triplane representations. Triplane representation not only improves memory efficiency and is more economical than dense voxels, but also provides a lightweight encoding method for the latent representation of 3D objects through further compression via variational autoencoders (VAEs). Furthermore, the triplane renderer of this invention is fully differentiable, allowing rendering losses to be directly used to supervise and optimize the encoding process of geometric and texture information, ensuring high-quality rendering results and consistency of the generated 3D objects across multiple viewpoints.

[0111] The latent diffusion model employed in this invention, trained on a three-plane latent space, can progressively learn and generate high-quality 3D object geometry and texture data, overcoming the limitations of traditional methods in handling complex 3D data. Through this end-to-end approach, this invention not only achieves a smooth workflow from concept to implementation but also ensures the compatibility of the generated 3D Gaussian field with existing rendering software, enabling its widespread application in fields such as 3D model design and game scene design. Furthermore, experimental results demonstrate the competitive performance of this invention in reconstruction quality and multi-view rendering, maintaining sharp edges and details even at moderate resolutions. In summary, this invention, through its advanced technology and methods, provides an efficient, flexible, and high-quality solution for the creation and editing of 3D content.

[0112] The following describes specific embodiments of the present invention.

[0113] Triplane representation

[0114] See Figure 1 Any object that needs to be modeled is represented as three mutually perpendicular planes with multiple channels. Given any point p, its features can be obtained by querying its projection position on the three planes. The features F_p can then be decoded by the three-plane renderer into the geometric properties TSDF of the point and various properties of the Gaussian field such as transparency and color.

[0115] Three-plane renderer

[0116] See Figure 2 The three-plane renderer comprises two branches, which decode the geometric and Gaussian field properties of the three planes respectively. The working principles of the two branches are as follows:

[0117] Geometric branch: By querying the grid coordinates at a preset resolution, grid features are obtained. A feature decoder based on a fully connected layer then obtains the corresponding fill state and TSDF value for each point. This allows the use of a differentiable Marching Cube algorithm to obtain the mesh of the 3D object. Point clouds are then obtained by sampling the mesh surface.

[0118] Gaussian field attribute branch: The point cloud output from the geometric branch is used to query the three planes to obtain grid features. The corresponding Gaussian field attributes of each point are obtained through a feature decoder based on a fully connected layer, thus obtaining the Gaussian field point cloud. At this point, different virtual camera positions can be set to render the Gaussian field to obtain multi-view images.

[0119] Training the three planes and their renderer: The parameters that need to be trained and updated include the three planes corresponding to each object in the dataset, and the decoder weights in the renderer shared by all objects. The trainable decoder parameters are randomly initialized, and the three plane parameters are initialized to zero values. During training, ground truth TSDF data and multi-view image data are used as supervision signals.

[0120] Generative framework based on three-plane scheme

[0121] See the generative framework based on the three-plane scheme. Figure 3 The main training steps of this framework are as follows:

[0122] 1. Represent each object in the dataset as a three-plane object using the training method described above, and obtain a shared three-plane renderer.

[0123] 2. Use a variational autoencoder (VAE) to further compress the three planes into a hidden code, and obtain a decoder for the hidden code.

[0124] 3. Use a diffusion model to train on this hidden code space.

[0125] The main reasoning steps of this framework are as follows:

[0126] 1. Generate a three-plane hidden code using a trained diffusion model.

[0127] 2. Use the decoder in VAE to decode the three-plane implicit code into a three-plane structure.

[0128] 3. Use a three-plane renderer to decode the three planes into a Gaussian field point cloud and render it.

[0129] Figure 2 The forward pass of the TriRenderer pipeline is shown: rendering three planes into a multi-view image. Figure 3 The overall training process of the 3D content generation framework of the present invention is shown. Figure 2 and Figure 3This paper presents an overview of the generative framework of this invention and its core component, TriRenderer. This invention uses three planes to represent 3D objects and designs a TriRenderer to decode the three planes into GS point clouds and render multi-view images.

[0130] Gaussian Scattering: Gaussian scattering

[17] uses “splats” of point clouds to describe 3D content, each splat being a 3D sphere with a Gaussian distribution shape and other properties such as opacity, color, or spherical harmonic parameters. GS can be rendered like a mesh through specially designed rasterization, which allows for the rendering of realistic scenes in real time. As a newly developed rendering technique, much research has focused on improving its reconstruction quality and optimizing its algorithms [49,7,10,20]. Another related active area includes 4D dynamic GS modeling [43,48,6], scene editing [11,5] and its applications such as SLAM [21,47].

[0131] Diffusion models: Diffusion models [14,9] are powerful generative frameworks in recent years, generating data stepwise through a denoising process. Stable diffusion

[33] is one of the most influential image generators using latent diffusion. In addition, various diffusion models [30,50] have significantly improved image generation performance. These image generators are widely used in the previously mentioned 3D generation methods based on multi-view images. Furthermore, some studies have explored the use of 3D diffusion to generate 3D objects [28,25,38,26] or scenes [16,45,29].

[0132] method

[0133] The training process of the generative framework in this embodiment of the invention can be carried out according to... Figure 3 The process is divided into three parts. First, the training data is encoded into three planes using TriRenderer. Second, the VAE is trained to compress the three planes into the latent space. Third, a diffusion model is trained on the three-plane latent code. During inference, the diffusion model generates the latent code, which can then be decoded by the VAE's decoder into its corresponding three planes. Finally, TriRenderer can decode the three planes into a standard GS point cloud, as described in

[17] , and then render it into an image.

[0134] Gaussian scattering

[17] uses a set of Gaussian points to describe 3D objects. Gaussian points are defined by the complete 3D covariance matrix Σ defined in world space. By point (mean) Centered on [0,1]. Each point has an opacity scalar α∈[0,1] for blending, and A series of spherical harmonic coefficients (SH) represent colors to accurately capture them. Scene perspective dependence. SH series The number of numbers can be calculated as 3 × (n + 1). 2, where n represents the order of SH. A higher n leads to more accurate viewpoint dependence.

[0135] In rendering, given the view transformation W, Gaussian points are projected into 2D, and the covariance transformation is Σ′=JWΣW T J T Σ, where J is the Jacobian matrix of the affine approximation of the projection transformation. To ensure that Σ remains positive semi-definite throughout the training process, Σ can be represented by the Cholesky decomposition as ∑ = RSS. T R T This can be saved as a vector for scaling. Sum represents the unnormalized quaternion of rotation. In summary, each Gaussian point is a joint set of its location p and its GS properties:

[0136]

[0137] GS point clouds can be quickly rendered by the rasterization-based splatting renderer mentioned in

[17] , which is fully differentiable. This renderer is called R.

[0138] GS Field and Three-Plane Renderer

[0139] The three-plane-based GS field: The three planes are three planes formed by each pair of axes (x, y, z) in 3D space, where any 3D point p can be orthogonally projected onto these planes to obtain the corresponding feature Fp = (Fxy, Fxz, Fyz). In the actual implementation, the three planes are a tensor with dimensions 3 × H × H × C, where H represents the resolution of each axis and C is the number of feature channels. A query for any continuous coordinate point should be a bilinear interpolation in the three-plane grid. Considering that the GS is a collection of point clouds with multiple channel attributes, it is natural to encode the GS using the three planes. If the three-plane features and GS attributes of a point are interconvertible, then the sparse GS point cloud of each object can be represented as a continuous three-plane GS field for further encoding. Using independent channels to encode 3D geometry and other GS attributes, as shown in Equation 1, is an experimental setup conducted to achieve better convergence during training.

[0140] There are three advantages to using a triplane to represent GS. First, it has higher memory efficiency compared to dense voxels. Second, experiments show that triplanes have sufficient expressive power, and the nonlinear decoder inside TriRenderer works well. Third, compared to the original sparse GS point cloud, triplanes are more suitable for use with convolution-based encoders. The original GS point cloud is usually concentrated on the surface of thin objects, occupying a small 3D space. Even with 3D convolution, the information cannot be effectively captured. Using triplanes to project 3D thin sheets onto a plane is a more reasonable compression method.

[0141] TriRenderer: The TriRenderer is key to converting three planes back to the original GS point cloud. For example... Figure 3 As shown, TriRenderer consists of a GS decoder and a GS renderer R. The GS decoder has a geometry decoding branch (Dgeo) and another GS attribute decoder. The code branch is Dgs. Dgeo is responsible for recovering the point cloud from the three planes, and then Dgs uses the obtained data... Point cloud query The three planes are used to obtain the GS attributes corresponding to each point. This allows the GS point cloud to be retrieved in its original format and rendered using the GS renderer. It's worth noting that all three planes of different objects share a common three-plane renderer, TriRenderer, which ensures that features on different three planes are distributed in a similar way.

[0142] The geometry is represented using a truncated signed distance function (TSDF) so that Dgeo decodes Fp into binary occupied logic op and the corresponding TSDF value sp. The TSDF volume of the original object can be reconstructed by querying all grid coordinates of a specified resolution L×L×L. To obtain the point cloud of GS, a triangular mesh is first extracted from the TSDF volume using the Marching Cube algorithm, and then the point cloud can be sampled at arbitrary density on the mesh surface. A fully differentiable Marching Cube

[42] and an interpolation-based point sampler are used so that gradients from the GS branches can be backpropagated to optimize the geometry. GS features are obtained by querying the three-plane GS channels using the sampled point cloud and then converted into GS attributes using Dgs. The GS renderer can then render the GS point cloud into a multi-view image. Since the different GS attributes have different numerical scales and distributions, their respective headers are customized to decode them separately, similar to

[52] .

[0143] Training: The trainable modules include the Triplanes and the decoder within the TriRenderer, with TSDF X and N multi-view images {Ii}Ni=1 as supervision. GS point clouds do not need to be pre-trained as in

[17] . In the first few epochs, L1 loss and cross-entropy are used to regress TSDF values ​​and binary occupancy O, where 1 indicates that the mesh is occupied. For the GS decoding branch, the point cloud is sampled using the real TSDF, and L1 pixel loss and SSIM loss

[41] are used as rendering losses. When the occupancy accuracy reaches a threshold, the rendering loss is enabled to backpropagate to the geometry branch to optimize the TSDF values. The geometry loss LGeo and the rendering loss LRender are as shown in Equation 2 and Equation 3.

[0144] As shown in Equation 3.

[0145]

[0146] in and These are the predicted binary occupied voxels and TSDF voxels, respectively. ⊙⊙ represents element-wise multiplication, and sg is obtained through... The gradient stopping function is used for the precision threshold operation. λ1, λ2, λ3, and β are fixed weights used for loss balancing.

[0147] Three-plane VAE

[0148] A VAE with a UNet-like structure was used to compress the three planes into the latent space. The training process is as follows: Figure 5 As shown in (a). Considering that the three planes of the triplane encode features from three orthogonal view directions, these three planes should be uniform data. Therefore, the triplanes with a batch size of B are reshaped into a new batch B×3×H×H×C to 3B×H×H×C for VAE training. Fully decoupled encoders and decoders are used for geometry and other GS attributes because experiments have shown that mixed encoding can lead to blurred rendering results. After the triplane reconstruction, the trained TriRenderer can be used as previously described to retrieve the GS point cloud and render it into an image. Figure 5 The three-plane VAE and two-stage potential diffusion are shown.

[0149] During training, the same loss function as when training the tri-plane renderer was used, but the L1 loss between the input and reconstructed tri-planes was not included. This is because the L1 loss equally penalizes the information and noise of the tri-planes, and using the same loss objective as in Section 3.3 is more efficient. In addition, a Kullback-Leibler divergence loss L was added. KL This constrains the latent space to prevent it from deviating too far from a normal distribution. The total loss can be summarized as follows:

[0150] LVAE =L Geo +L Render +γL KL (4)

[0151] Here, γ is a small weight. The encoder and decoder for geometry and GS attributes are trained in stages. First, the encoder and decoder for geometry are trained, and then the encoder and decoder for GS attributes are trained using rendering loss. Mesh extraction and point sampling are based on TSDF reconstruction, rather than on the real-world scenario.

[0152] Conditional two-stage potential diffusion generation

[0153] Since the implicit codes for geometry and other GS attributes are decoupled, a two-stage diffusion process is proposed to generate the geometric implicit code and the GS attribute implicit code sequentially, such as... Figure 5 As shown in (b). Thus, the geometric implicit code can serve as a new condition for the second stage. To simplify the theoretical content, the input samples for both stages are represented as v0, and all conditions are represented by y.

[0154] Latent diffusion is implemented under the condition of object category labels, following DDPM [14,33]. To better capture the relationship between different planes, the three-plane latent B×3×h×h×c with batch size B is unfolded into an image B×3h×h×c as input v0 to the diffusion model. The generated result is converted back to the original shape for TriRenderer decoding. Category labels are mapped to integers as diffusion conditions.

[0155] DDPM

[14] progressively mixes the input sample v0 with Gaussian noise until it becomes a white Gaussian noise volume vT ~ N(0,1) in step T. At each step t, the sample vt is added with independent and identically distributed Gaussian noise with variance βt and then... Scale the sample vt-1 from the previous step to obtain:

[0156]

[0157] This is also known as the forward direction. On the other hand, the reverse process can be described as:

[0158]

[0159] The diffusion model is trained to reverse the forward process and predict μ. θ for:

[0160]

[0161] in And ∈ θ It is a neural network with a parameter set θ, and it has a structure similar to UNet.

[0162] Mean squared error loss is used to supervise noise prediction:

[0163]

[0164] experiment

[0165] Experimental basis Figure 3 The proposed generation process is carried out step by step. First, the performance of the proposed three-plane representation of the GS field is tested. Then, VAE and diffusion models are trained to explore 3D generation. The OmniObject3D

[44] dataset is mainly used for evaluation, and in Figure 1 Supplementary visualizations are provided using a small subset of ShapeNetCore. Data preprocessing and implementation details are included in the supplementary materials.

[0166] Three-plane fitting

[0167] Figure 6 The channel visualization of the three planes of the sample is shown.

[0168] Tri-plane visualization: To better examine the rationale for using a convolution-based method to encode the three planes, the channel values ​​of the trained three planes are simply scaled to the pixel range and visualized, such as... Figure 7 As shown, the clear shapes can be observed in three different views.

[0169] Tri-plane rendering quality: The rendering quality of different methods is compared in Table 1, using Peak Signal-to-Noise Ratio (PSNR) as the evaluation metric. Tri-plane indicates that, with smaller data sizes, especially when differentiable Marching Cube (DiffMC) is enabled to involve geometrically supervised rendering loss, it outperforms voxel-based methods. Rendering loss is useful for geometric optimization because even the ground truth of the TSDF can be affected by quantization errors, which can be amplified during rendering. Voxelized GS means that the GS is trained on voxels at a fixed resolution. Thus, higher resolutions can yield better rendering results, but computational and memory complexity increase exponentially. Furthermore, the rendered samples... Figure 7 Listed for qualitative evaluation. The method of the present invention allows for better detail, such as sharper edges, even at moderate resolution.

[0170] Table 1: Three-plane representation of reconstruction quality in the GS field

[0171]

[0172] 3D object generation

[0173] Prepare three planes as training data: Train a VAE on the three planes, and then train a two-stage diffusion model in the latent space, as follows. Figure 5 As shown. In order to convert all 3D objects into a triplane and its corresponding shared TriRenderer, the triplane and TriRenderer must be trained together. However, the OmniObject3D dataset contains approximately 6000 objects. To speed up the training process, a strategy similar to

[35] was adopted. Initially, a mini-batch of 500 objects was randomly selected to train the TriRenderer, and then its parameters were frozen to train the triplane of the remaining objects.

[0174] Training the three-plane VAE: The three planes are compressed into the latent space using a downsampling factor of 4. Slight blurring is observed in the reconstructed image, which is unavoidable but acceptable.

[0175] Generated via diffusion: Some multi-view rendering sample columns in Figure 8 The results of this invention were compared with those of existing advanced image / text to 3D modeling methods such as One-2-3-45++

[18] and TripoSR

[37] . Figure 9 This shows the image quality degradation in a VAE. Left: Real-world scenario. Right: VAE reconstruction.

[0176] Ablation Research

[0177] 3D diffusion based on voxel-based Gaussian scattering: An attempt was made to generate GS properties under given voxel occupancy conditions using a 3D diffusion model. However, even toy experiments with overfitting to a single object failed, which may be attributed to complex multichannels with different distributions, particularly non-Euclidean ones such as quaternions. Figure 10 This indicates that GS attribute generation failed. Figure 10 As shown, the diffusion model learns to generate colors, but fails to generate splat scaling, opacity, and orientation.

[0178] Three-plane diffusion without VAE: Since VAE inherently involves reconstruction loss, we attempt to generate samples directly using diffusion in a three-plane space. Randomly generated samples are... Figure 11 Visualization in Chinese. Figure 11 The results show the noise generated by direct diffusion across the three planes (conditional categories are "train", "watermelon", and "box"). Experimental results show that this direct diffusion can lead to significant noise in the decoding geometry. A possible reason is that the multi-channels of the three planes contain considerable redundancy or noise, which is difficult for the diffusion model to capture or filter.

[0179] Inference efficiency: Inference efficiency is listed in Table 2. This experiment was conducted on a platform equipped with an RTX 3090 GPU with 24GB of memory.

[0180] Table 2: Inference efficiency of the generation process, with a batch size of 1

[0181]

[0182] Other experimental details

[0183] Dataset and Preprocessing: Experiments were conducted using the OmniObject3D

[44] dataset, which contains 6,000 scanned objects across 190 everyday categories. Since the proposed method is a generative model, all data were used as training data. The dataset contains 100 multi-view images and a colored mesh rendered by Blender for each object. In addition to these images, TriRenderer performed geometric regression via TSDF volumes. All object meshes were normalized by centering all object meshes and scaling them to [-1.2, 1.2] bounding boxes across all 3 axes before generating the TSDF volumes. To speed up TSDF computation, the high-resolution mesh was first simplified to a low-resolution mesh with fewer than 200,000 faces, and the SDF volumes were then derived using the open-source software SDFGen[1]. For efficiency, the fast scan algorithm within the truncated region was omitted. The voxelization resolution was 128×128×128, and the truncation distance was 3 voxels.

[0184] Implementation details: In TriRenderer training, all loss terms are weighted equally. In VAE training, a KL loss with a weight of 5e-8 is added. As for the implementation of the diffusion model, Diffusers

[39] is used as the code base. The time steps are encoded with α conditions[4] to stabilize training and enable parameter fine-tuning during inference. In training, a 2000-step cosine noise scheme is used for diffusion. In the generation process, the DDIM pipeline is used with 300 iterations.

[0185] Figure 12 More generated samples are shown.

[0186] Summarize

[0187] This invention proposes a novel Gaussian scattering field generation framework, TriGSGen. TriGSGen consists of three main parts: 1) a lightweight three-plane representation for 3D objects with Gaussian scattering format; 2) a fully differentiable TriRenderer that decodes the three planes back to the original Gaussian scattering field (GS) and renders it into a multi-view image; and 3) a three-plane VAE and a staged diffusion model for the entire generation process. By using TriGSGen, complex GS data can be efficiently encoded, enabling effective generation. Rendering quality depends on various factors, such as three-plane resolution and channel count, VAE compression ratio, and TSDF resolution.

[0188] This invention also provides a storage medium for storing a computer program, which, when executed, performs at least the methods described above.

[0189] This invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor executes the computer program by performing at least the method described above.

[0190] This invention also provides a processor that executes a computer program, at least performing the methods described above.

[0191] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk drive or magnetic tape drive. The storage media described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0192] In the several embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0193] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0194] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0195] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0196] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0197] The methods disclosed in the several method embodiments provided by this invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0198] The features disclosed in the several product embodiments provided by this invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0199] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0200] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or application, should be considered within the scope of protection of the present invention.

[0201] References

[0202] [1] Christopher Batty. Signed distance field generator for triangular meshes, 2015.

[0203] [2] Tianshi Cao, Karsten Kreis, Sanja Fidler, Nicholas Sharp, and Kangxue Yin. Texfusion: Synthesizing 3D Textures Using a Text-Guided Image Diffusion Model. Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 4169-4181, 2023.

[0204] [3] Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, ​​Zhuowen Tu, Lingjie Liu, and Hao Su. Single-stage diffuse neural fields: a unified approach to 3D generation and reconstruction. Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 2416-2425, 2023.

[0205] [4] Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan. Wavegrad: Estimating the gradient of waveform generation. arXiv preprint arXiv:2009.00713, 2020.

[0206] [5] Yiwen Chen, Zilong Chen, Chi Zhang, Feng Wang, Xiaofeng Yang, Yikai Wang, Zhongang Cai, Lei Yang, Huaping Liu, and Guosheng Lin. Gaussian editor: Fast and controllable 3D editing using Gaussian scattering. arXiv preprint arXiv:2311.14521, 2023.

[0207] [6] Yurui Chen, Chun Gu, Junzhe Jiang, Xiatian Zhu, and Li Zhang. Periodic Oscillating Gaussians: Dynamic Urban Scene Reconstruction and Real-Time Rendering. arXiv preprint arXiv:2311.18561, 2023.

[0208] [7] Kai Cheng, Xiaoxiao Long, Kaizhi Yang, Yao Yao, Wei Yin, Yuexin Ma, Wenping Wang, and Xuejin Chen. Gaussianpro: 3D Gaussian scattering with asymptotic propagation. arXiv preprint arXiv:2402.14650, 2024.

[0209] [8] Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam, Jaerin Lee, and Kyoung MuLee. Luciddreamer: Domain-unrestricted 3D Gaussian scattering scene generation. arXiv preprint arXiv:2311.13384, 2023.

[0210] [9] Prafulla Dhariwal and Alexander Nichol. Diffusion models outperform GANs in image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.

[0211]

[10] Zhiwen Fan, Kevin Wang, Kairun Wen, Zehao Zhu, Dejia Xu, and Zhangyang Wang. Lightgaussian: Unbounded 3D Gaussian compression with a compression ratio of up to 15x and over 200 fps. arXiv preprint arXiv:2311.17245, 2023.

[0212]

[11] Jiemin Fang, Junjie Wang, Xiaopeng Zhang, Lingxi Xie, and QiTian. Gaussian editor: Fine editing of 3D Gaussians using text commands. arXiv preprint arXiv:2311.16037, 2023.

[0213]

[12] Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: Generative models for learning high-quality 3D texture shapes from images. Advances in Neural Information Processing Systems, 35:31841–31854, 2022.

[0214]

[13] Anchit Gupta, Wenhan Xiong, Yixin Nie, Ian Jones, and Barlas 3DGen: Texture mesh generation using three-plane latent diffusion. arXiv preprint. arXiv:2303.05371,2023.

[0215]

[14] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probability model. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.

[0216]

[15] Lukas Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nieβner. Text2room: Extracting Textured 3D Meshes from 2D Text to Image Models. In the Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 7909–7920, 2023.

[0217]

[16] Xiaoliang Ju, Zhaoyang Huang, Yijin Li, Guofeng Zhang, Yu Qiao, and Hongsheng Li. Diffindscene: High-quality 3D indoor scene generation based on diffusion, 2023.

[0218]

[17] Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3D Gaussian scattering rendered in real time. ACM Graphics Trade, 42(4):1–14, 2023.

[0219]

[18] Minghua Liu, Ruoxi Shi, Linghao Chen, Zhuoyang Zhang, Chao Xu, Xinyue Wei, Hansheng Chen, Chong Zeng, Jiayuan Gu, and Hao Su. One-2-3-45++: Fast single-image to 3D object with consistent multi-view generation and 3D diffusion. arXiv preprint arXiv:2311.07885,2023.

[0220]

[19] Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero1-to-3: Zero-sample images to 3D objects. Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 9298–9309, 2023.

[0221]

[20] Tao Lu, Mulin Yu, Linning Xu, Yuanbo Xiangli, Limin Wang, Dahua Lin, and Bo Dai. Scaffold-gs: Structured 3D Gaussian for Adaptive View Rendering. arXiv preprint arXiv:2312.00109,2023.

[0222]

[21] Hidenobu Matsuki, Riku Murai, Paul HJ Kelly, and Andrew J Davison. Gaussian scattering SLAM. arXiv preprint arXiv:2312.06741,2023.

[0223]

[22] Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. Latent-NeRF for latent neural radiation fields in 3D shape and texture generation. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 12663–12673, 2023.

[23] Ben Mildenhall, Pratul PSrinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: View synthesis of scenes represented by neural radiation fields. ACM Communications, 65(1):99–106, 2021.

[0224]

[24] Paritosh Mittal, Yen-Chi Cheng, Maneesh Singh, and Shubham Tulsiani. AutoSDF: Shape Priors for 3D Completion, Reconstruction, and Generation. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 306–315, 2022.

[0225]

[25] Norman Müller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulò, Peter Kontschieder, and Matthias Nieβner. DiffRF: Render-guided 3D radiation field diffusion. In the Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 4328–4338, 2023.

[26] Gimin Nam, Mariem Khlifi, Andrew Rodriguez, Alberto Tono, Linqi Zhou, and Paul Guerrero. 3D-LDM: Neural implicit 3D shape generation using a latent diffusion model. arXiv preprint arXiv:2212.00842, 2022.

[0226]

[27] Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-E: A system for generating 3D point clouds from complex cues. arXiv preprint arXiv:2212.08751,2022.

[0227]

[28] Evangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang, Luc V Gool, and Sergey Tulyakov. Autodecoding latent 3D diffusion models. Advances in Neural Information Processing Systems, 36:67021–67047, 2023.

[0228]

[29] Ryan Po and Gordon Wetzstein. Combined 3D scene generation using local conditional diffusion. arXiv preprint arXiv:2303.12218,2023.

[0229]

[30] Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. SDXL: Improved latent diffusion model for high-resolution image synthesis. arXiv preprint arXiv:2307.01952,2023.

[0230]

[31] Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. DreamFusion: Text to 3D using 2D diffusion. arXiv preprint arXiv:2209.14988,2022.

[0231]

[32] Elad Richardson, Gal Metzer, Yuval Alaluf, Raja Giryes, and Daniel Cohen-Or. Texture: Text Guide Textures for 3D Shapes. Proceedings of ACM SIGGRAPH 2023, pp. 1–11, 2023.

[0232]

[33] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Ommer. High-resolution image synthesis and latent diffusion models. In the proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.

[0233]

[34] Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. MVDream: Multi-view diffusion for 3D generation. arXiv preprint arXiv:2308.16512,2023.

[0234]

[35] J Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein. Generation of 3D Neural Fields Using Triplanar Diffusion. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 20875–20886, 2023.

[0235]

[36] Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. DreamGaussian: Generative Gaussian Scattering for Efficient 3D Content Creation. arXiv preprint arXiv:2309.16653,2023.

[0236]

[37] Dmitry Tochilkin, David Pankratz, Zexiang Liu, Zixuan Huang, Adam Letts, Yangguang Li, Ding Liang, Christian Laforte, Varun Jampani, and Yan-PeiCao. TripoSR: Fast reconstruction of 3D objects from a single image. arXiv preprint arXiv:2403.02151,2024.

[0237]

[38] Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, Karsten Kreis, et al. Lion: A latent point diffusion model for 3D shape generation. Advances in Neural Information Processing Systems, 35:10021–10039, 2022.

[0238]

[39] Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf. Diffusers: State-of-the-art diffusion models. https: / / github.com / huggingface / diffusers, 2022.

[0239]

[40] Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: Generative Models for Sculpting 3D Digital Heads Using Diffusion. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 4563–4573, 2023.

[0240]

[41] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Image Processing Trade, 13(4):600–612, 2004.

[0241]

[42] Xinyue Wei, Fanbo Xiang, Sai Bi, Anpei Chen, Kalyan Sunkavalli, Zexiang Xu, and Hao Su. Neumanifold: Neural watertight manifold reconstruction with efficient and high-quality rendering support. arXiv preprint arXiv:2305.17134,2023.

[0242]

[43] Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, WeiWei, Wenyu Liu, Qi Tian, ​​and Xinggang Wang. 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering. arXiv preprint arXiv:2310.08528, 2023.

[0243]

[44] Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Liang Pan, Jiawei Ren, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, Dahua Lin, and Ziwei Liu. OmniObject3D: A large vocabulary of 3D object datasets for reality perception, reconstruction, and generation. IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023.

[0244]

[45] Zhennan Wu, Yang Li, Han Yan, Taizhang Shang, Weixuan Sun, Senbo Wang Ruikai Cui, Weizhe Liu, Hiroyuki Sato, Hongdong Li, et al. Blockfusion: Scalable 3D scene generation using latent triplanar extrapolation. arXiv preprint arXiv:2401.17053,2024.

[0245]

[46] Zijie Wu, Yaonan Wang, Mingtao Feng, He Xie, and Ajmal Mian. Sketch and text-guided diffusion models for color point cloud generation. Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 8929–8939, 2023.

[0246]

[47] Chi Yan, Delin Qu, Dong Wang, Dan Xu, Zhigang Wang, Bin Zhao, and Xuelong Li. GS-SLAM: Dense Visual SLAM Using 3D Gaussian Scattering. arXiv preprint arXiv:2311.11700,2023.

[0247]

[48] ​​Zeyu Yang, Hongye Yang, Zijie Pan, Xiatian Zhu, and Li Zhang. Real-time realistic dynamic scene representation and rendering with 4D Gaussian scattering. arXiv preprint arXiv:2310.10642,2023.

[0248]

[49] Zehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler, and Andreas Geiger. Mip-Splatting: Alias-free 3D Gaussian scattering. arXiv preprint arXiv:2311.16493, 2023.

[50] Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to a text-to-image diffusion model. Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 3836–3847, 2023.

[0249]

[51] Linqi Zhou, Yilun Du, and Jiajun Wu. 3D shape generation and completion via point-voxel diffusion. Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 5826–5835, 2021.

[0250]

[52] Zi-Xin Zou, Zhipeng Yu, Yuan-Chen Guo, Yangguang Li, Ding Liang, Yan-Pei Cao, and Song-Hai Zhang. Three planes encountering Gaussian scattering: fast and general single-view 3D reconstruction using a transformer. arXiv preprint.

Claims

1. A method for generating 3D Gaussian fields based on a triplane, characterized in that, Includes the following steps: S1. Input 3D object data and standardize it to meet the requirements of three-plane representation; S2. Use a three-plane encoder to encode the geometric and Gaussian properties of a 3D object onto three mutually perpendicular planes to form a three-plane mesh; S3. Store 3D object information captured from different perspectives on three vertical planes to construct a complete three-plane mesh; S4. Train the geometry rendering branch and Gaussian field property rendering branch of the triplane renderer to implement decoding from triplane mesh to 3D mesh and Gaussian property; S5. Use a three-plane decoder to decode the geometry and Gaussian properties of 3D objects from a three-plane mesh; S6. Train the encoder and decoder of the variational autoencoder (VAE), compress the three-plane representation into an implicit code, and reconstruct the three-plane representation from the implicit code; S7. Use the trained VAE encoder to compress the three-plane representation into a hidden code; S8. Construct and train a diffusion model in the implicit code space, including a conditional diffusion module and a noise injection mechanism, to learn the data generation process; S9. Using the trained diffusion model, the implicit code of a new 3D object is generated through a reverse diffusion process; S10. Use the VAE decoder to decode the generated implicit code back to the three-plane representation; S11. Use a three-plane renderer to render the decoded three-plane representation into a 3D Gaussian field point cloud, and further generate multi-view images.

2. The 3D Gaussian field generation method based on a three-plane model as described in claim 1, characterized in that, In step S2, the features of any 3D point are obtained by orthogonal projection using three planes formed by the x, y, and z axes in 3D space, and a continuous three-plane GS field is constructed. A tensor of size 3×H×H×C is implemented to store the feature information on the three planes, where H represents the resolution of each axis and C represents the number of feature channels.

3. The method for generating 3D Gaussian fields based on a three-plane model as described in claim 1 or 2, characterized in that, In step S3, bilinear interpolation is performed on any continuous coordinate point in a three-plane grid to obtain the point's characteristics.

4. The 3D Gaussian field generation method based on a three-plane plane as described in any one of claims 1 to 3, characterized in that, In step S4, the training of the two branches of the three-plane renderer specifically includes: S4.1 Geometric Branching Training: Grid point features are obtained using a grid point coordinate query system with a preset resolution; The grid features are processed by the first feature decoder based on the fully connected layer to obtain the filling state and TSDF value of each point. A differentiable Marching Cube algorithm is applied to generate the mesh of a 3D object based on the TSDF value; Perform sampling operations on the generated mesh surface to produce a point cloud for subsequent processing; S4.2 Gaussian field property branch training: Using the point cloud obtained in step S4.1, query the three planes to obtain the corresponding grid features; The Gaussian field properties of each point are obtained based on the grid features by using a second feature decoder based on a fully connected layer. Construct a Gaussian field point cloud based on the acquired Gaussian field properties; Multiple virtual camera positions are set, and a three-plane renderer is used to render the Gaussian field point cloud from multiple perspectives to generate multi-view images.

5. The 3D Gaussian field generation method based on a three-plane model as described in claim 4, characterized in that, In step S4, the training of the three-plane renderer also includes: S4.3 initializes the three-plane parameters corresponding to each object in the dataset, setting the initial value to zero; S4.4 randomly initializes the decoder weights in the common renderer; During training, S4.5 uses truncated symbolic distance function (TSDF) data and multi-view image data as supervision signals to train and update the three-plane parameters and decoder weights.

6. The method for generating 3D Gaussian fields based on a triplane as described in any one of claims 1 to 5, characterized in that, In step S4, a set of Gaussian spheres with positional mean, covariance, and three-axis scale are used to describe the 3D object, where each point is equipped with an opacity scalar and color information; the color of each Gaussian point is represented using spherical harmonic coefficients; during rendering, the Gaussian points are projected into two-dimensional space, and the covariance matrix is ​​adjusted to adapt to the observation transformation; the covariance matrix is ​​represented by Cholesky decomposition to ensure its positive semidefiniteness throughout the training process.

7. The 3D Gaussian field generation method based on a three-plane model as described in claim 6, characterized in that, In step S5, a decoder is used to decode the geometry and Gaussian properties of the 3D object from the three-plane mesh, including the opacity scalar and color information represented by spherical harmonics; the position and properties of each Gaussian point are integrated into a unified data structure; and the decoded Gaussian point cloud is rendered into an image using a rasterization-based splatting renderer that supports differentiable operations.

8. The method for generating 3D Gaussian fields based on a three-plane plane as described in any one of claims 1 to 7, characterized in that, In step S4, the geometry is represented by the truncated signed distance function (TSDF), and the triangular mesh is extracted from the TSDF volume and the point cloud is sampled using the Marching Cube algorithm. The gradient backpropagation is achieved by using a fully differentiable Marching Cube algorithm and an interpolation-based point sampler.

9. The method for generating a 3D Gaussian field based on a three-plane plane as described in any one of claims 1 to 8, characterized in that, In step S4, during the training of the three-plane renderer, the geometric loss is calculated using L1 loss and cross-entropy regression of TSDF value and binary occupancy, and the rendering loss is calculated using L1 pixel loss and SSIM loss; when the occupancy accuracy reaches a set threshold, backpropagation of rendering loss is enabled to optimize TSDF value.

10. The method for generating a 3D Gaussian field based on a three-plane plane as described in any one of claims 1 to 9, characterized in that, In step S6, a VAE structure similar to UNet is used to compress the three planes into the latent space, where the three planes are regarded as uniform data encoding features from three orthogonal view directions; the three plane data with a batch size of B is reshaped into a new batch format, changing from B×3×H×H×C to 3B×H×H×C.

11. The method for generating 3D Gaussian fields based on a triplane as described in any one of claims 1 to 10, characterized in that, In step S6, binary cross-entropy loss (BCE) and L1 loss are used to regress and predict binary occupied voxel grids and TSDF voxel grids. Rendering loss is calculated using L1 pixel loss and SSIM loss, and KL divergence loss is introduced to constrain the latent space, but L1 loss between the input and reconstructed three planes is not used. Mesh extraction and point sampling are performed using a reconstruction method based on truncated symbolic distance function (TSDF).

12. The method for generating 3D Gaussian fields based on a triplane as described in any one of claims 1 to 11, characterized in that, In step S8, based on the decoupled geometric and Gaussian property implicit coding, reverse diffusion is performed in two stages to generate geometric implicit coding and GS property implicit coding in sequence. The first stage generates geometric latent codes: During the first stage of training, the input samples are mixed with Gaussian noise through a diffusion model, and the UNet in the diffusion model is used to predict the noise and perform regression training. The second stage generates the Gaussian attribute implicit code: In the second stage, the geometric implicit code generated in the first stage is used as a condition to perform a diffusion process on the true Gaussian attribute implicit code. The UNet in the diffusion model is used to predict noise and perform regression training under this condition. The three-plane latent representation is unfolded into an image and used as input to the diffusion model, where the class labels are mapped to integers as diffusion conditions; In step S9, the model trained in two stages is used to perform the reverse diffusion process, predict noise and remove it step by step to reverse the forward diffusion process and generate the hidden code of the desired 3D object. In step S10, the generated hidden code will be converted back to its original shape and decoded using the TriRenderer trained in step S4 to render a multi-view image.

13. The 3D Gaussian field generation method based on a three-plane model as described in claim 12, characterized in that, In steps S8 and S9, mean squared error loss is used to supervise noise prediction. This loss measures the difference between the noise predicted by the model and the actual noise, and the model is trained by minimizing this difference.

14. A 3D Gaussian field generation method based on a triplane for use in the training phase of a model, characterized in that... Includes the following steps: T1. Input 3D object data and standardize it to meet the requirements of three-plane representation; T2. A three-plane encoder is used to encode the geometric and Gaussian properties of a 3D object onto three mutually perpendicular planes to form a three-plane mesh; T3. Store 3D object information captured from different perspectives on three vertical planes to construct a complete three-plane mesh; T4. Train the two branches of the triplane renderer: the geometry rendering branch and the Gaussian field property rendering branch, so that it can decode from the triplane mesh to the 3D mesh and Gaussian property; T5. Training the three-plane decoder to decode the three-plane mesh to accurately recover the geometry and Gaussian properties of the 3D object; T6. Train a VAE model, including an encoder and a decoder, to compress the three-plane representation into an implicit code and be able to reconstruct the three-plane representation from the implicit code; T7. Training the VAE encoder compresses the three-plane representation into implicit codes to prepare data for training the diffusion model; T8. Construct a diffusion model in the implicit code space, including a conditional diffusion module and a noise injection mechanism, so that the model can learn the data generation process; T9. Training the diffusion model generates the implicit code of new 3D objects through the reverse diffusion process; T10. The decoder trained on the VAE decodes the implicit code back to a three-plane representation and renders it using a three-plane renderer to generate 3D Gaussian field point clouds and multi-view images.

15. A 3D Gaussian field generation method based on a three-plane model for the inference stage, characterized in that... Includes the following steps: R1. Reverse Diffusion Process: Using the trained diffusion model, a new 3D object's implicit code is generated through the reverse diffusion process; R2. Implicit Code Decoding: Using the VAE decoder, the implicit code obtained during the inverse diffusion process is decoded back into a three-plane representation to reconstruct the geometric and Gaussian properties of the 3D object; R3. Triplane Rendering: Uses a triplane renderer to render the decoded triplane representation and generate a 3D Gaussian field point cloud; R4. Multi-view image generation: Generate multi-view images of objects based on the rendered 3D Gaussian field point cloud to achieve a comprehensive visual presentation; R5. Output: The generated 3D Gaussian field point cloud or multi-view image will be used as the final output.

16. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the 3D Gaussian field generation method based on any one of claims 1-10.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the 3D Gaussian field generation method based on any one of claims 1-10.