Multi-level surface scattering material rendering method based on conditional variation auto-encoder

The CVAE-based BSSRDF neural network addresses slow and inaccurate rendering of subsurface scattering by accurately modeling local geometry and multi-layer materials, enhancing rendering efficiency and accuracy.

CN120318398APending Publication Date: 2025-07-15SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510400685.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing computer rendering subsurface scattering method has slow calculation speed and high noise results, making it difficult to accurately simulate the local shape of the three-dimensional grid and the parameters of multi-layer materials, resulting in a large difference between the rendering effect and the real effect.

Method used

A multi-level surface scattering material rendering method based on conditional variational autoencoder is adopted, and a neural network is used as a BSSRDF model, combining the implicit SDF model and residual network structure, and by training the ray data set and the geometric feature data set, the transmittance and incident position of the rays are predicted to achieve efficient rendering.

Benefits of technology

The accurate expression of local geometric features of the three-dimensional mesh and effective rendering of multi-layer materials is achieved. The rendering results are close to the accuracy of the volume path tracking method, and the rendering efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318398A_ABST
    Figure CN120318398A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-level surface scattering material rendering method based on a conditional variation auto-encoder, and the method comprises the steps: training a neural network, which is called as a BSSRDF model, through employing simulation data, taking the trained neural network as a bidirectional subsurface reflection distribution function BSSRDF for path tracking rendering, and in the process of path tracking rendering, carrying out the reconstruction of the BSSRDF model. And when the light intersects with the three-dimensional grid with the multi-layer surface scattering material, inputting information of intersection points into the distribution function to predict an incidence point and transmissivity of the light in the next step, and reprojecting the predicted incidence point of the light by using an implicit SDF model to obtain an efficient and real rendering result. According to the method, the local geometric features of the three-dimensional grid can be accurately expressed, the BSSRDF model is trained by using the generated light data set, and the transmissivity, the incident position and the incident direction of the light in subsurface scattering are accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of subsurface scattering materials in computer rendering, and in particular to a method for rendering multi-level surface scattering materials based on conditional variational autoencoders. Background Art

[0002] There are many objects in reality with subsurface scattering materials, such as human skin, fruits, soap, clouds, etc. Subsurface scattering makes these objects present a soft and semi-transparent appearance. How to efficiently and accurately simulate this special phenomenon in computer rendering has always been an important issue.

[0003] Existing offline rendering schemes for subsurface scattering mainly adopt the method of volume ray tracing. This method accurately simulates the effect of subsurface scattering by tracking the specific scattering direction and position of light each time during the subsurface scattering process. However, the calculation speed of this rendering method is very slow, and at the same time, the noise of the rendering result is large, and a large number of repeated samplings are required to obtain better results. For real-time rendering of subsurface scattering, the main method is to use empirical BSSRDF models, but these models often deviate greatly from the real effect. In recent years, with the wide application of related technologies such as deep learning in computer graphics, there have also emerged some methods based on deep learning to accelerate volume ray tracing or train BSSRDF models. However, in terms of subsurface scattering, the following problems still exist:

[0004] 1. Subsurface scattering is affected by the local shape of the 3D mesh. In order to accurately predict the scattered position in the BSSRDF model, a geometric feature that can accurately express the local surface shape and spatial distribution is required;

[0005] 2. For some objects, such as human skin, there may be complex multi-layer materials, and the parameters of each layer of subsurface scattering material are different. There is a lack of a reasonable and effective method for generating a rendering scene. Summary of the Invention

[0006] The purpose of the present invention is to overcome the deficiencies of the prior art, and provide a method for rendering multi-level surface scattering materials based on conditional variational autoencoders, which can accurately express the local geometric features of the 3D mesh, and use the generated light dataset to train a BSSRDF model to accurately predict the transmittance, incident position and direction of light in subsurface scattering.

[0007] To achieve the above object, the technical solution provided by the present invention is: a multi-level surface scattering material rendering method based on a conditional variational autoencoder, which trains a neural network, called the BSSRDF model, using simulation data, and uses the trained neural network as the bidirectional subsurface reflectance distribution function BSSRDF for path tracing rendering. During the path tracing rendering process, when a ray intersects a three-dimensional mesh with a multi-level surface scattering material, the information of the intersection point is input into this distribution function to predict the next incident point and transmittance of the ray, and the predicted ray incident point is reprojected using the implicit SDF model to obtain an efficient and realistic rendering result; wherein, the feature input of the BSSRDF model is designed according to the characteristics of the simulation method, including ray information, geometric features, and material features, and its network structure is designed based on the conditional variational autoencoder CVAE framework and the residual network structure. The geometric features are calculated by designing an implicit SDF model, and this SDF model is also designed based on the residual network structure;

[0008] The specific implementation of this multi-level surface scattering material rendering method includes the following steps:

[0009] 1) Use the collected three-dimensional meshes to construct a dataset: perform random rotation and scaling transformations on each three-dimensional mesh to generate multiple transformed three-dimensional meshes. On this basis, construct a geometric feature dataset and a ray dataset: the geometric feature dataset samples multiple groups of feature points on each three-dimensional mesh and initializes a geometric feature latent vector for each group of feature points; the ray dataset is generated by randomly generating multiple pairs of "three-dimensional mesh, material parameter" data pairs. According to the thickness of each layer in the multi-layer material of each data pair, the corresponding three-dimensional mesh is shrunk inward to construct a scene for collecting simulation data. In these scenes, multiple ray data are collected using the simulation method as the ray dataset, and it is divided into a training set and a test set;

[0010] 2) Batch input the training set constructed in step 1) into the designed BSSRDF model to obtain the predicted incident point and transmittance, calculate the loss using the designed loss function, and use the stochastic gradient descent method to optimize the network weights. Adjust the weights of each loss and the learning rate according to the loss change and network performance, and finally obtain the optimal BSSRDF model;

[0011] 3) On the scenes of the test set in step 1), use the BSSRDF model and the SDF model to combine with the volume path tracing method for rendering, and obtain a rendering result close to the volume path tracing method in the scene with a multi-level surface scattering material.

[0012] Further, in step 1), when generating the geometric feature dataset, a plurality of target points are sampled on the transformed three-dimensional grid surface, a geometric feature latent vector is randomly initialized for each target point, and a plurality of feature points are further sampled within a cube centered on the target point, and the SDF value at each feature point is calculated to form a geometric feature dataset where each feature point corresponds to a geometric feature latent vector.

[0013] Further, in step 1), when generating the ray dataset, a set of material parameters is randomly generated for each transformed three-dimensional grid in sequence to form multiple pairs of "three-dimensional grid, material parameter" data pairs. Each material parameter includes the number of layers of the material and the thickness of each layer. According to the thickness of each layer, the three-dimensional grid is shrunk inward using the vertex normal offset method to obtain the three-dimensional grid corresponding to each layer, and a scene that can be used for simulation is constructed; in each scene, multiple ray data are collected using the simulation method of volume path tracing, and these data are divided into a training set and a test set according to a preset ratio.

[0014] Further, the feature inputs of the BSSRDF model include: a latent vector representing local geometric features as geometric features, the outgoing and incoming positions of the ray, the outgoing and incoming directions, and the normal vectors at the outgoing and incoming positions as ray information, and the material parameters as material features; the network structure of the BSSRDF model adopts a CVAE structure, which consists of an encoder and a decoder, and both the encoder and the decoder adopt a multi-layer residual network structure; the input of the encoder is the geometric feature latent vector z g of the ray outgoing position, the material parameter z m , the outgoing position x o of the ray, the normal vector n o , the direction ω o , and the incoming position x i of the ray, the normal vector n i and the direction ω i ; the output of the encoder is the mean μ and variance σ of the multi-dimensional normal distribution, which are used for the encoding reparameterization of the CVAE; the decoder is divided into two parts. The first part consists of 2 multi-layer residual networks. The geometric feature latent vector and the outgoing information of the ray are converted into a multi-dimensional tensor by one of the multi-layer residual networks, which is called the geometric tensor h g , and the material information is converted into a multi-dimensional tensor by another multi-layer residual network, which is called the material tensor h m ; the second part also consists of 2 multi-layer residual networks. One of the multi-layer residual networks inputs 2 tensors and the CVAE encoding z p and outputs the predicted incoming position Another multi-layer residual network inputs 2 tensors and the predicted incoming position and outputs the predicted transmittance

[0015] Furthermore, the latent vector representing the local geometric features is obtained by optimizing on an implicit SDF model, where the input of the SDF model is the geometric feature latent vector z g and the position Δp relative to the target point corresponding to the latent vector, and the output is the distance to the nearest surface of the 3D mesh, i.e., the SDF value; the SDF model adopts a multi-layer residual network with the same structure as the BSSRDF model. Before using the SDF model to optimize the latent vector, it needs to be pre-trained, and the training data used for pre-training is the generated geometric feature dataset; during pre-training, the parameters of the SDF model and the geometric feature latent vectors in the dataset will be optimized simultaneously. The loss function adopted contains two parts: the L1 loss between the predicted SDF value and the true value and the L2 norm of the geometric feature latent vector where N is the total number of data in the dataset, is the i-th SDF value predicted by the model, d i is the true value of the SDF value of the i-th data in the dataset, is the geometric feature latent vector of the i-th data in the dataset; after pre-training, it is necessary to calculate the geometric feature latent vectors for each 3D mesh in the dataset. This calculation is performed at each vertex of the 3D mesh. Taking each vertex as the target point, its geometric feature latent vector is randomly initialized and feature points are collected. Then, the parameters of the implicit SDF model are fixed, and the optimization method of deep learning is used to optimize each geometric feature latent vector. The obtained latent vectors are stored as the data of each vertex on the transformed 3D mesh

[0016] Furthermore, in step 2), the loss function contains five parts: the L2 loss between the predicted value and the true value of the incident position and the CD loss the L1 loss between the predicted value and the true value of the transmittance the predicted SDF value of the incident position and the KL divergence of the CVAE structure where, X i respectively represent the set of predicted incident positions and the set of true incident positions in the i-th batch of data used during training, and a represents the true value of the transmittance in the training data

[0017] Furthermore, in step 3), the rendering steps adopted are as follows

[0018] 3.1) The emitted ray intersects with the surface of the 3D mesh to obtain the intersection point

[0019] 3.2) According to the barycentric coordinates of the intersection point, the geometric feature latent vector of the intersection point is linearly interpolated using the geometric feature latent vectors of 3 vertices

[0020] 3.3) Input the geometric feature latent vector, outgoing position, normal vector, and direction into the decoder to obtain a geometric tensor, a material tensor, and the predicted incoming position;

[0021] 3.4) Input the geometric feature latent vector and the predicted incoming position into the SDF model to obtain the SDF value at that position, take the absolute value and calculate the partial derivative of the incoming position with respect to the SDF value as the direction of reprojection;

[0022] 3.5) Perform an intersection detection once in each of the positive and negative directions of the reprojection, calculate the distances between the two intersection points and the current predicted position, and take the intersection point with the closest distance as the result of reprojection;

[0023] 3.6) According to the normal vector of the three-dimensional grid at the intersection point after reprojection, randomly sample an incoming direction within the hemisphere along the positive direction of the normal vector using the cosine weighting algorithm;

[0024] 3.7) Input the geometric tensor, material tensor, the position obtained after reprojection, and the sampled incoming direction into the multi-layer residual network for predicting the transmittance in the second part of the decoder to obtain the predicted transmittance;

[0025] 3.8) Update the incoming position and incoming direction of the ray, multiply the obtained transmittances cumulatively to get the cumulative transmittance, and continue ray tracing rendering until the ray intersects with the light source. Multiply the cumulative transmittance by the brightness of the light source to obtain the rendering result.

[0026] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0027] 1. The present invention adopts a local geometric feature representation method based on an implicit SDF model, encoding the local geometric surface shape and thickness into the latent space, enabling the neural network to accurately predict the characteristics of light passing through subsurface scattering to reach adjacent surface areas or penetrate the surface.

[0028] 2. The present invention constructs a scene dataset with multi-level subsurface scattering materials (i.e., geometric feature dataset and ray dataset). These materials are closer to the material characteristics of objects such as human skin in reality, and the neural network trained on this dataset can better render objects with such multi-layer materials.

[0029] 3. The present invention provides a method for using a neural network as a BSSRDF to render subsurface scattering, which can bring rendering results close to volume path tracing, and is superior to the volume path tracing method in terms of rendering efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic logic flow diagram of the method of the present invention.

[0031] Figure 2 Schematic diagram of the network for the implicit SDF model.

[0032] Figure 3 Schematic diagram of the network for the BSSRDF model based on CVAE.

[0033] Figure 4 、 Figure 5 Rendering result image obtained from the multi-layer material test scene generated by the method of the present invention. Detailed implementation manners

[0034] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0035] As Figures 1 to 5 shown, this embodiment discloses a multi-level surface scattering material rendering method based on a conditional variational autoencoder. A geometric feature latent vector capable of describing local geometry is generated based on an implicit SDF model, and a multi-layer material scene is constructed using an in-shrinking method. Ray data is collected using a volumetric path tracing method, and a neural network that can be used as a BSSRDF is trained to achieve efficient rendering of the multi-level surface scattering effect, including the following steps:

[0036] 1) Construct a geometric feature dataset. Obtain the CAD and 3D scanned 3D meshes open-sourced by the Interactive Geometry Lab as the initial 3D meshes. Remove the non-water-tight 3D meshes from these 3D meshes from the dataset, and then randomly generate 3D scaling transformations and 3D rotation transformations for the remaining 3D meshes, and transform the 3D meshes in the order of scaling and rotation; on each transformed 3D mesh, uniformly sample 10,000 target points, and generate a 64-dimensional geometric feature latent vector for each target point, where each value of the latent vector is randomly generated within the range [0, 0.01); generate a local sampling template for each target point, which contains the positions of 4096 relative template points and is randomly sampled in a cube with a length of 10. The target point plus the relative positions in the template is the feature point set; use Open3D to calculate the SDF value of each feature point.

[0037] Concatenate the obtained target points, feature points, geometric feature latent vectors, and SDF values to form a geometric feature dataset, which includes two parts: a feature tensor dataset and a feature point dataset. The feature tensor dataset is the geometric feature tensors of each target point stored in sequence, and each item in the feature point dataset is the relative position, SDF value of the feature point, and the index value of its corresponding geometric feature tensor.

[0038] 2) Construct an implicit SDF model and perform pre-training. Use the pre-trained implicit SDF model to calculate the geometric feature latent vector of each 3D mesh in the dataset and record it in the 3D mesh file. The implicit SDF model is based on a residual network structure and an implicit encoder framework, and its network structure is as shown in Figure 2 and is trained using a geometric feature dataset. The model adopts an MLP structure and adds residual connections. Its input is the geometric feature latent vector z g and the position Δp relative to the target point corresponding to this latent vector, and outputs the SDF value, where both the relative position and the SDF value are normalized to the range of [-1, 1].

[0039] During training, first, the positions and SDF values of the feature points in the dataset need to be normalized, that is, both the relative positions and SDF values of the feature points are first divided by half of the sampling cube length, which is 5, and then truncated to the range of [-1, 1]. The loss function used for training consists of two parts: the L1 loss between the predicted SDF value and the true value and the L2 norm of the geometric feature latent vector where N is the total number of data in the dataset, is the i-th SDF value predicted by the model, d i is the true value of the SDF value of the i-th data in the dataset, is the geometric feature latent vector of the i-th data in the dataset. Next, use the Adam optimizer for training, set the learning rate to 1e-4, and train for 100 epochs to obtain the implicit SDF model.

[0040] Next, calculate the geometric feature latent vector of the 3D mesh. For each vertex of each 3D mesh, first randomly initialize a 64-dimensional geometric feature latent vector for it, and then randomly sample 4096 feature points within a cube centered at this vertex with a length of 10. Use Open3D to calculate the SDF value of each feature point, thus forming the geometric feature dataset of this vertex. Use this dataset for training. During training, fix the parameters of the implicit SDF model and only optimize the geometric feature latent vector. The loss function used for training is the same as that for pre-training. Also use the Adam optimizer, set the learning rate to 1e-3, and only train for one epoch. Finally, store the trained latent vector as per-vertex data in the 3D mesh file.

[0041] 3) Construct a ray dataset. For each 3D mesh for which the geometric feature latent vector has been calculated, randomly generate 100 hierarchical parameters. Each hierarchical parameter contains the number of layers N l ∈ {1, 2, 3}, and the thickness l of each layer i ∈ [0.01, 0.1], i = 1, 2,..., N l , where the thickness l of the last layer NlIt is always 0, indicating that this layer fills the remaining space of the three-dimensional grid. Then, according to the thickness of each layer, the three-dimensional grid of each layer is generated: the grid of the first layer is the original three-dimensional grid; for the i-th layer of the three-dimensional grid, each vertex of the (i - 1)-th layer of the three-dimensional grid is moved along its negative normal vector direction by a distance l to form the three-dimensional grid of this layer. Each group of multi-layer three-dimensional grids is written into a scene file. i-1 Figure Figure 4 Figure 4 and Figure Figure 5 are the results after rendering some scenes generated by this method.

[0042] On each scene file, 10,000 rays are sampled. Each ray is composed of the outgoing position x o , the outgoing normal vector n o , the outgoing direction ω o , the geometric feature z g and the material parameter z m . The outgoing position is obtained by randomly sampling on the outermost three-dimensional grid. The outgoing normal vector is the normal vector at this sampling point. The outgoing direction is obtained by randomly sampling a direction on the hemisphere. The geometric feature is obtained by linearly interpolating the geometric feature latent vector of each vertex on the triangle where the outgoing point is located according to the barycentric coordinates of this outgoing point. The material parameter includes the scattering parameters of each layer, including: the absorption coefficient σ a ∈[1, 100], the scattering coefficient σ s ∈[1, 100], the anisotropy parameter g ∈ [0, 1), the refractive index η ∈ [1, 1.5] and the layer thickness l. Except for l which is a known parameter, other parameters are randomly sampled within their respective domains. Then, using the volume path tracing renderer implemented by Mitsuba 3, after each ray undergoes subsurface scattering, the incident position x i , the incident normal vector n i , the incident direction ω i and the transmittance a when leaving the outermost layer of the three-dimensional grid are traced and recorded in the ray dataset.

[0043] Figure 3 4) Construct a BSSRDF model based on the CVAE framework and the residual network structure. The network structure of the model is as shown in Figure Figure 3 Figure 3 and is trained using the ray dataset. According to the CVAE framework, the model is divided into two parts: the encoder and the decoder. The input of the encoder is the geometric feature latent vector z g of the ray outgoing position, the material parameter z m , the outgoing position x o , the normal vector n o , the direction ω o of the ray, as well as the incident position x i , the normal vector n i and the direction ω i; The output is the mean μ and variance σ of the multi-dimensional normal distribution. The encoder uses an 8-layer MLP with residual connections added. The decoder is divided into two parts: The first part consists of 2 networks. The geometric feature latent vector and the outgoing information of the light ray are converted into a 64-dimensional tensor by a 4-layer MLP residual network, which is called the geometric tensor h g ; The material information is converted into a 64-dimensional tensor by another 4-layer MLP residual network, which is called the material tensor h m ; The second part also consists of 2 networks. A 4-layer MLP residual network takes in 2 tensors and the CVAE encoding z p and outputs the predicted incident position Another 4-layer MLP residual network takes in 2 tensors and the predicted incident position and outputs the predicted transmittance

[0044] The loss function used to train the BSSRDF model is divided into 5 parts: the L2 loss between the predicted value and the true value of the incident position and the CD loss the L1 loss between the predicted value and the true value of the transmittance the SDF value of the predicted incident position and the KL divergence of the CVAE structure Among them, X i respectively represent the set of predicted incident positions and the set of true incident positions in the i-th batch of data used during training. a represents the true value of the transmittance in the training data. The Adam optimizer is used, and the learning rate is set to 1e-5. The BSSRDF model after training is obtained by training for 100 epochs using the light ray dataset.

[0045] 5) Use the trained implicit SDF model and BSSRDF model to perform rendering on the test scene. The specific rendering steps are as follows:

[0046] 5.1) Shoot a light ray to intersect with the surface of the 3D mesh to obtain the intersection point;

[0047] 5.2) According to the barycentric coordinates of the intersection point, linearly interpolate the geometric feature latent vectors of 3 vertices to obtain the geometric feature latent vector of this intersection point;

[0048] 5.3) Input the geometric feature latent vector, outgoing position, normal vector, and direction into the decoder to obtain the geometric tensor, material tensor, and predicted incident position;

[0049] 5.4) Input the geometric feature latent vector and the predicted incident position into the SDF model to obtain the SDF value at this position. Take the absolute value and calculate the partial derivative of the incident position with respect to this SDF value as the re-projection direction;

[0050] 5.5) Perform intersection detection once in each of the positive and negative directions of reprojection, calculate the distances between the two intersection points and the current predicted position, and take the intersection point with the shortest distance as the reprojection result;

[0051] 5.6) According to the normal vector of the three-dimensional grid at the intersection point after reprojection, randomly sample an incident direction within the hemispherical surface along the positive direction of the normal vector using the cosine weighting algorithm;

[0052] 5.7) Input the geometric tensor, material tensor, the position obtained after reprojection, and the sampled incident direction into the multi-layer residual network that predicts the transmittance in the second part of the decoder to obtain the predicted transmittance;

[0053] 5.8) Update the incident position and incident direction of the ray, multiply the obtained transmittances cumulatively to obtain the cumulative transmittance, continue ray tracing rendering until the ray intersects with the light source, and multiply the cumulative transmittance by the brightness of the light source to obtain the rendering result.

[0054] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A multi-level surface scattering material rendering method based on a conditional variational autoencoder, characterized in that Train a neural network, called the BSSRDF model, using simulation data. Use the trained neural network as the bidirectional subsurface reflectance distribution function (BSSRDF) for path tracing rendering. During path tracing rendering, when a ray intersects a 3D mesh with a multi-level surface scattering material, input the information of the intersection point into this distribution function to predict the next incident point and transmittance of the ray, and use the implicit SDF model to reprojection the predicted ray incident point to obtain an efficient and realistic rendering result. Among them, the feature input of the BSSRDF model is designed according to the characteristics of the simulation method, including ray information, geometric features, and material features. Its network structure is designed based on the conditional variational autoencoder (CVAE) framework and the residual network structure. The geometric features are calculated by designing an implicit SDF model, which is also based on the residual network structure. The specific implementation of this multi-level surface scattering material rendering method includes the following steps: 1) Use the collected 3D meshes to construct a dataset: Perform random rotation and scaling transformations on each 3D mesh to generate multiple transformed 3D meshes. On this basis, construct a geometric feature dataset and a ray dataset: The geometric feature dataset samples multiple groups of feature points on each 3D mesh and initializes a geometric feature latent vector for each group of feature points. The ray dataset is generated by randomly generating multiple pairs of "3D mesh, material parameter" data pairs. According to the thickness of each layer in the multi-layer material of each data pair, the corresponding 3D mesh is shrunk inward to construct a scene for collecting simulation data. In these scenes, use the simulation method to collect multiple ray data as the ray dataset, and divide it into a training set and a test set. 2) Batch input the training set constructed in step 1) into the designed BSSRDF model to obtain the predicted incident point and transmittance. Use the designed loss function to calculate the loss, and use the stochastic gradient descent method to optimize the network weights. Adjust the weights of each loss and the learning rate according to the loss change and network performance. Finally, obtain the optimal BSSRDF model. 3) On the scenes of the test set in step 1), use the BSSRDF model and the SDF model to combine with the volume path tracing method for rendering, and obtain a rendering result close to the volume path tracing method in the scene with a multi-level surface scattering material.

2. The multi-level surface scattering material rendering method based on conditional variational autoencoder according to claim 1, wherein, In step 1), when generating the geometric feature dataset, sample multiple target points on the surface of the transformed 3D mesh, randomly initialize a geometric feature latent vector for each target point, and resample multiple feature points in a cube centered on the target point, and calculate the SDF value at each feature point to form a geometric feature dataset where each feature point corresponds to a geometric feature latent vector.

3. The multi-level surface scattering material rendering method based on conditional variational autoencoder according to claim 2, wherein In step 1), when generating the light dataset, a set of material parameters is randomly generated for each transformed 3D mesh in sequence to form multiple pairs of "3D mesh, material parameter" data pairs. Each material parameter includes the number of layers of the material and the thickness of each layer. According to the thickness of each layer, the vertex normal offset method is used to shrink the 3D mesh inward to obtain the corresponding 3D mesh for each layer, and a scene that can be used for simulation is constructed. In each scene, multiple light data are collected using the volume path tracing simulation method, and these data are divided into a training set and a test set according to a preset ratio.

4. The multi-level surface scattering material rendering method based on conditional variational autoencoder according to claim 3, characterized in that, The characteristic inputs of the BSSRDF model include: a latent vector representing local geometric features as geometric features, the outgoing and incoming positions of the light, the outgoing and incoming directions, and the normal vectors at the outgoing and incoming positions as light information, as well as material parameters as material features; the network structure of the BSSRDF model adopts a CVAE structure, consisting of an encoder and a decoder, and both the encoder and the decoder adopt a multi-layer residual network structure; the input of the encoder is the geometric feature latent vector z of the light outgoing position g , the material parameter z m , the outgoing position x of the light o , the normal vector n o , the direction ω o , and the incoming position x of the light i , the normal vector n i , and the direction ω i ; the output of the encoder is the mean μ and variance σ of a multi-dimensional normal distribution, used for the encoding reparameterization of the CVAE; the decoder is divided into two parts. The first part consists of 2 multi-layer residual networks. The geometric feature latent vector and the outgoing information of the light are converted by one of the multi-layer residual networks into a multi-dimensional tensor, called the geometric tensor h g , and the material information is converted by another multi-layer residual network into a multi-dimensional tensor, called the material tensor h m ; the second part also consists of 2 multi-layer residual networks. One of the multi-layer residual networks inputs 2 tensors and the CVAE encoding z p , and outputs the predicted incoming position Another multi-layer residual network inputs 2 tensors and the predicted incoming position and outputs the predicted transmittance 5. The multi-level surface scattering material rendering method based on conditional variational autoencoder according to claim 4, characterized in that, The latent vector representing the local geometric feature is obtained by optimizing on an implicit SDF model, where the input of the SDF model is the geometric feature latent vector z g and the position Δp relative to the corresponding target point of the latent vector, and the output is the distance to the nearest surface of the 3D mesh, i.e., the SDF value; the SDF model adopts a multi-layer residual network with the same structure as the BSSRDF model. Before using the SDF model to optimize the latent vector, pre-training is required. The training data used for pre-training is the generated geometric feature dataset; during pre-training, the parameters of the SDF model and the geometric feature latent vectors in the dataset will be optimized simultaneously. The loss function adopted contains two parts: the L1 loss between the predicted SDF value and the true value and the L2 norm of the geometric feature latent vector where N is the total number of data in the dataset, is the i-th SDF value predicted by the model, d i is the true value of the SDF value of the i-th data in the dataset, is the geometric feature latent vector of the i-th data in the dataset; After the pre-training is completed, it is necessary to calculate the geometric feature latent vector for each 3D mesh in the dataset. This calculation is performed at each vertex of the 3D mesh. Taking each vertex as the target point, its geometric feature latent vector is randomly initialized and feature points are collected. Then, the parameters of the implicit SDF model are fixed, and the deep learning optimization method is used to optimize each geometric feature latent vector. The obtained latent vector is stored as the data of each vertex on the transformed 3D mesh.

6. The multi-level surface scattering material rendering method based on conditional variational autoencoder according to claim 5, wherein In step 2), the loss function consists of five parts: the L2 loss between the predicted value and the true value of the incident position and the CD loss the L1 loss between the predicted value and the true value of the transmittance the SDF value of the predicted incident position and the KL divergence of the CVAE structure where X i respectively represent the set of predicted incident positions and the set of true incident positions in the i-th batch of data used during training, and a represents the true value of the transmittance in the training data.

7. The multi-level surface scattering material rendering method based on conditional variational autoencoder according to claim 6, wherein In step 3), the rendering steps adopted are as follows: 3.1) The emitted light intersects with the surface of the 3D mesh to obtain the intersection point; 3.2) According to the barycentric coordinates of the intersection point, the geometric feature latent vector of this intersection point is linearly interpolated using the geometric feature latent vectors of 3 vertices; 3.3) The geometric feature latent vector, the outgoing position, the normal vector, and the direction are input into the decoder to obtain the geometric tensor, the material tensor, and the predicted incoming position; 3.4) The geometric feature latent vector and the predicted incoming position are input into the SDF model to obtain the SDF value at this position, and the absolute value is taken to calculate the partial derivative of the incoming position with respect to this SDF value as the reprojection direction; 3.5) Intersection detection is performed once in each of the positive and negative directions of the reprojection, the distances between the two intersection points and the current predicted position are calculated, and the intersection point with the closest distance is taken as the reprojection result; 3.6) According to the normal vector of the 3D mesh at the intersection point after reprojection, a random incident direction is sampled within the hemisphere along the positive direction of this normal vector using the cosine weighted algorithm; 3.7) The geometric tensor, the material tensor, the position obtained after reprojection, and the sampled incident direction are input into the multi-layer residual network that predicts the transmittance in the second part of the decoder to obtain the predicted transmittance; 3.8) Update the incoming position and the incoming direction of the light, and multiply the obtained transmittances to obtain the cumulative transmittance. Continue the ray tracing rendering until the light intersects with the light source, and multiply the cumulative transmittance by the brightness of this light source to obtain the rendering result.

Citation Information

Cited By

  • Enhanced multi-linear mixed hyperspectral unmixing method, system and device

    CN122223458A