Simulation method and system based on three-dimensional Gaussian splashing

Through the simulation method based on three-dimensional Gaussian splashing, the problem that the existing technology cannot effectively support the re-renderable and re-illumination requirements is solved, and simulated image generation with high authenticity and real-time performance is achieved, providing strong support for the research and development and testing of autonomous driving systems.

CN119941947APending Publication Date: 2025-05-06UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202411868093.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When handling dynamic scenes, the prior art cannot effectively support the re-renderable and re-illumination requirements, resulting in difficulty in simulation testing of autonomous driving systems and insufficient image authenticity, real-timeness and dynamic scene generation.

Method used

Using a simulation method based on three-dimensional Gaussian splashing, we construct a three-dimensional scene, determine 3D Gaussian points, and project them to the two-dimensional image plane, calculate the depth in combination with the affine transformation matrix, and synthesize the final color using the Alpha Blending hybrid transparency algorithm, and optimize the rendered image through physical rendering or fractional distillation sampling.

Benefits of technology

It significantly improves the authenticity and real-time performance of simulated images, narrows the gap between simulation and reality, and improves the research and development and testing efficiency of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941947A_ABST
    Figure CN119941947A_ABST
Patent Text Reader

Abstract

The invention provides a simulation method and system based on three-dimensional Gaussian splashing, and relates to the technical field of computer vision, and the method comprises the steps: constructing a three-dimensional scene; determining a 3D Gaussian point in the three-dimensional scene; projecting the 3D Gaussian points to a two-dimensional image plane, and determining two-dimensional Gaussian points; according to the affine transformation matrix, the depth of each Gaussian body is calculated, and a Gaussian body sorting list is formed; a final color is synthesized and calculated through an Alpha Blending mixed transparency algorithm; rendering the three-dimensional scene according to the final color to obtain a first rendered image; and optimizing the first rendered image based on a physical rendering mode or a fractional distillation sampling reconstruction mode to obtain a second rendered image. According to the method, the authenticity of a simulation image can be improved, a data processing and rendering algorithm is optimized, the real-time performance is improved, and a simulation platform can perform more efficient real-time simulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a simulation method and system based on three-dimensional Gaussian splashing. Background Art

[0002] 3D Gaussian splashing is a new 3D model representation method that is closely integrated with deep learning and is often used in 3D reconstruction and model generation. 3D Gaussian splashing provides a new perspective synthesis and scene reconstruction method by using a set of 3D Gaussian functions to represent 3D scenes. The rendering method is a new computer graphics technology that aims to efficiently render 3D scenes and optimize rendering effects.

[0003] With the rapid development of autonomous driving technology, its potential in improving road safety, optimizing traffic flow, and reducing energy consumption has gradually emerged. However, the research and development and testing of autonomous driving systems face major challenges, especially in ensuring system safety and reliability. Although simulation technology plays an important role in the field of autonomous driving, existing simulation platforms such as CARLA and APOLLO have significant limitations. Therefore, optimizing data processing and rendering algorithms is of great significance in improving rendering efficiency, improving image quality, enhancing user experience, and reducing resource consumption.

[0004] The rendering engines currently in widespread use often fail to effectively reproduce the visual features of the real world when generating simulated images, resulting in significant differences between simulated images and actual images. This inter-domain difference affects the performance of autonomous driving algorithms in real environments and cannot meet the safety and reliability requirements of practical applications. Current realistic image synthesis methods, such as the NeRF method, often suffer from insufficient real-time performance due to complex calculation processes and cannot support real-time interaction and feedback in dynamic scenes. Existing reconstruction methods, including 3DGS technology, often face challenges when processing dynamic scenes and cannot effectively support re-rendering and re-lighting requirements, which makes simulation testing in complex traffic environments difficult, resulting in various deficiencies in image authenticity, real-time, and dynamic scene generation, limiting the effective verification and testing of autonomous driving systems. Summary of the invention

[0005] In order to solve the technical problems that the 3DGS technology in the prior art often faces challenges when processing dynamic scenes and cannot effectively support the requirements of re-rendering and re-lighting, which makes simulation testing in complex traffic environments difficult, resulting in various deficiencies in image authenticity, real-time and dynamic scene generation, and limits the effective verification and testing of autonomous driving systems, the present invention provides a simulation method and system based on three-dimensional Gaussian splashing.

[0006] The technical solution provided by the embodiment of the present invention is as follows:

[0007] First aspect

[0008] An embodiment of the present invention provides a simulation method based on three-dimensional Gaussian splashing, comprising:

[0009] S1: construct a three-dimensional scene;

[0010] S2: Determine 3D Gaussian points in the 3D scene;

[0011] S3: Project the 3D Gaussian points onto the 2D image plane to determine the 2D Gaussian points;

[0012] S4: Calculate the depth of each Gaussian body according to the affine transformation matrix to form a sorted list of Gaussian bodies;

[0013] S5: The final color is calculated by alpha blending algorithm.

[0014] S6: Rendering the three-dimensional scene according to the final color to obtain a first rendered image;

[0015] S7: Optimize the first rendered image based on a physical rendering method or a fractional distillation sampling reconstruction method to obtain a second rendered image.

[0016] Second aspect

[0017] An embodiment of the present invention provides a simulation system based on three-dimensional Gaussian splashing, comprising:

[0018] processor;

[0019] A memory having computer-readable instructions stored therein, wherein when the computer-readable instructions are executed by the processor, the simulation method based on three-dimensional Gaussian splashing as in the first aspect is implemented.

[0020] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0021] In the present invention, a three-dimensional scene is constructed based on a three-dimensional Gaussian splash algorithm, 3D Gaussian points in the three-dimensional scene are determined, and the 3D Gaussian points are projected onto a two-dimensional image plane to obtain two-dimensional Gaussian points. According to the two-dimensional Gaussian points, the depth of each Gaussian body is calculated through an affine transformation matrix. Through depth calculation and sorting, the final rendering result is real and consistent with physical logic. The two-dimensional Gaussian points are mixed according to the depth sorting, and the Alpha A blending transparency algorithm is used to generate the final image color. According to the final color, the three-dimensional scene is rendered to obtain a first rendered image. The first rendered image is optimized through a physically based rendering method (PBR) or a fractional distillation sampling method to eliminate defects and enhance the realism or artistic expression of the image. The present invention introduces PBR (physically based rendering) attributes into Gaussian points, so that the system can process rendering requirements under different lighting conditions in real time, ensure the reliability and accuracy of simulation results, optimize data processing and rendering algorithms, and enhance real-time performance, so that the simulation platform can perform more efficient real-time simulation, significantly improve the authenticity of simulated images, narrow the gap between simulation and reality, and provide strong support for the research and development and testing of autonomous driving systems, thereby promoting the further development of autonomous driving technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 A schematic flow chart of a simulation method based on three-dimensional Gaussian splashing provided by an embodiment of the present invention;

[0024] Figure 2 A schematic diagram of a 3DGS rendering representation provided by an embodiment of the present invention;

[0025] Figure 3 A schematic diagram of a 3DGS training process provided by an embodiment of the present invention;

[0026] Figure 4 A schematic diagram of a re-renderable 3DGS model training provided by an embodiment of the present invention;

[0027] Figure 5 A schematic diagram of a flow chart of a fractional distillation sampling method provided by an embodiment of the present invention;

[0028] Figure 6 A schematic diagram of a process flow of a semi-generative closed-loop simulation platform based on three-dimensional Gaussian splashing provided by an embodiment of the present invention;

[0029] Figure 7 A schematic structural diagram of a simulation system based on three-dimensional Gaussian splashing provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0031] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0032] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0033] Reference Manual Attached Figure 1 , shows a schematic flow chart of a simulation method based on three-dimensional Gaussian splashing provided in an embodiment of the present invention.

[0034] The embodiment of the present invention provides a simulation method based on three-dimensional Gaussian splashing, which can be implemented by a simulation device based on three-dimensional Gaussian splashing, and the simulation device based on three-dimensional Gaussian splashing can be a terminal or a server. The processing flow of the simulation method based on three-dimensional Gaussian splashing can include the following steps:

[0035] S1: Construct a 3D scene.

[0036] Among them, a three-dimensional scene refers to an environment with three-dimensional coordinates constructed in a virtual space, including information such as geometric models, light sources, camera perspectives, materials and textures. By accurately defining elements such as objects, lighting and camera perspectives in the three-dimensional scene, it ensures that there is a clear source of data input for subsequent steps. By integrating complex geometric structures and material information during the construction phase, a complete description of the scene is achieved, which not only supports flexible scene adjustment, but also provides high-quality input data for subsequent Gaussian point generation and projection.

[0037] Reference Manual Attached Figure 2 , shows a schematic diagram of a 3DGS rendering representation provided by an embodiment of the present invention.

[0038] Figure 2In the example, the coordinate Position represents the position of a point in three-dimensional space, [x, y, z] represents the position coordinates of the point on the X, Y, and Z axes, and the scaling Scale represents the scaling ratio of the object in each dimension, corresponding to the scaling factors in the X, Y, and Z axes, respectively. x ,S y ,S z ] indicates the degree to which a three-dimensional model is stretched or compressed in space. Rotation is used to describe the rotation state of an object in three-dimensional space in the form of a quaternion. w ,q x ,q y ,q z ] represents the components of the quaternion. Color spherical harmonics use spherical harmonics (SHs) to represent the distribution of ambient light and are used to process the color effects of global illumination in rendering. The 16 RGB values ​​represent the coefficients of the three colors red, green, and blue separated on the basis of spherical harmonics. Opacity represents the degree of transparency of an object, and usually ranges from 0 to 1, with 0 representing complete transparency and 1 representing complete opacity. [O] is a scalar value used to control the visual transparency of an object.

[0039] S2: Determine 3D Gaussian points in the 3D scene.

[0040] Among them, a 3D Gaussian point is a rendering unit represented by a three-dimensional Gaussian distribution function, which describes the spatial position, size, direction, shape and other properties of each point in the scene, and is mathematically defined by the mean, covariance matrix, scale matrix and rotation matrix.

[0041] It should be noted that converting the complex geometric information in the three-dimensional scene into 3D Gaussian points with clear mathematical expressions simplifies the rendering calculations and significantly improves the data processing efficiency. Gaussian points can not only flexibly adjust the scale and direction, but also conveniently describe the blur or gradient effects through mathematical properties, thereby more naturally simulating light interaction. At the same time, this abstract form reduces the complexity of traditional geometric processing, making subsequent projection and rendering algorithms more efficient and accurate.

[0042] In a possible implementation, the mathematical definition of a 3D Gaussian point is specifically:

[0043]

[0044] ∑=RSS T R T

[0045] Among them, G(x) represents the value of the Gaussian function, exp represents the natural exponential function, x represents the position in the multidimensional space, μ represents the mean of the center position (x, y, z), ∑ represents the 3D covariance matrix of the scaling degree, S represents the scale matrix, and R represents the rotation matrix that describes and transforms the direction of the Gaussian point. -1 represents the inverse matrix operation, T Represents a transpose operation.

[0046] S3: Project the 3D Gaussian points onto the 2D image plane to determine the 2D Gaussian points.

[0047] Among them, the two-dimensional image plane refers to a two-dimensional coordinate space used to represent the pixel distribution and color information in the image. It is the core intermediary for converting from a three-dimensional scene to an image during the rendering process and is the mathematical representation of the final image output. The two-dimensional Gaussian point is the projection result of the three-dimensional Gaussian point on the two-dimensional plane, which contains information such as position, size, and shape.

[0048] Specifically, in the rendering process after confirming the camera pose, the 3D Gaussian is first projected onto the image plane to obtain a 2D Gaussian. The mean of these two-dimensional Gaussian functions is determined by calculating the projection matrix. The three-dimensional Gaussian points are mapped to the two-dimensional image plane through an accurate projection algorithm, retaining depth information and geometric features. This process converts high-dimensional data into two-dimensional data that is easy to render, significantly reducing the computational complexity while maintaining the spatial characteristics and perspective effects of the object.

[0049] In a possible implementation, a specific projection method of projecting a 3D Gaussian point into a 2D Gaussian point is:

[0050] ∑′=JW∑W T J T

[0051] Wherein, ∑′ represents the covariance matrix of the two-dimensional Gaussian distribution, J represents the affine approximate Jacobian matrix, and W represents the affine transformation matrix.

[0052] Among them, the covariance matrix is ​​a mathematical tool used to describe the relationship between random variables. In three-dimensional Gaussian splash rendering, it represents the expansion and distribution of Gaussian distribution in different directions. The affine transformation matrix is ​​a geometric transformation matrix used to describe the linear mapping relationship between coordinate systems, maintaining the characteristics of points, lines and their parallelism. The affine approximation Jacobian matrix is ​​a linear approximation tool in affine transformation, which is used to describe the local change characteristics of points during coordinate changes.

[0053] It should be noted that the projection method using the affine transformation matrix makes the coordinate mapping process more flexible and mathematically precise, thereby ensuring the consistency of data transmission during the rendering process. This projection strategy lays an efficient and stable foundation for subsequent depth sorting and color synthesis, and is a key geometry processing step in the rendering process.

[0054] S4: According to the affine transformation matrix, the depth of each Gaussian body is calculated to form a sorted list of Gaussian bodies.

[0055] Among them, Gaussian body refers to the 3D Gaussian point in rendering, which is called Gaussian body, representing a three-dimensional Gaussian distribution with a specific position, size and direction. The Gaussian body sorting list is based on the depth value of the object or point, and is arranged from far to near or from near to far from the camera perspective, which is used to determine the order of rendering.

[0056] Specifically, through depth calculation and sorting, the front and back occlusion relationship of objects in the scene is resolved, ensuring the correctness of transparency blending and color synthesis. This process avoids visually unreasonable overlaps and perspective errors in the rendering results, thereby improving the realism and logical consistency of the image. At the same time, depth sorting can optimize the processing order of Gaussian bodies, so that computing resources are concentrated on rendering the visible parts, improving rendering efficiency.

[0057] It should be noted that, given a pixel position, the distance to all overlapping Gaussian bodies, i.e., the depth of these Gaussian bodies, can be calculated through view transformation to form a sorted list of Gaussian bodies.

[0058] S5: The final color is calculated by alpha blending algorithm.

[0059] Among them, the Alpha Blending algorithm is a color synthesis method based on pixel transparency value (Alpha), which is used to simulate the superposition effect of translucent objects. It generates the final mixed color by calculating the weighted average of the colors of each layer. The final color refers to the color value of each pixel in the synthesized image, taking into account the color and transparency of all layers.

[0060] It should be noted that the alpha blending algorithm can efficiently and realistically simulate the color overlay effect of translucent objects, so that the final image presents a more natural lighting and layering. This method combines the depth sorting results to ensure that each layer participates in color synthesis in the correct order, avoiding visual irrationality. At the same time, alpha blending has high computational efficiency and can quickly process multiple transparent effects in complex scenes.

[0061] In a possible implementation, S5 is specifically:

[0062] The final color is calculated by the Alpha Blending algorithm according to the following formula:

[0063]

[0064] Among them, C represents the final color, c i Indicates the color of the i-th layer, α i ′ represents the transparency value of the i-th layer, α j ′ represents the transparency value of the jth layer, and N represents the total number of layers.

[0065] Reference Manual Attached Figure 3 , showing a schematic diagram of a 3DGS training process provided by an embodiment of the present invention.

[0066] Figure 3 In the 3D space, the initialization point is a set of initial points initially defined in the 3D space, which serves as the basis for the subsequent generation of 3D Gaussian points (3DGS points). Initialization means converting the initial points into 3D Gaussian points (3DGS points), which have properties such as position, rotation, scale, and color. The camera pose describes the position and direction of the camera in the 3D space, which is used to determine the viewing angle. The projection calculation projects the 3D Gaussian points onto the 2D image plane according to the camera pose, and calculates their position and shape in the image. The adaptive density control optimizes the computational efficiency while ensuring the image quality, and reduces unnecessary point calculations. Differentiable rendering converts the projected 2D information into the final image, and allows the calculation of gradients for optimization. The rendered image contains the lighting, color, and geometric details of the scene, presenting the final rendering result. The forward process describes the generation process from 3D Gaussian points to the rendered image. The gradient backpropagation refers to the use of the error of the rendered image to calculate the gradient of the parameters and backpropagate, which is used to optimize the properties of the 3D Gaussian points (such as position, shape, color, etc.).

[0067] S6: Rendering the three-dimensional scene according to the final color to obtain a first rendered image.

[0068] Among them, the first rendered image is an image generated by the preliminary rendering, which includes transparency blending, color and lighting effects after Gaussian point projection, but has not been further optimized and enhanced. By integrating the results of previous steps such as transparency blending and Gaussian point projection, the first rendered image containing the main visual features of the scene is quickly generated, providing a solid foundation for subsequent optimization. This step can present a preliminary rendering effect that is close to reality while ensuring high efficiency, allowing users to quickly preview and adjust the scene. At the same time, the generated image already contains key information such as depth, transparency and lighting, ensuring the integrity of the rendering results.

[0069] In a possible implementation, the training method of the three-dimensional Gaussian splash algorithm specifically includes:

[0070] The optimization objective function of the three-dimensional Gaussian splash algorithm is constructed based on the one-norm loss function and the SSIM loss function:

[0071] L=(1-λ)L1+λL D-SSIM

[0072] Among them, L represents the optimization objective function, λ represents the weight parameter, L1 represents the one-norm loss function, and L D-SSIM Represents the SSIM loss function.

[0073] The three-dimensional Gaussian splashing algorithm is trained and optimized with the goal of minimizing the optimization objective function.

[0074] It should be noted that 3DGS has the advantages of continuous representation, efficient rendering, fast training and high-quality reconstruction. These characteristics make 3DGS have great potential in the field of autonomous driving simulation, and can significantly improve the realism of simulated images and narrow the gap between simulation and reality.

[0075] In the present invention, an adaptive Gaussian densification scheme is introduced in the Gaussian training process. When a small-scale geometric object is not fully covered, the algorithm copies two identical Gaussians; if the small-scale geometric object is covered and represented by a large Gaussian point, the algorithm splits the large Gaussian point into two small sub-Gaussian points. This copying and splitting process is performed after the loss calculation is optimized in each training round.

[0076] Reference Manual Attached Figure 4 , showing a schematic diagram of a re-renderable 3DGS model training provided by an embodiment of the present invention.

[0077] Figure 4In the figure, the first stage represents the preliminary rendering based on 3DGS. First, 3DGS points are input. Based on the representation of three-dimensional Gaussian points (3DGS points), different types of basic information are generated through differentiable rendering. RGB is the preliminary generated color image. Depth is the distance between each pixel in the scene and the camera, which is used to determine the front and back occlusion relationship. Normal represents the direction information of each surface, which is used for lighting calculation. The second stage represents rendering optimization based on PBR. Based on the results of the preliminary rendering, the material properties required for PBR rendering are extracted, including base color, roughness, metalness and visibility. The global light field simulates the global illumination effects in the scene, including the interaction of direct light, indirect light and ambient light, to ensure the physical consistency of the rendering results. GT stands for Ground Truth, that is, a real image or reference image. It is usually a high-quality image used as a target in the rendering optimization or learning process. It is used to compare with the generated rendering results to evaluate the performance of the model or rendering algorithm.

[0078] S7: Optimize the first rendered image based on a physical rendering method or a fractional distillation sampling reconstruction method to obtain a second rendered image.

[0079] Among them, Physically-Based Rendering is a rendering method based on real physical laws, which generates realistic images by simulating the interaction between light, material and environment. Fractional distillation sampling is an image generation technology based on probability distribution and optimized sampling. It reconstructs high-quality images from low-resolution or incomplete data. The second rendered image refers to the optimized final image with higher clarity, details and realism.

[0080] It should be noted that the first rendered image is deeply optimized through physical rendering or fractional distillation sampling reconstruction method, which significantly improves the details and realism of the image. The physical rendering method can accurately simulate the reflection, refraction and scattering of light, so that the material performance is more in line with the actual physical properties, while the fractional distillation sampling technology effectively supplements the details and accuracy of the image by refining the reconstruction of incomplete data. The two optimization methods provide flexibility and adaptability, and can generate high-quality images according to different scene requirements.

[0081] In a possible implementation, the attributes related to the physical rendering method include: base color, roughness, metalness, and normal.

[0082] Among them, the base color is the inherent color of the object's surface without the influence of light, reflecting the main color characteristics of the object's material. The roughness describes the microscopic unevenness of the object's surface and has a direct impact on the reflection characteristics of light. The metalness describes whether the material has metallic properties and affects the object's reflection characteristics to light. The normal is a vector perpendicular to the object's surface, which is used to describe the direction of the surface at a specific point.

[0083] In the present invention, it is considered to add base color, roughness, metalness, and normal as additional 3DGS attributes to 3DGS, and the error calculation between the rendering image and the real image is obtained according to the PBR rendering calculation method and the corresponding attributes are optimized. By adding the object PBR attributes to 3DGS and performing learning optimization, the object can better adapt to different ambient light fields to complete the scene editing work.

[0084] In a possible implementation manner, optimizing the first rendered image based on the physical rendering method in S7 specifically includes:

[0085] Based on the physical rendering method, the first rendering image is rendered according to the following formula:

[0086]

[0087] Among them, L o (p,ω o ) represents the direction ω from the surface point p o Outgoing light brightness, L e (p,ω o ) indicates that point p is in direction ω o The self-luminous brightness, L i (p,ω i ) represents the direction ω from which the point p comes i The incident light brightness, f r represents the bidirectional reflectance distribution function, n represents the normal, n·ω i represents the cosine of the angle between the incident light and the surface normal, and d represents the differential sign.

[0088] The bidirectional reflectance distribution function specifically adopts the Disney BRDF principle:

[0089]

[0090] F Schlick =F0+(1-F0)(1-cosθ d ) 5

[0091] F0=0.04(1-m)+bm

[0092] G(l,v,h)=G GGX (l)GGGX (v)

[0093]

[0094] α=(0.5+r / 2) 2

[0095] Among them, f r (l,v) represents the ratio of reflected light from the incident direction l to the observation direction v, f d Represents the diffuse reflection part, D(θ h ) represents the microfacet distribution function, θ h represents the angle between the half vector and the surface normal, F(θ d ) represents the Fresnel effect, θ d Represents the angle between the incident light direction and the surface normal, G(θ l ,θ v ) represents the geometric attenuation function, θ l and θ v They respectively represent the angles between the incident light direction and the observation direction and the surface normal, cos represents the cosine function, b represents the base color, m represents the metalness, n represents the normal, r represents the roughness parameter, F0 represents the specular reflectivity at vertical incidence, h represents the half vector, and α represents the transparency value.

[0096] It should be noted that the Disney BRDF principle (Bidirectional Reflectance Distribution Function) is a physically based reflection model developed by Disney Animation Studios. It is widely used in movies, games and other fields that require high-quality rendering. It aims to represent the lighting behavior of the surface of an object through a unified mathematical model, combined with easily adjustable parameters, to generate realistic material effects. It has become an important part of the modern rendering workflow and an important embodiment of the PBR (physically based rendering) concept.

[0097] In a possible implementation, the training method of the optimization process based on the physical rendering method specifically includes:

[0098] Construct the loss function of the optimization process based on physical rendering:

[0099] L′=(1-sum(λ))L PBR +λ n L n +λ smooth L smooth +λ light L light

[0100]

[0101] L smooth =||g material (p)-g material (p+ε)||2

[0102]

[0103] Among them, L′ represents the loss function of the optimization process, sum represents the summation symbol, and L PBR represents the loss function of physical rendering, λ n represents the weight of the normal loss function, λ smooth represents the weight of the smooth loss function, λ light Represents the weight of the illumination loss function, L n represents the normal loss function, N represents the normal of the Gaussian point, N represents the existing normal, L smooth represents the smooth loss function, g material represents the function used to calculate the material properties at a given position P, where P represents a point on the surface of the object, ε represents the noise, and L light represents the illumination loss function, L c Represents the light intensity in color channel c, where c represents the color channel, R represents the red channel, G represents the green channel, and B represents the blue channel.

[0104] The optimization process of physically based rendering is trained with the goal of minimizing the loss function.

[0105] It should be noted that the latest diffusion model has the ability to better generate consistent multi-view images, and multi-view images of the same object are inferred and generated, and then the 3DGS algorithm is used to reconstruct and distill the PBR attributes based on the many generated images. This method uses the existing multi-view diffusion model to infer multi-angle images of the same object, and then uses the existing data for reconstruction to generate a variety of 3D model resources, separating the dependencies of different models. The two parts of the algorithm can be replaced without affecting the other part.

[0106] Reference Manual Attached Figure 5 , showing a schematic flow chart of a fractional distillation sampling method provided in an embodiment of the present invention.

[0107] Figure 5 In, x t represents the noisy data diffused to time step t, x0 represents the original noise-free data, α t and σ t They represent the time t function of the control data and the time t function of the control noise ratio, ε represents noise, ∈ φrepresents the output of the diffusion model, and the input is the scene geometry information represented by three-dimensional Gaussian points (3DGS points). These Gaussian points are used as the initial conditions for model optimization. Then the diffusion model receives the noisy input and uses its learned features to gradually restore the high-quality target image through the reverse diffusion (ie, denoising) process. At the same time, the diffusion model is a generative model that generates high-quality images through a step-by-step denoising method.

[0108] In a possible implementation manner, optimizing the first rendered image based on the fractional distillation sampling reconstruction method in S7 specifically includes:

[0109] Based on fractional distillation sampling, the parameters of the three-dimensional Gaussian splash model are initialized, and the corresponding image is generated by rendering.

[0110] Select images rendered at different angles and calculate the similarity between the rendered images and the sample images of the diffusion model:

[0111]

[0112] Among them, L Diff (φ,x) represents the similarity loss function, t represents the time step, u(0,1) represents the uniform distribution, N(0,I) represents the multivariate standard normal distribution, E represents the expectation, w(t) represents the time-dependent weight function, α t and σ t They represent the time t function of the control data and the time t function of the control noise ratio, respectively. φ Represents the output of the diffusion model.

[0113] The three-dimensional Gaussian splash model parameters are optimized with the goal of minimizing the similarity between the rendered image and the sample image of the diffusion model:

[0114]

[0115] in, Denotes the loss function L SDS The gradient of the parameter θ, g(θ) represents the generating function, E t,∈ represents the expectation of all possible time steps and noise, It means that the diffusion model is under given conditions Z t , target y and output at time t, represents the partial derivative of the second rendering with respect to the parameter θ.

[0116] It should be noted that the advantage of optimizing the first rendered image based on fractional distillation sampling reconstruction is that it combines efficient computing and enhanced realism. It provides excellent quality assurance for the final rendered image by refining details, enhancing realism and optimizing similarity. This method is particularly suitable for application scenarios that require fine restoration and complex light and shadow simulation. It is resource-efficient and adaptable to multiple scenarios, and is an important tool for modern rendering optimization.

[0117] In actual application, by combining the high-quality scene reconstruction and real-time rendering capabilities of 3DGS and the flexibility of semi-generative methods to provide rich object materials in the scene, it is possible to simulate various complex traffic scenarios and environmental conditions, providing strong support for the research and development and testing of autonomous driving systems. Through this innovative simulation platform, the behavior and performance of autonomous driving systems in actual applications can be better predicted and evaluated, thereby promoting the further development and commercialization of autonomous driving technology.

[0118] Reference Manual Attached Figure 6 , showing a schematic flow chart of a semi-generative closed-loop simulation platform based on three-dimensional Gaussian splashing provided in an embodiment of the present invention.

[0119] Figure 6 In the process, environmental data such as RGB images, depth images and 2D radar images are collected from various sensors (such as cameras and LiDAR) through the acquisition module, and a virtual environment is used for simulation. The generated data includes RGB images, segmented images (Seg) and depth images (Depth). Environmental perception processing includes processing the simulated data and actual sensor data through specific algorithms (such as image segmentation and depth perception) to extract the key features of vehicles and roads. 3DGS (three-dimensional Gaussian point) technology is used to model the scene. Decouple Network refers to a neural network model for data decoupling, which is used to independently process various data streams. 3DGS rendering will render the 3D scene established by the model to obtain high-precision visual output. Combined with LiDAR point cloud data and images rendered by the 3DGS model, in-depth data fusion and analysis are carried out to improve the accuracy of environmental perception. Based on the above data and analysis results, path planning, obstacle recognition and driving decisions are carried out.

[0120] Specifically, the closed-loop simulation platform based on reusable plug-in functions is designed as follows:

[0121] First, we consider the design of the platform's visual scene composition: pure driving scene generation is too difficult, so a semi-generative scene design is adopted. The background is reconstructed from the dataset, and then the foreground is reconstructed from the dataset model and a wide range of foreground object materials are obtained through the generative model. Considering that the sky background often introduces interference and artifacts during reconstruction, and the sky is a different layer from the street scene and should be modeled separately, the sky is designed as a learnable environment map for learning through differentiable rendering. At the same time, such a design also supports replacing the environment map with other texture image materials for real-time rendering and imaging.

[0122] From the perspective of algorithm and display logic, the closed-loop simulation platform can be naturally divided into two parts: the offline training end and the real-time deduction end. The offline training end includes the process of training and reasoning the generated model materials (generating high-quality and suitable materials in advance and entering the material library, and using the material library content for subsequent real-time deduction), and the process of training the data set to reconstruct background and foreground materials (also added to the material library). The real-time deduction end is responsible for achieving the purpose of real-time deduction by using the trained materials, and can support the completion of various downstream application tasks, such as data generation, algorithm testing, and closed-loop simulation.

[0123] At the underlying data level, existing simulation platforms CARLA and BeamNG are used to collect data and process it into a specific format, or multiple types of data set formats such as waymo and kitti are supported as data layers to support the platform. On the one hand, it is used for offline training to obtain three-dimensional material resources, and on the other hand, it is also used as source data to support real-time deduction of scene object trajectory reproduction.

[0124] The simulation calculation end can be divided into an offline part and a real-time part according to whether it is running.

[0125] The offline part is divided into two parts. In the first part, the data is reconstructed through the 3DGS method based on the physics to obtain the 3D model material with foreground and background separation, and the algorithm is trained by dividing the foreground and background masks on each frame image in the data set. In the second part, a single object is generated using a pre-trained model based on the reconstruction-based generation method, and a single object generation model based on the adversarial generation network is trained in the data set. The single object model is generated using random noise. After the offline training and reasoning are completed, all generated and reconstructed materials are stored in the resource material library to support the real-time deduction function of the simulation platform.

[0126] The real-time part is dynamic deduction. The foreground, background and skybox resources in the scene are managed through the scene graph structure. The position of each foreground is updated according to different deduction modes when each frame is updated, and the internal time of the object is updated (for 4D models and physical simulation functions). In the process of management through the scene graph, scene control and information interaction functions such as collision detection between objects and event triggering can be added. At the same time, the visual display part of the real-time deduction uses a pluggable display interface UI design for offline training, real-time deduction effect display and interactive control, and control of scene editing object transformation.

[0127] The third level is the upper-level application, which can connect to the foreground dynamic objects in the scene through SOCKET to make control decisions and obtain images through camera simulation rendering. Various dynamic scenes can be constructed through preset scene scripts to test the robustness of the algorithm.

[0128] Finally, a general framework is provided to support more methods and more applications, which is sufficient to support and expand the implementation and operation of closed-loop simulation.

[0129] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0130] In the present invention, a three-dimensional scene is constructed based on a three-dimensional Gaussian splash algorithm, 3D Gaussian points in the three-dimensional scene are determined, and the 3D Gaussian points are projected onto a two-dimensional image plane to obtain two-dimensional Gaussian points. According to the two-dimensional Gaussian points, the depth of each Gaussian body is calculated through an affine transformation matrix. Through depth calculation and sorting, the final rendering result is real and consistent with physical logic. The two-dimensional Gaussian points are mixed according to the depth sorting, and the Alpha A blending transparency algorithm is used to generate the final image color. According to the final color, the three-dimensional scene is rendered to obtain a first rendered image. The first rendered image is optimized through a physically based rendering method (PBR) or a fractional distillation sampling method to eliminate defects and enhance the realism or artistic expression of the image. The present invention introduces PBR (physically based rendering) attributes into Gaussian points, so that the system can process rendering requirements under different lighting conditions in real time, ensure the reliability and accuracy of simulation results, optimize data processing and rendering algorithms, and enhance real-time performance, so that the simulation platform can perform more efficient real-time simulation, significantly improve the authenticity of simulated images, narrow the gap between simulation and reality, and provide strong support for the research and development and testing of autonomous driving systems, thereby promoting the further development of autonomous driving technology.

[0131] Reference Manual Attached Figure 7 , showing a structural schematic diagram of a simulation system based on three-dimensional Gaussian splashing provided by the present invention.

[0132] The present invention further provides a simulation system 20 based on three-dimensional Gaussian splashing, which is applied to the above-mentioned simulation method based on three-dimensional Gaussian splashing, and comprises:

[0133] Processor 201.

[0134] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201 , a simulation method based on three-dimensional Gaussian splashing as in the method embodiment is implemented.

[0135] The simulation system 20 based on three-dimensional Gaussian splashing provided by the present invention can execute the above-mentioned simulation method based on three-dimensional Gaussian splashing and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate on them.

[0136] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0137] In the present invention, a three-dimensional scene is constructed based on a three-dimensional Gaussian splash algorithm, 3D Gaussian points in the three-dimensional scene are determined, and the 3D Gaussian points are projected onto a two-dimensional image plane to obtain two-dimensional Gaussian points. According to the two-dimensional Gaussian points, the depth of each Gaussian body is calculated through an affine transformation matrix. Through depth calculation and sorting, the final rendering result is real and consistent with physical logic. The two-dimensional Gaussian points are mixed according to the depth sorting, and the Alpha A blending transparency algorithm is used to generate the final image color. According to the final color, the three-dimensional scene is rendered to obtain a first rendered image. The first rendered image is optimized through a physically based rendering method (PBR) or a fractional distillation sampling method to eliminate defects and enhance the realism or artistic expression of the image. The present invention introduces PBR (physically based rendering) attributes into Gaussian points, so that the system can process rendering requirements under different lighting conditions in real time, ensure the reliability and accuracy of simulation results, optimize data processing and rendering algorithms, and enhance real-time performance, so that the simulation platform can perform more efficient real-time simulation, significantly improve the authenticity of simulated images, narrow the gap between simulation and reality, and provide strong support for the research and development and testing of autonomous driving systems, thereby promoting the further development of autonomous driving technology.

[0138] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0139] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0140] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0141] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0142] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0143] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0144] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0146] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0147] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0148] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0149] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0150] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, a simulation method based on three-dimensional Gaussian splashing as in the method embodiment is implemented.

[0151] A computer-readable storage medium provided by the present invention can implement the steps and effects of the simulation method based on three-dimensional Gaussian splashing of the above method embodiment. To avoid repetition, the present invention will not go into details.

[0152] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0153] In the present invention, a three-dimensional scene is constructed based on a three-dimensional Gaussian splash algorithm, 3D Gaussian points in the three-dimensional scene are determined, and the 3D Gaussian points are projected onto a two-dimensional image plane to obtain two-dimensional Gaussian points. According to the two-dimensional Gaussian points, the depth of each Gaussian body is calculated through an affine transformation matrix. Through depth calculation and sorting, the final rendering result is real and consistent with physical logic. The two-dimensional Gaussian points are mixed according to the depth sorting, and the Alpha A blending transparency algorithm is used to generate the final image color. According to the final color, the three-dimensional scene is rendered to obtain a first rendered image. The first rendered image is optimized through a physically based rendering method (PBR) or a fractional distillation sampling method to eliminate defects and enhance the realism or artistic expression of the image. The present invention introduces PBR (physically based rendering) attributes into Gaussian points, so that the system can process rendering requirements under different lighting conditions in real time, ensure the reliability and accuracy of simulation results, optimize data processing and rendering algorithms, and enhance real-time performance, so that the simulation platform can perform more efficient real-time simulation, significantly improve the authenticity of simulated images, narrow the gap between simulation and reality, and provide strong support for the research and development and testing of autonomous driving systems, thereby promoting the further development of autonomous driving technology.

[0154] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0155] There are a few points to note:

[0156] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.

[0157] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.

[0158] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.

[0159] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A rendering method based on three-dimensional Gaussian splashing, characterized in that: include: S1: construct a three-dimensional scene; S2: Determine 3D Gaussian points in the three-dimensional scene; S3: Projecting the 3D Gaussian points onto a two-dimensional image plane to determine the two-dimensional Gaussian points; S4: Calculate the depth of each Gaussian body according to the affine transformation matrix to form a sorted list of Gaussian bodies; S5: The final color is calculated by alpha blending algorithm. S6: Rendering the three-dimensional scene according to the final color to obtain a first rendered image; S7: Optimize the first rendered image based on a physical rendering method or a fractional distillation sampling reconstruction method to obtain a second rendered image.

2. The rendering method based on three-dimensional Gaussian splashing according to claim 1, characterized in that: The mathematical definition of the 3D Gaussian point is specifically: ∑=RSS T R T Among them, G(p) represents the probability distribution of a three-dimensional Gaussian point at point p, exp represents the natural exponential function, x represents the position in the multidimensional space, μ represents the mean of the center position (x, y, z), and x, y, and z all represent component variables in the three-dimensional space coordinate system. ∑ The 3D covariance matrix represents the degree of scaling, S represents the scale matrix, and R represents the rotation matrix that describes and transforms the direction of the Gaussian points. -1 represents the inverse matrix operation, T Represents a transpose operation.

3. The rendering method based on three-dimensional Gaussian splashing according to claim 2, characterized in that: The specific projection method of projecting the 3D Gaussian point into the 2D Gaussian point is: ∑′=JWΣW T J T in, ∑ ′ represents the covariance matrix of the two-dimensional Gaussian distribution, J represents the affine approximation Jacobian matrix, and W represents the affine transformation matrix.

4. The rendering method based on three-dimensional Gaussian splashing according to claim 3, characterized in that: The S5 is specifically: The final color is calculated by the Alpha Blending algorithm according to the following formula: Among them, C represents the final color, c i Indicates the color of the i-th layer, α i ′ represents the transparency value of the i-th layer, α j ′ represents the transparency value of the jth layer, and N represents the total number of layers.

5. The rendering method based on three-dimensional Gaussian splashing according to claim 1, characterized in that: The training method of the three-dimensional Gaussian splash algorithm specifically includes: The optimization objective function of the three-dimensional Gaussian splash algorithm is constructed according to the one-norm loss function and the SSIM loss function: L=(1-λ)L1+λL D-SSIM Among them, L represents the optimization objective function, λ represents the weight parameter, L1 represents the one-norm loss function, and L D-SSIM represents the SSIM loss function; The three-dimensional Gaussian splashing algorithm is trained and optimized with the goal of minimizing the optimization objective function.

6. The rendering method based on three-dimensional Gaussian splashing according to claim 1, characterized in that: The properties related to the physical rendering method include: base color, roughness, metalness and normal.

7. The rendering method based on three-dimensional Gaussian splashing according to claim 1, characterized in that: The optimizing the first rendered image based on the physical rendering method in S7 specifically includes: Based on the physical rendering method, the first rendering image is rendered according to the following formula: Among them, L o (p,ω o ) represents the direction ω from the surface point p o The brightness of the emitted light, Ω + Indicates the spatial direction of the positive hemisphere along the normal direction, L e (p,ω o ) indicates that point p is in direction ω o The self-luminous brightness, L i (p,ω i ) represents the direction ω from which the point p comes i The incident light brightness, f r represents the bidirectional reflectance distribution function, n represents the normal, n·ω i represents the cosine of the angle between the incident light and the surface normal, and d represents the differential sign; The bidirectional reflectance distribution function specifically adopts the Disney BRDF principle: F Schlick =F0+(1-F0)(1-cosθ d ) 5 F0=0.04(1-m)+bm G(l,v,h)=G GGX (l)G GGX (v) α=(0.5+r / 2) 2 Among them, f r (l,v) represents the ratio of reflected light from the incident direction l to the observation direction v, f d Represents the diffuse reflection part, D(θ h ) represents the microfacet distribution function, θ h represents the angle between the half vector and the surface normal, F(θ d ) represents the Fresnel effect, θ d Represents the angle between the incident light direction and the surface normal, G(θ l ,θ v ) represents the geometric attenuation function, θ l and θ v They respectively represent the angles between the incident light direction and the observation direction and the surface normal, cos represents the cosine function, b represents the base color, m represents the metalness, n represents the normal, r represents the roughness parameter, F0 represents the specular reflectivity at vertical incidence, h represents the half vector, and α represents the transparency value.

8. The rendering method based on three-dimensional Gaussian splashing according to claim 7, characterized in that: The training methods of the optimization process based on physical rendering include: Construct the loss function of the optimization process based on physical rendering: L′=(1-sum(λ))L PBR +λ n L n +λ smooth L smooth +λ light L light L smooth D||g material (p)-g material (p+ε)||2 Among them, L′ represents the loss function of the optimization process, sum represents the summation symbol, and L PBR represents the loss function of physical rendering, λ n represents the weight of the normal loss function, λ smooth represents the weight of the smooth loss function, λ light Represents the weight of the illumination loss function, L n represents the normal loss function, N represents the normal of the Gaussian point, N Represents the existing normal, L smooth represents the smooth loss function, g material Indicates that it is used for Calculates the function of the material properties at a given position P, where P represents a point on the surface of the object, ε represents the noise, and L light represents the illumination loss function, L c Represents the light intensity in color channel c, where c represents the color channel, R represents the red channel, G represents the green channel, and B represents the blue channel; The optimization process of physically based rendering is trained with the goal of minimizing the loss function.

9. The rendering method based on three-dimensional Gaussian splashing according to claim 1, characterized in that: The step of optimizing the first rendered image based on the fractional distillation sampling reconstruction method in S7 specifically includes: Based on fractional distillation sampling, the parameters of the three-dimensional Gaussian splash model are initialized, and the corresponding image is generated by rendering; Select images rendered at different angles and calculate the similarity between the rendered images and the sample images of the diffusion model: Among them, L Diff (φ,IMg) represents the similarity loss function, t represents the time step, u(0,1) represents the uniform distribution, N(0,I) represents the multivariate standard normal distribution, E represents the expectation, w(t) represents the time-dependent weight function, α t and σ t They represent the time t function of the control data and the time t function of the control noise ratio, ∈ φ represents the output of the diffusion model; The three-dimensional Gaussian splash model parameters are optimized with the goal of minimizing the similarity between the rendered image and the sample image of the diffusion model: Among them, θ L SDS Denotes the loss function L SDS The gradient of the parameter θ, g(θ) represents the generating function, E t,∈ represents the expectation of all possible time steps and noise, It means that the diffusion model is under given conditions Z t , target y and output at time t, represents the partial derivative of the second rendering with respect to the parameter θ.

10. A rendering system based on three-dimensional Gaussian splashing, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the rendering method based on three-dimensional Gaussian splashing as described in any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Three-dimensional reconstruction method containing mirror reflection in dynamic scene based on dual-environment mapping

    CN120259517A

  • Image rendering method, system and device and readable storage medium

    CN120526024A

  • Pavement PBR material inversion method and system

    CN120726241A

  • Internet of Things visual monitoring method and system based on three-dimensional Gaussian splashing and computer equipment

    CN121120935A

  • A three-dimensional Gaussian splash-based Internet of Things visual monitoring method, system and computer device

    CN121120935B