Three-dimensional live-action model construction method, system and device under single light source and medium
By building and pre-training a material model under a single light source and integrating it into the Nerf or 3D Gaussian model, the problem of time-consuming reconstruction of highlight objects under sparse viewing angles is solved, and efficient 3D reconstruction is achieved, which is suitable for outdoor and indoor scenes.
Patent Information
- Application Number
- CN202510727281.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies cannot correctly perform three-dimensional reconstruction of the surface of specular objects on images with sparse perspectives. In particular, the Nerf and 3D Gaussian reconstruction methods require a large number of images with different observation perspectives to correctly fit the changes in the color of the specular object surface with the observation direction, resulting in a long reconstruction time and poor results.
By constructing a material model under a single light source and pre-training it, the pre-training dataset is used to train the trained material model, which is then fused with the Nerf or 3D Gaussian model to form an improved Nerf or 3D Gaussian initial model, reducing the number of images required and lowering the training and rendering overhead.
Under a single light source, it reduces the requirement for the number of images for 3D reconstruction, reduces the video memory overhead for training and rendering, and improves the efficiency of reconstruction of high-gloss object surfaces. It is suitable for 3D reconstruction of large outdoor and indoor scenes.
Smart Images

Figure CN120672943A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of three-dimensional digitization and relates to a three-dimensional real scene model construction technology, and specifically relates to a three-dimensional real scene model construction method, system, equipment and medium under a single light source. Background Art
[0002] 3D reconstruction is a technology that converts two-dimensional data into a three-dimensional model. This technology collects two-dimensional data of objects or scenes (such as images, videos, depth information, etc.) and uses computer vision, graphics, and machine learning technologies to restore their three-dimensional geometric structure and surface information.
[0003] When rendering a 3D model, the color of the pixels rendered on the screen is a mixture of three components: diffuse reflection, specular reflection, and specular reflection. The diffuse reflection component is independent of the viewing direction and is dependent only on the intensity of the light radiation received by the object's surface. The specular reflection component is dependent on the viewing direction, with the brightness increasing at angles closer to the surface's reflected direction of the incident light, even creating halos and bright spots. Specular reflection is also dependent on the viewing direction and is a clear reflection of other objects. The material parameters of different object surfaces determine the weighting of these three components, which ultimately results in the rendered image.
[0004] In the past, conventional 3D reconstruction methods generally fall into the following three categories:
[0005] 1. 3D reconstruction methods based on Structure from Motion (SFM) use feature point matching to calculate a feature point cloud that satisfies global constraints. This method then densifies the feature point cloud and performs surface reconstruction, ultimately yielding a textured 3D model. This method works well in scenes with rich, non-periodic texture features, but performance deteriorates or reconstruction fails when insufficient feature points are extracted. Furthermore, it can only model purely diffuse objects.
[0006] 2. 3D reconstruction methods based on stereo vision. This method calculates disparity and depth maps from images taken by multiple cameras in the same direction, blends the point clouds converted from the depth maps at multiple locations, and then reconstructs the surface of the point clouds to create a 3D model. This method can only model purely diffuse objects and fails in areas with sparse, solid colors or periodically repeating textures due to an inability to accurately estimate disparity.
[0007] 3. Active scanning (such as structured light / LiDAR) 3D reconstruction methods: This method fuses multiple frames of directly acquired 3D point clouds and then reconstructs the surface to create a 3D model. While this method can avoid the problems of feature point extraction failure and depth estimation failure, it can still only model purely diffuse reflective objects.
[0008] Currently, the more novel three-dimensional reconstruction methods are Nerf reconstruction and 3D Gaussian reconstruction, which can model high-light objects but have high requirements on the number of images.
[0009] 1. The Nerf reconstruction method uses implicit functions to model the color and opacity of a point in space in a specified viewing direction as input (x, y, z) coordinates and (dir x 、dir y 、dir z ) direction, and fit this implicit function through the back propagation of the neural network. When rendering, each pixel of the camera image is regarded as a ray, and a large number of sampling points are selected on this ray. The trained neural network is used to infer the color and opacity of each sampling point, and then the opacity of the color of the sampling point on the ray is mixed using the volume rendering formula to obtain the final rendered color. Although Nerf's method can model highlights and anisotropic objects, since the neural network is trained from scratch each time without prior knowledge, the reconstruction takes a long time, and a large number of images with different observation angles are required for highlight-rich objects to correctly fit the implicit function of the surface color of the highlight object as the observation direction changes. Otherwise, the change of the surface color of the highlight object as the observation direction changes will not be consistent with the actual situation, and even cause the problem of pure black and pure white in some observation directions.
[0010] The 3D Gaussian method replaces the neural network in Nerf with spherical harmonics. Only the coefficients of the spherical harmonics are learned, and a Gaussian kernel is used to store the sparse information of the sneakers. During 3D reconstruction, the Gaussian kernels are dynamically split and merged, and transparent Gaussian kernels are removed. While this method improves training speed, because spherical harmonics are less capable of expressing anisotropy than neural networks, reconstruction of specular objects still requires a large number of images from different perspectives. Summary of the Invention
[0011] In order to solve the problem that the Nerf 3D reconstruction method and the 3D Gaussian 3D reconstruction method cannot correctly model the surface of high-gloss objects in images with sparse perspectives, the present invention discloses a method for constructing a 3D real scene model under a single light source. This method uses a pre-trained material model to reduce the number of images required for real scene 3D reconstruction, thereby achieving 3D reconstruction of large scenes with fewer images, thereby indirectly reducing the cost of the reconstruction process. Specifically, the method includes the following steps:
[0012] S1. Construct a material model under a single light source through the object's surface property unit vector and material properties;
[0013] S2. Training the material model using a pre-training dataset to obtain a trained material model, wherein the pre-training dataset includes data on color rendering of a material with random material attributes under a random single light source direction and any valid observation direction;
[0014] S3, fusing the trained material model with the Nerf model to obtain an improved Nerf initial model, or fusing the trained material model with the 3D Gaussian model to obtain an improved 3D Gaussian initial model;
[0015] S4. Using large scene image data and estimated camera pose, the improved Nerf initial model is trained to obtain an improved Nerf model, or the improved 3D Gaussian initial model is trained to obtain an improved 3D Gaussian model.
[0016] After the improved Nerf model or the improved 3D Gaussian model is built, it can be loaded through the rendering engine and the image with known camera pose is used as input to reconstruct the three-dimensional scene.
[0017] An embodiment of the present invention further provides a system for constructing a three-dimensional real scene model under a single light source, comprising a material model establishment module, a data set establishment module, a model pre-training module, a fusion module and an improved model training module.
[0018] The material model building module is used to build a material model under a single light source through the surface property unit vector and material properties of the object;
[0019] The data set establishment module is used to establish data of rendering colors of materials with random material properties under random single light source directions and any effective observation directions to form a pre-training data set;
[0020] The model pre-training module is used to train the material model using a pre-training data set to obtain a trained material model;
[0021] The fusion module is used to fuse the trained material model with the Nerf model to obtain an improved Nerf initial model, or to fuse the trained material model with the 3D Gaussian model to obtain an improved 3D Gaussian initial model;
[0022] The improved model training module is used to use large scene image data and estimated camera pose to train the improved Nerf initial model to obtain an improved Nerf model, or to train the improved 3D Gaussian initial model to obtain an improved 3D Gaussian model.
[0023] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned methods for constructing a three-dimensional real scene model under a single light source, thereby solving the problem in the prior art that the Nerf three-dimensional reconstruction method and the 3D Gaussian three-dimensional reconstruction method cannot correctly model the surface of high-gloss objects in images with sparse perspectives.
[0024] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program for executing any of the above-mentioned methods for constructing a three-dimensional real scene model under a single light source, so as to solve the problem in the prior art that the Nerf three-dimensional reconstruction method and the 3D Gaussian three-dimensional reconstruction method cannot correctly model the surface of high-gloss objects in images with sparse perspectives.
[0025] Compared with the prior art, the at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects:
[0026] The present invention discloses a method for constructing a three-dimensional real scene model under a single light source. The method constructs a material model under a single light source, pre-trains the model using a pre-training data set to obtain a trained material model, and then uses the trained material model to modify a traditional Nerf model or a 3D Gaussian model to obtain an improved Nerf initial model or an improved 3D Gaussian initial model. After training, the method obtains a final improved Nerf model or an improved 3D Gaussian model. When the improved model is used to construct a three-dimensional scene under a single light source, the requirement for the number of images required for three-dimensional reconstruction of real scenes in outdoor / indoor scenes can be reduced. By reducing the number of necessary parameters stored during training, the video memory overhead for training and rendering is reduced, making the three-dimensional reconstruction of real scenes more applicable to applications in outdoor / indoor scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0028] Figure 1 A flowchart of a method for constructing a three-dimensional real scene model under a single light source;
[0029] Figure 2 The execution flow of the program for constructing a 3D real scene model under a single light source;
[0030] Figure 3 The architecture diagram of the system for building a 3D reality model under a single light source;
[0031] Figure 4 A schematic diagram of a computer device disclosed in an embodiment of the present invention;
[0032] Among them, 301, material model establishment module; 302, data set establishment module; 303, model pre-training module; 304, fusion module; 305, improved model training module; 401, memory; 402, processor. DETAILED DESCRIPTION
[0033] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0034] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, in the absence of conflict, the following embodiments and the features of the embodiments can be combined with each other. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.
[0035] The embodiment of the present invention discloses a method for constructing a three-dimensional real scene model under a single light source. Figure 1 and Figure 2 As shown, the method includes the following steps:
[0036] S1. Construct a material model under a single light source through the object's surface property unit vector and material properties;
[0037] S2. Training the material model using a pre-training dataset to obtain a trained material model, wherein the pre-training dataset includes data on color rendering of a material with random material attributes under a random single light source direction and any valid observation direction;
[0038] S3, fusing the trained material model with the Nerf model to obtain an improved Nerf initial model, or fusing the trained material model with the 3D Gaussian model to obtain an improved 3D Gaussian initial model;
[0039] S4. Using large scene image data and estimated camera pose, the improved Nerf initial model is trained to obtain an improved Nerf model, or the improved 3D Gaussian initial model is trained to obtain an improved 3D Gaussian model.
[0040] After the improved Nerf model or the improved 3D Gaussian model is built, it can be loaded through the rendering engine and the image with known camera pose is used as input to reconstruct the three-dimensional scene.
[0041] In implementation, whether it is an outdoor scene or an indoor scene, when there is only one global light source, for example, the light source of the outdoor scene comes from the parallel light of the sun, and the light source of the indoor scene is a single light source at a set position, when constructing the 3D scene, only the influence of the material properties of the object under a single light source on the rendered color needs to be considered. At this time, the material model under a single light source can be constructed by the surface property unit vector and material properties of the object. Specifically, the material model construction process includes the following steps:
[0042] S11. Using a multi-layer perceptron (MLP), implicit modeling is performed using the material properties of the object to obtain a light reflection characteristic function, where the material properties include the surface color, metallicity, roughness, and anisotropy intensity of the object. Using the light reflection characteristic function, a feature vector of the material itself can be obtained that is independent of the observation direction and extracted through the surface color, metallicity, roughness, and anisotropy intensity of the object.
[0043] During implementation, the illumination reflection characteristic function can be obtained by using three hidden layers in a multi-layer perceptron MLP, selecting 64 hidden neurons in each hidden layer, selecting a Softplus activation function, and using the material properties of the object for implicit modeling.
[0044] S12. Calculate a dot product using a single light source direction unit vector and a surface attribute unit vector to establish a lighting influence function in a local coordinate system. The surface attribute unit vectors include a normal direction unit vector, a tangent direction unit vector, a binormal direction unit vector, a viewing direction unit vector, and a half-angle direction unit vector. Using the lighting influence function, a unit vector representing the directional characteristics in the local coordinate system extracted from the normal, tangent, binormal, viewing direction, and half-angle direction can be obtained. This unit vector is not directly related to these quantities and can be generalized to any point on the surface of an object with any normal, tangent, or binormal direction.
[0045] During implementation, the single light source can be a light source for an outdoor scene or an indoor scene. The dot product calculation is performed using the single light source direction unit vector and the surface attribute unit vector to establish the lighting influence function in the local coordinate system, including:
[0046] S121, taking the single light source direction unit vector, the observation direction unit vector, and the half-angle direction unit vector as first parameters, and taking the normal direction unit vector, the tangent direction unit vector, and the binormal direction unit vector as second parameters;
[0047] S122: Perform dot product calculations on each of the first parameters and each of the second parameters, and establish a lighting influence function in a local coordinate system using all dot product calculation results.
[0048] Specifically, it can be expressed by the following relationship:
[0049] F2(N,L,S,H,V,Light))=(y1,y2,y3,y4,y5,y6,y7,y8,y9);
[0050] y1=dot(N,Light), where y1 is the dot product of the normal direction unit vector and the single light source direction unit vector;
[0051] y2=dot(L,Light), where y2 is the dot product of the tangent direction unit vector and the single light source direction unit vector;
[0052] y3=dot(S,Light), where y3 is the dot product of the unit vector in the direction of the binormal and the unit vector in the direction of the single light source;
[0053] y4=dot(N,V), where y1 is the dot product of the normal direction unit vector and the observation direction unit vector;
[0054] y5=dot(N,H), where y1 is the dot product of the normal direction unit vector and the half-angle direction unit vector;
[0055] y6=dot(L,V), where y1 is the dot product of the tangent direction unit vector and the observation direction unit vector;
[0056] y7=dot(L,H), where y1 is the dot product of the tangent direction unit vector and the half-angle direction unit vector;
[0057] y8=dot(S,V), where y1 is the dot product of the binormal direction unit vector and the observation direction unit vector;
[0058] y9=dot(S,H), where y1 is the dot product of the unit vector in the binormal direction and the unit vector in the half-angle direction.
[0059] S13. Through the multi-layer perceptron MLP, the illumination reflection characteristic function and the illumination influence function are used for implicit modeling to obtain a rendering function, thereby completing the construction of the material model under a single light source. The rendering function is used to render the final color by using the feature vector of the material itself that is independent of the direction in step S11 and the direction feature vector in the local coordinate system that is related to the direction in step S12.
[0060] During implementation, the rendering function can be obtained by using two hidden layers in a multi-layer perceptron MLP, selecting 64 hidden neurons in each hidden layer, selecting a Softplus activation function, and using a light reflection characteristic function and a light influence function for implicit modeling.
[0061] Optionally, the expression of the above material model is:
[0062] Color = F3(F1(C,M,R,I),F2(N,L,S,H,V,Light)), where Color is the rendering color, F1 is the lighting reflection characteristic function, F2 is the lighting influence function, F3 is the rendering function, C is the surface color, M is the metalness, R is the roughness, I is the anisotropy intensity value, N is the normal direction unit vector, L is the tangent direction unit vector, S is the binormal direction unit vector, V is the viewing direction unit vector, H is the half-angle direction unit vector, and Light is the single light source direction unit vector.
[0063] In addition, this material model can be pre-trained using the constructed pre-training dataset to learn how the color of the object surface changes with lighting and observation angle under different material parameters. Therefore, for points on the object surface, only a small amount of observation sampling data at key angles is needed to converge to the appropriate material parameters through backpropagation, and render normally at other unobserved angles. At this time, the trained material model can be obtained.
[0064] The above-mentioned trained material model is essentially equivalent to the shading model in the rendering engine, but it uses a different expression form to approximate it. Since this expression is fully differentiable, it can be added as a module to the existing Nerf model and 3D Gaussian model, which require backpropagation. The parameters of the trained material model are not updated in the backpropagation to improve the existing Nerf model and 3D Gaussian model.
[0065] During implementation, in the above step S2, the random parameters in the pre-training data set include color, metallicity, roughness, anisotropic intensity value, angle between parallel light source and the ground (ie, single light source direction) and observation angle (ie, camera direction).
[0066] During implementation, in the above step S3, the trained material model can be used to replace the sampling point implicit function modeling model in the Nerf model to obtain an improved Nerf initial model, and the trained material model can be used to replace the model for color modeling by spherical harmonic functions in the 3D Gaussian model to obtain an improved 3D Gaussian initial model.
[0067] For the traditional Nerf-type real-scene 3D reconstruction method, the neural network model modeled at a single sampling point can generally be expressed as:
[0068] RGB=F5(F4(x,y,z),dir x ,dir y ,dir z ), Density = f6(x,y,z);
[0069] Among them, x, y, z are the coordinates of a sampling point in the space in the X direction, Y direction and Z direction respectively; dir x 、dir y 、dir z where is the vector of the observation direction in the X, Y, and Z directions, i.e., the camera direction; RGB is the color of this sampling point, and Density is the density of this sampling point. Density can be further used to calculate the opacity for the color value of the sampling point along the opacity blending ray during volume rendering. When implementing the present invention, the trained material model can be used to improve the above neural network model to obtain an improved Nerf initial model, which is expressed as:
[0070] RGB1=F3(F1(C,M,R,I),F2(N,L,S,H,V,Light)), Density1=F6(x,y,z).
[0071] By using regularization terms to constrain N, L, and S, these three vectors are made orthogonal, and the direction of N tends to be consistent with the opposite direction of the gradient direction of F6 in (x, y, z). C, M, R, I, and Light in the above formula are parameters that need to be learned. In large outdoor scenes, Light is a consistent parallel light, so it is represented by a shared global learnable parameter. In indoor scenes with a single light source, Light is the direction vector from an unknown position Loc in space to the sampling point. Loc is represented by a global learnable parameter during training. C, M, R, and I are the color, metalness, roughness, and anisotropy values of the sampling point, which need to be predicted by a stacked MLP or stored in a dense grid.
[0072] For the traditional 3D Gaussian type real scene 3D reconstruction method, the spherical harmonics in the 3D Gaussian model can be replaced by the trained material model. The formula for rendering the color of each Gaussian kernel in the traditional 3D Gaussian model is:
[0073] Where A is the highest order of spherical harmonics, c l,m Represents the shoe coefficient related to direction, Y l,m (d) represents a basis function, and each RGB color needs to be calculated according to this formula. The expression of the improved 3D Gaussian initial model after replacement is:
[0074] Color rgb =F3(F1(C,M,R,I),F2(N,L,S,H,V,Light)). In the traditional 3D Gaussian model, each Gaussian kernel has its own orientation. Therefore, the shortest axis direction is used as the normal direction, and the tangent direction is learned using a set of parameters. In this case, each Gaussian kernel only needs to store the C, M, R, and I parameters, and the original sneaker coefficients do not need to be stored.
[0075] Compared to the traditional 3D Gaussian model, the improved 3D Gaussian initial model takes longer to compute per calculation. However, because the implicit function used in the material modeling is more consistent with the variations in illumination and viewing angles than the spherical harmonics, the number of iterations required for convergence of the parameters of a single Gaussian kernel can be reduced. Therefore, the actual training time is comparable to that of the traditional 3D Gaussian method. Furthermore, because the method of the present invention requires fewer images and significantly reduces the number of parameters stored for each Gaussian kernel during training, it is better suited for the reconstruction of large scenes.
[0076] In addition, the trained material model in the present invention is pre-trained from the synthetic data of the rendering engine, and the consistency of relevant material parameters is guaranteed. Therefore, the original shading model can be used as a substitute during rendering to reduce rendering overhead.
[0077] The present invention discloses a method for constructing a three-dimensional real scene model under a single light source. The method constructs a material model under a single light source, pre-trains the model using a pre-training data set to obtain a trained material model, and then uses the trained material model to modify a traditional Nerf model or a 3D Gaussian model to obtain an improved Nerf initial model or an improved 3D Gaussian initial model. After training, the method obtains a final improved Nerf model or an improved 3D Gaussian model. When the improved model is used to construct a three-dimensional scene under a single light source, the requirement for the number of images required for three-dimensional reconstruction of real scenes in outdoor / indoor scenes can be reduced. By reducing the number of necessary parameters stored during training, the video memory overhead for training and rendering is reduced, making the three-dimensional reconstruction of real scenes more applicable to applications in outdoor / indoor scenes.
[0078] Based on the same inventive concept, a system for constructing a three-dimensional real scene model under a single light source is also provided in an embodiment of the present invention, as described in the following embodiment. Since the principle of solving the problem by the system for constructing a three-dimensional real scene model under a single light source is similar to the method disclosed in the above embodiment, the implementation of the system for constructing a three-dimensional real scene model under a single light source can refer to the implementation of the method disclosed in the above embodiment, and the repeated parts will not be repeated. As used below, the term "unit" or "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0079] Figure 3 This is a structural block diagram of a system for constructing a three-dimensional real scene model under a single light source disclosed in an embodiment of the present invention. Figure 3 As shown, the system includes a material model building module 301, a data set building module 302, a model pre-training module 303, a fusion module 304 and an improved model training module 305, and the structure is described below.
[0080] The material model building module 301 is used to build a material model under a single light source through the surface property unit vector and material properties of the object;
[0081] The data set establishment module 302 is used to establish a pre-training data set for data on the rendering color of a material with random material properties under a random single light source direction and any valid observation direction;
[0082] The model pre-training module 303 is used to train the material model using a pre-training data set to obtain a trained material model;
[0083] The fusion module 304 is used to fuse the trained material model with the Nerf model to obtain an improved Nerf initial model, or to fuse the trained material model with the 3D Gaussian model to obtain an improved 3D Gaussian initial model;
[0084] The improved model training module 305 is used to use large scene image data and estimated camera pose to train the improved Nerf initial model to obtain an improved Nerf model, or to train the improved 3D Gaussian initial model to obtain an improved 3D Gaussian model.
[0085] Furthermore, the material model establishment module 301 includes an illumination reflection characteristic function establishment module, an illumination influence function establishment module and a rendering function establishment module.
[0086] The light reflection characteristic function establishment module is used to obtain the light reflection characteristic function by implicitly modeling the material properties of the object through a multi-layer perceptron (MLP), wherein the material properties include the surface color, metallicity, roughness, and anisotropy intensity of the object;
[0087] The illumination influence function establishment module is used to establish the illumination influence function in the local coordinate system by performing dot product calculation using a single light source direction unit vector and an object's surface attribute unit vector, wherein the surface attribute unit vector includes a normal direction unit vector, a tangent direction unit vector, a binormal direction unit vector, an observation direction unit vector, and a half-angle direction unit vector;
[0088] The rendering function establishment module is used to obtain the rendering function through the multi-layer perceptron MLP, using the light reflection characteristic function and the light influence function for implicit modeling, and complete the construction of the material model under a single light source.
[0089] In this embodiment, a computer device is provided, such as Figure 4 As shown, it includes a memory 401, a processor 402 and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, any of the above-mentioned methods for constructing a three-dimensional real scene model under a single light source is implemented.
[0090] Specifically, the computer device may be a computer terminal, a server or a similar computing device.
[0091] In this embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program for executing any of the above-mentioned methods for constructing a three-dimensional real scene model under a single light source.
[0092] Specifically, computer-readable storage media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable storage media does not include transitory media such as modulated data signals and carrier waves.
[0093] Obviously, those skilled in the art should understand that the various modules or steps of the above-mentioned embodiments of the present invention can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices. Alternatively, they can be implemented using program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into separate integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the embodiments of the present invention are not limited to any specific combination of hardware and software.
[0094] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for constructing a three-dimensional real scene model under a single light source based on a material model, characterized in that: include: Construct a material model under a single light source through the object's surface property unit vector and material properties; The material model is trained by using a pre-training data set to obtain a trained material model, wherein the pre-training data set includes data on the color rendering of a material with random material properties under a random single light source direction and any valid observation direction; The trained material model is fused with the Nerf model to obtain an improved Nerf initial model, or the trained material model is fused with the 3D Gaussian model to obtain an improved 3D Gaussian initial model; Using large scene image data and estimated camera pose, the improved Nerf initial model is trained to obtain an improved Nerf model, or the improved 3D Gaussian initial model is trained to obtain an improved 3D Gaussian model.
2. The method for constructing a three-dimensional real scene model under a single light source according to claim 1, characterized in that: The material model under a single light source is constructed through the object's surface property unit vector and material properties, including: Through the multi-layer perceptron (MLP), the material properties of the object are implicitly modeled to obtain the light reflection characteristic function, wherein the material properties include the surface color, metallicity, roughness and anisotropy strength of the object; The illumination influence function in the local coordinate system is established by performing dot product calculation using a single light source direction unit vector and a surface attribute unit vector, wherein the surface attribute unit vector includes a normal direction unit vector, a tangent direction unit vector, a binormal direction unit vector, an observation direction unit vector, and a half-angle direction unit vector; Through the multi-layer perceptron MLP, the lighting reflection characteristic function and the lighting influence function are implicitly modeled to obtain the rendering function, completing the construction of the material model under a single light source.
3. The method for constructing a three-dimensional real scene model under a single light source according to claim 2, wherein: The expression of the material model is Color=F3(F1(C,M,R,I),F2(N,L,S,H,V,Light)), where Color is the rendering color, F1 is the lighting reflection characteristic function, F2 is the lighting influence function, F3 is the rendering function, C is the surface color, M is the metalness, R is the roughness, I is the anisotropy intensity value, N is the normal direction unit vector, L is the tangent direction unit vector, S is the binormal direction unit vector, V is the observation direction unit vector, H is the half-angle direction unit vector, and Light is the single light source direction unit vector.
4. The method for constructing a three-dimensional real scene model under a single light source according to claim 3, wherein: Through the three hidden layers in the multi-layer perceptron MLP, each hidden layer uses 64 hidden neurons, selects the Softplus activation function, and uses the material properties of the object for implicit modeling to obtain the light reflection characteristic function; The rendering function is obtained by using two hidden layers in the multi-layer perceptron MLP, 64 hidden neurons in each hidden layer, the Softplus activation function, and the illumination reflection characteristic function and the illumination influence function for implicit modeling.
5. The method for constructing a three-dimensional real scene model under a single light source according to claim 3, characterized in that: The illumination influence function in the local coordinate system is established by performing dot product calculation using a single light source direction unit vector and a surface attribute unit vector, including: The single light source direction unit vector, the observation direction unit vector and the half-angle direction unit vector are used as the first parameter, and the normal direction unit vector, the tangent direction unit vector and the binormal direction unit vector are used as the second parameter; A dot product calculation is performed using each of the first parameters and each of the second parameters, and a lighting influence function in a local coordinate system is established using all dot product calculation results.
6. The method for constructing a three-dimensional real scene model under a single light source according to claim 1, characterized in that: The trained material model is used to replace the implicit function modeling model of the sampling points in the Nerf model to obtain an improved Nerf initial model, and the trained material model is used to replace the color modeling model through spherical harmonic functions in the 3D Gaussian model to obtain an improved 3D Gaussian initial model.
7. The method for constructing a three-dimensional real scene model under a single light source according to claim 1, characterized in that: The single light source is a light source for an outdoor scene or an indoor scene.
8. A system for constructing a three-dimensional real scene model under a single light source, characterized in that: include: A material model building module, wherein the material model building module is used to build a material model under a single light source by using the surface property unit vector and material properties of the object; A data set establishment module, wherein the data set establishment module is used to establish a pre-training data set for data on the rendering color of a material with random material properties under a random single light source direction and any valid observation direction; A model pre-training module, wherein the model pre-training module is used to train the material model using a pre-training data set to obtain a trained material model; A fusion module, wherein the fusion module is used to fuse the trained material model with the Nerf model to obtain an improved Nerf initial model, or to fuse the trained material model with the 3D Gaussian model to obtain an improved 3D Gaussian initial model; An improved model training module is used to use large scene image data and estimated camera pose to train the improved Nerf initial model to obtain an improved Nerf model, or to train the improved 3D Gaussian initial model to obtain an improved 3D Gaussian model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for constructing a three-dimensional real scene model under a single light source according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program for executing the method for constructing a three-dimensional real scene model under a single light source according to any one of claims 1 to 7.