Three-dimensional scene reconstruction method and apparatus, device, medium, and program product
By acquiring 3D scene data, using sparse point cloud and camera parameter information for model initialization and rendering, and combining neural radiation field or 3D Gaussian volume for self-supervised training, the problem of coarse quality of existing 3D mesh models is solved, and efficient and low-cost 3D scene reconstruction is achieved.
Patent Information
- Application Number
- PCT/CN2025/093166
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-28
- Filing Date
- 2025-05-07
- Publication Date
- 2026-01-02
AI Technical Summary
The existing 3D reconstruction methods produce relatively coarse 3D mesh models, which require further repair by professionals to meet the requirements of 3D reconstruction, and the cost is high.
By acquiring 3D scene data, determining sparse point cloud data and camera parameter information, performing model initialization and rendering, and using neural radiation fields or 3D Gaussian volumes for self-supervised training of the 3D data model, a higher quality target 3D data model is generated.
It enables automated, high-quality reconstruction of 3D scenes, reducing reconstruction costs and improving the detail and rendering quality of scene models.
Smart Images

Figure CN2025093166_02012026_PF_FP_ABST
Abstract
Description
Three-dimensional scene reconstruction method, device, equipment, medium and program product
[0001] This application claims priority to Chinese Patent Application No. 202410870567.8, filed on June 28, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to a three-dimensional scene reconstruction method, device, equipment, medium and program product. BACKGROUND
[0003] With the rapid development of computer technology, it is often necessary to reconstruct a scene using collected three-dimensional scene data to obtain a three-dimensional scene model. The three-dimensional scene model reconstructed by the current three-dimensional reconstruction method is usually a three-dimensional mesh model. However, such a three-dimensional mesh model has a relatively rough quality, and professional personnel need to further repair the three-dimensional mesh model to meet the three-dimensional reconstruction requirements. SUMMARY
[0004] The present disclosure provides a three-dimensional scene reconstruction method, device, equipment, medium and program product to realize automatic reconstruction of a three-dimensional scene with higher quality, reduce reconstruction cost, and improve the fineness and rendering quality of the scene model.
[0005] In a first aspect, embodiments of the present disclosure provide a three-dimensional scene reconstruction method, comprising:
[0006] obtaining three-dimensional scene data collected for a three-dimensional scene to be reconstructed, the three-dimensional scene data comprising a scene image of the three-dimensional scene;
[0007] determining sparse point cloud data corresponding to the three-dimensional scene and camera parameter information corresponding to the scene image according to the scene image;
[0008] performing model initialization according to the sparse point cloud data to obtain a current three-dimensional data model corresponding to the three-dimensional scene;
[0009] performing rendering according to the camera parameter information and the current three-dimensional data model to obtain a current color map, a current depth map and a current normal map under the view angle of the scene image, and determining a pseudo normal map under the view angle of the scene image according to the current depth map;
[0010] training the current three-dimensional data model according to the current color map, the current normal map, the pseudo normal map and an actual color map of the scene image to obtain a target three-dimensional data model after training is completed.
[0011] In a second aspect, embodiments of the present disclosure further provide a three-dimensional scene reconstruction device, comprising:
[0012] A 3D scene data acquisition module is used to acquire 3D scene data collected from the 3D scene to be reconstructed, wherein the 3D scene data includes scene images of the 3D scene;
[0013] The sparse point cloud data determination module is used to determine the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to the scene image based on the scene image.
[0014] The model initialization module is used to initialize the model based on the sparse point cloud data and obtain the current three-dimensional data model corresponding to the three-dimensional scene.
[0015] The model rendering module is used to render according to the camera parameter information and the current 3D data model, to obtain the current color map, current depth map and current normal map under the scene image view, and to determine the pseudo normal map under the scene image view based on the current depth map;
[0016] The model training module is used to train the current 3D data model based on the current color map, the current normal map, the pseudo normal map, and the actual color map of the scene image, so as to obtain the target 3D data model after training.
[0017] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0018] One or more processors;
[0019] Storage device for storing one or more programs.
[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional scene reconstruction method as described in any of the embodiments of this disclosure.
[0021] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the three-dimensional scene reconstruction method as described in any of the embodiments of this disclosure.
[0022] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the three-dimensional scene reconstruction method as described in any of the embodiments of this disclosure. Attached Figure Description
[0023] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0024] FIG. 1 is a flow diagram of a method for reconstructing a three-dimensional scene according to an embodiment of the present disclosure;
[0025] FIG. 2 is a flow diagram of another method for reconstructing a three-dimensional scene according to an embodiment of the present disclosure;
[0026] FIG. 3 is a flow diagram of yet another method for reconstructing a three-dimensional scene according to an embodiment of the present disclosure;
[0027] FIG. 4 is a structural diagram of a device for reconstructing a three-dimensional scene according to an embodiment of the present disclosure; and
[0028] FIG. 5 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.
[0030] It should be understood that the various steps of the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0031] As used herein, the term "comprises" and its variants are to be construed as open- ended, that is, "including but not limited to." The term "based on" is to be construed as "based at least in part on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related terms are to be construed accordingly.
[0032] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0033] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".
[0034] Names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0035] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and relevant provisions.
[0036] FIG. 1 is a flowchart of a three-dimensional scene reconstruction method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of reconstructing a real three-dimensional scene, and the reconstructed model can be directly applied to a virtual live room and other virtual application scenarios. The method can be executed by a three-dimensional scene reconstruction device, which can be implemented in the form of software and / or hardware, and can be implemented by an electronic device, which can be a mobile terminal, a PC terminal, or a server.
[0037] As shown in FIG. 1, the three-dimensional scene reconstruction method specifically includes the following steps:
[0038] S110, acquiring three-dimensional scene data collected for a three-dimensional scene to be reconstructed, the three-dimensional scene data including scene images of the three-dimensional scene.
[0039] The three-dimensional scene can be any three-dimensional entity scene that needs to be reconstructed in reality. The three-dimensional scene can be a closed scene or a non-closed scene. For example, the three-dimensional scene can be a tourist attraction, etc. The three-dimensional scene data refers to scene data collected for the three-dimensional scene by a collection device. The three-dimensional scene data can include but is not limited to scene images of the three-dimensional scene. The scene images can be real images of the three-dimensional scene collected by an image collection device. The image collection device can include but is not limited to at least one of a single-lens reflex, a mobile phone, a drone, a panoramic camera, a laser device, and a depth camera. The number of scene images is multiple, and all scene images contain all scene contents of the three-dimensional scene, so as to reconstruct a complete three-dimensional scene model based on the scene images.
[0040] Specifically, at least one acquisition device is used to collect the three-dimensional scene to be reconstructed according to the corresponding scene collection requirements of each acquisition device, and three-dimensional scene data collected is obtained. For example, at least one image acquisition device can be used to collect images of the three-dimensional scene to be reconstructed according to the corresponding scene collection requirements of each image acquisition device, and all collected scene images are obtained. The corresponding scene collection requirements of each acquisition device can include information such as the placement position and collection path of the acquisition device, so that scene images meeting the scene reconstruction requirements can be quickly obtained according to the scene collection requirements. By pre-configuring the corresponding simple and standard scene collection requirements of each acquisition device, non-professionals can also quickly collect high-quality three-dimensional scene data, thereby reducing the technical threshold for collection and improving the three-dimensional reconstruction effect.
[0041] S120, determining sparse point cloud data corresponding to the three-dimensional scene and camera parameter information corresponding to the scene image according to the scene image.
[0042] The sparse point cloud data can refer to data of some feature points in the three-dimensional scene. The feature points refer to three-dimensional space points with obvious features in the three-dimensional scene, such as corner and edge points of a building. The sparse point cloud data can be used to represent the sparse geometric structure of the three-dimensional scene. The camera parameter information can refer to current camera information when the scene image is collected. For example, the camera parameter information can include camera pose information and intrinsic parameter information. Different scene images can correspond to different camera parameter information.
[0043] Specifically, the Structure from Motion (SfM) algorithm and the Bundle Adjustment (BA) algorithm can be used to reconstruct the sparse point cloud and estimate the camera pose of all collected scene images, so as to obtain the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to each scene image.
[0044] S130, initializing a model according to the sparse point cloud data to obtain a current three-dimensional data model corresponding to the three-dimensional scene.
[0045] The current three-dimensional data model can refer to a three-dimensional data model at the current time. The current three-dimensional data model can change with the dynamic change of the model parameter value. Different model parameter values correspond to different three-dimensional data models. The three-dimensional data model can be a collection of spatial data used to describe the three-dimensional scene. The three-dimensional data model can be a model for reconstructing the three-dimensional scene into a light field. The light field is a function used to describe the amount of light through each point and each direction in the three-dimensional space.
[0046] Exemplarily, the current three-dimensional data model can include a neural radiance field (NeRF) or a three-dimensional Gaussian (3D Gaussian). The neural radiance field is to store radiance information from any spatial position to any direction by using a neural network. The radiance information can also be referred to as an energy distribution, which can refer to information such as color, brightness, and shadow. The neural radiance field is to reconstruct a three-dimensional scene into a radiance field represented by a neural network. The three-dimensional Gaussian is to reconstruct a three-dimensional scene into a Gaussian represented by a Gaussian distribution. The neural radiance field and the three-dimensional Gaussian have stronger scene expression capability and view rendering capability than a traditional three-dimensional mesh model, so as to improve the fineness and rendering quality of the scene model. The model parameters in the neural radiance field can include a volume density value, a color value, a depth value, and a normal value of a voxel in the three-dimensional scene. The model parameters in the three-dimensional Gaussian can include position information, a spherical harmonic function coefficient, transparency, a rotation value, a scale value, and a normal value of a voxel in the three-dimensional scene.
[0047] Specifically, initial values of each model parameter in the three-dimensional data model can be determined according to the sparse point cloud data, and the model parameters are assigned values, so as to realize model initialization and obtain an initialized current three-dimensional data model. For example, the volume density value, the color value, the depth value, and the normal value of each voxel in the neural radiance field are initialized according to the sparse point cloud data, and an initialized neural radiance field is obtained. Alternatively, the position information, the spherical harmonic function coefficient, the transparency, the rotation value, the scale value, and the normal value of each voxel in the three-dimensional Gaussian are initialized according to the sparse point cloud data, and an initialized three-dimensional Gaussian is obtained.
[0048] In S140, the current color map, the current depth map, and the current normal map under the scene image view are obtained by rendering according to the camera parameter information and the current three-dimensional data model, and the pseudo normal map under the scene image view is determined according to the current depth map.
[0049] The scene image view refers to the camera view when the scene image is taken. The current color map can include a current color value corresponding to each pixel. The current depth map can include a current depth value corresponding to each pixel. The current normal map can include a current normal value corresponding to each pixel. The current color map, the current depth map, and the current normal map have the same size as the scene image, and the pixels thereof are one-to-one corresponding. The current color value, the current depth value, and the current normal value corresponding to each pixel are obtained by rendering the current three-dimensional data model. The pseudo normal map can include a pseudo normal value corresponding to each pixel. The pseudo normal value can be regarded as the actual normal value, i.e., the true value of the normal, corresponding to the pixel.
[0050] Specifically, for each scene image, based on the camera parameter information corresponding to the scene image, the current three-dimensional data model is rendered to render the current three-dimensional data model into a two-dimensional current color map, a current depth map and a current normal map with the same perspective as the scene image, and a local gradient calculation is performed according to the current depth map to obtain a pseudo-normal map under the perspective of the scene image.
[0051] Illustratively, the "rendering according to the camera parameter information and the current three-dimensional data model to obtain the current color map, the current depth map and the current normal map under the perspective of the scene image" in step S140 can include: in response to the current three-dimensional data model being a neural radiance field, rendering according to the camera parameter information and the current three-dimensional data model by a neural rendering method to obtain the current color map, the current depth map and the current normal map under the perspective of the scene image; in response to the current three-dimensional data model being a three-dimensional Gaussian body, rendering according to the camera parameter information and the current three-dimensional data model by a Gaussian rasterization method to obtain the current color map, the current depth map and the current normal map under the perspective of the scene image.
[0052] Specifically, for each scene image, in the case that the current three-dimensional data model is a neural radiance field, the current color map, the current depth map and the current normal map under the perspective of the scene image can be obtained by rendering according to the camera parameter information corresponding to the scene image and the current three-dimensional data model by a neural rendering method. The neural rendering method is to sample the current neural radiance field by camera rays, and estimate the rendering information, i.e. color value, depth value and normal value, for each sampling point, so as to realize volume rendering of the neural radiance field and obtain the rendered image under the perspective of the scene image, i.e. the current color map, the current depth map and the current normal map. In the case that the current three-dimensional data model is a three-dimensional Gaussian body, the current color map, the current depth map and the current normal map under the perspective of the scene image can be obtained by rendering according to the camera parameter information corresponding to the scene image and the current three-dimensional data model by a Gaussian rasterization method. The Gaussian rasterization method is to project an ellipse to a camera plane and then rasterize to a pixel plane, and finally estimate the rendering information, i.e. color value, depth value and normal value, for each pixel, so as to rasterize the three-dimensional Gaussian body into a rendered image and realize the three-dimensional Gaussian splatting process.
[0053] S150, training the current three-dimensional data model according to the current color map, the current normal map, the pseudo-normal map and the actual color map of the scene image to obtain a target three-dimensional data model after training.
[0054] Specifically, since the actual color value of the scene image can be directly obtained, the pseudo normal map is obtained by the current rendering map calculation, so that the actual color value and the pseudo normal map are taken as the training label to perform self-supervised training on the current three-dimensional data model, without manual annotation, and the constraint on the depth and the normal is increased in the model training process, thereby improving the training speed and quality of the three-dimensional data model. By dynamically adjusting the model parameters in the current three-dimensional data model, the current three-dimensional data model is automatically iteratively trained, and the current three-dimensional data model after the training is completed is determined as the target three-dimensional data model corresponding to the three-dimensional scene. The target three-dimensional data model refers to the three-dimensional data model finally reconstructed by the three-dimensional scene, such as a neural radiance field or a three-dimensional Gaussian body. Since the target three-dimensional data model contains more model information than the traditional three-dimensional mesh model, the target three-dimensional data model has stronger model expression capability, such as the expression capability of the scene surface normal, the depth, and other geometric properties, so that the target three-dimensional data model can achieve super-fine and realistic reconstruction effect, and the geometric rendering quality is improved.
[0055] The technical scheme of the embodiment of the present disclosure determines the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to the scene image according to the scene image collected for the three-dimensional scene to be reconstructed; performs model initialization according to the sparse point cloud data to obtain an initialized current three-dimensional data model; performs rendering according to the camera parameter information and the current three-dimensional data model to obtain a current color map, a current depth map, and a current normal map under the view angle of the scene image, and determines a pseudo normal map under the view angle of the scene image according to the current depth map; and performs self-supervised training on the current three-dimensional data model according to the current color map, the current normal map, the pseudo normal map, and the actual color map of the scene image, so that a target three-dimensional data model with higher quality can be automatically obtained, and the reconstruction cost is reduced. Moreover, the target three-dimensional data model has stronger scene expression capability than the traditional three-dimensional mesh model, thereby improving the fineness and rendering quality of the scene model.
[0056] On the basis of the above technical scheme, step S120 can include: performing data processing on the scene image to obtain scene image data in a preset image format; performing preprocessing on the scene image data in the preset image format to obtain preprocessed scene image data; and performing sparse reconstruction on the preprocessed scene image data based on a motion recovery structure algorithm and a bundle adjustment algorithm to obtain the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to the scene image.
[0057] The preset image format can be a general image format set in advance. Specifically, since the scene images collected by different image collection devices have different image formats, it is necessary to unify all image formats to obtain scene image data in a general preset image format. For example, different imaging models are used for parameterization based on different image formats, such as using a fisheye imaging model, a panoramic imaging model, a pinhole imaging model, or a multi-camera group model, so as to process scene images in different image formats into a unified camera model data structure, and then obtain scene image data in a preset image format. Since the collected scene images are uniformly processed in format, image collection can be performed using different types of collection devices, reducing the collection cost and improving the adaptability of three-dimensional reconstruction. Through deep learning, the scene image data in the preset image format can be preprocessed, such as denoising, super-resolution, feature extraction and matching, so as to improve the success rate and accuracy of three-dimensional reconstruction, and also support feature extraction and matching between different data formats. By using a motion recovery structure SfM algorithm and a bundle adjustment BA algorithm, the preprocessed scene images are reconstructed into sparse point cloud and camera pose estimation, so as to more accurately obtain the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to each scene image, thereby further improving the three-dimensional reconstruction effect.
[0058] Exemplarily, the three-dimensional scene data further includes scene collection data in a non-image format. The scene collection data in a non-image format can refer to scene data collected by a non-image collection device, such as positioning data collected by a global positioning system GPS.
[0059] Specifically, for the three-dimensional scene to be reconstructed, in addition to collecting scene images, scene collection data in a non-image format can also be collected. In the case that there is scene collection data in a non-image format in the collected three-dimensional scene data, the scene collection data in a non-image format can be added to the unified camera model data structure in the form of point cloud, depth map, control point, etc. as additional collection data, so as to be additionally processed during sparse reconstruction.
[0060] Exemplarily, based on the motion recovery structure algorithm and the bundle adjustment algorithm, the preprocessed scene image data is sparse reconstructed to obtain the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to the scene image, which can include: adding the scene collection data in a non-image format as an algorithm constraint condition to the motion recovery structure algorithm and the bundle adjustment algorithm, sparse reconstructing the preprocessed scene image data to obtain the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to the scene image.
[0061] Specifically, in the case that non-image format scene acquisition data exists in the collected three-dimensional scene data, the non-image format scene acquisition data is introduced as an additional algorithm constraint condition into the motion recovery structure algorithm and the bundle adjustment algorithm for optimization, so that the reconstruction of the sparse point cloud and the estimation of the camera pose can be more accurate, and the scene reconstruction accuracy is further improved.
[0062] On the basis of the above technical solutions, in the case that the three-dimensional scene is a non-closed scene, the non-closed scene needs to be converted into a closed scene so as to be suitable for the reconstruction of the three-dimensional data model. For example, for the neural radiance field, since a ray needs to be emitted from the camera and sampled on the ray in the rendering process, the starting point and the ending point of the sampling need to be defined. For an outdoor large-scale scene, there is no predefined boundary (such as the sky), so the ending point cannot be defined. Therefore, for the non-closed scene (i.e. the unbounded scene), a mapping can be defined to convert the unbounded non-closed scene into a bounded scene, so as to be suitable for the reconstruction of a larger scene range. For example, a mapping is defined to convert the unbounded space into a sphere with a preset value (such as 2) as the radius. The sampling in the unit sphere is not modified, and the sampling outside the unit sphere is shrunk to adapt to the sphere with the preset value as the radius. Therefore, the calculation of all sampling points is to map a point in the space to the inside of the sphere with the preset value as the radius, and then the neural radiance field can be used to reconstruct the three-dimensional data model in the non-closed scene. For the three-dimensional Gaussian body, a spherical point cloud can be added to the outside of the reconstructed sparse point cloud, so as to perform three-dimensional reconstruction on the inside of the spherical surface. Therefore, the three-dimensional Gaussian body can also be used to reconstruct the three-dimensional data model in the non-closed scene.
[0063] FIG. 2 is a flowchart of another three-dimensional scene reconstruction method provided by the embodiment of the present disclosure. The embodiment of the present disclosure optimizes the steps of “determining the pseudo-normal map under the scene image view according to the current depth map” and “training the current three-dimensional data model according to the current color map, the current normal map, the pseudo-normal map and the actual color map of the scene image to obtain the target three-dimensional data model after the training is completed” on the basis of the above disclosed embodiments. The optimization of the two steps can exist separately or simultaneously, and the embodiment of the present disclosure does not limit this. The explanation of the same or corresponding terms as those in the above disclosed embodiments is not repeated here.
[0064] As shown in FIG. 2, the three-dimensional scene reconstruction method specifically includes the following steps:
[0065] S210, acquiring three-dimensional scene data collected for a three-dimensional scene to be reconstructed, the three-dimensional scene data including scene images of the three-dimensional scene.
[0066] S220, determining sparse point cloud data corresponding to the three-dimensional scene and camera parameter information corresponding to the scene images according to the scene images.
[0067] S230. Initialize the model based on sparse point cloud data to obtain the current 3D data model corresponding to the 3D scene.
[0068] S240. Render based on camera parameter information and the current 3D data model to obtain the current color map, current depth map and current normal map from the scene image perspective.
[0069] S250. Determine the spatial location information corresponding to each scene pixel in the scene image based on the current depth map.
[0070] The spatial location information corresponding to each scene pixel can be represented using three-dimensional coordinates (x, y, z) in the world coordinate system. Specifically, based on the current depth value corresponding to each pixel in the current depth map, the two-dimensional coordinates of each scene pixel in the view can be transformed into three-dimensional world space to obtain the spatial location information (x, y, z) corresponding to each scene pixel.
[0071] S260. Based on the spatial location information, determine the horizontal gradient information and vertical gradient information corresponding to each scene pixel.
[0072] The horizontal gradient information of a scene pixel refers to the gradient information in the horizontal direction of the scene pixel. The vertical gradient information refers to the gradient information in the vertical direction of the scene pixel. Specifically, for each scene pixel, the horizontal and vertical gradient information can be determined based on the spatial position information of the surrounding scene pixels. The surrounding scene pixels can refer to scene pixels within a neighborhood of a preset size (e.g., 3×3) centered on the scene pixel. The horizontal gradient information of the scene pixel is determined using the spatial position information of the surrounding scene pixels in the horizontal direction. The vertical gradient information of the scene pixel is determined using the spatial position information of the surrounding scene pixels in the vertical direction. For example, if there are 8 surrounding scene pixels within a 3×3 neighborhood, the surrounding scene pixels in the horizontal direction include: the surrounding scene pixel xyz located in the first row and first column. 00 The surrounding scene pixels xyz located in the 1st row and 3rd column 02 The surrounding scene pixels xyz located in the 2nd row and 1st column 10 The surrounding scene pixels xyz located in the 2nd row and 3rd column 12 The surrounding scene pixels xyz located in the 3rd row and 1st column 20 and the surrounding scene pixels xyz located in the 3rd row and 3rd column 22 The surrounding scene pixels in the vertical direction include: the surrounding scene pixels xyz located in row 1, column 1. 00 The surrounding scene pixels xyz located in the 1st row and 2nd column01 the surrounding scene pixel xyz at the 3rd column of the 1st row 02 the surrounding scene pixel xyz at the 1st column of the 3rd row 20 the surrounding scene pixel xyz at the 2nd column of the 3rd row 21 and the surrounding scene pixel xyz at the 3rd column of the 3rd row 22 For example, the scene pixel xyz 11 The corresponding horizontal gradient information is: -0.125×xyz 00 +0.125×xyz 02 -0.25×xyz 10 +0.25×xyz 12 -0.125×xyz 20 +0.125×xyz 22 The scene pixel xyz 11 The corresponding vertical gradient information is: -0.125×xyz 002 -0.25×xyz 01 -0.125×xyz 02 +0.125×xyz 20 +0.25×xyz 21 +0.125×xyz 22 It should be noted that the horizontal gradient information and the vertical gradient information corresponding to each scene pixel can be represented by a three-dimensional column vector.
[0073] S270, determining the normal corresponding to each scene pixel according to the horizontal gradient information and the vertical gradient information, and normalizing the normal to determine the pseudo-normal map under the scene image view angle.
[0074] Specifically, for each scene pixel, the horizontal gradient information and the vertical gradient information corresponding to the scene pixel can be vector cross-multiplied to obtain the normal corresponding to the scene pixel. The normal can also be represented by a three-dimensional column vector. Based on the maximum modulus length in all normals, the normal corresponding to each scene pixel is normalized, and the normalized normal is taken as the pseudo-normal value. Thus, based on the pseudo-normal values corresponding to all scene pixels, the pseudo-normal map under the scene image view angle is generated. By using the local gradient, the pseudo-normal map under each scene image view angle can be accurately obtained.
[0075] S280, determining the current color loss value according to the current color map and the actual color map of the scene image, and determining the current normal loss value according to the current normal map and the pseudo-normal map.
[0076] The current color loss value is used to represent the difference between the current color map and the actual color map. The current normal loss value is used to represent the difference between the current normal map and the pseudo normal map.
[0077] Specifically, based on the current color map and the actual color map corresponding to each scene image, the absolute difference value and the structural similarity between the current color map and the actual color map can be determined, and the absolute difference value and the structural pixel degree corresponding to each scene image are weighted and summed, and the sum result is taken as the current color loss value. Based on the current normal map and the pseudo normal map corresponding to each scene image, the mean square error between the current normal map and the pseudo normal map can be determined, and the mean square error is taken as the current normal loss value.
[0078] S290, based on the current color loss value and the current normal loss value, determine the current model loss value, and based on the current model loss value, adjust the model parameters in the current three-dimensional data model, until the preset convergence condition is met, and the training is ended, and the target three-dimensional data model after training is obtained.
[0079] Specifically, the current color loss value and the current normal loss value corresponding to each scene image can be added, and the addition result is taken as the current model loss value. The current model loss value is back propagated to the current three-dimensional data model, the model parameters in the current three-dimensional data model are adjusted, and the rendering operation and the parameter adjustment operation are performed based on the adjusted current three-dimensional data model, that is, the operations of steps S240-S290, until the preset convergence condition is met, such as the number of iterations is equal to the preset number, or the model loss value tends to be stable, it is determined that the model training is ended, and the current three-dimensional data model after training is determined as the final reconstructed target three-dimensional data model, so as to realize the full-automatic high-quality reconstruction of three-dimensional scene. For example, when the current three-dimensional data model is a neural radiation field, the iterative training is performed by adjusting the volume density value, color value, depth value and normal value corresponding to each voxel in the neural radiation field. When the current three-dimensional data model is a three-dimensional Gaussian body, the iterative training is performed by adjusting the position information, spherical harmonic function coefficient, transparency, rotation value, scale value and normal value corresponding to each voxel in the three-dimensional Gaussian body.
[0080] Exemplarily, in step S290, "determining the current model loss value based on the current color loss value and the current normal loss value" can include: determining the current smoothing loss value based on the current normal map, the current depth map and the actual color map of the scene image; determining the current model loss value based on the current color loss value, the current normal loss value and the current smoothing loss value.
[0081] Specifically, based on a preset smoothing function, a normal smoothing loss value is determined according to a current normal map and an actual color map of a scene image, a depth smoothing loss value is determined according to a current depth map and the actual color map of the scene image, and a weighted sum of the normal smoothing loss value and the depth smoothing loss value is obtained as a current smoothing loss value. The current color loss value, the current normal loss value and the current smoothing loss value corresponding to each scene image are added to obtain a current model loss value. By considering the smoothing loss, the model training effect can be further improved, thereby further improving the three-dimensional reconstruction quality.
[0082] The technical scheme of the embodiments of the present disclosure determines the spatial position information corresponding to each scene pixel in the scene image according to the current depth map, determines the horizontal gradient information and the vertical gradient information corresponding to each scene pixel according to the spatial position information, determines the normal corresponding to each scene pixel according to the horizontal gradient information and the vertical gradient information, and normalizes the normal, so that the pseudo-normal map under the scene image view can be accurately determined. The current color loss value is determined according to the current color map and the actual color map of the scene image, and the current normal loss value is determined according to the current normal map and the pseudo-normal map; the current model loss value is determined based on the current color loss value and the current normal loss value, and the model is trained based on the current model loss value, so that the trained target three-dimensional data model has stronger scene expression capability, and the fineness and rendering quality of the scene model are further improved.
[0083] FIG. 3 is a flowchart of another three-dimensional scene reconstruction method provided by the embodiments of the present disclosure. The embodiments of the present disclosure describe in detail the process of directly using the target three-dimensional data model for lighting and adding shadows based on the above disclosed embodiments. The explanations of the same or corresponding terms as those in the above disclosed embodiments are not repeated here.
[0084] As shown in FIG. 3, the three-dimensional scene reconstruction method specifically includes the following steps:
[0085] S310, acquiring three-dimensional scene data collected for a three-dimensional scene to be reconstructed, the three-dimensional scene data including scene images of the three-dimensional scene.
[0086] S320, determining sparse point cloud data corresponding to the three-dimensional scene and camera parameter information corresponding to the scene images according to the scene images.
[0087] S330, initializing a model according to the sparse point cloud data to obtain a current three-dimensional data model corresponding to the three-dimensional scene.
[0088] S340, rendering according to the camera parameter information and the current three-dimensional data model to obtain a current color map, a current depth map and a current normal map under the scene image view, and determining a pseudo-normal map under the scene image view according to the current depth map.
[0089] S350, training the current three-dimensional data model according to the current color map, the current normal map, the pseudo normal map and the actual color map of the scene image, and obtaining a target three-dimensional data model after training is completed.
[0090] S360, performing lighting and shadow adding processing according to the target three-dimensional data model, the current position information and the light source type information of the virtual light source and the current parameter information of the virtual camera, and obtaining a virtual scene image with lighting and shadow effects.
[0091] The virtual light source can be a virtual light-emitting device placed in the three-dimensional scene. The current position information of the virtual light source can be set in advance based on business requirements. The light source type information can include a parallel light source or a point light source. The virtual camera can be a virtual camera placed in the three-dimensional scene. The current parameter information of the virtual camera can include the current pose information and the current intrinsic information of the virtual camera. The virtual scene image can be a two-dimensional virtual image at a certain viewing angle in the three-dimensional scene with lighting and shadow.
[0092] Specifically, the user can set the current position information of the virtual light source and the current parameter information of the virtual camera based on lighting requirements and viewing angle imaging requirements. Since the reconstructed target three-dimensional data model has stronger scene expression capability, the target three-dimensional data model can be directly rendered with lighting and shadow adding based on the current position information and the light source type information of the virtual light source and the current parameter information of the virtual camera, so as to quickly obtain a high-quality virtual scene image with lighting and shadow effects without human intervention, thereby improving the rendering quality and speed of the virtual scene image.
[0093] Illustratively, in a virtual live room, the reconstructed target three-dimensional data model can be rendered with lighting and projection, etc. to obtain a virtual scene image with lighting and shadow effects, and the virtual scene image can be directly displayed as a live picture in the virtual live room, or the virtual scene image can be synthesized with the anchor picture as a live background picture to obtain a live picture of the anchor in the virtual three-dimensional scene, thereby realizing the virtual live effect.
[0094] Exemplarily, the step S360 can include: rendering according to the current position information of the virtual light source and the light source type information and the target three-dimensional data model to obtain a first depth map under the virtual light source perspective; rendering according to the current parameter information of the virtual camera and the target three-dimensional data model to obtain a target color map, a second depth map, a target normal map and a target opacity map under the virtual camera perspective; determining a target shadow map according to the first depth map and the second depth map; and determining a virtual scene image with lighting and shadow effects according to the target color map, the target normal map, the target opacity map and the target shadow map.
[0095] Specifically, the target projection mode can be determined according to the light source type information of the virtual light source, for example, if the light source type information is a parallel light source, the target projection mode is determined as an orthogonal projection; if the light source type information is a point light source, the target projection mode is determined as a perspective projection. The target three-dimensional data model is rendered according to the current position information of the virtual light source to obtain a first depth map with the target projection mode under the virtual light source perspective, such as a first depth map under the orthogonal projection or a first depth map under the perspective projection. In the case that the current three-dimensional data model is a neural radiance field, the first depth map with the target projection mode under the virtual light source perspective can be obtained by neural rendering according to the current position information of the virtual light source and the target three-dimensional data model, and the target color map, the second depth map, the target normal map and the target opacity map under the virtual camera perspective can be obtained by neural rendering according to the current parameter information of the virtual camera and the target three-dimensional data model. Alternatively, in the case that the current three-dimensional data model is a three-dimensional Gaussian body, the first depth map with the target projection mode under the virtual light source perspective can be obtained by Gaussian rasterization according to the current position information of the virtual light source and the target three-dimensional data model, and the target color map, the second depth map, the target normal map and the target opacity map under the virtual camera perspective can be obtained by Gaussian rasterization according to the current parameter information of the virtual camera and the target three-dimensional data model. The second depth map under the virtual camera perspective is mapped into the light source position space and compared with the depth value in the first depth map under the virtual light source perspective to generate a target shadow map. The final color map with lighting and added shadow under the virtual camera perspective, i.e., the virtual scene image with lighting and shadow effects, is determined by calculation according to the target color map, the target normal map, the target opacity map and the target shadow map under the virtual camera perspective, thereby improving the view rendering quality and the rendering speed.
[0096] The technical scheme of the embodiment of the present disclosure can quickly obtain a virtual scene image with lighting and shadow effects without human intervention, thereby improving the rendering quality and rendering speed of the virtual scene image, by performing lighting and shadow adding processing according to the target three-dimensional data model, the current position information and the light source type information of the virtual light source, and the current parameter information of the virtual camera.
[0097] FIG. 4 is a structural schematic diagram of a three-dimensional scene reconstruction device provided by an embodiment of the present disclosure. As shown in FIG. 4, the device specifically comprises: a three-dimensional scene data acquisition module 410, a sparse point cloud data determination module 420, a model initialization module 430, a model rendering module 440, and a model training module 450.
[0098] The three-dimensional scene data acquisition module 410 is configured to acquire three-dimensional scene data collected for a three-dimensional scene to be reconstructed, wherein the three-dimensional scene data comprises scene images of the three-dimensional scene. The sparse point cloud data determination module 420 is configured to determine sparse point cloud data corresponding to the three-dimensional scene and camera parameter information corresponding to the scene images according to the scene images. The model initialization module 430 is configured to perform model initialization according to the sparse point cloud data to obtain a current three-dimensional data model corresponding to the three-dimensional scene. The model rendering module 440 is configured to perform rendering according to the camera parameter information and the current three-dimensional data model to obtain a current color map, a current depth map, and a current normal map under the view angle of the scene images, and determine a pseudo normal map under the view angle of the scene images according to the current depth map. The model training module 450 is configured to perform training on the current three-dimensional data model according to the current color map, the current normal map, the pseudo normal map, and an actual color map of the scene images to obtain a target three-dimensional data model after the training is completed.
[0099] The technical scheme provided by the embodiment of the present disclosure can automatically obtain a target three-dimensional data model with higher quality by determining sparse point cloud data corresponding to a three-dimensional scene and camera parameter information corresponding to scene images according to the scene images collected for the three-dimensional scene to be reconstructed, performing model initialization according to the sparse point cloud data to obtain a current three-dimensional data model after initialization, performing rendering according to the camera parameter information and the current three-dimensional data model to obtain a current color map, a current depth map, and a current normal map under the view angle of the scene images, and determining a pseudo normal map under the view angle of the scene images according to the current depth map, thereby reducing the reconstruction cost. Moreover, the target three-dimensional data model has stronger scene expression capability than a traditional three-dimensional mesh model, thereby improving the fineness and rendering quality of the scene model.
[0100] On the basis of the above technical scheme, the sparse point cloud data determination module 420 comprises:
[0101] a data processing unit, configured to perform data processing on the scene image to obtain scene image data in a preset image format;
[0102] a preprocessing unit, configured to perform preprocessing on the scene image data in the preset image format to obtain preprocessed scene image data;
[0103] a sparse reconstruction unit, configured to perform sparse reconstruction on the preprocessed scene image data based on a motion structure from motion algorithm and a bundle adjustment algorithm to obtain sparse point cloud data corresponding to the three-dimensional scene and camera parameter information corresponding to the scene image.
[0104] On the basis of each of the above technical solutions, the three-dimensional scene data further includes scene acquisition data in a non-image format.
[0105] The sparse reconstruction unit is specifically configured to add the scene acquisition data in the non-image format as an algorithm constraint condition to the motion structure from motion algorithm and the bundle adjustment algorithm to perform sparse reconstruction on the preprocessed scene image data to obtain the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to the scene image.
[0106] On the basis of each of the above technical solutions, the current three-dimensional data model includes a neural radiance field or a three-dimensional Gaussian body.
[0107] The model parameters in the neural radiance field include a body density value, a color value, a depth value, and a normal value corresponding to a voxel in the three-dimensional scene.
[0108] The model parameters in the three-dimensional Gaussian body include position information, a spherical harmonic function coefficient, a transparency, a rotation value, a scale value, and a normal value corresponding to a voxel in the three-dimensional scene.
[0109] On the basis of each of the above technical solutions, the model rendering module 440 is specifically configured to:
[0110] in response to the current three-dimensional data model being a neural radiance field, performing rendering according to the camera parameter information and the current three-dimensional data model in a neural rendering manner to obtain a current color map, a current depth map, and a current normal map under the scene image view angle;
[0111] in response to the current three-dimensional data model being a three-dimensional Gaussian body, performing rendering according to the camera parameter information and the current three-dimensional data model in a Gaussian rasterization manner to obtain a current color map, a current depth map, and a current normal map under the scene image view angle.
[0112] On the basis of each of the above technical solutions, the model rendering module 440 includes:
[0113] a spatial position information determination unit configured to determine spatial position information corresponding to each scene pixel in the scene image according to a current depth map;
[0114] a gradient information determination unit configured to determine horizontal gradient information and vertical gradient information corresponding to each scene pixel according to the spatial position information;
[0115] a pseudo normal map determination unit configured to determine a normal corresponding to each scene pixel according to the horizontal gradient information and the vertical gradient information, and normalize the normal to determine a pseudo normal map under a view angle of the scene image.
[0116] In the above technical solutions, the model training module 450 comprises:
[0117] a current normal loss value determination unit configured to determine a current color loss value according to a current color map and an actual color map of the scene image, and determine a current normal loss value according to a current normal map and the pseudo normal map;
[0118] a model training unit configured to determine a current model loss value based on the current color loss value and the current normal loss value, and adjust model parameters in the current three-dimensional data model based on the current model loss value until a preset convergence condition is met, and obtain a target three-dimensional data model after training is completed.
[0119] In the above technical solutions, the model training unit is specifically configured to determine a current smoothness loss value based on the current normal map, the current depth map and the actual color map of the scene image, and determine the current model loss value based on the current color loss value, the current normal loss value and the current smoothness loss value.
[0120] In the above technical solutions, the device further comprises:
[0121] a virtual scene image determination module configured to perform lighting and shadow adding processing according to the target three-dimensional data model, current position information and light source type information of a virtual light source and current parameter information of a virtual camera, and obtain a virtual scene image with lighting and shadow effects.
[0122] In the above technical solutions, the virtual scene image determination module is specifically configured to:
[0123] According to the current position information and the light source type information of the virtual light source and the target three-dimensional data model, rendering is performed to obtain a first depth map under a virtual light source view angle; according to the current parameter information of the virtual camera and the target three-dimensional data model, rendering is performed to obtain a target color map, a second depth map, a target normal map and a target opacity map under a virtual camera view angle; according to the first depth map and the second depth map, a target shadow map is determined; and according to the target color map, the target normal map, the target opacity map and the target shadow map, a virtual scene image with lighting and shadow effects is determined.
[0124] The three-dimensional scene reconstruction apparatus provided by the embodiments of the present disclosure can perform the three-dimensional scene reconstruction method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of performing the method.
[0125] It is worth noting that each unit and module included in the above apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for convenient mutual distinction, and does not limit the protection scope of the embodiments of the present disclosure.
[0126] FIG. 5 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Referring to FIG. 5, a structural schematic diagram of an electronic device (for example, a terminal device or a server in FIG. 5) 500 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include but is not limited to a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle-mounted terminal (for example, a vehicle-mounted navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 5 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0127] As shown in FIG. 5, the electronic device 500 can include a processing apparatus (for example, a central processor, a graphics processor, and the like) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage apparatus 508 to a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing apparatus 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An editing / output (I / O) interface 505 is also connected to the bus 504.
[0128] In general, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 508 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG. 5 shows the electronic device 500 with various devices, it is understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0129] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processing devices 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0130] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0131] The electronic device provided by the embodiments of the present disclosure and the three-dimensional scene reconstruction method provided by the above-mentioned embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above-mentioned embodiments, and the present embodiment has the same beneficial effects as the above-mentioned embodiments.
[0132] The embodiments of the present disclosure provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the three-dimensional scene reconstruction method provided by the above-mentioned embodiments.
[0133] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0134] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0135] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and not be assembled into the electronic device.
[0136] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire three-dimensional scene data collected for a three-dimensional scene to be reconstructed, the three-dimensional scene data including a scene image of the three-dimensional scene; determine sparse point cloud data corresponding to the three-dimensional scene and camera parameter information corresponding to the scene image according to the scene image; perform model initialization according to the sparse point cloud data to obtain a current three-dimensional data model corresponding to the three-dimensional scene; perform rendering according to the camera parameter information and the current three-dimensional data model to obtain a current color map, a current depth map and a current normal map in a view of the scene image, and determine a pseudo normal map in the view of the scene image according to the current depth map; and train the current three-dimensional data model according to the current color map, the current normal map, the pseudo normal map and an actual color map of the scene image to obtain a target three-dimensional data model after training is completed.
[0137] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0138] The flow and block diagrams in the drawings show architectural, functional, and operational representations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0139] The units described in the embodiments of the present disclosure can be implemented by software, or can be implemented by hardware. In some cases, the names of the units do not constitute a limitation on the units themselves.
[0140] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0141] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0142] The embodiments of the present disclosure further provide a computer program product, comprising a computer program which, when executed by a processor, implements the three-dimensional scene reconstruction method provided by the above-described embodiments.
[0143] The computer program product, in the implementation, can be written in one or more programming languages or combinations of the same to implement the computer program code for performing the operations of the present disclosure, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on a user computer, partially on a user computer, as a separate software package, partially on a user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).
[0144] The above description is only preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features with similar functions disclosed in the present disclosure (but not limited to) can be formed.
[0145] Further, while operations are depicted in a particular order, this should not be understood as requiring such order nor the order that is illustrated. Rather, multi-tasking and parallel processing can be advantageous in certain circumstances. Likewise, while a number of specific implementation details have been included in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0146] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for reconstructing a three-dimensional scene, comprising: Acquire 3D scene data of the 3D scene to be reconstructed, wherein the 3D scene data includes scene images of the 3D scene; Based on the scene image, determine the sparse point cloud data corresponding to the 3D scene and the camera parameter information corresponding to the scene image; The model is initialized based on the sparse point cloud data to obtain the current 3D data model corresponding to the 3D scene. Rendering is performed based on the camera parameter information and the current 3D data model to obtain the current color map, current depth map and current normal map under the scene image view, and the pseudo normal map under the scene image view is determined based on the current depth map; Based on the current color map, current normal map, pseudo normal map, and the actual color map of the scene image, the current 3D data model is trained to obtain the target 3D data model after training.
2. The three-dimensional scene reconstruction method according to claim 1, wherein, The step of determining the sparse point cloud data corresponding to the 3D scene and the camera parameter information corresponding to the scene image based on the scene image includes: The scene image is processed to obtain scene image data in a preset image format; Preprocess the scene image data in the preset image format to obtain the preprocessed scene image data; Based on the motion recovery structure algorithm and the bundle adjustment algorithm, sparse reconstruction is performed on the preprocessed scene image data to obtain sparse point cloud data corresponding to the 3D scene and camera parameter information corresponding to the scene image.
3. The three-dimensional scene reconstruction method according to claim 2, wherein, The 3D scene data also includes: scene acquisition data in non-image format; The method based on the structure-of-motion (SOG) algorithm and the bundle adjustment algorithm performs sparse reconstruction on the preprocessed scene image data to obtain sparse point cloud data corresponding to the 3D scene and camera parameter information corresponding to the scene image, including: The non-image format scene acquisition data is added as an algorithm constraint to the motion recovery structure algorithm and the bundle adjustment algorithm. The preprocessed scene image data is sparsely reconstructed to obtain the sparse point cloud data corresponding to the 3D scene and the camera parameter information corresponding to the scene image.
4. The three-dimensional scene reconstruction method according to any one of claims 1-3, wherein, The current three-dimensional data model includes a neural radiation field or a three-dimensional Gaussian body; The model parameters in the neural radiation field include: the volume density value, color value, depth value, and normal value corresponding to the voxels in the three-dimensional scene; The model parameters in the three-dimensional Gaussian volume include: the position information of the voxels in the three-dimensional scene, the coefficients of the spherical harmonic function, the transparency, the rotation value, the scale value, and the normal value.
5. The three-dimensional scene reconstruction method according to any one of claims 1-4, wherein, The step of rendering based on the camera parameter information and the current 3D data model to obtain the current color map, current depth map, and current normal map from the scene image viewpoint includes: In response to the current 3D data model being a neural radiation field, a neural rendering method is used to render the current color map, current depth map, and current normal map from the scene image perspective. In response to the fact that the current 3D data model is a 3D Gaussian volume, the model is rendered using Gaussian rasterization based on the camera parameter information and the current 3D data model to obtain the current color map, current depth map, and current normal map from the perspective of the scene image.
6. The three-dimensional scene reconstruction method according to any one of claims 1-5, wherein, The step of determining the pseudo-normal map from the current depth map viewpoint of the scene image includes: Determine the spatial location information corresponding to each scene pixel in the scene image based on the current depth map; Based on the spatial location information, determine the horizontal gradient information and vertical gradient information corresponding to each scene pixel; Based on the horizontal gradient information and the vertical gradient information, the normal corresponding to each scene pixel is determined, and the normal is normalized to determine the pseudo normal map under the scene image view.
7. The three-dimensional scene reconstruction method according to any one of claims 1-6, wherein, The step of training the current 3D data model based on the current color map, the current normal map, the pseudo normal map, and the actual color map of the scene image to obtain the target 3D data model after training includes: The current color loss value is determined based on the current color map and the actual color map of the scene image, and the current normal loss value is determined based on the current normal map and the pseudo normal map; The current model loss value is determined based on the current color loss value and the current normal loss value. Based on the current model loss value, the model parameters in the current 3D data model are adjusted until the preset convergence condition is met, and the training ends, thus obtaining the target 3D data model after training.
8. The three-dimensional scene reconstruction method according to claim 7, wherein, The process of determining the current model loss value based on the current color loss value and the current normal loss value includes: The current smoothing loss value is determined based on the current normal map, the current depth map, and the actual color map of the scene image; The current model loss value is determined based on the current color loss value, the current normal loss value, and the current smoothing loss value.
9. The three-dimensional scene reconstruction method according to any one of claims 1-8 further includes: Based on the target 3D data model, the current position and type information of the virtual light source, and the current parameter information of the virtual camera, lighting and shadow addition processing is performed to obtain a virtual scene image with lighting and shadow effects.
10. The three-dimensional scene reconstruction method according to claim 9, wherein, The step of adding lighting and shadows based on the target 3D data model, the current position and type information of the virtual light source, and the current parameter information of the virtual camera to obtain a virtual scene image with lighting and shadow effects includes: Rendering is performed based on the current position information and light source type information of the virtual light source and the target 3D data model to obtain the first depth map from the perspective of the virtual light source; Rendering is performed based on the current parameter information of the virtual camera and the target 3D data model to obtain the target color map, second depth map, target normal map and target opacity map from the perspective of the virtual camera; Determine the target shadow map based on the first depth map and the second depth map; Based on the target color map, the target normal map, the target opacity map, and the target shadow map, a virtual scene image with lighting and shadow effects is determined.
11. A three-dimensional scene reconstruction device, comprising: The 3D scene data acquisition module is configured to acquire 3D scene data collected from the 3D scene to be reconstructed, wherein the 3D scene data includes scene images of the 3D scene; The sparse point cloud data determination module is configured to determine the sparse point cloud data corresponding to the three-dimensional scene and the camera parameter information corresponding to the scene image based on the scene image. The model initialization module is configured to initialize the model based on the sparse point cloud data to obtain the current 3D data model corresponding to the 3D scene. The model rendering module is configured to render based on the camera parameter information and the current 3D data model, obtain the current color map, current depth map and current normal map under the scene image view, and determine the pseudo normal map under the scene image view based on the current depth map; The model training module is configured to train the current 3D data model based on the current color map, the current normal map, the pseudo normal map, and the actual color map of the scene image, so as to obtain the target 3D data model after training.
12. An electronic device, comprising: One or more processors; A storage device is configured to store one or more programs, wherein, When the one or more programs are executed by the one or more processors, the one or more processors implement the three-dimensional scene reconstruction method as described in any one of claims 1-10.
13. A storage medium containing computer-executable instructions, wherein, The computer-executable instructions, when executed by a computer processor, are used to perform the three-dimensional scene reconstruction method as described in any one of claims 1-10.
14. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the three-dimensional scene reconstruction method as described in any one of claims 1-10.
Citation Information
Patent Citations
Three-dimensional object modelling fitting & tracking
CN103733227A
Outdoor unbounded scene three-dimensional reconstruction method and system based on neural radiation field
CN116051740A
Nerve radiation field three-dimensional reconstruction method and device based on adaptive mask
CN117934710A
Three-dimensional scene reconstruction method and device, equipment, medium and program product
CN118823234A
Method and device for gigapixel-level light field intelligent reconstruction of large-scale scene
US11908067B1
Cited By
Vehicle-mounted three-dimensional scene rendering method and device, storage medium and program product
CN121661216A
A vehicle-mounted three-dimensional scene rendering method, device, storage medium and program product
CN121661216B
Gaussian neural field dynamic scene reconstruction system based on depth consistency constraint
CN121708189A
Sparse view angle cultural relic three-dimensional reconstruction method, storage medium and computer equipment
CN121767572A