Three-dimensional reconstruction method, apparatus and device, and readable medium
By acquiring scene reconstruction data sets and using neural network models with implicit representations, combining depth and normal information to optimize the loss function, the problems of incomplete reconstruction and high computing resources in monocular camera three-dimensional reconstruction technology are solved, and high-precision indoor scene reconstruction is achieved, suitable for fields such as computer vision and virtual reality.
Patent Information
- Application Number
- CN202311866195.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing three-dimensional reconstruction technology of indoor scenes based on monocular cameras has problems such as incomplete reconstruction, high computing resource consumption, and poor reconstruction accuracy in general applications. Especially in textureless and repeated texture areas, the existing neural implicit representation methods cannot complete high-precision scene-level surface reconstruction.
By obtaining a scene reconstruction data set containing RGB images, prior images, pose information and camera parameters, the three-dimensional model is reconstructed using a neural implicit representation neural network model, combining the combined loss of depth information, normal information and RGB images to optimize network parameters, and using hybrid surface representation and volume rendering technology to reduce noise and error during the reconstruction process.
It improves the accuracy and efficiency of three-dimensional reconstruction of multi-view RGB images, reduces noise during the reconstruction process, realizes high-precision scene surface reconstruction, reduces computing resource requirements, and is suitable for fields such as computer vision, augmented reality and virtual reality.
Smart Images

Figure CN120236000A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a three-dimensional reconstruction method, apparatus, device, and readable medium. Background Art
[0002] In existing three-dimensional reconstruction methods, the scene three-dimensional reconstruction algorithms relying on depth or lidar sensors highly depend on professional devices and have high acquisition costs, which hinder large-scale use by users. Most traditional depth estimation algorithms have many constraints and poor effects in complex scenes; the reconstruction algorithms based on structure from motion (SFM) and multi-view stereo (MVS) can better complete sparse reconstruction, but are prone to generating holes when performing dense reconstruction. In contrast, the three-dimensional reconstruction technology based on a monocular camera presents many advantages such as low cost, fast running speed, and simple structure. Reconstructing an indoor scene from monocular multi-view images is a basic task in computer vision and graphics and plays an important role in various applications such as robots, games, virtual reality, and augmented reality. Summary of the Invention
[0003] The present disclosure provides a three-dimensional reconstruction method, apparatus, device, and readable medium.
[0004] In a first aspect of the present disclosure, a three-dimensional reconstruction method is provided, including:
[0005] Obtaining a set of scene reconstruction data, where the set of scene reconstruction data includes at least one set of sampling point data, and each set of sampling point data includes a two-dimensional RGB image, a prior image, pose information, and camera parameters of a sampling point, and the prior image includes depth information and normal information corresponding to the sampling point;
[0006] Based on the set of scene reconstruction data, reconstructing a three-dimensional model of the scene by using a neural network model with neural implicit representation.
[0007] In a second aspect of the present disclosure, a three-dimensional reconstruction apparatus is provided, including:
[0008] An obtaining module, configured to obtain a set of scene reconstruction data, where the set of scene reconstruction data includes at least one set of sampling point data, and each set of sampling point data includes a two-dimensional RGB image, a prior image, pose information, and camera parameters of a sampling point, and the prior image includes depth information and normal information corresponding to the sampling point;
[0009] A reconstruction module, configured to reconstruct a three-dimensional model of the scene by using a neural network model with neural implicit representation based on the set of scene reconstruction data.
[0010] The third aspect of the present disclosure provides an electronic device, including:
[0011] At least one processor;
[0012] A memory on which one or more programs are stored. When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to the first aspect;
[0013] At least one I / O interface, connected between the processor and the memory, configured to implement information interaction between the processor and the memory.
[0014] The fourth aspect of the present disclosure provides a computer-readable medium on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect is implemented.
[0015] The present disclosure has the following advantages:
[0016] Obtain a set of scene reconstruction data. Among them, the set of scene reconstruction data includes at least one set of sampling point data. Each set of sampling point data includes a two-dimensional RGB image, a prior image, pose information, and camera parameters of a sampling point. The prior image includes depth information and normal information corresponding to the sampling point; based on the set of scene reconstruction data, use a neural network model with neural implicit representation to reconstruct the three-dimensional model of the scene, so that when reconstructing the three-dimensional model of the scene, the pre-obtained prior image can be used to guide the reconstruction process, thereby improving the accuracy of three-dimensional scene reconstruction based on multi-view RGB images and improving the reconstruction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of a three-dimensional reconstruction method provided in an embodiment of the present disclosure;
[0018] Figure 2 It is a schematic diagram of the process of three-dimensional scene reconstruction provided in an embodiment of the present disclosure;
[0019] Figure 3 It is a schematic structural diagram of a three-dimensional reconstruction device provided in an embodiment of the present disclosure;
[0020] Figure 4 It is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following will describe the specific embodiments of the present disclosure in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and illustrating the present disclosure, and are not used to limit the present disclosure.
[0022] As used in this disclosure, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0023] The terms used in this disclosure are only used to describe specific embodiments and are not intended to limit this disclosure. As used in this disclosure, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0024] When the terms "comprising" and / or "consisting of" are used in this disclosure, it specifies the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof.
[0025] Unless otherwise defined, the meanings of all terms (including technical and scientific terms) used in this disclosure are the same as those commonly understood by a person of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning unless this disclosure clearly so defines.
[0026] The inventors found in the process of implementing this disclosure that:
[0027] However, the current monocular camera-based indoor scene 3D reconstruction technology has the following problems in terms of generalization, which limits its large-scale promotion and application:
[0028] (1) Scene reconstruction algorithms based on dense matching (such as COLMAP, OpenMVG): By feature extraction and matching, the scene depth map is restored, and then the scene surface model is fused. For areas containing a large amount of textureless, weakly textured, and repetitive textured areas, such as walls, floors, windows, etc., large areas will be missing during reconstruction, and the results are poor. Moreover, the calculation of generating a dense depth map and the subsequent Poisson surface reconstruction process consumes a large amount of computing resources and takes a long time.
[0029] (2) Based on the success of deep neural network and data-driven (depth-based and TSDF-based) methods, image features are extracted through multi-layer convolution, and voxel fusion is performed according to voxel costs, which can alleviate the surface missing problem in textureless and repetitive textured areas and restore a more complete surface. However, the 3D convolutional layers and Transformer structures it depends on have huge computational overheads and high computational resource requirements, and it is impossible to generate a high-precision surface model when reconstructing indoor scenes.
[0030] (3) The technology based on neural implicit representation for reconstruction can achieve high-quality reconstruction accuracy and restore fine details, becoming a highly potential alternative method after the convolutional-based 3D reconstruction methods. Some methods combine neural implicit representation with volume rendering technology to directly reconstruct through multi-view RGB images. Some methods use multi-layer perceptrons (MLPs) to learn implicit surface functions and optimize the sampling method of volume rendering to improve the reconstruction quality. However, these methods still need to provide additional object masks to reduce the network search space. Subsequent work attaches different volume rendering techniques to the reconstruction, thus eliminating the need for masks and further improving the reconstruction quality. However, due to the limitation of network capacity and the induced smoothing bias of multi-layer perceptrons, scene-level surface reconstruction cannot be completed. In addition, most of the existing 3D reconstruction methods of scenes based on implicit representation directly fit the scene surface function using multi-layer perceptrons, and cannot accurately reconstruct the complete real scene, with poor reconstruction accuracy.
[0031] Based on this, the embodiments of the present disclosure provide a 3D reconstruction method, which can be applied to any electronic device for 3D reconstruction of a scene.
[0032] Figure 1 The following is a schematic flow diagram of the 3D reconstruction method provided by the embodiments of the present disclosure. Refer to Figure 1 , the 3D reconstruction method flow provided by the embodiments of the present disclosure mainly includes the following steps:
[0033] Step 101, obtain a scene reconstruction data set, where the scene reconstruction data set includes at least one set of sampling point data, and each set of sampling point data includes a 2D RGB image, a prior image, pose information, and camera parameters of a sampling point. The prior image includes depth information and normal information corresponding to the sampling point.
[0034] In some embodiments, the obtaining of the scene reconstruction data set includes: collecting image data from at least two sampling points of the scene to obtain at least one set of original sampling data, where the original sampling data includes a 2D RGB image, pose information, and camera parameters of the sampling point; determining the prior image corresponding to the RGB image in each set of the original sampling data; and determining the scene reconstruction data set according to each set of the original sampling data and the prior image corresponding to the original sampling data.
[0035] In an exemplary embodiment, the camera parameters include an internal camera matrix and an external camera matrix.
[0036] In an exemplary embodiment, a two-dimensional RGB image of a scene is collected from multiple perspectives by an image acquisition device, and the camera parameters and pose information of the image acquisition device are recorded while collecting. Among them, the image acquisition device can be any device capable of collecting two-dimensional RGB images, such as an image acquisition application (APP) of a mobile phone, etc.
[0037] In an exemplary embodiment, format processing and alignment processing are performed on the two-dimensional RGB images, camera parameters, and pose information collected at each sampling point to obtain each group of sampling point data. Among them, format processing is mainly to perform equidistant frame extraction on the video and convert it into a unified RGB image format in the case of collecting a video, and align the RGB image and the corresponding pose information.
[0038] In an exemplary embodiment, the resolutions of the two-dimensional RGB images in each group of original sampling data are unified. Usually, the two-dimensional RGB images in each group of original sampling images are downsampled according to a predefined resolution to obtain the two-dimensional RGB images of each sampling point in the scene reconstruction data set. The data volume is reduced through downsampling, and the processing efficiency is improved.
[0039] In an exemplary embodiment, the camera parameters in the scene reconstruction data set are the camera parameters obtained after normalization processing, including the normalized camera internal parameter matrix and the normalized camera external parameter matrix.
[0040] The final scene reconstruction data set is obtained by preprocessing the original sampling data, so that the scene reconstruction data set for three-dimensional reconstruction can meet the accuracy and efficiency requirements of subsequent processing, laying a foundation for improving the reconstruction accuracy and efficiency.
[0041] In some embodiments, the scene reconstruction data set includes sampling point data of different perspectives of the complete scene area. Collecting a large number of images of the complete scene area from different perspectives helps to improve the network reconstruction accuracy.
[0042] In an exemplary embodiment, in the actual collection process, visual drastic changes are avoided, such as changes in light intensity, too fast movement or rotation of the shooting perspective, etc. Collecting a large number of images of the complete scene area from different perspectives helps to improve the network reconstruction accuracy. It is sufficient that there is no drastic jitter or too low clarity during shooting.
[0043] In an exemplary embodiment, the multi-perspective videos and camera parameters of the indoor scene collected by the staff are obtained and stored locally. Then, using ffmpeg to decompose, decode, and convert the multi-perspective videos into picture frames, and using a data preprocessing script to unify the camera poses and image formats to obtain the data for generating the prior image.
[0044] In some embodiments, determining the prior image corresponding to the RGB image in each group of the original sampling data includes: obtaining the prior image corresponding to the RGB image in each group of the original sampling data by using a pre-trained model. By generating a prior image and corresponding it to the RGB image, accurate depth information and normal information of the target scene are provided for the subsequent 3D reconstruction process to guide the optimization of the subsequent scene 3D reconstruction network. The introduction of the prior image greatly reduces the errors and noise generated in scene reconstruction, speeds up the network convergence speed, and improves the fineness and integrity of scene surface reconstruction.
[0045] In an exemplary embodiment, the prior image can be an image that includes both depth information and normal information, or a combination of a depth image including depth information and a normal image including normal information.
[0046] In an exemplary embodiment, in the prior image generation stage, a pre-trained model is used to predict the prior image corresponding to each RGB image. The prior image includes depth information and normal information, and the resolution and naming storage format of the prior image are adjusted to be aligned with the RGB image to facilitate the loading of the subsequent scene 3D reconstruction module. By reasonably selecting the pre-trained model for prior estimation and the resolution of the prior image, the accuracy and stability of scene reconstruction can be improved, the scene reconstruction time can be reduced, and thus the accuracy of the scene 3D reconstruction result.
[0047] In some embodiments, the pre-trained model includes an end-to-end 3D scene annotation network.
[0048] In an exemplary embodiment, the end-to-end 3D scene annotation network includes at least one of the following: Omnidata deep learning model, HiMODE deep learning model, DINOv2 deep learning model, NLL-AngMF deep learning model. In addition, other pre-trained models may also be used, and the protection scope of the embodiments of the present disclosure is not limited thereto.
[0049] Step 102, based on the scene reconstruction data set, use a neural network model with neural implicit representation to reconstruct the 3D model of the scene.
[0050] In some embodiments, reconstructing a three-dimensional model of the scene by using a neural network model with neural implicit representation based on the scene reconstruction data set includes: in one reconstruction process, constructing the surface of the three-dimensional model of the scene and performing color rendering on the surface by using the neural network model with neural implicit representation based on the scene reconstruction data set to obtain the three-dimensional model of the scene; determining a joint loss based on the scene reconstruction data set and the three-dimensional model of the scene, optimizing the parameters of the neural network model with neural implicit representation by using the joint loss, and performing the reconstruction process again until the number of times of performing the reconstruction process reaches a set number of times.
[0051] It should be noted that the neural network model with neural implicit representation learns the scene topology and texture, etc. from the supervision of RGB images and prior images to jointly generate a three-dimensional model of the scene. Among them, the voxel grid features of the three-dimensional model are optimized by using the joint loss of depth information, normal information, and RGB images, and the surface topology of the scene is learned by aggregating features. The color value of the sampling point is calculated by volume rendering the sampling points on the camera ray, and the loss is calculated in combination with the pixel color value of the actually sampled RGB image to optimize the neural network radiation field and learn the scene texture information.
[0052] In some embodiments, constructing the surface of the three-dimensional model of the scene and performing color rendering on the surface by using the neural network model with neural implicit representation based on the scene reconstruction data set to obtain the three-dimensional model of the scene includes: constructing the voxel feature corresponding to the sampling point in the surface grid of the three-dimensional model of the scene by using the neural network model with neural implicit representation based on the scene reconstruction data set; predicting the depth information and normal information corresponding to the sampling point by using the neural network model with neural implicit representation based on the voxel feature; predicting the pixel color value corresponding to the sampling point in the surface grid of the three-dimensional model by using the voxel feature, the predicted depth information and normal information of the sampling point, the pose information of the sampling point in the scene reconstruction data set, and the camera parameters.
[0053] In some embodiments, determining the joint loss based on the scene reconstruction data set and the three-dimensional model of the scene includes: determining a color loss according to the pixel color values of the RGB images of the sampling points in the scene reconstruction data set and the pixel color values of the corresponding sampling points in the surface mesh of the three-dimensional model; determining an eikonal loss according to the smoothness of the surface mesh of the three-dimensional model; determining a depth loss according to the depth information of the sampling points in the scene reconstruction data set and the predicted depth information of the sampling points; determining a normal loss according to the normal information of the sampling points in the scene reconstruction data set and the predicted normal information of the sampling points; and determining the joint loss according to the color loss, the eikonal loss, the depth loss, and the normal loss.
[0054] In some embodiments, the neural network model of the neural implicit representation includes an MLP.
[0055] In an exemplary embodiment, the process of training a three-dimensional scene model using an MLP mainly includes the following steps:
[0056] Step 1, represent the scene geometry as a signed distance function (SDF). The signed distance function is a continuous function f. For a given three-dimensional (3D) point, the SDF value returned by the signed distance function represents the distance from the point to the nearest surface:
[0057] f:R 3 →Rx→S=SDF(x) (1)
[0058] where x is a 3D point and S is the corresponding SDF value.
[0059] By mixing the use of an MLP and a single-resolution feature grid to learn the SDF function, the learnable parameters are represented as θ, and the scene surface is represented as the set of horizontal planes with an SDF value of 0:
[0060] S={x|f_θ (x)=0} (2)
[0061] The hybrid representation parameterization method directly stores trainable parameters in each cell of a discrete voxel G H ×R W ×R D with a resolution of R θ .
[0062] Step 2, use the voxel features of the sampling points as input, predict the SDF values of the grid through an MLP, and calculate the normal information and depth information according to voxel rendering.
[0063] SDF value: Use the feature aggregation MLP network f θ. Query the SDF value of any point x from the dense SDF grid using trilinear interpolation operation (denoted as interp):
[0064] S = f θ (λ(x), interp(x, Φ θ )) (3)
[0065] where λ(x) is the position encoding of point x, and Φ θ represents the feature vector in the voxel grid.
[0066] Optimize the hybrid surface representation using differentiable volume rendering. More specifically, to render a pixel, project a ray r from the camera center o, pass through the pixel along its view direction v, and sample M points along the ray where represents the ray passing through point x, represents the opacity of the current ray sampling point, and predict their SDF values. Convert the SDF value S to density value δ for volume rendering according to the following formula:
[0067]
[0068] where β is a learnable parameter.
[0069] Normal information: As a globally consistent geometric constraint to improve the reconstruction quality. Scene reconstruction uses the predicted anomaly prior as additional supervision. Specifically, use the following volume accumulation process to obtain the surface normal at the sampled viewpoints:
[0070]
[0071] where and are the opacity and α-channel value of sample point j along ray r respectively, is the three-dimensional normal of sample point j.
[0072] Depth information: The depth D(r) of the surface intersecting with the current ray is calculated by formula (6):
[0073]
[0074] Step 3, use the SDF value, normal information, depth information, view direction, and sampling point position encoding as inputs to predict the grid surface color value.
[0075] Color: In addition to 3D geometry, the neural network model can also predict color values, and the supervised model is optimized under the reconstruction loss. Therefore, the color function c is defined as:
[0076] c = f θ(x, m, n, z) (7)
[0077] m is the position encoding in the camera direction, n is the three-dimensional normal, where the three-dimensional normal n represents the SDF analytical gradient of the position x. The feature vector z is the output of the second linear layer of the f θ network. After that, the color value is converted into volume density for rendering.
[0078] The color can be generated by the following formula:
[0079]
[0080]
[0081]
[0082] where C(r) is the color of the sampling point, and are the transmittance and α-channel value of the sample point i along the ray r respectively, is the distance between adjacent sample points.
[0083] Step 4, calculate the joint loss according to the input data and the generated normal, depth, and color values corresponding to the viewing angle. The joint loss is defined as:
[0084] Color loss:
[0085]
[0086] where is the observed pixel color.
[0087] Eikonal loss: used to regularize the surface smoothness and make the surface SDF regularized orthogonal loss, defined as:
[0088]
[0089] where M is the set of uniformly sampled points around the surface.
[0090] Depth loss:
[0091]
[0092] where is the observed pixel depth.
[0093] Normal loss:
[0094]
[0095] where is the observed pixel normal.
[0096] The combined loss is defined as:
[0097]
[0098] where λ color , λ eik , λ normal and λ depth are hyperparameters, set to 1, 0.05, 0.05, 0.1.
[0099] Step Five, calculate the gradient through the combined loss to complete the update of the network parameters of the entire neural network model.
[0100] Step Six, when the training reaches the specified number of rounds, use the MC (Marching Cubes) algorithm to extract a high-precision three-dimensional surface model from the SDF field, calculate the normal direction of the surface mesh vertices in the reverse direction, sample the rays in this direction, input them into the neural radiance field to calculate the color and opacity, perform voxel rendering to determine the vertex colors, and finally output a scene mesh model with texture colors.
[0101] Among them, the MC algorithm is a surface reconstruction algorithm for three-dimensional voxel data, which converts discrete voxel data into a smooth three-dimensional mesh model.
[0102] SDF, the signed distance function, represents the distance and direction information from the object surface to a certain point.
[0103] Trilinear Interpolation (interp), trilinear interpolation is a method of linear interpolation on the tensor product grid of three-dimensional discrete sampling data.
[0104] In an exemplary embodiment, the three-dimensional model of the reconstructed scene is a triangular surface mesh file with vertex colors, which can be directly viewed in MeshLab.
[0105] In a specific embodiment, as Figure 2 shown is a schematic diagram of the process of three-dimensional scene reconstruction, mainly including:
[0106] Step 201, collect multi-view video and pose data of the scene;
[0107] Step 202, split the collected video and extract camera parameters;
[0108] Step 203, extract image frames from the video and adjust the resolution, and synchronously adjust the camera parameters;
[0109] Step 204, use a pre-trained model to generate prior images of depth and normal;
[0110] Step 205: Align the RGB image, camera parameters, and prior image to obtain a set of scene reconstruction data;
[0111] Step 206: Load the set of scene reconstruction data;
[0112] Step 207: Initialize the SDF feature voxel grid to identify the scene surface;
[0113] Step 208: Generate SDF, depth information, and normal information according to the hybrid representation;
[0114] Step 209: Render the pixel color according to the color representation network;
[0115] Step 210: Determine whether the number of training times has been reached. If not, execute Step 211; if so, execute Step 213;
[0116] Step 211: Calculate the joint loss;
[0117] Step 212: Optimize the network parameters according to the joint loss, and then go back to execute Steps 207 and 209;
[0118] Step 213: Output the rendered reconstructed 3D model.
[0119] The 3D reconstruction method provided by the embodiments of the present disclosure obtains a set of scene reconstruction data, where the set of scene reconstruction data includes at least one set of sampling point data, and each set of sampling point data includes a two-dimensional RGB image, a prior image, pose information, and camera parameters of a sampling point. The prior image includes the depth information and normal information corresponding to the sampling point. Based on the set of scene reconstruction data, a neural implicit representation neural network model is used to reconstruct the 3D model of the scene, so that the prior image obtained in advance can be used to guide the reconstruction process when reconstructing the 3D model of the scene, thereby improving the accuracy of 3D scene reconstruction based on multi-view RGB images and improving the reconstruction efficiency.
[0120] By adding prior-guided training and hybrid surface representation, the noise in the reconstruction process is reduced and the network capacity is improved.
[0121] The 3D reconstruction method provided by the embodiments of the present disclosure is applicable to a variety of computer vision tasks, such as running on Linux and Windows desktops or command-line operating systems that can run Python. For hardware, the method provided by the embodiments of the present disclosure has certain requirements for the computing power of the graphics card. Higher graphics card computing power can speed up the training process, but the computing power only affects the training time and does not affect the accuracy of the final inference result. Therefore, when using the method provided by the embodiments of the present disclosure, a graphics card with higher computing power can improve the efficiency, but does not affect the quality of the final inference result.
[0122] The 3D reconstruction method provided by the embodiments of the present disclosure does not require the introduction of new sensors, solves the problems of incomplete, noisy, and long training time when using multi-view images for scene reconstruction, improves the accuracy and quality of the reconstruction results, and provides an effective technical route for indoor scene reconstruction. It provides an efficient and accurate solution for the application of scene model generation technologies in the fields of computer vision, augmented reality, virtual reality, autonomous driving, etc., and has important application value.
[0123] The embodiments of the present disclosure adopt a hybrid neural surface representation, including steps such as initializing the construction of a single-resolution SDF grid coordinate system and predicting the surface SDF, which helps to reduce noise, improve network capacity, ultimately realize the generation of a more accurate 3D scene model, speed up the 3D reconstruction of the scene, and improve the real-time performance and efficiency of the system.
[0124] Calculate the gradient through the joint loss of pixels, depth, and normal vectors, optimize the network parameters, thereby completing the generation of the 3D scene model and texture, improving the stability and robustness of the reconstruction, reducing the errors and uncertainties in the reconstruction process, and enhancing the quality of the reconstruction results.
[0125] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of the present disclosure; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are all within the protection scope of the present disclosure.
[0126] An embodiment of the present disclosure provides a 3D reconstruction device. For the specific implementation of this device, reference can be made to the relevant description in the method embodiment, and details will not be repeated here. Figure 3 The following is a schematic structural diagram of the device, which mainly includes:
[0127] An acquisition module 301, configured to acquire a set of scene reconstruction data, where the set of scene reconstruction data includes at least one set of sampling point data, and each set of sampling point data includes a two-dimensional RGB image, a prior image, pose information, and camera parameters of a sampling point, and the prior image includes depth information and normal vector information corresponding to the sampling point;
[0128] A reconstruction module 302, configured to reconstruct a 3D model of the scene based on the set of scene reconstruction data by using a neural implicit representation neural network model.
[0129] The functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the method embodiments. For the specific implementation and technical effects, reference can be made to the description in the above method embodiments. For the sake of brevity, details will not be repeated here.
[0130] It should be noted that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present disclosure, units that are not closely related to solving the technical problems proposed by the present disclosure are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0131] Referring to Figure 4 , the embodiment of the present disclosure provides an electronic device, which includes:
[0132] At least one processor 401;
[0133] A memory 402, on which at least one program is stored. When the at least one program is executed by the at least one processor, the at least one processor implements the above method;
[0134] At least one I / O interface 403, connected between the processor and the memory, configured to implement information interaction between the processor and the memory.
[0135] Among them, the processor 401 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory 402 is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) 403 is connected between the processor 401 and the memory 402 and can implement information interaction between the processor 401 and the memory 402, including but not limited to a data bus (Bus), etc.
[0136] In some embodiments, the processor 401, the memory 402, and the I / O interface 403 are interconnected through a bus and then connected to other components of the computing device.
[0137] This embodiment also provides a computer-readable medium, on which a computer program is stored. When the program is executed by the processor, the method provided in this embodiment is implemented. To avoid repeated description, the specific steps of this method are not elaborated here.
[0138] Those of ordinary skill in the art can understand that all or some of the steps in the method, and the functional modules / units in the system and device described above can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cartridges, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0139] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0140] Those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments but not other features, the combination of the features of different embodiments means that it is within the scope of this embodiment and forms different embodiments.
[0141] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present disclosure. However, the present disclosure is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present disclosure, and these modifications and improvements are also regarded as the protection scope of the present disclosure.
Claims
1. A three-dimensional reconstruction method, characterized in that, Including: Obtain a set of scene reconstruction data, where the set of scene reconstruction data includes at least one set of sampled point data, and each set of the sampled point data includes a two-dimensional RGB image, a prior image, pose information, and camera parameters of a sampled point. The prior image includes depth information and normal information corresponding to the sampled point; Based on the set of scene reconstruction data, use a neural network model with neural implicit representation to reconstruct a three-dimensional model of the scene.
2. The method according to claim 1, wherein The step of using a neural network model with neural implicit representation to reconstruct a three-dimensional model of the scene based on the set of scene reconstruction data includes: In one reconstruction process, based on the set of scene reconstruction data, use a neural network model with neural implicit representation to construct the surface of the three-dimensional model of the scene and perform color rendering on the surface to obtain the three-dimensional model of the scene; Based on the set of scene reconstruction data and the three-dimensional model of the scene, determine a joint loss, use the joint loss to optimize the parameters of the neural network model with neural implicit representation, and perform the reconstruction process again until the number of times of performing the reconstruction process reaches a set number of times.
3. The method according to claim 2, characterized in that The step of using a neural network model with neural implicit representation to construct the surface of the three-dimensional model of the scene and perform color rendering on the surface to obtain the three-dimensional model of the scene includes: Based on the set of scene reconstruction data, use a neural network model with neural implicit representation to construct voxel features corresponding to the sampled points in the surface mesh of the three-dimensional model of the scene; Based on the voxel features, use the neural network model with neural implicit representation to predict the depth information and normal information corresponding to the sampled points; Use the voxel features, the predicted depth information and normal information of the sampled points, the pose information of the sampled points in the set of scene reconstruction data, and the camera parameters to predict the pixel color values corresponding to the sampled points in the surface mesh of the three-dimensional model.
4. The method according to claim 3, wherein The step of determining a joint loss based on the set of scene reconstruction data and the three-dimensional model of the scene includes: Determine a color loss according to the pixel color values of the RGB images of the sampled points in the set of scene reconstruction data and the pixel color values corresponding to the sampled points in the surface mesh of the three-dimensional model; Determine an eikonal loss according to the smoothness of the surface mesh of the three-dimensional model; Determine a depth loss according to the depth information of the sampled points in the set of scene reconstruction data and the predicted depth information of the sampled points; Determine a normal loss according to the normal information of the sampled points in the set of scene reconstruction data and the predicted normal information of the sampled points; Determine the joint loss according to the color loss, the eikonal loss, the depth loss, and the normal loss.
5. The method according to claim 1, characterized in that, The step of obtaining the set of scene reconstruction data includes: Collect image data from at least two sampled points of a scene to obtain at least one set of original sampled data. The original sampled data includes the two-dimensional RGB image, pose information, and camera parameters of the sampled points; Determine the prior image corresponding to the RGB image in each set of the original sampled data; Determine the set of scene reconstruction data according to the original sampling data of each group and the prior image corresponding to the original sampling data.
6. The method according to claim 5, characterized in that, The determination of the prior image corresponding to the RGB image in the original sampling data of each group includes: Obtain the prior image corresponding to the RGB image in the original sampling data of each group by using a pre-trained model.
7. The method according to claim 6, characterized in that, The pre-trained model includes an end-to-end three-dimensional scene annotation network.
8. A three-dimensional reconstruction device, characterized in that, It includes: An acquisition module for acquiring a set of scene reconstruction data, wherein the set of scene reconstruction data includes at least one group of sampling point data, and each group of sampling point data includes a two-dimensional RGB image, a prior image, pose information, and camera parameters of a sampling point, and the prior image includes depth information and normal information corresponding to the sampling point; A reconstruction module for reconstructing a three-dimensional model of the scene based on the set of scene reconstruction data by using a neural network model with neural implicit representation.
9. An electronic device, characterized in that, It includes: At least one processor; A memory having stored thereon at least one program, which when executed by the at least one processor causes the at least one processor to implement the method according to any one of claims 1-7; At least one I / O interface connected between the processor and the memory and configured to implement information interaction between the processor and the memory.
10. A computer-readable medium having stored thereon a computer program, which when executed by a processor implements the method according to any one of claims 1-7.