A Structure-Aware 3D Scene Reconstruction Method and Device
Through the sparse point cloud generation method combined with the Unet architecture and diffusion model, the shortcomings of the three-dimensional scene reconstruction method in the existing technology in capturing geometric details are solved, high-precision three-dimensional scene reconstruction is achieved, and the geometric accuracy and global consistency of sparse point clouds are improved.
Patent Information
- Application Number
- CN202510377177.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-28
AI Technical Summary
When existing three-dimensional scene reconstruction methods deal with large-scale and complex scenes, it is difficult to accurately capture geometric details in sparse texture areas, and based on Gaussian and GeoGaussian methods, sufficient geometric details cannot be captured in high-precision areas, resulting in the loss of local geometric details.
A sparse point cloud generation model is used to build a sparse point cloud generation model based on the Unet architecture. The camera's internal and external parameters are recovered through SIFT and ICP algorithms, sparse point cloud data is generated, and a rough three-dimensional structural grid is constructed. The Gaussian field is initialized through random sampling, and the Gaussian primitive parameters are optimized in combination with multiple loss functions to generate high-precision three-dimensional scenes.
It significantly improves the geometric accuracy and global consistency of sparse point clouds, and improves the spatial consistency and geometric accuracy of reconstruction results.
Smart Images

Figure CN119888133B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional scene reconstruction, and particularly to a structure-aware three-dimensional scene reconstruction method and apparatus. Background Art
[0002] In recent years, three-dimensional scene reconstruction technology has achieved rapid development in the fields of computer vision and graphics, and has shown broad application potential in the fields of virtual reality, augmented reality, film production, game development, architectural design, etc. In order to achieve a more realistic three-dimensional scene rendering effect, it is necessary to reconstruct a three-dimensional scene model with a high-precision geometric structure. However, scene reconstruction does face some significant challenges in terms of precise geometric structure, and the existing three-dimensional scene reconstruction methods have the following disadvantages: 1) Scene reconstruction methods based on neural radiance fields: When dealing with large-scale and complex scenes, the geometric detail expression ability for texture-sparse regions is limited, and for irregular geometric shapes, the reconstruction ability of the decoder may be restricted, making it difficult for the model to accurately capture the geometric relationships in these regions; 2) Scene reconstruction methods based on Gaussian: In some regions with high-precision requirements, insufficient geometric details can be captured, and fewer details are processed in the initial stage, resulting in some subtle geometric and texture details not being accurately captured. 3) GeoGaussian method: Its initialization depends on the smoothness assumption of the local point cloud region, which may lead to the loss of local geometric details when dealing with irregular surfaces.
[0003] Therefore, a structure-aware three-dimensional scene reconstruction method and apparatus are developed to solve the above problems. Summary of the Invention
[0004] The present invention proposes a structure-aware three-dimensional scene reconstruction method and apparatus to solve the problem that it is difficult to accurately capture geometric details in the prior art.
[0005] The present invention achieves the above object through the following technical solutions:
[0006] A structure-aware three-dimensional scene reconstruction method of the present invention includes:
[0007] Obtaining information, where the information includes a multi-view RGB image frame sequence of the scene;
[0008] Constructing a sparse point cloud generation model according to a diffusion model based on the Unet architecture;
[0009] Inputting the information into the sparse point cloud generation model to generate the internal and external camera parameters and sparse point cloud data of the RGB image frame corresponding to the view;
[0010] Constructing a rough three-dimensional structure grid according to the sparse point cloud data;
[0011] Randomly sample the rough three-dimensional structural grid, and initialize the Gaussian basis elements according to the random sampling results to obtain an initialized Gaussian field;
[0012] Randomly select a frame of RGB image from the information as the current image frame, project the initialized Gaussian field onto the imaging plane of the camera according to the current image frame and the internal and external camera parameters, generate an RGB color image rendered by the external camera parameters of the current image frame, calculate the loss between the current image frame and the RGB color image, and iteratively optimize the parameters of the Gaussian basis elements of all RGB image frames in the information in turn, and combine all the optimized Gaussian basis elements to obtain a reconstructed three-dimensional scene.
[0013] Specifically, input the information into the sparse point cloud generation model to generate the internal and external camera parameters and sparse point cloud data of the RGB image frames at the corresponding viewpoints, including:
[0014] Based on the geometric constraint relationship between the RGB image frames of multiple viewpoints of the scene, and based on the SIFT (Scale Invariant Feature Transform) feature point matching algorithm and the ICP (Iterative Closest Point) pose optimization algorithm, recover the internal and external camera parameters of each viewpoint;
[0015] Extract the multi-scale features of each RGB image frame based on the encoder of UNet to obtain multi-viewpoint features;
[0016] Map the multi-viewpoint features to three-dimensional space based on the diffusion model and the UNet decoder to generate the sparse point cloud data.
[0017] Specifically, construct a rough three-dimensional structural grid according to the sparse point cloud data, including:
[0018] Perform spatial analysis on the sparse point cloud data, extract the geometric position relationships and spatial distribution characteristics of each point in the sparse point cloud data to obtain the overall framework of the scene;
[0019] Use the triangulation algorithm or the surface fitting algorithm to transform the overall framework of the scene into a grid structure composed of vertices, edges and faces, and construct a rough three-dimensional structural grid.
[0020] Specifically, randomly sample the rough three-dimensional structural grid, and initialize the Gaussian basis elements according to the random sampling results to obtain an initialized Gaussian field, including:
[0021] Reconstruct the three-dimensional surface of the entire object in the scene according to the rough three-dimensional structural grid;
[0022] Randomly sample 70% of the triangular faces among all the facets on the three-dimensional surface of the entire object in the said scenario, and randomly initialize the weights at the three vertices of each triangular face;
[0023] Interpolate the coordinates of the sampling points inside the triangular face using the initialized weights for the said vertices;
[0024] Interpolate the normal vectors of the said vertices using the initialized weights to obtain the normal vectors of the sampling points;
[0025] Randomly offset the coordinates of the sampling points along the direction of the said normal vectors to obtain the central positions of the Gaussian basis elements;
[0026] Take the distance between the central position of the said Gaussian basis element and the nearest Gaussian basis element as the scaling value, and randomly initialize the rotation value, opacity, and spherical harmonic function of this Gaussian basis element according to the said scaling value to obtain the initialized Gaussian field.
[0027] Specifically, reconstruct the three-dimensional surface of the entire object in the said scenario based on the rough three-dimensional structure grid, including:
[0028] Divide the said rough three-dimensional grid into voxel grids with a fixed resolution;
[0029] Process each voxel cube in the said voxel grid one by one. According to the comparison results between the attribute values of the 8 vertices of each said voxel cube and the isosurface of the object, mark each vertex of each said voxel cube as internal or external to form an 8-bit binary number. The said binary number is used to look up the triangulation table, and the triangulation table predefines the number of triangles corresponding to each number and the connection method of its vertices;
[0030] Use the linear interpolation method to calculate the intersection positions of the surface of the entire object and the edges of the voxel cube. These intersection points will be used as the vertices of the triangles. Repeat this process to process all the cubes in the voxel grid one by one to generate local triangles;
[0031] Stitch all the said local triangles to form a complete mesh model to generate the three-dimensional surface of the entire object in the said scenario.
[0032] Specifically, randomly sample the said rough three-dimensional structure grid, and initialize the Gaussian field for the Gaussian basis elements according to the random sampling results to obtain the initialized Gaussian basis elements, including:
[0033]
[0034]
[0035]
[0036]
[0037] They are respectively the three vertices on each sampled triangular face Weights initialized randomly respectively, is the said sampled point, is the said normal vector, d is the randomly offset distance, is the central position of the Gaussian basis element, and the distance to the Gaussian basis element with the closest distance is used as the scaling value , where The central position of the Gaussian basis element with the closest distance, the initial rotation value of each Gaussian basis element is set to a quaternion , the spherical harmonic function h is initialized to a tensor of dimensions, where N is the number of Gaussian basis elements, the spherical harmonic function order d = 3, the value of the tensor at the position (N, 3, 0) is 0.5, and the rest are 0, the opacity of each Gaussian basis element is all initialized to 0.1.
[0038] Specifically, project the initialized Gaussian field onto the imaging plane of the camera according to the said current image frame and the internal and external camera parameters to generate the RGB color image rendered by the external camera parameters of the said current image frame, including:
[0039] Calculate the central position of the Gaussian basis element of the said current image frame according to the weights of the three said vertices on the sampled triangular face and the offset distance of the random offset;
[0040] By calculating the changes in the rotation and area scaling of the triangular faces corresponding to the mesh model of the said current image frame and the mesh model of the initialized scene, and attaching them to the attributes of the Gaussian surface element under the current scene geometry, obtain the position, scaling and rotation of the said Gaussian basis element of the current scene;
[0041] Render the color image, depth map and normal map of the said current image frame according to the position, scaling and rotation of the said Gaussian basis element of the said current image frame and the opacity.
[0042] Specifically, the Gaussian basis element of the current scene can be expressed as:
[0043]
[0044] where is the final position of the Gaussian basis element, is the final scaling of the Gaussian basis element, is the final rotation of the Gaussian basis element.
[0045] Specifically, calculate the loss between the current image frame and the RGB color image, and iteratively optimize the parameters of the Gaussian basis elements of all RGB image frames in the information according to the loss, including:
[0046] Perform loss, mean squared error loss and perceptual similarity loss between the color image of the current image frame and the RGB image rendered under the current frame's extrinsic camera parameters, Gaussian basis element scaling loss and normal consistency loss
[0047] where the
[0048]
[0049] loss is: where H, W, and C are the height, width, and number of channels of the image respectively, is the pixel value of the rendered image,
[0050] is:
[0051]
[0052] is the pixel value of the rendered image, is the pixel value of the image captured by the camera;
[0053] is:
[0054]
[0055] where represents the feature map of the i-th layer of the image I in the VGG deep network, N is the number of feature layers for calculating the loss, is the rendered image, is the image captured by the camera;
[0056] is:
[0057]
[0058] where is the maximum value of the Gaussian basis element scaling, is the minimum value of the Gaussian basis element scaling, is the specified maximum scaling value, is the maximum value of the ratio of the specified maximum scaling value to the minimum scaling value;
[0059] is:
[0060]
[0061] where N is the normal map derived from the rendered depth map, is the normal of the Gaussian basis element within the camera frustum, is the weight of the i-th Gaussian basis element.
[0062] The present invention also provides a structure-aware three-dimensional scene reconstruction device, including:
[0063] An acquisition module, which is used to acquire information, and the information includes a multi-view RGB image frame sequence of the scene;
[0064] A first construction module, which is used to construct a sparse point cloud generation model according to a diffusion model based on the Unet architecture;
[0065] A generation module, which is used to input the information into the sparse point cloud generation model to generate the internal and external camera parameters and sparse point cloud data of the RGB image frame corresponding to the view;
[0066] A second construction module, which is used to construct a rough three-dimensional structure grid according to the sparse point cloud data;
[0067] An initialization module, which is used to randomly sample the rough three-dimensional structure grid and perform Gaussian field initialization on the Gaussian basis elements according to the random sampling results to obtain an initialized Gaussian field;
[0068] An optimization module, which is used to randomly select a frame of RGB image frame from the information as the current image frame, project the initialized Gaussian field onto the imaging plane of the camera according to the current image frame and the internal and external camera parameters to generate an RGB color image rendered by the external camera parameters of the current image frame, calculate the loss between the current image frame and the RGB color image, and iteratively optimize the parameters of the Gaussian basis elements of all RGB image frames in the information in sequence, and combine all the optimized Gaussian basis elements to obtain a reconstructed three-dimensional scene.
[0069] The beneficial effects of the present invention are as follows:
[0070] A structure-aware three-dimensional scene reconstruction method and device proposed by the present invention solve the problem of being unable to accurately capture fine geometric and texture details, significantly improve the geometric accuracy and global consistency of the sparse point cloud through the diffusion model, and enhance the spatial consistency of the reconstruction result by constructing a rough three-dimensional structure grid. Description of the Drawings
[0071] Figure 1 This is the flowchart of a method for structure-aware 3D scene reconstruction in this application. Detailed implementation manners
[0072] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Components of the embodiments of the present invention usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0073] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0074] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0075] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", "left", "right", etc. is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship when the product of the present invention is in its usual placement, or the orientation or positional relationship commonly understood by those skilled in the art. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.
[0076] In addition, the terms "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0077] In the description of the present invention, it should also be noted that, unless otherwise clearly specified and limited, terms such as "arrangement", "connection", etc. should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0078] The following will specifically describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0079] As shown Figure 1 in the figure, a structure-aware three-dimensional scene reconstruction method includes:
[0080] S1: Obtain information, where the information includes a multi-view RGB image frame sequence of the scene;
[0081] S2: Construct a sparse point cloud generation model based on a diffusion model with a Unet architecture;
[0082] S3: Input the information into the sparse point cloud generation model to generate the internal and external camera parameters and sparse point cloud data of the RGB image frame corresponding to the viewing angle;
[0083] S4: Construct a rough three-dimensional structure mesh based on the sparse point cloud data;
[0084] S5: Randomly sample the rough three-dimensional structure mesh, and perform Gaussian field initialization on the Gaussian basis elements according to the random sampling results to obtain an initialized Gaussian field;
[0085] S6: Randomly select a frame of RGB image frame from the information as the current image frame, project the initialized Gaussian field onto the imaging plane of the camera according to the current image frame and the internal and external camera parameters, generate an RGB color image rendered by the external camera parameters of the current image frame, calculate the loss between the current image frame and the RGB color image, and sequentially iterate and optimize the parameters of the Gaussian basis elements of all RGB image frames in the information. Combine all the optimized Gaussian basis elements to obtain the reconstructed three-dimensional scene.
[0086] Through the above reconstructed three-dimensional scene, RGB image frames of a new viewing angle can be rendered. For example, at the provided new viewing angle, by using the known internal and external camera parameters, project the centers of the optimized Gaussian basis elements onto the camera imaging plane, calculate the two-dimensional position coordinates of each Gaussian basis element in the pixel space, and the shape characteristics of the Gaussian basis element are determined by its final scaling parameter and rotation parameter. During the new viewing angle rendering process, each Gaussian basis element is represented as a two-dimensional Gaussian distribution on the imaging plane, and the shape of this distribution is jointly determined by the scaling, rotation of the three-dimensional Gaussian basis element and the camera perspective relationship. By weighted accumulation of the influence regions of the Gaussian basis elements for each pixel, a preliminary image at the new viewing angle is generated, where the weight is related to the opacity, distribution characteristics, and depth information of the Gaussian basis element, and the calculation formula is as follows:
[0087]
[0088] is the weight of the i-th Gaussian basis element; is the opacity, which is related to the depth of the primitive. The transparency of the Gaussian kernel is usually higher for a farther distance, and its contribution to the reconstruction is smaller; the degree of influence of the Gaussian kernel on the final reconstructed model, is the mean of the pixel depths covered by the Gaussian kernel; is the mean of the Gaussian kernel distribution; is the standard deviation of the Gaussian distribution, which controls the width of the Gaussian kernel;
[0089] Spherical Harmonics (SH) is introduced to model the color information of the Gaussian primitives. The color distribution and light interaction of the Gaussian primitives are effectively represented by low-order polynomials, and the final color value of each pixel is generated through the continuous mixing of color and opacity.
[0090]
[0091] is the color distribution on the sphere based on i; L, l are the orders of the spherical harmonics, is the spherical harmonic coefficient, which represents the distribution of color at different orders and azimuth angles; is the spherical harmonic basis function, which represents the directional change of color.
[0092] The color of the final rendering is calculated through the following formula:
[0093]
[0094] is the opacity of the i-th Gaussian primitive, is the color of the i-th Gaussian primitive, is the color of the final rendering at the j-th point in the scene.
[0095] In some embodiments, the information is input into the sparse point cloud generation model to generate the internal and external camera parameters and sparse point cloud data of the RGB image frame corresponding to the perspective, including:
[0096] Through the geometric constraint relationship between the RGB image frames of multiple perspectives of the scene, based on the SIFT (Scale Invariant Feature Transform) feature point matching algorithm and the ICP (Iterative Closest Point) pose optimization algorithm, the internal and external camera parameters of each perspective are restored;
[0097] Based on the encoder of UNet, the multi-scale features of each RGB image frame are extracted to obtain multi-perspective features;
[0098] Based on the diffusion model and the UNet decoder, the multi-perspective features are mapped into three-dimensional space to generate the sparse point cloud data.
[0099] In some embodiments, constructing a rough three-dimensional structure grid based on the sparse point cloud data includes:
[0100] Performing a spatial analysis on the sparse point cloud data to extract the geometric position relationships and spatial distribution characteristics of each point in the sparse point cloud data, obtaining the overall framework of the scene;
[0101] Using a triangulation algorithm or a surface fitting algorithm, transforming the overall framework of the scene into a grid structure composed of vertices, edges, and faces, and constructing a rough three-dimensional structure grid.
[0102] In some embodiments, randomly sampling the rough three-dimensional structure grid and initializing a Gaussian basis element with a Gaussian field according to the random sampling result to obtain an initialized Gaussian basis element, including:
[0103] Reconstructing the three-dimensional surface of the entire object in the scene according to the rough three-dimensional structure grid;
[0104] Randomly sampling 70% of the triangular faces among all the face elements on the three-dimensional surface of the entire object in the scene, and randomly initializing weights at the three vertices of each triangular face;
[0105] Interpolating the vertices using the initialized weights to obtain the coordinates of the sampled points inside the triangular faces;
[0106] Interpolating the normal vectors of the vertices using the initialized weights to obtain the normal vectors of the sampled points;
[0107] Randomly offsetting the coordinates of the sampled points along the direction of the normal vector to obtain the central position of the Gaussian basis element;
[0108] Taking the distance between the central position of the Gaussian basis element and the nearest Gaussian basis element as a scaling value, and randomly initializing the rotation value, opacity, and spherical harmonic function of this Gaussian basis element according to the scaling value to obtain an initialized Gaussian field.
[0109] In some embodiments, reconstructing the three-dimensional surface of the entire object in the scene according to the rough three-dimensional structure grid includes:
[0110] Dividing the rough three-dimensional grid into voxel grids with a fixed resolution;
[0111] Processing each voxel cube in the voxel grid one by one, and marking each vertex of each voxel cube as internal or external according to the comparison result between the attribute values of the 8 vertices of each voxel cube and the isosurface of the object, forming an 8-bit binary number, and the binary number is used to look up a triangulation table, and the triangulation table predefines the number of triangles corresponding to each number and the connection method of its vertices;
[0112] The intersection positions of the entire object surface and the edges of the voxel cube are calculated using linear interpolation, and these intersections will serve as the vertices of the triangles. This process is repeated to process each cube in all voxel grids one by one to generate local triangles;
[0113] All the local triangles are stitched together to form a complete mesh model, generating the three-dimensional surface of the entire object in the scene.
[0114] In some embodiments, random sampling is performed on the rough three-dimensional structure mesh, and Gaussian field initialization is performed on the Gaussian basis elements according to the random sampling results to obtain the initialized Gaussian basis elements, including:
[0115]
[0116]
[0117]
[0118]
[0119] The three vertices on each sampled triangular face respectively The weights randomly initialized respectively, For the sampled point, For the normal vector, d is the random offset distance, Is the center position of the Gaussian basis element, and the distance to the Gaussian basis element with the closest distance is used as the scaling value where The center position of the Gaussian basis element with the closest distance, the initial rotation value of each Gaussian basis element Is set to a quaternion The spherical harmonic function h is initialized to A tensor of dimension, where N is the number of Gaussian basis elements, the spherical harmonic function order d = 3, the value of the tensor at the position (N, 3, 0) is 0.5, and the rest are 0, and the opacity of each Gaussian basis element Is initialized to 0.1.
[0120] In some embodiments, the initialized Gaussian field is projected onto the imaging plane of the camera according to the current image frame and the internal and external camera parameters to generate an RGB color image rendered with the external camera parameters of the current image frame, including:
[0121] Calculate the center position of the Gaussian basis element of the current image frame according to the weights of the three vertices on the sampled triangular face and the offset distance of the random offset;
[0122] By calculating the changes in the rotation and area scaling of the triangular faces corresponding to the mesh model of the current image frame and the mesh model of the initialized scene, and attaching them to the attributes of the Gaussian surface elements under the current scene geometry, the positions, scalings, and rotations of the Gaussian basis elements of the current scene are obtained;
[0123] According to the positions, scalings, and rotations of the Gaussian basis elements of the current image frame and the opacity, a color image, a depth map, and a normal map of the current image frame are rendered.
[0124] In some embodiments, the Gaussian basis elements of the current scene are represented as:
[0125]
[0126] where is the final position of the Gaussian basis element, is the final scaling of the Gaussian basis element, is the final rotation of the Gaussian basis element.
[0127] In some embodiments, the loss between the current image frame and the RGB color image is calculated, and according to the loss, the parameters of the Gaussian basis elements of all RGB image frames in the information are iteratively optimized in sequence, including:
[0128] The color image of the current image frame is compared with the RGB image rendered under the extrinsic parameters of the current frame camera to calculate a loss, mean squared error loss a perceptual similarity loss a Gaussian basis element scaling loss and a normal consistency loss where:
[0129] The loss is:
[0130]
[0131] where are the height, width, and number of channels of the image respectively, is the pixel value of the rendered image, is the pixel value of the image captured by the camera;
[0132] is:
[0133]
[0134] is the pixel value of the rendered image, is the pixel value of the image captured by the camera;
[0135] is:
[0136]
[0137] wherein represents the feature map of the i-th layer of the image I in the VGG deep network, and N is the number of feature layers for calculating the loss, is the rendered picture, is the picture taken by the camera;
[0138] is:
[0139]
[0140] wherein is the maximum value of the scaling of the Gaussian basis element, is the minimum value of the scaling of the Gaussian basis element, is the specified maximum scaling value, is the maximum value of the ratio of the specified maximum scaling value to the minimum scaling value;
[0141] is:
[0142]
[0143] where N is the normal map derived from the rendered depth map, is the normal of the Gaussian basis element within the camera frustum, is the weight of the i-th Gaussian basis element. This loss can make the Gaussian achieve a smoother surface geometry-consistent direction. By continuously optimizing the parameter weights u, v, offset distance d, scaling parameter s, rotation parameter q, spherical harmonic function h, and opacity , the final scene model can be obtained after the training ends.
[0144] The total loss of the training is:
[0145] +
[0146] wherein, the picture rendered from the scene and the RGB picture of the i-th frame are used to calculate the loss, is the mean squared error loss, is the perceptual similarity loss, is the Gaussian basis element scaling loss, is the normal consistency loss, are all constants.
[0147] Based on the RGB image frame of the said information, the corresponding internal and external camera parameters, the initialization parameters of the Gaussian basis elements, the randomly initialized weights of the sampling points, and the random offset distance of the sampling points along the normal vector direction, a training strategy for adding random data perturbations includes:
[0148] Randomly select an RGB image frame and the corresponding internal and external camera parameters from the said information.
[0149] Obtain the current frame scene rendering picture according to the Gaussian field and the internal and external camera parameters.
[0150] Optimize the parameters of the Gaussian field model by minimizing the RGB pixel loss between the current frame and the human body rendering picture of the current frame.
[0151] In some instances, looping the above steps 30,000 times can output a scene model with better quality.
[0152] An embodiment of the present invention also provides a structure-aware three-dimensional scene reconstruction device, including:
[0153] An acquisition module, which is used to acquire information, and the information includes a multi-view RGB image frame sequence of the scene;
[0154] A first construction module, which is used to construct a sparse point cloud generation model according to a diffusion model based on the Unet architecture;
[0155] A generation module, which is used to input the information into the sparse point cloud generation model to generate the internal and external camera parameters and sparse point cloud data corresponding to the RGB image frame in the corresponding view;
[0156] A second construction module, which is used to construct a rough three-dimensional structure grid according to the sparse point cloud data;
[0157] An initialization module, which is used to randomly sample the rough three-dimensional structure grid and perform Gaussian field initialization on the Gaussian basis elements according to the random sampling results to obtain the initialized Gaussian basis elements;
[0158] An optimization module, which is used to randomly select a frame of RGB image from the information as the current image frame, project the initialized Gaussian basis elements of the current RGB image frame onto the imaging plane of the camera according to the current image frame and the internal and external camera parameters, generate the color image, depth map and normal map of the current image frame, use the color image, depth map and normal map of the RGB image frame of the first frame of the scene in the information as a reference, calculate the loss between the color image, depth map and normal map of the current image frame and the color image, depth map and normal map of the RGB image frame of the first frame of the scene, and iteratively optimize the parameters of the initialized Gaussian basis elements of all RGB image frames in the information according to the loss to obtain a reconstructed three-dimensional scene.
[0159] The advantages of the present invention compared with the prior art are as follows:
[0160] 1. Introduce a diffusion model and a UNet architecture for sparse point cloud generation.
[0161] By combining the generation characteristics of gradually denoising of the diffusion model and the multi-scale feature extraction ability of the UNet network, the present invention realizes high-precision reconstruction from multi-view images to three-dimensional sparse point clouds. The diffusion model gradually restores the scene structure during the noise optimization process, significantly improving the geometric accuracy and global consistency of the sparse point clouds.
[0162] 2. Initialization of Gaussian basis elements guided by a rough grid model.
[0163] The present invention constructs a rough three-dimensional structure grid, randomly samples on the triangular facets of the grid, and uses the vertex normal vectors of the facets for interpolation and offset to initialize the center, scaling and rotation parameters of the Gaussian basis elements. This method ensures the rationality of the initial position of the Gaussian basis elements and improves the spatial consistency of the reconstruction results.
[0164] 3. Optimization of Gaussian basis elements under multiple loss constraints.
[0165] The present invention introduces multiple loss functions for color maps, depth maps and normal maps, and combines normal consistency loss and scaling constraint loss to iteratively optimize the position, scaling, rotation and opacity of Gaussian basis elements. Through the joint constraints of these geometric and color losses, the surface smoothness and geometric accuracy of Gaussian field reconstruction are ensured.
[0166] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A structure-aware three-dimensional scene reconstruction method, characterized in that: include: Acquiring information, the information comprising a multi-view RGB image frame sequence of a scene; Construct a sparse point cloud generation model based on the diffusion model based on the Unet architecture; Inputting the information into the sparse point cloud generation model to generate camera internal and external parameters and sparse point cloud data of the RGB image frame under the corresponding viewing angle; constructing a rough three-dimensional structure grid according to the sparse point cloud data; Randomly sampling the rough three-dimensional structure grid, and performing Gaussian field initialization on the Gaussian basis element according to the random sampling result to obtain an initialized Gaussian field; A frame of RGB image frame is randomly selected from the information as the current image frame, the initialized Gaussian field is projected onto the imaging plane of the camera according to the current image frame and the camera internal and external parameters, an RGB color image rendered by the camera external parameters of the current image frame is generated, the loss between the current image frame and the RGB color image is calculated, the parameters of the Gaussian basis elements of all RGB image frames in the information are iteratively optimized in sequence according to the loss, all the optimized Gaussian basis elements are combined to obtain a reconstructed three-dimensional scene; The process of randomly sampling the rough three-dimensional structure grid and initializing the Gaussian field of the Gaussian basis element according to the random sampling result to obtain the initialized Gaussian field includes: Reconstructing the three-dimensional surface of the entire object in the scene according to the rough three-dimensional structure grid; Randomly sample 70% of the triangular faces of all face elements in the three-dimensional surface of the entire object in the scene, and randomly initialize weights of the three vertices on each triangular face; Using the initialized weights, the vertices are interpolated to obtain the coordinates of the sampling points within the triangle surface; Interpolate the normal of the vertex using the initialized weight to obtain the normal vector of the sampling point; Randomly shifting the coordinates of the sampling points along the direction of the normal vector to obtain the center position of the Gaussian primitive; The distance between the center position of the Gaussian primitive and the nearest Gaussian primitive is used as a scaling value, and the rotation value, opacity and spherical harmonic function of the Gaussian primitive are randomly initialized according to the scaling value to obtain an initialized Gaussian field.
2. The structure-aware 3D scene reconstruction method according to claim 1, characterized in that: The information is input into the sparse point cloud generation model to generate camera internal and external parameters and sparse point cloud data of the RGB image frame under the corresponding viewing angle, including: Through the geometric constraint relationship between the RGB image frames of the scene's multiple perspectives, based on the feature point matching algorithm and the pose optimization algorithm, the camera's internal and external parameters of each perspective are restored; Extracting multi-scale features of each of the RGB image frames based on a UNet encoder to obtain multi-view features; The multi-view features are mapped to a three-dimensional space based on a diffusion model and a UNet decoder to generate the sparse point cloud data.
3. The structure-aware 3D scene reconstruction method according to claim 1, characterized in that: Constructing a rough three-dimensional structure grid according to the sparse point cloud data, including: Performing spatial analysis on the sparse point cloud data, extracting the geometric position relationship and spatial distribution characteristics of each point in the sparse point cloud data, and obtaining the overall framework of the scene; A triangulation algorithm or a surface fitting algorithm is used to transform the overall framework of the scene into a grid structure consisting of vertices, edges and faces, thereby constructing a rough three-dimensional structure grid.
4. The structure-aware 3D scene reconstruction method according to claim 1, characterized in that: Reconstructing the three-dimensional surface of the entire object in the scene according to the rough three-dimensional structure grid, comprising: segmenting the coarse three-dimensional grid into a voxel grid of fixed resolution; Processing each voxel cube in the voxel grid one by one, marking each vertex of each voxel cube as internal or external according to the comparison result of the attribute values of the eight vertices of each voxel cube with the object isosurface, forming an 8-bit binary number, wherein the binary number is used to look up a triangulation table, wherein the triangulation table predefines the number of triangles corresponding to each number and the connection mode of their vertices; The intersection positions of the entire object surface and the edges of the voxel cube are calculated using a linear interpolation method. These intersection points will be used as the vertices of the triangles. This process is repeated to process the cubes in all voxel grids one by one to generate local triangles. All the local triangles are stitched together to form a complete mesh model, generating a three-dimensional surface of the entire object in the scene.
5. The structure-aware 3D scene reconstruction method according to claim 4, characterized in that: The rough three-dimensional structure grid is randomly sampled, and Gaussian field initialization is performed on the Gaussian basis element according to the random sampling result to obtain the initialized Gaussian basis element, including: , , , , The three vertices on each sampling triangle face The weights are randomly initialized, is the sampling point, is the normal vector, d is the random offset distance, is the center position of the Gaussian primitive, and uses the distance to the nearest Gaussian primitive as the scaling value ,in The center position of the nearest Gaussian primitive and the initial rotation value of each Gaussian primitive Set to quaternion , the spherical harmonic function h is initialized to dimensional tensor, where N is the number of Gaussian basis elements, the spherical harmonic function order d=3, and the tensor is ( ) position has a value of 0.5, and the rest are 0. The opacity of each Gaussian primitive are initialized to 0.
1.
6. The structure-aware 3D scene reconstruction method according to claim 5, characterized in that: The method projects the initialized Gaussian field onto an imaging plane of a camera according to the current image frame and the camera extrinsic parameters to generate an RGB color image rendered by the camera extrinsic parameters of the current image frame, including: Calculate the center position of the Gaussian primitive of the current image frame according to the weights of the three vertices on the sampling triangle and the offset distance of the random offset; The position, scaling and rotation of the Gaussian basis element of the current scene are obtained by calculating the rotation and scaling changes of the triangular face corresponding to the mesh model of the current image frame and the mesh model of the initialization scene, and adding them to the properties of the Gaussian basis element under the geometric form of the current scene; The color image, the depth map and the normal map of the current image frame are rendered according to the position, scaling and rotation of the Gaussian primitives and the opacity of the current image frame.
7. The structure-aware 3D scene reconstruction method according to claim 6, characterized in that: The Gaussian primitive representation of the current scene is: , , , in is the final position of the Gaussian element, is the final scaling of the Gaussian basis, is the final rotation of the Gaussian primitive, , Respectively represent the changes in rotation and area scaling of the triangular faces corresponding to the mesh model of the current scene and the mesh model of the initialization scene.
8. The structure-aware 3D scene reconstruction method according to claim 7, characterized in that: Calculating the loss between the current image frame and the RGB color image, and iteratively optimizing the parameters of the Gaussian primitives of all RGB image frames in the information according to the loss, including: Compare the color image of the current image frame with the RGB image rendered under the camera external parameters of the current frame Loss, mean square error loss , Perceptual Similarity Loss , Gaussian scaling loss and normal consistent loss ,in: The loss is: , Where H, W, and C are the height, width, and number of channels of the image respectively. is the pixel value of the rendered image, The pixel value of the picture taken by the camera; for: , is the pixel value of the rendered image, The pixel value of the picture taken by the camera; for: , in represents the feature map of the i-th layer of image I in the VGG deep network, N is the number of feature layers for calculating the loss, To render the image, Take pictures for the camera; for: , in is the maximum value of the Gaussian basis element scaling, is the minimum value of the Gaussian basis element scaling, To specify the maximum scaling value, The maximum value of the ratio of the maximum scaling value to the minimum scaling value is specified; for: , Where N is the normal map derived from the depth map, is the normal of the Gaussian primitive inside the camera frustum, is the weight of the i-th Gaussian basis.
9. A structure-aware three-dimensional scene reconstruction device, characterized in that: include: An acquisition module, the acquisition module is used to acquire information, the information including a multi-view RGB image frame sequence of a scene; A first building module, wherein the first building module is used to build a sparse point cloud generation model according to a diffusion model based on a Unet architecture; A generation module, the generation module is used to input the information into the sparse point cloud generation model to generate camera internal and external parameters and sparse point cloud data of the RGB image frame under the corresponding viewing angle; A second construction module, the second construction module is used to construct a rough three-dimensional structure grid according to the sparse point cloud data; An initialization module, wherein the initialization module is used to randomly sample the rough three-dimensional structure grid, and initialize the Gaussian field of the Gaussian basis element according to the random sampling result to obtain an initialized Gaussian field; An optimization module, the optimization module is used to randomly select a frame of RGB image frames from the information as the current image frame, project the initialized Gaussian field to the imaging plane of the camera according to the current image frame and the camera internal and external parameters, generate an RGB color image rendered by the camera external parameters of the current image frame, calculate the loss between the current image frame and the RGB color image, iteratively optimize the parameters of the Gaussian basis units of all RGB image frames in the information according to the loss, and combine all the optimized Gaussian basis units to obtain a reconstructed three-dimensional scene; The process of randomly sampling the rough three-dimensional structure grid and initializing the Gaussian field of the Gaussian basis element according to the random sampling result to obtain the initialized Gaussian field includes: Reconstructing the three-dimensional surface of the entire object in the scene according to the rough three-dimensional structure grid; Randomly sample 70% of the triangular faces of all face elements in the three-dimensional surface of the entire object in the scene, and randomly initialize weights of the three vertices on each triangular face; Using the initialized weights, the vertices are interpolated to obtain the coordinates of the sampling points within the triangle surface; Interpolate the normal of the vertex using the initialized weight to obtain the normal vector of the sampling point; Randomly shifting the coordinates of the sampling points along the direction of the normal vector to obtain the center position of the Gaussian primitive; The distance between the center position of the Gaussian primitive and the nearest Gaussian primitive is used as a scaling value, and the rotation value, opacity and spherical harmonic function of the Gaussian primitive are randomly initialized according to the scaling value to obtain an initialized Gaussian field.
Citation Information
Patent Citations
Target driving movement grabbing control method used in multi-scene environment
CN119427349A
Planar mesh reconstruction using images from multiple camera poses
WO2024238237A1