3D Mesh Reconstruction Method and Device Enhanced by Multi-Dimensional Features and Tri-Plane Representations
By designing a multi-dimensional feature extraction and fusion module and a high-resolution three-plane representation building module in single-image-driven three-dimensional grid generation, the problems of insufficient accuracy of the generated grid geometry and surface blur or artifact are solved, and high-precision three-dimensional grid generation is achieved.
Patent Information
- Application Number
- CN202510104129.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In the three-dimensional grid generation driven by single image, the multi-dimensional information in the image is not sufficiently extracted, resulting in insufficient accuracy of the geometric structure of the generated grid and blurring or artifacts on the surface.
A multi-dimensional feature extraction and fusion module is designed to obtain information of multiple dimensions through a single image, including spatial information, global information and local information. At the same time, a high-resolution three-plane representation building block is designed to build a three-plane representation with strong expression capabilities through dual-stream Transformer, convolutional layer and pixel rearrangement operations.
By integrating multi-dimensional feature extraction and fusion modules and high-resolution three-plane representation building modules, a high-precision three-dimensional mesh is generated to reduce blur and artifacts on the mesh surface and improve the accuracy of geometric structure.
Smart Images

Figure CN119540497B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional mesh reconstruction, and in particular relates to a three-dimensional mesh reconstruction method and device with enhanced multi-dimensional features and three-plane representation. Background Art
[0002] In the field of computer graphics and computer vision, 3D reconstruction is one of the fundamental problems, which aims to build a corresponding 3D model from input data such as images or text. The key technologies for 3D object modeling are mainly divided into traditional methods and deep learning-based methods. Traditional methods include using stereo vision, structured light and other technologies to reconstruct 3D meshes from multiple images, or manually created by professionals using tools such as Maya. However, traditional methods have defects such as limited reconstruction accuracy, high learning cost, and low production efficiency. Although physical scanning technology can obtain 3D data of real objects with high precision through professional equipment, its equipment is expensive and has poor portability, which limits its widespread application.
[0003] In recent years, the rapid development of deep learning has promoted 3D reconstruction into a new paradigm, and a series of excellent 3D generation models have emerged, which can generate matching 3D models based on multiple modal conditions. Images and text are the two main driving conditions. Compared with text, images can provide richer and more accurate 3D model information and provide more detailed guidance for the generation of 3D models. At the same time, single-image-driven 3D mesh generation technology is more in line with current industry needs. For example, in application scenarios such as game modeling, modelers need to construct 3D meshes based on original paintings. Single-image-driven 3D mesh generation technology has achieved good results. For example, the LRM (Large Reconstruction Model) model uses the DINO model (a deep learning model for self-supervised visual learning) to extract image features through the input of a single image, and combines the Transformer-based three-plane representation construction network to construct a three-plane representation from image features. Finally, the Marching Cubes algorithm (MC, isosurface extraction algorithm) is used to construct a 3D mesh from the three-plane representation. This method provides a network design paradigm for single-image driven 3D mesh construction solutions. Based on this work, InstantMesh (a tool for quickly generating 3D meshes based on a single image) introduces a multi-view image diffusion model to obtain more 3D priors, and uses FlexiCubes (an open source project for high-quality isosurface representation of gradient-optimized meshes) to obtain and display 3D meshes from three-plane representations. During training, supervision is performed on the 3D mesh to complete 3D mesh reconstruction with higher accuracy.
[0004] However, most of the work in this field usually considers the global information of the driving image and ignores the local information, which makes it impossible to reconstruct sufficient three-dimensional prior knowledge; in addition, since the three-plane representation construction network is usually based on the Transformer architecture, the computational storage complexity and the resolution of the three-plane representation are in a quadratic relationship, and a trade-off needs to be made between the efficiency of the network architecture and the expressiveness of the three-plane representation. These two main reasons lead to the insufficient accuracy of the generated mesh geometry and the presence of blur or artifacts on the surface, that is, the three-dimensional mesh is not accurate enough. Summary of the invention
[0005] In view of the problems existing in the prior art field, the purpose of the present invention is to provide a three-dimensional mesh reconstruction method and device with enhanced multi-dimensional features and three-plane representation, which specifically includes: designing a multi-dimensional feature extraction and fusion module to obtain information of multiple dimensions from a single image, which contains spatial information, global information and local information. At the same time, a high-resolution three-plane representation construction module is also designed. The calculation and storage complexity of this module is linearly related to the resolution of the three-plane representation, and it can efficiently construct a three-plane representation with strong expression ability. By integrating the above two parts of technology, the present invention can generate high-precision three-dimensional meshes based on a single image, effectively reduce the blur and artifacts on the mesh surface, and improve the accuracy of its geometric structure. The generated high-precision three-dimensional mesh can be widely used in the modeling process of the game and film and television industries, simplifying the traditional modeling steps, and improving work efficiency, thereby helping the industry to achieve cost reduction and efficiency improvement.
[0006] To achieve the above-mentioned object of the invention, an embodiment provides a three-dimensional mesh reconstruction method with enhanced multi-dimensional features and three-plane representation, comprising the following steps:
[0007] Acquisition of three-dimensional prior knowledge: After generating a multi-view image sequence based on a single image, a multi-dimensional feature extraction and fusion module is used to extract multi-dimensional features of the image sequence, specifically including: allowing the modified DINO model to accept camera parameter input, and considering spatial information when extracting image features in combination with the image sequence, to obtain image global features containing global information and spatial information, and also encoding based on the image sequence to obtain image local features containing local information, and splicing the image global features and image local features to obtain multi-dimensional features as prior knowledge;
[0008] Three-plane representation construction: Use a high-resolution three-plane representation construction module consisting of a two-stream Transformer, a convolutional layer, and a pixel rearrangement operation to fuse the multi-dimensional features and the initial three-plane features to organize the initial three-plane features into a high-resolution three-plane representation;
[0009] Texture and mesh generation: Use the 3D mesh construction module to generate a 3D mesh based on the three-plane representation. Use the color decoder to generate the color texture of the spatial points based on the three-plane representation decoding. Combine the 3D mesh and color texture to get a textured 3D mesh.
[0010] Preferably, generating a multi-view image sequence based on a single image includes:
[0011] A pre-trained multi-view image diffusion model is used to generate a multi-view image sequence based on a single image, where the multi-view image diffusion model includes a Zero123++ model.
[0012] Preferably, the modified DINO model accepts camera parameter input, and simultaneously considers spatial information when extracting image features in combination with the image sequence, so as to obtain image global features containing global information and spatial information, including:
[0013] A camera condition injection module is added to the DINO model to obtain a modified DINO model. The camera condition injection module obtains a scaling vector and an offset vector based on camera parameters. The scaling vector and the offset vector are used to scale and offset the features extracted based on the image sequence in the DINO model, thereby injecting spatial information into it, so that the obtained global features of the image contain spatial information and global information.
[0014] Preferably, encoding is performed based on the image sequence to obtain local features of the image containing local information, and the global features of the image and the local features of the image are spliced and fused to obtain multi-dimensional features as prior knowledge, including:
[0015] A lightweight convolutional network including convolutional layers is used to encode the image sequence. Specifically, the convolutional layers are used to perform sliding encoding on different image blocks in the image sequence to obtain local image features containing local information. The local image features are concatenated with the global image features and then mapped into multi-dimensional features as prior knowledge through layer normalization and a fully connected layer.
[0016] Preferably, a high-resolution three-plane representation building module including a two-stream Transformer, a convolutional layer, and a pixel rearrangement operation is used to perform feature fusion on the multi-dimensional features and the initial three-plane features to organize the initial three-plane features into a high-resolution three-plane representation, including:
[0017] The attention mechanism in the two-stream Transformer adopts a fast attention mechanism. The two-stream Transformer takes the input initial three-plane features as one stream, and the input multi-dimensional features and randomly initialized features as another stream, and divides the two-stream features into blocks through the feature fusion basic block of the fast attention mechanism to achieve feature fusion. The obtained fused features are further feature encoded by the convolutional layer, and then the resolution is adjusted by the pixel rearrangement operation to obtain a high-resolution three-plane representation, which contains the geometric structure and texture detail information of the three-dimensional object.
[0018] Preferably, generating a three-dimensional mesh based on a three-plane representation using a three-dimensional mesh construction module comprises:
[0019] A three-dimensional mesh is generated based on a three-plane representation using a three-dimensional mesh construction module including a signed distance field (SDF) decoder, a weight decoder, an offset decoder, and a FlexiCubes model. Specifically, a set of spatial points is initialized, and the three-plane representation features of the spatial points are queried from the three-plane representation. The SDF decoder, the weight decoder, and the offset decoder are respectively decoded based on the three-plane representation of the spatial points to obtain SDF values, weight values, and offset values, which constitute FlexiCubes parameters. The FlexiCubes model constructs a three-dimensional mesh based on the FlexiCubes parameters.
[0020] Preferably, the multi-dimensional feature extraction and fusion module, the high-resolution three-plane representation construction module, the three-dimensional grid construction module, and the color decoder are trained in two stages before being applied. The training process does not require the use of a multi-view image diffusion model. The training process includes:
[0021] The first stage of training is as follows: the multi-dimensional feature extraction and fusion module accepts the input of multi-view images and extracts the corresponding multi-dimensional features from them. The high-resolution three-plane representation construction module accepts the multi-dimensional features to construct a three-plane representation, initializes the spatial point set, and queries the three-plane representation features of the spatial points from the three-plane representation; the neural radiation field model including the density decoder and the color decoder is used to decode the three-plane representation features of the spatial points to obtain the density value and the color value, and the volume rendering is performed based on the density value and the color value and the camera parameters to obtain the first rendered color image, the first rendered depth image, and the first rendered mask image corresponding to the camera perspective, and the mean square error loss and the perceptual loss of the first rendered color image and the mean square error loss of the first rendered mask image are constructed to form the first stage training loss and conduct training;
[0022] The second stage of training is as follows: the multi-dimensional feature extraction and fusion module accepts the input of multi-view images and extracts the corresponding multi-dimensional features from them. The high-resolution three-plane representation construction module accepts the multi-dimensional features to construct a three-plane representation, initializes the spatial point set, and queries the three-plane representation features of the spatial points from the three-plane representation; the SDF decoder, weight decoder, and offset decoder are used to decode the features of the spatial point set to obtain the SDF value, weight value, and offset value, which constitute the FlexiCubes parameters. The FlexiCubes model constructs a three-dimensional grid based on the FlexiCubes parameters; the Nvdiffrast algorithm is used to render based on the three-dimensional grid, color texture, and camera parameters to obtain the second rendered color image, the second rendered depth image, the second rendered mask image, and the rendered normal image, and the mean square error loss and perceptual loss of the second rendered color image, the mean square error loss of the second rendered mask image, the L1 loss of the second rendered depth image, the Cosine loss of the rendered normal image, and the regularization loss of the FlexiCubes parameters are constructed to constitute the second stage training loss and train.
[0023] Preferably, the first stage training loss It is expressed as:
[0024] ;
[0025] in, Represents a sequence of rendered color images, with superscript represents an image sequence consisting of M images, express The corresponding color image true value sequence, represents the loss weight corresponding to the mask image, represents the sequence of mask images for rendering, express The corresponding mask image true value sequence, represents the weight parameter corresponding to the perceptual loss, It means using VGG network to extract deep features of rendered color image and true value of color image, and use them to calculate perceptual loss. represents mean square error;
[0026] The loss in the second training phase It is expressed as:
[0027] ;
[0028] in, represents the rendered second depth image sequence, express The corresponding second depth image true value sequence, Denote the rendered normal image sequence, Denote The corresponding ground truth sequence of normal images, the symbol Denote element-wise multiplication, Denote the L1 loss of the second rendered depth image, Denote the weight parameter corresponding to the L1 loss, Denote the cosine loss of the rendered normal image, Denote the weight parameter corresponding to the cosine loss, Denote the regularization loss of the FlexiCubes parameter, Denote the weight parameter of the regularization loss;
[0029] When calculating the first training loss, Take the value of the first color image sequence , Is the first mask image sequence When calculating the second training loss, Take the value of the second color image sequence , Is the second mask image sequence .
[0030] To achieve the above object of the invention, the embodiment also provides a three-dimensional mesh reconstruction device with enhanced multi-dimensional features and three-plane representations, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the above three-dimensional mesh reconstruction method with enhanced multi-dimensional features and three-plane representations.
[0031] To achieve the above object of the invention, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above three-dimensional mesh reconstruction method with enhanced multi-dimensional features and three-plane representations.
[0032] Compared with the prior art, the beneficial effects of the present invention at least include:
[0033] Combined with the image sequence generation part and the multi-dimensional feature extraction and fusion module, the present invention can obtain more comprehensive multi-dimensional features as prior knowledge through a single image input, covering the spatial information, global information, and local information of multiple perspectives of a three-dimensional object, providing sufficient and rich conditions for subsequent three-dimensional mesh reconstruction. Secondly, the high-resolution tri-plane representation construction module designed based on the dual-stream Transformer, convolutional layer, and pixel rearrangement operation not only significantly improves the resolution of the tri-plane representation but also ensures that the computational and storage complexities only increase linearly. Through this optimized design, while ensuring the efficiency of the network model, the present invention significantly improves the expressive ability of the tri-plane representation, enabling the subsequent three-dimensional mesh construction module to construct a high-precision three-dimensional mesh from the tri-plane representation. The improved resolution of the tri-plane representation effectively enhances the accuracy of the generated three-dimensional mesh, reduces blur and artifacts, and improves the accuracy of the geometric structure, further enhancing the performance of the three-dimensional mesh in terms of geometric structure and texture details. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 is the overall process framework diagram of the three-dimensional mesh reconstruction method with enhanced multi-dimensional features and tri-plane representation provided by the embodiment;
[0036] Figure 2 is the structural diagram of the multi-dimensional feature extraction and fusion module provided by the embodiment;
[0037] Figure 3 is the structural diagram of the camera condition injection module provided by the embodiment;
[0038] Figure 4 is the structural diagram of the high-resolution tri-plane representation construction module provided by the embodiment;
[0039] Figure 5 is the process framework diagram of the first-stage training of the network model provided by the embodiment;
[0040] Figure 6 is the process framework diagram of the second-stage training of the network model provided by the embodiment;
[0041] Figure 7 is the three-dimensional mesh reconstruction example diagram provided by the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0043] The inventive concept of the present invention is as follows: The existing single-image-driven 3D mesh generation technology has the problem of low precision in the generation result, that is, the geometric structure accuracy of the mesh is insufficient and its surface is prone to blurring and artifacts. This is mainly because the existing technology fails to fully extract multi-dimensional information in the image and the expressive ability of the used triplane representation is insufficient, resulting in poor mesh precision. To solve this problem, the present invention provides a 3D mesh reconstruction solution with enhanced multi-dimensional features and triplane representation. This solution is improved from two aspects: feature extraction and triplane representation. On the one hand, a multi-dimensional feature extraction and fusion module is used to fully obtain 3D prior knowledge, and on the other hand, a high-resolution triplane representation construction module is used to generate a high-resolution triplane representation while being efficient. This significantly improves the precision of the generated 3D mesh, ensuring that the mesh is more delicate and realistic in terms of geometric structure and texture details. Finally, the present invention can better meet the needs of individual creators and the requirements of industry application scenarios, providing a high-precision and high-quality 3D mesh reconstruction solution.
[0044] As Figure 1 shown, a 3D mesh reconstruction method with enhanced multi-dimensional features and triplane representation provided by the embodiment can obtain sufficient 3D prior knowledge under the input of a single image, construct a high-resolution triplane representation, and finally generate a high-precision 3D mesh, which specifically includes the following steps:
[0045] Step 1, 3D prior knowledge acquisition: After generating a multi-view image sequence based on a single image, use a multi-dimensional feature extraction and fusion module to extract the multi-dimensional features of the image sequence.
[0046] In the embodiment, a pre-trained multi-view image diffusion model is used to generate a multi-view image sequence based on the input single image where the multi-view image diffusion model can be the Zero123++ model. Specifically, the single image is input into the Zero123++ model to generate a 6-view image sequence , and the image sequence is used to provide richer 3D prior knowledge. This image sequence is input into the multi-dimensional feature extraction and fusion module for feature extraction to obtain multi-dimensional features . Among them, before the single image is input, it also undergoes preprocessing, including background removal, cropping, and scaling, to make it meet the input format of the multi-view image diffusion model.
[0047] In the embodiment, the following process is carried out in the multi-dimensional feature extraction and fusion module:
[0048] For a single image , where H is the image height, W is the image width, and C is the number of image channels, usually 3. As Figure 2 shown, this image is first processed by a convolutional layer and divided into non-overlapping image patches for encoding to obtain an initial image embedding vector , is the total number of single image patches, P is the side length of each patch, represents the feature dimension after encoding each patch; then the camera parameters are mapped to a camera embedding vector through a fully connected layer to provide the spatial information of the three-dimensional object. Among them, the camera parameters include the internal and external parameters of the camera, and the specific settings are as follows:
[0049] The external parameters are, the azimuth angles are set to {30, 90, 150, 210, 270, 330} degrees respectively, the pitch angles are set to {20, -10, 20, -10, 20, -10} degrees respectively, the distance of the camera from the origin is set to 4. It should be emphasized that the object in the input image is default at the origin of the world coordinate system and scaled within the cube. When both the azimuth angle and the pitch angle are 0, it matches the pose of the object in the input single image. The internal parameter field of view angle is set to 30 degrees, the focal length is about 1.866, and the principal point (image center) is at the position of the normalized coordinate .
[0050] To effectively fuse the spatial information and the image features, a camera condition injection module is added to the DINO model to obtain a modified DINO model. Among them, the DINO model can adopt the DINOv2 model. The structure of the camera condition injection module is as Figure 3 shown. The camera embedding vector is mapped to twice the image coding dimension through two fully connected layers and split into a scaling vector and an offset vector for channel-level scaling and translation operations on the image features, so as to obtain an image global feature containing spatial information and global information. The modified DINO model can not only focus on the semantic-level information (i.e., global information) of the image, but also the output image global feature contains spatial information and global information.
[0051] However, for the 3D reconstruction task, the local information of the image is also crucial. Existing methods often ignore the representation of local information, resulting in the lack of geometric structure and texture details in the generated 3D mesh. Therefore, in order to capture the local information in the image, a lightweight convolutional network with convolutional layers is designed to encode the image sequence, ensuring that the setting of the convolutional layer is consistent with the convolutional layer in the DINO model, so as to achieve alignment with the global features of the image in the channel dimension, and adopt a Patchify operation similar to ViT to extract local information of different regions by sliding the convolution kernel on the image block to obtain the local features of the image. Finally, the local features of the image and the global features of the image are spliced in the channel dimension, and after layer normalization and full connection layer, the final multi-dimensional features are obtained. As prior knowledge.
[0052] The obtained multi-dimensional features The spatial information, global image information and local image information are integrated to significantly improve the accuracy and detail expression of 3D reconstruction. The present invention injects camera parameters into the DINO model to obtain spatial information, and extracts local information of the image by combining lightweight convolutional layers and Patchify operations to generate feature vectors containing multi-dimensional information, which provides a solid foundation for the subsequent three-plane representation construction and high-precision 3D mesh (3D mesh) reconstruction.
[0053] Step 2, three-plane representation construction: Use a high-resolution three-plane representation construction module consisting of a two-stream Transformer, a convolutional layer, and a pixel reordering operation to fuse the multi-dimensional features and the initial three-plane features to organize the initial three-plane features into a high-resolution three-plane representation.
[0054] In getting the feature vector After that, it is necessary to use the high-resolution three-plane representation construction module to construct the three-plane representation. The constructed three-plane representation is used as the intermediate representation of three-dimensional generation. The shape of the three-plane representation is set to , where 3 represents three mutually perpendicular planes, Represents the height and width of each plane, that is, the resolution, and also represents the number of features stored in each plane. The three-plane representation is composed of the XY plane, XZ plane and YZ plane of the Cartesian coordinate system, and stores various information of the spatial point.
[0055] In order to obtain the three-plane representation features of the spatial point, it is necessary to query in the three-plane representation. First, the coordinates of the spatial point are scaled to within the range of the cube; project the scaled spatial point coordinates onto the three planes represented by the three planes to obtain the two-dimensional coordinates on different planes; since the plane contains The feature vectors cannot cover every coordinate point on the plane, so it is necessary to use bilinear interpolation to obtain the corresponding feature vectors; finally, the feature vectors obtained from each plane are spliced to obtain the feature vectors of the three-plane representation of the spatial point. It can be seen that the resolution of the three-plane representation determines the number of features on each plane. If the resolution is too low, it will affect the expressiveness of the spatial point features. Existing methods usually use a Transformer-based network to construct a three-plane representation, in which the three-plane representation is used as the query matrix (Query), and the computational storage complexity of the self-attention operation is in a square relationship with the sequence length, that is, the resolution of the three-plane representation ( ) will lead to a quadratic increase in the computational storage complexity of the Transformer structure, which will seriously affect the training and reasoning efficiency of the model. Therefore, the three-plane representation resolution of existing methods is usually low. For example, the three-plane representation resolution of the InstantMesh and LDM models is , however, this also makes the results generated by these models less accurate.
[0056] To this end, the structure of the high-resolution three-plane representation building block constructed in the embodiment of the present invention is as follows: Figure 4 As shown in the figure, it includes a two-stream Transformer, a convolutional layer, and a pixel rearrangement operation. The two-stream Transformer splits the self-attention mechanism and the cross-attention mechanism in the traditional Transformer structure into two independent streams, and significantly reduces the computational and storage complexity by adjusting the execution order of the attention mechanism, so that when the resolution of the three-plane representation is improved, the complexity only increases linearly, thereby supporting higher-resolution three-plane representations. Specifically, the attention mechanism in the two-stream Transformer adopts a fast attention mechanism, including a fast attention mechanism 2 (i.e., Flash Attention2), i.e. Figure 4 The two-stream Transformer uses the input initial three-plane features as one stream, and the input multi-dimensional features and randomly initialized features as another stream. The two-stream features are processed in blocks by using the feature fusion basic block of the fast attention mechanism to achieve feature fusion. The obtained fused features are further encoded by the convolution layer, and then the resolution is adjusted by the pixel rearrangement operation to obtain a high-resolution three-plane representation, which effectively improves the resolution and detail expression ability of the three-plane representation. The three-plane representation contains all information related to the three-dimensional object, including spatial information, geometric structure, and texture detail information. Specifically:
[0057] The input includes multi-dimensional features , randomly initialize the features and the initial tri-plane features . The shape of the initial tri-plane features is , , which is the number of features in the tri-plane representation, is the dimension of the features, which is set to 1024 in the model, and are concatenated in the sequence dimension after layer normalization and fully connected layers to obtain the conditional feature vector , with the shape of . This conditional feature vector is used to inject the spatial information, geometric structure, and texture information of the 3D mesh into the initial tri-plane features .
[0058] The two-stream Transformer contains 4 feature fusion basic blocks. Each feature fusion basic block consists of 2 fusion blocks and 3 calculation blocks. The attention operation in each block is based on the fast attention mechanism. Through block processing, the efficient fusion of tri-plane features and image features is gradually completed, thus significantly reducing the overall computational complexity. In the first fusion block, there is a fast cross-attention operation using the fast attention mechanism. At this time serves as the query matrix (Query) for the attention operation, serves as the key-value matrix (Key, Value). In the calculation block, fast self-attention and fast cross-attention operations are performed on the intermediate features. In the fast cross-attention operation of the calculation block, the intermediate features serve as the query matrix (Query), and the multi-dimensional features will be input into the calculation block as the key-value matrix (Key, Value) for the fast cross-attention operation using the fast attention mechanism. Subsequently, in the second fusion block, will serve as the query matrix for the fast cross-attention, and the feature vector from the calculation block will serve as the key-value matrix. Through the above attention operation organization, the attention operation of the traditional Transformer is placed in two different streams, while achieving feature fusion, maintaining the efficiency of the structure. In the above calculation process, the computational complexity of the first fusion block is , and the storage complexity is , where d is the feature dimension; the computational complexity of the calculation block is , and the storage complexity is ; the computational complexity of the second fusion block is , and the storage complexity is . It can be seen that when the resolution of the tri-plane representation increases (i.e., When improved, the computational complexity and storage complexity in the dual-stream Transformer only increase linearly. Compared with the prior art, the overall computational storage complexity is significantly reduced.
[0059] where the resolution of the initial three-plane feature is set to , after passing through the dual-stream Transformer, which contains the spatial, geometric structure, and texture information of the three-dimensional object, a three-plane representation is obtained . Then, through the convolutional layer, the resolution is maintained unchanged for further feature encoding, and the shape of the three-plane representation obtained is , there are 4 layers in the convolutional layer, and the size of the convolutional kernel in each layer is , the stride is 1 for all, and the padding is 1 for all. The input dimension and output dimension of the first three layers are 1024, and the input dimension of the last layer is 1024, and the output dimension is 1280. Finally, through the pixel rearrangement operation, the three-plane representation is further improved T to a resolution of . The fineness of the feature expression is significantly enhanced, ensuring that the generated three-dimensional mesh has higher precision, reducing blur and artifacts, and improving the accuracy of the geometric structure. Through the dual-stream Transformer structure, the computational and storage complexity is effectively reduced, so that while increasing the resolution of the three-plane representation, the time and space complexity only increase linearly. Combining the convolutional layer and the pixel rearrangement operation further improves the fineness of the feature expression and the accuracy of the model generation.
[0060] Step 3, texture and mesh generation: Use the three-dimensional mesh construction module to generate a three-dimensional mesh based on the three-plane representation, use the color decoder to decode and generate the color texture of the spatial points based on the three-plane representation, and combine the three-dimensional mesh and the color texture to obtain a textured three-dimensional mesh.
[0061] After obtaining the three-plane representation , it is necessary to generate a three-dimensional mesh. Specifically, use the three-dimensional mesh construction module containing the SDF decoder, weight decoder, offset decoder, and FlexiCubes model to generate a three-dimensional mesh based on the three-plane representation. Specifically: Initialize the spatial point set, represent the spatial points using the coordinates in the Cartesian coordinate system, and query the features of the spatial points from the three-plane representation. Use the SDF decoder, weight decoder, and offset decoder to decode the signed distance field (i.e., the SDF value) , weight value , and offset value respectively based on the three-plane representation of the spatial points to form the FlexiCubes parameters. Among them, is used to determine the relationship between the vertex and the inner and outer spaces, A method for regulating interpolation and quadrilateral splitting in the Dual Marching Cubes algorithm (an improved isosurface extraction algorithm), offset For deforming the vertices of the reference grid, input these three FlexiCubes parameters into the FlexiCubes model to complete grid construction, obtain the constructed 3D grid, and store it in OBJ format or GLB format.
[0062] In the embodiment, a four-layer fully connected layer combined with the ReLU activation function is uniformly adopted in the design of the three decoders, namely the SDF decoder, the weight decoder, and the offset decoder. The dimension of each hidden layer is 64, and the final output channels are 1 ( ), 21 ( ), and 3 ( ), respectively. In the core process of the FlexiCubes model, first, according to the preset resolution and scaling factor, the three-dimensional space is discretized into a regular voxel grid, so as to obtain the vertices and cube units of the reference grid. On this basis, the offset value is added to each reference vertex one by one to form the coordinates of the deformed grid vertices; at the same time, by parsing the obtained weight value , the interpolation position and grid topology (especially the quadrilateral splitting strategy) can be controlled in the subsequent extraction process, so as to endow the entire grid construction process with sufficient differentiable freedom. Subsequently, the FlexiCubes model adopts the idea of Dual Marching Cubes to discriminate the symbol information (positive and negative values) of each vertex, and accurately identifies the surface voxels and surface grid edges. At each grid edge intersecting the surface, based on the SDF zero-crossing point and the weight value , linear interpolation or least squares calculation is performed to solve for the dual vertex, and the triangles or quadrilaterals of the grid surface are stitched together accordingly to obtain the 3D grid.
[0063] In the embodiment, it is also necessary to obtain the texture information (i.e., color texture) of the grid surface. For the triplanar representation features corresponding to the spatial points, use a color decoder (consisting of 4 fully connected layers) to decode the triplanar representation to obtain the color texture corresponding to the spatial points. Finally, use the trimesh library to combine the 3D grid and the color texture to output a textured 3D grid.
[0064] In the above multi-dimensional feature and three-plane representation enhanced 3D mesh reconstruction method, the multi-dimensional feature extraction and fusion module, high-resolution three-plane representation construction module, 3D mesh construction module, and color decoder are trained in two stages before being applied to optimize the model parameters. Among them, both the first-stage training and the second-stage training require three parts, namely 3D prior knowledge acquisition, three-plane representation construction, and differentiable rendering. Among them, 3D prior knowledge acquisition is different from that during inference. At this time, there is no need to introduce a pre-trained multi-view image diffusion model, and the input is multi-view images; the three-plane representation construction is the same as that during inference; the last step during inference is the generation of textures and meshes. However, since supervision is performed on image data during training, a differentiable rendering part is introduced to obtain the rendered images. The difference between the first-stage training and the second-stage training lies in the differentiable rendering technology. The first stage uses volume rendering technology, and the second stage uses grid-based differentiable rendering (Nvdiffrast).
[0065] The goal of the first-stage training is to enable the network to capture approximate 3D mesh information under image drive, so that the network model can initially construct the mapping relationship from images to three-plane representations. To conduct more stable training, an explicit 3D mesh is not constructed first, and the NeRF model is introduced to replace the FlexiCubes part. The specific process framework of the first-stage training is as Figure 5 shown, the multi-view image sequence passes through the multi-dimensional feature extraction and fusion module to obtain multi-dimensional features , combined with the initial three-plane features to construct the three-plane representation and obtain a high-resolution three-plane representation . Using the density decoder and color decoder in the neural radiance field model (NeRF model), based on the three-plane representation decode the spatial point features into density values and color values. Based on the density values, color values, and camera parameters, the rendered images of the camera parameters specified perspectives can be rendered, including the first rendered color image, the first rendered depth image, and the first rendered mask image, and the mean squared error loss (MSE loss) and perceptual loss of the first rendered color image, and the mean squared error loss of the first rendered mask image are constructed to form the first-stage training loss and perform training. Among them, the MSE loss is used to provide supervision for geometric structures and texture details, and the perceptual loss is used to provide supervision at the semantic level. The specific first training stage loss is expressed as:
[0066] ;
[0067] Among them, represents the sequence of rendered color images, and the superscript represents an image sequence consisting of M images, express The corresponding color image true value sequence, represents the loss weight corresponding to the mask image, represents the sequence of mask images for rendering, express The corresponding mask image true value sequence, represents the weight parameter corresponding to the perceptual loss, It means using VGG network to extract deep features of rendered color image and true value of color image, and use them to calculate perceptual loss. Represents the mean square error.
[0068] The purpose of the second stage of training is to enable the network to capture more detailed 3D mesh information under the input of multi-view images and map it into a more informative three-dimensional plane representation, thereby improving the final generation effect. At this time, FlexiCubes is used instead of NeRF to construct the 3D mesh (the 3D mesh generation process is the same as the inference stage). The process framework of the second stage of training is as follows: Figure 6 As shown, the multi-view image sequence After multi-dimensional feature extraction and fusion module, multi-dimensional features are obtained After that, combined with the initial three-plane features Perform three-plane representation construction to obtain a high-resolution three-plane representation Initialize the spatial point set, query the features of the spatial point set from the three-plane representation, and decode the spatial point features into SDF values using the SDF decoder, weight decoder, and offset decoder respectively. , weight value and offset value , which constitutes FlexiCubes parameters. Combining these FlexiCubes parameters, the FlexiCubes model can construct a 3D mesh (without texture). Using a color decoder to decode the three-plane representation of the spatial points on the mesh can obtain the color texture. The Nvdiffrast algorithm can render the camera parameters based on the 3D mesh, color texture and camera parameters. The rendered image of the specified perspective includes a second rendered color image, a second rendered depth image, a second rendered mask image, and a rendered normal image. The mean square error loss and perceptual loss of the second rendered color image and the mean square error loss of the second rendered mask image, that is, the loss of the first training stage, are constructed. In addition, the L1 loss of the second rendered depth image, the Cosine loss of the rendered normal image, and the regularization loss of the FlexiCubes parameters are constructed to constitute the second stage training loss and are trained. The regularization loss makes the generated mesh shape more reasonable.
[0069] The training loss in the second stage is expressed as:
[0070] ;
[0071] wherein, represents the second depth image sequence rendered, represents the corresponding ground truth sequence of the second depth image, represents the rendered normal image sequence, represents the corresponding ground truth sequence of the normal image, the symbol represents element-wise multiplication, represents the L1 loss of the second rendered depth image, represents the weight parameter corresponding to the L1 loss, represents the Cosine loss of the rendered normal image, represents the weight parameter corresponding to the Cosine loss, represents the regularization loss of the FlexiCubes parameter, represents the weight parameter of the regularization loss;
[0072] When calculating the first training loss, takes the value of the first color image sequence , is the first mask image sequence , when calculating the second training loss, takes the value of the second color image sequence , is the second mask image sequence .
[0073] After the above two-stage training, it is possible to input a single image through the process framework shown in Figure 1 and output a high-precision 3D mesh. Figure 7 shows the final generation result, and it can be seen that the generation result has no blur or artifacts and the geometric structure is accurate.
[0074] Based on the same inventive concept, the embodiment also provides a 3D mesh reconstruction device with enhanced multi-dimensional features and three-plane representations, including a memory and one or more processors. When the one or more processors execute the executable code stored in the memory, it is used to implement the above-mentioned 3D mesh reconstruction method with enhanced multi-dimensional features and three-plane representations, specifically including the following steps:
[0075] Step 1, obtaining 3D prior knowledge: After generating a multi-view image sequence based on a single image, use the multi-dimensional feature extraction and fusion module to extract the multi-dimensional features of the image sequence;
[0076] Step 2, Tri-Plane Representation Construction: Use a high-resolution tri-plane representation construction module that includes a dual-stream Transformer, convolutional layers, and pixel rearrangement operations to perform feature fusion on the multi-dimensional features and the initial tri-plane features, so as to organize the initial tri-plane features into a high-resolution tri-plane representation;
[0077] Step 3, Texture and Mesh Generation: Use a 3D mesh construction module to generate a 3D mesh based on the tri-plane representation, use a color decoder to decode and generate the color texture of spatial points based on the tri-plane representation, and combine the 3D mesh and the color texture to obtain a textured 3D mesh.
[0078] The 3D mesh reconstruction device with enhanced multi-dimensional features and tri-plane representation provided by the embodiment belongs to hardware equipment. At the hardware level, in addition to including a processor and a memory, it also includes an internal bus, a network interface, memory, and other hardware required for other services. The memory is a non-volatile memory. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the 3D mesh reconstruction method with enhanced multi-dimensional features and tri-plane representation in the above Steps 1 - Step 3. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0079] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above 3D mesh reconstruction method with enhanced multi-dimensional features and tri-plane representation, specifically including the following steps:
[0080] Step 1, 3D Prior Knowledge Acquisition: After generating a multi-view image sequence based on a single image, use a multi-dimensional feature extraction and fusion module to extract the multi-dimensional features of the image sequence;
[0081] Step 2, Tri-Plane Representation Construction: Use a high-resolution tri-plane representation construction module that includes a dual-stream Transformer, convolutional layers, and pixel rearrangement operations to perform feature fusion on the multi-dimensional features and the initial tri-plane features, so as to organize the initial tri-plane features into a high-resolution tri-plane representation;
[0082] Step 3, Texture and Mesh Generation: Use a 3D mesh construction module to generate a 3D mesh based on the tri-plane representation, use a color decoder to decode and generate the color texture of spatial points based on the tri-plane representation, and combine the 3D mesh and the color texture to obtain a textured 3D mesh.
[0083] In an embodiment, a computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.
[0084] The specific embodiments described above have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above are only the most preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representation, characterized in that: The following steps are involved: Acquisition of three-dimensional prior knowledge: After generating a multi-view image sequence based on a single image, a multi-dimensional feature extraction and fusion module is used to extract multi-dimensional features of the image sequence, specifically including: allowing the modified DINO model to accept camera parameter input, and considering spatial information when extracting image features in combination with the image sequence, to obtain image global features containing global information and spatial information, and also encoding based on the image sequence to obtain image local features containing local information, and splicing the image global features and image local features to obtain multi-dimensional features as prior knowledge; Three-plane representation construction: Use a high-resolution three-plane representation construction module including a two-stream Transformer, a convolutional layer, and a pixel rearrangement operation to fuse the multi-dimensional features and the initial three-plane features to organize the initial three-plane features into a high-resolution three-plane representation, including: the attention mechanism in the two-stream Transformer adopts a fast attention mechanism, the two-stream Transformer takes the input initial three-plane features as one stream, takes the input multi-dimensional features and the randomly initialized features as another stream, and performs block processing on the two-stream features through the feature fusion basic block using the fast attention mechanism to achieve feature fusion. The obtained fused features are further feature encoded by the convolutional layer, and then the resolution is adjusted by the pixel rearrangement operation to obtain a high-resolution three-plane representation, which contains the geometric structure and texture detail information of the three-dimensional object; Texture and mesh generation: Use the 3D mesh construction module to generate a 3D mesh based on the three-plane representation. Use the color decoder to generate the color texture of the spatial points based on the three-plane representation decoding. Combine the 3D mesh and color texture to get a textured 3D mesh.
2. The three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representation according to claim 1 is characterized in that: Generate multi-view image sequences based on a single image, including: A pre-trained multi-view image diffusion model is used to generate a multi-view image sequence based on a single image, where the multi-view image diffusion model includes a Zero123++ model.
3. The three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representation according to claim 1 is characterized in that: The modified DINO model accepts camera parameter input and considers spatial information when extracting image features in combination with the image sequence, thereby obtaining image global features containing global information and spatial information, including: A camera condition injection module is added to the DINO model to obtain a modified DINO model. The camera condition injection module obtains a scaling vector and an offset vector based on camera parameters. The scaling vector and the offset vector are used to scale and offset the features extracted based on the image sequence in the DINO model, thereby injecting spatial information into it, so that the obtained global features of the image contain spatial information and global information.
4. The three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representation according to claim 1, characterized in that: Based on the image sequence, the local features of the image containing local information are obtained by encoding, and the global features of the image and the local features of the image are spliced to obtain multi-dimensional features as prior knowledge, including: A lightweight convolutional network including convolutional layers is used to encode the image sequence. Specifically, the convolutional layers are used to perform sliding encoding on different image blocks in the image sequence to obtain local image features containing local information. The local image features are concatenated with the global image features and then mapped into multi-dimensional features as prior knowledge through layer normalization and a fully connected layer.
5. The three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representation according to claim 1, characterized in that: Generate a 3D mesh based on a tri-planar representation using the 3D Mesh Building Blocks, including: A three-dimensional grid is generated based on a three-plane representation using a three-dimensional grid construction module including a signed distance field decoder, a weight decoder, an offset decoder, and a FlexiCubes model, specifically: a set of spatial points is initialized, and the three-plane representation features of the spatial points are queried from the three-plane representation, and the signed distance field value, the weight value, and the offset value are obtained by decoding the three-plane representation features of the spatial points using a signed distance field decoder, a weight decoder, and an offset decoder, respectively, to form FlexiCubes parameters, and the FlexiCubes model constructs a three-dimensional grid based on the FlexiCubes parameters.
6. The three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representation according to claim 5, characterized in that: The multi-dimensional feature extraction and fusion module, the high-resolution three-plane representation construction module, the three-dimensional mesh construction module, and the color decoder are trained in two stages before being applied, including: The first stage of training is as follows: the multi-dimensional feature extraction and fusion module accepts the input of multi-view images and extracts the corresponding multi-dimensional features from them. The high-resolution three-plane representation construction module accepts the multi-dimensional features to construct a three-plane representation, initializes the spatial point set, and queries the three-plane representation features of the spatial points from the three-plane representation; the neural radiation field model including the density decoder and the color decoder is used to decode the three-plane representation features of the spatial points to obtain the density value and the color value, and the volume rendering is performed based on the density value and the color value and the camera parameters to obtain the first rendered color image, the first rendered depth image, and the first rendered mask image corresponding to the camera perspective, and the mean square error loss and the perceptual loss of the first rendered color image and the mean square error loss of the first rendered mask image are constructed to form the first stage training loss and conduct training; The second stage of training is as follows: the multi-dimensional feature extraction and fusion module accepts the input of multi-view images and extracts the corresponding multi-dimensional features from them. The high-resolution three-plane representation construction module accepts the multi-dimensional features to construct a three-plane representation, initializes the spatial point set, and queries the three-plane representation features of the spatial points from the three-plane representation. The three-plane representation features of the spatial points are decoded using the signed distance field decoder, the weight decoder, and the offset decoder to obtain the SDF value, weight value, and offset value. These values constitute the FlexiCubes parameters. The FlexiCubes model constructs a three-dimensional grid based on the FlexiCubes parameters; the Nvdiffrast algorithm is used to render based on the three-dimensional grid, color texture, and camera parameters to obtain the second rendered color image, the second rendered depth image, the second rendered mask image, and the rendered normal image. The mean square error loss and perceptual loss of the second rendered color image, the mean square error loss of the second rendered mask image, the L1 loss of the second rendered depth image, the cosine loss of the rendered normal image, and the regularization loss of the FlexiCubes parameters are constructed to constitute the second stage training loss and training is performed.
7. The three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representation according to claim 6, characterized in that: First stage training loss It is expressed as: ; in, Represents a sequence of rendered color images, with superscript represents an image sequence consisting of M images, express The corresponding color image true value sequence, represents the loss weight corresponding to the mask image, represents the sequence of mask images for rendering, express The corresponding mask image true value sequence, represents the weight parameter corresponding to the perceptual loss, It means using VGG network to extract deep features of rendered color image and true value of color image, and use them to calculate perceptual loss. represents mean square error; The loss in the second training phase It is expressed as: ; in, represents the rendered second depth image sequence, express The corresponding second depth image true value sequence, represents a sequence of rendered normal images, express The corresponding normal image true value sequence, symbol represents element-wise multiplication, represents the L1 loss of the second rendered depth image, represents the weight parameter corresponding to the L1 loss, represents the cosine loss of the rendered normal image, Represents the weight parameter corresponding to the cosine loss, represents the regularization loss of FlexiCubes parameters, The weight parameter representing the regularization loss; When calculating the first training loss, The value is the first color image sequence , is the first mask image sequence , when calculating the second training loss, The value is the second color image sequence , is the second mask image sequence .
8. A three-dimensional mesh reconstruction device enhanced by multi-dimensional features and three-plane representation, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the three-dimensional mesh reconstruction method with enhanced multi-dimensional features and three-plane representations as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the three-dimensional mesh reconstruction method enhanced by multi-dimensional features and three-plane representations as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Arbitrary track three-dimensional scene construction and roaming video generation method and system guided by plain text
CN117853686A
Monocular indoor scene reconstruction method based on triangular three planes
CN118447182A