Three-dimensional reconstruction method and device, equipment and storage medium

Through adaptive voxel updating and differentiable rendering optimization, the resolution of spatial grids and three-plane features is gradually improved, solving the problems of low efficiency and insufficient accuracy of 3D reconstruction in existing technologies, and achieving efficient and high-quality 3D model reconstruction.

CN120612447APending Publication Date: 2025-09-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410278451.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In existing 3D reconstruction technology, the volume rendering process requires a large number of parameters and time, resulting in low geometric model accuracy and low training efficiency.

Method used

Through adaptive voxel updating and differentiable rendering optimization, the resolution of spatial grids and three-plane features is gradually improved, and high-resolution three-dimensional models are generated using geometric iteration and differentiable rendering optimization, reducing data volume and improving reconstruction efficiency.

Benefits of technology

While ensuring the quality of the 3D model, the amount of data in the 3D reconstruction process is significantly reduced, and the reconstruction efficiency and model resolution are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612447A_ABST
    Figure CN120612447A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a three-dimensional reconstruction method and device, equipment and a storage medium, and belongs to the technical field of three-dimensional reconstruction. The method comprises the steps that self-adaptive voxel updating is carried out on an nth-level space grid corresponding to a target object, an (n + 1) th-level space grid is obtained, and the voxel distribution density of the three-dimensional surface in the (n + 1) th-level space grid is higher than the voxel distribution density of the three-dimensional surface in the nth-level space grid; performing up-sampling processing on the nth-level three-plane feature corresponding to the nth-level space grid to obtain an (n + 1) th-level three-plane feature; grid vertex features of all grid vertexes in the (n + 1) th-level space grid are extracted from the (n + 1) th-level three-plane features; obtaining an (n + 1) th-level three-dimensional model corresponding to the target object through geometric iteration and micro rendering optimization based on the grid vertex characteristics of each grid vertex in the (n + 1) th-level space grid; the data volume in the three-dimensional reconstruction process is reduced, and the three-dimensional reconstruction efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of three-dimensional reconstruction technology, and in particular to a three-dimensional reconstruction method, apparatus, device, and storage medium. Background Art

[0002] Three-dimensional reconstruction refers to the establishment of a mathematical model of three-dimensional objects suitable for computer representation and processing. It is the basis for processing, operating and analyzing the properties of three-dimensional objects in a computer environment, and is also a key technology for establishing virtual reality in computers to express the objective world.

[0003] In related technologies, scenes are represented by learning color and density functions in a continuous three-dimensional space, and three-dimensional objects and textures are rendered through volume rendering, thereby optimizing the MLP (Multilayer Perceptron) network through supervision of the rendered image.

[0004] The volume rendering process requires counting the predicted values ​​of hundreds of sampling points on each viewing ray, and each iteration requires running the MLP network inference hundreds of thousands of times, which has a large number of parameters and consumes a lot of time. In addition, the trained MLP network needs to run the Marching Cube algorithm to obtain the geometric expression, resulting in low accuracy of the geometric model. Summary of the Invention

[0005] The embodiments of the present application provide a 3D reconstruction method, apparatus, device, and storage medium that can reduce the amount of data in the 3D reconstruction process and improve the efficiency of 3D reconstruction. The technical solution is as follows:

[0006] In one aspect, an embodiment of the present application provides a three-dimensional reconstruction method, the method comprising:

[0007] Adaptively updating voxels on an n-th level spatial grid corresponding to the target object to obtain an n+1-th level spatial grid, wherein a density of voxel distribution of the three-dimensional surface in the n+1-th level spatial grid is higher than a density of voxel distribution of the three-dimensional surface in the n-th level spatial grid;

[0008] Upsampling the n-th level three-plane features corresponding to the n-th level spatial grid to obtain the n+1-th level three-plane features;

[0009] Extracting mesh vertex features of each mesh vertex in the n+1th level spatial mesh from the n+1th level three-plane features;

[0010] Based on the mesh vertex features of each mesh vertex in the n+1-level spatial mesh, through geometric iteration and differentiable rendering optimization, an n+1-level three-dimensional model corresponding to the target object is obtained, and the model resolution of the n+1-level three-dimensional model is higher than the model resolution of the n-level three-dimensional model. The geometric iteration is used to generate the n+1-level patch mesh corresponding to the target object, and the differentiable rendering is used to perform differentiable rendering on the n+1-level patch mesh.

[0011] On the other hand, an embodiment of the present application provides a three-dimensional reconstruction device, comprising:

[0012] a voxel updating module configured to adaptively update the nth-level spatial grid corresponding to the target object to obtain an n+1th-level spatial grid, wherein the density of voxel distribution of the three-dimensional surface in the n+1th-level spatial grid is higher than the density of voxel distribution of the three-dimensional surface in the nth-level spatial grid;

[0013] a sampling processing module, configured to perform upsampling processing on the n-th level three-plane features corresponding to the n-th level spatial grid to obtain the n+1-th level three-plane features;

[0014] a feature extraction module, configured to extract mesh vertex features of each mesh vertex in the n+1th level spatial mesh from the n+1th level three-plane features;

[0015] A model generation module is configured to obtain, based on mesh vertex features of each mesh vertex in the n+1-th level spatial mesh, a level n+1 three-dimensional model corresponding to the target object through geometric iteration and differentiable rendering optimization, wherein the model resolution of the n+1-th level three-dimensional model is higher than the model resolution of the n-th level three-dimensional model, the geometric iteration is configured to generate the level n+1 patch mesh corresponding to the target object, and the differentiable rendering is configured to perform differentiable rendering on the level n+1 patch mesh.

[0016] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory; the memory stores at least one computer instruction, and the at least one computer instruction is used to be executed by the processor to implement the three-dimensional reconstruction method described in the above aspect.

[0017] On the other hand, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer instruction is stored. The at least one computer instruction is loaded and executed by a processor to implement the three-dimensional reconstruction method as described in the above aspects.

[0018] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; the processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the three-dimensional reconstruction method described in the above aspects.

[0019] In an embodiment of the present application, in order to achieve three-dimensional reconstruction of the target object, iterative optimization corresponding to multiple levels of resolution is set, and the first-level spatial grid and first-level three-plane features corresponding to the target object are constructed starting from low resolution. The first-level three-dimensional model is obtained through geometric iteration and differentiable rendering optimization, so that the first-level spatial grid is adaptively updated with voxels, so that the voxel distribution density of the three-dimensional surface in the spatial grid is gradually increased. At the same time, the first-level three-plane features are upsampled, so that the three-plane features are gradually refined. Then, based on the second-level spatial grid and the second-level three-plane features, the second-level three-dimensional model is obtained through geometric iteration and differentiable rendering optimization, so that the model resolution of the three-dimensional model is gradually improved, and an optimization reconstruction process from low resolution to high resolution and from coarse to fine is realized. While ensuring the quality of the three-dimensional model reconstruction, the amount of data in the three-dimensional model reconstruction process is reduced, and the three-dimensional reconstruction efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A schematic diagram showing an implementation environment provided by an exemplary embodiment of the present application is shown;

[0022] Figure 2 A flowchart of a three-dimensional reconstruction method provided by an exemplary embodiment of the present application is shown;

[0023] Figure 3 A schematic diagram of a three-dimensional model from low resolution to high resolution provided by an exemplary embodiment of the present application is shown;

[0024] Figure 4 A flowchart of a three-dimensional reconstruction method provided by another exemplary embodiment of the present application is shown;

[0025] Figure 5 A schematic diagram showing real-life photographs corresponding to multiple camera poses provided by an exemplary embodiment of the present application is shown;

[0026] Figure 6 A schematic diagram of a three-dimensional model from low resolution to high resolution provided by another exemplary embodiment of the present application is shown;

[0027] Figure 7 A schematic diagram showing a three-dimensional model at the highest level of resolution provided by an exemplary embodiment of the present application is shown;

[0028] Figure 8 A flowchart of a three-dimensional reconstruction method provided by another exemplary embodiment of the present application is shown;

[0029] Figure 9 A flowchart of a single-resolution 3D reconstruction method provided by an exemplary embodiment of the present application is shown;

[0030] Figure 10 A flowchart of a three-dimensional reconstruction method provided by another exemplary embodiment of the present application is shown;

[0031] Figure 11 A structural block diagram of a three-dimensional reconstruction device provided by an exemplary embodiment of the present application is shown;

[0032] Figure 12 A schematic structural diagram of a computer device provided by an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION

[0033] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0034] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0035] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0036] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.

[0037] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual reality devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, smart virtual characters in games, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0038] The solutions provided in the embodiments of this application involve technologies such as machine learning of artificial intelligence, which are specifically illustrated by the following embodiments.

[0039] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by one embodiment of the present application. The implementation environment includes a terminal 120 and a server 140. Data communication is performed between the terminal 120 and the server 140 via a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network.

[0040] The terminal 120 is a computer device installed with an application program having a 3D reconstruction function. The 3D reconstruction function can be a function of a native application in the terminal 120 or a function of a third-party application. The terminal 120 can be a smartphone, tablet computer, laptop computer, desktop computer, smart TV, wearable device, or vehicle-mounted terminal, etc. Figure 1 In the description, the terminal 120 is taken as a desktop computer as an example, but this is not a limitation.

[0041] Server 140 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In the embodiment of the present application, server 140 can be a backend server for an application with 3D reconstruction capabilities.

[0042] In one possible implementation, Figure 1 As shown, there is data interaction between the server 140 and the terminal 120. When receiving a 3D reconstruction task for a target object, the terminal 120 sends the acquired real-world images of the target object under multiple camera poses to the server 140. The server 140 initializes the first-level spatial grid and first-level three-plane features corresponding to the target object, and obtains a first-level 3D model through geometric iteration and differentiable rendering optimization based on the mesh vertex features of the first-level spatial grid extracted from the first-level three-plane features. Then, the spatial grid is adaptively updated with voxels, the three-plane features are upsampled, and the n-level 3D model is obtained based on the mesh vertex features of the n-th level spatial grid through geometric iteration and differentiable rendering optimization. Finally, after completing iterative optimization from low to high resolution, the server 140 can return the 3D models corresponding to each level of resolution to the terminal 120.

[0043] Optionally, the 3D reconstruction method provided in the embodiments of the present application can be applied to the fields of game rendering, augmented reality, virtual reality (AR), virtual reality (VR), and artificial intelligence generated content (AIGC). For example, the generated 3D model can be added to the game pipeline as a 3D asset to generate effects such as skinning, driving, and animation. The 3D model can be rendered from a new arbitrary perspective to achieve immersive browsing and experience in AR / VR, thereby improving the user experience.

[0044] In one possible implementation, the three-dimensional reconstruction method provided in the embodiment of the present application is applied to the construction of a virtual character model for a game as an example. In order to perform three-dimensional reconstruction of a character model created in the real world, the computer device first needs to initialize the low-resolution spatial grid and three-plane features corresponding to the character model, obtain a low-resolution three-dimensional model through geometric iteration and differentiable rendering optimization, and then perform step-by-step adaptive voxel updates on the low-resolution spatial grid, while step-by-step upsampling of the low-resolution three-plane features, and complete multiple rounds of iterative optimization corresponding to each level of resolution. A high-resolution three-dimensional model can be obtained, so that the computer device can apply the high-resolution three-dimensional model to the virtual scene of the game, improve the geometry and texture effects of the virtual character in the virtual scene of the game, and make the virtual character more vivid.

[0045] Please refer to Figure 2 , which shows a flowchart of a three-dimensional reconstruction method provided by an exemplary embodiment of the present application. This embodiment takes the method applied to a computer device as an example for explanation. The computer device may be Figure 1 The terminal 120 or server 140 shown in FIG. 1 includes the following steps:

[0046] In step 201 , adaptive voxel update is performed on the nth level spatial grid corresponding to the target object to obtain the n+1th level spatial grid, wherein the voxel distribution density of the three-dimensional surface in the n+1th level spatial grid is higher than that in the nth level spatial grid.

[0047] Unlike related technologies, in order to obtain a high-resolution three-dimensional model, the high-resolution spatial grid and high-resolution three-plane features corresponding to the target object are directly initialized, and then the high-resolution three-plane model corresponding to the target object is obtained by extracting features from the high-resolution three-plane features, and performing geometric iteration and differentiable rendering optimization. This results in a large amount of data in the three-dimensional reconstruction process and requires more computing resources. In the embodiment of the present application, considering that there are a large number of non-object surface areas in the three-dimensional reconstruction process, the three-dimensional reconstruction of the object only requires iterative optimization of the features on the object surface. Therefore, in order to achieve the goal of reducing the amount of calculation in the three-dimensional reconstruction process while ensuring the quality of the three-dimensional reconstruction, the target object can be reconstructed in three dimensions by iterative optimization from low resolution to high resolution and from coarse to fine.

[0048] In some embodiments, the computer device can first initialize a low-resolution spatial grid corresponding to the target object, and then adaptively update the voxels of the low-resolution spatial grid so that the voxels near the object surface in the spatial grid are gradually denser, and the voxels near the non-object surface are gradually sparser.

[0049] Optionally, the spatial grid is composed of a plurality of voxels connected in a tree-like manner. In one possible implementation, the computer device performs voxelization processing on the space where the target object is located to obtain voxels in the three-dimensional space, and then connects the voxels based on an octree data structure to obtain voxels connected in a tree-like manner.

[0050] Voxels can be considered as pixels in three-dimensional space, the smallest unit for segmenting three-dimensional space. Voxels are used to represent spatial units. Voxels have a certain size and position and can be used to store specific attributes. An octree is a tree-like data structure used to describe three-dimensional space. An octree can recursively divide space into eight equal cubic sub-regions until a certain termination condition is reached. Each node in an octree represents a volume element of a cube. Each node has eight child nodes. The sum of the volume elements represented by the eight child nodes equals the volume of the parent node.

[0051] To achieve iterative optimization from low to high resolution, the computer only needs to initialize the spatial grid corresponding to the first-level resolution, thereby obtaining a low-resolution first-level spatial grid. Then, before iterative optimization at the second-level resolution, the computer only needs to perform an adaptive voxel update on the first-level spatial grid to obtain a second-level spatial grid. Furthermore, during the adaptive voxel update, by subdividing voxels near the object surface and merging voxels near non-object surfaces within the first-level spatial grid, the voxel density of the three-dimensional surface in the second-level spatial grid is made higher than that in the first-level spatial grid. This allows for greater focus on vertex features near the object surface during iterative optimization at the second-level resolution. Similarly, by performing adaptive voxel updates on the n-th level spatial grid, a level n+1 spatial grid is obtained, and the voxel density of the three-dimensional surface in the n+1 level spatial grid is higher than that in the n-th level spatial grid.

[0052] Step 202 : Up-sampling the n-th level three-plane features corresponding to the n-th level spatial grid to obtain the (n+1)-th level three-plane features.

[0053] Unlike the related art, in which the mesh vertex features of the mesh vertices of the spatial tetrahedron are maintained through a hash grid, and the mesh vertex features are encoded into a three-dimensional data format, resulting in the mesh vertex features having a large parameter amount, resulting in low training efficiency of the three-dimensional reconstruction network, and hash conflicts are prone to exist in the hash grid. In the embodiment of the present application, by converting the mesh vertex features in the three-dimensional space into planar features on three two-dimensional planes, the parameter amount of the mesh vertex features can be reduced, thereby improving the network training efficiency.

[0054] In some embodiments, a computer device encodes the vertices of each spatial tetrahedron and projects the three-dimensional vertices (x, y, z) onto a two-dimensional plane where the three axes intersect. That is, the two-dimensional projection points of the three-dimensional vertices on the (x, y), (x, z) and (y, z) planes can be obtained respectively, so that the vertex features of the three-dimensional vertices can be represented as the projection point features of the three-dimensional projection points.

[0055] In an illustrative example, there are H in the three-dimensional space with a resolution of H. 3 vertices. If each vertex is individually encoded, C*H 3 The feature variables, where C is the feature dimension of each vertex. By projecting the three-dimensional vertices onto the two-dimensional planes where the three axes intersect, and using three two-dimensional plane features to represent the vertex features, only 3*C*H are needed. 2 The characteristic variables of , it can be seen that the number of parameters is significantly reduced.

[0056] Optionally, a tri-plane feature is a feature representation of a spatial tetrahedron on a two-dimensional plane where three axes intersect. In one possible embodiment, to achieve three-dimensional reconstruction of a target object, a computer device may first randomly initialize a tri-plane feature corresponding to the target object. Then, during an iterative optimization process, by optimizing the tri-plane feature, the tri-plane feature may be obtained to continuously approximate the true features of the target object.

[0057] To achieve iterative optimization from low resolution to high resolution, the computer device also only needs to initialize the first-level three-plane features to obtain the first-level three-plane features for iterative optimization of the first-level resolution. After completing the iterative optimization of the first-level resolution, the iteratively optimized first-level three-plane features are obtained. Then, when entering the iterative optimization of the second-level resolution, the computer device only needs to upsample the first-level three-plane features to obtain the second-level three-plane features. Similarly, by upsampling the n-th level three-plane features, the computer device can obtain the n+1-th level three-plane features, and the feature resolution of the n+1-th level three-plane features is greater than the feature resolution of the n-th level three-plane features.

[0058] Step 203 : extracting mesh vertex features of each mesh vertex in the n+1th level spatial mesh from the n+1th level three-plane features.

[0059] In some embodiments, during the iterative optimization process for each level of resolution, the computer device first extracts mesh vertex features of each mesh vertex in the spatial mesh from the three-plane features, thereby using the mesh vertex features for three-dimensional reconstruction.

[0060] In a possible implementation, in order to accurately extract mesh vertex features from three-plane features, the computer device can also first obtain the mesh vertex coordinates of each mesh vertex in the spatial grid, and then extract the two-dimensional vertex features of the mesh vertices on each two-dimensional plane from the three-plane features based on the mesh vertex coordinates, and then obtain the mesh vertex features of the mesh vertices by accumulating and summing the three two-dimensional vertex features.

[0061] Optionally, the mesh vertex coordinates in the three-dimensional space can be expressed as v = (x, y, z), and the mesh vertex features can be expressed as f(v), then the computer device can extract the features from the three two-dimensional planes (x, y), (x, z), and (y, z) respectively. as well as The process of obtaining mesh vertex features by accumulating and summing three two-dimensional vertex features can be expressed as

[0062] In step 204, based on the mesh vertex features of each mesh vertex in the n+1-th level spatial mesh, a level n+1 three-dimensional model corresponding to the target object is obtained through geometric iteration and differentiable rendering optimization. The model resolution of the n+1-th level three-dimensional model is higher than the model resolution of the n-th level three-dimensional model. The geometric iteration is used to generate the n+1-th level patch mesh corresponding to the target object, and the differentiable rendering is used to perform differentiable rendering on the n+1-th level patch mesh.

[0063] In some embodiments, after obtaining the mesh vertex features of each mesh vertex in the spatial mesh, the computer device can obtain a three-dimensional model corresponding to the target object through geometric iteration and differentiable rendering optimization based on the mesh vertex features of each mesh vertex.

[0064] Among them, geometric iteration is used to generate a patch mesh corresponding to the target object through a moving cube algorithm based on the mesh vertex features of each mesh vertex; differentiable rendering is used to perform differentiable rendering on the patch mesh to obtain a three-dimensional model corresponding to the target object.

[0065] For the iterative optimization of each level of resolution, since the voxel distribution density of the three-dimensional surface in the n+1-level spatial grid is higher than that of the three-dimensional surface in the n-level spatial grid, the feature resolution of the n+1-level three-plane features is higher than that of the n-level three-plane features. Therefore, after completing the iterative optimization of each level of resolution, the model resolution of the obtained n+1-level three-dimensional model is higher than the model resolution of the n-level three-dimensional model.

[0066] Indicative, such as Figure 3As shown, the computer device iteratively optimizes each resolution step by step, and can obtain the n-th level three-dimensional model 302 and the n+1-th level three-dimensional model 303 corresponding to the target object 301, wherein the model resolution of the n+1-th level three-dimensional model 303 is higher than the model resolution of the n-th level three-dimensional model 302.

[0067] To sum up, in the embodiment of the present application, in order to achieve three-dimensional reconstruction of the target object, by setting iterative optimization corresponding to multiple levels of resolution, the first-level spatial grid and the first-level three-plane features corresponding to the target object are constructed starting from low resolution, and the first-level three-dimensional model is obtained through geometric iteration and differentiable rendering optimization, so that the first-level spatial grid is adaptively updated, so that the voxel distribution density of the three-dimensional surface in the spatial grid is gradually increased, and at the same time, the first-level three-plane features are upsampled, so that the three-plane features are gradually refined, and then based on the second-level spatial grid and the second-level three-plane features, the second-level three-dimensional model is obtained through geometric iteration and differentiable rendering optimization, so that the model resolution of the three-dimensional model is gradually improved, and an optimization reconstruction process from low resolution to high resolution and from coarse to fine is realized. While ensuring the quality of the three-dimensional model reconstruction, the amount of data in the three-dimensional model reconstruction process is reduced, and the three-dimensional reconstruction efficiency is improved.

[0068] In some embodiments, after dividing the 3D reconstruction process into 3D reconstruction sub-processes corresponding to multiple levels of resolution, in order to improve the 3D reconstruction efficiency, during the step-by-step resolution transition process, the computer device also needs to optimize the process of adaptive voxel updating of the spatial grid and the process of upsampling three-plane features. In addition, in the 3D reconstruction process of a single level of resolution, it is also necessary to use a grid reconstruction network for grid reconstruction and a differentiable rendering network for differentiable rendering.

[0069] The following embodiments will illustrate the process of step-by-step resolution transition and the iterative optimization process corresponding to a single-level resolution in the embodiments of the present application.

[0070] Please refer to Figure 4 , which shows a flow chart of a three-dimensional reconstruction method provided by another exemplary embodiment of the present application. This embodiment takes the method applied to a computer device as an example for explanation. The computer device may be Figure 1 The terminal 120 or server 140 shown in FIG. 1 includes the following steps:

[0071] Step 401: Determine the signed distance function value of each grid vertex in the n-th level spatial grid.

[0072] In some embodiments, in order to make the voxel distribution on the object surface in the adaptively updated spatial grid denser and the voxel distribution on the non-object surface more sparse, the computer device can determine the spatial position relationship between each grid vertex and the object surface based on the signed distance function (SDF) value of each grid vertex.

[0073] The signed distance function value is used to characterize the correspondence between the mesh vertices and the object surface in three-dimensional space. A positive sign indicates that the mesh vertex is inside the shape of the target object, and a negative sign indicates that the mesh vertex is outside the shape of the target object.

[0074] Alternatively, the computer device may input the mesh vertex features of each mesh vertex in the n-th level spatial grid into a multilayer perceptron, and use the multilayer perceptron to output the signed distance function value of each mesh vertex. In one possible implementation, during the geometric iteration process for each level of spatial grid, since the spatial grid is reconstructed, the computer device can obtain the SDF value corresponding to each mesh vertex in each level of spatial grid.

[0075] Step 402 : Determine the subdividable voxels and mergeable voxels contained in the n-th level spatial grid based on the signed distance function value of each grid vertex in the n-th level spatial grid.

[0076] In some embodiments, after determining the signed distance function value of each mesh vertex, the computer device can determine the positional relationship between each mesh vertex and the object surface based on the signed distance function value, thereby determining whether to subdivide or merge the voxels belonging to the mesh vertex, that is, determining the subdividable voxels and mergeable voxels in the spatial grid.

[0077] Among them, the subdividable voxels refer to voxels that can be subdivided during the adaptive voxel update process, and the mergeable voxels refer to voxels that can be merged during the adaptive voxel update process. The subdividable voxels are determined based on a pre-set subdivision threshold, and the mergeable voxels are determined based on a pre-set merge threshold.

[0078] Optionally, based on the definition of the signed distance function, the computer device can determine the distance between a mesh vertex and the object surface by determining the absolute value of the signed distance function value corresponding to each mesh vertex. When the absolute value is 0, it indicates that the mesh vertex is a point on the object surface. Based on this, the computer device can divide the voxels within the spatial grid through the following steps.

[0079] Step 402A: Based on the signed distance function values ​​of the mesh vertices in the n-th level spatial mesh, determine the minimum absolute value among the absolute values ​​corresponding to the signed distance function values.

[0080] Optionally, the computer device determines the absolute value corresponding to the signed distance function value of each grid vertex in the n-th level spatial grid according to the signed distance function value of each grid vertex, and determines the minimum absolute value from each absolute value.

[0081] Illustratively, when the absolute value corresponding to the SDF value of a mesh vertex is the minimum absolute value, it means that the mesh vertex is the point closest to the three-dimensional surface of the target object.

[0082] Step 402B: if the minimum absolute value is less than the subdivision threshold and the voxel where the mesh vertex corresponding to the minimum absolute value is located has at least two mesh vertices whose signed distance function values ​​have opposite signs, the voxel is determined to be a subdividable voxel.

[0083] The subdivision threshold is a pre-set threshold used to characterize whether a voxel can be used as a subdividable voxel. Optionally, the subdivision threshold can be determined based on at least one factor of the resolution of the spatial grid and the scale of the target object. Different subdivision thresholds can be set for different target objects. Schematically, the subdivision threshold can be expressed as T sub .

[0084] In one possible implementation, when the minimum absolute value is less than the subdivision threshold, it means that the mesh vertex corresponding to the minimum absolute value is closest to the object surface, and when the signs of the signed distance function values ​​of at least two mesh vertices in the voxel where the mesh vertex corresponding to the minimum absolute value is located are opposite, it means that the voxel can pass through the object surface of the target object, so the computer device can determine the voxel as a subdividable voxel.

[0085] In another possible embodiment, in order to improve the accuracy of mesh adaptive updating and subdivide as many voxels near the object surface as possible, the computer device can also determine the subdividable voxels based only on the condition that the minimum absolute value is less than the subdivision threshold, or based only on the condition that the signs of the signed distance function values ​​of at least two mesh vertices in the voxel where the mesh vertex corresponding to the minimum absolute value is located are opposite.

[0086] For example, when the minimum absolute value is less than the subdivision threshold, the computer device determines the voxel where the mesh vertex corresponding to the minimum absolute value is located as a subdividable voxel; or, when the signs of the signed distance function values ​​of at least two mesh vertices in the voxel where the mesh vertex corresponding to the minimum absolute value is located are opposite, the voxel is determined to be a subdividable voxel.

[0087] Alternatively, the signed distance function values ​​of the mesh vertices of the subdividable voxels can satisfy the following formula:

[0088]

[0089] Among them, sdf i Represents the signed distance function value of the i-th mesh vertex of the voxel; |sdfi i | represents the absolute value of the signed distance function value corresponding to the i-th grid vertex of the voxel; The minimum absolute value among the absolute values ​​corresponding to the signed distance function values ​​of the 8 mesh vertices representing the voxel; T sub Represents the subdivision threshold; sign represents the sign of the signed distance function value; & represents simultaneous satisfaction; The signs of the signed distance function values ​​indicating that a voxel exists between at least two mesh vertices are opposite.

[0090] In step 402C, the voxels other than the subdividable voxels are determined as non-subdividable voxels.

[0091] Optionally, after determining the subdividable voxels in the n-th level spatial grid, the computer device may determine other voxels in the n-th level spatial grid except the subdividable voxels as non-subdividable voxels.

[0092] Step 402D: When the absolute value of the signed distance function value of the mesh vertex of the non-subdividable voxel is greater than the merging threshold, the non-subdividable voxel is determined as a mergeable voxel.

[0093] In some embodiments, for each undividable voxel, the computer device may determine whether the undividable voxel is a mergeable voxel based on the absolute value of the signed distance function values ​​of each mesh vertex of the undividable voxel and a merging threshold.

[0094] Optionally, the merge threshold is a pre-set threshold used to characterize whether a voxel can be used as a mergeable voxel. Similar to the subdivision threshold, the merge threshold can be determined based on at least one factor of the resolution of the spatial grid and the scale of the target object. Different merge thresholds can be set for different target objects. Schematically, the merge threshold can be expressed as T merge .

[0095] In one possible implementation, if the absolute value of the signed distance function values ​​of the mesh vertices of an unsubdividable voxel is greater than a merging threshold, indicating that the unsubdividable voxel is far from the surface of the object, the computer device may determine the unsubdividable voxel as a merging voxel. The determination condition that the absolute value of the signed distance function values ​​of the mesh vertices of the unsubdividable voxel is greater than the merging threshold can be that the absolute value of the signed distance function values ​​of all mesh vertices of the unsubdividable voxel is greater than the merging threshold, or that the absolute value of the signed distance function value of at least one mesh vertex in the unsubdividable voxel is greater than the merging threshold, which is not limited in this embodiment of the present application.

[0096] Optionally, the signed distance function values ​​of the mesh vertices of the mergeable voxels can satisfy the following formula:

[0097]

[0098] Among them, sdf i Represents the signed distance function value of the i-th mesh vertex of the voxel; |sdf i | represents the absolute value of the signed distance function value corresponding to the i-th grid vertex of the voxel; The minimum absolute value among the absolute values ​​corresponding to the signed distance function values ​​of the 8 vertices of the voxel; T merge Indicates the merge threshold.

[0099] Step 403 : Subdivide the subdividable voxels and merge the mergeable voxels to obtain an n+1th level spatial grid.

[0100] In some embodiments, after determining the subdividable voxels and mergeable voxels contained in the n-th level spatial grid, the computer device can subdivide the subdividable voxels and merge the mergeable voxels respectively to obtain the n+1-th level spatial grid, so that the voxel distribution density of the three-dimensional surface in the n+1-th level spatial grid is higher than the voxel distribution density of the three-dimensional surface in the n-th level spatial grid, and the voxel distribution density of the non-object surface in the n+1-th level spatial grid is lower than the voxel distribution density of the non-object surface in the n-th level spatial grid.

[0101] Optionally, the subdivision process may be to divide the subdividable voxels into 8 small voxels; and the merging process may be to merge the mergable voxels connected to the same parent node based on the tree-like connection relationship of the voxels.

[0102] In one possible implementation, after obtaining the n+1th level spatial grid, in order to facilitate subsequent feature processing of the mesh vertices within the n+1th level spatial grid, the computer device also needs to re-determine the mesh vertex coordinates, mesh vertex attributes, and position codes for each mesh vertex within the n+1th level spatial grid based on the voxel update status. The position code is used to represent the position of each mesh vertex in the octree, the number of levels, and the multi-layer parent-child relationships between each mesh vertex and other voxels in a tree-like manner.

[0103] Step 404 : determining the voxel partitioning granularity in the process of adaptively updating the n-th level spatial grid.

[0104] In some embodiments, after adaptively updating the n-th level spatial grid, in order to extract the mesh vertex features of the subdivided voxels from the three-plane features, the computer device also needs to upsample the three-plane features.

[0105] Optionally, in order to improve the efficiency of extracting mesh vertex features and to ensure as much as possible that the mesh vertex features of each mesh vertex in the n+1th level spatial grid can be extracted from the three-plane features, the computer device can first determine the voxel division granularity in the process of adaptive voxel update of the nth level spatial grid, that is, determine the voxel subdivision granularity of the subdividable voxels.

[0106] Illustratively, when the subdividable voxel is subdivided into 8 small voxels, the computer device can determine that the voxel subdivision granularity is increased by a multiple of 2, and the upsampling ratio corresponding to the n-th level three-plane feature can be determined to be 2.

[0107] In step 405 , based on the voxel division granularity, upsample the nth level three-plane feature to obtain the n+1th level three-plane feature. The feature resolution of the n+1th level three-plane feature is consistent with the grid resolution of the n+1th level spatial grid.

[0108] In some embodiments, after determining the voxel division granularity corresponding to the n-th level spatial grid, the computer device can upsample the n-th level three-plane features based on the voxel division granularity to obtain the n+1-th level three-plane features, so that the feature resolution of the n+1-th level three-plane features is consistent with the grid resolution of the n+1-th level spatial grid.

[0109] Schematically, when the voxel subdivision granularity increases by a multiple of 2, the computer device can perform a two-fold upsampling process on the n-th level three-plane features, that is, perform a two-fold upsampling process on the features on the two-dimensional plane where the three axes intersect, thereby obtaining the n+1-th level three-plane features.

[0110] Step 406 : extracting mesh vertex features of each mesh vertex in the n+1th level spatial mesh from the n+1th level three-plane features.

[0111] The specific implementation of step 406 can refer to step 203, which will not be described in detail in this embodiment.

[0112] Step 407 : Based on the mesh vertex features of each mesh vertex in the n+1th level spatial mesh, the target object is reconstructed through a mesh reconstruction network to obtain the n+1th level patch mesh corresponding to the target object.

[0113] In some embodiments, after obtaining the mesh vertex features of each mesh vertex in the n+1th level spatial mesh, the computer device can use the mesh reconstruction network to reconstruct the mesh of the target object, thereby generating the n+1th level patch mesh corresponding to the target object.

[0114] Optionally, in order to ensure that the reconstructed patch mesh has the same geometric features as the target object, the mesh reconstruction network may include a first multilayer perceptron, so that the geometric features of the target object are learned through the first multilayer perceptron.

[0115] In one possible implementation, a computer device inputs mesh vertex features of each mesh vertex in the n+1th level spatial mesh into a first multilayer perceptron in a mesh reconstruction network, and outputs signed distance function values ​​of each mesh vertex in the n+1th level spatial mesh through the first multilayer perceptron. The signed distance function values ​​are used to characterize the correspondence between the mesh vertices and the object surface in three-dimensional space, so that the computer device can generate the n+1th level patch mesh corresponding to the target object using a marching cubes algorithm based on the signed distance function values ​​of each mesh vertex.

[0116] Optionally, the process of outputting the signed distance function values ​​of the mesh vertices through the first multilayer perceptron can be expressed as s=MLP geom (f(v)), where the mesh vertex is represented as v and the signed distance function value is represented as s.

[0117] Optionally, considering that in an embodiment of the present application, after the spatial grid corresponding to each level of resolution is adaptively updated, the voxels in the spatial grid are in an unevenly distributed state, in order to reconstruct the grid based on the unevenly distributed voxels, the computer device can determine the dual vertices corresponding to each voxel and construct a dual grid to achieve grid reconstruction of the target object. This process can be implemented based on the following steps.

[0118] Step 407A: Determine the dual vertex corresponding to each voxel based on the position code of the grid vertex corresponding to each voxel in the (n+1)th level spatial grid.

[0119] In one possible implementation, considering that the voxels in the spatial grid in the embodiment of the present application are connected in a tree-like manner, that is, the grid vertex of each voxel corresponds to a unique position code, the position code can characterize the position of the grid vertex in the octree, the number of layers, and the multi-layer parent-child relationship connected in a tree-like manner with other voxels, so that the computer device can determine the dual vertex corresponding to each voxel based on the position code of the grid vertex of each voxel in the n+1th level spatial grid.

[0120] Optionally, each voxel corresponds to a dual vertex, where the dual vertex is a point obtained by dualizing the voxel to a point. Optionally, the dual vertex can be the voxel center point of the voxel, or can be an interpolation point obtained by performing an interpolation operation based on the signed distance function values ​​of each mesh vertex of the voxel, which is not limited in this embodiment of the application.

[0121] Step 407B: Connect the dual vertex corresponding to each voxel with the dual vertex of at least one adjacent voxel to obtain the (n+1)th level dual grid corresponding to the target object, where the adjacent voxels are other voxels that have common vertices with the voxel.

[0122] In one possible implementation, after determining the dual vertices corresponding to each voxel in the n+1th level spatial grid, the computer device can connect the dual vertices corresponding to each voxel with the dual vertices of at least one adjacent voxel, thereby obtaining the n+1th level dual grid corresponding to the target object.

[0123] Optionally, when two voxels have a common mesh vertex, the two voxels can be determined as adjacent voxels, that is, the adjacent voxels of a voxel refer to other voxels that have a common mesh vertex with the voxel. In the embodiment of the present application, since the voxels are connected using an octree data structure, a voxel can have a maximum of 8 adjacent voxels.

[0124] Optionally, a dual grid is a grid structure formed by connecting dual vertices.

[0125] Step 407C: Based on the signed distance function value of each mesh vertex in the (n+1)th level dual mesh, the marching cubes algorithm is used to map and obtain the isosurface of the (n+1)th level dual mesh.

[0126] Unlike related technologies, in which the spatial grid is composed of uniformly distributed spatial tetrahedrons, the marching cubes (MC) algorithm can be directly used to reconstruct the grid after extracting the grid vertex features. In the embodiment of the present application, since the voxels in the spatial grid are unevenly distributed, it is necessary to determine the dual vertices corresponding to each voxel and generate a dual grid, thereby using the Lewiner marching cube algorithm (MC33 algorithm) to map the isosurface of the dual grid based on the signed distance function value of each grid vertex in the dual grid.

[0127] In one possible implementation, after obtaining the n+1th level dual grid corresponding to the target object, the computer device can map the isosurface of the n+1th level dual grid according to the signs of the signed distance function values ​​of each grid vertex in the n+1th level dual grid, wherein the signed distance function value corresponding to the point on the isosurface is 0.

[0128] Step 407D: Generate the n+1th level patch mesh corresponding to the target object based on the isosurface of the n+1th level dual mesh.

[0129] In one possible implementation, after obtaining the isosurface of the n+1th level dual grid, since the signed distance function value corresponding to the points on the isosurface is 0, the isosurface can be determined as the three-dimensional surface of the patch grid, thereby generating the n+1th level patch grid corresponding to the target object.

[0130] Step 408 : Based on the n+1th level patch mesh, perform differentiable rendering through a differentiable rendering network to obtain a rendering prediction image of the n+1th level patch mesh at each camera pose.

[0131] In some embodiments, after obtaining the n+1th level facet mesh, the geometric features of the target object are determined. In order to improve the reconstruction quality of the facet mesh, the computer device also needs to train and optimize the mesh reconstruction network through reconstruction loss. Since it is difficult to obtain the real three-dimensional structural data of the target object, the computer device can convert the three-dimensional facet mesh into a two-dimensional image, and generate a rendering prediction image of the n+1th level facet mesh under the corresponding camera pose according to the camera pose corresponding to the real shot image of the target object, thereby replacing the reconstruction loss with the prediction loss between the real shot image and the rendering prediction image.

[0132] Optionally, the real shot image is obtained by shooting the target object based on various camera positions. In addition, in order to improve the quality of 3D reconstruction, the computer device needs to obtain the real shot images corresponding to the target object at as many angles as possible, for example, using a camera to shoot the target object in all directions. Figure 5As shown in FIG, by shooting the target object from multiple angles, the real shooting images of the target object under multiple camera postures can be obtained.

[0133] In one possible implementation, after obtaining the n+1th level patch mesh, the computer device can perform differentiable rendering through a differentiable rendering network according to the camera pose corresponding to each real shot image, and obtain a rendering prediction image of the n+1th level patch mesh under each camera pose. This process can be implemented based on the following steps.

[0134] Step 408A, based on the camera pose, multi-view rasterization processing is performed on the n+1th level patch mesh to obtain a dense pixel map of the n+1th level patch mesh at each camera pose, where the dense pixel map represents the surface vertex information of the surface vertices located on the mesh patch at the camera pose.

[0135] In one possible implementation, after obtaining the n+1th level patch grid, in order to determine the mesh surface contained in the predicted rendering image under different camera poses, the computer device can perform multi-perspective rasterization processing on the n+1th level patch grid according to the camera pose, first obtaining a sparse pixel map of the n+1th level patch grid under each camera pose, and then interpolating the sparse pixel map through barycentric coordinate interpolation to obtain a dense pixel map under each camera pose.

[0136] The dense pixel map represents surface vertex information of surface vertices located on a mesh of a patch mesh at a camera pose, wherein the surface vertex information may be information about the mesh patch to which the surface vertex belongs. The sparse pixel map represents mesh vertex information of mesh vertices located on a mesh patch of a patch mesh at a camera pose, i.e., a computer device performs rasterization processing on the patch mesh at each camera pose, and obtains three mesh vertex information corresponding to each mesh patch at the camera pose, wherein the mesh vertex information may be information about the mesh patch to which the mesh vertex belongs.

[0137] Step 408B: determining surface vertex features of surface vertices based on the dense pixel map and the (n+1)th level patch mesh.

[0138] In one possible implementation, after obtaining the dense pixel map, the computer device can back-project the dense pixel map to the n+1-th level patch mesh, and then obtain the surface vertex coordinates of each surface vertex based on the projected n+1-th level patch mesh, thereby extracting the surface vertex features of each surface vertex from the n+1-th level three-plane features based on the surface vertex coordinates.

[0139] Optionally, a surface vertex can be represented as v c , the surface vertex features can be expressed as

[0140] Step 408C: Input the surface vertex features into the second multi-layer perceptron in the differentiable rendering network, and output the vertex color features of the surface vertices through the second multi-layer perceptron.

[0141] In one possible implementation, in order to generate a rendering prediction graph, after determining the surface vertex features of each surface vertex, the computer device can input the surface vertex features into a second multi-layer perceptron in the differentiable rendering network, and output the vertex color features of the surface vertices through the second multi-layer perceptron.

[0142] Optionally, the process of outputting vertex color features through the second multi-layer perceptron can be expressed as color=MLP color (f(v c )).

[0143] Step 408D: Generate a rendering prediction map of the n+1th level patch mesh at each camera pose based on the vertex color features.

[0144] In one possible implementation, after obtaining the vertex color features of each surface vertex at each camera pose, the computer device can obtain a rendering prediction image of the n+1th level face mesh at each camera pose through differentiable rendering based on the vertex color features of the surface vertices.

[0145] Step 409 : Based on the prediction loss between the rendering prediction image and the real shot image, the n+1th level three-plane features, the mesh reconstruction network, and the differentiable rendering network are trained. The real shot image is obtained by shooting the target object based on various camera poses.

[0146] In some embodiments, after obtaining the rendering prediction images corresponding to different camera poses, the computer device can calculate the prediction loss based on the rendering prediction image and the real shooting image, and then train the n+1-th level three-plane features, mesh reconstruction network and differentiable rendering network based on the prediction loss, and update the feature parameters in the n+1-th level three-plane features, mesh reconstruction network and differentiable rendering network.

[0147] Optionally, the prediction loss can be the color difference of each pixel between the rendered prediction image and the real captured image, or it can be a data calculation result based on the color difference, such as the norm calculation result of the color difference, etc. The embodiment of the present application is not limited to this.

[0148] In some embodiments, in order to improve the efficiency of three-dimensional reconstruction at various levels of resolution, the computer device needs to continue to re-extract the mesh vertex features of the mesh vertices in the n+1th spatial mesh based on the updated n+1th level three-plane features, mesh reconstruction network, and differentiable rendering network, and perform mesh reconstruction and differentiable rendering, and continue to update the parameters of the n+1th level three-plane features, mesh reconstruction network, and differentiable rendering network based on the new prediction loss. That is, the n+1th level three-plane features, mesh reconstruction network, and differentiable rendering network need to undergo multiple rounds of training to minimize the prediction loss.

[0149] In one possible implementation, the computer device can first calculate the pixel prediction difference between the rendering prediction image and the real shooting image based on the rendering prediction image and the real shooting image corresponding to each camera pose, and then calculate the norm of the pixel prediction difference and accumulate and sum the calculation results to obtain the predicted total loss, so as to train the n+1-th level three-plane features, mesh reconstruction network and differentiable rendering network based on the predicted total loss.

[0150] Optionally, the real shot image can be represented as I gt , the predicted rendering can be expressed as I pred , so the pixel prediction difference between the predicted rendering image and the real shooting image can be expressed as I pred -I gt , the camera pose can be expressed as T, the number of real shots is N, and the total prediction loss can be expressed as

[0151] In addition, considering that the 3D reconstruction process is divided into 3D reconstruction sub-processes corresponding to multiple levels of resolution in the embodiment of the present application, and the real photographic images of the target object at each camera pose are fixed-resolution images, the prediction loss between the rendering prediction image and the real photographic image generated by the 3D reconstruction process at low resolution is significantly greater than the prediction loss between the rendering prediction image and the real photographic image generated in the 3D reconstruction process at high resolution. Moreover, since the resolution itself is low, the prediction loss cannot be effectively reduced even through multiple rounds of iterative optimization. Therefore, in order to improve the 3D reconstruction efficiency at each level of resolution, the computer device can also downsample the real photographic image to obtain the real photographic image corresponding to each level of resolution.

[0152] In one possible implementation, for the three-dimensional reconstruction process corresponding to the n+1th level resolution, the computer device can obtain the n+1th level real shooting image by downsampling the real shooting image, so that the image resolution of the n+1th level real shooting image is consistent with the grid resolution of the n+1th level spatial grid. Then, in each round of training, the computer device can calculate the pixel prediction difference of the current training round based on the rendering prediction image corresponding to each camera pose and the n+1th level real shooting image.

[0153] Step 410 , when the training is completed, based on the mesh vertex features of each mesh vertex in the n+1-th level spatial mesh, the target object is reconstructed through the mesh reconstruction network and the differentiable rendering network to obtain the n+1-th level 3D model.

[0154] In some embodiments, for the three-dimensional reconstruction process corresponding to a single resolution, when the training is completed, the computer device can extract the mesh vertex features of each mesh vertex of the n+1-level spatial grid from the n+1-level three-plane features after parameter update, and use the trained mesh reconstruction network to perform mesh reconstruction based on the mesh vertex features of each mesh vertex to obtain the n+1-level face mesh, and then use the differentiable rendering network to output the vertex color features of each surface vertex and mesh vertex in the n+1-level face mesh, and generate the n+1-level three-dimensional model corresponding to the target object through rendering based on the vertex color features.

[0155] Optionally, considering that in the embodiment of the present application, the three-dimensional reconstruction process is divided into three-dimensional reconstruction sub-processes corresponding to multiple levels of resolution, and the number of iterative optimizations required in the three-dimensional reconstruction processes corresponding to different resolutions may be different, for example, for a low-resolution three-dimensional reconstruction process, since the resolution requirement is low, the number of iterative optimizations can be appropriately reduced; and for a high-resolution three-dimensional reconstruction process, since the resolution requirement is high, the number of iterative optimizations can be appropriately increased.

[0156] Optionally, for the three-dimensional reconstruction processes corresponding to different resolutions, the computer device can set different training completion conditions, such as different loss thresholds corresponding to different resolutions, or different iteration thresholds corresponding to different resolutions.

[0157] For the three-dimensional reconstruction process corresponding to the n+1th level resolution, in one possible implementation, the computer device can set the n+1th level loss threshold, and the n+1th level loss threshold is less than the nth level loss threshold, so that when the prediction loss between the rendered prediction image and the real shot image meets the n+1th level loss threshold, the computer device reconstructs the model of the target object based on the mesh vertex features of each mesh vertex in the n+1th level spatial grid through the mesh reconstruction network and the differentiable rendering network to obtain the n+1th level three-dimensional model.

[0158] In another possible embodiment, the computer device can set an n+1-level iteration threshold, and the n+1-level iteration threshold is greater than the n-level iteration threshold, so that when the number of training iterations meets the n+1-level iteration threshold, the computer device reconstructs the model of the target object based on the mesh vertex features of each mesh vertex in the n+1-level spatial grid through the mesh reconstruction network and the differentiable rendering network to obtain the n+1-level three-dimensional model.

[0159] In the above embodiment, in the process of adaptive voxel updating of the spatial grid, the voxels in the spatial grid are divided according to the signed distance function value of each grid vertex to obtain subdividable voxels located near the object surface and mergeable voxels located near the non-object surface, so that the subdividable voxels are subdivided and the mergeable voxels are merged, thereby achieving the goal of increasing the voxel distribution density on the object surface in the spatial grid and reducing the voxel distribution density on the non-object surface in the spatial grid in the process of gradually increasing the resolution, thereby achieving the goal of reducing the amount of calculation for three-dimensional reconstruction while improving the model accuracy and efficiency of the three-dimensional model.

[0160] Furthermore, considering that the voxels in the spatial grid are unevenly distributed as the voxels are continuously updated adaptively, in the process of grid reconstruction for a single-level resolution, the embodiment of the present application adopts a method of generating a dual grid and uses the MC33 algorithm to perform grid reconstruction, thereby obtaining the patch grid corresponding to the target object, thereby improving the efficiency and accuracy of grid reconstruction for the target object.

[0161] In addition, during the multiple iterative optimization processes for a single level of resolution, considering that the real-life photographic image corresponding to the target object has a fixed resolution, while there is a significant difference in the image resolution of the rendered prediction images generated at different levels of resolution, the real-life photographic image is downsampled for different levels of resolution. Based on the downsampled real-life photographic image and the rendered prediction image, the prediction loss in each iterative optimization process is calculated, thereby improving the accuracy of the loss calculation, optimizing the iterative process for each level of resolution, and improving the quality of the three-dimensional model corresponding to each level of resolution.

[0162] In some embodiments, when multiple levels of resolution are divided, the computer device repeats the steps in the above embodiment for each level of resolution, and thus can sequentially obtain a three-dimensional model corresponding to each level of resolution. Figure 6As shown, when there are four levels of resolution, the computer device can obtain a first-level three-dimensional model 601, a second-level three-dimensional model 602, a third-level three-dimensional model 603 and a fourth-level three-dimensional model 604 in sequence by executing the three-dimensional reconstruction method proposed in the embodiment of the present application, and the model resolution of the three-dimensional model is gradually improved.

[0163] For example, the resolution is increased from [3, C, H1, W1] to [3, C, Hm, Wm] by multiples of 2, Hm = H1 × 2 m-1 , where H1=W1=16, m=4, Hm=Wm=128, and C=32 are taken as an example. In the embodiment of the present application, in order to obtain a three-dimensional model of [3, 32, 128, 128], firstly, the first-level spatial grid with a resolution of [3, 32, 16, 16] and the three-plane features corresponding to the first-level spatial grid are initialized, and then a three-dimensional model with a resolution of [3, 32, 16, 16] is obtained through geometric iteration and differentiable rendering optimization. Then, from coarse to fine, after gradually completing the optimization of each level of resolution, the following can be obtained. Figure 7 The resolution shown is [3, 32, 128, 128] and the three-dimensional model 702. It can be seen that in the embodiment of the present application, through the step-by-step optimization of the resolution, when reconstructing the three-dimensional model with a resolution of [3, 32, 128, 128], it is only necessary to perform feature extraction, geometric iteration and differentiable rendering based on the spatial grid after adaptive voxel update, that is, it is only necessary to process the refined mesh vertices of the three-dimensional surface and other merged mesh vertices. Compared with directly constructing a spatial grid with a resolution of [3, 32, 128, 128], making the mesh vertex distribution density of the three-dimensional surface and the non-three-dimensional surface in the three-dimensional space the same, and then performing geometric iteration and differentiable rendering to generate the amount of data, the amount of data generated by the step-by-step iterative optimization from low resolution to high resolution in the embodiment of the present application is significantly reduced.

[0164] Please refer to Figure 8 , which shows a flow chart of a three-dimensional reconstruction method provided by an exemplary embodiment of the present application.

[0165] First, in order to perform three-dimensional reconstruction of the target object, the computer device first needs to initialize a low-resolution spatial grid and three-plane features, namely the first-level spatial grid (Octgrid_L) and the first-level three-plane features (Tri-plane_L). For the three-dimensional reconstruction process of the first-level resolution, the computer device extracts the mesh vertex features of each mesh vertex in the first-level spatial grid from the first-level three-plane features, and based on the mesh vertex features, performs geometric iteration and differentiable rendering optimization until multiple rounds of optimization of the first-level resolution are completed. After obtaining the first-level three-dimensional model, the computer device adaptively updates the first-level spatial grid with voxels and upsamples the first-level three-plane features to obtain the second-level spatial grid (Octgrid_L+1) and the second-level three-plane features (Tri-plane_L+1), and then performs the three-dimensional reconstruction process of the second-level resolution.

[0166] Furthermore, by completing the three-dimensional reconstruction process for each level of resolution, the three-dimensional model corresponding to each level of resolution can be obtained. When the maximum resolution is reached, the three-dimensional model of the highest level of resolution corresponding to the target object can be obtained, thereby completing the three-dimensional reconstruction of the target object.

[0167] Please refer to Figure 9 , which shows a flowchart of a single-resolution corresponding three-dimensional reconstruction method provided by an exemplary embodiment of the present application.

[0168] Step 901: Determine the grid vertices of the (n+1)th level spatial grid.

[0169] First, the computer device obtains the n+1th level spatial grid by performing adaptive voxel update on the nth level spatial grid, and determines the grid vertex coordinates, vertex attributes, position coding and other information of each grid vertex in the n+1th level spatial grid based on the adaptive voxel update process.

[0170] Step 902 : extracting mesh vertex features of mesh vertices from the n+1th level three-plane features.

[0171] To extract higher-resolution mesh vertex features, the computer device also needs to upsample the n-th level tri-planar features based on the voxel update granularity of the adaptive voxel update for the n-level spatial grid to obtain the n+1-th level tri-planar features 912. When performing adaptive voxel updates on a level-by-level basis, the voxel update granularity increases by a factor of 2, i.e., the computer device upsamples the n-th level tri-planar features by a factor of 2, thereby obtaining the n+1-th level tri-planar features 912.

[0172] After obtaining the n+1th level three-plane feature 912, the computer device can extract the mesh vertex features of each mesh vertex in the n+1th level spatial grid from the n+1th level three-plane feature 912 according to the mesh vertex coordinates of each mesh vertex in the n+1th level spatial grid.

[0173] Step 903: Input the mesh vertex features into the MLP network and output the SDF value.

[0174] In order to reconstruct the mesh of the target object, the computer device can input the mesh vertex features of each mesh vertex in the n+1th level spatial grid into the MLP network, and output the SDF value of each mesh vertex through the MLP network.

[0175] Step 904: Obtain the n+1th level facet mesh through mesh reconstruction.

[0176] After obtaining the SDF value of each grid vertex, the computer device can determine the dual vertex of each voxel in the n+1-level spatial grid, connect the dual vertex corresponding to each voxel with the dual vertex of at least one adjacent voxel, and obtain the n+1-level dual grid corresponding to the target object. Then, based on the signed distance function value of each grid vertex in the n+1-level dual grid, the marching cube algorithm is used to map the isosurface of the n+1-level dual grid, and then generate the n+1-level patch grid based on the isosurface of the n+1-level dual grid.

[0177] Step 905 , performing multi-view rasterization processing on the n+1th level facet mesh.

[0178] After obtaining the n+1th level patch mesh, the geometric features of the target object are determined. In order to improve the reconstruction quality of the patch mesh, the computer equipment also needs to train and optimize the mesh reconstruction network through reconstruction loss. Therefore, in order to determine the prediction loss, the computer equipment needs to perform multi-view rasterization processing on the n+1th level patch mesh to obtain a sparse pixel map of the n+1th level patch mesh under each camera pose.

[0179] Step 906 : Obtain the surface vertices on the mesh patch of the (n+1)th level mesh through interpolation processing.

[0180] Furthermore, after obtaining the sparse pixel map of the n+1th level patch mesh at each camera pose, the computer device can interpolate the sparse pixel map by barycentric coordinate interpolation to obtain a dense pixel map at each camera pose, which represents the surface vertex information of the surface vertices located on the mesh patch of the patch mesh at the camera pose.

[0181] Step 907 : Obtain the surface vertex coordinates of each surface vertex by back-projecting to the n+1th level facet mesh.

[0182] After obtaining the dense pixel map, the computer device can back-project the dense pixel map to the n+1th level patch grid, and then obtain the surface vertex coordinates of each surface vertex based on the projected n+1th level patch grid.

[0183] Step 908 : extracting surface vertex features of surface vertices from the n+1th level three-plane features.

[0184] Based on the surface vertex coordinates of each surface vertex, the computer device can extract the surface vertex features of each surface vertex from the (n+1)th level three-plane features 912 .

[0185] Step 909: Obtain vertex color features through the MLP network.

[0186] In order to generate a rendering prediction image, after determining the surface vertex features of each surface vertex, the computer device can input the surface vertex features into the MLP network in the differentiable rendering network, output the vertex color features of the surface vertices through the MLP network, and obtain the rendering prediction image of the n+1th level patch mesh under various camera poses through differentiable rendering.

[0187] Step 910: Obtain the real shot images under each camera posture.

[0188] To determine the prediction loss during the multiple iterations of optimization at the n+1 resolution level, the computer needs to obtain the real-world images at each camera pose. To improve the accuracy of the prediction loss, the computer can downsample the real-world images so that the rendered predictions and the real-world images have the same resolution.

[0189] Step 911: Determine the prediction loss based on the rendered prediction image and the real captured image.

[0190] According to the pixel prediction difference between the rendering prediction image and the real shooting image under each camera pose, the norm of the pixel prediction difference is calculated and the calculation results are accumulated and summed to obtain the total prediction loss of the current iteration round. Based on the total prediction loss, the n+1-th level three-plane feature 912, the mesh reconstruction network and the differentiable rendering network are trained until the training completion condition is met, that is, the n+1-th level three-dimensional model corresponding to the n+1-th level resolution can be generated.

[0191] Please refer to Figure 10 , which shows a flow chart of a three-dimensional reconstruction method provided by an exemplary embodiment of the present application.

[0192] In order to reconstruct the target object in three dimensions, the computer equipment divides the resolution into m levels from coarse to fine, and increases the resolution by a multiple of 2, so that the model resolution of the first-level three-dimensional model is [3, C, H0, W0], and the model resolution of the last-level three-dimensional model is [3, C, H×2 m-1 ,W×2 m-1 By performing iterative optimization at m levels of resolution, the computer device can obtain an m-level LOD 3D model. The iterative optimization process of the mth level is the most refined, but it has more meshes and consumes more computing resources. The iterative optimization process of the first level is the coarsest, but it has fewer meshes and consumes fewer computing resources. The iterative optimization process of the intermediate level is somewhere in between.

[0193] Please refer to Figure 11 , which shows a structural block diagram of a three-dimensional reconstruction device provided by an exemplary embodiment of the present application, the device includes:

[0194] a voxel updating module 1101 configured to adaptively update the nth-level spatial grid corresponding to the target object to obtain an n+1th-level spatial grid, wherein the density of voxel distribution of the three-dimensional surface in the n+1th-level spatial grid is higher than the density of voxel distribution of the three-dimensional surface in the nth-level spatial grid;

[0195] A sampling processing module 1102 is configured to perform upsampling processing on the n-th level three-plane features corresponding to the n-th level spatial grid to obtain the n+1-th level three-plane features;

[0196] A feature extraction module 1103 is configured to extract mesh vertex features of each mesh vertex in the n+1-th level spatial mesh from the n+1-th level three-plane features;

[0197] The model generation module 1104 is used to obtain the n+1th level three-dimensional model corresponding to the target object based on the mesh vertex features of each mesh vertex in the n+1th level spatial mesh through geometric iteration and differentiable rendering optimization. The model resolution of the n+1th level three-dimensional model is higher than the model resolution of the nth level three-dimensional model. The geometric iteration is used to generate the n+1th level patch mesh corresponding to the target object, and the differentiable rendering is used to perform differentiable rendering on the n+1th level patch mesh.

[0198] Optionally, the sampling processing module 1102 is configured to:

[0199] Determining a voxel partitioning granularity in a process of adaptively updating the n-th level spatial grid;

[0200] Based on the voxel division granularity, the nth level three-plane feature is upsampled to obtain the n+1th level three-plane feature, and the feature resolution of the n+1th level three-plane feature is consistent with the grid resolution of the n+1th level spatial grid.

[0201] Optionally, the model generation module 1104 includes:

[0202] A mesh reconstruction unit, configured to reconstruct the mesh of the target object through a mesh reconstruction network based on mesh vertex features of each mesh vertex in the n+1th level spatial mesh, to obtain an n+1th level patch mesh corresponding to the target object;

[0203] A differentiable rendering unit is configured to perform differentiable rendering based on the n+1th level patch mesh through a differentiable rendering network to obtain a rendering prediction image of the n+1th level patch mesh under various camera poses;

[0204] a training unit, configured to train the n+1th level three-plane features, the mesh reconstruction network, and the differentiable rendering network based on a prediction loss between the rendering prediction image and a real shot image obtained by photographing the target object based on various camera poses;

[0205] A model generation unit is used to, after training is completed, reconstruct a model of the target object based on the mesh vertex features of each mesh vertex in the n+1-th level spatial mesh through the mesh reconstruction network and the differentiable rendering network to obtain the n+1-th level three-dimensional model.

[0206] Optionally, the training unit is used to:

[0207] Based on the rendering prediction image and the real shooting image corresponding to each camera pose, calculating the pixel prediction difference between the rendering prediction image and the real shooting image;

[0208] Performing norm calculation and accumulation processing on the pixel prediction differences to obtain a total prediction loss;

[0209] Based on the predicted total loss, the n+1th level three-plane features, the mesh reconstruction network, and the differentiable rendering network are trained.

[0210] Optionally, the training unit is further configured to:

[0211] Downsampling the real photographic image to obtain an n+1th level real photographic image, where the image resolution of the n+1th level real photographic image is consistent with the grid resolution of the n+1th level spatial grid;

[0212] The pixel prediction difference is calculated based on the rendering prediction image corresponding to each camera pose and the n+1th level real shooting image.

[0213] Optionally, the model generating unit is used to:

[0214] When the prediction loss between the rendered prediction image and the real photographed image satisfies the n+1th level loss threshold, based on the mesh vertex features of each mesh vertex in the n+1th level spatial mesh, the target object is model reconstructed by the mesh reconstruction network and the differentiable rendering network to obtain the n+1th level three-dimensional model, and the n+1th level loss threshold is less than the nth level loss threshold; or

[0215] When the number of training iterations meets the n+1th level iteration threshold, the target object is modeled reconstructed through the mesh reconstruction network and the differentiable rendering network based on the mesh vertex features of each mesh vertex in the n+1th level spatial grid to obtain the n+1th level three-dimensional model, and the n+1th level iteration threshold is greater than the nth level iteration threshold.

[0216] Optionally, the grid reconstruction unit is used to:

[0217] Inputting mesh vertex features of each mesh vertex in the n+1th level spatial mesh into a first multilayer perceptron in the mesh reconstruction network, and outputting a signed distance function value of each mesh vertex in the n+1th level spatial mesh through the first multilayer perceptron, wherein the signed distance function value is used to represent a correspondence between the mesh vertex and the object surface in three-dimensional space;

[0218] Based on the signed distance function value of each mesh vertex, a marching cubes algorithm is used to generate an n+1th level patch mesh corresponding to the target object.

[0219] Optionally, the grid reconstruction unit is further configured to:

[0220] Determining a dual vertex corresponding to each voxel based on a position code of a grid vertex corresponding to each voxel in the n+1th level spatial grid;

[0221] Connecting the dual vertex corresponding to each voxel with the dual vertex of at least one adjacent voxel to obtain an n+1th level dual grid corresponding to the target object, where the adjacent voxels are other voxels that have a common vertex with the voxel;

[0222] Based on the signed distance function value of each grid vertex in the n+1-th level dual grid, the marching cube algorithm is used to map and obtain the isosurface of the n+1-th level dual grid;

[0223] Based on the isosurface of the n+1th level dual grid, the n+1th level patch grid corresponding to the target object is generated.

[0224] Optionally, the differentiable rendering unit is used to:

[0225] Based on the camera pose, performing multi-view rasterization processing on the n+1th level patch mesh to obtain a dense pixel map of the n+1th level patch mesh at each camera pose, wherein the dense pixel map represents surface vertex information of surface vertices located on the mesh patch at the camera pose;

[0226] determining surface vertex features of the surface vertices based on the dense pixel map and the n+1th level patch mesh;

[0227] Inputting the surface vertex features into a second multilayer perceptron in the differentiable rendering network, and outputting vertex color features of the surface vertices through the second multilayer perceptron;

[0228] Based on the vertex color features, the rendering prediction graph of the n+1th level patch mesh under each camera pose is generated.

[0229] Optionally, the voxel updating module 1101 is configured to:

[0230] Determining a signed distance function value of each grid vertex in the n-th level spatial grid;

[0231] Determining, based on signed distance function values ​​of respective grid vertices in the n-th level spatial grid, subdividable voxels and mergeable voxels contained in the n-th level spatial grid;

[0232] The subdividable voxels are subdivided, and the mergeable voxels are merged to obtain the n+1th level spatial grid.

[0233] Optionally, the voxel updating module 1101 is further configured to:

[0234] Determining, based on the signed distance function values ​​of the grid vertices in the n-th level spatial grid, a minimum absolute value among the absolute values ​​corresponding to the signed distance function values;

[0235] When the minimum absolute value is smaller than the subdivision threshold and the voxel where the mesh vertex corresponding to the minimum absolute value is located has at least two mesh vertices whose signed distance function values ​​have different signs, the voxel is determined as the subdividable voxel;

[0236] determining other voxels except the subdividable voxels as non-subdividable voxels;

[0237] When the absolute value of the signed distance function value of the mesh vertex of the non-subdividable voxel is greater than a merging threshold, the non-subdividable voxel is determined as the mergable voxel.

[0238] To sum up, in the embodiment of the present application, in order to achieve three-dimensional reconstruction of the target object, by setting iterative optimization corresponding to multiple levels of resolution, the first-level spatial grid and the first-level three-plane features corresponding to the target object are constructed starting from low resolution, and the first-level three-dimensional model is obtained through geometric iteration and differentiable rendering optimization, so that the first-level spatial grid is adaptively updated, so that the voxel distribution density of the three-dimensional surface in the spatial grid is gradually increased, and at the same time, the first-level three-plane features are upsampled, so that the three-plane features are gradually refined, and then based on the second-level spatial grid and the second-level three-plane features, the second-level three-dimensional model is obtained through geometric iteration and differentiable rendering optimization, so that the model resolution of the three-dimensional model is gradually improved, and an optimization reconstruction process from low resolution to high resolution and from coarse to fine is realized. While ensuring the quality of the three-dimensional model reconstruction, the amount of data in the three-dimensional model reconstruction process is reduced, and the three-dimensional reconstruction efficiency is improved.

[0239] It should be noted that the apparatus provided in the above embodiments is merely exemplified by the division of the above functional modules. In actual applications, the above functions can be distributed among different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.

[0240] It should be noted that before collecting relevant user data such as real-life photos and during the process of collecting relevant user data such as target audio and video content, this application can display a prompt interface, pop-up window or output voice prompt information. The prompt interface, pop-up window or voice prompt information is used to remind the user that its relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining user-related data after obtaining the user's confirmation operation on the prompt interface or pop-up window. Otherwise (that is, when the user's confirmation operation on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining user-related data are terminated, that is, the user's relevant data is not obtained. In other words, the information involved in this application (including but not limited to user device information, user personal information, etc., user corresponding operation data), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the real-life photos and other data involved in this application are all obtained with full authorization.

[0241] Please refer to Figure 12 , which shows a schematic diagram of the structure of a computer device provided by an exemplary embodiment of the present application. Specifically, the computer device 1200 includes a central processing unit (CPU) 1201, a system memory 1204 including a random access memory 1202 and a read-only memory 1203, and a system bus 1205 connecting the system memory 1204 and the CPU 1201. The computer device 1200 also includes a basic input / output system (I / O system) 1206 that helps transmit information between various components within the computer, and a mass storage device 1207 for storing an operating system 1213, application programs 1214, and other program modules 1215.

[0242] The basic input / output system 1206 includes a display 1208 for displaying information and an input device 1209, such as a mouse and keyboard, for user input. Both the display 1208 and the input device 1209 are connected to the central processing unit 1201 via an input / output controller 1210 connected to the system bus 1205. The basic input / output system 1206 may also include an input / output controller 1210 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1210 also provides output to a display screen, printer, or other types of output devices.

[0243] The mass storage device 1207 is connected to the central processing unit 1201 via a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer-readable media provide non-volatile storage for the computer device 1200. In other words, the mass storage device 1207 may include a computer-readable medium (not shown) such as a hard disk or drive.

[0244] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cassette, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above-mentioned ones. The above-mentioned system memory 1204 and mass storage device 1207 can be collectively referred to as memory.

[0245] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 1201. The one or more programs contain instructions for implementing the above-mentioned method. The central processing unit 1201 executes the one or more programs to implement the three-dimensional reconstruction method provided by the above-mentioned various method embodiments.

[0246] According to various embodiments of the present application, the computer device 1200 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1200 may be connected to the network 1211 via the network interface unit 1212 connected to the system bus 1205, or the network interface unit 1212 may be used to connect to other types of networks or remote computer systems (not shown).

[0247] An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer instruction is stored. The at least one computer instruction is loaded and executed by a processor to implement the three-dimensional reconstruction method described in the above embodiment.

[0248] Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSDs) or optical disks, etc. Among them, RAM may include resistance random access memory (ReRAM) and dynamic random access memory (DRAM).

[0249] The present invention provides a computer program product comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the three-dimensional reconstruction method described in the above embodiment.

[0250] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0251] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A three-dimensional reconstruction method, characterized in that: The method comprises: Adaptively updating voxels on an n-th level spatial grid corresponding to the target object to obtain an n+1-th level spatial grid, wherein a density of voxel distribution of the three-dimensional surface in the n+1-th level spatial grid is higher than a density of voxel distribution of the three-dimensional surface in the n-th level spatial grid; Upsampling the n-th level three-plane features corresponding to the n-th level spatial grid to obtain the n+1-th level three-plane features; Extracting mesh vertex features of each mesh vertex in the n+1th level spatial mesh from the n+1th level three-plane features; Based on the mesh vertex features of each mesh vertex in the n+1-level spatial mesh, through geometric iteration and differentiable rendering optimization, an n+1-level three-dimensional model corresponding to the target object is obtained, and the model resolution of the n+1-level three-dimensional model is higher than the model resolution of the n-level three-dimensional model. The geometric iteration is used to generate the n+1-level patch mesh corresponding to the target object, and the differentiable rendering is used to perform differentiable rendering on the n+1-level patch mesh.

2. The method according to claim 1, characterized in that The upsampling process is performed on the n-th level three-plane features corresponding to the n-th level spatial grid to obtain the n+1-th level three-plane features, including: Determining a voxel partitioning granularity in a process of adaptively updating the n-th level spatial grid; Based on the voxel division granularity, the nth level three-plane feature is upsampled to obtain the n+1th level three-plane feature, and the feature resolution of the n+1th level three-plane feature is consistent with the grid resolution of the n+1th level spatial grid.

3. The method according to claim 1, characterized in that The method of obtaining the n+1th level three-dimensional model corresponding to the target object based on the mesh vertex features of each mesh vertex in the n+1th level spatial mesh through geometric iteration and differentiable rendering optimization includes: Based on the mesh vertex features of each mesh vertex in the n+1th level spatial mesh, the target object is reconstructed through a mesh reconstruction network to obtain an n+1th level patch mesh corresponding to the target object; Based on the n+1th level patch mesh, performing differentiable rendering through a differentiable rendering network to obtain a rendering prediction image of the n+1th level patch mesh under various camera poses; Training the n+1th level three-plane features, the mesh reconstruction network, and the differentiable rendering network based on the prediction loss between the rendering prediction image and the real shot image, wherein the real shot image is obtained by shooting the target object based on various camera poses; When the training is completed, the target object is reconstructed based on the mesh vertex features of each mesh vertex in the n+1-th level spatial mesh through the mesh reconstruction network and the differentiable rendering network to obtain the n+1-th level three-dimensional model.

4. The method according to claim 3, characterized in that The training of the n+1th level three-plane features, the mesh reconstruction network, and the differentiable rendering network based on the prediction loss between the rendering prediction image and the real photographic image includes: Based on the rendering prediction image and the real shooting image corresponding to each camera pose, calculating the pixel prediction difference between the rendering prediction image and the real shooting image; Performing norm calculation and accumulation processing on the pixel prediction differences to obtain a total prediction loss; Based on the predicted total loss, the n+1th level three-plane features, the mesh reconstruction network, and the differentiable rendering network are trained.

5. The method according to claim 4, characterized in that The calculating, based on the rendering prediction image and the real shot image corresponding to each camera pose, a pixel prediction difference between the rendering prediction image and the real shot image, comprises: Downsampling the real photographic image to obtain an n+1th level real photographic image, where the image resolution of the n+1th level real photographic image is consistent with the grid resolution of the n+1th level spatial grid; The pixel prediction difference is calculated based on the rendering prediction image corresponding to each camera pose and the n+1th level real shooting image.

6. The method according to claim 3, characterized in that When the training is completed, based on the mesh vertex features of each mesh vertex in the n+1-th level spatial mesh, the target object is reconstructed through the mesh reconstruction network and the differentiable rendering network to obtain the n+1-th level three-dimensional model, including: When the prediction loss between the rendered prediction image and the real photographed image satisfies the n+1th level loss threshold, based on the mesh vertex features of each mesh vertex in the n+1th level spatial mesh, the target object is model reconstructed by the mesh reconstruction network and the differentiable rendering network to obtain the n+1th level three-dimensional model, and the n+1th level loss threshold is less than the nth level loss threshold; or When the number of training iterations meets the n+1th level iteration threshold, the target object is modeled reconstructed through the mesh reconstruction network and the differentiable rendering network based on the mesh vertex features of each mesh vertex in the n+1th level spatial grid to obtain the n+1th level three-dimensional model, and the n+1th level iteration threshold is greater than the nth level iteration threshold.

7. The method according to claim 3, characterized in that The method of reconstructing the target object by a mesh reconstruction network based on the mesh vertex features of each mesh vertex in the n+1-th level spatial mesh to obtain the n+1-th level patch mesh corresponding to the target object includes: Inputting mesh vertex features of each mesh vertex in the n+1th level spatial mesh into a first multilayer perceptron in the mesh reconstruction network, and outputting a signed distance function value of each mesh vertex in the n+1th level spatial mesh through the first multilayer perceptron, wherein the signed distance function value is used to represent a correspondence between the mesh vertex and the object surface in three-dimensional space; Based on the signed distance function value of each mesh vertex, a marching cubes algorithm is used to generate an n+1th level patch mesh corresponding to the target object.

8. The method according to claim 7, characterized in that The step of generating an n+1th level patch mesh corresponding to the target object using a marching cubes algorithm based on the signed distance function values ​​of the mesh vertices includes: Determining a dual vertex corresponding to each voxel based on a position code of a grid vertex corresponding to each voxel in the n+1th level spatial grid; Connecting the dual vertex corresponding to each voxel with the dual vertex of at least one adjacent voxel to obtain an n+1th level dual grid corresponding to the target object, where the adjacent voxels are other voxels that have a common vertex with the voxel; Based on the signed distance function value of each grid vertex in the n+1-th level dual grid, the marching cube algorithm is used to map and obtain the isosurface of the n+1-th level dual grid; Based on the isosurface of the n+1th level dual grid, the n+1th level patch grid corresponding to the target object is generated.

9. The method according to claim 3, characterized in that The method of performing differentiable rendering based on the n+1th level patch mesh through a differentiable rendering network to obtain a rendering prediction graph of the n+1th level patch mesh under various camera poses includes: Based on the camera pose, performing multi-view rasterization processing on the n+1th level patch mesh to obtain a dense pixel map of the n+1th level patch mesh at each camera pose, wherein the dense pixel map represents surface vertex information of surface vertices located on the mesh patch at the camera pose; determining surface vertex features of the surface vertices based on the dense pixel map and the n+1th level patch mesh; Inputting the surface vertex features into a second multilayer perceptron in the differentiable rendering network, and outputting vertex color features of the surface vertices through the second multilayer perceptron; Based on the vertex color features, the rendering prediction graph of the n+1th level patch mesh under each camera pose is generated.

10. The method according to claim 1, characterized in that The adaptive voxel updating of the nth level spatial grid corresponding to the target object to obtain the n+1th level spatial grid includes: Determining a signed distance function value of each grid vertex in the n-th level spatial grid; Determining, based on signed distance function values ​​of respective grid vertices in the n-th level spatial grid, subdividable voxels and mergeable voxels contained in the n-th level spatial grid; The subdividable voxels are subdivided, and the mergeable voxels are merged to obtain the n+1th level spatial grid.

11. The method according to claim 10, characterized in that The determining, based on the signed distance function value of each grid vertex in the n-th level spatial grid, the subdividable voxels and the mergeable voxels contained in the n-th level spatial grid comprises: Determining, based on the signed distance function values ​​of the grid vertices in the n-th level spatial grid, a minimum absolute value among the absolute values ​​corresponding to the signed distance function values; When the minimum absolute value is smaller than the subdivision threshold and the voxel where the mesh vertex corresponding to the minimum absolute value is located has at least two mesh vertices whose signed distance function values ​​have different signs, the voxel is determined as the subdividable voxel; determining other voxels except the subdividable voxels as non-subdividable voxels; When the absolute value of the signed distance function value of the mesh vertex of the non-subdividable voxel is greater than a merging threshold, the non-subdividable voxel is determined as the mergable voxel.

12. A three-dimensional reconstruction device, characterized in that: The device comprises: a voxel updating module configured to adaptively update the nth-level spatial grid corresponding to the target object to obtain an n+1th-level spatial grid, wherein the density of voxel distribution of the three-dimensional surface in the n+1th-level spatial grid is higher than the density of voxel distribution of the three-dimensional surface in the nth-level spatial grid; a sampling processing module, configured to perform upsampling processing on the n-th level three-plane features corresponding to the n-th level spatial grid to obtain the n+1-th level three-plane features; a feature extraction module, configured to extract mesh vertex features of each mesh vertex in the n+1th level spatial mesh from the n+1th level three-plane features; A model generation module is configured to obtain, based on mesh vertex features of each mesh vertex in the n+1-th level spatial mesh, a level n+1 three-dimensional model corresponding to the target object through geometric iteration and differentiable rendering optimization, wherein the model resolution of the n+1-th level three-dimensional model is higher than the model resolution of the n-th level three-dimensional model, the geometric iteration is configured to generate the level n+1 patch mesh corresponding to the target object, and the differentiable rendering is configured to perform differentiable rendering on the level n+1 patch mesh.

13. A computer device, characterized in that: The computer device includes a processor and a memory; the memory stores at least one computer instruction, and the at least one computer instruction is used to be executed by the processor to implement the three-dimensional reconstruction method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that The readable storage medium stores at least one computer instruction, and the at least one computer instruction is loaded and executed by a processor to implement the three-dimensional reconstruction method according to any one of claims 1 to 11.

15. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the three-dimensional reconstruction method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Sparse view scene super-resolution reconstruction method based on three-dimensional Gaussian representation and wavelet domain constraint

    CN121616461A