3D Reconstruction and Model Optimization Method, Device, Storage Medium and Program Product
By introducing microgrid reconstruction and microrendering technologies into the three-dimensional reconstruction technology, the initial mesh model is optimized, and the problems of irregular geometric structure and topological structure of the three-dimensional model are solved, achieving higher quality three-dimensional model reconstruction.
Patent Information
- Application Number
- CN202411375726.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-09-29
AI Technical Summary
In the existing three-dimensional reconstruction technology, the reconstructed three-dimensional model may have poor geometric structure accuracy and irregular topological structure, which cannot meet the requirements of model accuracy in application scenarios.
By obtaining the initial mesh model of the target object, a microgrid reconstruction technology is introduced to optimize the initial mesh model, a target reconstruction model with differentiable properties is generated, and a microrendering technology is used to further optimize the target reconstruction model to obtain the target mesh model using the loss function between the original image and the rendered image as the optimization goal.
The geometric accuracy and topological regularity of the three-dimensional model are improved, thereby improving the overall quality of the three-dimensional model.
Smart Images

Figure CN118941742B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model processing, and in particular, to a three-dimensional reconstruction and model optimization method, device, storage medium, and program product. Background Art
[0002] Three-dimensional reconstruction refers to the use of computer vision and graphics technologies to digitally restore real three-dimensional objects on a computer to obtain a digital model suitable for computer representation and processing, which has a relatively wide range of applications. For example, in the e-commerce field, by creating a three-dimensional digital model of a product, the online shopping experience of consumers can be greatly enriched. For example, in the medical field, through the three-dimensional reconstruction of the human bronchus, disease research and teaching analysis can be carried out more accurately.
[0003] With the development of three-dimensional reconstruction technology, more and more three-dimensional reconstruction schemes have emerged, such as three-dimensional reconstruction based on structured light, three-dimensional reconstruction based on laser scanning method, three-dimensional reconstruction based on stereo vision method, and three-dimensional reconstruction based on deep learning method, etc. These three-dimensional reconstruction methods can reconstruct the three-dimensional model of an object, but the reconstructed three-dimensional model may have problems with unsatisfactory quality, such as poor geometric structure accuracy or irregular topological structure, and cannot meet the requirements of the application scenario for model accuracy. Summary of the Invention
[0004] Multiple aspects of this application provide a three-dimensional reconstruction and model optimization method, device, storage medium, and program product to improve the accuracy of the geometric structure and the regularity of the topological structure of the three-dimensional model, and improve the overall model quality.
[0005] An embodiment of this application provides a model optimization method, including: obtaining an initial mesh model of a target object, where the initial mesh model is obtained by performing three-dimensional reconstruction on a raw image including the target object; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; using the loss function between the raw image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target, and optimizing the target reconstruction model to obtain a target mesh model.
[0006] An embodiment of this application also provides a three-dimensional reconstruction method, including: obtaining a raw image including a target object, and performing three-dimensional reconstruction on the target object based on the raw image to obtain an initial mesh model; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; using the loss function between the raw image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target, and optimizing the reconstruction model to obtain a target mesh model.
[0007] An embodiment of the present application further provides a model optimization method, including: receiving an initial mesh model of a target object sent by a terminal device, where the initial mesh model is obtained by performing three-dimensional reconstruction on an original image including the target object; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties; using the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as an optimization target, optimizing the reconstruction model to obtain a target mesh model; and sending the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0008] An embodiment of the present application further provides a three-dimensional reconstruction method, including: receiving an original image including a target object sent by a terminal device, and performing three-dimensional reconstruction on the target object based on the original image to obtain an initial mesh model; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties; using the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as an optimization target, optimizing the target reconstruction model to obtain a target mesh model; and sending the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0009] An embodiment of the present application further provides an electronic device, including: a memory and a processor; the memory is used for storing one or more computer instructions; the processor is used for executing the one or more computer instructions to: execute the steps in the model optimization method or the three-dimensional reconstruction method.
[0010] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to be able to implement the steps in the model optimization method or the three-dimensional reconstruction method.
[0011] An embodiment of the present application further provides a computer program product, including computer program / instructions, which, when executed by a processor, execute the steps in the model optimization method or the three-dimensional reconstruction method.
[0012] In this embodiment, either obtain the initial mesh model of the target object or perform three-dimensional reconstruction on the original image of the target object to obtain the initial mesh model; on the basis of the initial mesh model, introduce the differentiable mesh reconstruction technology to perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties; on this basis, introduce the differentiable rendering technology, and optimize the target reconstruction model with the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model to obtain the target mesh model. In this way, based on the differentiable mesh reconstruction and the model optimization process based on the differentiable rendering technology, the geometric structure accuracy of the three-dimensional model can be higher, and the topological structure can be more regular, thereby improving the overall quality of the three-dimensional model. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0014] Figure 1 It is a schematic flowchart of a model optimization method provided by an exemplary embodiment of the present application;
[0015] Figure 2 It is a schematic flowchart of a model optimization method in an actual scenario provided by an exemplary embodiment of the present application;
[0016] Figure 3 It is an effect diagram of topological structure optimization provided by an exemplary embodiment of the present application;
[0017] Figure 4 It is a schematic flowchart of a three-dimensional reconstruction method provided by an exemplary embodiment of the present application;
[0018] Figure 5 It is a schematic flowchart of a model optimization method provided by an exemplary embodiment of the present application;
[0019] Figure 6 It is a schematic flowchart of a three-dimensional reconstruction method provided by an exemplary embodiment of the present application;
[0020] Figure 7 It is a schematic structural diagram of a model optimization device provided by an exemplary embodiment of the present application;
[0021] Figure 8 It is a schematic structural diagram of a three-dimensional reconstruction device provided by an exemplary embodiment of the present application;
[0022] Figure 9 It is a schematic structural diagram of a model optimization device provided by an exemplary embodiment of the present application;
[0023] Figure 10 Schematic structural diagram of a three-dimensional reconstruction device provided by an exemplary embodiment of the present application;
[0024] Figure 11 Schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject. Additionally, various models (including but not limited to language models or large models) involved in the present application comply with relevant laws and standards.
[0027] In related application scenarios of three-dimensional reconstruction, although some three-dimensional reconstruction methods can reconstruct the three-dimensional model of an object, the reconstructed three-dimensional model may have problems with unsatisfactory quality. For example, the geometric structure accuracy is poor, the topological structure of the generated three-dimensional model is often uncontrollable and very messy, and it is also easy to generate blurring problems when performing texture mapping on the three-dimensional model later, and the complexity of UV (two-dimensional texture coordinates corresponding to the vertex information of geometric graphics) mapping expansion also becomes correspondingly higher. And these problems usually can only be solved by manual processing in the later stage, which greatly increases the cost of three-dimensional reconstruction.
[0028] For the above technical problems, in the embodiments of the present application, an initial mesh model of the target object can be obtained, or an initial mesh model can be obtained by performing three-dimensional reconstruction on the original image of the target object; on the basis of the initial mesh model, a differentiable mesh reconstruction technology is introduced to perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can express the initial mesh model and has differentiable properties; on this basis, a differentiable rendering technology is introduced, and the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model is used as the optimization target to optimize the target reconstruction model to obtain the target mesh model. In this way, based on differentiable mesh reconstruction and optimization based on differentiable rendering technology, the geometric structure accuracy of the three-dimensional model is higher, and the topological structure is more regular, thereby improving the overall quality of the three-dimensional model.
[0029] The following will describe in detail the technical solutions provided by each embodiment of the present application with reference to the accompanying drawings.
[0030] Figure 1 FIG. is a schematic flowchart of a model optimization method provided by an exemplary embodiment of the present application. The model optimization method can be executed by a model optimization platform. In terms of implementation form, the model optimization platform can be implemented as an electronic device. The electronic device can be a terminal device such as a mobile phone, a computer, and a tablet computer, or a server such as a cloud server or a general server. The type of the electronic device is not limited in this embodiment. As Figure 1 shown, the method includes the following steps:
[0031] Step 11: Obtain an initial mesh model of the target object, where the initial mesh model is obtained by performing three-dimensional reconstruction on the original image containing the target object.
[0032] Step 12: Perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can express the initial mesh model and has differentiable properties.
[0033] Step 13: Use the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target to optimize the target reconstruction model to obtain the target mesh model.
[0034] The specific type of the target object is not limited in this embodiment and can be an object, a commodity, a person, or an animal or a plant, etc. In the embodiments of the present application, an initial mesh model of the target object can be obtained, and the initial mesh model is obtained by performing three-dimensional reconstruction on the original image containing the target object. The specific implementation manner of the three-dimensional reconstruction is not limited in this embodiment. In some exemplary embodiments, it can be obtained by the model optimization platform for three-dimensional reconstruction, or can be reconstructed by a three-dimensional reconstruction platform provided by a third party. This embodiment does not make any restrictions.
[0035] Among them, the mesh model is composed of a series of interconnected triangular planes (faces), and each triangular plane is a mesh. Each triangular plane can be composed of a set of vertices and edges, and these triangular planes together constitute the surface of a complete mesh model. Correspondingly, the initial mesh model refers to the model obtained by three-dimensional reconstruction of the original image and expressed based on meshes / triangular planes. Optionally, the initial mesh model does not have differentiable properties. Among them, the differentiable property refers to the model property that can calculate continuous and differentiable rates of change on the surface or parameter space of the three-dimensional model. Based on the differentiable property of the three-dimensional model, gradient information can be used to optimize model parameters. The fact that the initial mesh model does not have differentiable properties means that model parameters such as the surface normal, surface curvature, vertex position, and topology of the initial mesh model are not continuously differentiable, and the initial mesh model does not support the optimization of model parameters using optimization algorithms such as gradient descent.
[0036] After obtaining the initial mesh model, the initial mesh model can be further subjected to differentiable mesh reconstruction to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties. Among them, the target reconstruction model refers to the reconstructed differentiable three-dimensional model. Differentiable mesh reconstruction refers to the process of reconstructing the initial mesh model into a differentiable model. In other words, it is to re-express the initial mesh model expressed in meshes in another way. The specific implementation method of differentiable mesh reconstruction in this embodiment is not limited. For example, FlexiCubes (a technology for realizing differentiable rendering) can be used to perform differentiable mesh reconstruction on the initial mesh model, or other technologies such as DMTet (Deep Marching Tetrahedra) or DMesh (Differentiable Mesh) can be used to perform differentiable mesh reconstruction on the initial mesh model, and this implementation does not make restrictions. Among them, FlexiCubes is an algorithm for efficiently representing and operating three-dimensional meshes. It can use a series of deformable cube units (cubes) to approximate the surface of the three-dimensional model. Each unit can be stretched or compressed according to the details and needs of the model, so as to form a highly adaptable mesh representation. Among them, DMTet uses a series of stacked tetrahedral units to represent the three-dimensional model in space, and adapts to different geometric structures by changing the properties such as the deformation and distance of each unit. DMesh is an algorithm for adaptively expressing differentiable three-dimensional meshes based on tetrahedral unit representation and using Delaunay Triangulation.
[0037] In this embodiment, a differentiable mesh reconstruction technique is introduced. The target reconstruction model obtained by performing differentiable mesh reconstruction on the initial mesh model can be used to represent the initial mesh model, and the target reconstruction model has differentiable properties. That is to say, the model parameters of the target reconstruction model, such as surface normals, surface curvatures, vertex positions, topological structures, etc., are continuously differentiable, which is simply referred to as these model parameters being differentiable. For differentiable model parameters, optimization algorithms such as gradient descent can be used to adjust these model parameters to achieve specific optimization goals. Based on this, in this embodiment, a differentiable rendering technique is further introduced, which can perform differentiable rendering on the target reconstruction model to obtain a rendered image of the target object. Among them, differentiable rendering is a technique used to optimize 3D models in computer graphics. It is a rendering process that can be differentiated. It can be divided into forward and reverse processes. In the forward process, a model and related model parameters can be input to obtain a rendered image. The reverse process refers to the process of calculating the gradient of the rendered image with respect to the model parameters. In the reverse process, the partial derivatives of the loss function with respect to each model parameter can be calculated through backpropagation, so as to determine how each model parameter affects the final rendering result based on the calculated partial derivatives. After determining the gradient of the rendered image with respect to the scene parameters through the reverse process, optimization algorithms such as gradient descent can be further used to adjust the model parameters to update the model, making the model parameters more accurate and the rendering result closer to the target image or the desired effect.
[0038] In this embodiment, the implementation method of the rendered image is not limited. For example, the rendered image may include a reconstructed image of the original image, or may include a silhouette image corresponding to the reconstructed image, or may also include both the reconstructed image and the silhouette image. Among them, the reconstructed image is the reconstruction result of the original image and is a complete image, which can be an RGB (Red-Green-Blue, a color model) image or an RGBD image. An RGB image refers to an image expressed based on the three colors of red, green, and blue, and an RGBD image is an image expressed based on the three colors of red, green, and blue and depth information; the silhouette image refers to an image used to describe the contour of the target object, which can be obtained by performing binary processing on the RGB image. Binary processing refers to marking the foreground object (i.e., the target object) of the RGB image as white (usually 1), and the rest of the background is marked as black (usually 0).
[0039] After obtaining the rendered image of the target object, the loss function between the original image and the rendered image can be used as the optimization objective to optimize the target reconstruction model, so as to obtain the target mesh model. Among them, the loss function between the original image and the rendered image can be converged to a preset range as the optimization objective to optimize the target reconstruction model to obtain the target network model. Among them, the preset range can be set to any range according to actual design requirements, and this embodiment does not make any restrictions.
[0040] In the embodiment of the present application, the implementation manner of optimizing the target reconstruction model is not limited. For example, it can be that the loss function between the original image and the rendered image is converged to a preset range as the optimization objective, and the target reconstruction model is optimized in a single stage; or, it can also be that the loss function between the original image and the rendered image is converged to a preset range as the optimization objective, and the target reconstruction model can be optimized in two stages from coarse to fine to obtain the target mesh model; of course, according to the accuracy requirements of the target reconstruction model, the target reconstruction model can also be optimized in more stages (such as three stages or more than three stages). Among them, regardless of the number of stages of the optimization method, the loss function between the original image and the rendered image is converged to a preset range as the optimization objective, and the optimization objects are different according to the different optimization stages. For example, in the single-stage optimization process, the optimization objects can include one or more of the texture information, camera pose, and geometric structure of the target reconstruction model. Preferably, the texture information, camera pose, and geometric structure of the target reconstruction model can be optimized simultaneously. For another example, in the two-stage optimization process, the optimization objects in the first stage can include the texture information and camera pose of the target reconstruction model, and the optimization objects in the second stage can include the geometric structure of the target reconstruction model; further optionally, the optimization objects in the first stage can also include the geometric structure of the target reconstruction model. For another example, in the three-stage optimization process, the optimization object in the first stage can include the texture information of the target reconstruction model, the optimization object in the second stage can include the camera pose of the target reconstruction model, and the optimization object in the third stage can include the geometric structure of the target reconstruction model, that is, different objects can be optimized in three stages. The above examples of which optimization objects are optimized in each stage are only examples and are not limited thereto.
[0041] In an embodiment, an initial mesh model of a target object is obtained. Based on the initial mesh model, a differentiable mesh reconstruction technique is introduced to perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes. On this basis, a differentiable rendering technique is introduced, and the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model is used as the optimization objective to optimize the target reconstruction model to obtain a target mesh model. In this way, based on differentiable mesh reconstruction and optimization based on the differentiable rendering technique, the geometric structure accuracy of the 3D model is higher and the topological structure is more regular, thereby improving the overall quality of the 3D model.
[0042] The embodiments of the present application do not limit the specific implementation manner of the foregoing step 12 "performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes". In an exemplary embodiment, a solid unit can be used to perform differentiable mesh reconstruction on the initial mesh model, and the reconstruction process can be implemented based on the following steps 121 - 123:
[0043] Step 121: Determine the initial differentiable parameters and the initial solid unit required to represent differentiable meshes. The initial solid unit has a default shape.
[0044] Among them, differentiable meshes are an advanced representation method used in the fields of computer graphics and machine learning, and can perform differentiation on 3D geometric shapes during the optimization and training processes. Differentiable meshes are usually composed of a set of vertices and the faces connecting these vertices, and the vertex positions and other geometric attributes can be continuously differentiable functions, so that optimization algorithms such as gradient descent can be performed based on differentiable meshes to perform subsequent optimization on the target reconstruction model.
[0045] Among them, the initial solid unit refers to the smallest unit used to fit the initial mesh model. The fitting of the initial mesh model can be completed by adjusting the relative positions of the individual initial solid units. The initial solid unit has a default shape, and the specific implementation manner of the default shape is not limited in this embodiment, and it can be a cuboid, a cube, or the shape of other tetrahedrons. Preferably, a cuboid can be used as the initial solid unit. The cuboid has a simple geometric shape, is easy to understand and process, and is conducive to segmentation and combination. Therefore, using a cuboid as the initial solid unit can improve the efficiency of fitting the initial mesh model.
[0046] Among them, the initial differentiable parameters required for the differentiable grid are used to represent the geometry of the initial solid elements and the initial grid model. That is to say, the initial differentiable parameters include both the parameters of each initial solid element and the parameters of the overall initial reconstruction model composed of the initial solid elements. These parameters work together to represent complex geometries in a flexible and detailed manner. Among them, the parameters of each initial solid element may include at least one of the following: position parameters, deformation parameters, and level-of-detail parameters. The position parameters are used to describe the position of each initial solid element in its space, usually represented by three-dimensional coordinates. The deformation parameters are used to control the ability of each initial solid element to deform from its original default shape, so that each initial solid element adapts to the local details of the initial grid model. The deformation methods may include translation, rotation, scaling, and more complex non-linear deformations. The level-of-detail parameters are used to control the resolution of each initial solid element, that is, the fineness of the details represented by the initial reconstruction model. Among them, the parameters at the level of the entire initial reconstruction model composed of the initial solid elements may include at least one of the following: camera pose parameters, size ratio parameters, connectivity parameters, texture parameters, optimization parameters, and constraint parameters. The size ratio parameters are used to define the size and proportional relationship of the entire initial reconstruction model. The size ratio parameters can affect the relative size and position of each initial solid element. The connectivity parameters are used to control the connection method between the initial solid elements to ensure the continuity and integrity of the initial reconstruction model. The camera pose parameters are used to describe the position and pose of the camera when taking the original image. The texture parameters are used to define the visual appearance of the initial reconstruction model, including color, reflection characteristics, transparency, etc. The optimization parameters refer to the parameters related to the fitting process of the initial grid model, and may include at least one of the following: the number of iterations, the convergence threshold, and the regularization term. The optimization parameters affect how the initial solid elements approximate the initial grid model. The constraint parameters are used to constrain the fitting process of the initial grid model. For example, the constraint parameters can keep certain parts of the initial grid model unchanged, or limit the type and degree of deformation of certain parts of the initial grid model to meet specific engineering or aesthetic requirements.
[0047] Step 122: Adjust the shape of the initial solid elements to obtain target solid elements that fit the shape of the target object. Among them, the shape refers to the proportional relationship of the various components of the initial solid element. If the initial solid element is a cuboid, then its corresponding shape is the proportional relationship of length, width, and height.
[0048] In this embodiment, the specific shape adjustment method is not limited. In some exemplary embodiments, the initial three-dimensional unit can be enlarged or reduced proportionally. In other exemplary embodiments, the size information of the initial mesh model can be obtained. The size information is used to describe the length, width, and height of the initial mesh model. According to the size information, the shape of the initial three-dimensional unit is adjusted to obtain a target three-dimensional unit that fits the shape of the target object. Among them, when adjusting the shape of the initial three-dimensional unit according to the size information, the first ratio between the length, width, and height of the initial mesh model can be calculated, and the initial three-dimensional unit is adjusted using the first ratio so that the second ratio between the length, width, and height of the initial three-dimensional unit is the same as or close to the first ratio. In this case, "close" means that the first ratio and the second ratio are within a certain error range. By adjusting the shape of the initial three-dimensional unit, the degree of fit between the initial three-dimensional unit and the shape of the target object can be improved, and finally the target three-dimensional unit can be obtained. Compared with directly using the initial three-dimensional unit to fit the initial mesh model, using the target three-dimensional unit that fits the shape of the target object can more accurately and efficiently fit the initial mesh model.
[0049] Step 123: Fit the initial mesh model according to the initial differentiable parameters and the target three-dimensional unit to obtain a target reconstruction model. The target reconstruction model obtained through this fitting process has target differentiable parameters. The process of fitting the initial mesh model is also the process of continuously adjusting the initial differentiable parameters of the initial mesh model and finally obtaining the target differentiable parameters. The target differentiable parameters refer to the differentiable parameters that satisfy the fitting termination condition after adjusting the initial differentiable parameters.
[0050] This embodiment does not limit the specific implementation manner of this step 123. In an exemplary embodiment, the initial mesh model can be covered with the target three-dimensional unit according to the initial differentiable parameters to obtain an initial reconstruction model. Among them, the number, size, and position of the target three-dimensional units can be adjusted according to the initial differentiable parameters, and the target three-dimensional units are combined to cover the initial mesh model. Then, taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization goal, the gradient descent algorithm can be used to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model.
[0051] In the embodiments of the present application, the implementation manner of expressing the similarity between the initial mesh model and the initial reconstruction model is not limited. For example, it can include, but is not limited to: the chamfer distance, Hausdorff distance, average point-to-point distance, or the similarity between geometric features between the initial mesh model and the initial reconstruction model, etc. The following is an example:
[0052] Optionally, the chamfer distance between the initial reconstruction model and the initial mesh model can be calculated as the similarity error therebetween. Among them, the average value of the distances from each position point located in the initial reconstruction model to the nearest position point of the initial mesh model can be calculated as the chamfer distance. In the case where the chamfer distance does not meet the set first distance condition, the differentiable parameters of the initial reconstruction model are iteratively optimized according to the gradient vector of the chamfer distance until the chamfer distance meets the first set distance condition, and the initial reconstruction model when the chamfer distance meets the first set distance condition is used as the target reconstruction model. Among them, the first set distance condition can be set to any condition, which is not limited in this embodiment. Among them, the iterative optimization of the differentiable parameters of the initial reconstruction model includes but is not limited to: optimizing the position parameters, deformation parameters, and detail level parameters of each solid unit in the initial reconstruction model, and optimizing the parameters at the entire model level, such as size ratio parameters, connectivity parameters, optimization parameters, and constraint parameters.
[0053] Optionally, the Hausdorff distance between the initial reconstruction model and the initial mesh model can be calculated as the similarity error therebetween. Among them, the Hausdorff distance is a measure describing the similarity degree between two sets of point sets. The distance between the farthest point pair in the first point set of the initial reconstruction model and the second point set of the initial mesh model can be calculated as the Hausdorff distance. Among them, the first point set is the point set composed of each position point of the initial reconstruction model, and the second point set is the point set composed of each position point of the initial mesh model. In the case where the Hausdorff distance does not meet the second distance condition, the differentiable parameters of the initial reconstruction model are iteratively optimized according to the gradient vector of the Hausdorff distance until the Hausdorff distance meets the second distance condition, and the initial reconstruction model when the Hausdorff distance meets the second distance condition is used as the target reconstruction model. Among them, the second set distance condition can be set to any condition, which is not limited in this embodiment.
[0054] Optionally, the average point-to-point distance between the initial reconstruction model and the initial mesh model can be calculated as the similarity error therebetween. Among them, the average Euclidean distance between the first point set of the initial reconstruction model and the second point set of the initial mesh model can be calculated as the average point-to-point distance. In the case where the average point-to-point distance does not meet the third set distance condition, the differentiable parameters of the initial reconstruction model are iteratively optimized according to the gradient vector of the average point-to-point distance until the chamfer distance meets the third set distance condition, and the initial reconstruction model when the average point-to-point distance meets the third set distance condition is used as the target reconstruction model. Among them, the third set distance condition can be set to any condition, which is not limited in this embodiment.
[0055] Optionally, a similarity error between geometric features of the initial reconstruction model and the initial mesh model may be calculated. The geometric features may be curvature, normal vectors, etc., and this embodiment does not make any limitations. In the case where the geometric features do not meet the set feature conditions, the differentiable parameters of the initial reconstruction model are iteratively optimized according to the gradient vector of the geometric features until the geometric features meet the set feature conditions, and the initial reconstruction model when the geometric features meet the set feature conditions is used as the target reconstruction model. Among them, the set feature conditions may be set to any conditions, and this embodiment does not make any limitations.
[0056] Based on each of the above optional implementation manners, when iteratively optimizing the differentiable parameters of the initial reconstruction model, the differentiable parameters may be regularized in each round of iteration to more effectively prevent overfitting, thereby further improving the accuracy of the target reconstruction model.
[0057] In this embodiment, based on the above steps 121 - 123, the initial mesh model can be fitted according to the target solid unit adapted to the shape of the target object and the initial differentiable parameters to obtain the target reconstruction model more accurately.
[0058] In the embodiment of the present application, the specific implementation manner of the foregoing step 13 "using the loss function between the original image and the rendered image obtained by differentiable rendering of the target reconstruction model as the optimization target to optimize the target reconstruction model to obtain the target mesh model" is not limited. In an exemplary embodiment, it may be implemented based on the following steps 131 - 135:
[0059] Step 131: Input the target reconstruction model, its corresponding texture information, and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image. Among them, the texture information of the target reconstruction model is used to describe the details / physical properties on the surface of the target reconstruction model, and may include at least one of the following: color, brightness, roughness, and smoothness. The camera pose is used to describe the position and pose of the camera when taking the original image.
[0060] Optionally, before inputting the target reconstruction model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering, the camera pose can be obtained in the following way: the camera pose of the original image can be estimated to obtain the camera pose corresponding to the target reconstruction model. The camera pose estimation can be implemented based on the following steps: First, the key points in the original image can be identified. Since these key points can correspond to the feature points on the target reconstruction model, the detected key points can be matched with the feature points on the target reconstruction model. According to the matching result, the corresponding points between the original image and the target reconstruction model can be determined, and the camera pose corresponding to the target reconstruction model can be determined based on the corresponding points between the original image and the target reconstruction model. The accuracy of the camera pose determined here may be relatively low. Therefore, further camera pose optimization is required, and the relevant content has been described in detail later, so it will not be elaborated here.
[0061] Optionally, before inputting the target reconstruction model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering, the texture information can be obtained in the following way: based on the texture prediction network corresponding to the target object, the texture information of the target reconstruction model can be predicted to obtain the texture information corresponding to the target reconstruction model. Among them, the texture prediction network can have the ability to predict the texture information of the three-dimensional model based on the pre-training process. Specifically, the coordinates of any position point on the target reconstruction model can be input into the texture prediction network. The texture prediction network can output the texture information corresponding to the arbitrary position point according to the mapping relationship between the pre-learned model position point and the texture information and the coordinates of any position point on the target reconstruction model, so as to finally obtain the texture information corresponding to the target reconstruction model.
[0062] Among them, the images of the target object from multiple perspectives can be collected as the sample images with texture information; the initial network can be trained based on the sample images and the target reconstruction model for the ability to output texture information based on the position coordinates on the target reconstruction model to obtain the texture prediction network. Specifically, the sample images and the target reconstruction model can be input into the initial network. Under the supervision of the sample images, the initial network can be iteratively trained with the goal of converging the loss function of the initial network to a specified range to obtain the texture prediction network. Among them, in any iteration round, the coordinates of any position point on the target reconstruction model can be input into the initial network to obtain the texture information corresponding to the coordinates. If the error between the texture information corresponding to the coordinates and the texture information on the sample image does not converge to the specified range, the parameters of the initial network are updated and the next iteration round is entered; if the error converges to the specified range, the texture prediction network is output. Based on the above training process, the texture prediction network can relatively accurately output the corresponding texture information based on the position points on the target reconstruction model.
[0063] Step 132: Optimize at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image, so as to obtain a first optimized model and its corresponding texture information and camera pose. The topological structure of the target reconstruction model refers to the relationship between vertices, edges, and faces in the target reconstruction model. In this embodiment, the specific implementation manner of the first loss function is not limited and can be any function for calculating the error between the first rendered image and the original image, including but not limited to: Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), or Mean Squared Error (MSE).
[0064] Specifically, when the first loss function does not meet the first optimization end condition, the topological structure of the target reconstruction model, as well as at least one of the geometric structure of the target reconstruction model, the camera pose, and the texture prediction network, can be jointly optimized until the first loss function meets the first optimization end condition, so as to obtain a first optimized model and its corresponding texture information and camera pose. The first optimization end condition can be set to any condition according to actual design requirements, and this embodiment does not make any restrictions. Among them, different from the topological structure, the geometric structure refers to the spatial layout and shape characteristics of the model, which can be characterized by the position coordinates in the three-dimensional mesh of the model. In other words, when the position coordinates in the three-dimensional mesh of the model change, the geometric structure of the model will also change accordingly. Therefore, the accuracy of the position coordinates of the vertices in the three-dimensional mesh directly affects the quality of the model shape appearance. The first optimized model refers to the model obtained after the first-stage optimization. In this embodiment, when optimizing the topological structure of the target reconstruction model, the parameters related to the topological structure in the target differentiable parameters of the target reconstruction model can be adjusted, such as connectivity parameters, etc. By optimizing the topological structure of the target reconstruction model, the complexity of the topological structure of the target reconstruction model can be reduced, and the topological structure of the obtained first optimized model is more regular. By optimizing the geometric structure of the target reconstruction model, the shape of the target reconstruction model can be made closer to the actual shape of the target object, and the shape of the obtained first optimized model is more accurate. By optimizing the camera pose, the error from the actual camera pose when shooting the original image can be reduced, and the accuracy of the camera pose can be improved. By optimizing the texture prediction network, the texture information predicted by the texture prediction network can be made closer to the actual surface material of the target object. In this way, at least the topological structure and the camera pose can be optimized more accurately, and the problems of irregular model topological structure and inaccurate camera pose can be more effectively solved.
[0065] Step 133: Input the first optimized model, its corresponding texture information, and the camera pose into a differentiable renderer for differentiable rendering to obtain a second rendered image. In this embodiment, the type of the differentiable renderer and the specific implementation manner of its differentiable rendering are not limited. In an exemplary embodiment, the differentiable renderer can sequentially perform vertex processing, rasterization, and texture interpolation according to the first optimized model, its corresponding texture information, and the camera pose, thereby completing differentiable rendering. Among them, vertex processing refers to the process of transforming and operating on the vertices of the first optimized model, usually including steps such as vertex position transformation, lighting calculation, and texture coordinate calculation. Rasterization refers to the process of converting the first optimized model into a two-dimensional image. Specifically, it can determine which pixels are covered by which geometric elements, and calculate the colors of these pixels based on the colors and lighting information of these geometric elements, thereby completing rasterization. Texture interpolation refers to interpolating the color values on the texture map during the rendering process to obtain the color of the pixel. In this way, the differentiable renderer can perform differentiable rendering more efficiently based on the first optimized model, its corresponding texture information, and the camera pose.
[0066] Step 134: Optimize at least the geometric structure of the first optimized model according to the second loss function between the second rendered image and the original image to obtain a second optimized model, its corresponding texture information, and the camera pose. Among them, the second optimized model refers to the model obtained after the second-stage optimization.
[0067] It should be noted here that the above steps 131 - 132 describe the model optimization process in the first stage, and the above steps 133 - 134 describe the model optimization process in the second stage. Based on the model optimization process in the first stage, at least the topological structure and camera pose of the model can be optimized more accurately. Based on the model optimization process in the second stage, at least the geometric structure of the model can be optimized more accurately. In step 134, the topological structure and camera pose of the first optimized model can be fixed, and no further adjustments are made to the topological structure and camera pose, that is, the topological structure and camera pose of the first optimized model are non - adjustable. In this case, when the second loss function does not meet the second optimization end condition, the geometric structure of the first optimized model can be optimized. Further optionally, in the second stage, the geometric structure and texture information of the first optimized model can be optimized simultaneously to obtain a second optimized model with high - quality texture information and geometric structure. The specific implementation manner of the second loss function in this embodiment is not limited and can be any function for calculating the error between the second rendered image and the original image, including but not limited to: mean absolute error, root mean square error, or mean square error. Among them, the second optimization end condition can be set to any condition according to actual needs, and this embodiment does not make any restrictions. Among them, the number, size, and position of the target solid units can be adjusted to optimize the geometric structure of the first optimized model. When adjusting the position of the target solid units, the vertex position of the target solid units can be offset within a first preset range near the normal of the target solid units to fine - tune the geometric structure. Among them, the texture information of the first optimized model can be offset within a second preset range to fine - tune the texture information. By optimizing the geometric structure and texture information of the first optimized model, the adjusted second optimized model and its texture information can be made to match the actual situation of the target object in the original image as much as possible, improving the quality of the second optimized model.
[0068] Step 135: Generate a target mesh model according to the second optimized model and its corresponding texture information and camera pose. Among them, the texture can be added to the second optimized model according to the texture information corresponding to the second optimized model, and the target mesh model under the camera view can be generated according to the camera pose.
[0069] Based on the above steps 131 - 135, when the topological structure, geometric structure, camera pose, and texture prediction network of the target reconstruction model are all adjustable, a first - stage joint optimization of at least one of them can be performed based on the first loss function between the first rendered image and the original image. After that, the topological structure and camera pose can be fixed, and a second - stage optimization of the geometric structure and texture information of the first optimized model can be carried out specifically. Through the above two - stage optimization process from coarse to fine, the geometric structure accuracy of the three - dimensional model can be higher, and the topological structure can be more regular, thereby improving the overall quality of the three - dimensional model.
[0070] The following will further illustrate the above process of this solution in conjunction with Figure 2 and Figure 3 the above process of this solution will be further described.
[0071] As Figure 2 shown, in an actual scenario, the initial mesh model can be covered by the target solid unit, that is, the random initialization process, to obtain the initial reconstruction model, that is, Figure 2 the random initialization model in ; taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization objective, the gradient descent algorithm is used to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model.
[0072] Based on the texture prediction network corresponding to the target object, texture information prediction is performed on the target reconstruction model to obtain the texture information corresponding to the target reconstruction model; the target reconstruction model, its corresponding texture information, and the camera pose are input into the differentiable renderer for differentiable rendering to obtain the first rendered image; the first rendered image may include an RGB image and a silhouette image.
[0073] After that, according to the first loss function between the first rendered image and the original image, that is, Figure 2 the loss function between the RGB image shown in and the original image and the loss function between the silhouette image and the original image, the topological structure of the target reconstruction model, the geometric structure of the target reconstruction model, the camera pose, and the texture prediction network are jointly optimized until the first loss function meets the first optimization end condition to obtain the first optimized model, its corresponding texture information, and the camera pose. Inputting the first optimized model, its corresponding texture information, and the camera pose into the differentiable renderer, the differentiable renderer can sequentially perform vertex processing, rasterization, and texture interpolation to obtain the second rendered image. The above is the "coarse optimization" process of the first stage.
[0074] Next, the "fine optimization" process of the second stage can be continued: fixing the topological structure and the camera pose of the first optimized model, and optimizing the geometric structure and texture information of the first optimized model in the case where the second loss function does not meet the second optimization end condition to obtain the second optimized model, its texture information, and the camera pose. Generating the target mesh model according to the second optimized model, its corresponding texture information, and the camera pose.
[0075] In this way, based on differentiable mesh reconstruction and the two-stage optimization process, the geometric structure accuracy of the three-dimensional model can be higher, and the topological structure can be more regular, thereby improving the overall quality of the three-dimensional model. See Figure 3The optimization effect of the shown topological structure. On the left is a product that has not been optimized according to the optimization method proposed in the embodiments of the present application, and the topological structure on its surface is chaotic. On the right is a product optimized according to the optimization method proposed in the embodiments of the present application, and the topological structure on the surface is more regular and closer to the original image.
[0076] The following will combine Figure 4 , Figure 5 and Figure 6 to further illustrate the above process of this solution in other scenarios.
[0077] As Figure 4 shown, embodiments of the present application further provide a 3D reconstruction method. This method can be executed by a 3D reconstruction platform. In terms of implementation form, the 3D reconstruction platform can be implemented as an electronic device, or can be implemented as a 3D reconstruction tool or component running on an electronic device. The electronic device can be a terminal device such as a mobile phone, a computer, and a tablet computer, or can be a server such as a cloud server or a general server. The type of the electronic device is not limited in this embodiment. This method may include:
[0078] Step 41, obtain an original image containing a target object, and perform 3D reconstruction on the target object based on the original image to obtain an initial mesh model.
[0079] Step 42, perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes.
[0080] Step 43, use the loss function between the original image and the rendered image obtained by differentiable rendering of the target reconstruction model as the optimization target, and optimize the target reconstruction model to obtain a target mesh model.
[0081] In this embodiment, the specific acquisition method of the original image containing the target object is not limited. A user or a merchant can upload the original image to the 3D reconstruction platform through the upload interface provided by the terminal device they hold, or can directly upload the original image to the 3D reconstruction platform through a web (World Wide Web) page, or can also draw the original image on the terminal device they hold and send it to the 3D reconstruction platform. This embodiment does not make any restrictions.
[0082] In this embodiment, an original image containing the target object can be obtained, and the target object can be three-dimensionally reconstructed based on the original image to obtain an initial mesh model. In this embodiment, the detailed implementation manner of three-dimensional reconstruction based on the original image is not limited. Based on the initial mesh model, a differentiable mesh reconstruction technique is introduced to perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties. On this basis, a differentiable rendering technique is introduced, and the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model is used as the optimization target to optimize the target reconstruction model to obtain a target mesh model. In this way, based on the differentiable mesh reconstruction and the model optimization process based on the differentiable rendering technique, the geometric structure accuracy of the three-dimensional model can be higher, and the topological structure can be more regular, thus improving the overall quality of the reconstructed three-dimensional model.
[0083] Optionally, performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties includes: determining the initial differentiable parameters and initial solid units required to represent the differentiable mesh, where the initial solid units have a default shape; adjusting the shape of the initial solid units to obtain target solid units adapted to the shape of the target object; and fitting the initial mesh model according to the initial differentiable parameters and the target solid units to obtain a target reconstruction model, where the target reconstruction model has target differentiable parameters.
[0084] Optionally, fitting the initial mesh model according to the initial differentiable parameters and the target solid units to obtain a target reconstruction model includes: covering the initial mesh model with the target solid units according to the initial differentiable parameters to obtain an initial reconstruction model; and using the similarity error between the initial mesh model and the initial reconstruction model as the optimization target, and using the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain a target reconstruction model.
[0085] Optionally, using the similarity error between the initial mesh model and the initial reconstruction model as the optimization target and using the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain a target reconstruction model includes: calculating the chamfer distance between the initial reconstruction model and the initial mesh model as the similarity error between the two; in the case where the chamfer distance does not meet the set distance condition, iteratively optimizing the differentiable parameters of the initial reconstruction model according to the gradient vector of the chamfer distance; until the chamfer distance meets the set distance condition, and taking the initial reconstruction model when the chamfer distance meets the set distance condition as the target reconstruction model.
[0086] Optionally, taking the loss function between the original image and the rendered image obtained by differentiable rendering of the target reconstruction model as the optimization objective, optimize the target reconstruction model to obtain the target mesh model, including: inputting the target reconstruction model, its corresponding texture information, and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; optimizing at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model, its corresponding texture information, and camera pose; inputting the first optimized model, its corresponding texture information, and camera pose into a differentiable renderer for differentiable rendering to obtain a second rendered image; optimizing at least the geometric structure of the first optimized model according to the second loss function between the second rendered image and the original image to obtain a second optimized model, its corresponding texture information, and camera pose; generating the target mesh model according to the second optimized model, its corresponding texture information, and camera pose. Among them, the model optimization process based on differentiable rendering technology can make the geometric structure of the 3D model more accurate and the topological structure more regular, thus improving the overall quality of the reconstructed 3D model.
[0087] Optionally, before inputting the target reconstruction model, its corresponding texture information, and camera pose into a differentiable renderer for differentiable rendering, it further includes: predicting the texture information of the target reconstruction model based on the texture prediction network corresponding to the target object to obtain the texture information corresponding to the target reconstruction model; estimating the camera pose of the original image to obtain the camera pose corresponding to the target reconstruction model.
[0088] Optionally, optimizing at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model, its corresponding texture information, and camera pose, including: when the first loss function does not meet the first optimization end condition, jointly optimizing at least one of the topological structure of the target reconstruction model, the geometric structure of the target reconstruction model, the camera pose, and the texture prediction network until the first loss function meets the first optimization end condition to obtain a first optimized model, its corresponding texture information, and camera pose.
[0089] Optionally, it further includes: collecting images of the target object from multiple perspectives as sample images with texture information; training the initial network according to the sample images and the ability of the target reconstruction model to output texture information based on the position coordinates on the target reconstruction model to obtain the texture prediction network.
[0090] Optionally, if the topological structure and camera pose of the first optimization model are not adjustable, then at least optimize the geometric structure of the first optimization model according to the second loss function between the second rendered image and the original image, so as to obtain the second optimization model and its corresponding texture information and camera pose, including: when the second loss function does not meet the second optimization end condition, optimize the geometric structure and texture information of the first optimization model to obtain the second optimization model and its texture information and camera pose.
[0091] The detailed implementation manners and beneficial effects of the steps in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0092] As Figure 5 shown, an embodiment of the present application further provides a model optimization method, which can be executed by an electronic device. The electronic device can be a terminal device such as a mobile phone, a computer, and a tablet computer, or a server such as a cloud server or a general server. The type of the electronic device is not limited in this embodiment. The method includes the following steps:
[0093] Step 51: Receive the initial mesh model of the target object sent by the terminal device. The initial mesh model is obtained by performing three-dimensional reconstruction on the original image including the target object;
[0094] Step 52: Perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties;
[0095] Step 53: Take the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target, and optimize the reconstruction model to obtain the target mesh model;
[0096] Step 54: Send the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0097] In this embodiment, the terminal device can be a merchant terminal or a user terminal. When the terminal device is a merchant terminal, the merchant terminal can send the initial mesh model of the target object to the electronic device. The initial mesh model can be generated locally by the merchant terminal based on the original image, manually made by the merchant's service staff using the merchant terminal, or received by the merchant terminal from other terminal devices. This embodiment does not make any restrictions. After receiving the initial mesh model, the electronic device can execute the above steps 52 - 53 to obtain the target mesh model. Then, the electronic device can send the target mesh model to the merchant terminal for the merchant terminal to use. Optionally, the merchant terminal can publish the target mesh model to the product details page for the user terminal to view, or use the target mesh model and the corresponding 3D scene information to construct a 3D product scene to obtain the mesh model in the target scene for the user terminal to view. Among them, the 3D scene information is used to describe the scene where the target mesh model is located. For example, if the target mesh model is a glass of milk, then the 3D scene information corresponding to the target mesh model can be the relevant information of the grassland scene. Using the target mesh model and the corresponding 3D scene information to construct a 3D product scene, milk in the grassland scene can be obtained. Optionally, the merchant terminal can also publish the target mesh model to the product details page for the user terminal to view.
[0098] When the terminal device is a user terminal, the user terminal can send the initial mesh model of the target object to the electronic device. The initial mesh model can be generated locally by the user terminal based on the original image, manually made by the user using the user terminal, or received by the user terminal from other terminal devices. This embodiment does not make any restrictions. After receiving the initial mesh model, the electronic device can execute the above steps 52 - 53 to obtain the target mesh model. Then, the electronic device can send the target mesh model to the user terminal for the user terminal to use. Optionally, in the makeup try-on scenario, the target object can be the user himself, or the user's friends, family, etc. Then the target mesh model is the 3D model of the user, the user's friends, family, etc. The user terminal can display the target mesh model to the user for the user to add makeup effects to the target mesh model to simulate makeup try-on. Optionally, in the metaverse scenario, the user terminal can display the target mesh model to the user for the user to control the target mesh model to perform mobile interactions in the virtual world.
[0099] Optionally, perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties, including: determining initial differentiable parameters and initial solid elements required to represent a differentiable mesh, where the initial solid elements have a default shape; adjusting the shape of the initial solid elements to obtain target solid elements adapted to the shape of the target object; fitting the initial mesh model according to the initial differentiable parameters and the target solid elements to obtain a target reconstruction model, and the target reconstruction model has target differentiable parameters.
[0100] Optionally, fit the initial mesh model according to the initial differentiable parameters and the target solid elements to obtain a target reconstruction model, including: covering the initial mesh model with the target solid elements according to the initial differentiable parameters to obtain an initial reconstruction model; taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization objective, and using the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain a target reconstruction model.
[0101] Optionally, taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization objective, and using the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain a target reconstruction model, including: calculating the chamfer distance between the initial reconstruction model and the initial mesh model as the similarity error between the two; in the case where the chamfer distance does not meet the set distance condition, iteratively optimize the differentiable parameters of the initial reconstruction model according to the gradient vector of the chamfer distance; until the chamfer distance meets the set distance condition, take the initial reconstruction model when the chamfer distance meets the set distance condition as the target reconstruction model.
[0102] Optionally, taking the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization objective, optimize the target reconstruction model to obtain a target mesh model, including: inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; optimizing at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model and its corresponding texture information and camera pose; inputting the first optimized model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a second rendered image; optimizing at least the geometric structure of the first optimized model according to the second loss function between the second rendered image and the original image to obtain a second optimized model and its corresponding texture information and camera pose; generating a target mesh model according to the second optimized model and its corresponding texture information and camera pose. Among them, the model optimization process based on differentiable rendering technology can make the geometric structure accuracy of the three-dimensional model higher and the topological structure more regular, thus improving the overall quality of the reconstructed three-dimensional model.
[0103] Optionally, before inputting the target reconstruction model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering, it further includes: predicting the texture information of the target reconstruction model based on the texture prediction network corresponding to the target object to obtain the texture information corresponding to the target reconstruction model; estimating the camera pose of the original image to obtain the camera pose corresponding to the target reconstruction model.
[0104] Optionally, according to the first loss function between the first rendered image and the original image, at least optimize the topological structure of the target reconstruction model to obtain the first optimized model and its corresponding texture information and camera pose, including: when the first loss function does not meet the first optimization end condition, jointly optimize at least one of the topological structure of the target reconstruction model, the geometric structure of the target reconstruction model, the camera pose, and the texture prediction network until the first loss function meets the first optimization end condition to obtain the first optimized model and its corresponding texture information and camera pose.
[0105] Optionally, it further includes: collecting images of the target object from multiple perspectives as sample images with texture information; training the initial network according to the sample images and the target reconstruction model for the ability to output texture information based on the position coordinates on the target reconstruction model to obtain the texture prediction network.
[0106] Optionally, if the topological structure and camera pose of the first optimized model are not adjustable, then according to the second loss function between the second rendered image and the original image, at least optimize the geometric structure of the first optimized model to obtain the second optimized model and its corresponding texture information and camera pose, including: when the second loss function does not meet the second optimization end condition, optimize the geometric structure and texture information of the first optimized model to obtain the second optimized model and its texture information and camera pose.
[0107] The detailed implementation manners and beneficial effects of the steps in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0108] In this embodiment, an initial mesh model of a target object sent by a terminal device can be received. The initial mesh model is obtained by performing three-dimensional reconstruction on an original image containing the target object. Based on the initial mesh model, a differentiable mesh reconstruction technique is introduced to perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties. On this basis, a differentiable rendering technique is introduced, and the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model is used as the optimization objective to optimize the target reconstruction model to obtain a target mesh model. The target mesh model is sent to the terminal device for the terminal device to use the target mesh model. In this way, based on the differentiable mesh reconstruction and the model optimization process based on the differentiable rendering technique, the geometric structure accuracy of the three-dimensional model can be higher, and the topological structure can be more regular, thus improving the overall quality of the three-dimensional model.
[0109] As Figure 6 shown, an embodiment of the present application also provides a three-dimensional reconstruction method. This method can be executed by a three-dimensional reconstruction platform. In terms of implementation form, the three-dimensional reconstruction platform can be implemented as an electronic device, or can be implemented as a three-dimensional reconstruction tool or component running on an electronic device. The electronic device can be a terminal device such as a mobile phone, a computer, and a tablet computer, or can be a server such as a cloud server or a general server. The type of the electronic device is not limited in this embodiment. The method includes the following steps:
[0110] Step 61: Receive an original image containing a target object sent by a terminal device, and perform three-dimensional reconstruction on the target object based on the original image to obtain an initial mesh model;
[0111] Step 62: Perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties;
[0112] Step 63: Use the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization objective to optimize the reconstruction model to obtain a target mesh model;
[0113] Step 64: Send the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0114] In this embodiment, the terminal device can be a user terminal or a merchant terminal. The three-dimensional reconstruction platform can provide a three-dimensional reconstruction service and an original image upload entry for the terminal device. The original image upload entry can be implemented as a web page or an email. For example, the three-dimensional reconstruction platform can provide a web page for the three-dimensional reconstruction service, in which an original image upload control is set, and a user or a merchant can upload an original image based on the original image upload control. In addition, a user or a merchant can also send the original image to the three-dimensional reconstruction platform through an email.
[0115] After receiving the initial mesh model, the 3D reconstruction platform can execute the above steps 62 - 63 to obtain the target mesh model. After that, the 3D reconstruction platform can send the target mesh model to the user terminal for the user terminal to use. Optionally, in the makeup trial scenario, the target object can be the user himself / herself or the user's friends, family members, etc., then the target mesh model is the 3D model of the user, the user's friends, family members, etc., and the user terminal can display the target mesh model to the user for the user to add makeup effects to the target mesh model to simulate makeup trial. Optionally, in the metaverse scenario, the user terminal can display the target mesh model to the user for the user to control the target mesh model to perform movement interactions in the virtual world.
[0116] Optionally, perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties, including: determining the initial differentiable parameters and initial solid elements required to represent the differentiable mesh, where the initial solid elements have a default shape; adjusting the shape of the initial solid elements to obtain target solid elements adapted to the shape of the target object; fitting the initial mesh model according to the initial differentiable parameters and the target solid elements to obtain the target reconstruction model, and the target reconstruction model has target differentiable parameters.
[0117] Optionally, fitting the initial mesh model according to the initial differentiable parameters and the target solid elements to obtain the target reconstruction model, including: covering the initial mesh model with the target solid elements according to the initial differentiable parameters to obtain the initial reconstruction model; taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization target, and using the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model.
[0118] Optionally, taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization target, and using the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model, including: calculating the chamfer distance between the initial reconstruction model and the initial mesh model as the similarity error between the two; in the case where the chamfer distance does not meet the set distance condition, iteratively optimize the differentiable parameters of the initial reconstruction model according to the gradient vector of the chamfer distance; until the chamfer distance meets the set distance condition, and taking the initial reconstruction model when the chamfer distance meets the set distance condition as the target reconstruction model.
[0119] Optionally, with the loss function between the original image and the rendered image obtained by differentiable rendering of the target reconstruction model as the optimization objective, the target reconstruction model is optimized to obtain the target mesh model, including: inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; optimizing at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model and its corresponding texture information and camera pose; inputting the first optimized model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a second rendered image; optimizing at least the geometric structure of the first optimized model according to the second loss function between the second rendered image and the original image to obtain a second optimized model and its corresponding texture information and camera pose; generating the target mesh model according to the second optimized model and its corresponding texture information and camera pose. Among them, the model optimization process based on differentiable rendering technology can make the geometric structure of the 3D model more accurate and the topological structure more regular, thus improving the overall quality of the reconstructed 3D model.
[0120] Optionally, before inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering, it further includes: predicting the texture information of the target reconstruction model based on the texture prediction network corresponding to the target object to obtain the texture information corresponding to the target reconstruction model; estimating the camera pose of the original image to obtain the camera pose corresponding to the target reconstruction model.
[0121] Optionally, optimizing at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model and its corresponding texture information and camera pose includes: when the first loss function does not meet the first optimization end condition, jointly optimizing at least one of the topological structure of the target reconstruction model, the geometric structure of the target reconstruction model, the camera pose, and the texture prediction network until the first loss function meets the first optimization end condition to obtain the first optimized model and its corresponding texture information and camera pose.
[0122] Optionally, it further includes: collecting images of the target object from multiple perspectives as sample images with texture information; training the initial network according to the sample images and the target reconstruction model for the ability of the target reconstruction model to output texture information based on the position coordinates on the model to obtain the texture prediction network.
[0123] Optionally, if the topological structure and camera pose of the first optimization model are not adjustable, then at least optimize the geometric structure of the first optimization model according to the second loss function between the second rendered image and the original image to obtain the second optimization model and its corresponding texture information and camera pose, including: when the second loss function does not meet the second optimization end condition, optimize the geometric structure and texture information of the first optimization model to obtain the second optimization model and its texture information and camera pose.
[0124] The detailed implementation manners and beneficial effects of the steps in the method of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0125] In this embodiment, an original image containing a target object sent by a terminal device can be received, and a three-dimensional reconstruction of the target object is performed based on the original image to obtain an initial mesh model; on the basis of the initial mesh model, a differentiable mesh reconstruction technology is introduced to perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; on this basis, a differentiable rendering technology is introduced, and the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model is used as the optimization target to optimize the target reconstruction model to obtain a target mesh model; the target mesh model is sent to the terminal device for the terminal device to use the target mesh model. In this way, based on the differentiable mesh reconstruction and the model optimization process based on the differentiable rendering technology, the geometric structure accuracy of the three-dimensional model can be higher, and the topological structure can be more regular, thereby improving the overall quality of the reconstructed three-dimensional model.
[0126] It should be noted that the execution subject of each step of the method provided in the foregoing embodiment can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 11 to 13 can be device A; for another example, the execution subject of steps 11 and 12 can be device A, and the execution subject of step 13 can be device B; and so on.
[0127] In addition, in some processes described in the foregoing embodiments and the accompanying drawings, a plurality of operations appear in a specific order, but it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 11 and 12 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.
[0128] Figure 7 The structural schematic diagram of a model optimization device provided for another exemplary embodiment of the present application. As Figure 7 shown, the device includes:
[0129] An acquisition module 701, configured to: acquire an initial mesh model of a target object, where the initial mesh model is obtained by performing three-dimensional reconstruction on a raw image including the target object; a reconstruction module 702, configured to: perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; an optimization module 703, configured to: optimize the target reconstruction model with the loss function between the raw image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target, to obtain a target mesh model.
[0130] Optionally, when the reconstruction module 702 performs differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes, it is specifically configured to: determine an initial differentiable parameter and an initial solid unit required to represent a differentiable mesh, where the initial solid unit has a default shape; adjust the shape of the initial solid unit to obtain a target solid unit adapted to the shape of the target object; fit the initial mesh model according to the initial differentiable parameter and the target solid unit to obtain a target reconstruction model, where the target reconstruction model has target differentiable parameters.
[0131] Optionally, when the reconstruction module 702 fits the initial mesh model according to the initial differentiable parameter and the target solid unit to obtain the target reconstruction model, it is specifically configured to: cover the initial mesh model with the target solid unit according to the initial differentiable parameter to obtain an initial reconstruction model; take the similarity error between the initial mesh model and the initial reconstruction model as the optimization target, and use the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model.
[0132] Optionally, when the reconstruction module 702 takes the similarity error between the initial mesh model and the initial reconstruction model as the optimization target and uses the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model, it is specifically configured to: calculate the chamfer distance between the initial reconstruction model and the initial mesh model as the similarity error between the two; in the case where the chamfer distance does not meet the set distance condition, iteratively optimize the differentiable parameters of the initial reconstruction model according to the gradient vector of the chamfer distance; until the chamfer distance meets the set distance condition, take the initial reconstruction model when the chamfer distance meets the set distance condition as the target reconstruction model.
[0133] Optionally, when the optimization module 703 optimizes the target reconstruction model with the loss function between the original image and the rendered image obtained by differentiable rendering of the target reconstruction model as the optimization objective to obtain the target mesh model, it is specifically configured to: input the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; optimize at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model and its corresponding texture information and camera pose; input the first optimized model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering to obtain a second rendered image; optimize at least the geometric structure of the first optimized model according to the second loss function between the second rendered image and the original image to obtain a second optimized model and its corresponding texture information and camera pose; generate the target mesh model according to the second optimized model and its corresponding texture information and camera pose.
[0134] Optionally, before inputting the target reconstruction model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering, the optimization module 703 is further configured to: predict the texture information of the target reconstruction model based on the texture prediction network corresponding to the target object to obtain the texture information corresponding to the target reconstruction model; estimate the camera pose of the original image to obtain the camera pose corresponding to the target reconstruction model.
[0135] Optionally, when the optimization module 703 optimizes at least the topological structure of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model and its corresponding texture information and camera pose, it is specifically configured to: when the first loss function does not meet the first optimization end condition, jointly optimize at least one of the topological structure of the target reconstruction model, the geometric structure, camera pose, and texture prediction network of the target reconstruction model until the first loss function meets the first optimization end condition to obtain the first optimized model and its corresponding texture information and camera pose.
[0136] Optionally, the optimization module 703 is further configured to: collect images of the target object from multiple perspectives as sample images with texture information; train the initial network according to the sample images and the ability of the target reconstruction model to output texture information based on the position coordinates on the target reconstruction model to obtain the texture prediction network.
[0137] Optionally, if the topological structure and camera pose of the first optimization model are not adjustable, when the optimization module 703 optimizes at least the geometric structure of the first optimization model according to the second loss function between the second rendered image and the original image to obtain the second optimization model and its corresponding texture information and camera pose, it is specifically configured to: when the second loss function does not meet the second optimization end condition, optimize the geometric structure and texture information of the first optimization model to obtain the second optimization model and its texture information and camera pose.
[0138] Figure 8 The structural schematic diagram of a 3D reconstruction device provided by another exemplary embodiment of the present application. As Figure 8 shown, the device includes:
[0139] An acquisition module 801, configured to: acquire an original image including a target object, and perform 3D reconstruction on the target object based on the original image to obtain an initial mesh model; a reconstruction module 802, configured to: perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; an optimization module 803, configured to: use the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target, and optimize the target reconstruction model to obtain a target mesh model.
[0140] The detailed implementation manners and beneficial effects of each module in the device of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0141] Figure 9 The structural schematic diagram of a model optimization device provided by another exemplary embodiment of the present application. As Figure 9 shown, the device includes:
[0142] A receiving module 901, configured to: receive the initial mesh model of the target object sent by the terminal device, where the initial mesh model is obtained by performing 3D reconstruction on the original image including the target object; a reconstruction module 902, configured to: perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; an optimization module 903, configured to: use the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target, and optimize the target reconstruction model to obtain a target mesh model; a sending module 904, configured to: send the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0143] The detailed implementation manners and beneficial effects of each module in the device of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0144] Figure 10 The following is a schematic structural diagram of a three-dimensional reconstruction device provided in another exemplary embodiment of this application. As Figure 10 shown, the device includes:
[0145] A receiving module 1001, configured to: receive an original image containing a target object sent by a terminal device, and perform three-dimensional reconstruction on the target object based on the original image to obtain an initial mesh model; a reconstruction module 1002, configured to: perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; an optimization module 1003, configured to: use the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as an optimization target, and optimize the target reconstruction model to obtain a target mesh model; a sending module 1004, configured to: send the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0146] The detailed implementation manners and beneficial effects of each module in the device of this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated herein.
[0147] The internal functions and structures of the model optimization device and the three-dimensional reconstruction device are described above. As Figure 11 shown, in practice, each of the above devices can be implemented as an electronic device, including: a memory 1101, a processor 1102, and a communication component 1103.
[0148] The memory 1101 is used to store computer programs and can be configured to store various other data to support operations on a computing platform. Examples of these data include instructions for any application program or method for operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.
[0149] The memory 1101 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0150] In some alternative embodiments, the processor 1102, which is coupled to the memory 1101, is configured to execute the computer program in the memory 1101 for: obtaining an initial mesh model of a target object, where the initial mesh model is obtained by performing three-dimensional reconstruction on an original image containing the target object; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties; and optimizing the target reconstruction model with the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model to obtain a target mesh model.
[0151] In other alternative embodiments, the processor 1102, which is coupled to the memory 1101, is configured to execute the computer program in the memory 1101 for: obtaining an original image containing a target object and performing three-dimensional reconstruction on the target object based on the original image to obtain an initial mesh model; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties; and optimizing the target reconstruction model with the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model to obtain a target mesh model.
[0152] In other alternative embodiments, the processor 1102, which is coupled to the memory 1101, is configured to execute the computer program in the memory 1101 for: receiving an initial mesh model of a target object sent by a terminal device, where the initial mesh model is obtained by performing three-dimensional reconstruction on an original image containing the target object; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties; optimizing the reconstruction model with the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model to obtain a target mesh model; and sending the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0153] In some other alternative embodiments, the processor 1102, which is coupled to the memory 1101, is configured to execute the computer program in the memory 1101 for: receiving the original image containing the target object sent by the terminal device, and performing three-dimensional reconstruction on the target object based on the original image to obtain an initial mesh model; performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties; optimizing the target reconstruction model with the loss function between the original image and the rendered image obtained by differentiable rendering of the target reconstruction model as the optimization target to obtain a target mesh model; and sending the target mesh model to the terminal device for the terminal device to use the target mesh model.
[0154] Optionally, in each of the above embodiments, when the processor 1102 performs differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable properties, it is specifically configured to: determine the initial differentiable parameters and the initial solid unit required to represent the differentiable mesh, where the initial solid unit has a default shape; adjust the shape of the initial solid unit to obtain a target solid unit adapted to the shape of the target object; and fit the initial mesh model according to the initial differentiable parameters and the target solid unit to obtain a target reconstruction model, where the target reconstruction model has target differentiable parameters.
[0155] Optionally, in each of the above embodiments, when the processor 1102 fits the initial mesh model according to the initial differentiable parameters and the target solid unit to obtain the target reconstruction model, it is specifically configured to: cover the initial mesh model with the target solid unit according to the initial differentiable parameters to obtain an initial reconstruction model; and use the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model with the similarity error between the initial mesh model and the initial reconstruction model as the optimization target to obtain the target reconstruction model.
[0156] Optionally, in each of the above embodiments, when the processor 1102 uses the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model with the similarity error between the initial mesh model and the initial reconstruction model as the optimization target to obtain the target reconstruction model, it is specifically configured to: calculate the chamfer distance between the initial reconstruction model and the initial mesh model as the similarity error between the two; in the case where the chamfer distance does not meet the set distance condition, iteratively optimize the differentiable parameters of the initial reconstruction model according to the gradient vector of the chamfer distance; and until the chamfer distance meets the set distance condition, take the initial reconstruction model when the chamfer distance meets the set distance condition as the target reconstruction model.
[0157] Optionally, in each of the above embodiments, when the processor 1102 optimizes the target reconstruction model with the loss function between the original image and the rendered image obtained by differentiable rendering of the target reconstruction model to obtain the target mesh model, it is specifically configured to: input the target reconstruction model, its corresponding texture information, and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; optimize at least the topology of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model, its corresponding texture information, and camera pose; input the first optimized model, its corresponding texture information, and camera pose into the differentiable renderer for differentiable rendering to obtain a second rendered image; optimize at least the geometric structure of the first optimized model according to the second loss function between the second rendered image and the original image to obtain a second optimized model, its corresponding texture information, and camera pose; generate the target mesh model according to the second optimized model, its corresponding texture information, and camera pose. Among them, the model optimization process based on differentiable rendering technology can make the geometric structure accuracy of the three-dimensional model higher and the topology more regular, thereby improving the overall quality of the reconstructed three-dimensional model.
[0158] Optionally, in each of the above embodiments, before the processor 1102 inputs the target reconstruction model, its corresponding texture information, and camera pose into the differentiable renderer for differentiable rendering, it is further configured to: predict the texture information of the target reconstruction model based on the texture prediction network corresponding to the target object to obtain the texture information corresponding to the target reconstruction model; estimate the camera pose of the original image to obtain the camera pose corresponding to the target reconstruction model.
[0159] Optionally, in each of the above embodiments, when the processor 1102 optimizes at least the topology of the target reconstruction model according to the first loss function between the first rendered image and the original image to obtain a first optimized model, its corresponding texture information, and camera pose, it is specifically configured to: in the case where the first loss function does not meet the first optimization end condition, jointly optimize at least one of the topology of the target reconstruction model, the geometric structure, camera pose, and texture prediction network of the target reconstruction model until the first loss function meets the first optimization end condition to obtain the first optimized model, its corresponding texture information, and camera pose.
[0160] Optionally, in each of the above embodiments, the processor 1102 is further configured to: collect images of the target object from multiple perspectives as sample images with texture information; and train the initial network based on the sample images and the target reconstruction model for the ability to output texture information based on the position coordinates on the target reconstruction model, so as to obtain the texture prediction network.
[0161] Optionally, if the topological structure and camera pose of the first optimization model are not adjustable, when the processor 1102 optimizes at least the geometric structure of the first optimization model according to the second loss function between the second rendered image and the original image to obtain the second optimization model and its corresponding texture information and camera pose, it is specifically configured to: when the second loss function does not meet the second optimization end condition, optimize the geometric structure and texture information of the first optimization model to obtain the second optimization model and its texture information and camera pose.
[0162] Further, as Figure 11 shown, the electronic device further includes: a display 1104, a power supply component 1105, an audio component 1106, and other components. Figure 11 Only some components are schematically shown, which does not mean that the electronic device only includes Figure 11 the components shown. Additionally, Figure 11 the components within the dashed box in Figure 11 are optional components, rather than essential components, and can be determined according to the product form of the working node. The working node of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the working node of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include Figure 11 the components within the dashed box in
[0163] In this embodiment, either obtain the initial mesh model of the target object or perform 3D reconstruction on the original image of the target object to obtain the initial mesh model; on the basis of the initial mesh model, introduce the differentiable mesh reconstruction technology to perform differentiable mesh reconstruction on the initial mesh model to obtain a target reconstruction model that can represent the initial mesh model and has differentiable attributes; on this basis, introduce the differentiable rendering technology, and use the loss function between the original image and the rendered image obtained by performing differentiable rendering on the target reconstruction model as the optimization target to optimize the target reconstruction model to obtain the target mesh model. In this way, based on the differentiable mesh reconstruction and the model optimization process based on the differentiable rendering technology, the geometric structure accuracy of the 3D model can be higher and the topological structure can be more regular, thereby improving the overall quality of the reconstructed 3D model.
[0164] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step executable by an electronic device in the above method embodiment.
[0165] The above memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0166] The above communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.
[0167] The above display includes a screen, and the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0168] The above power supply component provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device where the power supply component is located.
[0169] The above audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operation mode, such as a call mode, a recording mode and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0170] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, compact disc read-only memory (CD-ROM), optical memory, etc.) that contain computer-usable program code.
[0171] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0172] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0174] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and a memory.
[0175] Memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0176] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0177] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0178] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A model optimization method, characterized in that: include: Acquire an initial mesh model of the target object, wherein the initial mesh model is obtained by three-dimensionally reconstructing an original image containing the target object; Performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstructed model that can express the initial mesh model and has differentiable properties; Inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; According to a first loss function between the first rendered image and the original image, at least the topological structure of the target reconstruction model is optimized to obtain a first optimized model and its corresponding texture information and camera pose; Inputting the first optimized model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering to obtain a second rendered image; According to a second loss function between the second rendered image and the original image, at least the geometric structure of the first optimization model is optimized to obtain a second optimization model and its corresponding texture information and camera pose; A target mesh model is generated according to the second optimized model and its corresponding texture information and camera pose.
2. The method according to claim 1, characterized in that Performing a differentiable mesh reconstruction on the initial mesh model to obtain a target reconstructed model that can represent the initial mesh model and has differentiable properties, comprising: Determining initial differentiable parameters and initial solid units required to represent a differentiable grid, wherein the initial solid units have a default shape; Adjusting the shape of the initial three-dimensional unit to obtain a target three-dimensional unit that matches the shape of the target object; The initial mesh model is fitted according to the initial differentiable parameters and the target stereo unit to obtain a target reconstruction model, wherein the target reconstruction model has target differentiable parameters.
3. The method according to claim 2, characterized in that Fitting the initial mesh model according to the initial differentiable parameters and the target stereo unit to obtain the target reconstruction model includes: According to the initial differentiable parameters, the initial mesh model is covered with the target stereo unit to obtain an initial reconstructed model; Taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization target, a gradient descent algorithm is used to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model.
4. The method according to claim 3, characterized in that Taking the similarity error between the initial mesh model and the initial reconstruction model as the optimization target, adopting the gradient descent algorithm to iteratively optimize the differentiable parameters of the initial reconstruction model to obtain the target reconstruction model, including: Calculating the chamfer distance between the initial reconstruction model and the initial mesh model as a similarity error between the two; When the chamfer distance does not meet the set distance condition, iteratively optimizing the differentiable parameters of the initial reconstruction model according to the gradient vector of the chamfer distance; Until the chamfer distance meets the set distance condition, the initial reconstruction model when the chamfer distance meets the set distance condition is used as the target reconstruction model.
5. The method according to claim 1, characterized in that Before inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering, the method further includes: Based on the texture prediction network corresponding to the target object, predicting the texture information of the target reconstruction model, to obtain the texture information corresponding to the target reconstruction model; The camera pose of the original image is estimated to obtain the camera pose corresponding to the target reconstruction model.
6. The method according to claim 5, characterized in that According to a first loss function between the first rendered image and the original image, at least the topological structure of the target reconstruction model is optimized to obtain a first optimized model and its corresponding texture information and camera pose, including: When the first loss function does not satisfy the first optimization termination condition, the topological structure of the target reconstruction model and at least one of the geometric structure, camera pose and texture prediction network of the target reconstruction model are jointly optimized until the first loss function meets the first optimization termination condition to obtain the first optimized model and its corresponding texture information and camera pose.
7. The method according to claim 5, characterized in that Also includes: Collecting images of the target object at multiple viewing angles as sample images with texture information; According to the sample image and the target reconstruction model, the initial network is trained to output texture information based on the position coordinates on the target reconstruction model to obtain the texture prediction network.
8. The method according to claim 1, characterized in that If the topological structure and camera pose of the first optimization model are not adjustable, then according to a second loss function between the second rendered image and the original image, at least the geometric structure of the first optimization model is optimized to obtain a second optimization model and its corresponding texture information and camera pose, including: When the second loss function does not satisfy the second optimization termination condition, the geometric structure and texture information of the first optimization model are optimized to obtain the second optimization model and its texture information and camera pose.
9. A three-dimensional reconstruction method, characterized in that: include: Acquire an original image containing a target object, and perform three-dimensional reconstruction of the target object based on the original image to obtain an initial mesh model; Performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstructed model that can express the initial mesh model and has differentiable properties; Inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; According to a first loss function between the first rendered image and the original image, at least the topological structure of the target reconstruction model is optimized to obtain a first optimized model and its corresponding texture information and camera pose; Inputting the first optimized model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering to obtain a second rendered image; According to a second loss function between the second rendered image and the original image, at least the geometric structure of the first optimization model is optimized to obtain a second optimization model and its corresponding texture information and camera pose; A target mesh model is generated according to the second optimized model and its corresponding texture information and camera pose.
10. A model optimization method, characterized in that: include: Receiving an initial mesh model of a target object sent by a terminal device, wherein the initial mesh model is obtained by three-dimensionally reconstructing an original image containing the target object; Performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstructed model that can express the initial mesh model and has differentiable properties; Inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; According to a first loss function between the first rendered image and the original image, at least the topological structure of the target reconstruction model is optimized to obtain a first optimized model and its corresponding texture information and camera pose; Inputting the first optimized model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering to obtain a second rendered image; According to a second loss function between the second rendered image and the original image, at least the geometric structure of the first optimization model is optimized to obtain a second optimization model and its corresponding texture information and camera pose; Generate a target mesh model according to the second optimized model and its corresponding texture information and camera pose; The target grid model is sent to the terminal device so that the terminal device uses the target grid model.
11. A three-dimensional reconstruction method, characterized in that: include: Receiving an original image containing a target object sent by a terminal device, and performing three-dimensional reconstruction of the target object based on the original image to obtain an initial mesh model; Performing differentiable mesh reconstruction on the initial mesh model to obtain a target reconstructed model that can express the initial mesh model and has differentiable properties; Inputting the target reconstruction model and its corresponding texture information and camera pose into a differentiable renderer for differentiable rendering to obtain a first rendered image; According to a first loss function between the first rendered image and the original image, at least the topological structure of the target reconstruction model is optimized to obtain a first optimized model and its corresponding texture information and camera pose; Inputting the first optimized model and its corresponding texture information and camera pose into the differentiable renderer for differentiable rendering to obtain a second rendered image; According to a second loss function between the second rendered image and the original image, at least the geometric structure of the first optimization model is optimized to obtain a second optimization model and its corresponding texture information and camera pose; Generate a target mesh model according to the second optimized model and its corresponding texture information and camera pose; The target grid model is sent to the terminal device so that the terminal device uses the target grid model.
12. An electronic device, characterized in that: include: A memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to: execute the steps in the method described in any one of claims 1-8, claim 9, claim 10 or claim 11.
13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 8, claim 9, claim 10 or claim 11.
14. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, performs the steps of the method according to any one of claims 1 to 8, claim 9, claim 10 or claim 11.
Citation Information
Patent Citations
Three-dimensional reconstruction method and device, equipment and storage medium
CN117456128A