Training method, device, equipment and storage medium for three-dimensional reconstruction model

By combining real and synthetic training images in the training of the three-dimensional reconstruction model, the generalization and fidelity effect of the three-dimensional reconstruction model is achieved, and the problem of insufficient generalization of the three-dimensional reconstruction model in the existing technology is solved.

CN115115805BActive Publication Date: 2025-05-13SHENZHEN TENCENT COMP SYST CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210869094.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-05-13
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

In the training process of the three-dimensional reconstruction model in the prior art, the training image is too single, resulting in insufficient generalization and poor fidelity in actual use.

Method used

By acquiring multiple training images, including real training images and synthetic training images, the three-dimensional reconstruction model is used to perform semi-supervised learning based on these images, combining the advantages of real training images and synthetic training images, overcome their respective defects and improve training performance.

Benefits of technology

Provide rich image details through real training images and synthesize training images to provide accurate labels, improving the generalization and reconstruction performance of the three-dimensional reconstruction model, and the generated three-dimensional geometric configuration is higher fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115805B_ABST
    Figure CN115115805B_ABST
Patent Text Reader

Abstract

The present application discloses a training method, device, equipment and storage medium for a three-dimensional reconstruction model, and relates to the field of artificial intelligence technology. The method includes: obtaining multiple training images and three-dimensional reconstruction labels corresponding to each training image; wherein the multiple training images include at least one real training image and at least one synthetic training image, the real training image refers to an image obtained by photographing a real target object, and the synthetic training image refers to an image generated according to a three-dimensional model of a synthetic target object; obtaining three-dimensional reconstruction information corresponding to the training image through a three-dimensional reconstruction model according to the training image, and the three-dimensional reconstruction information is used to determine the three-dimensional geometric configuration of the target object in the training image in three-dimensional space; training the three-dimensional reconstruction model according to the three-dimensional reconstruction information and three-dimensional reconstruction labels corresponding to the training image. Training with real images and synthetic images helps to improve the generalization and fidelity of the three-dimensional reconstruction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a training method, apparatus, device, and storage medium for a three-dimensional reconstruction model. Background Art

[0002] By acquiring information related to the three-dimensional space from the two-dimensional image of the target object and generating a three-dimensional geometric configuration corresponding to the target object, it helps to shorten the creation time of the three-dimensional geometric configuration.

[0003] In related technologies, a 3D reconstruction model can be used to obtain the 3D geometric configuration of a target object in an image. Before a 3D reconstruction model can be put into use, it must be trained. This is typically done using synthetic images with real-world labels. This involves fully supervised training of the model's output against the real-world labels of the synthetic images.

[0004] However, when using this method to train a 3D reconstruction model, due to limitations such as the training images being too single, the trained 3D reconstruction model has insufficient generalization and poor fidelity in actual use. Summary of the Invention

[0005] The present invention provides a method, apparatus, device, and storage medium for training a 3D reconstruction model. The technical solution is as follows:

[0006] According to one aspect of an embodiment of the present application, a three-dimensional model reconstruction method is provided, the method comprising:

[0007] Acquire multiple training images and 3D reconstruction labels corresponding to each of the training images; wherein the multiple training images include at least one real training image and at least one synthetic training image, wherein the real training image refers to an image obtained by photographing a real target object, and the synthetic training image refers to an image generated based on a synthetic 3D model of the target object;

[0008] Obtaining three-dimensional reconstruction information corresponding to the training image using the three-dimensional reconstruction model based on the training image, wherein the three-dimensional reconstruction information is used to determine a three-dimensional geometric configuration of a target object in the training image in a three-dimensional space;

[0009] The three-dimensional reconstruction model is trained according to the three-dimensional reconstruction information and the three-dimensional reconstruction label corresponding to the training image.

[0010] According to one aspect of an embodiment of the present application, a training device for a three-dimensional reconstruction model is provided, the device comprising:

[0011] an image acquisition module, configured to acquire a plurality of training images and a 3D reconstruction label corresponding to each of the training images; wherein the plurality of training images include at least one real training image and at least one synthetic training image, wherein the real training image is an image obtained by photographing a real target object, and the synthetic training image is an image generated based on a synthetic 3D model of the target object;

[0012] an information generation module, configured to obtain, based on the training image and using the 3D reconstruction model, 3D reconstruction information corresponding to the training image, wherein the 3D reconstruction information is used to determine a 3D geometric configuration of a target object in the training image in a 3D space;

[0013] The model training module is used to train the 3D reconstruction model according to the 3D reconstruction information and 3D reconstruction labels corresponding to the training images.

[0014] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above method.

[0015] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above method.

[0016] According to one aspect of an embodiment of the present application, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal device to perform the above method.

[0017] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:

[0018] Real training images are used to generate 3D reconstruction labels and then used in semi-supervised learning of the 3D reconstruction model with synthetic training images. On the one hand, because real training images provide richer image details and are easily accessible, incorporating real training images into model training improves the generalization of the network model and enables the reconstruction of high-fidelity 3D geometric configurations. On the other hand, because the 3D reconstruction labels corresponding to real training images are predicted, while those corresponding to synthetic training images are fixed, training the 3D reconstruction model with a mixture of real and synthetic training images helps overcome the shortcomings of each type of training sample and improves the reconstruction performance of the trained 3D reconstruction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a schematic diagram of an implementation environment for a solution provided by an embodiment of the present application;

[0020] Figure 2 This is a schematic diagram of an application scenario of a three-dimensional reconstruction model provided by an embodiment of the present application;

[0021] Figure 3 This is a flowchart of a training method for a three-dimensional reconstruction model provided by one embodiment of the present application;

[0022] Figure 4 is a schematic diagram of a three-dimensional reconstruction model training method provided by an exemplary embodiment of the present application;

[0023] Figure 5 is a schematic diagram of a three-dimensional reconstruction label generation process provided by an exemplary embodiment of the present application;

[0024] Figure 6 is a schematic diagram of a 3D reconstruction model training process provided by an exemplary embodiment of the present application;

[0025] Figure 7 is a schematic diagram of a training result provided by an embodiment of the present application;

[0026] Figure 8 is a schematic diagram of a training result provided by an embodiment of the present application;

[0027] Figure 9 This is a block diagram of a training device for a three-dimensional reconstruction model provided by one embodiment of the present application;

[0028] Figure 10 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0030] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0031] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0032] Computer vision (CV) is the science of making machines "see." Specifically, it refers to using cameras and computers to replace the human eye in identifying and measuring objects, and then further processing them to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0033] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning-based teaching.

[0034] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart wearable devices, virtual assistants, smart marketing, smart medical care, smart creation of 3D models, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0035] The solutions provided in the embodiments of this application involve technologies such as artificial intelligence computer vision, which are specifically illustrated by the following embodiments.

[0036] Before introducing the embodiments of the present application, in order to facilitate understanding of the present solution, the following explanations are given for the terms appearing in the present solution.

[0037] Synthetic Data: The synthetic data described in the examples of this application is not based on real objects, but rather simulates real-world data. In scenarios where collecting real data would be dangerous, using synthetic data to train the model can significantly reduce the risk.

[0038] Real data: The real data described in the embodiments of the present application can be understood as data obtained based on real objects. For example, if you take a photo of the person in front of you, the photo can be considered as real data.

[0039] Ground Truth Labeling: The label of the annotated data, which is rendered based on the synthetic data.

[0040] Pseudo Labeling: refers to labels that are not actually marked by humans. For example, they can be the results of predictions from another trained model and used as supervisory signals during training.

[0041] Orthographic projection transformation: Use a rectangular box as the frame and project the scene onto the front of this box. This projection does not produce a perspective contraction effect (farther objects appear smaller on the image plane) because it ensures that parallel lines remain parallel after the transformation, which also keeps the relative distances between objects unchanged. Simply put, the orthographic projection transformation ignores the size changes of objects at different distances and projects the objects at their original proportions onto a cross-section (such as a display screen). A camera that achieves this effect is called an orthographic projection camera, also known as an orthographic camera.

[0042] Perspective projection transformation: Like orthographic projection, it projects a spatial volume (a perspective pyramid with the projection center as its vertex) onto a two-dimensional image plane. However, it exhibits a foreshortening effect: objects farther away appear smaller on the image plane than objects of the same size closer. Unlike orthographic projection, perspective projection does not preserve the relative size of distances and angles, so the projections of parallel lines are not necessarily parallel. In other words, the perspective projection transformation can make an object appear larger when close to the player and smaller when farther away. A camera that achieves this effect is called a telephoto camera. Telephoto cameras are often used in 3D game development. They work by scaling the projection (i.e., the size of the cross-section) based on the distance between the camera and the object. Perspective projection closely resembles the principle by which the human eye or camera lens produces images of the three-dimensional world. The essential difference between the two projection methods is that the distance between the projection center and the projection plane is finite in perspective projection, while it is infinite in parallel projection.

[0043] Depth Image: Also known as Range Image, it refers to the distance (depth) from the image collector to each point in the scene as a pixel value, which directly reflects the geometry of the visible surface of the scene. Depth images can be calculated as point cloud data through coordinate conversion. Point cloud data with regularity and necessary information can also be inversely calculated into depth image data. Each pixel in the depth image represents the distance from the coordinate of the specific object to the part of the object closest to the camera plane in the depth sensor's field of view.

[0044] Voxel: Short for volume pixel, a volume containing voxels can be represented by volume rendering or by extracting polygonal isosurfaces with a given threshold outline. As the name suggests, voxels are the smallest unit of digital data used in three-dimensional space. They are used in fields such as 3D imaging, scientific data, and medical imaging. Conceptually, they are similar to the smallest unit of two-dimensional space, the pixel, used in image data in 2D computer graphics. Some true 3D displays use voxels to describe their resolution; for example, a display capable of displaying 512×512×512 voxels.

[0045] Feature Volume: Similar to a feature vector, each grid (i.e., voxel) in three-dimensional space has its own corresponding feature vector. In the embodiment of the present application, the feature vector is obtained through a deep neural network.

[0046] Multilayer Perceptron (MLP): It is a forward-structured artificial neural network that maps a set of input vectors to a set of output vectors.

[0047] Signed distance field: A signed distance function (or signed distance function) for a set Ω in metric space determines the distance of a given point x from the boundary of Ω, with its sign depending on whether x is inside Ω. The function has a positive value at a point x inside Ω, decreases as x approaches the boundary of Ω (where the signed distance function is zero), and takes negative values ​​outside Ω.

[0048] Truncated Signed Distance Function (TSDF): A truncated signed distance field. Compared to a signed distance field, this function has a maximum and minimum value. When the function value exceeds or falls below a certain value, the function value will be replaced.

[0049] 3D geometry: Also known as 3D human geometry or 3D human mesh. This untextured 3D human model contains only the geometric information of the human surface, represented by points and triangular meshes.

[0050] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application. The implementation environment of the solution may include: a terminal device 10 and a server 20.

[0051] The terminal device 10 includes but is not limited to mobile phones, tablet computers, intelligent voice interaction devices, game consoles, wearable devices, multimedia playback devices, PCs (Personal Computers), vehicle-mounted terminals, smart home appliances, and other electronic devices. The client of the target application can be installed in the terminal device 10.

[0052] In an embodiment of the present application, the above-mentioned target application can be any application that can provide image processing functions. Typically, the application is an image processing application. Of course, in addition to image processing applications, image processing services can also be provided in other types of applications, such as news applications, shopping applications, social applications, interactive entertainment applications, browser applications, shopping applications, content sharing applications, virtual reality (VR) applications, augmented reality (AR) applications, etc., which are not limited in this embodiment of the present application. In addition, for different applications, the types of images processed may be different, and the corresponding functions may also be different, which can be pre-configured according to actual needs, which are not limited in this embodiment of the present application. Optionally, the client of the above-mentioned application runs in the terminal device 10.

[0053] The server 20 is used to provide background services for the client of the target application in the terminal device 10. For example, the server 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but is not limited thereto.

[0054] The server 20 has at least data reception and processing capabilities, enabling communication between the terminal device 10 and the server 20 via a network. The network can be either a wired or wireless network. The server 20 receives the image to be processed from the terminal device 10 and processes the image to obtain the corresponding three-dimensional geometric configuration.

[0055] In some embodiments, before using the 3D reconstruction model to generate the 3D geometric configuration corresponding to the image, the 3D reconstruction model needs to be trained. In the method provided in the embodiment of the present application, the execution subject of each step can be a computer device. The computer device can be any electronic device with data storage and processing capabilities. For example, the computer device can be Figure 2 The server 20 in the example may be Figure 2 The terminal device 10 in the example may also be another device other than the terminal device 10 and the server 20.

[0056] Please refer to Figure 2 , which shows a schematic diagram of an application scenario of a three-dimensional reconstruction model provided by an embodiment of the present application.

[0057] like Figure 2 As shown, a real-life human image is captured using a webcam or camera. The model trained using this method quickly and efficiently reconstructs high-fidelity 3D human geometry with clothing. The reconstructed 3D human geometry has the following potential applications: one-click generation of digital humans in the metaverse, virtual clothing reconstruction, one-click costume changes, and film, television, and animation production.

[0058] In the process of 3D film and television and animation production, modelers are required to build 3D human body models from scratch, and the production cycle of 3D human body models is long and the cost is correspondingly high. According to the technical solution provided in the embodiment of the present application, it is only necessary to simply take a human body image to quickly obtain 3D human body geometry. Subsequently, the modeler only needs to make slight adjustments to the human body geometry to obtain a high-quality 3D human body model, which greatly shortens the production cycle. With the development of related technologies, the image quality in many large-scale 3D games has gradually approached the real world. This method can reconstruct high-fidelity three-dimensional human body geometry, increasing the feasibility of projecting real-life users into the virtual game world.

[0059] In the embodiments of the present application, training the 3D reconstruction model using a mixture of real and synthetic data helps overcome the poor fit of the 3D reconstruction model resulting from training the model solely with synthetic data. Furthermore, synthetic data is also used during the training of the 3D reconstruction model. Since synthetic data can more accurately label data, training the 3D reconstruction model using both real and mixed data helps improve the accuracy of the 3D human body geometry obtained by the trained 3D reconstruction model.

[0060] In some embodiments, the target object is a woman wearing a skirt. If only synthetic data is used to simulate the image of the woman, the representation of many details such as the folds of the skirt will be insufficient, because there are still differences in details between the simulated human body and the real human body. Therefore, in the embodiment of the present application, real data taken by the camera is used to obtain a three-dimensional human body mesh. Considering that real data can include objects that cannot be represented by synthetic data, and real data can more accurately represent the target object, the network model trained with real data (referring to the three-dimensional reconstruction model in the embodiment of the present application) has relatively better generalization and accuracy. The three-dimensional human body mesh obtained by the network model is certainly better than other network models trained with synthetic data in terms of details.

[0061] Please refer to Figure 3 , which shows a flow chart of a training method for a three-dimensional reconstruction model provided by an embodiment of the present application. The execution subject of each step of the method can be Figure 1 The terminal device 10 in the implementation environment of the solution shown in the figure can be a client of the target application or a Figure 1 The server 20 in the implementation environment of the solution shown. In the following method embodiment, for ease of description, only the execution subject of each step is introduced as a "computer device". The method can include at least one of the following steps (310-330):

[0062] Step 310: Acquire multiple training images and three-dimensional reconstruction labels corresponding to each training image; wherein the multiple training images include at least one real training image and at least one synthetic training image, the real training image refers to an image obtained by photographing a real target object, and the synthetic training image refers to an image generated based on a synthetic three-dimensional model of the target object.

[0063] In some embodiments, the training images are used to train the 3D reconstruction model, so that the trained 3D reconstruction model can obtain a relatively accurate 3D geometric structure from the images.

[0064] In some embodiments, the 3D reconstruction labels of the training images are used to represent surface information of the geometric configuration of the target object in the training images in 3D space. The 3D reconstruction labels can be used to assess deviations in the output of the 3D reconstruction model during training, thereby adjusting parameters in the 3D reconstruction model.

[0065] The target object is any object in the real world, including but not limited to at least one of the following: human body, human face, animal, scenery, virtual character, etc.

[0066] In some embodiments, the target objects in different training images are not identical, and their geometric configurations in 3D space are not identical. Therefore, the 3D reconstruction labels corresponding to training images with different target objects differ. In some embodiments, the method for obtaining the 3D reconstruction labels depends on the type of training image.

[0067] In some embodiments, real training images refer to images obtained by photographing. For example, real training images obtained by photographing people in a real environment can also be referred to as real person images. In some embodiments, a computer device can obtain real training images through external input or by downloading images stored on a server as real training images.

[0068] In some embodiments, different real training images are captured using different cameras or with different camera parameters. Using real training images obtained using different capture methods in the training process of the 3D reconstruction model helps improve the trained 3D reconstruction model's ability to process images captured from different angles, thereby improving the generalization of the trained 3D reconstruction model.

[0069] In some embodiments, the 3D reconstructed labels corresponding to the real training images are obtained by processing the real training images. In some embodiments, the 3D reconstructed labels corresponding to the real training images are obtained using a trained deep learning model. Because the 3D reconstructed labels obtained through the deep learning model differ from the surface of the target object in 3D space in the real training images, the 3D reconstructed labels corresponding to the real training images can only approximately represent the 3D geometric configuration of the target object. In some embodiments, the 3D reconstructed labels corresponding to the real training images are referred to as pseudo labels.

[0070] The 3D reconstruction labels corresponding to the real training images can be generated by a computer device. For example, after the computer device acquires the real training images, it processes the real training images to obtain the 3D reconstruction labels corresponding to the real training images. For details on this process, please refer to the examples below. The above method allows arbitrary real training images to be selected for training during the model training process, which helps improve the completeness of the model training method.

[0071] The 3D reconstruction labels corresponding to the real training images can also be generated by other devices. For example, other devices process a batch of real training images, obtain the 3D reconstruction labels corresponding to each real training image, and send the real training images and the corresponding 3D reconstruction labels to the computer device. The above method helps to reduce the calculation steps of the computer device in the model training process and speed up the training of the 3D reconstruction model. It should be noted that the execution entity for generating the 3D reconstruction labels of the real training images and the generation time of the 3D reconstruction labels are determined according to actual conditions, and this application does not limit them here.

[0072] In some embodiments, a synthetic training image refers to an image generated using a virtual synthesis method. In some embodiments, the synthetic training image can be obtained by generating a virtual human mesh and rendering the virtual human mesh. In some of the above, the 3D reconstructed labels corresponding to the synthetic training images can be calculated using the virtual human mesh. For details on this process, please refer to the embodiments below. In some embodiments, the 3D reconstructed labels corresponding to the synthetic training images can be referred to as true labels.

[0073] In some embodiments, a computer device obtains a synthetic training image and a three-dimensional reconstruction label corresponding to the synthetic training image through a synthetic image generating device; wherein the synthetic image generating device can be a computer device or other device.

[0074] In step 320 , three-dimensional reconstruction information corresponding to the training image is obtained based on the training image using a three-dimensional reconstruction model. The three-dimensional reconstruction information is used to determine the three-dimensional geometric configuration of the target object in the training image in the three-dimensional space.

[0075] In some embodiments, 3D reconstruction information refers to information in 3D space, used to determine the 3D geometric configuration of a target object in 3D space. Specifically, the 3D reconstruction information is the distribution of the target object's surface (or critical surface) in 3D space, as estimated by the 3D reconstruction model.

[0076] Optionally, the 3D reconstruction information includes coordinate information of points in 3D space and distance information from the points to the surface of the object. Based on this distance information, the 3D geometric configuration of the target object in 3D space can be further determined. In some embodiments, the computer device uses an isosurface extraction algorithm to extract the surface of the target object based on the 3D reconstruction information, thereby determining the 3D geometric configuration.

[0077] In some embodiments, the 3D reconstruction information and the 3D reconstruction labels corresponding to the training images are represented in the same manner. For example, the representation of the 3D reconstruction information and the 3D reconstruction labels includes, but is not limited to, at least one of the following: an occupancy field, a signed distance field, or a directed truncated distance field. For example, the 3D reconstruction information and the 3D reconstruction labels are both represented in the occupancy field manner. For another example, the 3D reconstruction information and the 3D reconstruction labels are represented in the directed truncated distance field manner.

[0078] In some embodiments, the 3D reconstruction model can estimate surface information of the target object's geometric configuration in 3D space based on the 2D information in the training image. For details on how the 3D reconstruction model estimates the training image to obtain the corresponding 3D reconstruction information, please refer to the following embodiments.

[0079] Step 330 : Training the 3D reconstruction model according to the 3D reconstruction information and 3D reconstruction labels corresponding to the training images.

[0080] In some embodiments, after generating 3D reconstruction information corresponding to training images, parameters in the 3D reconstruction model are adjusted by calculating the difference between the 3D reconstruction information corresponding to the training images and the 3D reconstruction labels. In some embodiments, the computer device uses regularization constraints to minimize the difference between the 3D reconstruction information obtained during training and the corresponding 3D reconstruction labels.

[0081] In some embodiments, when the gap between the 3D reconstruction information and the 3D reconstruction label meets a preset condition, the training process of the 3D reconstruction model is completed.

[0082] In some embodiments, the computer device trains the 3D reconstruction model according to training batches. In some training batches, the computer device selects at least one training image to input into the 3D reconstruction model and obtains 3D reconstruction information corresponding to each of the at least one training images. In some embodiments, the at least one training image is a real training image. In other embodiments, the at least one training image is a synthesized training image. In still other embodiments, the at least one training image includes both real and synthesized training images. For details on this process, please refer to the following examples.

[0083] Figure 4It is a schematic diagram of a training method for a three-dimensional reconstruction model provided by an exemplary embodiment of the present application. The computer device uses a mixture of synthetic training images and real training images to train the three-dimensional reconstruction model. Before training the three-dimensional reconstruction model, the computer device obtains at least one training image and training labels corresponding to the training images. For example, the computer device can obtain different types of training data at the same time, or can obtain different types of training data one after another, and the present application does not limit this. In the process of training the three-dimensional reconstruction model, the computer device inputs at least one training image into the three-dimensional reconstruction model, and the three-dimensional reconstruction model processes the training image to obtain three-dimensional reconstruction information corresponding to the training image. The training loss of the model is determined by the three-dimensional reconstruction information and the three-dimensional reconstruction label corresponding to the training image, and the three-dimensional reconstruction model is adjusted by the training loss of the model. When the three-dimensional reconstruction information and the three-dimensional reconstruction label corresponding to the training image meet the training conditions, a trained three-dimensional reconstruction model is obtained.

[0084] In some embodiments, the trained three-dimensional reconstruction model is used for at least one of the following: generating a three-dimensional geometric configuration of a human body in a three-dimensional virtual scene based on a real color image of the human body; generating a three-dimensional geometric configuration of a digital human in a three-dimensional virtual scene based on a real color image of the digital human; generating a three-dimensional geometric configuration of clothes in a three-dimensional virtual scene based on a real color image of clothes; generating a three-dimensional geometric configuration of a human face in a three-dimensional virtual scene based on a real color image of the human face.

[0085] In some embodiments, such as a human body simulation game, a real human body needs to be projected into a virtual environment. Therefore, a photo of the human body in the real world can be taken first. After taking the photo of the human body in the real world and obtaining a depth image of the human body, a three-dimensional geometric configuration of the human body can be generated based on the depth image through a three-dimensional reconstruction model. That is, a three-dimensional geometric configuration of the human body corresponding to the real human body is generated in the virtual environment, so that the user's gaming experience can also be better.

[0086] In some embodiments, such as a metaverse digital human scene, multiple digital humans need to be generated. Similarly, the geometric configuration of the digital humans can be generated in the virtual scene through a three-dimensional reconstruction model based on the real image of the digital humans.

[0087] In some embodiments, such as in scenarios such as changing clothes, the three-dimensional geometric configuration of the clothes can be generated through a three-dimensional reconstruction model based on the real color image of the clothes. Different clothes correspond to different three-dimensional geometric configurations, so clothes can be changed.

[0088] In some embodiments, such as artificial intelligence face-changing technology, a real human face can also be used to reconstruct a three-dimensional model to obtain the three-dimensional geometric configuration of the face, and then apply it to where the face needs to be changed.

[0089] The three-dimensional reconstruction model trained in the embodiment of the present application can be applied to a wide range of scenarios. The three-dimensional reconstruction model can generate a three-dimensional geometric configuration corresponding to the real target object, which can be widely used in various scenarios such as games and animation production. It can not only improve the fineness of the generated geometric configuration, but also improve the user experience.

[0090] In summary, the technical solution provided in this application uses real training images to generate 3D reconstruction labels and conducts semi-supervised learning of the 3D reconstruction model with synthetic training images. On the one hand, because real training images can provide richer image details and are easy to obtain, adding real training images to the model training improves the generalization of the network model and reconstructs high-fidelity 3D geometric configurations. On the other hand, because the 3D reconstruction labels corresponding to real training images are obtained through prediction, while the 3D reconstruction labels corresponding to synthetic training images are determined, training the 3D reconstruction model by mixing real training images with synthetic training images helps overcome the shortcomings of each type of training sample and improves the reconstruction performance of the trained 3D reconstruction model.

[0091] The following describes the training process of the 3D reconstruction model through several embodiments.

[0092] In some embodiments, a computer device trains a three-dimensional reconstruction model based on the three-dimensional reconstruction information and three-dimensional reconstruction labels corresponding to the training images, including: the computer device calculates a first training loss based on the three-dimensional reconstruction information and three-dimensional reconstruction labels corresponding to the real training images, and the first training loss is used to indicate the difference between the three-dimensional reconstruction information and the three-dimensional reconstruction labels corresponding to the real training images; the computer device calculates a second training loss based on the three-dimensional reconstruction information and the three-dimensional reconstruction labels corresponding to the synthetic training images, and the second training loss is used to indicate the difference between the three-dimensional reconstruction information and the three-dimensional reconstruction labels corresponding to the synthetic training images; the computer device adjusts the parameters of the three-dimensional reconstruction model based on the first training loss and the second training loss.

[0093] In some embodiments, the first training loss and the second training loss are calculated similarly. In some embodiments, the first training loss and the second training loss can both be referred to as training losses, and are used to represent the difference between the 3D reconstruction information highlighted by the training and the 3D reconstruction label.

[0094] In some embodiments, the computer device determines the training loss based on the 3D reconstruction information and 3D reconstruction labels corresponding to the training images. Alternatively, in some embodiments, the computer device determines the training loss of the model based on the mean square error or the absolute value of the difference between the 3D reconstruction information and the 3D reconstruction labels corresponding to the training images.

[0095] In some embodiments, regularization is used to constrain the 3D reconstruction information and 3D reconstruction labels corresponding to the training images. Using regularization to constrain the 3D reconstruction information and 3D reconstruction labels corresponding to the training images can effectively improve generalization capabilities while preventing overfitting.

[0096] In some embodiments, the computer device performs regularization processing on the three-dimensional reconstruction information and three-dimensional reconstruction labels corresponding to the training images to determine the training loss. In some embodiments, the regularization processing includes but is not limited to: using the L1 norm regularization method or the L2 norm regularization method. L1 regularization and L2 regularization can be regarded as penalty terms of the loss function. The so-called penalty refers to some restrictions on certain parameters in the loss function. The specific L1 norm regularization method or the L2 norm regularization method are not described in detail in this application. Based on the technical solution provided in the embodiment of the present application, the result of the L2 norm regularization method for model training is slightly better than the result of the L1 norm regularization method for model training. Using L1 and L2 norms to calculate the model loss helps to speed up the training speed of the model and increase the time taken for the three-dimensional reconstruction model to reach a convergence state.

[0097] It should be noted that this application does not limit the function for determining the training loss, nor does it limit the method for adjusting the model parameters based on the loss.

[0098] In some embodiments, the computer device adjusts the parameters of the three-dimensional reconstruction model based on the first training loss and the second training loss, including: the computer device performs a weighted summation of the first training loss and the second training loss to obtain a total training loss; and adjusts the parameters of the three-dimensional reconstruction model based on the total training loss.

[0099] In some embodiments, real training images and synthetic training images have different image characteristics. For example, real training images are more diverse and easier to obtain, and the target objects in real training images are richer in detail. For example, if the target object is a person, the texture details of the target object's clothing in the real training images obtained through photography are more realistic. However, synthetic training images may have fewer target object details. The 3D reconstruction labels corresponding to synthetic training images are calculated, so the 3D reconstruction labels corresponding to synthetic training images are more accurate. Therefore, different weightings can be set for real training images and synthetic training images.

[0100] In some embodiments, the first training loss and the second training loss are weightedly summed to obtain a total training loss, including determining a first weight corresponding to the real training image and a second weight corresponding to the synthetic training image; wherein the first weight is not equal to the second weight, the first training loss is weighted by the first weight, and the total training loss is obtained by weighting the second training loss by the second weight.

[0101] In some embodiments, there is a correspondence between the 3D reconstruction label and the weighted weight, and the computer device determines the corresponding weighted weight based on the 3D reconstruction label. It should be noted that the first weight and the second weight can be determined according to actual needs and are not limited in this application.

[0102] In some embodiments, the computer device adjusts the model parameters of the three-dimensional reconstruction model based on the total training loss. In some embodiments, the values ​​of the first weight and the second weight are related to the degree of adjustment of the model parameters. In some embodiments, the first weight and the second weight are positively correlated with the degree of adjustment of the model parameters, that is, the larger the values ​​of the first weight and the second weight, the larger the adjustment parameters of the model parameters. The smaller the values ​​of the first weight and the second weight, the smaller the adjustment parameters of the model parameters.

[0103] By using different weighting weights to process the first training loss and the second training loss, different types of training images can have different degrees of influence on the model parameters. During the actual training process, the values ​​of the weighting weights can be set according to the performance requirements of the 3D reconstruction model after training. This method helps to improve the flexibility of the 3D reconstruction model training process.

[0104] In some embodiments, the three-dimensional reconstruction information and the three-dimensional reconstruction label are represented in the form of a signed distance field, and the signed distance field is used to characterize the distance between at least one spatial point and the three-dimensional geometric configuration surface corresponding to the target object; the method also includes: a computer device obtains three-dimensional reconstruction information and three-dimensional reconstruction labels represented in the form of an occupancy field based on the three-dimensional reconstruction information and three-dimensional reconstruction labels represented in the form of a signed distance field; wherein the occupancy field is used to characterize the internal and external relationship between the three-dimensional geometric configuration surface corresponding to the at least one spatial point and the target object; the computer device trains the three-dimensional reconstruction model based on the three-dimensional reconstruction information and three-dimensional reconstruction labels corresponding to the training images, including: training the three-dimensional reconstruction model based on the three-dimensional reconstruction information and three-dimensional reconstruction labels represented in the form of a signed distance field corresponding to the training images, and the three-dimensional reconstruction information and three-dimensional reconstruction labels represented in the form of an occupancy field corresponding to the training images.

[0105] In some embodiments, the 3D reconstruction information and 3D reconstruction labels are represented using a signed distance field. Specifically, the 3D reconstruction information and 3D reconstruction labels are represented using a directed truncated distance field. In some embodiments, the signed distance field is used to represent the distance between at least one spatial point in the 3D space corresponding to the training image and a surface (critical surface) of the 3D geometric configuration of the target object.

[0106] In some embodiments, a signed distance field can be understood as follows: Ω represents the surface (critical surface) of the target object in three-dimensional space. The distance value corresponding to a spatial point x within Ω is positive, and the distance value decreases as the spatial point x approaches the boundary of Ω. When the spatial point x is outside Ω, the distance value of the spatial point is negative. In some embodiments, the distance value of the spatial point is represented by a distance function.

[0107] When the 3D reconstruction information and 3D reconstruction labels are represented using a directed phased distance field, if the distance function determines that the distance of a model null point exceeds the distance range, the computer device replaces the distance function to ensure that the distance function best fits the distance between each spatial point and the surface of the 3D geometric configuration of the target object. In some embodiments, the distance range of the directed truncated distance field is set to -0.8 to 0.8. This application does not limit the value of the directed truncated distance field.

[0108] In some embodiments, classifying the 3D reconstruction information represented in the form of a signed distance field refers to classifying at least one spatial point in the 3D space corresponding to the training image based on the 3D reconstruction information in the form of a signed distance field. In some embodiments, the spatial points are classified based on their positional relationship with the surface of the 3D geometric configuration of the target object. For example, a Boolean classification is performed on the at least one spatial point, where a classification value of 0 is assigned to a spatial point within the 3D geometric configuration of the target object, and a classification value of 1 is assigned to a spatial point outside the 3D geometric configuration of the target object.

[0109] In some embodiments, the 3D reconstruction information represented by a signed distance field obtained by the above classification method is represented by an occupancy field. It should be noted that other classification methods can also be used to classify the 3D reconstruction information and 3D reconstruction labels represented by a signed distance field, such as assigning a classification value of -1 to spatial points within the target object's 3D geometric configuration and a classification value of 1 to spatial points outside the target object's 3D geometric configuration, and this application does not limit this.

[0110] In some embodiments, the computer device trains the three-dimensional reconstruction model based on the three-dimensional reconstruction information and three-dimensional reconstruction labels represented in the form of occupancy fields corresponding to the training images. For the specific process, please refer to the above embodiments and will not be repeated here.

[0111] In some embodiments, a computer device calculates a model loss using the 3D reconstruction information and 3D label information represented in the form of a signed distance field, and trains the 3D reconstruction model based on the model loss. For details, please refer to the above embodiments. In this case, the 3D reconstruction information of the training image can be represented in the form of a signed distance field.

[0112] In other embodiments, a computer device calculates a model loss using the 3D reconstruction information and 3D label information represented in occupancy form, and trains the 3D reconstruction model based on the model loss. For details, please refer to the above embodiments. In this case, the 3D reconstruction information of the training image can be represented in the form of an occupancy field. Alternatively, the 3D reconstruction information of the training image can be represented in the form of a signed distance field (or a directed truncated distance field). When the 3D reconstruction information can be represented in the form of a signed distance field (or a directed truncated distance field), the 3D reconstruction information is classified and processed to obtain 3D reconstruction information represented in the form of an occupancy field. The same applies to the 3D reconstruction labels. For the details of this process, please refer to the above embodiments and will not be described in detail here.

[0113] In other embodiments, in order to improve the training quality, the computer device adjusts the model parameters of the three-dimensional reconstruction model twice in one training batch to speed up the training of the model.

[0114] For example, during the first training process, the computer device trains the 3D reconstruction model based on the 3D reconstruction information and 3D reconstruction labels corresponding to the training images, which are represented in the form of signed distance fields. During the second training process, the computer device trains the 3D reconstruction model based on the 3D reconstruction information and 3D reconstruction labels corresponding to the training images, which are represented in the form of occupancy fields.

[0115] For another example, during the first training process, the computer device trains the 3D reconstruction model based on the 3D reconstruction information and 3D reconstruction labels corresponding to the training images, which are represented in the form of occupancy fields. During the second training process, the computer device trains the 3D reconstruction model based on the 3D reconstruction information and 3D reconstruction labels corresponding to the training images, which are represented in the form of signed distance fields.

[0116] When the 3D reconstruction information and the 3D reconstruction label are represented in the form of a directed (truncated) distance field, during the first training process, the training loss is calculated by the difference between the 3D reconstruction information and the 3D reconstruction label. Due to the distance sign in the directed distance field, suppose that the distance value of spatial point A in the 3D reconstruction information is -0.2, and the distance value of spatial point A in the 3D reconstruction label is -0.1. Then, for spatial point A, the difference between the 3D reconstruction information and the 3D reconstruction label is 0.1. Suppose there is a spatial point B, the distance value of spatial point B in the 3D reconstruction information is -0.05, and the distance value of spatial point B in the 3D reconstruction label is 0.05. Then, for spatial point B, the difference between the 3D reconstruction information and the 3D reconstruction label is 0.1.

[0117] However, for spatial point A, whether in the 3D reconstruction information or the 3D reconstruction label, this spatial point A is outside the surface of the 3D geometric configuration of the target object, while for spatial point B, in the 3D reconstruction information, this point is outside the surface of the 3D geometric configuration of the target object; in the 3D reconstruction label, this point is inside the surface of the 3D geometric configuration of the target object. By classifying the 3D reconstruction information and 3D reconstruction labels, it helps to avoid loss calculation defects caused by the sign of the numerical value in the signed distance field representation, helps to improve the training effect of the 3D reconstruction model, and improves the accuracy of the 3D geometric configuration generated by the trained 3D reconstruction model.

[0118] The following describes the process of obtaining training images and 3D reconstruction labels corresponding to the training images through several embodiments.

[0119] First, we introduce a process for obtaining 3D reconstruction labels corresponding to real training images. As we can see from the above content, real training images are images obtained by photographing real target objects. Therefore, it is necessary to process the real training images to obtain the 3D reconstruction labels corresponding to the real training images.

[0120] In some embodiments, a computer device obtains multiple training images and three-dimensional reconstruction labels corresponding to each training image, including: the computer device uses a depth map prediction model to generate a predicted depth image corresponding to the real training image; the computer device performs spatial transformation on the predicted depth image to obtain three-dimensional point cloud data corresponding to the real training image; wherein the spatial transformation is used to achieve conversion from two-dimensional space to three-dimensional space; the computer device samples the three-dimensional point cloud data to obtain a three-dimensional reconstruction label corresponding to the real training image.

[0121] Optionally, the real training image is an RGB image, which contains position information and color information of each pixel.

[0122] In some embodiments, a depth map prediction model is a model that processes an input image to generate a depth map corresponding to the input image. In some embodiments, the depth map prediction model is a machine learning model. The depth map prediction model includes a pixel-to-pixel algorithm model.

[0123] In some embodiments, the depth map prediction model includes more than one network layer. In one embodiment, the depth map prediction model includes a first conversion network layer and a second conversion network layer; wherein the first conversion network layer is used to convert the real training image into an intermediate conversion image, and the second conversion network layer is used to obtain a predicted depth image based on the intermediate conversion image.

[0124] In some embodiments, the intermediate conversion image is used to represent the normal information of the contours in the real training image. In some embodiments, the intermediate conversion image is a predicted normal image corresponding to the real training image.

[0125] In some embodiments, the first conversion network layer and the second conversion network layer are connected in series, that is, the input of the first conversion network layer serves as the output of the second conversion network layer. For example, a computer device inputs a real training image into the depth map prediction model, processes the real training image through the first conversion network layer to obtain a predicted normal phase image, and the first conversion network layer passes the predicted normal phase image to the second conversion network layer. The second network conversion layer processes the predicted normal phase image and the real training image to obtain a predicted depth image.

[0126] In some embodiments, the predicted depth image is used to represent the depth information of at least one spatial point in the actual training image. For example, the predicted depth image is used to represent the distance information between the at least one spatial point and the camera. In some embodiments, the first and second conversion network layers also belong to the network structure of the (pixel2pixel) algorithm.

[0127] In some embodiments, after the computer device uses the depth map prediction model to generate a predicted depth image corresponding to the real training image, it also includes: the computer device upsampling the predicted depth image to obtain the upsampled predicted depth image; wherein, during the upsampling process, the depth value of the pixel at the edge position of the predicted depth image is maintained; the computer device spatially transforms the predicted depth image to obtain three-dimensional point cloud data corresponding to the real training image, including: the computer device spatially transforms the upsampled predicted depth image to obtain three-dimensional point cloud data corresponding to the real training image.

[0128] After obtaining the predicted depth image, upsampling is performed on the predicted depth image to improve the resolution of the predicted depth image. The upsampling process includes but is not limited to at least one of the following: interpolation (such as bilinear interpolation), deconvolution, and depooling.

[0129] In some embodiments, during the upsampling process, processing of pixels at the edge of the predicted depth image is reduced to avoid large errors in edge data of the upsampled predicted depth image. In some embodiments, the resolution of the upsampled predicted depth image is higher than the resolution of the predicted depth image.

[0130] In some embodiments, the three-dimensional point cloud data is used to represent the coordinate information corresponding to each pixel point in the upsampled predicted depth image in the regularized space.

[0131] In some embodiments, a computer device performs spatial transformation on the upsampled predicted depth image to obtain three-dimensional point cloud data corresponding to the real training image, including: processing the upsampled predicted depth image through a projection matrix to obtain three-dimensional point cloud data; wherein the projection matrix is ​​used to determine the depth value corresponding to each pixel point in the sampled predicted depth image.

[0132] In an example, the target object in the real training image is a human body (in this case, the real training image can be called a real human body image). The following steps are used to introduce and illustrate the process of generating a three-dimensional reconstruction label corresponding to the real human body image.

[0133] Figure 5 3D reconstruction label generation process provided by an exemplary embodiment of the present application is a schematic diagram.

[0134] Step 1: Input the human body image into the relevant deep learning algorithm model to obtain the human body depth image. The process is as follows:

[0135] M(I)→"D"

[0136] Here, M() represents a general pixel-to-pixel algorithm model, i.e., a deep learning algorithm model that takes an image as input and generates another image. In this case, a human image is input and a corresponding human depth image is generated. I represents the input real human image, and D represents the predicted depth image estimated by M().

[0137] Step 2: Use the upsampling method to process the predicted depth image obtained in step 1 to obtain an upsampled predicted depth image; convert the upsampled predicted depth image into three-dimensional point cloud data through the projection matrix. The process is:

[0138] UP Sampling(I)→I up ,;Projection(I up )→PC;

[0139] Among them, UP Sampling represents the upsampling algorithm used in coherent technology, such as bilinear interpolation, which requires fixing the depth information of the edge of the predicted depth image during the upsampling process; I up "Projection" represents the high-resolution depth image obtained after upsampling. "Projection" represents the back-projection operation of the projection matrix, which is the transformation matrix between image space and regularized space. The depth information of a point (n) on a two-dimensional image can be represented as (X, Y) and the corresponding depth value, i.e., a three-dimensional vector. By back-projecting the projection matrix, a new three-dimensional vector is obtained, representing the coordinates of point n in regularized space. "PC" represents the coordinates of all pixels in the depth image in regularized space, which is called a three-dimensional point cloud.

[0140] Step 3: Sample the point cloud obtained in step 2 to get pseudo labels for training:

[0141] Sampling(PC)→L pseudo

[0142] Among them, Sampling represents a custom spatial sampling algorithm. For any point m(x, y, z) in the PC, several points m are sampled around it. i (+Δx,y+Δy,z+Δz), where Δx=α*Δz, Δy=β*Δz, where Δz is a custom parameter. Here, Δz∈(-2,2), α, β are random variables and satisfy 0<α<0.2, 0<β<0.2. Here, the upper limit of 0.2 can be adjusted according to actual conditions; L pseudo Indicates the 3D reconstruction label corresponding to the real human image.

[0143] Next, we introduce a method for generating 3D reconstruction labels corresponding to synthetic training data.

[0144] In some embodiments, a computer device obtains multiple training images and three-dimensional reconstruction labels corresponding to each training image, including: for a synthetic training image, the computer device obtains a virtual geometric configuration corresponding to the synthetic training image; wherein the virtual geometric configuration refers to the three-dimensional geometric configuration of the synthesized target object; the computer device renders and samples the virtual geometric configuration corresponding to the synthetic training image to obtain a three-dimensional reconstruction label corresponding to the synthetic training image.

[0145] In some embodiments, the virtual geometric configuration is modeled in software. The virtual geometric configuration includes texture information. By rendering and sampling the virtual geometric configuration, a synthetic training image and a corresponding 3D reconstruction label can be obtained.

[0146] The following describes a method for generating three-dimensional reconstruction information through several embodiments.

[0147] In some embodiments, the computer device obtains three-dimensional reconstruction information corresponding to the training image based on the training image through a three-dimensional reconstruction model, including: the computer device obtains feature voxels corresponding to the training image based on the training image through a three-dimensional reconstruction model, and the feature voxels include feature information of voxels corresponding to the target object in the training image; the computer device obtains three-dimensional reconstruction information corresponding to the training image based on the feature voxels through a three-dimensional reconstruction model.

[0148] In some embodiments, a computer device obtains three-dimensional reconstruction information corresponding to a training image based on feature voxels through a three-dimensional reconstruction model, including: the computer device samples points in the three-dimensional space where a target object in the training image is located to obtain multiple sampling points; the computer device determines feature information corresponding to multiple sampling points from the feature voxels through interpolation; the computer device obtains three-dimensional reconstruction information corresponding to the training image based on the feature information corresponding to the multiple sampling points through a three-dimensional reconstruction model.

[0149] In some embodiments, the sampling point is random, any point in the space where the target object is located.

[0150] In some embodiments, the x-axis and y-axis in the three-dimensional space are normalized according to the size of the feature voxel, that is, the space is normalized to the length and width range of the feature voxel, and points in this space are sampled.

[0151] In some embodiments, the computer device determines feature information of the sampling point in space based on the feature voxels. The feature information refers to the coordinates corresponding to the spatial point and the depth information corresponding to the spatial point.

[0152] In some embodiments, linear interpolation is performed on the feature information of the voxels in different directions to determine the feature information corresponding to the multiple sampling points. In some embodiments, bilinear sampling is performed on the feature voxels on the x-axis and y-axis to obtain the feature vectors of the sampling points.

[0153] The interpolation method described in the embodiment of the present application can be a spatial bilinear interpolation method. The present application does not limit the specific interpolation method. Any method of determining the characteristic information of the sampling point in space based on the characteristic voxels is included in the protection scope of the present application.

[0154] In some embodiments, the three-dimensional reconstruction model includes: a feature voxel extraction sub-model and a three-dimensional reconstruction sub-model; the feature voxel extraction sub-model is used to obtain feature voxels corresponding to the training image based on the training image; the three-dimensional reconstruction sub-model is used to obtain three-dimensional reconstruction information corresponding to the training image based on the feature voxels.

[0155] In some embodiments, the feature voxel extraction submodel is an encoder model within a convolutional neural network, which can be used to extract feature voxels corresponding to the depth image. In some embodiments, the 3D reconstruction submodel is a multilayer perceptron, optionally including fully connected layers, which can map feature vectors of spatial sampling points derived from feature voxels to signed distance fields in space, i.e., 3D reconstruction information.

[0156] The technical solutions provided in the embodiments of this application extract feature voxels using a feature voxel extraction submodel and obtain 3D reconstruction information using a 3D reconstruction submodel. The implicit field function can be considered to include both the feature voxel extraction submodel and the 3D reconstruction submodel. In some embodiments, the implicit field function can be considered to include both the feature voxel extraction submodel and a multilayer perceptron.

[0157] Figure 6 It is a schematic diagram of a three-dimensional reconstruction model training process provided by an exemplary embodiment of the present application.

[0158] Step 1: Mix real and synthetic human body images into the deep learning algorithm model to obtain the corresponding feature voxels. The process is as follows:

[0159] M1(I)→F V ;

[0160] Among them, it is a deep learning algorithm model that generates another image by inputting an image. M1() represents a feature voxel extraction model, that is, inputting an image and obtaining the corresponding feature voxel through the model. It needs additional explanation that M1 can contain multiple general pixel2pixel algorithm models, each pixel2pixel algorithm model can generate a normal map or a depth map, and then obtain the feature voxel from the normal map or the depth map or the input human body image. I represents the input human body image, F V Represents the feature voxels predicted by M1(). Generally, an image consists of three channels, RGB, and can be represented as a three-dimensional array, where the first and second dimensions represent the length and width of the image respectively, and the third dimension represents RGB.

[0161] Step 2: Use the spatial sampling interpolation method to obtain the feature vector of the entire space from the feature voxels obtained in step 1. Each spatial sampling point will have a corresponding feature vector, and use the multi-layer perceptron to predict the directed truncated distance field corresponding to each sampling point in the space to obtain the directed truncated distance field of the entire space. The process is:

[0162] Interpolation(F V , P i )→F i ,M2(F s )→F;

[0163] Among them, P i represents the points to be sampled in space, F i Represents the feature vector corresponding to each sampling point in the space, F S The feature vector of the entire space, M2() represents the multi-layer perceptron, F represents the directed truncated distance field of the entire space, and Interpolation represents the spatial bilinear interpolation method.

[0164] Step 3: By regularizing the difference between the predicted directed truncated distance field and the mixed label in step 2, the parameters P to be optimized of the feature voxel extraction model M1 in step 1 and the multilayer perceptron M2 in step 2 are optimized. M1 and P M2 :

[0165] P M1,M2 ((λ1 or λ2)*L(F,TSDF(L GT or L pseudo )));

[0166] Among them, λ1 and λ2 represent the weights of the real label and the pseudo label respectively, and their values ​​are proportional to the change of the model parameters. L represents the regularization constraint of the L1 norm. The L2 norm or similar norm can also be used for constraint. P M1,M2 Represents the parameters to be optimized in the feature voxel extraction model M1 and the multi-layer perceptron M2, TSDF(L GT or L pseudo ) represents the conversion of the mixed label into an approximate directed truncated distance field, L GT represents the true label corresponding to the synthetic human image, L pseudo Represents the pseudo label corresponding to the real human image. When the input is real data, use λ1 and L pseudo ; When the input is synthetic data, use λ2 and L GT .

[0167] Step 4: Convert the mixed label and directed truncated distance field into occupancy field forms respectively, use the difference between the two after regularization constraint conversion, and further optimize the parameter P to be optimized of the feature voxel extraction model M1 in step 1 and the multi-layer perceptron M2 in step 2 M1 and P M2 :

[0168] P M1,M2 ((λ3 or λ4)*L(OCC(F),OCC(TSDF(L GT or L pseudo ))));

[0169] λ3 and λ4 represent the weights of the true label and pseudo label, respectively. Their values ​​are proportional to the change in model parameters. L represents the regularization constraint of the L1 norm. The L2 norm can also be used. Since the occupancy field contains only 0 and 1, L here includes the BCE constraint, that is, the binary cross entropy loss function (Binary Cross Entropy). OCC() represents the conversion of the directed truncated distance field into the occupancy field.

[0170] Step 5: Use the optimized feature voxel extraction model and multi-layer perceptron to obtain the directed truncated distance field and use the Marching Cube algorithm to obtain the 3D geometry of the human body:

[0171] MC(F)→S

[0172] Here, S represents the reconstruction of the 3D human geometry based on the input human image, and MC represents the Marching Cube algorithm. The directed truncated distance field of the entire space contains the 3D human mesh, and the MC algorithm is required to extract the 3D human mesh separately.

[0173] Through the above method, by sampling points in space, the feature vectors corresponding to the sampling points are obtained based on the feature voxels. That is, the two-dimensional plane image is converted into three-dimensional point data information through the feature voxels. The three-dimensional point data information can be used to outline the three-dimensional geometric configuration of the target object.

[0174] Furthermore, the 3D reconstruction model is divided into a feature voxel extraction sub-model and a 3D reconstruction sub-model. This layered design effectively and accurately locates errors when they occur. Simultaneously, the model is trained simultaneously with the feature voxel extraction sub-model and the 3D reconstruction sub-model. This simultaneous training further improves the final training accuracy.

[0175] Figure 7 This is a diagram of the training results provided by an embodiment of the present application, comparing the present method with related methods. Figure 4 The first column shows a real-world human image. The second and third columns show the 3D human mesh and local area magnification reconstructed by related methods, respectively. The fourth and fifth columns show the 3D human mesh and local area magnification reconstructed by our method, respectively. The results show that the 3D human geometry reconstructed by our method has more realistic details and is closer to the real world.

[0176] Figure 8 This is a schematic diagram of the training results provided by an embodiment of the present application. Figure 5The first column shows real-world human images, the second column shows 3D human meshes reconstructed using related methods, and the third column shows the 3D human meshes reconstructed using our method. The results show that the 3D human geometry reconstructed by our method is more complete, without any "defects," and closer to the real-world image input to the network.

[0177] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0178] Please refer to Figure 9 , which shows a block diagram of a training device for a three-dimensional reconstruction model provided by an embodiment of the present application. The device has the function of implementing the above-mentioned method example, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device described above, or it can be set in a computer device. Figure 9 As shown, the device 900 may include: an image acquisition module 910, an information generation module 920 and a model training module 930.

[0179] The image acquisition module 910 is used to acquire multiple training images and three-dimensional reconstruction labels corresponding to each of the training images; wherein the multiple training images include at least one real training image and at least one synthetic training image, the real training image refers to an image obtained by photographing a real target object, and the synthetic training image refers to an image generated based on a synthetic three-dimensional model of the target object.

[0180] The information generation module 920 is used to obtain three-dimensional reconstruction information corresponding to the training image based on the training image through the three-dimensional reconstruction model, and the three-dimensional reconstruction information is used to determine the three-dimensional geometric configuration of the target object in the training image in the three-dimensional space.

[0181] The model training module 930 is configured to train the 3D reconstruction model according to the 3D reconstruction information and 3D reconstruction labels corresponding to the training images.

[0182] In some embodiments, the model training module 930 includes: a loss calculation unit, used to calculate a first training loss based on the three-dimensional reconstruction information and the three-dimensional reconstruction label corresponding to the real training image, and the first training loss is used to indicate the difference between the three-dimensional reconstruction information and the three-dimensional reconstruction label corresponding to the real training image; a second training loss is calculated based on the three-dimensional reconstruction information and the three-dimensional reconstruction label corresponding to the synthetic training image, and the second training loss is used to indicate the difference between the three-dimensional reconstruction information and the three-dimensional reconstruction label corresponding to the synthetic training image; a total loss determination unit, used to adjust the parameters of the three-dimensional reconstruction model based on the first training loss and the second training loss.

[0183] In some embodiments, the total loss determination unit includes: performing weighted summation of the first training loss and the second training loss to obtain a total training loss; and adjusting the parameters of the three-dimensional reconstruction model according to the total training loss.

[0184] In some embodiments, the 3D reconstruction information and the 3D reconstruction label are represented in the form of a signed distance field, where the signed distance field is used to characterize the distance between at least one spatial point and a 3D geometric configuration surface corresponding to a target object. The apparatus 900 further includes: an information classification model for obtaining the 3D reconstruction information and the 3D reconstruction label represented in the form of an occupancy field based on the 3D reconstruction information and the 3D reconstruction label represented in the form of the signed distance field; wherein the occupancy field is used to characterize the internal and external relationship between the at least one spatial point and the target object and the 3D geometric configuration surface.

[0185] The model training module 930 is used to train the 3D reconstruction model based on the 3D reconstruction information and 3D reconstruction labels corresponding to the training images and represented in the form of the signed distance field, and the 3D reconstruction information and 3D reconstruction labels corresponding to the training images and represented in the form of the occupancy field.

[0186] In some embodiments, the image acquisition module 910 includes: a depth prediction unit, which is used to generate a predicted depth image corresponding to the real training image using a depth map prediction model; a spatial conversion module, which is used to perform spatial conversion on the predicted depth image to obtain three-dimensional point cloud data corresponding to the real training image; wherein the spatial conversion is used to realize the conversion from two-dimensional space to three-dimensional space; and a point cloud sampling module, which is used to sample the three-dimensional point cloud data to obtain a three-dimensional reconstruction label corresponding to the real training image.

[0187] In some embodiments, the device 900 further includes a pixel sampling module for upsampling the predicted depth image to obtain the upsampled predicted depth image; wherein, during the upsampling process, the depth values ​​of the pixels at the edge positions of the predicted depth image are maintained; and the point cloud sampling module is used to perform spatial conversion on the upsampled predicted depth image to obtain three-dimensional point cloud data corresponding to the real training image.

[0188] In some embodiments, the image acquisition module 910 includes: a configuration acquisition unit, used to acquire a virtual geometric configuration corresponding to the synthetic training image; wherein the virtual geometric configuration refers to a three-dimensional geometric configuration corresponding to the synthesized target object; and a label generation unit, used to render and sample the virtual geometric configuration corresponding to the synthetic training image to obtain a three-dimensional reconstruction label corresponding to the synthetic training image.

[0189] In some embodiments, the information generation module 920 includes: a voxel generation unit, which is used to obtain feature voxels corresponding to the training image based on the training image through the three-dimensional reconstruction model, and the feature voxels include feature information of voxels corresponding to the target object in the training image; an information generation unit, which is used to obtain three-dimensional reconstruction information corresponding to the training image based on the feature voxels through the three-dimensional reconstruction model.

[0190] In some embodiments, the information generation unit is used to sample points in the three-dimensional space where the target object in the training image is located to obtain multiple sampling points; determine the feature information corresponding to the multiple sampling points from the feature voxels by interpolation; and obtain the three-dimensional reconstruction information corresponding to the training image based on the feature information corresponding to the multiple sampling points through the three-dimensional reconstruction model.

[0191] In some embodiments, the three-dimensional reconstruction model includes: a feature voxel extraction sub-model and a three-dimensional reconstruction sub-model; the feature voxel extraction sub-model is used to obtain the feature voxels corresponding to the training image based on the training image; the three-dimensional reconstruction sub-model is used to obtain the three-dimensional reconstruction information corresponding to the training image based on the feature voxels.

[0192] In some embodiments, the trained three-dimensional reconstruction model is used for at least one of the following: generating a three-dimensional geometric configuration of a human body in a three-dimensional virtual scene based on a real color image of the human body; generating a three-dimensional geometric configuration of a digital human in a three-dimensional virtual scene based on a real color image of the digital human; generating a three-dimensional geometric configuration of clothes in a three-dimensional virtual scene based on a real color image of clothes; generating a three-dimensional geometric configuration of a human face in a three-dimensional virtual scene based on a real color image of the human face.

[0193] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0194] Please refer to Figure 10 , which shows a structural block diagram of a computer device 1000 provided in one embodiment of the present application.

[0195] Typically, the computer device 1000 includes a processor 1001 and a memory 1002 .

[0196] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI processor for processing computing operations related to machine learning.

[0197] Memory 1002 may include one or more computer-readable storage media, which may be non-transitory. Memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 1002 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned three-dimensional reconstruction model training method.

[0198] Those skilled in the art will understand that Figure 10 The structure shown in the figure does not constitute a limitation on the computer device 1000, and the computer device 1000 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.

[0199] In an exemplary embodiment, a computer-readable storage medium is further provided, wherein a computer program is stored in the storage medium. When the computer program is executed by a processor, the computer program is used to implement the training method of the three-dimensional reconstruction model.

[0200] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or an optical disk, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0201] In an exemplary embodiment, a computer program product is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a terminal device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the terminal device to perform the above-described method for training a 3D reconstruction model.

[0202] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.

[0203] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A training method for a three-dimensional reconstruction model, characterized in that: The method comprises: Acquire a plurality of training images; wherein the plurality of training images include at least one real training image and at least one synthetic training image, the real training image refers to an image obtained by photographing a real target object, and the synthetic training image refers to an image generated based on a synthetic three-dimensional model of the target object; For the real training image, a depth map prediction model is used to generate a predicted depth image corresponding to the real training image; the predicted depth image is spatially transformed to obtain three-dimensional point cloud data corresponding to the real training image; wherein the spatial transformation is used to achieve the conversion from two-dimensional space to three-dimensional space; the three-dimensional point cloud data is sampled to obtain a three-dimensional reconstruction label corresponding to the real training image; For the synthetic training image, a virtual geometric configuration corresponding to the synthetic training image is obtained; wherein the virtual geometric configuration refers to the three-dimensional geometric configuration of the synthesized target object; the virtual geometric configuration corresponding to the synthetic training image is rendered and sampled to obtain a three-dimensional reconstruction label corresponding to the synthetic training image; Obtaining three-dimensional reconstruction information corresponding to the training image through the three-dimensional reconstruction model according to the training image, wherein the three-dimensional reconstruction information is used to determine the three-dimensional geometric configuration of the target object in the training image in the three-dimensional space; The 3D reconstruction model is trained according to the 3D reconstruction information and the 3D reconstruction label corresponding to the training image.

2. The method according to claim 1, characterized in that: The step of training the 3D reconstruction model according to the 3D reconstruction information and the 3D reconstruction label corresponding to the training image includes: Calculating a first training loss according to the 3D reconstruction information and the 3D reconstruction label corresponding to the real training image, where the first training loss is used to indicate a difference between the 3D reconstruction information and the 3D reconstruction label corresponding to the real training image; Calculating a second training loss according to the 3D reconstruction information and the 3D reconstruction label corresponding to the synthetic training image, where the second training loss is used to indicate a difference between the 3D reconstruction information and the 3D reconstruction label corresponding to the synthetic training image; The parameters of the three-dimensional reconstruction model are adjusted according to the first training loss and the second training loss.

3. The method according to claim 2, characterized in that The adjusting the parameters of the three-dimensional reconstruction model according to the first training loss and the second training loss includes: Performing a weighted summation on the first training loss and the second training loss to obtain a total training loss; The parameters of the three-dimensional reconstruction model are adjusted according to the total training loss.

4. The method according to claim 1, characterized in that: The three-dimensional reconstruction information and the three-dimensional reconstruction label are represented in the form of a signed distance field, and the signed distance field is used to characterize the distance between at least one spatial point and a three-dimensional geometric configuration surface corresponding to the target object; The method further comprises: According to the three-dimensional reconstruction information and the three-dimensional reconstruction label represented by the signed distance field, the three-dimensional reconstruction information and the three-dimensional reconstruction label represented by the occupancy field are obtained; wherein the occupancy field is used to characterize the internal and external relationship between at least one spatial point and the three-dimensional geometric configuration surface corresponding to the target object; The step of training the 3D reconstruction model according to the 3D reconstruction information and the 3D reconstruction label corresponding to the training image includes: The 3D reconstruction model is trained according to the 3D reconstruction information and 3D reconstruction labels corresponding to the training images and the 3D reconstruction information and 3D reconstruction labels corresponding to the training images and the 3D reconstruction information and 3D reconstruction labels corresponding to the training images and the occupancy field forms.

5. The method according to claim 1, characterized in that: After the depth map prediction model is used to generate the predicted depth image corresponding to the real training image, the method further includes: Upsampling the predicted depth image to obtain the upsampled predicted depth image; wherein, during the upsampling process, the depth values ​​of the pixels at the edge positions of the predicted depth image are maintained; The performing spatial transformation on the predicted depth image to obtain three-dimensional point cloud data corresponding to the real training image includes: The upsampled predicted depth image is spatially transformed to obtain three-dimensional point cloud data corresponding to the real training image.

6. The method according to claim 1, characterized in that The obtaining, according to the training image, the three-dimensional reconstruction information corresponding to the training image by using the three-dimensional reconstruction model includes: Obtaining feature voxels corresponding to the training image according to the training image through the three-dimensional reconstruction model, wherein the feature voxels include feature information of voxels corresponding to the target object in the training image; The three-dimensional reconstruction information corresponding to the training image is obtained through the three-dimensional reconstruction model according to the characteristic voxels.

7. The method according to claim 6, characterized in that The obtaining, by using the three-dimensional reconstruction model according to the characteristic voxels, the three-dimensional reconstruction information corresponding to the training image includes: Sampling points in the three-dimensional space where the target object in the training image is located to obtain a plurality of sampling points; Determining feature information corresponding to each of the plurality of sampling points from the feature voxels by interpolation; The three-dimensional reconstruction information corresponding to the training image is obtained through the three-dimensional reconstruction model according to the feature information respectively corresponding to the multiple sampling points.

8. The method according to claim 6, characterized in that The three-dimensional reconstruction model includes: a feature voxel extraction sub-model and a three-dimensional reconstruction sub-model; The feature voxel extraction sub-model is used to obtain the feature voxels corresponding to the training image according to the training image; The 3D reconstruction sub-model is used to obtain 3D reconstruction information corresponding to the training image according to the characteristic voxels.

9. The method according to any one of claims 1 to 8, characterized in that: The trained three-dimensional reconstruction model is used for at least one of the following: Generating a three-dimensional geometric configuration of the human body in a three-dimensional virtual scene according to a real color image of the human body; Generating a three-dimensional geometric configuration of the digital human in a three-dimensional virtual scene according to the real color image of the digital human; Generating a three-dimensional geometric configuration of the clothes in a three-dimensional virtual scene according to a real color image of the clothes; A three-dimensional geometric configuration of the human face is generated in a three-dimensional virtual scene according to a real color image of the human face.

10. A training device for a three-dimensional reconstruction model, characterized in that: The device comprises: An image acquisition module, used to acquire multiple training images; wherein the multiple training images include at least one real training image and at least one synthetic training image, the real training image refers to an image obtained by photographing a real target object, and the synthetic training image refers to an image generated based on a three-dimensional model of a synthetic target object; for the real training image, a depth map prediction model is used to generate a predicted depth image corresponding to the real training image; the predicted depth image is spatially transformed to obtain three-dimensional point cloud data corresponding to the real training image; wherein the spatial transformation is used to achieve conversion from two-dimensional space to three-dimensional space; the three-dimensional point cloud data is sampled to obtain a three-dimensional reconstruction label corresponding to the real training image; for the synthetic training image, a virtual geometric configuration corresponding to the synthetic training image is obtained; wherein the virtual geometric configuration refers to the three-dimensional geometric configuration of the synthetic target object; the virtual geometric configuration corresponding to the synthetic training image is rendered and sampled to obtain a three-dimensional reconstruction label corresponding to the synthetic training image; An information generation module, configured to obtain three-dimensional reconstruction information corresponding to the training image according to the training image through the three-dimensional reconstruction model, wherein the three-dimensional reconstruction information is used to determine a three-dimensional geometric configuration of a target object in the training image in a three-dimensional space; The model training module is used to train the 3D reconstruction model according to the 3D reconstruction information and 3D reconstruction labels corresponding to the training images.

11. A computer device, characterized in that: The computer device comprises a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The computer program product comprises computer instructions, which are stored in a computer-readable storage medium. A processor reads and executes the computer instructions from the computer-readable storage medium to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • High-precision three-dimensional face reconstruction method

    CN111402403A