Face Reconstruction Method, Device, Electronic Device and Storage Medium
By integrating the three-dimensional shape information of the face in the map generation model, the problem of inconsistency between shape reconstruction and texture map in the existing technology is solved, and the effect of three-dimensional face reconstruction is improved.
Patent Information
- Application Number
- CN202210473877.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The three-dimensional face reconstruction method based on a single RGB image in the prior art separates shape reconstruction and texture generation, resulting in the lack of shape information when generating the three-dimensional face texture map, resulting in poor effect on the side face area and inconsistent with the texture map, affecting the overall face reconstruction effect.
By obtaining the face image and face three-dimensional shape information to be reconstructed, input it into the map generation model, integrating the three-dimensional shape information to generate three-dimensional face texture maps, optimizing the texture map generation process, and improving the three-dimensional face reconstruction effect.
The three-dimensional shape information of the face is integrated during the map generation process, which avoids the inconsistency between the shape reconstruction results and the texture map, optimizes the generation effect of the three-dimensional face texture map, and improves the overall effect of the three-dimensional face reconstruction.
Smart Images

Figure CN114820941B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of 3D face reconstruction, and particularly relates to a face reconstruction method, device, electronic device, and storage medium. Background Art
[0002] Face reconstruction has always been an important research direction in the field of computer graphics. Its purpose is to reconstruct a face model that is most similar to the face in the image. Face reconstruction technology plays a key role in fields such as games, movies, and social entertainment. In related face reconstruction methods, when performing face reconstruction, shape reconstruction and texture generation are separated, resulting in poor overall face reconstruction effects. Summary of the Invention
[0003] In view of the above problems, this application proposes a face reconstruction method, device, electronic device, and storage medium to improve the above problems.
[0004] In a first aspect, an embodiment of this application provides a face reconstruction method, which includes: obtaining a face image to be reconstructed and corresponding 3D face shape information of the face image to be reconstructed; inputting the face image to be reconstructed and the 3D face shape information into a texture mapping generation model to obtain a 3D face texture map output by the texture mapping generation model; and generating a 3D reconstruction image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
[0005] In a second aspect, an embodiment of this application provides a face reconstruction device, which includes: a data acquisition unit for obtaining a face image to be reconstructed and corresponding 3D face shape information of the face image to be reconstructed; a texture mapping generation unit for inputting the face image to be reconstructed and the 3D face shape information into a texture mapping generation model to obtain a 3D face texture map output by the texture mapping generation model; and an image generation unit for generating a 3D reconstruction image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
[0006] In a third aspect, an embodiment of this application provides an electronic device, including one or more processors and a memory; one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above method.
[0007] In a fourth aspect, an embodiment of this application provides a computer-readable storage medium, in which program code is stored. When the program code runs, the above method is executed.
[0008] The embodiments of the present application provide a face reconstruction method, apparatus, electronic device, and storage medium. First, an image of a face to be reconstructed and the three-dimensional shape information of the face corresponding to the image of the face to be reconstructed are obtained. Then, the image of the face to be reconstructed and the three-dimensional shape information of the face are input into a texture map generation model to obtain the three-dimensional face texture map output by the texture map generation model. Finally, based on the three-dimensional shape information of the face and the three-dimensional face texture map, a three-dimensional reconstruction image corresponding to the image of the face to be reconstructed is generated. Through the above method, during the generation of the texture map, the three-dimensional shape information of the face is fused, avoiding the inconsistency between the three-dimensional shape reconstruction result and the texture map, optimizing the effect of generating the three-dimensional face texture map, and further improving the effect of three-dimensional face reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0010] Figure 1 FIG. shows a schematic diagram of an application scenario of a face reconstruction method proposed in an embodiment of the present application;
[0011] Figure 2 FIG. shows a schematic diagram of an application scenario of a face reconstruction method proposed in an embodiment of the present application;
[0012] Figure 3 FIG. shows a flowchart of a face reconstruction method proposed in an embodiment of the present application;
[0013] Figure 4 FIG. shows a schematic diagram of the implementation of steps S110 - S130 in an embodiment of the present application;
[0014] Figure 5 FIG. shows a schematic diagram of the three-dimensional reconstruction image generated in an embodiment of the present application;
[0015] Figure 6 FIG. shows a flowchart of a face reconstruction method proposed in another embodiment of the present application;
[0016] Figure 7 FIG. shows a schematic diagram of the implementation of steps S210 - S250 in another embodiment of the present application;
[0017] Figure 8 FIG. shows a schematic diagram of the implementation of steps S210 - S250 in another embodiment of the present application;
[0018] Figure 9Shows the flowchart of a face reconstruction method proposed in another embodiment of the present application;
[0019] Figure 10 Shows the schematic diagram of the implementation of steps S310 - S350 in another embodiment of the present application;
[0020] Figure 11 Shows the flowchart of a face reconstruction method proposed in another embodiment of the present application;
[0021] Figure 12 Shows the flowchart of step S410 in another embodiment of the present application;
[0022] Figure 13 Shows the schematic diagram of the implementation of steps S411 - S414 in another embodiment of the present application;
[0023] Figure 14 Shows the schematic diagram of the implementation of steps S410 - S460 in another embodiment of the present application;
[0024] Figure 15 Shows the flowchart of a face reconstruction method proposed in another embodiment of the present application;
[0025] Figure 16 Shows the schematic diagram of the implementation of steps S510 - S560 in another embodiment of the present application;
[0026] Figure 17 Shows the flowchart of a face reconstruction method proposed in another embodiment of the present application;
[0027] Figure 18 Shows the schematic diagram of the implementation of steps S610 - S670 in another embodiment of the present application;
[0028] Figure 19 Shows the structural block diagram of a face reconstruction device proposed in an embodiment of the present application;
[0029] Figure 20 Shows the structural block diagram of an electronic device or server for executing the face reconstruction method according to an embodiment of the present application;
[0030] Figure 21 Shows the storage unit for storing or carrying the program code for implementing the face reconstruction method according to an embodiment of the present application. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] In the fields of computer vision and computer graphics, 3D face reconstruction is a topic that is gradually gaining popularity because it can be widely applied in fields such as face recognition, face editing, human-computer interaction, expression-driven animation, augmented reality, virtual reality, etc. And 3D face reconstruction based on a single RGB image - that is, reconstructing the 3D face shape and texture using a single RGB image - is an enduring topic.
[0033] However, the inventors found in the research on related face reconstruction methods that in the current texture generation method for 3D face reconstruction based on a single RGB image, by separating shape reconstruction and 3D face texture mapping generation, there is a lack of shape information during the 3D face texture mapping generation process, resulting in poor effects in the side face area of the generated 3D face texture mapping. In addition, separating the two processes of shape reconstruction and 3D face texture mapping generation easily leads to inconsistencies between the shape reconstruction result and the 3D face texture mapping generation result, resulting in poor overall face reconstruction results.
[0034] Therefore, the inventors proposed the face reconstruction method, device, electronic device, and storage medium in the present application. First, obtain the face image to be reconstructed and the corresponding 3D face shape information of the face image to be reconstructed, then input the face image to be reconstructed and the 3D face shape information into the texture mapping generation model to obtain the 3D face texture mapping output by the texture mapping generation model, and finally, based on the 3D face shape information and the 3D face texture mapping, generate the 3D reconstruction image corresponding to the face image to be reconstructed. Through the above method, during the texture mapping generation process, the 3D face shape information is fused, avoiding inconsistencies between the 3D shape reconstruction result and the texture mapping, optimizing the effect of 3D face texture mapping generation, and thus improving the effect of 3D face reconstruction.
[0035] In the embodiments of the present application, the provided face reconstruction method can be executed by an electronic device. In this manner executed by an electronic device, all steps in the face reconstruction method provided in the embodiments of the present application can be executed by the electronic device. For example, as Figure 1As shown, the data acquisition device of the electronic device 100 can acquire the face image to be reconstructed and the three-dimensional face shape information corresponding to the face image to be reconstructed, and then transmit the acquired face image to be reconstructed and the three-dimensional face shape information corresponding to the face image to be reconstructed to the processor, so that the processor can input the face image to be reconstructed and the three-dimensional face shape information into the texture mapping generation model in real time to obtain the three-dimensional face texture mapping output by the texture mapping generation model; based on the three-dimensional face shape information and the three-dimensional face texture mapping, a three-dimensional reconstructed image corresponding to the face image to be reconstructed is generated.
[0036] Furthermore, the face reconstruction method provided in the embodiments of the present application can also be executed by a server (cloud). Correspondingly, in this manner executed by the server, the electronic device can acquire the face image to be reconstructed and the three-dimensional face shape information corresponding to the face image to be reconstructed, and synchronously send the face image to be reconstructed and the three-dimensional face shape information corresponding to the face image to be reconstructed to the server, and then the server inputs the face image to be reconstructed and the three-dimensional face shape information into the texture mapping generation model in real time to obtain the three-dimensional face texture mapping output by the texture mapping generation model; based on the three-dimensional face shape information and the three-dimensional face texture mapping, a three-dimensional reconstructed image corresponding to the face image to be reconstructed is generated.
[0037] In addition, it can also be executed collaboratively by the electronic device and the server. In this manner executed collaboratively by the electronic device and the server, some steps in the face reconstruction method provided in the embodiments of the present application are executed by the electronic device, while other steps are executed by the server.
[0038] Exemplarily, as Figure 2 shown, the electronic device 100 can execute the following steps included in the face reconstruction method: acquire the face image to be reconstructed and the three-dimensional face shape information corresponding to the face image to be reconstructed, and then the server 200 executes inputting the face image to be reconstructed and the three-dimensional face shape information into the texture mapping generation model to obtain the three-dimensional face texture mapping output by the texture mapping generation model; based on the three-dimensional face shape information and the three-dimensional face texture mapping, a three-dimensional reconstructed image corresponding to the face image to be reconstructed is generated.
[0039] It should be noted that, in this manner executed collaboratively by the electronic device and the server, the steps respectively executed by the electronic device and the server are not limited to the manner introduced in the above example, and in practical applications, the steps respectively executed by the electronic device and the server can be dynamically adjusted according to the actual situation.
[0040] It should be noted that in addition to being Figure 1 and Figure 2In addition to the smart phone shown, it can also be a vehicle-mounted device, a wearable device, a tablet computer, a notebook computer, a smart speaker, etc. The server 120 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers.
[0041] The embodiments of the present application will be specifically described below with reference to the accompanying drawings.
[0042] Please refer to Figure 3 , a face reconstruction method provided by an embodiment of the present application is applied to an electronic device or a server as shown in Figure 1 or Figure 2 . The method includes:
[0043] Step S110: Obtain a face image to be reconstructed and three-dimensional shape information of the face corresponding to the face image to be reconstructed.
[0044] In the embodiment of the present application, the face image to be reconstructed can be a face image including a face for which three-dimensional face reconstruction is required. The three-dimensional shape information of the face can include face contour information, face expression information, and three-dimensional coordinate information of each vertex in the face, etc.
[0045] As a way, a face image in RGB, YUV, RAW and other formats can be collected by an image acquisition device. After collecting a face image in RGB, YUV, RAW and other formats by the image acquisition device, the face image in RGB, YUV, RAW and other formats can be preprocessed to obtain a preprocessed face image, and the preprocessed face image is used as the face image to be reconstructed. After obtaining the face image to be reconstructed, obtain the three-dimensional shape information of the face included in the face image to be reconstructed. Among them, the image acquisition device can be a smart device with a camera, such as a smart phone, a camera, etc.
[0046] When obtaining the three-dimensional shape information of the face, the three-dimensional shape information of the face can be obtained by a hardware acquisition device. For example, the three-dimensional shape information of the face can be obtained by a 3D point cloud device or a multi-camera system. Optionally, the three-dimensional shape information of the face included in the face image to be reconstructed can also be obtained by a three-dimensional face shape acquisition model.
[0047] As another way, the face image to be reconstructed can also be an image obtained from a cloud server or a preset storage area (such as the photo album of a smart phone). In response to an operation of triggering the acquisition of the face image to be reconstructed, the face image to be reconstructed is obtained from the cloud server or the preset storage area. Among them, the operation of triggering the acquisition of the face image to be reconstructed can be an operation of detecting the start of a specified application program, or a preset touch operation on the specified application program, etc., which is not specifically limited here.
[0048] Step S120: Input the face image to be reconstructed and the 3D face shape information into a texture map generation model, and obtain the 3D face texture map output by the texture map generation model.
[0049] In an embodiment of the present application, the texture map generation model may be a pre-trained model based on a convolutional neural network algorithm, such as Unet, StarGAN, StyleGAN and other models. Models based on this type of algorithm usually include an Encoder texture encoding module and a Decoder texture map generation module. The Encoder texture encoding module is used to obtain the texture features of the face image to be reconstructed and the 3D face shape information; the Decoder texture map generation module is used to generate a 3D face texture map based on the texture features extracted by the Encoder texture encoding module.
[0050] As a way, the steps to obtain the texture map generation model may include: obtaining an image training data set; training the model to be trained based on the image training data set and a preset loss function until the training end condition is met, and obtaining the texture map generation model. Among them, the image training data set may include multiple face images and the texture maps or stylized texture maps corresponding to each face image under UV coordinates; the preset loss function may be a total loss function obtained by weighted summation based on the L1 loss function, the L2 loss function, and the generative adversarial ADV loss function; the training end condition may be that the loss value meets the preset loss value (the loss value no longer changes) or the number of training iterations reaches the preset number of iterations, where the preset loss value and the preset number of iterations may be preset.
[0051] Step S130: Generate a 3D reconstruction image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
[0052] In an embodiment of the present application, after obtaining the 3D face shape information and the 3D face texture map, a 3D reconstruction image corresponding to the face image to be reconstructed may be generated based on the 3D face shape information and the 3D face texture map. Among them, the 3D reconstruction image corresponding to the face image to be reconstructed includes the texture information generated based on the 3D face texture map, that is, the generated 3D reconstruction image is a textured 3D face image.
[0053] As a way, after obtaining the 3D face shape information and the 3D face texture map, the 3D face shape information and the 3D face texture map may be combined to generate a 3D reconstruction image corresponding to the face image to be reconstructed. Specifically, the generated 3D face texture map is corresponding to the 3D face shape information in a preset format, and a textured 3D face reconstruction image can be obtained. Among them, the preset format may be the UV coordinate format.
[0054] The schematic diagram of the implementation of the above steps S110 - S130 can be as follows Figure 4 shown. First, a face image and three - dimensional face shape information are obtained; then, the face image is pre - processed to obtain the face image to be reconstructed; the face image to be reconstructed and the three - dimensional face shape information are input into the fusion module of the texture mapping generation model (the Encoder texture encoding module as described above), and texture features are obtained through the fusion module. Then, the texture features are input into the texture mapping generation module of the texture mapping generation model (the Decoder texture mapping generation module as described above) to generate a three - dimensional face texture map; finally, based on the three - dimensional face texture map and the three - dimensional face shape information, a three - dimensional reconstructed image is generated, as Figure 5 shown Figure 5 In, the image on the left side of the '+' sign in is the three - dimensional face shape information Figure 5 In, the image on the right side of the '+' sign in is the three - dimensional face texture map Figure 5 In, the image on the right side of the arrow in is the three - dimensional reconstructed image
[0055] A face reconstruction method provided by this application first obtains a face image to be reconstructed and the corresponding three - dimensional face shape information of the face image to be reconstructed, then inputs the face image to be reconstructed and the three - dimensional face shape information into a texture mapping generation model to obtain the three - dimensional face texture map output by the texture mapping generation model, and finally generates a three - dimensional reconstructed image corresponding to the face image to be reconstructed based on the three - dimensional face shape information and the three - dimensional face texture map. Through the above method, during the texture mapping generation process, the three - dimensional face shape information is fused, avoiding the inconsistency between the three - dimensional shape reconstruction result and the texture map, optimizing the effect of three - dimensional face texture map generation, and thus improving the effect of three - dimensional face reconstruction
[0056] Please refer to Figure 6 , a face reconstruction method provided by an embodiment of this application is applied to an electronic device or a server as shown in Figure 1 or Figure 2 shown, and the method includes:
[0057] Step S210: Obtain a face image to be reconstructed and the corresponding three - dimensional face shape information of the face image to be reconstructed
[0058] Step S220: Generate a map in the form of UV coordinates based on the three - dimensional face shape information
[0059] As a way, the step of generating a map in the form of UV coordinates based on the three - dimensional face shape information includes: generating a corresponding normal map based on the three - dimensional face shape information; unfolding the normal map into a map in the form of UV coordinates; or unfolding the three - dimensional face shape information into a map in the form of UV coordinates
[0060] In the embodiments of the present application, when generating a graph in UV coordinate format, it is possible to first generate a corresponding normal map through the three-dimensional shape information of the human face, and then expand the normal map into a graph in UV coordinate format. That is to say, first calculate the normal vector of each vertex based on the three-dimensional coordinate information of each vertex in the three-dimensional shape information of the human face to determine the corresponding normal map, and then convert the normal map into two-dimensional coordinate information in UV coordinate format; it is also possible to directly expand the three-dimensional shape information of the human face into a graph in UV coordinate format, that is, convert the three-dimensional coordinate information of each vertex in the three-dimensional shape information of the human face into two-dimensional coordinate information in UV coordinate format. Among them, the normal map is a three-channel normal image.
[0061] Among them, the method of first generating a corresponding normal map through the three-dimensional shape information of the human face and then expanding the normal map into a graph in UV coordinate format results in a more accurate graph in UV coordinate format; the method of directly expanding the three-dimensional shape information of the human face into a graph in UV coordinate format can obtain a graph in UV coordinate format more quickly. Therefore, when choosing which method to obtain a graph in UV coordinate format, it can be selected according to actual needs. If a graph in UV coordinate format is needed more quickly, then directly expand the three-dimensional shape information of the human face into a graph in UV coordinate format; if a graph in UV coordinate format with higher accuracy is needed, then first generate a corresponding normal map through the three-dimensional shape information of the human face, and then expand the normal map into a graph in UV coordinate format.
[0062] Step S230: Perform a splicing or addition operation on the graph in the UV coordinate form and the face image to be reconstructed to obtain a fused face image to be reconstructed.
[0063] In the embodiments of the present application, the face image to be reconstructed is a three-channel image (RGB image / YUV image / RAW image), and the graph in UV coordinate format is also a three-channel image (x / y / z three channels). Performing a splicing or addition operation on the graph in UV coordinate form and the face image to be reconstructed means adding or splicing the channels of the graph in UV coordinate form and the face image to be reconstructed. That is to say, the fused face image to be reconstructed obtained is a six-channel image.
[0064] Step S240: Input the fused face image to be reconstructed into the texture encoding module to obtain the three-dimensional face texture features output by the texture encoding module.
[0065] Among them, the texture mapping generation model includes a texture encoding module and a texture mapping generation module.
[0066] Step S250: Input the three-dimensional face texture features into the texture mapping generation module to obtain the three-dimensional face texture map output by the texture mapping generation module.
[0067] As a way, the schematic diagram of the implementation of steps S210 - S250 can be as follows Figure 7 shown. First, obtain the face image to be reconstructed and the 3D face shape information, then calculate the normal map based on the 3D face shape information to obtain the corresponding normal map, then unfold the normal map into an image in UV coordinate format, and then perform an image channel splicing or addition operation on the face image to be reconstructed and the image in UV coordinate format to obtain the fused face image to be reconstructed. Furthermore, the fused face image to be reconstructed can be input into the texture encoding module for feature extraction.
[0068] As another way, the schematic diagram of the implementation of steps S210 - S250 can also be as follows Figure 8 shown. First, obtain the face image to be reconstructed and the 3D face shape information, then directly unfold the 3D face shape information into an image in UV coordinate format, and then perform an image channel splicing or addition operation on the face image to be reconstructed and the image in UV coordinate format to obtain the fused face image to be reconstructed. Furthermore, the fused face image to be reconstructed can be input into the texture encoding module for feature extraction.
[0069] Optionally, after obtaining the image in UV coordinate format, if the size of the image in UV coordinate format is different from the size of the face image to be reconstructed, it is necessary to first scale the size of the image in UV coordinate format to be the same as the size of the face image to be reconstructed, and then perform an image channel addition or splicing operation on the scaled image in UV coordinate format and the face image to be reconstructed.
[0070] Step S260: Generate a 3D reconstruction image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
[0071] A face reconstruction method provided by the present application first obtains a face image to be reconstructed and the 3D face shape information corresponding to the face image to be reconstructed, generates an image in UV coordinate format based on the 3D face shape information, then performs a splicing or addition operation on the image in UV coordinate form and the face image to be reconstructed to obtain the fused face image to be reconstructed, and then inputs the fused face image to be reconstructed into the texture encoding module to obtain the 3D face texture features output by the texture encoding module. Then, the 3D face texture features are input into the texture map generation module to obtain the 3D face texture map output by the texture map generation module. Finally, based on the 3D face shape information and the 3D face texture map, a 3D reconstruction image corresponding to the face image to be reconstructed is generated. Through the above method, during the texture map generation process, the 3D face shape information is fused, avoiding the inconsistency between the 3D shape reconstruction result and the texture map, optimizing the effect of 3D face texture map generation, and thus improving the effect of 3D face reconstruction.
[0072] Please refer to Figure 9, A face reconstruction method provided by an embodiment of the present application is applied to an electronic device or a server as shown in Figure 1 or Figure 2 . The method includes:
[0073] Step S310: Obtain a face image to be reconstructed and corresponding 3D face shape information.
[0074] Step S320: Generate a map in the form of UV coordinates based on the 3D face shape information.
[0075] As a way, the step of generating a map in the form of UV coordinates based on the 3D face shape information includes: generating a corresponding normal map based on the 3D face shape information; unfolding the normal map into a map in the form of UV coordinates; or unfolding the 3D face shape information into a map in the form of UV coordinates.
[0076] Step S330: Input the face image to be reconstructed into the first texture encoding module and obtain the first texture feature output by the first texture encoding module.
[0077] Wherein, the texture mapping generation model includes a first texture encoding module, a second texture encoding module, and a texture mapping generation module.
[0078] In an embodiment of the present application, the texture mapping generation model may include two texture encoding modules, and each texture encoding module is used to extract different features. Among them, the first texture encoding module is used to extract the texture features of the face image to be reconstructed, and the second texture encoding module is used to extract the texture features of the map in the form of UV coordinates. Among them, the first texture feature and the second texture feature come from different types of images, and the number of channels of the texture features output by the first texture encoding module and the second texture encoding module is the same, and the size of the feature map is also the same. Therefore, when training the first texture encoding module and the second texture encoding module, different training image data sets will be used to train the first texture encoding module and the second texture encoding module.
[0079] Step S340: Input the map in the form of UV coordinates into the second texture encoding module and obtain the second texture feature output by the second texture encoding module.
[0080] Step S350: Perform a splicing or addition operation on the first texture feature and the second texture feature to obtain a third texture feature.
[0081] In an embodiment of the present application, the schematic diagram of the implementation of steps S310 - S350 may be as shown in Figure 10As shown, first, obtain the face image to be reconstructed and the 3D face shape information. Then, based on the 3D face shape information, calculate the corresponding normal map, and then unfold the normal map into an image in UV coordinate format.
[0082] Then, input the face image to be reconstructed into the first texture encoding module respectively, and input the image in UV coordinate format into the second texture encoding module. Extract the first texture feature and the second texture feature through the first texture encoding module and the second texture encoding module respectively.
[0083] Furthermore, the first texture feature and the second texture feature can be first subjected to image channel splicing or addition operations to obtain the third texture feature.
[0084] Step S360: Input the third texture feature into the texture mapping generation module to obtain the 3D face texture map output by the texture mapping generation module.
[0085] Step S370: Generate the 3D reconstruction image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
[0086] A face reconstruction method provided by this application first extracts features from the face image to be reconstructed and the image in UV coordinate format through the first texture encoding module and the second texture encoding module respectively to obtain the first texture feature and the second texture feature. Then, perform a splicing operation on the first texture feature and the second texture feature to achieve the fusion of shape information. Then, input the fused third texture feature into the texture mapping generation module to generate a 3D face texture map, and generate a 3D reconstruction result according to the 3D face texture map and the 3D face shape information, improving the texture generation effect of the side face area of the single-view face, enhancing the consistency between the 3D shape information and the 3D face texture map, and thus improving the overall effect of 3D face reconstruction.
[0087] Please refer to Figure 11 , a face reconstruction method provided by an embodiment of this application is applied to an electronic device or a server as shown in Figure 1 or Figure 2 shown, and the method includes:
[0088] Step S410: Obtain the face image to be reconstructed and the 3D face shape information corresponding to the face image to be reconstructed.
[0089] In an embodiment of this application, as shown in Figure 12 shown, the specific steps of step S410 may include:
[0090] Step S411: Obtain the initial face image.
[0091] In the embodiments of the present application, the initial face image is the original face image that has not been preprocessed. Among them, the initial face image can be the one captured by the image acquisition device in real time, or the original face image obtained from the cloud server or the preset storage area, and no specific limitation is made here.
[0092] Step S412: According to the face direction, perform an image rotation operation on the initial face image to obtain the rotated initial face image.
[0093] In the embodiments of the present application, the face direction is the pre-set face direction.
[0094] Sometimes, the direction of the face included in the initial face image is different from the pre-set face direction. Therefore, it is necessary to perform an image rotation operation on the initial face image according to the pre-set face direction to adjust the face direction in the initial face image to be the same as the pre-set face direction.
[0095] Step S413: Perform face detection on the rotated initial face image to obtain the initial face image of a preset size including the face area.
[0096] In the embodiments of the present application, the face detection model can be used to perform face detection on the rotated initial face image to obtain the initial face image that only includes the face area, and then perform a size scaling operation on the initial face image that only includes the face area to obtain the initial face image of a preset size.
[0097] Step S414: Perform normalization processing on the initial face image of a preset size including the face area to obtain the face image to be reconstructed.
[0098] In the embodiments of the present application, the normalization processing is to perform the operation of first subtracting the mean value (usually taking the value of 127.5) and then dividing by the variance (usually taking the value of 127.5) on the RGB three-channel values of each pixel point in the initial face image of a preset size including the face area. The calculation formula is as follows: (X - 127.5) / 127.5, where X is the RGB three-channel value of the pixel point. Among them, the mean value and the variance can be pre-set or obtained based on experience.
[0099] Optionally, other methods can also be used for normalization processing, such as directly dividing the RGB three-channel values of each pixel point by 255.
[0100] The schematic diagram of the implementation of the above steps S411 - S414 can be as Figure 13 shown.
[0101] Step S420: Generate a graph in the form of UV coordinates based on the three-dimensional face shape information.
[0102] Step S430: Input the face image to be reconstructed into the first texture encoding module to obtain the first texture feature output by the first texture encoding module.
[0103] Step S440: Input the image in the form of UV coordinates into the second texture encoding module to obtain the second texture feature output by the second texture encoding module.
[0104] Step S450: Perform a splicing or addition operation on the first texture feature and the second texture feature to obtain a third texture feature.
[0105] Step S460: Input the third texture feature into the third texture encoding module to obtain the fourth texture feature output by the third texture encoding module.
[0106] In the embodiment of the present application, the schematic diagram of the implementation of steps S410 - S460 can be as Figure 14 shown. First, obtain the face image to be reconstructed and the 3D face shape information, then first calculate the corresponding normal map based on the 3D face shape information, and then unfold the normal map into an image in the UV coordinate format.
[0107] Then, input the face image to be reconstructed into the first texture encoding module respectively, and input the image in the UV coordinate format into the second texture encoding module, and extract the first texture feature and the second texture feature through the first texture encoding module and the second texture encoding module respectively.
[0108] Furthermore, the first texture feature and the second texture feature can be first subjected to an image channel splicing or addition operation to obtain a third texture feature, and finally the third texture feature is input into the third texture encoding module for feature extraction.
[0109] Step S470: Input the fourth texture feature into the texture mapping generation model to obtain the 3D face texture map output by the texture mapping generation module.
[0110] Step S480: Generate a 3D reconstruction image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
[0111] A face reconstruction method provided by the present application, through the above method, in the process of texture mapping generation, integrates the 3D face shape information, avoids the inconsistency between the 3D shape reconstruction result and the texture map, optimizes the effect of 3D face texture map generation, and further improves the effect of 3D face reconstruction.
[0112] Please refer to Figure 15 , a face reconstruction method provided by the embodiment of the present application is applied to such as Figure 1 or Figure 2The electronic device or server shown, the method includes:
[0113] Step S510: Obtain the face image to be reconstructed and the three-dimensional face shape information corresponding to the face image to be reconstructed.
[0114] Step S520: Generate a graph in the form of UV coordinates based on the three-dimensional face shape information.
[0115] Step S530: Add or exchange the network parameters of the first texture encoding module and the second texture encoding module to obtain the first texture encoding module with updated parameters and the second texture encoding module with updated parameters.
[0116] In the embodiment of the present application, since the first texture encoding module and the second texture encoding module are trained through different training image data sets, the network parameters of the second texture encoding module and the second texture encoding module are different.
[0117] In order to further enhance the fusion of shape features, the network parameters of the first texture encoding module and the second texture encoding module can be added or exchanged, so as to obtain the first texture encoding module with updated parameters and the second texture encoding module with updated parameters.
[0118] Since the first texture encoding module with updated parameters contains the network parameters of the second texture encoding module, the first texture features extracted by the first texture encoding module with updated parameters also have the characteristics of the second texture features at the same time; similarly, since the second texture encoding module with updated parameters contains the network parameters of the first texture encoding module, the second texture features extracted by the second texture encoding module with updated parameters also have the characteristics of the first texture features at the same time.
[0119] Step S540: Input the face image to be reconstructed into the first texture encoding module with updated parameters, and obtain the first texture features output by the first texture encoding module with updated parameters.
[0120] Step S550: Input the graph in the form of UV coordinates into the second texture encoding module with updated parameters, and obtain the second texture features output by the second texture encoding module with updated parameters.
[0121] Step S560: Perform a splicing or adding operation on the first texture features and the second texture features to obtain third texture features.
[0122] In the embodiment of the present application, the schematic diagram of the implementation of steps S510 - S560 can be as Figure 16As shown, first, obtain the face image to be reconstructed and the 3D face shape information. Then, based on the 3D face shape information, calculate the corresponding normal map, and then unfold the normal map into an image in UV coordinate format.
[0123] Then, input the face image to be reconstructed into the first texture encoding module with updated parameters, and input the image in UV coordinate format into the second texture encoding module with updated parameters. Extract the first texture feature and the second texture feature through the first texture encoding module with updated parameters and the second texture encoding module with updated parameters respectively.
[0124] Furthermore, the first texture feature and the second texture feature can be first subjected to an image channel splicing or addition operation to obtain a third texture feature.
[0125] Step S570: Input the third texture feature into the texture mapping generation module to obtain the 3D face texture map output by the texture mapping generation module.
[0126] Step S580: Generate a 3D reconstruction image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
[0127] A face reconstruction method provided by the present application, through the above method, in the process of texture mapping generation, integrates the 3D face shape information, avoids the inconsistency between the 3D shape reconstruction result and the texture map, optimizes the effect of 3D face texture map generation, and further improves the effect of 3D face reconstruction.
[0128] Please refer to Figure 17 , a face reconstruction method provided by an embodiment of the present application, is applied to an electronic device or a server as shown in Figure 1 or Figure 2 shown, and the method includes:
[0129] Step S610: Obtain a face image to be reconstructed and the 3D face shape information corresponding to the face image to be reconstructed.
[0130] Step S620: Generate a map in the form of UV coordinates based on the 3D face shape information.
[0131] Step S630: Add or exchange the network parameters of the first texture encoding module and the second texture encoding to obtain the first texture encoding module with updated parameters and the second texture encoding module with updated parameters.
[0132] Step S640: Input the face image to be reconstructed into the first texture encoding module with updated parameters to obtain the first texture feature output by the first texture encoding module with updated parameters.
[0133] Step S650: Input the graph in the form of UV coordinates into the second texture encoding module after updating the parameters, and obtain the second texture feature output by the second texture encoding module after updating the parameters.
[0134] Step S660: Perform a splicing or addition operation on the first texture feature and the second texture feature to obtain a third texture feature.
[0135] Step S670: Input the third texture feature into the third texture encoding module, and obtain the fourth texture feature output by the third texture encoding module.
[0136] In the embodiment of the present application, the schematic diagram of the implementation of steps S610 - S670 can be as Figure 18 shown. First, obtain the face image to be reconstructed and the three-dimensional shape information of the face. Then, based on the three-dimensional shape information of the face, calculate the corresponding normal map, and then expand the normal map into an image in the UV coordinate format.
[0137] Then, input the face image to be reconstructed into the first texture encoding module after updating the parameters, and input the graph in the UV coordinate format into the second texture encoding module after updating the parameters. Extract the first texture feature and the second texture feature through the first texture encoding module after updating the parameters and the second texture encoding module after updating the parameters respectively.
[0138] Furthermore, the first texture feature and the second texture feature can be first subjected to an image channel splicing or addition operation to obtain a third texture feature, and then the third texture feature is input into the third texture encoding module for texture feature extraction.
[0139] Step S680: Input the fourth texture feature into the texture mapping generation model, and obtain the three-dimensional face texture mapping output by the texture mapping generation module.
[0140] Step S690: Generate the three-dimensional reconstruction image corresponding to the face image to be reconstructed based on the three-dimensional shape information of the face and the three-dimensional face texture mapping.
[0141] A face reconstruction method provided by the present application, through the above method, in the process of texture mapping generation, integrates the three-dimensional shape information of the face, avoids the inconsistency between the three-dimensional shape reconstruction result and the texture mapping, optimizes the effect of three-dimensional face texture mapping generation, and further improves the effect of three-dimensional face reconstruction.
[0142] Please refer to Figure 19 , a face reconstruction device 700 provided by the embodiment of the present application, the device 700 includes:
[0143] A data acquisition unit 710 is configured to acquire a face image to be reconstructed and corresponding three-dimensional face shape information of the face image to be reconstructed.
[0144] As a method, the data acquisition unit 710 is specifically configured to acquire an initial face image; perform an image rotation operation on the initial face image according to the face direction to obtain a rotated initial face image; perform face detection on the rotated initial face image to acquire an initial face image of a preset size including a face region; and perform normalization processing on the initial face image of the preset size including the face region to obtain the face image to be reconstructed.
[0145] A texture map generation unit 720 is configured to input the face image to be reconstructed and the three-dimensional face shape information into a texture map generation model to acquire a three-dimensional face texture map output by the texture map generation model.
[0146] As a method, the texture map generation unit 720 is specifically configured to generate a map in the form of UV coordinates based on the three-dimensional face shape information; perform a splicing or addition operation on the map in the form of UV coordinates and the face image to be reconstructed to obtain a fused face image to be reconstructed; input the fused face image into the texture encoding module to acquire three-dimensional face texture features output by the texture encoding module; and input the three-dimensional face texture features into the texture map generation module to acquire a three-dimensional face texture map output by the texture map generation module.
[0147] As another method, the texture map generation unit 720 is specifically configured to generate a map in the form of UV coordinates based on the three-dimensional face shape information; input the face image to be reconstructed into the first texture encoding module to acquire first texture features output by the first texture encoding module; input the map in the form of UV coordinates into the second texture encoding module to acquire second texture features output by the second texture encoding module; perform a splicing or addition operation on the first texture features and the second texture features to obtain third texture features; and input the third texture features into the texture map generation module to acquire a three-dimensional face texture map output by the texture map generation module.
[0148] Optionally, the texture map generation unit 720 is further specifically configured to input the third texture features into the third texture encoding module to acquire fourth texture features output by the third texture encoding module; and input the fourth texture features into the texture map generation model to acquire a three-dimensional face texture map output by the texture map generation module.
[0149] Optionally, the texture map generation unit 720 is further specifically configured to add or exchange the network parameters of the first texture encoding module and the second texture encoding module to obtain the first texture encoding module with updated parameters and the second texture encoding module with updated parameters; input the face image to be reconstructed into the first texture encoding module with updated parameters to obtain the first texture feature output by the first texture encoding module with updated parameters; and input the map in the form of UV coordinates into the second texture encoding module with updated parameters to obtain the second texture feature output by the second texture encoding module with updated parameters.
[0150] Optionally, the texture map generation unit 720 is further specifically configured to generate a corresponding normal map based on the three-dimensional face shape information; unfold the normal map into a map in the form of UV coordinates; or unfold the three-dimensional face shape information into a map in the form of UV coordinates.
[0151] The image generation unit 730 is configured to generate a three-dimensional reconstructed image corresponding to the face image to be reconstructed based on the three-dimensional face shape information and the three-dimensional face texture map.
[0152] It should be noted that the device embodiments in this application correspond to the foregoing method embodiments. For the specific principles in the device embodiments, reference may be made to the content in the foregoing method embodiments, which will not be elaborated here.
[0153] Next, Figure 20 an electronic device or a server provided in this application will be described.
[0154] Please refer to Figure 20 , based on the foregoing face reconstruction method and device, another electronic device or server 800 capable of executing the foregoing face reconstruction method is further provided in an embodiment of this application. The electronic device or server 800 includes one or more (only one is shown in the figure) processors 802, a memory 804, and a network module 806 that are coupled to each other. Among them, a program that can execute the content in the foregoing embodiments is stored in the memory 804, and the processor 802 can execute the program stored in the memory 804.
[0155] Among them, the processor 802 may include one or more processing cores. The processor 802 connects various parts within the entire server 800 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 804, and by calling the data stored in the memory 804, it executes various functions of the server 800 and processes data. Optionally, the processor 802 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 802 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the displayed content; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 802 and may be implemented separately through a communication chip.
[0156] The memory 804 may include random access memory (RAM) and may also include read-only memory. The memory 804 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 804 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the electronic device or the server 800 (such as phone books, audio and video data, chat record data, etc.).
[0157] The network module 806 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 806 may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The network module 806 can communicate with various networks such as the Internet, intranets, wireless networks or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. For example, the network module 806 can interact with a base station.
[0158] Please refer to Figure 21 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable storage medium 900, and the program code can be called by a processor to execute the method described in the above method embodiment.
[0159] The computer-readable storage medium 900 can be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. Optionally, the computer-readable storage medium 900 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 900 has a storage space for program code 910 for executing any method step in the above method. These program codes can be read out from or written into one or more computer program products. The program code 910 can be compressed in an appropriate form, for example.
[0160] A face reconstruction method, device, electronic device and storage medium provided by the present application first obtain a face image to be reconstructed and corresponding three-dimensional face shape information of the face image to be reconstructed, then input the face image to be reconstructed and the three-dimensional face shape information into a texture map generation model to obtain a three-dimensional face texture map output by the texture map generation model, and finally generate a three-dimensional reconstruction image corresponding to the face image to be reconstructed based on the three-dimensional face shape information and the three-dimensional face texture map. Through the above method, during the texture map generation process, the three-dimensional face shape information is fused, avoiding the inconsistency between the three-dimensional shape reconstruction result and the texture map, optimizing the effect of generating the three-dimensional face texture map, and thus improving the effect of three-dimensional face reconstruction.
[0161] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims, and all of them fall within the protection scope of the present invention.
Claims
1. A face reconstruction method, characterized in that, The method includes: Obtaining a face image to be reconstructed and corresponding three-dimensional face shape information of the face image to be reconstructed, where the three-dimensional face shape information includes face contour information, face expression information, and three-dimensional coordinate information of each vertex in the face; Inputting the face image to be reconstructed and the three-dimensional face shape information into a texture mapping generation model to obtain a three-dimensional face texture map output by the texture mapping generation model; The texture mapping generation model includes a texture encoding module and a texture mapping generation module; the step of inputting the face image to be reconstructed and the three-dimensional face shape information into the texture mapping generation model to obtain the three-dimensional face texture map output by the texture mapping generation model includes: Generating a map in the form of UV coordinates based on the three-dimensional face shape information; Performing a splicing or adding operation on the map in the form of UV coordinates and the face image to be reconstructed to obtain a fused face image to be reconstructed, where the fused face image to be reconstructed is a six-channel image; Inputting the fused face image to be reconstructed into the texture encoding module to obtain three-dimensional face texture features output by the texture encoding module; Inputting the three-dimensional face texture features into the texture mapping generation module to obtain a three-dimensional face texture map output by the texture mapping generation module; Generating a three-dimensional reconstruction image corresponding to the face image to be reconstructed based on the three-dimensional face shape information and the three-dimensional face texture map.
2. The method according to claim 1, characterized in that, The texture mapping generation model further includes a first texture encoding module, a second texture encoding module, and a texture mapping generation module; the step of inputting the face image to be reconstructed and the three-dimensional face shape information into the texture mapping generation model to obtain the three-dimensional face texture map output by the texture mapping generation model includes: Generating a map in the form of UV coordinates based on the three-dimensional face shape information; Inputting the face image to be reconstructed into the first texture encoding module to obtain first texture features output by the first texture encoding module; Inputting the map in the form of UV coordinates into the second texture encoding module to obtain second texture features output by the second texture encoding module; Performing a splicing or adding operation on the first texture features and the second texture features to obtain third texture features; Inputting the third texture features into the texture mapping generation module to obtain a three-dimensional face texture map output by the texture mapping generation module.
3. The method according to claim 2, characterized in that, The texture mapping generation model further includes a third texture encoding module, and after performing the splicing or adding operation on the first texture features and the second texture features to obtain third texture features, it further includes: Inputting the third texture features into the third texture encoding module to obtain fourth texture features output by the third texture encoding module; The step of inputting the third texture features into the texture mapping generation module to obtain a three-dimensional face texture map output by the texture mapping generation module includes: Inputting the fourth texture features into the texture mapping generation model to obtain a three-dimensional face texture map output by the texture mapping generation module.
4. The method according to claim 2, wherein The step of inputting the face image to be reconstructed into the first texture encoding module to obtain first texture features output by the first texture encoding module includes: Add or exchange the network parameters of the first texture encoding module and the second texture encoding module to obtain the first texture encoding module with updated parameters and the second texture encoding module with updated parameters; Input the face image to be reconstructed into the first texture encoding module with updated parameters to obtain the first texture feature output by the first texture encoding module with updated parameters; The step of inputting the map in the form of UV coordinates into the second texture encoding module to obtain the second texture feature output by the second texture encoding module includes: Input the map in the form of UV coordinates into the second texture encoding module with updated parameters to obtain the second texture feature output by the second texture encoding module with updated parameters.
5. The method according to claim 1 or 2, characterized in that, The step of generating a map in the form of UV coordinates based on the 3D face shape information includes: Generate a corresponding normal map based on the 3D face shape information; Unfold the normal map into a map in the form of UV coordinates; or, Unfold the 3D face shape information into a map in the form of UV coordinates.
6. The method according to claim 1, characterized in that, The step of obtaining the face image to be reconstructed includes: Obtain an initial face image; Perform an image rotation operation on the initial face image according to the face direction to obtain the rotated initial face image; Perform face detection on the rotated initial face image to obtain an initial face image of a preset size including the face region; Perform normalization processing on the initial face image of the preset size including the face region to obtain the face image to be reconstructed.
7. A face reconstruction device, characterized in that, The apparatus includes: A data acquisition unit, configured to acquire a face image to be reconstructed and the 3D face shape information corresponding to the face image to be reconstructed, where the 3D face shape information includes face contour information, face expression information, and 3D coordinate information of each vertex in the face; A texture map generation unit, configured to input the face image to be reconstructed and the 3D face shape information into a texture map generation model to obtain a 3D face texture map output by the texture map generation model; the texture map generation model includes a texture encoding module and a texture map generation module; the step of inputting the face image to be reconstructed and the 3D face shape information into the texture map generation model to obtain the 3D face texture map output by the texture map generation model includes: generating a map in the form of UV coordinates based on the 3D face shape information; performing a splicing or adding operation on the map in the form of UV coordinates and the face image to be reconstructed to obtain a fused face image to be reconstructed, where the fused face image to be reconstructed is a six-channel image; inputting the fused face image to be reconstructed into the texture encoding module to obtain the 3D face texture feature output by the texture encoding module; inputting the 3D face texture feature into the texture map generation module to obtain the 3D face texture map output by the texture map generation module; An image generation unit, configured to generate a 3D reconstructed image corresponding to the face image to be reconstructed based on the 3D face shape information and the 3D face texture map.
8. An electronic device, characterized in that, Comprising one or more processors; one or more programs are stored in the memory and configured to be executed by the one or more processors to perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, wherein, when the program code is run by a processor, the method according to any one of claims 1-6 is performed.
Citation Information
Patent Citations
Model head portrait creation method and device, electronic equipment and storage medium
CN112669447A
Living body face detection model training method, device, equipment and storage medium
CN114093006A