3D face reconstruction

By fitting the combination of neural networks and image conversion networks, a high-resolution three-dimensional face model is generated, which solves the problem of unrealistic three-dimensional face texture in the existing technology, and achieves a more realistic skin rendering effect.

CN114746904BActive Publication Date: 2025-08-08HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180006744.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-21
Filing Date
2021-02-20
Publication Date
2025-08-08
Estimated Expiration
2041-02-20

AI Technical Summary

Technical Problem

When the existing methods reconstruct three-dimensional face texture, the generated quality is unreal and lacks details, and are prone to falling into the ‘uncanny valley’ phenomenon.

Method used

A fitted neural network is used to generate three-dimensional shape models and low-resolution texture maps, combining super-resolution models and de-light and shadow images to image conversion neural networks to generate high-resolution diffuse albedo maps, and detail improvement is achieved by rendering high-resolution three-dimensional models.

Benefits of technology

It improves the details quality of three-dimensional face rendering, generates more realistic skin textures, overcomes the ‘uncanny valley’ phenomenon, and enhances the realism of the texture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114746904B_ABST
    Figure CN114746904B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for reconstructing a three-dimensional face model from a two-dimensional face image. A computer-implemented method for generating a three-dimensional face rendering from a two-dimensional image containing a face image is disclosed. The method comprises: using a fitting neural network to generate a three-dimensional shape model of the face image and a low-resolution two-dimensional texture map of the face image from the two-dimensional image (2.1); applying a super-resolution model to the low-resolution two-dimensional texture map to generate a high-resolution two-dimensional texture map (2.2); using a de-shaded image-to-image conversion neural network to generate a two-dimensional diffuse albedo map from the high-resolution texture map (2.3); and rendering a high-resolution three-dimensional model of the face image using the two-dimensional diffuse albedo map and the three-dimensional shape model (2.4).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the benefit of priority to UK patent application number GB2002449.3 filed on February 21, 2020, entitled “Three-dimensional Facial Reconstruction”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This specification describes a method and system for reconstructing a three-dimensional face model from a two-dimensional face image. Background Art

[0003] Reconstructing three-dimensional (3D) faces and textures from two-dimensional (2D) images is one of the most extensively and deeply studied areas at the intersection of computer vision, graphics, and machine learning. This is because, in addition to its countless applications, it represents a new advancement in learning, inferring, and synthesizing the geometry of 3D objects. Recently, thanks largely to the advent of deep learning, significant progress has been made in reconstructing smooth 3D facial geometry even from 2D images captured under arbitrary recording conditions (also known as "in the wild").

[0004] However, while the geometry can be inferred accurately to a certain extent, the quality of the generated texture is still unrealistic, and 3D face renderings generated by existing methods usually lack details and fall into the "uncanny valley". Summary of the Invention

[0005] According to a first aspect, this specification discloses a computer-implemented method for generating a three-dimensional face rendering from a two-dimensional image containing a face image. The method comprises: using one or more fitted neural networks to generate a three-dimensional shape model of the face image and a low-resolution two-dimensional texture map of the face image from the two-dimensional image; applying a super-resolution model to the low-resolution two-dimensional texture map to generate a high-resolution two-dimensional texture map; using a de-shaded image-to-image conversion neural network to generate a two-dimensional diffuse albedo map from the high-resolution texture map; and rendering a high-resolution three-dimensional model of the face image using the two-dimensional diffuse albedo map and the three-dimensional shape model.

[0006] The two-dimensional diffuse albedo map may be a high-resolution two-dimensional diffuse albedo map.

[0007] The method may further comprise determining a two-dimensional normal map of the facial image based on the three-dimensional shape model, wherein, additionally, the two-dimensional diffuse albedo map is generated using the two-dimensional normal map.

[0008] The method may further include: generating a two-dimensional specular albedo map from the two-dimensional diffuse albedo map using a specular albedo image-to-image conversion neural network, wherein the high-resolution three-dimensional model of the facial image is further rendered based on the two-dimensional specular albedo map. The method may further include: generating a grayscale two-dimensional diffuse albedo map from the two-dimensional diffuse albedo map; and inputting the grayscale two-dimensional diffuse albedo map into the specular albedo image-to-image conversion neural network. The method may further include: determining a two-dimensional normal map for the facial image based on the three-dimensional shape model, wherein, in addition, the two-dimensional specular albedo map is generated from the two-dimensional normal map using the specular albedo image-to-image conversion neural network.

[0009] The method may further include: determining a two-dimensional normal map for the facial image based on the three-dimensional shape model; generating a two-dimensional specular normal map based on the two-dimensional diffuse albedo map and the two-dimensional normal map using a specular normal image to image conversion neural network, wherein the high-resolution three-dimensional model of the facial image is further rendered based on the two-dimensional specular normal map. Generating the two-dimensional specular normal map using the specular normal image to image conversion neural network may include: generating a grayscale two-dimensional diffuse albedo map based on the two-dimensional diffuse albedo map; and inputting the grayscale two-dimensional diffuse albedo map and the two-dimensional normal map into the specular normal image to image conversion neural network.

[0010] The two-dimensional normal map may be a two-dimensional normal map in tangent space. Generating the two-dimensional normal map in tangent space based on the three-dimensional shape model may include: generating a two-dimensional normal map in object space based on the three-dimensional shape model; and applying a high-pass filter to the two-dimensional normal map in object space.

[0011] The method may further include: determining a two-dimensional normal map in object space for the facial image based on the three-dimensional shape model; generating a two-dimensional diffuse normal map based on the two-dimensional diffuse albedo map and the two-dimensional normal map in tangent space using a diffuse normal image to image conversion neural network, wherein the high-resolution three-dimensional model of the facial image is further rendered based on the two-dimensional diffuse normal map. Generating the two-dimensional diffuse normal map using the diffuse normal image to image conversion neural network may include: generating a grayscale two-dimensional diffuse albedo map based on the two-dimensional diffuse albedo map; and inputting the grayscale two-dimensional diffuse albedo map and the two-dimensional normal map in tangent space into the diffuse normal image to image conversion neural network.

[0012] The method may further include, for each image-to-image conversion neural network: dividing the input two-dimensional map into a plurality of overlapping input blocks; generating an output block for each of the input blocks using the image-to-image conversion neural network; and generating a complete output two-dimensional map by combining the plurality of output blocks.

[0013] The fitting neural network and / or the image-to-image conversion network may be a generative adversarial network.

[0014] The method may further include generating a three-dimensional model of a head from the high-resolution three-dimensional model of the face image using a combined face and head model.

[0015] One or more of the two-dimensional maps may include a UV map.

[0016] According to another aspect, the present specification discloses a system comprising one or more processors and a memory, wherein the memory comprises computer-readable instructions that, when executed by the one or more processors, cause the system to perform any one or more of the methods disclosed herein.

[0017] According to another aspect, the present specification discloses a computer program product comprising computer-readable instructions, which, when executed by a computing system, cause the computing system to perform any one or more of the methods disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various embodiments will now be described by way of non-limiting examples with reference to the following drawings, in which:

[0019] Figure 1 A schematic diagram illustrating an example method for generating a three-dimensional face rendering from a two-dimensional image;

[0020] Figure 2 A flow chart illustrating an example method for generating a three-dimensional face rendering from a two-dimensional image;

[0021] Figure 3 A schematic diagram illustrating another example method for generating a three-dimensional face rendering from a two-dimensional image;

[0022] Figure 4 A schematic diagram illustrating an example method for training an image-to-image neural network;

[0023] Figure 5 Illustrative examples of systems / apparatus for performing any of the methods described herein are shown. DETAILED DESCRIPTION

[0024] To achieve realistic human skin rendering, diffuse reflectance (albedo) is modeled. Given a low-resolution 2D texture map (e.g., UV map) and the basic geometry reconstructed from a single unconstrained face image as input, the diffuse albedo A is inferred by applying a super-resolution model to the low-resolution 2D texture map to generate a high-resolution texture map. D , and then pass it through the deshading network to obtain a high-resolution diffuse albedo. The diffuse albedo shows the color of the light "emitted" by the skin. The diffuse albedo, high-resolution texture map, and base geometry can be used to render a high-quality 3D face model. Other components (e.g., diffuse normal, specular albedo, and / or specular normal) can be inferred from the diffuse albedo in combination with the base geometry, and these components can be used to render a high-quality 3D face model.

[0025] Figure 1 A schematic diagram illustrates an example method 100 for generating a three-dimensional rendering of a human face from a two-dimensional image. The method can be implemented on a computer. A 2D image 102 including a human face is input into one or more fitted neural networks 104, which generate a low-resolution 2D texture map 106 of the human face's texture and a 3D model 108 of the human face's geometry. A super-resolution model 110 is applied to the low-resolution 2D texture map 106 to upscale the low-resolution 2D texture map 106 to a high-resolution 2D texture map 112. An image-to-image conversion neural network 114 (also referred to herein as a "delighted image-to-image conversion network") is used to generate a 2D diffuse albedo map 116 from the high-resolution 2D texture map 112. The 2D diffuse albedo map 116 is used to render the 3D model 108 of the human face's geometry to generate a high-resolution 3D model 118 of the human face in the input image 102.

[0026] The input 2D image 102 (I) comprises a set of pixel values in an array. For example, in a color image Where H is the height of the image (in pixels), W is the width of the image (in pixels), and the image has three color channels (e.g., RGB or CIELAB). Alternatively, the input 2D image 102 may be a black and white or grayscale image. The input image may be cropped from the larger image based on the detection of a face in the larger image.

[0027] One or more fitted neural networks 104 generate 3D face shapes 108 and low-resolution 2D texture maps 106 Where N is the number of vertices in the 3D face shape mesh, H LR and W LRare the height and width of the low-resolution 2D texture map 106, respectively. In some embodiments, a single fitted neural network is used to generate the 3D face shape 108 and the low-resolution 2D texture map 106. This can be represented symbolically as:

[0028]

[0029] in, is a fitting neural network. The fitting neural network can be based on a Generative Adversarial Network (GAN) architecture. An example of such a network is described in “GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction” (B. Gecer et al., Transactions of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1155-1164, 2019), the contents of which are incorporated herein by reference. However, any neural network or model trained to fit a 3D face shape 108 to an image and / or generate a 2D texture map from a 2D image 102 can be used. In some embodiments, a separate fitting neural network is used to generate each of the 3D face shape 108 and the low-resolution 2D texture map 106.

[0030] The low-resolution 2D texture map 106 can be any 2D map that can represent a 3D texture. An example of such a map is a UV map. A UV map is a 2D representation of a 3D surface or mesh. Points in 3D space (e.g., described by (x, y, z) coordinates) are mapped to 2D space (described by (u, v) coordinates). The UV map can be formed by unfolding a 3D mesh in 3D space onto a uv plane in 2D UV space and storing parameters associated with the 3D surface at each point in the UV space. The texture UV map 110 can be formed by storing the color values of the vertices of the 3D surface / mesh in 3D space at corresponding points in the UV space.

[0031] The super-resolution model 110 converts the low-resolution texture map 106 As input, and generate a high-resolution texture map 112 from it Among them, H HR and W HR are the height and width of the high-resolution 2D texture map 112, H HR >H LR And W HR >W LR . This can be expressed symbolically as:

[0032]

[0033] in, is a super-resolution model. The super-resolution model 110 can be a neural network. The super-resolution model 110 can be a convolutional neural network. An example of such a super-resolution neural network is the RCAN described in “Image super-resolution using very deep residual channel attention networks” (Y. Zhang et al., Transactions of the European Conference on Computer Vision (ECCV), pp. 286-301, 2018), the contents of which are incorporated herein by reference, but any example of a super-resolution neural network can be used. The super-resolution neural network can be trained on data including low-resolution texture maps, each low-resolution texture map having a corresponding high-resolution texture map.

[0034] The high-resolution 2D texture map 112 may be any 2D map capable of representing a 3D texture, such as a UV map (as described above in conjunction with the low-resolution 2D texture map 106 ).

[0035] The de-shaded image-to-image conversion network 114 converts the high-resolution texture map 112 As input, and generate a 2D diffuse albedo map from it 116 Among them, H D and W D are the height and width of the high-resolution 2D diffuse albedo map 116, respectively. Typically, low-resolution textures generated by fitting a neural network contain baked lighting (e.g., reflections, shadows) because the fitted neural network has been trained on a large dataset of objects captured under near-constant lighting generated by ambient light and three-point light sources. Consequently, the captured objects contain sharp highlights and shadows that cannot achieve photorealistic rendering.

[0036] A satisfactory neural network can be pre-trained to generate the de-shaded diffuse albedo from the high-resolution texture map 112, as described below in conjunction with Figure 4 described.

[0037] The satisfactory image-to-image translation network 116 can be represented symbolically as:

[0038]

[0039] in, In some embodiments, a 2D normal map derived from the 2D input image may additionally be input into the image conversion network 116 to create a satisfactory image, as described below in conjunction with Figure 3 In some embodiments, the satisfactory image is fed into the image conversion network 116 (in embodiments using this map together with the 2D normal map). Figure 1 The high-resolution 2D texture map 112 is normalized to the range [-1, 1]. The normalized high-resolution texture can be represented as

[0040] Image-to-image conversion refers to the task of converting an input image into a specified target domain (for example, converting a sketch into an image, or converting a daytime scene into a nighttime scene). Image-to-image conversion typically utilizes a generative adversarial network (GAN) based on the input image. The image-to-image conversion network disclosed in this article (for example, a satisfactory image-to-image conversion network / specular albedo / diffuse normal / specular normal image-to-image conversion network) can utilize such a GAN. The GAN architecture includes: a generator network for generating a converted image based on an input image; and a discriminator network for determining whether the converted image is a reasonable conversion of the input image. The generator and the discriminator are trained in an adversarial manner; the purpose of training the discriminator is to distinguish the converted image from the corresponding ground truth image, while the purpose of training the generator is to generate a converted image to deceive the discriminator. The following text is combined with Figure 4 Describes an example of training an image-to-image translation network.

[0041] An example of an image-to-image translation network is pix2pixHD, which can be found in “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs” (TC Wang et al., Transactions of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8798–8807, 2018), which is incorporated herein by reference. Variants of pix2pixHD can be trained to perform tasks such as deshading and extracting diffuse and specular components from ultra-high resolution data. The pix2pixHD network can be modified to take as input a 2D map and a shape normal map. In the global generator, the pix2pixHD network can have nine residual blocks. In the local generator, the pix2pixHD network can have three residual blocks.

[0042] The 2D diffuse albedo map 116 is used to render the 3D model 108 of the face geometry to generate a high-resolution 3D model 118 of the face in the input image 102. To render the 3D model 118 under different lighting conditions, the 2D diffuse albedo map 116 can be re-lit using an arbitrary lighting environment.

[0043] Figure 2 A flow chart of an example method 200 for generating a three-dimensional face rendering from a two-dimensional image is shown. The method can be implemented on a computer.

[0044] In operation 2.1, one or more fitting neural networks are used to generate a 3D shape model of the face image and a low-resolution 2D texture map of the face image based on the 2D image. The fitting neural network may be a generative adversarial network.

[0045] In some embodiments, one or more 2D normal maps for the facial image may be generated based on the 3D shape model. The one or more 2D normal maps may include a normal map in object space and / or a normal map in tangent space. The normal map in tangent space may be generated by applying a high-pass filter to the normal map in object space.

[0046] In operation 2.2, a super-resolution model is applied to the low-resolution 2D texture map to generate a high-resolution 2D texture map. The super-resolution model may be a super-resolution neural network. The super-resolution neural network may include one or more convolutional layers.

[0047] In operation 2.3, a 2D diffuse albedo map is generated from the high-resolution texture map using a de-shaded image-to-image conversion neural network. The 2D diffuse albedo map can be a high-resolution 2D diffuse albedo map. The de-shaded image-to-image conversion neural network can be a GAN. Alternatively, the 2D diffuse albedo map can be generated using the 2D normal map.

[0048] Furthermore, one or more other 2D maps may be generated using a corresponding image-to-image translation network.

[0049] A specular albedo image-to-image conversion neural network can be used to generate a 2D specular albedo map from the 2D diffuse albedo map (or a grayscale version of the 2D diffuse albedo map). Alternatively, the specular albedo image-to-image conversion neural network can be used to generate the 2D specular albedo map from the 2D normal map, i.e., the 2D normal map and the 2D diffuse albedo map can be input into the specular albedo image-to-image conversion neural network.

[0050] A diffuse normal image to image conversion neural network may be used to generate a 2D diffuse normal map from the 2D diffuse albedo map (or a grayscale version of the 2D diffuse albedo map) and a 2D normal map in tangent space.

[0051] A specular normal image to image conversion neural network may be used to generate a two-dimensional specular normal map from the two-dimensional diffuse albedo map (or a grayscale version of the 2D diffuse albedo map) and the two-dimensional normal map.

[0052] In operation 2.4, a high-resolution 3D model of the facial image is rendered using the 2D diffuse albedo map and the 3D shape model. The 3D model of the facial image may also be rendered using one or more other texture maps. A 3D head model may be generated from the high-resolution 3D model of the facial image using a combined face and head model. During the rendering process, different lighting environments may be applied to the 2D diffuse albedo map.

[0053] Figure 3 A schematic diagram of another example method 300 for generating a three-dimensional face rendering based on a two-dimensional image is shown. The method 300 can be implemented on a computer. Figure 1 , a 2D image 302 including a face is input into one or more fitted neural networks 304, which generate a low-resolution 2D texture map 306 of the face's texture and a 3D model 308 of the face's geometry. A super-resolution model 310 is applied to the low-resolution 2D texture map 306 to upscale the low-resolution 2D texture map 306 to a high-resolution 2D texture map 312. An image-to-image conversion neural network 314 is used to generate a 2D diffuse albedo map 316 from the high-resolution 2D texture map 112.

[0054] The 3D model 308 of the face geometry can be used to generate one or more 2D normal maps 324, 330 for the face. The 2D normal map 324 in object space can be generated directly from the 3D model 308 of the face geometry. A high-pass filter can be applied to the 2D normal map 324 in object space to generate the 2D normal map 324 in tangent space. The normal for each vertex of the 3D model can be calculated as a vector perpendicular to two vectors of a "face" (e.g., a triangle) of the 3D mesh. The normals can be stored in an image format using UV map parameterization. Interpolation can be used to create a smooth normal map.

[0055] In some embodiments, when generating the diffuse albedo map 316, one or more of the 2D normal maps 324, 330 may be input to the image network 314 in addition to the high-resolution texture map 312. Specifically, the 2D normal map 324 in tangent space may be input. The used 2D normal maps 324, 330 may be concatenated with the high-resolution texture map 312 (or a normalized version thereof) and input to the diffuse albedo image. Including the 2D normal maps 324, 330 in the input may reduce the variation in edit shading in the output diffuse albedo map 316. Because lighting occlusion on the skin surface is geometry-dependent, the quality of the albedo map may be improved when the network is fed with both the texture and geometry of the 3D MIME. The shape normals may serve as a geometry "guide" for the image-to-image translation network.

[0056] Other 2D maps can be generated from the 2D diffuse albedo map 316. One example is the specular albedo map 322, which can be generated from the diffuse albedo map 316 using the specular albedo image-to-image conversion neural network 320. Specular albedo 322 acts as a multiplier for the intensity of reflected light and is independent of color. Specular albedo 322 is defined by the composition and roughness of the skin. Therefore, its value can be inferred by differentiating between skin locations (e.g., facial hair, bare skin).

[0057] In principle, specular albedo can be calculated from a baked-lighting texture, as long as the texture includes baked specular lighting. However, the specular component derived using this approach can be significantly biased by ambient lighting and occlusion. Inferring specular albedo from diffuse albedo can result in higher-quality specular albedo maps.

[0058] To generate the specular albedo map 322 (A s ), the diffuse albedo map 316 is input into the image-to-image conversion network 320. The diffuse albedo map 316 may be pre-processed before being input into the image-to-image conversion network 320. For example, the diffuse albedo map 316 may be converted into a grayscale diffuse albedo map. (For example, using In some embodiments, a shape normal map (eg, a shape normal map N in object space) O ) is also input into the image conversion network 320.

[0059] The specular albedo image to image conversion network 320 processes its input through multiple layers and outputs a specular albedo map 322. In an embodiment where only a diffuse albedo map is used, this process can be represented symbolically as:

[0060] A s =ψ(A D ).

[0061] In embodiments where a shape normal map in object space is also input and the diffuse albedo map is converted to a grayscale map, this can be represented symbolically as:

[0062]

[0063] in, H s and W s denote the height and width of the specular albedo map 112, respectively. In some embodiments, H s and W s Equal to H D and W D The generated 2D specular albedo map 322 may be a UV map.

[0064] The generated 2D specular albedo map 322 is used together with the diffuse albedo map 316 and the 3D model 308 of the facial geometry to render the 3D face model 318 .

[0065] Alternatively or additionally, a diffuse normal image to image conversion network 326 can be used to generate a diffuse normal map 328. The diffuse normal is highly correlated with the shape normal because diffuse reflections are evenly distributed across the skin. Scars and wrinkles change the distribution of the diffuse reflections and some non-skin features, such as hair, produce less subsurface scattering.

[0066] To generate the diffuse normal map 328(N D ), the diffuse albedo map 316 is input into the image-to-image conversion network 326 along with the shape normal maps 324 and 330. The diffuse albedo map 316 may be pre-processed before being input into the image-to-image conversion network 320. For example, the diffuse albedo map 316 may be converted into a grayscale diffuse albedo map, as described above with respect to the specular albedo map 322. The shape normal map may be the shape normal map 324 (N in object space) o ).

[0067] The diffuse normal image to image conversion network 326 processes its input through multiple layers and outputs a diffuse normal map 328. In an embodiment where a shape normal map in object space is input and the diffuse albedo map is converted to a grayscale map, this can be represented symbolically as:

[0068]

[0069] in, H ND and W ND denote the height and width of the diffuse normal map 328, respectively. In some embodiments, H ND and W ND Equal to H D and W D The generated 2D diffuse normal map 328 may be a UV map.

[0070] The generated 2D diffuse normal map 328 is used together with the diffuse albedo map 316 of the face geometry and the 3D model 308 to render the 3D face model 318. Additionally, a 2D specular albedo map 322 may be used.

[0071] Alternatively or additionally, a specular normal image-to-image conversion network 332 can be used to generate a specular normal map 334. Specular normals exhibit sharp surface details, such as fine lines and skin pores, and are difficult to estimate because some high-frequency details do not appear in the lighting texture or the estimated diffuse albedo. While a high-resolution texture map 312 can be used to generate the specular normal map 334, it includes sharp highlights that may be incorrectly interpreted by the network as facial features. Diffuse albedo, even when stripped from specular reflections, contains texture information that defines mid- and high-frequency details, such as pores and wrinkles.

[0072] To generate the specular normal map 334(N s ), the diffuse albedo map 316 is input into the image-to-image conversion network 332 along with the shape normal maps 324 and 330. The diffuse albedo map 316 may be pre-processed before being input into the image-to-image conversion network 320. For example, the diffuse albedo map 316 may be converted into a grayscale diffuse albedo map, as described above with respect to the specular albedo map 322. The shape normal map may be a shape normal map 330 in tangent space (N T ).

[0073] The specular normal image to image conversion network 332 processes its input through multiple layers and outputs a specular normal map 334. In an embodiment where a shape normal map in tangent space is input and the diffuse albedo map is converted to a grayscale map, this can be represented symbolically as:

[0074]

[0075] in, H Ns and W NsD denote the height and width of the specular normal map 334, respectively. In some embodiments, HNs and W Ns Equal to H D and W D The generated 2D specular normal map 334 may be a UV map. In some embodiments, the specular normal map 334 is passed through a high pass filter to constrain it to tangent space.

[0076] The generated 2D specular normal map 334 is used along with the diffuse albedo map 316 of the face geometry and the 3D model 308 to render the 3D face model 318. Additionally, the 2D specular albedo map 322 and / or the diffuse normal map 328 may be used.

[0077] The inferred normal (i.e. N D and N s ) can be used to enhance the base reconstructed geometry by refining its mid-range frequencies and adding reasonable high frequency detail. The specular normals 334 can be integrated in tangent space to produce a detailed displacement map, which can then be imprinted on the subdivided base geometry.

[0078] A high-resolution 3D face model 318 is generated based on the 3D model of facial geometry 308 and one or more of the 2D maps 316 , 322 , 328 , 334 .

[0079] In some embodiments, an entire head model can be generated from the face model 318. The face mesh can be projected onto the subspaces, and the latent head parameters can be regressed according to the learned regression matrix that performs alignment between the subspaces. An example of such a model is the combined face and head model described in "Combining 3D morphable models: A large scale face-and-head model" (S. Ploumpis et al., IEEE Transactions on Computer Vision and Pattern Recognition, pp. 10934-10943, 2019), the contents of which are incorporated herein by reference.

[0080] Figure 4A schematic diagram of a method 400 for training an image-to-image translation network is shown. An input 2D map 402 (and in some embodiments, a 2D normal map 404) s from a training dataset is input to a generator neural network 406G. The generator neural network 406 generates a translated 2D map 408G(s) based on the input 2D map 402 (and in some embodiments, the 2D normal map 404). The input 2D map 402 and the translated 2D map 408 are input to a discriminator neural network 410D to generate a score 412D(s, G(s)) indicating how well the discriminator 410 finds the translated 2D map 408. Furthermore, the input 2D map 402 and the corresponding ground-truth translated 2D map 414x are input to the discriminator neural network 410 to generate a score 412D(s, x) indicating how well the discriminator 410 finds the ground-truth translated 2D map 414. The parameters of the discriminator 410 are determined according to the discriminator objective function 416 (Compare these scores 412) to update. The parameters of the generator 406 are updated according to the generator objective function 418 (Comparing these scores 412 and comparing the generated transformed 2D map 408 to the ground truth transformed 2D map 414) is updated. This process can be iterated on the training dataset until a threshold condition is met, such as reaching a threshold number of training stages or achieving a balance between the generator 406 and the discriminator 410. After training, the generator 406 can be used in an image-to-image translation network.

[0081] The training dataset includes a plurality of training examples 420. The training dataset can be divided into a plurality of training batches, each batch including a plurality of training examples. Each training example includes an input 2D map 402 and a corresponding ground-truth converted 2D map 414 of the input 2D map 402. The type of input 2D map 402 and the type of ground-truth converted 2D map 414 in the training example depend on the type of image-to-image conversion network 406 being trained. For example, if a satisfactory image-to-image conversion network is being trained, the input 2D map 402 is a high-resolution texture map and the ground-truth converted 2D map 414 is a ground-truth diffuse albedo map. If a specular albedo image-to-image conversion network is being trained, the input 2D map 402 is a diffuse albedo map (or a grayscale diffuse albedo map) and the ground-truth converted 2D map 414 is a ground-truth specular albedo map. If a diffuse normal image-to-image conversion network is being trained, the input 2D map 402 is a diffuse albedo map (or grayscale diffuse albedo map) and the ground truth converted 2D map 414 is a ground truth diffuse normal map. If a specular normal image-to-image conversion network is being trained, the input 2D map 402 is a diffuse albedo map (or grayscale diffuse albedo map) and the ground truth converted 2D map 414 is a ground truth specular normal map.

[0082] Each training example may also include a normal map 404 corresponding to the image from which the input 2D map 402 was derived. The normal map 404 may be a normal map in object space or a normal map in tangent space. The normal map 404 may be the same as the input 2D map. Figure 1 The normal map is input into the generator neural network 406 to generate the converted 2D map 408. In some embodiments, it can also be input into the discriminator neural network 410 when determining the authenticity score 412. In some embodiments, the normal map is input into the generator neural network 406 instead of the discriminator neural network 410.

[0083] Training examples can be captured using any method known in the art. For example, training examples can be captured from objects illuminated by a polarized LED sphere using the method described in "Multiview face capture using polarized spherical gradient illumination" (A. Ghosh et al., ACM Transactions on Graphics (TOG), vol. 30, p. 129, ACM, 2011) to capture high-resolution pore-level geometry and reflectance maps of faces.

[0084] Half of the LEDs on the sphere can be vertically polarized (for parallel polarization), while the other half can be horizontally polarized (for cross polarization), in a staggered pattern. When using an LED sphere, a multi-view facial capture method can be used, such as the one described in "Multi-view facial capture using binary spherical gradient illumination" (A. Lattas et al., ACM SIGGRAPH 2019 Poster, page 59, ACM, 2019), which separates diffuse and specular components based on color space analysis. Compared to other methods, these methods produce very clear results, require less data capture (thus reducing capture time), and have a simpler setup (no polarizers), enabling the capture of large datasets.

[0085] To generate the ground truth diffuse albedo map, the lighting conditions of the dataset can be modeled using a model of the cornea of the eye, and then a 2D map with the same lighting can be synthesized in order to train an image-to-image translation network from textures with baked lighting to de-lighted diffuse albedo. Using the cornea model of the eye, the average direction of three point light sources relative to the object is determined. In addition, an environment map for the texture is determined. The environment map is a good estimate of the color of the scene, while the three point light sources help simulate highlights. Physically based renderings of each object captured from all viewpoints are generated using the predicted environment map and the predicted light sources (optionally including random variations in their positions), and an illuminated (normalized) texture map is generated. The simulation process can be represented symbolically as: It converts the diffuse albedo into a texture distribution

[14] as follows:

[0086]

[0087] The generator 406 may have a U-Net architecture. The discriminator 410 may be a convolutional neural network. The discriminator 410 may have a fully convolutional structure.

[0088] Using the Discriminator Objective Function 416 To train the discriminator neural network 410, a loss function compares the scores 412 {s, x} generated by the discriminator 410 from the training examples with the scores 412 {s, G(x)} generated by the discriminator 410 from the output of the discriminator. The discriminator objective function 416 can be based on the difference between the expected values of these scores obtained in the training batch. An example of such a loss function is shown below:

[0089]

[0090] An optimization process such as stochastic gradient descent or the Adam optimization algorithm (eg, β1 = 0.5 and β2 = 0.999) may be applied to the discriminator objective function 416 with the goal of maximizing the objective function to determine parameter updates.

[0091] Using Generator Objective Function 418 To train the generator neural network 406, the function compares the scores 412 {s, x} generated by the discriminator 410 from the training examples with the scores 412 {s, G(x)} generated by the discriminator 410 from the output of the discriminator. The generator objective function 418 may include a term that compares the scores 412 {s, x} generated by the discriminator 410 from the training examples with the scores 412 {s, G(x)} generated by the discriminator 410 from the output of the discriminator 410 (i.e., may include the same terms as the discriminator loss 416). For example, the generator objective function 418 may include the term The generator objective function 418 may also include a term that compares the converted 2D map 408 to the ground truth converted 2D map 414. For example, the generator objective function 418 may include a norm (e.g., an L1 or L2 norm) of the difference between the converted 2D map 408 and the ground truth converted 2D map 414. An optimization process such as stochastic gradient descent or an Adam optimization algorithm (e.g., β1 = 0.5 and β2 = 0.999) may be applied to the discriminator objective function 418 with the goal of minimizing the objective function to determine parameter updates.

[0092] During training, high-resolution data can be split into blocks (e.g., blocks of size 512x512 pixels) to increase the number of data samples and avoid overfitting. For example, using a stride of a given size (e.g., 128 pixels), partially overlapping blocks can be derived by passing through each original 2D map (e.g., UV map) horizontally and vertically. Block-based methods can also help overcome hardware limitations (e.g., some high-resolution images cannot be processed even with a 32GB memory graphics card).

[0093] Preferably, the term "neural network" as used herein is used to refer to a model comprising a plurality of node layers, each node being associated with one or more parameters. The parameters of each node of the neural network may include one or more weights and / or biases. The node takes as input one or more outputs of a node in the previous layer of the network (or the values of the input data in the initial layer). One or more outputs of a node in the previous layer are used by the node to generate an activation value using an activation function and the parameters of the neural network. One or more layers of the neural network may be convolutional layers, each layer being used to apply one or more convolution filters. One or more layers of the neural network may be fully connected layers. The neural network may include one or more skip connections.

[0094] Figure 5 Schematic examples of systems / devices for performing any of the methods described herein are shown. The systems / devices shown are examples of computing devices. Those skilled in the art will appreciate that other types of computing devices / systems, such as distributed computing systems, may alternatively be used to implement the methods described herein.

[0095] The device (or system) 500 includes one or more processors 502. The one or more processors control the operation of other components of the system / device 500. For example, the one or more processors 502 may include a general-purpose processor. The one or more processors 502 may be a single-core device or a multi-core device. The one or more processors 502 may include a central processing unit (CPU) or a graphics processing unit (GPU). Alternatively, the one or more processors 502 may include dedicated processing hardware, such as a RISC processor or programmable hardware with embedded firmware. Multiple processors may be included.

[0096] The system / device includes a working memory or volatile memory 504. The one or more processors can access the volatile memory 504 to process data and can control the storage of data in the memory. The volatile memory 504 can include any type of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), or can include flash memory (such as an SD card).

[0097] The system / device includes a non-volatile memory 506. The non-volatile memory 506 stores a set of operating instructions 508 in the form of computer-readable instructions for controlling the operation of the processor 502. The non-volatile memory 506 can be any type of memory, such as read-only memory (ROM), flash memory, or magnetic drive memory.

[0098] The one or more processors 502 are configured to execute operating instructions 508 to cause the system / device to perform any of the methods described herein. The operating instructions 508 may include code associated with the hardware components of the system / device 500 (i.e., drivers), as well as code associated with the basic operation of the system / device 500. Generally speaking, the one or more processors 502 use the volatile memory 504 to execute one or more instructions in the operating instructions 508 (these instructions are permanently or semi-permanently stored in the non-volatile memory 506) to temporarily store data generated during the execution of the operating instructions 508.

[0099] The methods described herein can be implemented as digital electronic circuitry, integrated circuits, specially designed application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These can include computer program products (e.g., software stored on a disk, optical disk, memory, programmable logic device, etc.) comprising computer-readable instructions that, when executed by a computer, for example, in combination with Figure 5 The computer is caused to execute one or more methods described herein.

[0100] Any system features described herein may also be provided as method features, and vice versa. Alternatively, the apparatus and functional features used herein may be represented according to their corresponding structures. Specifically, method aspects may be applied to system aspects, and vice versa.

[0101] In addition, any, some and / or all features of one aspect may be applied to any, some and / or all features of any other aspect in any appropriate combination. It should also be understood that the specific combination of various features described and defined in any aspect of the present invention may be independently implemented and / or provided and / or used.

[0102] While several embodiments have been shown and described, it will be appreciated by those skilled in the art that changes may be made to these embodiments without departing from the principles of the invention, the scope of which is defined in the claims.

Claims

1. A computer-implemented method for generating a three-dimensional face rendering from a two-dimensional image containing a face image, characterized in that: The method comprises: generating a three-dimensional shape model of the facial image and a low-resolution two-dimensional texture map of the facial image based on the two-dimensional image using one or more fitting neural networks, wherein the fitting neural networks are generative adversarial networks; Applying a super-resolution model to the low-resolution two-dimensional texture map to generate a high-resolution two-dimensional texture map; generating a two-dimensional diffuse albedo map from the high-resolution texture map using a deshaded image-to-image conversion neural network; A high-resolution three-dimensional model of the facial image is rendered using the two-dimensional diffuse albedo map and the three-dimensional shape model.

2. The method according to claim 1, characterized in that The two-dimensional diffuse albedo map is a high-resolution two-dimensional diffuse albedo map.

3. The method according to claim 1 or 2, characterized in that The method further comprises: determining a two-dimensional normal map of the face image based on the three-dimensional shape model, Wherein, additionally, the two-dimensional diffuse albedo map is generated using the two-dimensional normal map.

4. The method according to claim 1 or 2, characterized in that The method further comprises: generating a two-dimensional specular albedo map from the two-dimensional diffuse albedo map using a specular albedo image-to-image conversion neural network, Wherein, the high-resolution three-dimensional model of the facial image is also rendered based on the two-dimensional specular reflection albedo map.

5. The method according to claim 4, characterized in that The method further comprises: generating a grayscale two-dimensional diffuse albedo map according to the two-dimensional diffuse albedo map; The grayscale two-dimensional diffuse albedo map is input into the specular albedo image into an image conversion neural network.

6. The method according to claim 4, characterized in that The method further comprises: determining a two-dimensional normal map of the face image based on the three-dimensional shape model, Wherein, additionally, the two-dimensional specular albedo map is generated from the two-dimensional normal map using the specular albedo image-to-image conversion neural network.

7. The method according to claim 1 or 2, characterized in that The method further comprises: determining a two-dimensional normal map of the facial image based on the three-dimensional shape model; generating a two-dimensional specular normal map from the two-dimensional diffuse albedo map and the two-dimensional normal map using a specular normal image-to-image conversion neural network, Wherein, the high-resolution three-dimensional model of the facial image is also rendered based on the two-dimensional specular reflection normal map.

8. The method according to claim 7, characterized in that Generating the two-dimensional specular normal map using the specular normal image to image conversion neural network includes: generating a grayscale two-dimensional diffuse albedo map according to the two-dimensional diffuse albedo map; The grayscale two-dimensional diffuse albedo map and the two-dimensional normal map are input into the specular normal image into an image conversion neural network.

9. The method according to claim 3, characterized in that The two-dimensional normal map is a two-dimensional normal map in tangent space.

10. The method according to claim 1 or 2, characterized in that The method further comprises: determining a two-dimensional normal map in object space of the facial image based on the three-dimensional shape model; generating a two-dimensional diffuse normal map from the two-dimensional diffuse albedo map and the two-dimensional normal map in tangent space using a diffuse normal image-to-image conversion neural network, Wherein, the high-resolution three-dimensional model of the facial image is also rendered based on the two-dimensional diffuse reflection normal map.

11. The method according to claim 10, characterized in that Generating a 2D diffuse normal map using a diffuse normal image-to-image conversion neural network involves: generating a grayscale two-dimensional diffuse albedo map according to the two-dimensional diffuse albedo map; The grayscale two-dimensional diffuse albedo map and the two-dimensional normal map in the tangent space are input into the diffuse normal image into an image conversion neural network.

12. The method according to claim 1 or 2, characterized in that The method further comprises, for each image-to-image translation neural network: Divide the input 2D map into multiple overlapping input blocks; generating an output block for each of the input blocks using the image-to-image conversion neural network; A complete output two-dimensional map is generated by combining the plurality of output blocks.

13. The method according to claim 12, characterized in that The image-to-image translation neural network is a generative adversarial network.

14. The method according to claim 1 or 2, characterized in that The method also includes generating a three-dimensional model of a head from the high-resolution three-dimensional model of the facial image using a combined face and head model.

15. The method according to claim 12, characterized in that One or more of the two-dimensional maps include a UV map.

16. A system comprising one or more processors and a memory, characterized in that: The memory includes computer readable instructions which, when executed by the one or more processors, cause the system to perform a method according to any one of the preceding claims.

17. A computer program product comprising computer-readable instructions, characterized in that When the computer readable instructions are executed by a computing system, the computing system is caused to perform the method according to any one of claims 1 to 15.