3D Facial Scan Enhancement

Through the joint training of generator and discriminator neural network, the problem of converting low-quality three-dimensional facial data into high-quality data is solved, and the effect of improving the quality of three-dimensional facial data is achieved.

CN113454678BActive Publication Date: 2025-06-10HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080015378.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-06
Filing Date
2020-03-05
Publication Date
2025-06-10
Estimated Expiration
2040-03-05

AI Technical Summary

Technical Problem

The prior art is difficult to effectively perform shape-to-shape conversion of three-dimensional facial data, especially in the case of low-quality depth camera output, especially in nonlinear 3D facial data.

Method used

Convert low-quality 3D facial scans to high-quality 3D facial scans by jointly training generator neural networks and discriminator neural networks. The specific method includes applying the generator neural network to low-quality spatial UV maps, generating candidate high-quality spatial UV maps, and reconstructing and optimizing them through the discriminator neural network, and finally updating the parameters of the generator and discriminator.

Benefits of technology

A method of converting low-quality 3D facial scanning into high-quality scanning is realized, improving the quality of 3D facial data, especially in nonlinear 3D facial data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113454678B_ABST
    Figure CN113454678B_ABST
Patent Text Reader

Abstract

This specification describes methods for enhancing 3D facial data using neural networks, as well as methods for training neural networks to enhance 3D facial data. According to a first aspect of the present invention, a method for training a generator neural network to convert low-quality 3D facial scans into high-quality 3D facial scans is described, the method comprising: applying the generator neural network to a low-quality spatial UV map to generate a candidate high-quality spatial UV map; applying a discriminator neural network to the candidate high-quality spatial UV map to generate a reconstructed candidate high-quality spatial UV map; applying the discriminator neural network to a high-quality ground truth spatial UV map to generate a reconstructed high-quality ground truth spatial UV map, wherein the high-quality ground truth spatial UV map corresponds to the low-quality spatial UV map; updating the parameters of the generator neural network based on a comparison of the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map; and updating the parameters of the discriminator neural network based on a comparison of the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map and a comparison of the high-quality ground truth spatial UV map and the reconstructed high-quality ground truth spatial UV map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification describes methods for enhancing three-dimensional facial data using neural networks, and methods for training neural networks to enhance three-dimensional facial data. Background Art

[0002] Image-to-image translation is a prevalent problem in image processing, in which an input image is translated into a synthetic image that preserves certain attributes of the original input image. Examples of image-to-image translation include converting an image from black and white to color, or converting a daytime scene to a nighttime scene, thereby improving the image quality and / or processing the facial attributes of the image. However, current methods for performing image-to-image translation are limited to two-dimension (2D) texture images.

[0003] With the introduction of depth cameras, the capture and use of three-dimension (3D) image data has become increasingly common. However, the use of shape-to-shape translation (the 3D analogue of image-to-image translation) on such 3D image data is limited by several factors, including the low-quality output of many depth cameras. This is especially true for 3D facial data where non-linearity often exists. Summary of the Invention

[0004] According to a first aspect of the present invention, a method for training a generator neural network to convert low-quality three-dimensional facial scans into high-quality three-dimensional facial scans is described. The method includes jointly training a discriminator neural network and a generator neural network. The joint training includes: applying the generator neural network to a low-quality spatial UV map to generate a candidate high-quality spatial UV map; applying the discriminator neural network to the candidate high-quality spatial UV map to generate a reconstructed candidate high-quality spatial UV map; applying the discriminator neural network to a high-quality ground truth spatial UV map to generate a reconstructed high-quality ground truth spatial UV map, where the high-quality ground truth spatial UV map corresponds to the low-quality spatial UV map; updating the parameters of the generator neural network based on a comparison between the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map; and updating the parameters of the discriminator neural network based on a comparison between the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map, and a comparison between the high-quality ground truth spatial UV map and the reconstructed high-quality ground truth spatial UV map. When updating the parameters, a comparison between the candidate high-quality spatial UV map and the corresponding ground truth high-quality spatial UV map may also be used.

[0005] The generator neural network and / or the discriminator neural network may include a set of encoding layers and a set of decoding layers, wherein the encoding layers are used to convert an input space UV map into an embedding, and the decoding layers are used to convert the embedding into an output space UV map. During the joint training of the generator neural network and the discriminator neural network, the parameters of one or more of the decoding layers may be fixed. The decoding layers of the generator neural network and / or the discriminator neural network may include one or more skip connections in the initial layer of the decoding layers.

[0006] The generator neural network and / or the discriminator neural network may include a plurality of convolutional layers. The generator neural network and / or the discriminator neural network may include one or more fully connected layers. The generator neural network and / or the discriminator neural network may include one or more upsampling layers and / or subsampling layers. The network structures of the generator neural network and / or the discriminator neural network may be the same.

[0007] Updating the parameters of the generator neural network may also be based on the comparison between the candidate high-quality space UV map and the corresponding high-quality ground truth space UV map.

[0008] Updating the parameters of the generator neural network may include: calculating a generator loss using a generator loss function based on the difference between the candidate high-quality space UV map and the corresponding reconstructed candidate high-quality space UV map; applying an optimization process to the generator neural network to update the parameters of the generator neural network according to the calculated generator loss. The generator loss function may also calculate the generator loss based on the difference between the candidate high-quality space UV map and the corresponding high-quality ground truth space UV map.

[0009] Updating the parameters of the discriminator neural network may include: calculating a discriminator loss using a discriminator loss function based on the difference between the candidate high-quality space UV map and the reconstructed candidate high-quality space UV map and the difference between the high-quality ground truth space UV map and the reconstructed high-quality ground truth space UV map; applying an optimization process to the discriminator neural network to update the parameters of the discriminator neural network according to the calculated discriminator loss.

[0010] The method may further include pre-training the discriminator neural network to reconstruct a high-quality ground truth space UV map based on an input high-quality ground truth space UV map.

[0011] According to another aspect of the present invention, a method for converting a low-quality 3D face scan into a high-quality 3D face scan is described, the method comprising: receiving a low-quality spatial UV map of the face scan; applying a neural network to the low-quality spatial UV map; outputting from the neural network a high-quality spatial UV map of the face scan, wherein the neural network is a generator neural network trained using any one of the training methods described herein.

[0012] According to another aspect of the present invention, an apparatus is described, comprising: one or more processors; a memory, wherein the memory includes computer-readable instructions that, when executed by the one or more processors, cause the apparatus to perform one or more of the methods described herein.

[0013] According to another aspect of the present invention, a computer program product comprising computer-readable instructions is described, which, when executed by a computer, cause the computer to perform one or more of the methods described herein.

[0014] The term "quality" as used herein may preferably be used to represent any one or more of the following: noise level (e.g., peak signal-to-noise ratio); texture quality; error relative to a ground truth scan; 3D shape quality (e.g., may refer to the degree of retention of high-frequency details such as eyelid and / or lip variations in 3D face data). BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Embodiments are now described by way of non-limiting examples with reference to the accompanying drawings, wherein:

[0016] Figure 1 An overview of an exemplary method for enhancing 3D face data using a neural network is shown;

[0017] Figure 2 An overview of an exemplary method for training a neural network to enhance 3D face data is shown;

[0018] Figure 3 A flowchart of an exemplary method for training a neural network to enhance 3D face data is shown;

[0019] Figure 4 An overview of an exemplary method for preprocessing 3D face data is shown;

[0020] Figure 5 An overview of an exemplary method for pre-training a discriminator neural network is shown;

[0021] Figure 6 An example of the structure of a neural network for enhancing 3D face data is shown;

[0022] Figure 7Shows an example schematic diagram of a computing system. Detailed implementation

[0023] The original 3D face scans captured by some 3D camera systems are usually of low quality. For example, they have less surface detail and / or are noisier. For example, this may be the result of the method used by the camera to capture the 3D face scan, or due to the technical limitations of the 3D camera system. However, applications that use face scans may require a higher quality of scan than the face scans captured by the 3D camera system.

[0024] Figure 1 Shows an overview of an exemplary method for enhancing 3D face data using a neural network. The method also includes receiving low-quality 3D face data 102 and using a neural network 106 to generate high-quality 3D face data 104 based on the low-quality 3D face data 102.

[0025] The low-quality 3D face data 102 may include a UV map of a low-quality 3D face scan. Alternatively, the low-quality 3D face data 102 may include a 3D mesh representing a low-quality 3D face scan. In a preprocessing step 108, the 3D mesh can be converted into a UV map. Examples of such preprocessing steps are described below in conjunction with Figure 4 Describe examples of such preprocessing steps.

[0026] A spatial UV map is a two-dimensional representation of a 3D surface or mesh. Each point in 3D space (e.g., described by coordinates (x, y, z)) is mapped onto a two-dimensional space (described by coordinates (u, v)). The UV map can be formed by unfolding a 3D mesh in 3D space onto the u-v plane in two-dimensional UV space. In some embodiments, the coordinates (x, y, z) of the 3D mesh in 3D space are stored as the RGB values of the corresponding points in UV space. Using a spatial UV map can facilitate the use of two-dimensional convolution when improving the quality of 3D scans, rather than using geometric deep learning methods, which tend to mainly retain the low-frequency details of the 3D mesh.

[0027] The neural network 106 includes multiple layers of nodes, and each node is associated with one or more parameters. The parameters of each node in the neural network may include one or more weights and / or biases. A node takes one or more outputs of the nodes in the previous layer as inputs. One or more outputs of the nodes in the previous layer are used by the node to generate activation values through an activation function and the parameters of the neural network.

[0028] The neural network 106 can have an autoencoder architecture. Examples of the neural network architecture are described below in conjunction with Figure 6 Describe various examples of the neural network architecture.

[0029] The parameters of the neural network 106 can be trained using generative adversarial training, and the neural network 106 can thus be referred to as a Generative Adversarial Network (GAN). The neural network 106 can be a generative network for generative adversarial training. The following describes various examples of the training method in conjunction with Figures 3 to 5 each example of the training method.

[0030] The neural network uses the UV map of the low-quality 3D face scan to generate high-quality 3D face data 104. The high-quality 3D face data 104 can include a high-quality UV map. The high-quality UV map can be converted into a high-quality 3D spatial grid in the post-processing step 110.

[0031] Figure 2 An overview of an exemplary method for training a neural network to enhance 3D face data is shown. Method 200 includes jointly training a generator neural network 202 and a discriminator neural network 204 in an adversarial manner. During training, the purpose of the generator neural network 202 is to learn to generate high-quality UV face maps 206 based on the input low-quality UV face maps (also referred to herein as low-quality spatial UV maps) 208, which are close to the corresponding ground truth UV face maps (also referred to herein as real high-quality UV face maps and / or high-quality ground truth spatial UV maps) 210. The pair set {(x, y)} of the low-quality spatial UV map x and the high-quality ground truth spatial UV map y can be referred to as the training set / data. A training data set can be constructed from the original face scan using a preprocessing method, as described in more detail below in conjunction with Figure 4 more detailed description.

[0032] During training, the purpose of the discriminator neural network 204 is to learn to distinguish between the ground truth UV face map 210 and the generated high-quality UV face map 206 (also referred to herein as a fake high-quality UV face map or a candidate high-quality spatial UV map). The discriminator neural network 204 can have an autoencoder structure.

[0033] In some embodiments, the discriminator neural network 204 can be pre-trained on pre-training data, as described below in conjunction with Figure 5 described. In embodiments where the structures of the discriminator neural network 204 and the generator neural network 202 are the same, the parameters of the pre-trained discriminator neural network 204 can be used to initialize both the discriminator neural network 204 and the generator neural network 202.

[0034] During the training process, the generator neural network 202 and the discriminator neural network 204 compete with each other until they reach a threshold / equilibrium condition. For example, the generator neural network 202 and the discriminator neural network 204 compete with each other until the discriminator neural network 204 can no longer distinguish between real and fake UV face maps.

[0035] During the training process, the generator neural network 202 is applied to the low-quality spatial UV map 208, x obtained from the training data. The output of the generator neural network is the corresponding candidate high-quality spatial UV map 206, G(x).

[0036] The discriminator neural network 204 is applied to the candidate high-quality spatial UV map 206 to generate the reconstructed candidate high-quality spatial UV map 212, D(G(x)). The discriminator neural network 204 is also applied to the high-quality ground truth spatial UV map 210, y corresponding to the low-quality spatial UV map 208, x to generate the reconstructed high-quality ground truth spatial UV map 214, D(y).

[0037] The candidate high-quality spatial UV map 206, G(x) and the reconstructed candidate high-quality spatial UV map 212, D(G(x)) are compared, and the parameters of the generator neural network are updated using the comparison result. The high-quality ground truth spatial UV map 210, y and the reconstructed high-quality ground truth spatial UV map 214, D(y) can also be compared, and the candidate high-quality spatial UV map 206 and the reconstructed candidate high-quality spatial UV map 212 are compared. The two comparison results are used together to update the parameters of the discriminator neural network. One or more loss functions can be used to perform the comparison. In some embodiments, a loss function is calculated using the results of applying the generator neural network 202 and the discriminator neural network 204 to multiple pairs of low-quality spatial UV maps 208 and high-quality ground truth spatial UV maps 210.

[0038] In some embodiments, an adversarial loss function can be used. An example of the adversarial loss is the BEGAN loss. The loss function of the generator neural network ( also referred to as the generator loss in this document) and the loss function of the discriminator neural network ( also referred to as the discriminator loss in this document) can be given by the following formula:

[0039]

[0040]

[0041]

[0042] where t denotes the update iteration (e.g., for the first update of the network, t = 0, for the second set of updates of the network, t = 1), represents a metric for comparing the input z of the discriminator with the corresponding output D(z), k t is a parameter that controls how much weight should be imposed on £(G(x)), λ k is k tThe learning rate, γ ∈ [0, 1], is a hyperparameter that controls the balance. The hyperparameters can take values γ = 0.5 and λ = 10, and the value of k t is initialized to 0.001. However, other values can also be used. In some embodiments, the metric £(z) is given by although it will be understood that other examples are possible. represents the expected value of the training data set.

[0043] During training, the discriminator neural network 204 is trained to minimize while the generator neural network 202 is trained to minimize Effectively, the generator neural network 202 is trained to “fool” the discriminator neural network 204.

[0044] In some embodiments, updating the parameters of the generator neural network can also be based on a comparison between the candidate high-quality spatial UV map 206 and the high-quality ground truth spatial UV map 210. The comparison can be performed using an additional term in the generator loss, which is referred to herein as the reconstruction loss. Then, the total generator loss can be given by the following formula:

[0045]

[0046] where λ is a hyperparameter that controls the degree of emphasis on the reconstruction loss. An example of the reconstruction loss is It will be understood that other examples are also possible.

[0047] The comparison can be used to update the parameters of the generator and / or discriminator neural networks using an optimization process / method aimed at minimizing the above loss function. An example of such a method is the gradient descent algorithm. A characteristic of the optimization method can be the learning rate, which characterizes the “size” of the step used during each iteration of the algorithm. In some embodiments using gradient descent, the learning rate can initially be set to 5e(-5) for both the generator neural network and the discriminator neural network.

[0048] During training, the learning rate of the training process can be changed after a threshold number of epochs and / or iterations. After every N iterations, the learning rate may be decreased by a given factor. For example, after every 30 training epochs, the learning rate may be decreased by 5%.

[0049] Different learning rates can be used for different layers of neural networks 202, 204. For example, in an embodiment where the discriminator neural network 204 has been pre-trained, one or more layers of the discriminator neural network 204 and / or the generator neural network 202 can be frozen during training (i.e., the learning rate is 0). The decoder layers of the discriminator neural network 204 and / or the generator neural network 202 can be frozen during training. The encoder and bottleneck portions of the neural networks 202, 204 can have small learning rates to prevent their values from deviating significantly from the values found during pre-training. These learning rates can reduce the training time and improve the accuracy of the trained generator neural network 106.

[0050] The training process can be iterative until a threshold condition is met. For example, the threshold condition can be a threshold number of iterations and / or epochs. For example, training can be performed for 300 epochs. Optionally or additionally, the threshold condition can be that the loss functions are each optimized within a threshold of their minimum values.

[0051] Figure 3 Flowchart of an exemplary method for training a neural network to convert a low-quality 3D face scan into a high-quality 3D face scan. The flowchart corresponds to the method described above in connection with Figure 2 the method described.

[0052] In operation 3.1, the generator neural network is applied to the low-quality spatial UV map to generate a candidate high-quality spatial UV map. The generator neural network can have an autoencoder structure and includes a set of encoder layers and a set of decoder layers, where the encoder layers are used to generate an embedding of the low-quality spatial UV map, and the decoder layers are used to generate a candidate high-quality spatial UV map based on the embedding. The generator neural network is described by a set of generator neural network parameters (e.g., the weights and biases of the neural network nodes in the generator neural network).

[0053] In operation 3.2, the discriminator neural network is applied to the candidate high-quality spatial UV map to generate a reconstructed candidate high-quality spatial UV map. The discriminator neural network can have an autoencoder structure and includes a set of encoder layers and a set of decoder layers, where the encoder layers are used to generate an embedding of the input spatial UV map, and the decoder layers are used to generate an output high-quality spatial UV map based on the embedding. The discriminator neural network is described by a set of discriminator neural network parameters (e.g., the weights and biases of the neural network nodes in the discriminator neural network).

[0054] In operation 3.3, the discriminator neural network is applied to a high-quality ground truth spatial UV map to generate a reconstructed high-quality ground truth spatial UV map, where the high-quality ground truth spatial UV map corresponds to a low-quality spatial UV map. The high-quality ground truth spatial UV map and the low-quality spatial UV map can be a training pair from a training dataset, both representing the same object but captured at different qualities (e.g., captured by different 3D camera systems).

[0055] In operation 3.4, the parameters of the generator neural network are updated based on the comparison between the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map. The comparison can be performed by a generator loss function. An optimization process, such as gradient descent, can be applied to the loss function to determine the update to the generator neural network parameters. When updating the parameters of the generator neural network, the comparison between the candidate high-quality spatial UV map and the corresponding ground truth high-quality spatial UV map can also be used.

[0056] In operation 3.5, the parameters of the discriminator neural network are updated based on the comparison between the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map and the comparison between the high-quality ground truth spatial UV map and the reconstructed high-quality ground truth spatial UV map. The above comparisons can be performed by a discriminator loss function. An optimization process, such as gradient descent, can be applied to the loss function to determine the update to the discriminator neural network parameters.

[0057] Operations 3.1 to 3.5 can be iterative until a threshold condition is met. During each iteration, different spatial UV maps from the training dataset can be used.

[0058] Figure 4 An overview of an exemplary method for preprocessing 3D facial data is shown. In some embodiments, the raw 3D facial data cannot be directly processed by the neural networks described herein. The preprocessing step 400 is used to convert this raw 3D data into a UV facial map that can be processed by the generator and / or discriminator neural networks. The following description will describe the preprocessing method from the perspective of preprocessing training data, but it is obvious that the elements of this method can equally be applied to methods for enhancing 3D facial data, e.g., as described above in conjunction with Figure 1 described.

[0059] Before training the neural network, high-quality raw scans (y r ) and low-quality raw scans (x r)402 pairs. The corresponding mesh pairs depict the same subject but have different structures (e.g., topologies) in terms of the number of vertices and triangulation. Note that the number of vertices in a high-quality original scan is not necessarily greater than that in a low-quality original scan; correctly representing facial features is the key to determining the overall scan quality. For example, some scanners that generate scans with a large number of vertices may use methods that cause unnecessary points to overlap each other, thus generating complex graphics with less surface detail. In some embodiments, high-quality original scans (y r ) are also preprocessed in this way. 3D

[0060] high-quality original scans (y r ) and low-quality original scans (x r )(402) are each mapped to a template 404 (T) that describes both scans in the same topology. An example of such templates is the LSFM model. The template includes multiple vertices sufficient to depict high-level facial details (in the example of the LSFM model, there are 54,000 vertices).

[0061] During training, the original scans of high-quality original scans (y r ) and low-quality original scans (x r ) 402 have a correspondence by non-rigidly deforming the template mesh 404 into each original scan. The non-rigid deformation of the template mesh can be performed using algorithms such as the Optimal Step Non-rigid Iterative Closest Point (NICP) algorithm. For example, vertices can be weighted according to the Euclidean distance measured from a given feature (e.g., the tip of the nose) in the facial scan. For example, the greater the distance from the tip of the nose to a given vertex, the greater the weight assigned to that vertex. This can help remove the noisy information recorded in the facial scan within the outer regions of the original scan.

[0062] Then, the mesh of the facial scan is converted into a sparse spatial UV map 406. UV maps are typically used to store texture information. In this method, the spatial position of each vertex of the mesh is represented as an RBG value in the UV space. The mesh will be unfolded into the UV space to obtain the UV coordinates of the mesh vertices. For example, the best cylindrical unfolding technique can be used to unfold the mesh.

[0063] In some embodiments, before storing the 3D coordinates in the UV space, the mesh is aligned by performing General Procrustes Analysis (GPA). The mesh can also be normalized to the [-1,1] scale.

[0064] Then, the sparse spatial UV map 406 is converted to an interpolated UV map 408 with a greater number of vertices. Two-dimensional interpolation can be used in the UV domain to fill in the missing regions, resulting in a densified representation of the originally sparse UV map 406. Examples of such interpolation methods include two-dimensional nearest-point interpolation or barycentric interpolation.

[0065] In embodiments where the number of vertices is greater than 50,000, the UV map size can be selected as 256×256×3, which can help retrieve a high-precision point cloud with negligible resampling error.

[0066] Figure 5 An overview of an exemplary method of a pre-trained discriminator neural network 500 is shown. In some embodiments, the discriminator neural network 204 is pre-trained before adversarial training with the generator neural network 202. Pre-training the discriminator neural network 204 can reduce the occurrence of mode collapse in the generative adversarial training.

[0067] The discriminator neural network 204 is pre-trained on high-quality real facial UV maps 502. The real high-quality spatial UV map 502 is input into the discriminator neural network 204, and the discriminator neural network 204 generates an embedding of the real high-quality spatial UV map 502 and generates a reconstructed real high-quality spatial UV map 504 based on the embedding. Based on the comparison between the real high-quality spatial UV map 502 and the reconstructed real high-quality spatial UV map 504, the parameters of the discriminator neural network 204 are updated. A discriminator loss function 506 can be used to compare the real high-quality spatial UV map 502 and the reconstructed real high-quality spatial UV map 504. For example,

[0068] The data on which the discriminator neural network 204 is pre-trained (i.e., the pre-training data) can be different from the training data used in the above-mentioned adversarial training. For example, the batch size used during pre-training can be 16.

[0069] Pre-training can be performed until a threshold condition is met. The threshold condition can be a threshold number of training epochs. For example, pre-training can be performed for 300 epochs. The learning rate may change after a number of epochs less than the threshold, for example, every 30 epochs.

[0070] The initial parameters of the discriminator neural network 204 and the generator neural network 202 can be selected based on the parameters of the pre-trained discriminator neural network.

[0071] Figure 6 An example of the structure of a neural network for enhancing 3D facial data is shown. Such a neural network architecture can be used for the discriminator neural network 204 and / or the generator neural network 202.

[0072] In this example, the neural network 106 is in the form of an autoencoder. The neural network includes a set of encoder layers 600 for generating an embedding 602 based on the input UV map 604 of a face scan. The neural network also includes a set of decoder layers 608 for generating an output UV map 610 of the face scan based on the embedding 602.

[0073] Each of the encoder layer 600 and the decoder layer 608 includes a plurality of convolutional layers 612. Each convolutional layer 612 can be used to apply one or more convolutional filters to the input of the convolutional layer 612. For example, one or more convolutional layers 612 can apply a two-dimensional convolutional block with a kernel size of 3, a stride of 1, and a padding size of 1. However, other kernel sizes, strides, and padding sizes can be selected or alternatively used. In the example shown, there are a total of 12 convolutional layers 612 in the encoding layer 600 and a total of 13 convolutional layers 612 in the decoding layer 608. Other numbers of convolutional layers 612 can also be used.

[0074] Interleaved with the convolutional layers 612 of the encoder layer 600 are a plurality of subsampling layers 614 (also referred to herein as downsampling layers). One or more convolutional layers 612 can be located between the respective subsampling layers 614. In the example shown, two convolutional layers 612 are placed between each subsampling layer 614. Each subsampling layer 614 can be used to reduce the size of the input to that subsampling layer. For example, one or more subsampling layers can apply average two-dimensional pooling with a kernel size and a stride size of 2. However, other subsampling methods and / or subsampling parameters can be selected or alternatively used.

[0075] One or more fully connected layers 616 can also be present in the encoder layer 600, for example, as the last layer of the encoder layer that outputs the embedding 602 (i.e., at the bottleneck of the autoencoder). The fully connected layer 616 projects the input tensor onto eigenvectors or vice versa.

[0076] By performing a series of convolutional and subsampling operations and then generating an embedding 602 of size h (the bottleneck size is h) by the fully connected layer 616, the encoder layer 600 acts on the input UV map 604 of the face scan (in this example, including a 256×256×3 tensor, i.e., 256×256 RGB values, although other sizes are also possible). In the example shown, h is equal to 128.

[0077] Interleaved with the convolutional layer 612 of the decoder layer 608 are a plurality of upsampling layers 618. One or more convolutional layers 612 may be located between the respective upsampling layers 618. In the example shown, two convolutional layers 612 are applied between the upsampling layers 618. Each upsampling layer 618 is used to increase the input dimension of that upsampling layer. For example, one or more of the upsampling layers in the upsampling layer 618 may apply the nearest neighbor method with a scale factor of 2. However, other upsampling methods and / or upsampling parameters (e.g., scale factor) may be selected or alternatively used.

[0078] One or more fully connected layers 616 may also be present in the decoder layer 608, such as the initial layer of the encoder layer that takes the embedding 602 as input (i.e., at the bottleneck of the autoencoder).

[0079] The decoder layer 608 may also include one or more skip connections 620. The skip connection 620 injects the output / input of a given layer into the input of a subsequent layer. In the example shown, the skip connection injects the output of the initial fully connected layer 616 into the first upsampling layer 618a and the second upsampling layer 618b. When using the output UV map 602 of the neural network 106, more compelling visual results can be produced.

[0080] One or more activation functions are used in the various layers of the neural network 106. For example, the ELU activation function may be used. Additionally or alternatively, the Tanh activation function may be used in one or more layers. In some embodiments, the last layer of the neural network may use the Tanh activation function. Additionally or alternatively, other activation functions may be used.

[0081] Figure 7 A schematic example of a system / apparatus for performing any of the methods described herein is shown. The system / apparatus shown is an example of a computing device. Those skilled in the art will understand that other types of computing devices / systems may alternatively be used to implement the methods described herein, such as a distributed computing system.

[0082] The apparatus (or system) 700 includes one or more processors 702. The one or more processors control the operation of the other components of the system / apparatus 700. The one or more processors 702 may include, for example, general-purpose processors. The one or more processors 702 may be a single-core device or a multi-core device. The one or more processors 702 may include a central processing unit (CPU) or a graphical processing unit (GPU). Alternatively, the one or more processors 702 may include dedicated processing hardware, such as a RISC processor or programmable hardware with embedded firmware. Multiple processors may be included.

[0083] The system / apparatus includes a working or volatile memory 704. One or more processors can access the volatile memory 704 to process data and can control the storage of data in the memory. The volatile memory 704 can include any type of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), or can include flash memory, such as an SD card.

[0084] The system / apparatus includes a non-volatile memory 706. The non-volatile memory 706 stores a set of operation instructions 708 for controlling the operation of the processor 702 in the form of computer-readable instructions. The non-volatile memory 706 can be any type of memory, such as read only memory (ROM), flash memory, or magnetic drive memory.

[0085] One or more processors 702 are used to execute the operation instructions 408 to cause the system / apparatus to perform any of the methods described herein. The operation instructions 708 can include code related to the hardware components of the system / apparatus 700 (i.e., drivers), as well as code related to the basic operation of the system / apparatus 700. Generally, one or more processors 702 use the volatile memory 704 to temporarily store data generated during the execution of the operation instructions 708, thereby executing one or more instructions of the operation instructions 708 permanently or semi-permanently stored in the non-volatile memory 706.

[0086] The implementation of the methods described herein can be implemented in digital electronic circuits, integrated circuits, application specific integrated circuits (ASICs) specifically designed, computer hardware, firmware, software, and / or combinations thereof, which can include computer program products (e.g., software stored on, for example, a disk, an optical disc, a memory, a programmable logic device), including computer-readable instructions that, when executed by a computer (e.g., in combination with Figure 7 the computer described), cause the computer to perform one or more of the methods described herein.

[0087] Any system features described herein can also be provided as method features, and vice versa. As used herein, means-plus-function features can be expressed in accordance with their corresponding structures. Specifically, method aspects can be applied to system aspects, and vice versa.

[0088] In addition, any, some, and / or all features in one aspect can be applied to any, some, and / or all features in any other aspect in any suitable combination. It should also be understood that specific combinations of the various features described and defined in any aspect of the present invention can be implemented and / or provided and / or used independently.

[0089] Although several embodiments have been shown and described, those skilled in the art will understand that changes can be made in these embodiments without departing from the principles of the disclosure, the scope of which is defined in the claims.

Claims

1. A method for training a generator neural network to convert a low-quality 3D face scan into a high-quality 3D face scan, the method comprising jointly training a discriminator neural network and a generator neural network, characterized in that, the joint training includes: applying the generator neural network to a low-quality spatial UV map to generate a candidate high-quality spatial UV map; applying the discriminator neural network to the candidate high-quality spatial UV map to generate a reconstructed candidate high-quality spatial UV map; applying the discriminator neural network to a high-quality ground truth spatial UV map to generate a reconstructed high-quality ground truth spatial UV map, wherein the high-quality ground truth spatial UV map corresponds to the low-quality spatial UV map; updating the parameters of the generator neural network according to a comparison between the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map; updating the parameters of the discriminator neural network according to a comparison between the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map and a comparison between the high-quality ground truth spatial UV map and the reconstructed high-quality ground truth spatial UV map; wherein updating the parameters of the generator neural network is also based on a comparison between the candidate high-quality spatial UV map and the corresponding high-quality ground truth spatial UV map.

2. The method according to claim 1, characterized in that, the generator neural network and / or the discriminator neural network includes a set of encoding layers and a set of decoding layers, wherein the encoding layers are used to convert an input spatial UV map into an embedding, and the decoding layers are used to convert the embedding into an output spatial UV map.

3. The method according to claim 2, characterized in that, during the joint training of the generator neural network and the discriminator neural network, the parameters of one or more of the decoding layers are fixed.

4. The method according to claim 2 or 3, characterized in that, the decoding layer of the generator neural network and / or the discriminator neural network includes one or more skip connections in the initial layer of the decoding layer.

5. The method according to any one of claims 1 to 3, characterized in that, the generator neural network and / or the discriminator neural network includes a plurality of convolutional layers.

6. The method according to any one of claims 1 to 3, characterized in that, the generator neural network and / or the discriminator neural network includes one or more fully connected layers.

7. The method according to any one of claims 1 to 3, characterized in that, the generator neural network and / or the discriminator neural network includes one or more upsampling layers and / or subsampling layers.

8. The method according to any one of claims 1 to 3, characterized in that, the network structures of the generator neural network and / or the discriminator neural network are the same.

9. The method according to any one of claims 1 to 3, characterized in that, updating the parameters of the generator neural network includes: Calculate a generator loss using a generator loss function based on the difference between the candidate high-quality spatial UV map and the corresponding reconstructed candidate high-quality spatial UV map; Apply an optimization process to the generator neural network to update the parameters of the generator neural network according to the calculated generator loss.

10. The method according to claim 9, wherein, the generator loss function further calculates the generator loss based on the difference between the candidate high-quality spatial UV map and the corresponding high-quality ground truth spatial UV map.

11. The method according to any one of claims 1 to 3, wherein, updating the parameters of the discriminator neural network includes: Calculate a discriminator loss using a discriminator loss function based on the difference between the candidate high-quality spatial UV map and the reconstructed candidate high-quality spatial UV map, and the difference between the high-quality ground truth spatial UV map and the reconstructed high-quality ground truth spatial UV map; Apply an optimization process to the discriminator neural network to update the parameters of the discriminator neural network according to the calculated discriminator loss.

12. The method according to any one of claims 1 to 3, wherein, further includes pre-training the discriminator neural network to reconstruct a high-quality ground truth spatial UV map based on an input high-quality ground truth spatial UV map.

13. A method for converting a low-quality 3D face scan into a high-quality 3D face scan, wherein, the method includes: Receiving a low-quality spatial UV map of a face scan; Applying a neural network to the low-quality spatial UV map; Outputting a high-quality spatial UV map of the face scan from the neural network, wherein, the neural network is a generator neural network trained using the method according to any one of claims 1 to 12.

14. An apparatus for training a generator neural network to convert a low-quality 3D face scan into a high-quality 3D face scan, wherein, includes: One or more processors; A memory, wherein, the memory includes computer-readable instructions that, when executed by the one or more processors, cause the apparatus to perform the method according to any one of claims 1 to 13.

15. A computer program product including computer-readable instructions, wherein, when executed by a computer, the computer-readable instructions cause the computer to perform the method according to any one of claims 1 to 13.