Facial image generation with wrinkles
By using the concept of grid tension to gather wrinkle textures in expression scanning, the problem of difficulty in generating realistic facial synthetic images in the prior art is solved, and efficient and fidelity wrinkle generation is achieved, which significantly improves the authenticity of facial synthetic images and the performance of downstream tasks.
Patent Information
- Application Number
- CN202380058125.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-01
- Filing Date
- 2023-08-21
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to efficiently generate realistic facial synthetic images, especially when generating expressions with wrinkles. The existing methods require great efforts from artists and are difficult to reconstruct high-frequency skin details with fidelity.
By using the concept of grid tension, possible wrinkles from high-quality expression scans are aggregated into wrinkle textures and using these textures to generate facial images with wrinkles, even for expressions not represented in the source data.
Realize the generation of realistic wrinkles on a large and diverse population of digital people, improve the authenticity and fidelity of facial synthetic images, and significantly improve the performance of downstream computer vision tasks.
Smart Images

Figure CN120051804A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Generating synthetic images of human and animal faces in an efficient manner and having the resulting images be realistic is extremely difficult to achieve. Synthetic images of human and animal faces are useful for a wide range of tasks such as video games, telepresence, movie production, augmented and virtual reality, machine learning, and computer vision.
[0002] The embodiments described below are not limited to implementations that solve any or all of the disadvantages of known face image generation processes. SUMMARY OF THE INVENTION
[0003] A simplified summary of the present disclosure is presented below in order to provide a basic understanding to the reader. This summary of the invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Its sole purpose is to present a selection of concepts disclosed herein in a simplified form as a prelude to the more detailed description that is presented later.
[0004] Generate a photon image of a face, where the image depicts a face with a wrinkled expression. By using a wrinkle texture calculated by aggregating maps of faces with different expressions, wrinkles can be generated for an expression even beyond what is represented in the input data.
[0005] In various examples, there is a method of computing an image depicting a face with a wrinkled expression. A 3D polygonal mesh model of the face has a non-neutral expression. A tension map is computed from the 3D polygonal mesh model. A neutral wrinkle texture, a compressed wrinkle texture, and an extended wrinkle texture are computed or obtained from a library. The neutral texture includes a map of a first face with a neutral expression. The compressed wrinkle texture is a map of the first face formed by aggregating maps of the first face with different expressions using the tension map, and the extended wrinkle texture includes a map of the first face formed by aggregating maps of the first face with different expressions using the tension map. The wrinkle texture is applied to the 3D model according to the tension map. An image is rendered from the 3D model.
[0006] Many attendant features will be more readily appreciated as they become better understood by reference to the following detailed description considered in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] This specification will be better understood from the following detailed description when read in conjunction with the accompanying drawings, in which:
[0008] Figure 1 is a schematic diagram of a wrinkled face synthesizer deployed in a communication network;
[0009] Figure 2Schematic diagrams of a synthetic image of a wrinkle - free face and the same synthetic image with wrinkles;
[0010] Figure 3 Flowchart of a method for generating a synthetic image of a face with wrinkles;
[0011] Figure 4 Flowchart of a method for calculating wrinkle textures of a human or an animal;
[0012] Figure 5 Schematic diagram of another method for calculating wrinkle textures of a human or an animal;
[0013] Figure 6 is Figure 5 Flowchart of the method of;
[0014] Figure 7 Flowchart of a method for training and using a face synthesizer; and
[0015] Figure 8 Shows an exemplary computing - based device in which an embodiment of a wrinkled face synthesizer is implemented.
[0016] In the drawings, the same reference numerals are used to denote the same components. Detailed Description
[0017] The detailed description provided below in conjunction with the accompanying drawings is intended as a description of this example and is not intended to represent the only form in which this example may be constructed or utilized. This specification sets forth the functions of the example and the sequence of operations for constructing and operating the example. However, the same or equivalent functions and sequences may be achieved by different examples.
[0018] As described above, it is extremely difficult to generate synthetic images of human and animal faces in an efficient manner and for the resulting images to be realistic. The inventors have discovered a way to enhance the realism of synthetic faces: by using wrinkle textures formed from empirical data and introducing dynamic skin wrinkles in response to facial expressions. Since the wrinkle textures are formed from empirical data, this technique is considered data - driven. As a result, significant performance improvements in downstream computer vision tasks such as facial landmark detection have been found.
[0019] Alternative methods for generating such wrinkles either require great effort from artists to span identities and expressions or cannot reconstruct high - frequency skin details with fidelity.
[0020] This technique produces realistic wrinkles on a large and diverse digital population. The inventors have formalized the concept of mesh tension and used it to cluster possible wrinkles from high - quality expression scans into wrinkle textures. To synthesize facial images, these wrinkle textures are used to generate wrinkles, even for expressions not represented in the source data.
[0021] A wrinkle texture is a map in UV space, where the map can be represented as a two-dimensional numerical array, and each numerical value is an albedo or a displacement. Albedo is a numerical value representing color, and displacement is a numerical value representing the displacement between the surface of the face depicted by the array element and the surface defined by the 3D mesh of the face. The wrinkle texture stores information from multiple maps in UV space depicting the same person or animal with different facial expressions. The UV space is unaware of the camera viewpoint.
[0022] Figure 1 FIG. 6 is a schematic diagram of a wrinkled face synthesizer 100 for calculating a synthetic image of a facially expressive face with wrinkles. In some cases, the wrinkled face synthesizer 100 is deployed as a web service in a communication network 124. In some cases, the wrinkled face synthesizer 100 is deployed at a personal computer or other computing device in communication with a head-mounted computer 114, such as a head-mounted display device. In some cases, the wrinkled face synthesizer 100 is deployed in an accompanying computing device of the head-mounted computer 114.
[0023] The wrinkled face synthesizer 100 includes a wrinkle texture generator 102, at least one processor 104, a memory 106, and a graphics engine 108. The wrinkle texture generator 102 is computer-implemented and calculates a wrinkle texture from images of faces 134 with different expressions, as described in more detail below. The wrinkle texture can be stored in a library 130 for use by the wrinkled face synthesizer to render an image depicting a face with wrinkles. In some examples, the wrinkle texture is calculated from an image of a specified identity, i.e., a specified person or animal. In some cases, the wrinkle texture generator 102 is capable of calculating the wrinkle texture by using a wrinkle texture from the library 130 (in a process called wrinkle grafting) rather than calculating the wrinkle texture from scratch.
[0024] The wrinkled face synthesizer 100 has access to a three-dimensional (3D) mesh 132 of a face, where the mesh is a polygonal mesh. The 3D mesh is stored at any location accessible to the wrinkled face synthesizer 100. In some examples, the mesh depicts a generic face with a neutral expression. A neutral expression is the expression of the face when at rest, where the wrinkles on the face are minimal. In an example, the neutral expression of a human face is the case where the eyes are open and the mouth is closed.
[0025] The graphics engine 108 is any well-known computer graphics engine that obtains a 3D mesh 132, applies at least one wrinkle texture to the 3D mesh 132 according to a tension map, and renders an image from the 3D mesh with the applied wrinkle texture. Thus, the graphics engine 108 calculates an output image depicting a face with wrinkles. In some examples, the graphics engine is a well-known rasterization engine that renders an image from a 3D model using rasterization, or the graphics engine is a well-known ray tracing engine that renders an image from a 3D model using ray tracing. Examples of suitable commercially available ray tracing engines are: Blender Cycles (trademark), Autodesk Arnold (trademark). Examples of suitable commercially available rasterization engines are: Unity (trademark), Unreal (trademark).
[0026] The wrinkled face synthesizer 100 is configured to receive a request 118 from a client device such as a smart phone 122, a computer game device 110, a head-mounted computer 114, a movie production device 120, or other client devices. The request is sent from the client device to the wrinkled face synthesizer 100 via a communication network 124.
[0027] In an example, the request from the client device includes an image of a face in a neutral expression and a value of an expression parameter of the 3D mesh 132. In response to the request, the wrinkled face synthesizer calculates a synthesized output image 116 of a face with an expression and having wrinkles suitable for the expression. This is achieved even when the expression does not exist in the data for creating the wrinkle texture 130.
[0028] In another example, the request from the client device includes the identity (human or animal) of at least one wrinkle texture present in the library 130 and a value of an expression parameter of the 3D mesh 132. In response to the request, the wrinkled face synthesizer uses the wrinkle texture for the identity to calculate a synthesized image of the face. The synthesized output image 116 depicts a face with an expression according to the expression parameter and having wrinkles suitable for the expression and the identity. This is achieved even when the expression does not exist in the data for creating the wrinkle texture of the identity.
[0029] In another example, the request 118 from the client device is a request to generate an image of a default or random face with a default or random expression. In this case, the wrinkled face synthesizer 100 uses the default or random expression parameter value of the 3D mesh 132 and randomly or by default selects a wrinkle texture from the library 130. The synthesized output image 116 depicts a face with an expression according to the expression parameter and having wrinkles according to the selected wrinkle texture.
[0030] The wrinkled face synthesizer 100 receives a request 118 and, in response, generates a synthesized output image 116 and sends it to the client device. The client device uses the output image 116 for one of various useful purposes, including but not limited to: generating a virtual webcam stream, generating video for a computer video game, generating a hologram for display by a mixed reality head-mounted computing device, generating a movie. The wrinkled face synthesizer 100 is capable of calculating synthesized images of dynamic faces with varying expressions and wrinkles as needed for a particular specified expression and a particular specified viewpoint. In an example, the dynamic scene is the face of a speaker. The wrinkled face synthesizer 100 is capable of calculating synthesized images of the face from multiple viewpoints and with any specified dynamic content. Non-limiting examples of the specified viewpoint and dynamic content are: front view, eyes closed, face tilted upward, smiling; perspective view, eyes open, mouth open, angry expression. Note that the wrinkled face synthesizer 100 is capable of calculating synthesized images of facial expressions that do not exist in the data used to calculate the wrinkle texture 130.
[0031] In some examples, the wrinkled face synthesizer is used to generate training data 128 that includes images depicting faces with different expressions and identities. The training data 128 is used to train machine learning systems, such as for generating realistic images depicting faces or other tasks.
[0032] In an example, the face tracker 126 tracks the values of the parameters of the 3D mesh of the video from a person's face, where the person has given appropriate consent to use their data. The parameter values from the face tracker 126 are used in the 3D mesh so that a synthesized image of the person's face can be rendered by the wrinkled face synthesizer. The synthesized image is used for the person's avatar, such as for telepresence, video conferencing, or other applications.
[0033] The wrinkle texture of the present disclosure operates in a non-conventional manner to achieve the rendering of images depicting faces with expressions and appropriate wrinkles.
[0034] Using the wrinkle texture improves the functionality of the underlying computing device by enabling the rendering of images depicting faces with expressions and appropriate wrinkles.
[0035] Alternatively or additionally, the functionality of the wrinkled face synthesizer 100 is performed at least in part by one or more hardware logic components. For example but not limited to, illustrative types of hardware logic components that may optionally be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), graphics processing units (GPUs).
[0036] Figure 2Schematic diagrams of a synthetic image 200 of a wrinkle - free face and the same synthetic image 204 with wrinkles calculated using the techniques described herein. Synthetic image 204 has wrinkles between and below the eyes, and these wrinkles are adapted to the facial expression when the eyes are closed. Note that Figure 2 the line drawings in Figure 2 are created by tracing over a panchromatic photorealistic image and thus do not have a lot of detail, although the wrinkles are shown at least schematically. Included
[0037] Figure 3 is a flowchart of a method for generating a synthetic image of a face with wrinkles. Figure 3 The method of
[0038] is performed by a wrinkled face synthesizer 100. A graphics engine 306 is available and takes as input a 3D face mesh 300, a plurality of wrinkle textures 302, and a tension map 304.
[0039] The 3D mesh has a face and is a polygon mesh, such as a triangular mesh or other polygon mesh. The 3D mesh is a model of a face in a non - neutral expression, for example, where parameter values have been applied to the 3D mesh to apply that non - neutral expression to the 3D mesh.
[0039] The wrinkle textures 302 include at least a neutral texture, a compressed wrinkle texture, and an extended wrinkle texture, as explained in more detail later in this document. In some examples, there are five wrinkle textures 302 for a single identity (i.e., a person or an animal): a neutral texture, a compressed albedo wrinkle texture, an extended albedo wrinkle texture, a compressed displacement wrinkle texture, and an extended displacement wrinkle texture. The neutral texture is a map of a face with a neutral expression, where the map can be a UV map. The compressed wrinkle texture is a map of a face formed by aggregating maps of a first face with different expressions using a tension map. The extended wrinkle texture is also a map of a first face formed by aggregating maps of a face with different expressions using a tension map.
[0040] The tension map can be a UV map, where the numerical values in the map represent the amount of compression or extension. In an example, the tension map is calculated from the 3D polygon mesh model 300, and for an individual vertex of the 3D polygon mesh, the tension map includes the amount of compression to be applied to move the vertex to the position modeling a face of a person with a neutral expression, and the amount of extension to be applied to move the vertex to the position modeling a face of a person with a neutral expression. The compressed wrinkle texture has negative values from the tension map, and the extended wrinkle texture has positive values from the tension map.
[0041] The graphics engine 306 applies the wrinkle texture 302 to the 3D face mesh 300 according to the strain map 304. That is, the strain map 304 serves as the weight for blending the wrinkle texture. Thus, for any arbitrary expression other than the expression represented in the source image, and this process uses the concept of strain in the face mesh to blend the wrinkle texture to obtain a dynamic wrinkling effect. The graphics engine renders an image 308 from the 3D face mesh 300 with the applied wrinkle texture. Thus, the graphics engine renders a realistic image of a face with expressions and wrinkles suitable for those expressions. In Figure 3 the example of, the image 308 is rendered (schematically shown as a line drawing in Figure 3 ), and depicts a face with the expression of the 3D mesh model 300 and having wrinkles on the forehead, between the eyes, and from the nose towards the mouth from the 3D mesh model 300. In this way, dynamic, expression-dependent wrinkles are achieved in an efficient manner. By using the wrinkle texture, scalability is achieved because the wrinkle texture enables generalization to wrinkles not observed in the data used to create the wrinkle texture. Figure 3 The process of is an efficient and effective method for incorporating expression-based wrinkles into a synthetic image of a face. The alternative method of using artist-defined bump maps or normal maps has three drawbacks. First, bump maps and normal maps only simulate underlying geometric changes; the contours and shadows related to face-related tasks such as landmark localization remain unaffected. Second, the bump map or normal map method does not affect the albedo or diffuse texture. Finally, the most critical drawback is scale. The bump map or normal map method requires manual definition of the wrinkle map and mask for each deformation of each character. In contrast, the automatic mesh strain-driven method of the present invention scales naturally with the number of identities and expressions, while incorporating real wrinkles from both scanned albedo and displacement textures. Additionally, this technique handles identities without expression scans, transferring reasonable wrinkles from the most similar neutral texture, as now explained with reference to Figure 4 .
[0042] The inventors have found that using strain in the face mesh enables automatic scaling of the method described herein with identities and expressions, which is a bottleneck for other wrinkle-adding methods that rely on artist effort. In addition to eliminating manual effort, this data-driven method also enables capturing real wrinkles from scans, which does not require artistic judgment.
[0043] By enhancing the realism of synthetic faces with dynamic wrinkles, the methods described herein present a clear case for photorealistic synthesis from an empirical perspective: the methods yield improved performance for models on downstream tasks such as facial landmark recognition. Additionally, the broader approach is also able to mitigate undesirable biases in the model from a social perspective. Synthesizing a dataset that includes diverse faces of different races and genders involves far less manual effort compared to collecting a well-representative dataset in the wild. As a result, downstream realistic systems developed using such synthetic data are less likely to be affected by unfair biases of these sensitive variables.
[0044] Figure 4 is a flowchart of a method for calculating wrinkle textures of a human or animal using a so-called "wrinkle transplantation". As explained in reference Figure 1 There is a wrinkle texture library 130. The pleated face synthesizer 100 receives 402 a neutral image 402 of a target identity without wrinkle texture. The target identity is the human or animal for which a facial image is to be synthesized. The target identity is a real or synthetic human or animal. The received neutral image is an image depicting the face of the target identity with a neutral expression. At this point in the process, no wrinkle texture in the wrinkle texture library 130 is available for the target identity.
[0045] The pleated face synthesizer 100 selects a wrinkle texture from the library 130 using a similarity metric 404. The similarity metric between the received neutral image of the target identity and each neutral texture in the library 130 is calculated. Based on the similarity metric data, one of the neutral textures is selected, for example, the neutral texture found to be most similar to the received neutral image. Any suitable similarity metric can be used, such as mean squared error, structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), similarity. The similarity metric between any characteristic of the neutral texture and the received image (such as pixel color) is calculated. In some cases, the similarity metric between the mesh from the library and the mesh for the target identity is calculated.
[0046] The change amount (delta) between the wrinkle texture of the selected identity and the neutral texture of the selected identity is calculated 406, the change amount is applied 408 to the received neutral image 408, and the resulting wrinkle texture is stored in association with the target identity.
[0047] At operation 412, it is checked according to a rule whether to repeat for more types of wrinkle textures. In some cases, wrinkle textures are calculated only for compressed albedo and extended albedo in addition to neutral. In some cases, wrinkle textures are calculated for compressed albedo, extended albedo, compressed displacement, and extended displacement in addition to neutral. The process ends at operation 414.
[0048] Figure 5Schematic diagram of another method for calculating wrinkle textures of a human or an animal. Here, wrinkle textures are calculated from real images 500, 502, 504 of people with different facial expressions. Note that in Figure 5 a line graph is used to schematically represent the real images to meet the requirements of the patent office. Figure 5 Shows the calculation of albedo wrinkle textures with three original expression images 500, 502, 504 and a neutral expression image (not shown). Each of the original expression images 500, 502, 504 and the neutral expression image has Figure 5 a plurality of other associated images not shown in
[0049] Tension maps 506, 508, 510 corresponding to scans 500, 502, 504 are calculated, so there is one tension map for each scan. To calculate the tension map corresponding to scan 500, scan 500 is used to calculate the identity and expression parameter values of the 3D mesh model of the face by using optimized model fitting or by using machine learning or by using retopology of expert industrial software. The parameter values are applied to the 3D mesh model. As will be explained now, the tension map is then calculated from the 3D mesh model.
[0050] The inventors have formalized mesh tension to capture the amount of compression or expansion at each vertex of a 3D polygonal mesh caused by deformation. More specifically, mesh tension is expressed as a function of the average change in the length of the edges connected to the vertex due to deformation. Consider which has a vertex sequence and an edge sequence undergoing deformation to obtain a mesh X = (V, E). Only consider the deformation such that and X have the same topology. For vertex v i ∈ V, let (e 1, ..., e K ) represents the sequence of K edges connected to v i , where represents the connection to the corresponding edge in. Define the mesh tension as
[0051]
[0052] where [K] = {1, ..., K} and ||.|| represents the edge length. Equation one is expressed in words as the tension at vertex i of the 3D facial mesh is set to the value one minus the following value: the reciprocal of the number of edges originating from the vertex multiplied by the sum of the lengths of the edges originating from the vertex divided by the length of the corresponding edge in the version of the 3D facial mesh representing a face with a neutral expression.
[0053] Note that the process subtracts 1 so that positive values indicate compression, negative values indicate expansion, and a value of 0 indicates no change.
[0054] In fact, for finer manual control, an intensity parameter s is introduced to scale the tension, and a bias b is introduced to artificially favor expansion or compression. The weighted tension at v i is calculated as Furthermore, it is allowed to artificially propagate the expansion and compression effects through the mesh. For each effect, a parameter representing the number of iterations of a morphological dilation (positive value) or erosion (negative value) operation is introduced. First, the propagation of each effect is performed independently on the mesh, and the resulting tension values are added for vertices that end with both expansion and compression.
[0055] In the case of a neutral expression, the tension map has a tension value of zero at each vertex.
[0056] In the example, once the tension values have been calculated for each vertex of the 3D facial mesh, the UV tension map is calculated.
[0057] The facial image 500 and its associated images are used to calculate the albedo texture, which is the UV map as shown at 512 in Figure 5 . The facial image 502 and its associated images are used to calculate the second albedo texture 514. The facial image 504 and its associated images are used to calculate the third albedo texture 516.
[0058] Reference Figure 5, the corresponding strain maps are used as weights to combine the albedo textures 512, 514, 516 and a neutral albedo texture (not shown) to obtain an expanded albedo wrinkle texture 518 and a compressed albedo wrinkle texture 520. In this example, three albedo textures and a neutral albedo texture are used to create one compressed albedo texture and one expanded albedo texture. However, more than four albedo textures can be used.
[0059] In the example, the strain at each vertex is used as a weight in a linear combination of the cleaned textures across expressions, where zero strain corresponds to the neutral texture. In the example, linear aggregation is used by linearly combining the albedo textures using the normalized strain as a weight to obtain the expanded and compressed wrinkle textures 518, 520. Other aggregation methods are possible, such as softmax or max aggregation, where the weighted textures are compared and the maximum value at each image element location is selected for output in the wrinkle texture used. In other examples, as now explained, the weights to be used are learned. For each scan depicting a non-neutral expression, a 3D face mesh and the computed wrinkle texture are used to render a synthetic version. The rendered image is then compared with the real image to obtain a loss and backpropagated to update the weights used to form the wrinkle texture. When rendering using the computed wrinkle map, this type of supervised training produces weights that give the most realistic reproduction of the input scan; in theory, the most accurate wrinkle map is based on observation.
[0060] Figure 5 A process for albedo textures is shown. The same process is applied to obtain displacement wrinkle textures. In this case, the facial image 500 and its associated image are used to compute the distance from the camera viewpoint to the facial surface. Since wrinkles cause the skin to protrude or retract, there is a difference in displacement from the 3D mesh surface compared to the neutral expression without wrinkles. The distance forms a UV displacement texture (not shown). The same operation is performed on other expressions to obtain displacement textures. The displacement textures are combined to form a compressed displacement wrinkle texture and an expanded displacement wrinkle texture.
[0061] Figure 5 The process is data-driven such that the wrinkle texture is computed from empirical data.
[0062] Figure 6 is Figure 5Flowchart of the method. Receive an image depicting a face with an expression. The image is a plurality of images captured at the same moment by different capture devices, and these images have different viewpoints of the face. The capture device viewpoints are known. Register the images 601 to a common topology, for example, by registering the images to a 3D face mesh that models the face depicted in the images. When very accurate, the registration process is time-consuming and computationally expensive. Each image depicting the face must be registered such that each individual vertex in the 3D mesh represents the same surface of the face in each image. As part of the registration process, photogrammetry is used to calculate one or both of a scan and an albedo texture and a displacement texture.
[0063] After the registration process, clean 602 the albedo texture and / or the displacement texture. In an example, cleaning includes removing the depiction of hair. In some cases, sensor noise is removed as part of the cleaning process. In an example, an automatic cleaning process is used. Using an automatic cleaning process is a significant benefit because manual cleaning is error-prone and expensive. In an example, the automatic preprocessing includes: calculating the difference between the texture of a first face with a neutral expression that has been manually preprocessed and the texture of a first face with an expression; calculating a fine mask based on the difference; and using the fine mask as part of the automatic preprocessing. In an example, the automatic preprocessing is a two-stage process, where the first stage includes applying a coarse mask to the texture of the first face to filter out artifacts outside a specified area of the face, and then using a second stage, whereby a fine mask is applied to the texture of the first face.
[0064] Manual cleaning of the scan is a labor-intensive process. To automate the process of masking noise and / or hair artifacts from expression scans, various examples utilize the difference between the original neutral scan and the manually cleaned neutral scan. Specifically, a two-stage masking process is employed. First, a coarse mask that is independent of identity is applied to filter out most of the artifacts outside the hockey-mask and neck regions where expression-based wrinkles occur. Next, to capture the manual changes made by an artist during the cleaning of each neutral scan, a background subtraction technique based on a Gaussian mixture model or other background subtraction techniques is used. Treating the cleaned neutral texture as the background and the original neutral texture as the foreground, a identity-specific mask of noise and hair artifacts for each identity is obtained. The fine mask is applied to clean the texture from the corresponding expression scan for each identity.
[0065] The registered and cleaned texture is used to calculate a strain map 605. As described above with reference to Figure 5 Calculating the strain map. In this example, one expression is depicted in the received image, so there is one strain map.
[0066] The registered and cleaned texture is also used to calculate the expression map 604. The expression map is a UV map formed from the received images. In this example, an expression is depicted in the received image, and thus there is an expression map.
[0067] At decision 606, it is checked whether to repeat the process for more expressions, for example, by checking whether a threshold number of expression maps have been calculated or by checking whether more received images are available. If the decision is to repeat, the process returns to operation 600 and an image depicting the same face with another facial expression different from the first facial expression is received. The received image is registered 601 and cleaned 602, and another strain map is calculated 605 together with another expression map 604. Operations 605 and 604 are optionally performed in parallel.
[0068] At least one of the expression maps calculated by operation 604 is for a neutral expression, and in this case, the corresponding strain map calculated by operation 605 has a zero strain value at each vertex.
[0069] When decision point 606 determines that a sufficient number of expression maps 604 have been calculated, the process proceeds to operation 608, where the expression maps including the neutral expression are combined 608 in a manner weighted by the corresponding strain maps. As described above, as Figure 6 a result of the process, each expression map has an associated strain map. The individual expression maps are weighted by their corresponding strain maps, such as by calculating so as to take the value with the maximum strain value or by calculating a weighted sum based on the strain values (linearly), or by calculating the softmax between the strain values. Then the weighted expression maps are aggregated by addition or in any other suitable manner.
[0070] The result of operation 608 is the wrinkle texture. The wrinkle texture is stored 610.
[0071] Figure 6 The entire method of is optionally repeated 612 to form different types of wrinkle textures, such as one or more of the following: compressed albedo wrinkle texture, expanded albedo wrinkle texture, compressed displacement wrinkle texture, expanded displacement wrinkle texture. To form a compressed albedo wrinkle texture, operation 604 is configured to calculate an albedo expression map, and operation 608 is configured to use an aggregation method, such as any of the following: linear aggregation, taking the maximum value, calculating the softmax, using learned weights for weighted aggregation. The compressed albedo wrinkle texture maintains an absolute compression value. As described above, A positive value indicates compression, a negative value indicates expansion, and a value of 0 indicates no change. The compressed albedo wrinkle texture remains absolutely compressed to. To form the expanded albedo wrinkle texture, operation 604 is configured to calculate the albedo atlas, and operation 608 is configured to use an aggregation method, such as any of the following: linear aggregation, taking the maximum value, calculating softmax, using learned weights for weighted aggregation. The expanded albedo wrinkle texture remains at an absolute expansion value. To form the compressed displacement wrinkle texture, operation 604 is configured to calculate the displacement atlas, and operation 608 is configured to use an aggregation method, such as any of the following: linear aggregation, taking the maximum value, calculating softmax, using learned weights for weighted aggregation. The compressed displacement wrinkle texture remains at an absolute compression value. To form the expanded displacement wrinkle texture, operation 604 is configured to calculate the displacement atlas, and operation 608 is configured to use an aggregation method, such as any of the following: linear aggregation, taking the maximum value, calculating softmax, using learned weights for weighted aggregation. The expanded displacement wrinkle texture remains at an absolute expansion value.
[0072] It is not necessary to use both the albedo wrinkle texture and the displacement wrinkle texture. In some examples, only the albedo compression and expansion wrinkle textures are used. In some examples, only the displacement compression and expansion wrinkle textures are used.
[0073] Figure 5 and Figure 6 The process gives an efficient and effective method for capturing the complex wrinkling effect of an identity from a high-resolution scan (received image) of the pose of the face.
[0074] Figure 7 is a flowchart of a method for training and using a face synthesizer. Figure 3 The process is used to calculate multiple synthetic images of a face with wrinkles, and these images are stored in database 700. The synthetic images of the face with wrinkles from memory 700 are used to train the face image synthesizer 702.
[0075] The trained face image synthesizer 704 is capable of receiving an input including expression parameter values and identity parameter values 706, and generating a synthetic image 708 of the face according to the parameter values and using wrinkles.
[0076] By using the synthetic training data from storage 700, it was unexpectedly found that the trained face image synthesizer 704 has improved performance compared to when using real images for training. In the example,
[0077] As explained now, the present technology has been empirically tested. A set of 208 high-quality commercially available 3D scans of individuals was obtained. All 208 identities included scans with neutral expressions, while 52 identities included additional scans for posed expressions. The neutral scans for each identity were manually cleaned to remove noise and hair and registered to the topology of the 3D face model, resulting in a mesh with 7667 vertices and 7414 polygons.
[0078] 3D scans were used as described herein to generate synthetic images of faces with expressions and appropriate wrinkles. When the training dataset of 100k synthetic images was rendered, it consisted of 20k identities, each with 5 frames (different viewpoints, expressions, and environments). Ground truth annotations of 703 dense 2D landmarks were generated from the face meshes to accompany each image.
[0079] Then, the synthetic images of faces in the training dataset were used to train a neural network to detect face landmarks. The neural network was an off-the-shelf ResNet101. 256x256 pixel red-green-blue (RGB) images were used as input to predict dense face landmarks.
[0080] Another version of the same neural network was trained for the same task using real images of faces.
[0081] It has been found that for the eye region results on datasets called 300W, 300W-winks, and Pexels, the method trained only on synthesis outperforms the model based on real data. The following table shows different methods in rows and the performance on different datasets in columns. The performance of the trained neural network on the landmark detection task is represented as a numerical value. The following table gives the eye opening error for the Pexels dataset and the eyelid point-to-polyline error for the 300W dataset and the blink subset. In all cases, normalization was done by the bounding box diagonal. The lower, the better. The error for the eyelid landmarks was calculated by computing the point-to-line distance between each predicted eyelid landmark and the corresponding polyline defining the true eyelid..
[0082]
[0083] The Pexels dataset contains 318 images of fully closed eyes (due to blinking, frowning, or compressed faces) and 105 images of only one eye closed (blinking). This allows the evaluation of the model performance under such conditions that are rare in other datasets. Knowing which images contain fully closed eyes or only one eye closed allows the measurement of eyelid accuracy without explicit landmark annotations. The eye opening error was defined as the average eye aperture of both eyes in the case of closed eyes and the eye aperture of the closed eye in the case of blinking.
[0084] The 300W dataset is a commercially available facial dataset. A small subset of 30 images from 300W was identified as containing winks and compressed facial expressions (300W - winks) to provide a more detailed indication of performance under such deformations.
[0085] In another example, synthetic facial images with wrinkles were used to train a machine learning system to predict surface normals. Surface normals can be used to infer 3D information about a surface from a 2D image and can be used for clothing and facial shape reconstruction and relighting. In the example, a U - Net was trained with a ResNet18 encoder to predict the camera surface normals of a face. As input, 256x256 pixel RGB images (10k identities, 5 frames per identity) from a dataset of 50k synthetic images were used. The network was trained using PyTorch for 200 epochs with a learning rate of 1e - 3 using a cosine similarity loss. The camera space surface normal images rendered as part of the synthetic data pipeline described herein were used as the ground truth. The network trained using synthetic images with mesh - tension - driven wrinkles resulted in predictions with significantly more high - frequency details on the face compared to a network trained using data without mesh - tension - driven wrinkles.
[0086] Figure 8 Exemplary components of a computing - based device 800 implemented as any form of computing and / or electronic device are shown, and in some examples, an embodiment of a wrinkled face synthesizer 802 is implemented therein.
[0087] The computing - based device 800 includes one or more processors 814, which are microprocessors, controllers, or any other suitable type of processor for processing computer - executable instructions to control the operation of the device to compute a wrinkle texture 804 and use the wrinkle texture to render a synthetic image of a face with an expression and wrinkles suitable for the expression. In some examples, such as in the case of using a system - on - chip architecture, the processor 814 includes one or more fixed - function blocks (also known as accelerators) that are implemented in hardware (rather than software or firmware) Figures 3 to 7 as part of any of the methods described herein. The wrinkled face synthesizer 802 includes a wrinkle texture 804 and a graphics engine 806. Platform software including an operating system 808 or any other suitable platform software is provided at the computing - based device to enable an application software 810 to execute on the device. A data repository 822 stores parameter values, images, expression maps, 3D face mesh models, and other data.
[0088] Provide computer-executable instructions using any computer-readable medium accessible by a computing-based device 800. Computer-readable media include, for example, computer storage media such as memory 812 and communication media. Computer storage media such as memory 812 includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, etc. Computer storage media includes but is not limited to random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media for storing information accessible by a computing device. In contrast, communication media embodies computer-readable instructions, data structures, program modules, etc. in a modulated data signal such as a carrier wave or other transmission mechanism. As defined herein, computer storage media does not include communication media. Thus, computer storage media should not be construed as propagating signals themselves. Although computer storage media (memory 812) is shown within computing-based device 800, it should be understood that in some examples, storage is distributed or remote and accessed via a network or other communication link (e.g., using communication interface 816). Computing-based device 800 also includes an optional capture device 818 for capturing images of a face or other scene.
[0089] Alternatively or in addition to the other examples described herein, examples include any combination of the following:
[0090] Clause A. A computer-implemented method for computing an image depicting a face of a person or animal, the method comprising:
[0091] Access a 3D polygon mesh model of a face of a person or animal having a non-neutral expression;
[0092] Compute a tension map from the 3D polygon mesh model, where for a single vertex of the 3D polygon mesh, the tension map includes a compression amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression, and an expansion amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression;
[0093] For a first face, obtain the following:
[0094] A neutral texture, a compressed wrinkle texture, and an expanded wrinkle texture,
[0095] Among them, the neutral texture includes a graph of a first face with a neutral expression, the compressed wrinkle texture is a graph of the first face formed by aggregating graphs of the first face with different expressions using the tension graph, and the extended wrinkle texture includes a graph of the first face formed by aggregating graphs of the first face with different expressions using the tension graph;
[0096] Apply the wrinkle texture to the 3D polygon mesh model according to the tension graph; and
[0097] Render an image from the 3D polygon mesh model.
[0098] Clause B. The method according to Clause A includes applying the wrinkle texture to the 3D model and rendering an image using a graphics engine, where the graphics engine is a rasterization engine that uses rasterization to render an image from the 3D model, or the graphics engine is a ray tracing engine that uses ray tracing to render an image from the 3D model.
[0099] Clause C. The method according to any of the preceding clauses, wherein the first face is the face of a human or an animal, and the graph of the first face and the 3D polygon mesh model are registered to a common topology.
[0100] Clause D. The method according to any of the preceding clauses, wherein at least one of the graphs of the first face is preprocessed using automatic preprocessing to remove hair.
[0101] Clause E. The method according to Clause D, wherein the automatic preprocessing includes: calculating the difference between a manually preprocessed graph of the first face with a neutral expression and a graph of the first face with an expression; calculating a fine mask from the difference; and using the fine mask as part of the automatic preprocessing.
[0102] Clause F. The method according to Clause E, wherein the automatic preprocessing is a two-stage process, where the first stage includes applying a coarse mask to the graph of the first face to filter out artifacts outside a specified area of the face, and then using a second stage, whereby the fine mask is applied to the graph of the first face.
[0103] Clause G. The method according to any of the preceding claims, wherein the first face is a different face, and the neutral texture, the compressed wrinkle texture, and the extended wrinkle texture are obtained from a wrinkle texture library.
[0104] Clause H. The method according to Clause G includes: receiving an image of a target identity with a neutral expression; selecting a neutral texture from the library using a similarity metric, the neutral texture having an associated compressed wrinkle texture and an associated extended wrinkle texture.
[0105] Clause I. The method according to Clause H includes calculating a change amount between the associated compressed wrinkle texture or the associated extended wrinkle texture and the selected neutral texture, and
[0106] applying the change amount to the received image of the target identity with a neutral expression to form a wrinkle texture for the target identity; and
[0107] applying the change amount to the received image of the target identity to form a compressed or extended wrinkle texture of the target identity respectively.
[0108] Clause J. The method according to any of the foregoing clauses, wherein the compressed wrinkle texture is an albedo texture, and wherein the extended wrinkle texture is an albedo texture, and the graph is a color graph.
[0109] Clause K. The method according to Clause J includes: for the first face, obtaining a compressed wrinkle texture and an extended wrinkle texture, the compressed wrinkle texture being a displacement texture, the extended wrinkle texture being a displacement texture, and wherein a graphics engine is used to apply the displacement texture to the 3D polygon mesh model according to the tension map.
[0110] Clause L. The method according to any one of Clauses A to I, wherein the compressed wrinkle texture is a displacement wrinkle texture, the extended wrinkle texture is a displacement texture, and the graph is a displacement map.
[0111] Clause M. The method according to any of the foregoing clauses includes aggregating the graph by any one of the following: linear aggregation, max aggregation, softmax aggregation, weighted aggregation using learned weights.
[0112] Clause N. The method according to any of the foregoing clauses includes: for different non-neutral expressions of the face, repeating the method so as to render a plurality of images depicting the face with wrinkles, and using the plurality of images to train a machine learning model.
[0113] Clause O. The method according to Clause N, wherein the machine learning model is a facial image synthesizer or a facial landmark recognition system or a normal map prediction system.
[0114] Clause P. An apparatus includes:
[0115] at least one processor;
[0116] A memory storing instructions which, when executed by the at least one processor, perform a method for computing an image depicting a face, the method comprising:
[0117] Accessing a 3D polygon mesh model of the face of a person or animal with a non-neutral expression;
[0118] Computing a strain map from the 3D polygon mesh model, wherein for a single vertex of the 3D polygon mesh model, the strain map includes a compression amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression, and an expansion amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression;
[0119] For a first face that is the face of the person or animal or a different face, obtaining:
[0120] A neutral texture, a compressed wrinkle texture, and an expanded wrinkle texture,
[0121] wherein the neutral texture includes a color map of the first face with a neutral expression, the compressed wrinkle texture is a map of the first face formed by aggregating maps of the first face with different expressions, and the expanded wrinkle texture includes a map of the first face formed by aggregating maps of the first face with different expressions;
[0122] Applying the wrinkle textures to the 3D model according to the strain map; and
[0123] Rendering an image from the 3D model.
[0124] Clause Q. A computer-implemented method, comprising:
[0125] Accessing a 3D polygon mesh model of a face with a non-neutral expression;
[0126] Computing a strain map from the 3D polygon mesh model, wherein for a single vertex of the 3D polygon mesh, the strain map includes a compression amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression, and an expansion amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression;
[0127] Accessing a plurality of maps of the face in different expressions, the maps being registered to the topology of the 3D polygon mesh;
[0128] Computing a weighted combination of the maps to produce a compressed wrinkle texture, wherein the weights used in the weighted combination are negative weights from the strain map;
[0129] Compute a weighted combination of the maps to produce an extended wrinkle texture, where the weights used in the weighted combination are positive weights from a tension map;
[0130] Store the compressed wrinkle texture and the extended wrinkle texture for use in computing the rendering of a face under an expression with wrinkles.
[0131] Clause R. The method according to Clause Q, wherein the weighted combination includes any one of the following: linear aggregation, max aggregation, softmax aggregation, weighted aggregation using learned weights.
[0132] Clause S. The method according to Clause Q or Clause R, wherein the map is a color map, and the extended wrinkle texture is an extended albedo wrinkle texture, and the compressed wrinkle texture is a compressed albedo wrinkle texture.
[0133] Clause T. The method according to Clause S, including: accessing a plurality of displacement maps of the face under different expressions, the displacement maps being registered to the topology of the 3D polygon mesh; and
[0134] Compute a weighted combination of the displacement maps to produce a compressed displacement wrinkle texture, where the weights used in the weighted combination are negative weights from a tension map;
[0135] Compute a weighted combination of the maps to produce an extended displacement wrinkle texture, where the weights used in the weighted combination are positive weights from a tension map;
[0136] Store the compressed displacement wrinkle texture and the extended displacement wrinkle texture.
[0137] The term "computer" or "computing-based device" is used herein to refer to any device having processing capabilities such that it can execute instructions. Those skilled in the art will recognize that such processing capabilities are incorporated into many different devices, and thus the terms "computer" and "computing-based device" each include personal computers (PCs), servers, mobile phones (including smart phones), tablet computers, set-top boxes, media players, game consoles, personal digital assistants, wearable computers, and many other devices.
[0138] In some examples, the methods described herein are performed by software in machine-readable form on a tangible storage medium, such as in the form of a computer program including computer program code means adapted to perform all the operations of one or more of the methods described herein when the program is run on a computer, and wherein the computer program can be embodied on a computer-readable medium. The software is adapted to be executed on a parallel processor or a serial processor such that the method operations can be performed in any suitable order or simultaneously.
[0139] Those skilled in the art will recognize that storage devices for storing program instructions are optionally distributed over a network. For example, a remote computer can store examples of processes described as software. A local or terminal computer can access the remote computer and download part or all of the software to run the program. Alternatively, the local computer can download software segments as needed, or execute some software instructions at the local terminal and some software instructions at the remote computer (or computer network). Those skilled in the art will also recognize that all or part of the software instructions can be executed by dedicated circuitry, such as a digital signal processor (DSP), programmable logic array, etc., by using conventional techniques known to those skilled in the art.
[0140] It will be apparent to those skilled in the art that any range or device value given herein can be extended or changed without losing the desired effect.
[0141] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims need not be limited to the specific features or acts described above. Rather, the above specific features and acts are disclosed as example forms of implementing the claims.
[0142] It should be understood that the above benefits and advantages may relate to one embodiment or may relate to several embodiments. Embodiments are not limited to embodiments that solve any or all of the stated problems or have any or all of the stated benefits and advantages. It will also be understood that references to "one" refer to one or more of those items.
[0143] The operations of the methods described herein can be performed in any suitable order or, where appropriate, simultaneously. Additionally, individual blocks can be deleted from any method without departing from the scope of the subject matter described herein. Aspects of any of the above-described examples can be combined with aspects of any of the other examples described to form other examples without losing the desired effect.
[0144] The term "comprising" is used herein to mean including the identified method blocks or elements, but such blocks or elements do not comprise an exclusive list, and a method or apparatus can contain additional blocks or elements.
[0145] It should be understood that the above description is given by way of example only, and that various modifications can be made by those skilled in the art. The above specification, examples, and data provide a complete description of the structure and use of the exemplary embodiments. Although the various embodiments have been described above with a certain degree of particularity or with reference to one or more individual embodiments, those skilled in the art can make many changes to the disclosed embodiments without departing from the scope of this specification.
Claims
1. A computer-implemented method for computing an image depicting the face of a person or animal, the method comprising: accessing a 3D polygonal mesh model of the face of the person or animal with a non-neutral expression; computing a tension map from the 3D polygonal mesh model, wherein for a single vertex of the 3D polygonal mesh model, the tension map includes a compression amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression, and an expansion amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression; for a first face, obtaining: a neutral texture, a compressed wrinkle texture, and an expanded wrinkle texture, wherein the neutral texture includes a map of the first face with a neutral expression, and the compressed wrinkle texture is a map of the first face formed by aggregating maps of the first face with different expressions using the tension map, and the expanded wrinkle texture includes a map of the first face formed by aggregating maps of the first face with different expressions using the tension map; applying the wrinkle texture to the 3D polygonal mesh model according to the tension map; and rendering the image from the 3D polygonal mesh model.
2. The method according to claim 1, comprising applying the wrinkle texture to the 3D polygonal mesh model and rendering the image using a graphics engine, wherein, the graphics engine is a rasterization engine that renders the image from the 3D polygonal mesh model using rasterization, or the graphics engine is a ray tracing engine that renders the image from the 3D polygonal mesh model using ray tracing.
3. The method according to claim 1, wherein, the first face is the face of the person or animal, and the map of the first face and the 3D polygonal mesh model are registered to a common topology.
4. The method according to claim 1, wherein, at least one of the maps of the first face is preprocessed using automatic preprocessing to remove hair.
5. The method according to claim 4, wherein, the automatic preprocessing includes: computing the difference between a manually preprocessed map of the first face with a neutral expression and a map of the first face with an expression; computing a fine mask from the difference; and using the fine mask as part of the automatic preprocessing.
6. The method according to claim 5, wherein, the automatic preprocessing is a two-stage process, wherein the first stage includes applying a coarse mask to the map of the first face to filter out artifacts outside a specified region of the face, and then using a second stage, whereby the fine mask is applied to the map of the first face.
7. The method according to claim 1, wherein, the first face is a different face, and the neutral texture, the compressed wrinkle texture, and the expanded wrinkle texture are obtained from a wrinkle texture library.
8. The method according to claim 7, comprising: receiving an image of a target identity with a neutral expression; Select neutral textures from the library using a similarity metric, the neutral textures having an associated compressed wrinkle texture and an associated expanded wrinkle texture.
9. The method according to claim 8, comprising: calculating a change amount between the associated compressed wrinkle texture or the associated expanded wrinkle texture and the selected neutral texture, and applying the change amount to an image of the target identity with a neutral expression to respectively form a compressed wrinkle texture or an expanded wrinkle texture for the target identity.
10. The method according to claim 1, wherein, the compressed wrinkle texture is an albedo texture, and wherein the expanded wrinkle texture is an albedo texture, and the image is a color image.
11. The method according to claim 10, comprising: for the first face, obtaining a compressed wrinkle texture and an expanded wrinkle texture, the compressed wrinkle texture being a displacement texture, the expanded wrinkle texture being a displacement texture, and wherein a graphics engine is used to apply the displacement texture to the 3D polygon mesh model according to the tension map.
12. The method according to claim 1, wherein, the compressed wrinkle texture is a displacement wrinkle texture, the expanded wrinkle texture is a displacement texture, and the image is a displacement map.
13. The method according to claim 1, comprising: aggregating the image by any one of the following: linear aggregation, max aggregation, softmax aggregation, weighted aggregation using learned weights.
14. The method according to claim 1, comprising: repeating the method for different non-neutral expressions of the face so as to render a plurality of images depicting the face with wrinkles, and using the plurality of images to train a machine learning model.
15. An apparatus, comprising: at least one processor; a memory storing instructions which, when executed by the at least one processor, perform a method for calculating an image depicting a face, the method comprising: accessing a 3D polygon mesh model of a face of a person or animal with a non-neutral expression; calculating a tension map from the 3D polygon mesh model, for a single vertex of the 3D polygon mesh model, the tension map including a compression amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression, and an expansion amount to be applied to move the vertex to a position modeling the face of the person with a neutral expression; for a first face that is the face of the person or animal or a different face, obtaining: a neutral texture, a compressed wrinkle texture, and an expanded wrinkle texture, wherein the neutral texture includes a color image of the first face with a neutral expression, and the compressed wrinkle texture is an image of the first face formed by aggregating images of the first face with different expressions, and the expanded wrinkle texture includes an image of the first face formed by aggregating images of the first face with different expressions; applying the wrinkle texture to the 3D polygon mesh model according to the tension map; and Render the image from the 3D polygon mesh model.