Single-Image Large-Pose 3D Color Face Reconstruction Method Based on UV Position Map and CGAN
Through the UV position map and CGAN method, UV texture maps are generated and completed, and the self-occlusion problem in large-pose face image reconstruction is solved, and high-precision three-dimensional face reconstruction and texture detail recovery is achieved, which is suitable for face recognition and multi-angle image generation.
Patent Information
- Application Number
- CN202110290418.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-03-18
AI Technical Summary
The existing three-dimensional face reconstruction technology has self-occlusion problems in large-pose face images, resulting in reduced reconstruction accuracy and missing textures, and it is impossible to generate a complete color three-dimensional face model.
Using a method based on UV position map and conditional generation adversarial network (CGAN), a three-dimensional point cloud model is recorded by generating UV position maps, and a coding-decoder network and conditional generation adversarial network is used to complete the incomplete UV texture map, and finally fit a complete color 3D face model.
It realizes the generation of a complete three-dimensional color face model from a single image, solves the problem of self-occlusion in large postures, improves reconstruction accuracy and texture details, and is suitable for face recognition and multi-angle face image generation, reducing the complexity of data acquisition.
Smart Images

Figure CN113052976B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular to a single image large posture three-dimensional color face reconstruction method based on UV position map and CGAN Background Art
[0002] Biometrics is a type of information feature that has been widely used and paid attention to recently. The technology for reconstructing the corresponding model is also developing with the changes in social needs. The rich feature information contained in the face makes it an important carrier for human identity recognition, expression recognition, age and gender judgment, etc., so the processing of facial information has always been an important research topic in the field of computer vision. However, the facial information that can be retained in a two-dimensional image is very limited, and it will be affected by shooting angle, object occlusion and lighting angle. The recently popular three-dimensional reconstruction technology has also been greatly improved with the development of machine learning technology. Therefore, using this technology to reconstruct a complete three-dimensional face model from a two-dimensional image can alleviate the impact of the above problems and give the model more information. Summary of the invention
[0003] Existing 3D face reconstruction technology can use a 3D model obtained from a single image, but the large angle of the face in the image causes large errors in the reconstruction result, and the model lacks complete surface texture, resulting in reduced authenticity. The present invention provides a single-image large-pose 3D color face reconstruction method based on UV position map and CGAN, which targets a large number of faces in a single image that are invisible due to self-occlusion of large-pose faces, resulting in reduced accuracy of 3D face reconstruction and a lack of a large amount of facial color texture in the final result. The present invention mainly uses UV position records to generate a 3D point cloud model, and then uses a network designed based on CGAN to complete the incomplete face, and finally obtains a complete color 3D face model.
[0004] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0005] A single-image large-pose 3D color face reconstruction method based on UV position map and CGAN includes the following steps:
[0006] S1: Collecting Data
[0007] Use active vision method to obtain a large number of 3D models of human faces. At the same time, take photos with the front face as 0° and the rotation range of [-90°, -90°] in 5° steps, classify them and save them in the set format;
[0008] S2: Generate UV position map
[0009] The UV map is a two-dimensional image plane converted from three-dimensional surface parameters. A three-dimensional model uses the (X, Y, Z) coordinate system, and its structure is a polygon model with point cloud coordinates as vertices. The work of the UV coordinate system is to correspond the vertices of the polygon to the pixels on the two-dimensional image. In this way, the UV coordinates define the position information of each point on the picture, and these points are interconnected with the three-dimensional model. Image smoothing interpolation is performed on the gaps between points, so that the UV texture map can be mapped onto the three-dimensional model. According to this principle, the three-dimensional point cloud data is recorded into a two-dimensional image by constructing a UV position map;
[0010] S3: Generate the UV texture map
[0011] After obtaining the UV position map, a bilinear sampler is used to resample the vertices of the three-dimensional model and their related UV coordinates, and the color texture information in the photo is rendered into the position map to obtain the required UV texture map; However, due to self-occlusion, a large area of the face is invisible, resulting in the mutilation of the output texture map, and the mutilated part is filled with black;
[0012] S4: Construct an encoder-decoder network
[0013] The 256*256*3 image input in the encoder part first passes through a convolutional layer with a kernel of 4, and then 10 residual blocks are used to obtain its 8*8*512 features. Here, it is not directly compressed into a one-dimensional feature vector because for a three-dimensional face model, the information about the relative positions of points in space. Although this will increase the training difficulty, retaining the spatial position information will improve the accuracy of the reconstruction result; In the decoder part, 17 transposed convolutions are used to predict and generate a 256*256*3 UV position map;
[0014] S5: Construct a loss function
[0015] The mean square error is used to calculate the error between the position map P(u, v) obtained by 3DMM-STN and the UV position between the map. However, when calculating the mean square error, all points in the map have the same weight, while the reconstruction accuracy requirements for different regions of the face are different. For example, the neck part in the image has less information and little significance for reconstruction, so its weight needs to be reduced; For parts of the face such as eyes, nose, mouth, ears, etc. that contain a lot of useful information, the corresponding weights need to be increased; A mask is used to highlight the important parts, and the weights are changed by setting different gray values for different parts and then normalizing;
[0016]
[0017] Among them, (u, v) represents a point in the UV coordinate system, and P(u, v) represents the position of the point in the real target image. represents the point position generated by the network, and W(u, v) represents the weight assigned to the corresponding point.
[0018] S6: Train the encoder-decoder network
[0019] Use the UV position map as the target value, and the face photos at various angles as the input to the encoder-decoder network for training. Use the Adam optimizer, set the learning rate to 0.0002, and set the batch size to 16. The final network output is the UV position map; then use a simple convolutional neural network to reconstruct the three-dimensional shape of the face from the UV position map, but at this time, texture details have not been added to its surface yet.
[0020] S7: Construct a conditional generative adversarial network
[0021] The main inspiration of GAN comes from the idea of zero-sum games in game theory. Applied to the neural network in deep learning, it is to continuously play games between the generator G and the discriminator D. G is used as the generator, and the input is a random noise x, and an image is generated through this random noise; D is used as the discriminator to determine whether the picture is real, and its input is the picture; during the training process, G needs to try its best to generate real pictures to deceive D, while D needs to distinguish the true and false of the pictures generated by G, forming a game process, and finally reaching the Nash equilibrium point.
[0022]
[0023] Among them, x is the input noise, and its range is the probability distribution p z (x), y is the real picture, and its range is the real data p data (y), G represents the generator, and D represents the generator.
[0024] In the constructed GAN, use the incomplete UV texture map to replace the noise input to the generator. The generator part adopts an encoder-decoder structure. The encoder part has 8 convolutional layers, and the decoder part has 8 transposed convolutional layers. Their convolutional kernels are all 4, and the stride is 2; the discriminator part adopts 4 convolutional layers, and connects the input picture with the label to obtain their features. However, when training GAN, problems such as instability, gradient disappearance, and mode collapse often occur. Therefore, according to the CGAN idea, make certain improvements to GAN to obtain better results. A deep convolutional neural adversarial network with certain result constraints. Compared with GAN, its generator uses fractional-strided convolution, and the discriminator uses strided convolution to replace all pooling layers, removes the fully connected hidden layers for deeper architectures, uses ReLU as the activation function in the generator, and uses LeakyReLU as the activation function in the discriminator;
[0025] S8: Construct the adversarial loss function
[0026] To improve the realism and texture details of the generated UV texture map, multiple loss functions are set to take the weighted sum, which are the pixel-level loss function L1, the face feature-level loss function L f , the symmetry loss function L sym and the adversarial loss function L d ;
[0027] The pixel-level loss function L1 adopts the mean square error to make the generated image close to the target image at the pixel level, and a mask P is added to increase the weights of the eyes, nose, and mouth parts. As a key part to improve performance, it will be given a higher weight than other loss functions;
[0028]
[0029] Among them, W and H are the width and height of the image respectively, j represents the pixel position in the width, k represents the pixel position in the length, and x and y are the input image and the real image respectively;
[0030] Introduce the deepface module, denoted as F here, to obtain and compare the features of the faces in the generated image and the label, determine the face contour, eye, nose, and mouth positions from a global perspective, and maintain the different features of each person's data, so that the output result will not be an average and similar UV texture map.
[0031]
[0032] Among them, N represents the number of features obtained, F represents the result obtained by inputting the image into the deepface module, and x and y represent the input image and the real image respectively.
[0033] Due to the symmetric characteristics of the human face, the symmetry loss function is adopted. Using the prior knowledge of the visible part can effectively solve the self-occlusion problem caused by large poses and complement the parts that cannot be seen in a single image; in reality, self-occlusion may cause the left or right side to be invisible, so both left occlusion and right occlusion exist in the input images during training; however, using the symmetry loss function may misjudge the different brightness levels on both sides of the face caused by lighting, so it is necessary to adjust the weight ratio with other loss functions well and not give the symmetry loss function a very large weight.
[0034]
[0035] Among them, W and H are the width and height of the image respectively, j represents the pixel position in the width, k represents the pixel position in the length, and x and y are the input image and the real image respectively.
[0036] Calculate the loss value of the generated face image discriminated from the label using the adversarial loss function, which helps improve the realism of the generated image and reduce the blurriness.
[0037]
[0038] Among them, G represents the generator, D represents the discriminator, W and H are the width and height of the image respectively, j represents the pixel position in the width direction, k represents the pixel position in the length direction, and x is the input image.
[0039] The final generation loss function takes the weighted sum of the above loss functions.
[0040] L g =λ1L1+λ f L f +λ sym L sym +λ d L d (7)
[0041] Among them, L1 is the pixel-level loss function, λ1 is the weight of the pixel-level loss function, L f is the face feature-level loss function, λ f is the weight of the face feature-level loss function, L sym is the symmetry loss function, λ sym is the weight of the symmetry loss function, L d is the adversarial loss function, λ d is the weight of the adversarial loss function.
[0042] S9: Train the conditional adversarial generation network
[0043] Using the scanned complete UV texture map as the generation target, use the incomplete UV texture map to replace the noise and input it into the network for training. Use the Adam optimizer with the learning rate set to 0.0002. The obtained model can complete the incomplete part in the UV texture map.
[0044] S10: Fit the generated three-dimensional face shape model with the UV texture map to obtain the final complete colored three-dimensional face model.
[0045] The beneficial effects of the present invention are: solving the self-occlusion problem that often appears in the reconstruction of three-dimensional face models from large-pose single face images, and realizing the generation of a complete and realistic three-dimensional face model directly from a single two-dimensional image. It can solve the problem of reduced recognition accuracy caused by large poses in face recognition, or can be used to generate multi-angle face images from single face images to increase experimental data and reduce complex data collection. Description of the Drawings
[0046] Figure 1It is the overall structure diagram of the 3D color face model generation network.
[0047] Figure 2 It is the schematic diagram of recording 3D information in the UV position map. Specific implementation manners
[0048] The following further describes the present invention.
[0049] Refer to Figure 1 and Figure 2 A single-image large-pose 3D color face reconstruction method based on the UV position map and CGAN includes the following steps:
[0050] S1: Data acquisition
[0051] Use a laser scanner to obtain 3D models of a large number of faces, and at the same time take photos with the front face as 0°, rotating in steps of 5° within the range of [-90°, -90°], classify and save them according to the set format;
[0052] S2: Generate the UV position map
[0053] A 3D model uses the (X, Y, Z) coordinate system, and its structure is a polygon model with point cloud coordinates as vertices. The work of the UV coordinate system is to correspond the vertices of the polygon to the pixels on the 2D image. In this way, the UV coordinates define the position information of each point on the picture, and these points are interconnected with the 3D model. Smooth interpolation processing of the image is performed at the gaps between points, so that the UV texture map can be mapped onto the 3D model. According to this principle, we can record the 3D point cloud data into a 2D image by constructing the UV position map.
[0054] S3: Generate the UV texture map
[0055] After obtaining the UV position map, use a bilinear sampler to resample the vertices of the 3D model and their related UV coordinates, and render the color texture information in the photo into the position map to obtain the required UV texture map.
[0056] S4: Construct an encoder-decoder network
[0057] In the encoder part, the input 256*256*3 image first passes through a convolutional layer with a kernel of 4, and then uses 10 residual blocks to obtain its 8*8*512 features. Here, it is not directly compressed into a one-dimensional feature vector because for the 3D face model, the information about the relative positions of each point in space, although this will increase the training difficulty, retaining the spatial position information will improve the accuracy of the reconstruction result. In the decoder part, 17 transposed convolutions are used to predict and generate a 256*256*3 UV position map.
[0058] S5: Construct the loss function
[0059] Calculate the error between the position map P(u, v) obtained by 3DMM-STN and the UV positions output by the network using the mean squared error. However, when calculating the mean squared error, all points in the map have the same weight, while the reconstruction accuracy requirements for different regions of the human face are different. For example, the neck part in the image has less information and little significance for reconstruction, so its weight needs to be reduced. For parts of the human face such as eyes, nose, mouth, and ears that contain a large amount of useful information, the corresponding weights need to be increased; use a mask to highlight the important parts, and change the weights by setting different gray values for different parts and then normalizing them. Among them, (u, v) represents a point in the UV coordinate system, P(u, v) represents the position of the point in the real target map,
[0060]
[0061] represents the point position generated by the network, and W(u, v) represents the weight assigned to the corresponding point. represents the point position generated by the network, and W(u, v) represents the weight assigned to the corresponding point.
[0062] S6: Train the encoder-decoder network
[0063] Use the UV position map as the target value, and input face photos at various angles into the encoder-decoder network for training. Use the Adam optimizer, set the learning rate to 0.0002, and set the batch size to 16. The final network output is the UV position map. Then use a simple convolutional neural network to reconstruct the three-dimensional shape of the human face from the UV position map, but at this time, texture details have not been added to its surface.
[0064] S7: Construct a conditional generative adversarial network
[0065] The main inspiration of GAN comes from the idea of zero-sum games in game theory. Applied to neural networks in deep learning, it is through continuous games between the generator G and the discriminator D. G, as the generator, takes a random noise x as input in the original paper and generates an image through this random noise; D, as the discriminator, needs to judge whether the picture is real. Its input is the picture. If y is the labeled picture, the output is the probability that the picture is real. During the training process, G needs to generate real pictures as much as possible to deceive D, while D has to distinguish the authenticity of the pictures generated by G, forming a game process that ultimately reaches the Nash equilibrium point.
[0066]
[0067] Among them, x is the input noise whose range is the probability distribution p z (x), y is the real picture whose range is the real data p data(y), G represents the generator, and D represents the generator.
[0068] In the constructed GAN, the incomplete UV texture map is used to replace the noise input generator. The generator part adopts an encoder-decoder structure. The encoder part is 8 convolutional layers, and the decoder part is 8 deconvolutional layers. Their convolution kernels are all 4 and the stride is 2. The discriminator part uses a 4-layer convolutional layer to connect the input image with the label to obtain their features. However, when training GAN, problems such as instability, gradient disappearance, and mode collapse often occur. Therefore, according to the CGAN idea, GAN is improved to obtain better results. A deep convolutional neural adversarial network with certain result constraints. Compared with GAN, its generator uses fractional strided convolution, and the discriminator uses strided convolution to replace all pooling layers. The fully connected hidden layer is removed for a deeper architecture. ReLU is used as the activation function in the generator, and LeakyReLU is used as the activation function in the discriminator.
[0069] S8: Constructing adversarial loss function
[0070] In order to improve the realism and texture details of the generated UV texture map, multiple loss functions are set to take the weighted sum, namely the pixel level loss function L1, the face feature level loss function L f , symmetric loss function L sym And the adversarial loss function L d .
[0071] The pixel-level loss function L1 uses mean square error to make the generated image close to the target image at the pixel level, and adds a mask P to increase the weight of the eyes, nose and mouth. As a key part to improve performance, it will be given a higher weight than other loss functions.
[0072]
[0073] Where W and H are the width and height of the image, j represents the pixel position on the width, k represents the pixel position on the length, and x and y are the input image and the real image, respectively.
[0074] The deepface module is introduced and represented by F. The features of the face in the generated image and the label are obtained and compared. The face contour, eye, nose and mouth positions are determined from a global perspective and the different features of each person in the data are maintained, so that the output result will not be an average similar UV texture map.
[0075]
[0076] Among them, N represents the number of features obtained, F represents the result obtained by inputting the image into the deepface module, and the x and y distributions represent the input image and the real image.
[0077] Due to the symmetric characteristics of the human face, a symmetric loss function is adopted. Utilizing the prior knowledge of the visible part can effectively solve the self-occlusion problem caused by large poses and complete the parts that cannot be seen in a single image. In reality, self-occlusion may cause the left or right side to be invisible, so both left-occluded and right-occluded images exist in the training input. However, using the symmetric loss function may misjudge the different brightness levels on both sides of the face caused by lighting. Therefore, it is necessary to adjust the weight ratio with other loss functions and not assign a large weight to the symmetric loss function.
[0078]
[0079] Among them, W and H are the width and height of the image respectively, j represents the pixel position in the width direction, k represents the pixel position in the length direction, and x and y are the input image and the real image respectively.
[0080] Calculate using the adversarial loss function to determine the loss value of the generated face image from the label, which is beneficial to improving the realism of the generated image and reducing the blur degree.
[0081]
[0082] Among them, G represents the generator, D represents the discriminator, W and H are the width and height of the image respectively, j represents the pixel position in the width direction, k represents the pixel position in the length direction, and x is the input image.
[0083] The final generation loss function takes the weighted sum of the above loss functions.
[0084] L g = λ1L1 + λ f L f + λ sym L sym + λ d L d (7)
[0085] Among them, L1 is the pixel-level loss function, λ1 is the weight of the pixel-level loss function, L f is the face feature-level loss function, λ f is the weight of the face feature-level loss function, L sym is the symmetric loss function, λ sym is the weight of the symmetric loss function, L d is the adversarial loss function, λ d is the weight of the adversarial loss function.
[0086] S9: Train the conditional adversarial generation network
[0087] Using the complete UV texture map obtained by scanning as the generation target, using the incomplete UV texture map to replace the noise and input it into the network for training, and using the Adam optimizer with the learning rate set to 0.0002, the obtained model can complete the incomplete part in the UV texture map.
[0088] S10: Fit the generated three-dimensional face shape model with the UV texture map to obtain the final complete colored three-dimensional face model.
[0089] The single-image large-pose three-dimensional colored face reconstruction method based on the UV position map and CGAN in this embodiment solves the self-occlusion problem that often occurs in the reconstruction of the three-dimensional face model from large-pose single-face images, and realizes the generation of a complete and realistic three-dimensional face model directly from a single two-dimensional image. It overcomes the deficiencies of the decrease in accuracy or even the inability to correctly recognize the face during single-image large-pose face reconstruction. Therefore, it can be used to solve the problem of the decrease in recognition accuracy caused by large poses in face recognition, or can be used to generate multi-angle face images from single-face images to increase experimental data and reduce complex data acquisition.
Claims
1. A single-image large-pose three-dimensional color face reconstruction method based on a UV position map and CGAN, characterized in that The method comprises the following steps: S1: Collecting Data Use active vision method to obtain a large number of 3D models of human faces, and take photos with the frontal face as 0°, 5° as the step length, and the rotation range of [-90°, 90°], classify them and save them in the set format; S2: Generate UV position map The 3D model uses the (X, Y, Z) coordinate system. Its structure is a polygonal model with point cloud coordinates as vertices. The UV coordinate system matches the vertices of the polygon with the pixels on the 2D image. The UV coordinates define the position information of each point on the image. These points are interconnected with the 3D model. The image smoothing interpolation is performed in the gaps between points, so that the UV texture map can be mapped to the 3D model. The 3D point cloud data is recorded in the 2D image by constructing the UV position map. S3: Generate UV texture map After obtaining the UV position map, a bilinear sampler is used to resample the vertices of the 3D model and their related UV coordinates, and the color texture information in the photo is rendered into the UV position map to obtain the required UV texture map; the incomplete parts of the UV texture map are filled with black; S4: Constructing the encoder-decoder network The 256*256*3 image input in the encoder part first passes through a convolution layer with a kernel of 4, and then uses 10 residual blocks to obtain its 8*8*512 features. In the decoder part, 17 deconvolution predictions are used to generate a 256*256*3 UV position map; S5: Construct loss function Use a mask to highlight important parts, set different grayscale values for different parts, and then change the weights after normalization; Among them, (u, v) represents a point in the UV coordinate system, and P(u, v) represents the position of the point in the real target image. represents the position of the point generated by the network, and W(u, v) represents the weight assigned to the corresponding point; S6: Training the encoder-decoder network The UV position map is used as the target value, and the face photos at various angles are input into the encoder-decoder network for training using the Adam optimizer. The final network output is the UV position map. Then a simple convolutional neural network is used to reconstruct the 3D shape of the face from the UV position map. S7: Constructing Conditional Generative Adversarial Networks The generator G and the discriminator D continuously engage in a game. As the generator, G takes a random noise as input and generates an image through this random noise. As the discriminator, D needs to determine whether the picture is real. Its input is the picture. During the training process, G needs to try its best to generate real pictures to deceive D, while D has to distinguish the authenticity of the pictures generated by G, forming a game process that ultimately reaches the Nash equilibrium point; wherein, is random noise, whose range is the probability distribution p z ( ), y is the real picture, whose range is the real data p data (y), G represents the generator, and D represents the discriminator; S8: Constructing adversarial loss function Set multiple loss functions to take the weighted sum, which are the pixel-level loss function L1, the face feature-level loss function L f , the symmetry loss function L sym and the adversarial loss function L d ; S9: Training Conditional Generative Adversarial Networks The scanned complete UV texture map is used as the generation target, and the incomplete UV texture map is used instead of the noise input network for training. The Adam optimizer is used, and the learning rate is set to 0.0002. The trained conditional adversarial generation network completes the incomplete part of the UV texture map. S10: fitting the generated three-dimensional shape of the human face with the UV texture map to obtain a final complete color three-dimensional human face model; In step S8, the pixel-level loss function L1 adopts mean square error to make the generated image close to the target image at the pixel level, and adds a mask P to increase the weight of the eyes, nose and mouth. As a key part to improve performance, it will be given a higher weight than other loss functions; Where W and H are the width and height of the image, j represents the pixel position on the width, k represents the pixel position on the length, and x and y are the input image and the real image, respectively. The deepface module is introduced to obtain and compare the features of the generated image and the face in the label, determine the position of the eyes, nose and mouth of the face contour from a global perspective, and save the different features of each person, so that the output result will not be an average similar UV texture map; Among them, N represents the number of features obtained, F represents the result obtained by inputting the image into the deepface module, and the distribution of x and y represents the input image and the real image; Since the face has a symmetrical characteristic, the symmetric loss function is used. By using the prior knowledge of the visible part, it can effectively solve the self-occlusion problem caused by large postures and complete the parts that cannot be seen in a single image. Where W and H are the width and height of the image, j represents the pixel position on the width, k represents the pixel position on the length, and x and y are the input image and the real image, respectively. Using the adversarial loss function to calculate the loss value of the generated face image from the label, which helps to improve the realism of the generated image and reduce the blur; Where G represents the generator, D represents the discriminator, W and H represent the width and height of the image respectively, j represents the pixel position on the width, k represents the pixel position on the length, and x represents the input image; The final generated loss function is the weighted sum of the above loss functions: L g = λ1L1 + λ f L f + λ sym L sym + λ d L d (7) Among them, L1 is the pixel-level loss function, λ1 is the weight of the pixel-level loss function, L f is the face feature-level loss function, λ f is the weight of the face feature-level loss function, L sym is the symmetry loss function, λ sym is the weight of the symmetry loss function, L d is the adversarial loss function, λ d is the weight of the adversarial loss function.
2. The single-image large-pose three-dimensional color face reconstruction method based on the UV position map and CGAN according to claim 1, characterized in that In step S6, the learning rate is set to 0.0002 and the batch size is set to 16.
3. The method for single-image large-pose three-dimensional color face reconstruction based on UV position map and CGAN according to claim 1 or 2, characterized in that In step S7, the noise input generator is replaced by an incomplete UV texture map in the constructed CGAN, wherein the generator part adopts an encoder-decoder structure, the encoder part is 8 convolutional layers, and the decoder part is 8 deconvolutional layers, their convolution kernels are all 4, and the step size is 2; the discriminator part adopts a 4-layer convolutional layer, and the input image is connected with the label to obtain their features; ReLU is used as the activation function in the generator, and LeakyReLU is used as the activation function in the discriminator.
Citation Information
Patent Citations
Multi-metric three-dimensional face reconstruction method based on parameterized model and position map
CN112184912A
Real-time avatars using dynamic textures
US20200051303A1