A system and method for synthesizing human faces based on line drawings
Patent Information
- Application Number
- CN202310837204.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-07-10
AI Technical Summary
[0009]2.1)2017年发表于ACM Transactions On Graphics的“DeepSketch2Face:A DeepLearning Based Sketching System for 3D Face and Caricature Modeling”提出了使用线稿合成漫画风格的三维人脸的方法,但由于其基于人脸参数化网格(Mesh网格)模型来合成人脸,真实感较差,且不具备眼睛和头发等区域
Smart Images

Figure CN116883568B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer graphics and computer vision, specifically to the field of hardware acceleration for neural network model computation, and more specifically, to a system and method for synthesizing human faces based on line drawings. Background Technology
[0002] 3D face synthesis technology is an important research topic in the field of computer graphics. Synthesizing high-quality geometric and textured face models has wide applications in digital character design, virtual video conferencing, and other areas. Existing 3D face synthesis techniques are typically based on mesh representations with added texture maps. To synthesize highly realistic 3D faces, professional modelers need to use specialized 3D modeling software such as Maya, ZBrush, or NVIDIA Omniverse, which is extremely time- and financially costly. Furthermore, obtaining highly realistic face images from any viewpoint requires complex rendering software and algorithms, further increasing the cost of face synthesis and limiting the application scenarios of these methods.
[0003] With the development of deep learning, Neural Radiance Fields (NeRF) have attracted widespread attention as a novel 3D representation. Synthesizing 3D models based on NeRF requires multiple perspective 2D images. The neural network predicts the color and density of spatial points and synthesizes images from any perspective using volumetric rendering techniques. Compared to traditional network representations, this method has lower modeling costs and produces more realistic rendering results. However, this method can only reconstruct existing objects or scenes based on input images and is difficult to intuitively edit details. While existing methods using generative adversarial networks can synthesize high-quality 3D faces based on NeRF, these methods also struggle to achieve fine-grained control over the detailed structure of the face.
[0004] In recent years, some researchers have also been working on face synthesis technology solutions. Some illustrative technical solutions are as follows:
[0005] 1) In the area of 2D image generation and editing, some related work on compositing and editing facial images and videos at the 2D image level based on line drawings includes the following:
[0006] 1.1) The paper "DeepFaceDrawing: DeepGeneration of Face Images from Sketches," published in ACM Transactions on Graphics in 2020, proposed a method for synthesizing 2D face images based on line drawings. However, due to the use of manifold projection algorithms, this method suffers from poor synthesis quality for personalized features such as hairstyles. Furthermore, this method is designed for the image domain and cannot be extended to 3D face synthesis.
[0007] 1.2) The paper "SketchEdit: Mask-Free Local Image Manipulation with Partial Sketches," published in the Proceedings of the IEEE / CVF International Conference on Computer Vision in 2022, is based on line art for editing 2D images. However, because it does not optimize for hand-drawn line art, the method has poor robustness. Furthermore, this method is only designed for 2D images and cannot be applied to 3D face editing.
[0008] 2) Regarding the generation and editing of 3D faces, some related work is as follows:
[0009] 2.1) The paper “DeepSketch2Face: A DeepLearning Based Sketching System for 3D Face and Caricature Modeling” published in ACM Transactions on Graphics in 2017 proposed a method for synthesizing cartoon-style 3D faces using line art. However, because it synthesizes faces based on a parametric mesh model of the face, the realism is poor and it does not have areas such as eyes and hair.
[0010] 2.2) The paper “IDE-3D: Interactive Disentangled Editing for High-Resolution 3D-aware Portrait Synthesis” published in ACM Transactions on Graphics in 2022 uses semantic masks to edit 3D faces using NeRF. However, semantic masks and other methods have difficulty controlling details such as hair direction and wrinkles. At the same time, since this method only considers single-view information, it cannot maintain the 3D consistency of non-edited areas during the editing process.
[0011] 2.3) The paper “NeRF Face Editing: Disentangled Face Editing in Neural Radiance Fields” published in SIGGRAPH Asia in 2022 also uses semantic mask to generate and edit 3D faces using NeRF. However, for the generation problem, similar to the technology in (2.2), this method requires users to draw a mask map with semantic labels from scratch. The steps are cumbersome and time-consuming. At the same time, it requires high drawing skills and is difficult to meet the needs of amateur users.
[0012] It is evident that existing technologies suffer from the following shortcomings that urgently require improvement:
[0013] (1) Existing technology that uses two-dimensional line drawings to synthesize faces can only be applied to the level of two-dimensional image editing. It only uses the fixed angle corresponding to the two-dimensional line drawings to synthesize faces, which makes it difficult for users to synthesize faces conveniently and flexibly according to their needs, resulting in a poor user experience.
[0014] (2) For existing technologies for generating and editing 3D faces, either the skin of the face can only be synthesized using line drawings at fixed angles and parametric mesh models of the face, resulting in poor realism; or specific masks can be used to directly edit the 3D face NeRF. Drawing masks on the 3D face NeRF is not only cumbersome, but also requires users to have strong professional skills.
[0015] (3) Existing technologies make it difficult to maintain the three-dimensional consistency of non-editing areas during the editing process. Summary of the Invention
[0016] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a system and method for synthesizing human faces based on line drawings.
[0017] The objective of this invention is achieved through the following technical solution:
[0018] According to a first aspect of the present invention, a system for synthesizing a human face based on a line drawing is provided, comprising: a three-plane-based implicit code synthesis module configured to: acquire a line drawing and an appearance reference image in pairs; inject appearance information into the line drawing using the appearance reference image to synthesize a three-plane feature corresponding to the line drawing; project the three-plane feature into the latent space of an EG3D model to obtain an implicit code corresponding to the line drawing, wherein the line drawing is a two-dimensional line drawing containing a character's face, and the appearance reference image is a color image containing a character's face; a line drawing rendering module configured to: generate a corresponding three-dimensional line drawing of a human face based on the implicit code; an image rendering module configured to: generate a corresponding color three-dimensional human face based on the implicit code, which uses a pre-trained EG3D model; and a user operation module configured to: acquire a first line drawing and a first appearance reference image, and use the implicit code synthesis module, the line drawing rendering module, and the image rendering module to obtain a first three-plane feature, a first implicit code, a first three-dimensional line drawing of a human face, and a first three-plane feature corresponding to the first line drawing. A color 3D face; providing users with an interface to select different perspectives to observe the first 3D line drawing of a face; providing a 2D editable custom perspective line drawing based on the user-selected perspective and the first 3D line drawing of the face; using the image rendering module to obtain a second appearance reference image; obtaining a second line drawing obtained after the user edits the face on the custom perspective line drawing and a mask used to distinguish the edited area and the non-edited area; based on the second line drawing and the second appearance reference image, using the implicit code synthesis module to obtain the second three-plane features corresponding to the second line drawing, and obtaining the fusion implicit code obtained by projecting the result of fusing the first three-plane features and the second three-plane features according to the mask, wherein the mask is used to indicate changes to the three-plane features of the edited area and suppress changes to the three-plane features of the non-edited area; based on the fusion implicit code, using the line drawing rendering module and the image rendering module, generating the second 3D line drawing of the face and the second color 3D face obtained after the user edits the face.
[0019] Optionally, the three-plane-based implicit code synthesis module includes a trained line drawing three-plane prediction network, which includes an appearance encoder, a transformation network, and a convolutional network. The trained line drawing three-plane prediction network is trained as follows: a first training set is obtained, which includes multiple first samples, each first sample including a set of line drawings, appearance reference images, and three-plane feature ground values for training; the preset line drawing three-plane prediction network is iteratively trained using the multiple first samples to obtain the trained line drawing three-plane prediction network, wherein: the appearance editor is used to extract appearance information based on the appearance reference image corresponding to the input line drawing, the appearance information including the color and texture of each part of the face; the transformation network is used to extract color feature maps based on the line drawing and the corresponding appearance information; the convolutional network is used to output the three-plane features corresponding to the line drawing based on the voxel features of the three planes constructed from the color feature maps; the training loss is calculated using a preset loss function and the gradient is calculated and backpropagated to update the trainable parameters of the line drawing three-plane prediction network.
[0020] Optionally, the loss of the training line drawing three-plane prediction network can be calculated according to the following loss function:
[0021] L(E c G c G v )=β1L1(p s ,p)+β2L1(I s ,I t )+β3L VGG (I s ,I t )
[0022] Among them, L1(p s ,p) represents the output three-plane feature p s The L1 distance between the L1 distance and the corresponding three-plane feature ground value p, L1(I s ,I t ) represents the image I generated by the hidden code corresponding to the output three-plane features. s and the corresponding image truth value I t The L1 distance between them, L VGG (I s ,I t ) represents image I s and the corresponding image truth value I t Perceptual distance between them, image ground truth I t Using the appearance reference image corresponding to the line drawing, β1, β2, and β3 are respectively L1(p s ,p),L1(I s ,I t L VGG (I s ,I tPreset weighting coefficients.
[0023] Optionally, the three-plane-based latent code synthesis module includes a trained 2D encoder, which is the encoder in a trained pSp framework image autoencoder. The trained pSp framework image autoencoder is trained as follows: the pSp framework image autoencoder is iteratively trained using the multiple first samples to obtain the trained image autoencoder, wherein: the encoder of the image autoencoder is used to project the input three-plane features into latent codes in the latent space of EG3D; the decoder of the image autoencoder is used to reconstruct the three-plane features based on the latent codes projected by the encoder; the loss for training the image autoencoder is determined according to the loss function of the pSp framework; the gradient is calculated based on the loss of the image autoencoder and backpropagation is used to update the trainable parameters of the image autoencoder.
[0024] Optionally, the line art rendering module is implemented using an EG3D model that converts the output RGB three channels to a single channel. The line art rendering module includes a StyleGAN backbone, a low-resolution line art decoder with a single-channel output, and a super-resolution decoder with a single-channel output.
[0025] Optionally, the line art rendering module is trained as follows: a second training set is obtained, which includes multiple second samples. Each second sample includes a training latent code, a line art ground value at a first resolution corresponding to the latent code, and a line art ground value at a second resolution obtained by downsampling from the line art ground value at the first resolution. The low-resolution line art decoder and super-resolution decoder of the line art rendering module are trained using the second training set to generate line art of the corresponding resolution, wherein: the StyleGAN backbone is used to convert the latent code into three-plane features; the low-resolution line art decoder is used to generate color features and density information of sampling points in space based on the converted three-plane features, and uses volume rendering combined with camera parameters to generate multi-channel feature maps and line art at the second resolution; the super-resolution line art decoder is used to generate line art at the first resolution based on the multi-channel feature maps and the line art at the second resolution; the sub-loss calculated based on the reconstruction loss function of the line art ground value plus the sub-loss calculated based on the regularization loss function used to constrain the three-dimensional consistency of the line art is used to obtain the total loss, the gradient is calculated, and the trainable parameters of the low-resolution line art decoder and the super-resolution line art decoder are updated by backpropagation.
[0026] Optionally, the total loss during training of the line art rendering module is determined as follows:
[0027]
[0028] Among them, L recon L represents the sub-loss calculated using the reconstruction loss function. viewThe sub-loss calculated by the regularized loss function, L1(S′,S GT () represents the generated line drawing S′ at the first resolution and the true value S of the line drawing at the first resolution. GT The L1 distance between them, L1(S′) raw ,S raw ) represents the generated second-resolution line drawing S r ′ aw True value S of the second resolution line drawing raw The L1 distance between them, L VGG (S′,S GT ) represents the line art S′ and the line art truth value S. GT The perceived distance between them, L VGG (S′ raw ,S raw ) represents the line drawing S′ raw And line art truth value S raw The perceptual distance between them, |C| represents the number of pixels in the set of pixels randomly sampled from the line art, S′[i,j] represents the pixel value at coordinates [i,j] in the line art S′, Render(r ij This represents the true value S of the line drawing at the second resolution, obtained using a ray projection method. raw The light rays obtained during volume rendering ij The rendered pixel values, ||·||1 represents the L1 distance, and α1, α2, α3, α4, and α5 represent the weight coefficients of the corresponding terms.
[0029] Optionally, the fusion hidden code is iteratively optimized multiple times using a preset optimization method to make the line drawing generated by the optimized fusion hidden code more compatible with the second line drawing; the second 3D line drawing of the face and the second color 3D face are generated based on the optimized fusion hidden code.
[0030] According to a second aspect of the present invention, a method for generating and editing a face based on the system described in the first aspect is provided. The method includes: providing a function to generate a three-dimensional face based on a line drawing, including: acquiring a first line drawing and a first appearance reference image, and using a hidden code synthesis module, a line drawing rendering module, and an image rendering module to obtain a first three-plane feature, a first hidden code, a first three-dimensional face line drawing, and a first colored three-dimensional face corresponding to the first line drawing; providing a function to edit a three-dimensional face based on the line drawing, including: providing a user interface for selecting different perspectives to observe the first three-dimensional face line drawing, providing a two-dimensional editable custom perspective line drawing based on the user-selected perspective and the first three-dimensional face line drawing; and acquiring a third face obtained after the user edits the face on the custom perspective line drawing. The second line drawing and a mask for distinguishing between the editing and non-editing areas are used; using the image rendering module, a second appearance reference image is obtained, which is obtained by observing the first colored 3D face from the same perspective as the custom view line drawing; based on the second line drawing and the second appearance reference image, the second three-plane features corresponding to the second line drawing are obtained using the implicit code synthesis module, and a fusion implicit code is obtained by projecting the result of fusing the first three-plane features and the second three-plane features according to the mask, wherein the mask is used to indicate changes to the three-plane features of the editing area and to suppress changes to the three-plane features of the non-editing area; the fusion implicit code is iteratively optimized multiple times using a preset optimization method, and the second 3D face line drawing and the second colored 3D face are generated based on the optimized fusion implicit code.
[0031] Optionally, the method further includes: generating a two-dimensional output line drawing from a user-defined perspective based on the second three-dimensional face line drawing; and / or generating a two-dimensional output color face image from a user-defined perspective based on the second color three-dimensional face.
[0032] Optionally, the method further includes: using the second 3D face line drawing and the second colored 3D face as a new first 3D face line drawing and a new first colored 3D face, providing users with the function of continuously editing 3D faces based on line drawings.
[0033] According to a third aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions; wherein the one or more processors are configured to implement the steps of the method described in the second aspect by executing the executable instructions. Attached Figure Description
[0034] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0035] Figure 1 This is a schematic diagram illustrating the process of generating a 3D face and editing a 3D face using a system based on line drawing synthesis according to an embodiment of the present invention.
[0036] Figure 2 This is a schematic diagram illustrating the principle of a system for synthesizing human faces based on line drawings according to an embodiment of the present invention;
[0037] Figure 3 This is a schematic diagram of the structure of a three-plane prediction network for line drawings according to an embodiment of the present invention;
[0038] Figure 4 This is a schematic diagram of a mask according to an embodiment of the present invention;
[0039] Figure 5 This is a schematic diagram of a user interface according to an embodiment of the present invention;
[0040] Figure 6 This is a schematic diagram illustrating the effect of line drawing-based 3D face generation according to an embodiment of the present invention;
[0041] Figure 7 This is a schematic diagram illustrating the effect of line drawing-based 3D face editing according to an embodiment of the present invention;
[0042] Figure 8 This is a schematic diagram illustrating the effect of multi-view editing based on line art according to an embodiment of the present invention;
[0043] Figure 9 This is a schematic diagram illustrating the effect of first generating a line drawing and then performing detailed editing according to an embodiment of the present invention.
[0044] Figure 10 This is a schematic diagram illustrating the effect of partial appearance editing based on line art according to an embodiment of the present invention;
[0045] Figure 11 This is a schematic diagram illustrating the effect of 3D face synthesis based on line drawings of various styles according to an embodiment of the present invention;
[0046] Figure 12 This is a schematic diagram illustrating the effect of editing a face using a system according to an embodiment of the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.
[0048] As mentioned in the background section, the existing technology has the aforementioned defects (1), (2) and (3). In response, the applicant has made improvements. First, the user only needs to provide a first line drawing and a first appearance reference image, which can generate the corresponding first three-dimensional line drawing of the face and the first color three-dimensional face. The user can freely choose the viewing angle of the first three-dimensional line drawing of the face, and automatically generate the corresponding three-dimensional line drawing of the face and the color three-dimensional face. The user can freely adjust and select the viewing angle. Based on the user's selected viewing angle and the first three-dimensional line drawing of the face, a corresponding two-dimensional editable custom viewing angle line drawing is provided (equivalent to converting the two-dimensional line drawing into a three-dimensional line drawing, and providing an editable line drawing based on the three-dimensional line drawing with another viewing angle that the user can freely define. This allows the user to conveniently and flexibly synthesize faces from various angles as needed when the input is only a two-dimensional image, thereby improving the efficiency of face synthesis and user experience). Second, the user only needs to edit on the line drawing. The difficulty of modifying the line drawing is low, and no excessive professional skills are required. Finally, by setting the mask, the three-dimensional consistency of the non-editing area can be maintained during the editing process, avoiding the significant impact of changes to a certain local area on other areas of the face.
[0049] Before describing the embodiments of the present invention in detail, some of the terms used therein are explained as follows:
[0050] S represents line art (two-dimensional line art);
[0051] p s This represents a three-plane feature, including feature maps of the three planes, respectively. Among them, the three planes refer to the three abstract planes of space, namely XY, YZ and ZX, which are orthogonal to the z-axis, x-axis and y-axis of the three-dimensional coordinate system, respectively;
[0052] I represents an image;
[0053] f c Represents a color feature map;
[0054] E c This represents the appearance encoder of the three-plane prediction network for line drawings;
[0055] G c The transformation network representing the three-plane prediction network for line art;
[0056] G v A convolutional network representing a three-plane prediction network for line art;
[0057] L1 represents the distance L1;
[0058] L VGGRepresents the perceptual distance, which is usually calculated based on the feature map after the input features (feature map or image) are transformed again by the VGG model (such as VGG16 or VGG19);
[0059] I t Represents the true value of an image;
[0060] E proj Indicates a 2D encoder;
[0061] S′ represents the generated line drawing at the first resolution;
[0062] S GT Represents the truth value of the line drawing at the first resolution;
[0063] S r ′ aw This indicates the generated line art at a second resolution, which is smaller than the first resolution.
[0064] S raw Indicates the truth value S based on the line drawing. GT The ground truth value of the line drawing at the second resolution obtained by downsampling;
[0065] |C| represents the number of pixels in the set of pixels randomly sampled from the line art;
[0066] S′[i,j] represents the pixel value at coordinates [i,j] in the generated first-resolution line drawing S′;
[0067] Render(r i,j This represents the true value S of the line drawing at the second resolution, obtained using a ray projection method. raw The light rays obtained during volume rendering ij The rendered pixel values;
[0068] M represents a mask (two-dimensional mask);
[0069] M 3D A 3D mask is represented by M. xy M xz M yz ;
[0070] w indicates a hidden code;
[0071] o i,j It represents the origin of light rays in Euclidean space;
[0072] d i,j Indicates the direction of light rays in Euclidean space;
[0073] D[i,j] represents the depth value of [i,j] in the depth map;
[0074] Sedit This indicates a revised line drawing (corresponding to the second line drawing);
[0075] This represents the fused three-plane features, that is, the result of fusing the first and second three-plane features according to the mask; in the fusion formula, p represents the first three-plane feature. xy Indicates the features of the second and third planes;
[0076] w edit Indicates a fused implicit code;
[0077] ⊙ indicates pixel-by-pixel multiplication;
[0078] R x This represents the image rendering module;
[0079] R s This refers to the line art rendering module;
[0080] N represents the number of sampling points;
[0081] Indicates the non-editing area;
[0082] r(i) represents the i-th sampling point rendered along ray r in the non-editing area;
[0083] This indicates the point-by-point feature calculation process;
[0084] β1, β2, β3, α1, α2, α3, α4, α5, γ1, γ2, γ3 represent the corresponding weight coefficients, which are hyperparameters set by the user.
[0085] See Figure 1 According to one embodiment of the present invention, a system for synthesizing human faces based on line drawings is provided, including a three-plane-based implicit code synthesis module, a line drawing rendering module, an image rendering module, and a user operation module. Figure 1 (Not shown), mask fusion module, and implicit code optimization module. Among them, the three-plane implicit code synthesis module includes a three-plane prediction network and a 2D encoder.
[0086] To more easily understand the overall usage flow of the system in this embodiment of the invention, and to gain a more macroscopic understanding of the technical solution, the following will first combine... Figure 1 This is to simplify the explanation of the process of generating and editing 3D faces. For ease of understanding, the corresponding terms used in the claims are given in parentheses.
[0087] A schematic diagram of the process of generating a 3D human face is as follows: Figure 1 As shown in a, it includes:
[0088] The input line drawing and appearance reference image are input into the line drawing three-plane prediction network, and the output three-plane feature 1 (corresponding to the first three-plane feature);
[0089] Three-plane feature 1 input 2D encoder, output hidden code (corresponding to the first hidden code);
[0090] The hidden code is processed by the line drawing rendering module to generate 3D face line drawing 1 (corresponding to the first 3D face line drawing);
[0091] The hidden code is processed by the image rendering module to generate a color 3D face 1 (corresponding to the first color 3D face);
[0092] Based on the user's selected perspective, line drawing 1 (corresponding to a two-dimensional editable custom perspective line drawing) and image 1 (corresponding to a second appearance reference image) can be generated from the 3D face line drawing 1 and the color 3D face 1.
[0093] The process of editing a 3D face is as follows Figure 1 As shown in b, it includes:
[0094] The user modifies the generated line art 1 to obtain the modified line art (corresponding to the second line art);
[0095] Input the user-modified line art and the generated image 1 into the line art three-plane prediction network, and output three-plane feature 2 (corresponding to the second three-plane feature);
[0096] Input the three-plane feature 1 and the three-plane feature 2 into the mask fusion module. The mask fusion module then edits the mask of the indicated area. Figure 1 (not shown in b) Merge the three-plane features 1 and 2, combining the features of the non-editing area in three-plane feature 1 with the features of the editing area in three-plane feature 2 to obtain the merged three-plane features;
[0097] The fused three-plane features are input into a 2D encoder, and the corresponding hidden code is output (corresponding to the fused hidden code);
[0098] The integrated hidden code input hidden code optimization module, after one or more optimizations, yields the optimized integrated hidden code;
[0099] The optimized fusion hidden code is processed by the line drawing rendering module to generate a 3D face line drawing 2 (corresponding to the second 3D face line drawing);
[0100] The optimized fusion hidden code is processed by the image rendering module to generate a color 3D face 2 (corresponding to the second color 3D face);
[0101] If the user needs (e.g., the user has achieved the desired editing effect), the user can adjust the viewpoint according to their desired angle, and then adjust the viewpoint to the corresponding angle to generate line drawing 2 and image 2; or, if the user is not satisfied with the editing effect, the user can also edit again based on line drawing 2 (and image 2 as an appearance reference image).
[0102] To explain in detail the various sub-components and application scenarios of the aforementioned system, the following sections will provide a detailed description of the three-plane-based steganography synthesis module (three-plane prediction network, 2D encoder), line drawing rendering module, image rendering module, user operation module, mask fusion module, steganography optimization module, and application scenarios.
[0103] I. Three-plane based steganography synthesis module
[0104] According to an embodiment of the present invention, a three-plane-based latent code synthesis module is configured to: acquire line drawings and appearance reference images in pairs; inject appearance information into the line drawings using the appearance reference images to synthesize three-plane features corresponding to the line drawings; project the three-plane features into the latent space of the EG3D model to obtain the latent code corresponding to the line drawings, wherein the line drawings are two-dimensional line drawings containing a character's face, and the appearance reference images are color images containing a character's face.
[0105] The authenticity of a character should be determined based on the user's needs. According to one embodiment of the invention, the character can be a user-created virtual person, such as the protagonist of an animated film, a movie protagonist, or the protagonist of a promotional poster. Alternatively, the character can be a real person. However, if it is a real person, their consent must be obtained when synthesizing their face, and the process must comply with relevant laws and regulations. Alternatively, when synthesizing the face of a real person according to an embodiment of the invention, a watermark or other type of identifier can be added to the synthesized face image to clearly indicate that the face image is synthesized.
[0106] It should be noted that the terms "line art" and "appearance reference image" are used generically here. For example, they could refer to a first line art and a first appearance reference image; or a second line art and a second appearance reference image, etc. In a pair of line art and appearance reference images, the appearance of the appearance reference image does not need to be absolutely identical to the line art in terms of facial contours, perspective, and other visual characteristics. This is because the appearance reference image primarily provides visual information, such as color, texture, lighting, and / or style. This requirement for absolute consistency is also reflected in the second line art and the second appearance reference image, as the second line art is edited by the user on a custom perspective line art corresponding to the second appearance reference image. The user may modify facial features such as hair shape, accessories, eye shape, mouth shape, expression, and eyes, resulting in differences in appearance between the second line art and the second appearance reference image.
[0107] According to one embodiment of the present invention, in the three-plane-based implicit code synthesis module, the synthesis of the three-plane features corresponding to the line drawing can be accomplished using a three-plane prediction network; implicit code generation can be accomplished by a 2D encoder. The three-plane prediction network and the 2D encoder are described below.
[0108] (1A) Three-plane prediction network
[0109] According to one embodiment of the present invention, see Figure 2 The three-plane-based hidden code synthesis module includes a trained line drawing three-plane prediction network, which includes an appearance encoder E c , conversion network G c and convolutional network G v The trained line art three-plane prediction network is trained as follows: A first training set is obtained, comprising multiple first samples, each first sample including a set of line art for training, appearance reference images, and three-plane feature ground values; the preset line art three-plane prediction network is iteratively trained using the multiple first samples to obtain the trained line art three-plane prediction network, wherein: Appearance Editor E c This is used to extract appearance information from the appearance reference image corresponding to the input line drawing, whereby the appearance information includes the color and texture (or color, texture, and / or style) of various parts of the face; the transformation network C c Used to extract color feature maps based on the line drawing and corresponding appearance information; convolutional network G v The voxel features of the three planes constructed based on the color feature map are used to output the three-plane features corresponding to the line drawing; the training loss is calculated using a preset loss function and the gradient backpropagation is used to update the trainable parameters of the line drawing three-plane prediction network. To construct the training set, this embodiment of the invention generates different 3D faces based on a series of randomly sampled hidden codes from a pre-trained EG3D model, and synthesizes the ground truth p of the three-plane features; furthermore, from each 3D face generated by EG3D, multiple camera positions are randomly sampled to synthesize face images as appearance reference images (also supervision data), and the corresponding line drawings are extracted as input data; thus forming multiple first samples to train the line drawing three-plane prediction network. Therefore, the line drawing three-plane prediction network can be trained using a synthesized multi-view dataset. The three-plane generation based on the line drawing includes: given a 2D line drawing and an appearance reference image, using adaptive normalization to fuse the two types of information to obtain a color feature map, assigning color information to the line drawing. Further, voxels are constructed in 3D space, and the features of each spatial point are retrieved from the color feature map after projection onto the 3D coordinates. The voxel features are transformed into two-dimensional feature maps by performing shape transformation, and three-plane features are generated based on 2D convolution.
[0110] According to one embodiment of the present invention, a schematic structure of a three-plane prediction network is as follows: Figure 3 As shown, where, Figure 3 a, Figure 3 b and Figure 3 c represents the appearance editor E. c , conversion network G c and convolutional network G v A schematic structure. Figure 3 In the diagram, the "nx" (e.g., 4x, 5x) on the left side of some convolutional blocks refers to n stacked convolutional blocks of that type; convolutional blocks in the form of A-BxB (e.g., convolution 16-7×7) indicate that the number of channels in the output feature map is A and the kernel size is BxB; convolutional blocks in the form of residual convolutional blocks C (e.g., residual convolutional block 256) indicate that the number of channels in the output feature map is C; BN, IN, and AdaIN are the names of the normalization layers, and ReLU and tanh are the names of the activation functions used. It should be understood that... Figure 3 The structure shown can be adjusted, for example, by changing the size of some convolution kernels, or by adjusting the number of convolution layers, the position of residual convolution blocks, etc., to form other alternative implementation methods.
[0111] According to one embodiment of the present invention, the loss of the training line drawing three-plane prediction network is calculated according to the following loss function:
[0112] L(E c G c G v )=β1L1(p s ,p)+β2L1(I s ,I t )+β3L VGG (I s ,I t )
[0113] Among them, L1(p s ,p) represents the output three-plane feature p s The L1 distance between the L1 distance and the corresponding three-plane feature ground value p, L1(I s ,I t ) represents the image I generated by the hidden code corresponding to the output three-plane features. s and the corresponding image truth value I t The L1 distance between them, L VGG (I s ,I t ) represents image I s and the corresponding image truth value I t Perceptual distance between them, image ground truth I t Using the appearance reference image corresponding to the line drawing, β1, β2, and β3 are respectively L1(p s ,p),L1(I s ,It L VGG (I s ,I t Preset weight coefficients. During training, given a single-viewline drawing as input, the three-plane prediction network synthesizes three-plane features p. s Image I is generated by volumetric rendering using the decoder of a pre-trained EG3D model (employing existing trained models). s This is used to assist in the training of the line drawing three-plane prediction network. These rendered images I s Comparing the input line art with different perspectives t, the three-plane prediction network of the line art is forced to imagine the face from other perspectives to enhance 3D information, resulting in better performance. It should be understood that some formulas in this embodiment or later are preferred implementations, and those skilled in the art can optimize or fine-tune the formulas as needed to obtain other theoretically feasible alternative implementations. For example, if performance impact is not considered, implementers may also delete β2L1(I s ,I t ) and β3L VGG (I s ,I t One or two of the following can be used to obtain other implementation methods. For example, the loss function can be adjusted, such as by changing L1(p)... s ,p) and / or L1(I s ,I t The similarity can be changed to cosine similarity to obtain other implementation methods; other formulas are similar. Preferably, in this embodiment of the invention, β1 = 0.01, β2 = 1.0, and β3 = 0.1 are set. It should be understood that the magnitudes of the weight coefficients in the loss function may be different in different scenarios and can be adjusted as needed according to the implementer's needs or on-site experience. For example, β1 = 0.02, β2 = 0.9, and β3 = 0.09, etc. This embodiment of the invention does not limit this.
[0114] According to one example of the present invention, given a single-viewline line drawing S, a three-plane prediction network for the line drawing is designed to transform S into a three-dimensional three-plane representation p. s Including feature maps The generated structure has the same distribution of the three planes as the EG3D model.
[0115] Since the input line drawing only contains geometric information, but a 3D face has different appearance information, the input sketch S is first converted into a color feature map f. c This involves injecting color, lighting, and texture information. The present invention trains an appearance encoder E. c Appearance information is extracted from reference image I. To integrate geometric and appearance information, a transformation network G is designed. cColor feature maps are generated using Adaptive Instance Normalization (AdaIN):
[0116] f c =G c (S,W c (I))
[0117] In order to extract from the two-dimensional feature map f c This example predicts a 3D three-plane representation and constructs 3D voxel features in the volumetric rendering space: First, the camera's intrinsic and extrinsic parameters are estimated based on the line art, and each point x in the voxel features is projected onto the input image space to obtain 2D coordinates π(x); then, for each point x, the 2D coordinates are retrieved from the color feature map and bivariate interpolation is performed to obtain the corresponding feature vector f. c (π(x)). Based on the above description, a feature voxel with a resolution of 128 was constructed, with a shape of 128×128×128×3, where 3 is the number of channels in the feature map.
[0118] Next, shape transformations are performed on the three-dimensional voxel features along the z, y, and x axes respectively, generating three 128×128×384 feature maps, denoted as V. xy V xz V yz The above feature maps are concatenated along the feature channels and then processed using a 2D convolutional network G. v The image was upsampled and converted into a 256×256×96 line drawing feature map. This feature map was then segmented by feature channels, ultimately forming three 32-channel feature planes. Abbreviated as three-plane feature p s :
[0119] p s =G v (V xy V xz V yz )
[0120] As can be seen, although the three-plane prediction network analyzes 3D information in detail, it only uses 2D convolutions to improve memory and time efficiency.
[0121] (1B) 2D encoder
[0122] According to one embodiment of the present invention, the three-plane-based latent code synthesis module includes a trained 2D encoder, which is the encoder in a trained pSp framework (refer to the paper Encoding in Style: aStyleGAN Encoder for Image-to-Image Translation) image autoencoder. The trained pSp framework image autoencoder is trained as follows: the pSp framework image autoencoder is iteratively trained using the plurality of first samples to obtain the trained image autoencoder, wherein: the encoder of the image autoencoder projects the input three-plane features into latent codes in the latent space of EG3D; the decoder of the image autoencoder reconstructs the three-plane features based on the latent codes projected by the encoder; the loss for training the image autoencoder is determined according to the loss function of the pSp framework; the gradient is calculated based on the loss of the image autoencoder, and the trainable parameters of the image autoencoder are updated by backpropagation. To train the 2D encoder E... proj This embodiment of the invention uses the same training strategy as the pSp method. The training loss function includes pixel L2 distance, perceptual distance, identity distance, and regularization constraints. A step-by-step training strategy can be used, first training the line drawing three-plane prediction network, and then training the 2D encoder E. proj .
[0123] II. Line Art Rendering Module
[0124] According to one embodiment of the present invention, the line art rendering module is implemented by converting the output RGB three channels into a single channel EG3D model. The line art rendering module includes a StyleGAN backbone, a low-resolution line art decoder with a single channel output, and a super-resolution decoder with a single channel output.
[0125] According to one embodiment of the present invention, the line art rendering module is trained as follows: a second training set is obtained, which includes multiple second samples, each second sample including a training occultation, a line art ground truth value at a first resolution corresponding to the occultation, and a line art ground truth value at a second resolution obtained by downsampling the line art ground truth value at the first resolution; the low-resolution line art decoder and super-resolution decoder of the line art rendering module are trained using the second training set to generate line art of corresponding resolution, wherein: the StyleGAN backbone is used to convert the occultation into three-plane features; the low-resolution line art decoder is used to generate line art of corresponding resolution based on the converted three-plane features. The system generates color features and density information of sampling points in the spatial region; and uses volume rendering combined with camera parameters to generate multi-channel feature maps and line art at a second resolution; the super-resolution line art decoder is used to generate a first resolution line art based on the multi-channel feature maps and the second resolution line art; the total loss is obtained by adding the sub-loss calculated by the reconstruction loss function of the EG3D model itself to the sub-loss calculated by the regularization loss function used to constrain the three-dimensional consistency of the line art, the gradient is calculated and backpropagated to update the trainable parameters of the low-resolution line art decoder and the super-resolution line art decoder (the parameters of the StyleGAN backbone are not updated).
[0126] Preferably, since the image rendering module uses an existing pre-trained EG3D model, and the StyleGAN backbone of the line art rendering module in this embodiment of the invention also uses the existing pre-trained EG3D model's StyleGAN backbone, and the parameters of the StyleGAN backbone are not updated during training, in order to reduce the number and size of system parameters, the StyleGAN backbone of the line art rendering module can be shared with the image rendering module, or in other words, the StyleGAN backbone of the line art rendering module is implemented by calling the StyleGAN backbone of the image rendering module.
[0127] In other words, an additional line art generation branch is added to the pre-trained EG3D model to generate line art corresponding to the implicit code in any view. The line art rendering branch and the original image rendering branch of the pre-trained EG3D model share the same StyleGAN backbone network, but have different low-resolution decoders (corresponding to the low-resolution line art decoder) and super-resolution modules (corresponding to the super-resolution line art decoder): This invention uses the new line art decoder to decode the features of spatial points into line art features, and uses the density of the image branch to generate a low-resolution line art feature map based on volume rendering. The first channel of the line art feature map corresponds to the low-resolution line art (128×128), denoted as S. r ′ aw To obtain high-resolution results, a line art super-resolution module similar to the original super-resolution module in the EG3D model was used to synthesize the final arbitrary viewpoint line art S′.
[0128] According to one embodiment of the present invention, the total loss during training of the line art rendering module is determined as follows:
[0129]
[0130] Among them, L recon L represents the sub-loss calculated using the reconstruction loss function. view The sub-loss calculated by the regularized loss function, L1(S′,S GT () represents the generated line drawing S′ at the first resolution and the true value S of the line drawing at the first resolution. GT The L1 distance between them, L1(S′) raw ,S raw ) represents the generated second-resolution line drawing S r ′ aw True value S of the second resolution line drawing raw The L1 distance between them, L VGG (S′,S GT ) represents the line art S′ and the line art truth value S. GT The perceived distance between them, L VGG (S′ raw ,S raw ) represents the line drawing S′ raw And line art truth value S raw The perceptual distance between them, |C| represents the number of pixels in the set of pixels randomly sampled from the line art, S′[i,j] represents the pixel value at coordinates [i,j] in the line art S′, Render(r ij This represents the true value S of the line drawing at the second resolution, obtained using a ray projection method. raw The light rays obtained during volume rendering ij The rendered pixel values, ||·||1 represent the L1 distance, and α1, α2, α3, α4, and α5 represent the weight coefficients of the corresponding terms. Where L... recon ==α1L1(S′,S GT )+α2L1(S r ′ aw ,S raw )+α3L VGG (S′,S GT )+α4L VGG (S′ raw ,S raw ),
[0131] According to one example of the present invention, in order to train the line art generation branch, a pre-trained pix2pixHD is used to convert a face image into a ground truth line art image S. GTTraining data is constructed, but the ground truth is not three-dimensionally consistent. This invention designs a training loss to ensure consistency of line drawings from different perspectives, synthesizing high-quality structures. The loss function first uses a reconstruction loss to match the original line drawing distribution:
[0132] L recon =α1L1(S′,S GT )+α2L1(S′ raw ,S raw )+α3L VGG (S′,S GT )
[0133] +α4L VGG (S′ raw ,S raw )
[0134] Among them, S raw Indicates the truth value S of the line art GT The structure after downsampling. In this example, we set α1 = α2 = 3.0, α3 = α4 = 4.0.
[0135] This example uses a regularization term to constrain the three-dimensional consistency of the line art. In the training data, the line art truth value S... GT Since the perspective is inconsistent, the perspective consistency property of volume rendering is used to constrain the final rendered line art result. This loss term compares the sampled pixels in the rendered line art result with the pixels generated by NeRF:
[0136]
[0137] Where α5 = 3.0 and |C| = 8192.
[0138] The final optimized loss function is:
[0139] L sketch =L recon +L view
[0140] During the training process of line drawing generation from any viewpoint, the network weights of the StyleGAN backbone, neural renderer, and image generation branch of the pre-trained EG3D model are fixed, and only the network weights of the line drawing decoder and line drawing super-resolution module are updated.
[0141] It should be understood that, without considering the performance impact, implementers can also adjust the total loss when training the line art rendering module to obtain other results, such as letting L... sketch =L recon .
[0142] III. Image Rendering Module
[0143] According to one embodiment of the present invention, the image rendering module is configured to generate a corresponding color three-dimensional face based on the hidden code, which uses a pre-trained EG3D model.
[0144] According to one embodiment of the present invention, based on the fusion implicit code, an image rendering module is used to generate a second color 3D face obtained after the user edits the face. Optionally, the color 3D face can be directly generated based on the initially obtained fusion implicit code.
[0145] In some cases, directly generating a color 3D face from the initially obtained fusion implicit code may yield unsatisfactory results. Further optimization is possible to achieve better results. According to one embodiment of the present invention, the image rendering module is configured to generate a second 3D face line drawing and a second color 3D face based on the optimized fusion implicit code. The fusion implicit code is iteratively optimized multiple times using a preset optimization method to ensure a better match between the line drawing generated by the optimized fusion implicit code and the second line drawing. The optimization process of the fusion implicit code will be further explained in the subsequent implicit code optimization module section.
[0146] IV. User Operation Module
[0147] According to an embodiment of the present invention, a user operation module is configured to: acquire a first line drawing and a first appearance reference image, and use a steganography synthesis module, a line drawing rendering module, and an image rendering module to obtain a first three-plane feature, a first steganography, a first three-dimensional face line drawing, and a first colored three-dimensional face corresponding to the first line drawing; provide a user interface for selecting different perspectives to observe the first three-dimensional face line drawing, and provide a two-dimensional editable custom perspective line drawing based on the user-selected perspective and the first three-dimensional face line drawing; acquire a second line drawing obtained after the user edits the face on the custom perspective line drawing and a mask for distinguishing the edited area and the non-edited area; and use the image rendering module to... A second appearance reference image is obtained, which is obtained by observing the first color 3D face from the perspective of the custom viewpoint line drawing. Based on the second line drawing and the second appearance reference image, the second three-plane features corresponding to the second line drawing are obtained using the implicit code synthesis module, and the fusion implicit code is obtained by projecting the result of fusing the first three-plane features and the second three-plane features according to the mask. The mask is used to indicate changes to the three-plane features of the editing area and to suppress changes to the three-plane features of the non-editing area. Based on the fusion implicit code, the second 3D line drawing of the face and the second color 3D face are generated after the user edits the face using the line drawing rendering module and the image rendering module.
[0148] V. Mask Fusion Module
[0149] According to one embodiment of the present invention, the mask fusion module is configured to fuse the first three-plane features and the second three-plane features according to the mask to obtain the fused three-plane features.
[0150] The following is a schematic illustration of the fusion process performed by the mask fusion module:
[0151] After the user performs editing operations through the interface of this invention, the two-dimensional mask M indicates the area to be edited (editing area). For the implicit code w to be edited, the image generation branch based on the EG3D model obtains the original three-plane representation (p xy ,p xz ,p yz And depth map D. Based on depth map D, each pixel M in the two-dimensional mask is back-projected to a 3D position in Euclidean space to form a set of 3D points, representing the three-dimensional mask.
[0152] Indicative, such as Figure 4 For a ray located at pixel (i,j) in the depth map, the origin o of the ray in Euclidean space can be obtained. i,j and direction d i,j And further obtain the corresponding 3D position. i,j +d i,j ·D[i,j]. Based on the above method, all pixels in the 2D mask M are converted into 3D positions based on the depth map, forming a set of 3D points M. 3D :
[0153]
[0154] The set of 3D points is further projected onto a three-plane space to synthesize a set of three projected 2D points. D[i,j] represents the depth value at [i,j] in the depth map, used to correlate with the origin of the ray. i,j and direction d i,j The points are projected back together to show their 3D positions.
[0155] Next, based on the modified line art S edit (Corresponding to the second line drawing), using the line drawing three-plane prediction network, the corresponding three planes can be obtained. By obtaining a new 3D face and a new depth map, another set of 2D points on three planes can be obtained. By merging the two sets of 2D points on each plane, a guided 3D fusion mask is obtained. This mask is derived by taking the union of the 2D masks on the three planes, denoted as M. xy M xz M yz Furthermore, the projected 2D mask undergoes morphological dilation to obtain smooth boundaries and fill internal holes, which are caused by using discrete sampling points to represent the 3D mask.
[0156] The original three planes (first three-plane features) and the three planes generated from the line drawing (second three-plane features) are fused separately on different planes. Taking the xy plane as an example, the fusion is based on the following process:
[0157]
[0158] The above formula is explained using the xy plane. The other two planes are handled in the same way, resulting in a new fused three-plane feature representation.
[0159] It's important to note that swapping and fusing one region of a three-plane feature affects the entire orthogonal 3D cylindrical space; therefore, the fused three-plane feature cannot be directly used for volume rendering. To address this issue, a 2D encoder, E, can be used. proj The fused three-plane features are then further projected into the W+ space of the EG3D model:
[0160]
[0161] Thus, the fusion hidden code is obtained.
[0162] VI. Hidden Code Optimization Module
[0163] According to one embodiment of the present invention, the steganography optimization module is configured to: perform multiple iterative optimizations on the fused steganography using a preset optimization method, so that the line drawing generated by the optimized fused steganography matches the second line drawing better. The fused steganography is updated by calculating the loss using a preset target loss function during the multiple iterative optimizations, with the aim of minimizing the loss calculated by the target loss function.
[0164] To ensure better line drawing correspondence while preserving the characteristics of the original image, this embodiment of the invention further refines w. edit Optimization is then performed. The following is a schematic explanation of the optimization and fusion of the hidden code:
[0165] Given the image rendering branch R x and the line art rendering branch R s Optimize the following sub-loss terms:
[0166] L edit =L VGG (R s (w edit )⊙M,S⊙M)
[0167]
[0168] in, S represents the non-edited region, ⊙ represents pixel-by-pixel multiplication, and w is the implicit code of the original input. The above formula constrains the edited region to be the same as S, while other regions retain their original features;
[0169] The aforementioned sub-loss terms can solve 2D image editing problems, but are insufficient for 3D face editing in NeRF because 3D editing requires preserving the stereo features of unedited areas and ensuring consistency across different viewpoints. One possible approach is to add a multi-view loss, but selecting the viewpoint and obtaining the corresponding 2D mask is very difficult. Therefore, this embodiment of the invention proposes a spatial loss term to calculate the similarity of the sampling point features used in volume selection:
[0170]
[0171] Where r(i) is the i-th sample point rendered along ray r in the unedited region, and N is the number of sample points. This represents the point-by-point feature calculation process, including three-plane projection and feature decoding. NeRF rendering uses layered volume sampling, while this embodiment of the invention only calculates the spatial loss of coarsely sampled points, because they correspond to different faces with the same spatial location.
[0172] Therefore, the preferred and final target loss function (or optimization function) is:
[0173] L(w edit )=γ1L edit +γ2L img +γ3L space
[0174] Among them, γ1, γ2, and γ3 are hyperparameters set by the user. The default value of this invention is γ1 = 40, γ2 = 20, and γ3 = 0.2. They can also be set to γ1 = 20, γ2 = 30, and γ3 = 0.2. The number of optimization iterations is set to 10 steps to balance time efficiency and optimization quality.
[0175] VII. Application Scenarios
[0176] According to one embodiment of the present invention, a UI (User Interface) is designed, such as... Figure 5 As shown, it supports 3D face generation based on line drawings and detailed editing of 3D faces.
[0177] As an example, based on this UI interface, users can perform the following operations:
[0178] A1: To generate a NeRF face, the user needs to draw a line drawing on the canvas on the left. The system then composites and displays a highly realistic 3D face result in the window on the right. The user can drag the Angle Yaw and Angle Pitch sliders to control the perspective of the rendered result. Simultaneously, the user can select a reference image and drag it onto the canvas on the right to composite a result with a new appearance.
[0179] A2: To edit the NeRF of a face, the system automatically generates a corresponding 3D line drawing. When the user changes the viewpoint using a slider, the generated face image and the corresponding line drawing rotate simultaneously. Thanks to the 3D line drawing synthesis method, the rendered line drawing maintains high consistency during viewpoint changes, improving the user's interactive experience. During editing, the user can erase unwanted lines and draw new lines depicting the desired structure. These operations provide sufficient information to infer the mask M, representing the local area being edited. Specifically, this invention expands the newly drawn lines and combines them with the erased area to generate an initial mask. Then, a connectivity detection algorithm is used to find and fill small holes within the editing area, and the mask boundaries are smoothed using polygon curves. Based on the input face, the modified line drawing, and the predicted mask, this invention generates a new edited face NeRF, which is then rendered and displayed in the system. After editing from a single viewpoint, the user can rotate the face and continue editing from other viewpoints, supporting refined face model modifications.
[0180] According to an embodiment of the present invention, a method for generating and editing faces based on the system of the foregoing embodiments is provided. The method includes: providing a function to generate a 3D face based on line art, and providing a function to edit a 3D face based on line art. The method involves iteratively optimizing a fusion implicit code using a preset optimization method, and generating a second 3D face line art and a second colored 3D face based on the optimized fusion implicit code. Editing a 3D face based on line art is performed on a 2D line art (i.e., a custom viewpoint line art) by erasing existing lines and / or adding new lines. This reduces the difficulty of editing and allows users to edit faces more simply, efficiently, and conveniently.
[0181] According to one embodiment of the present invention, the method further includes: generating a two-dimensional output line drawing from a user-defined perspective based on the second three-dimensional face line drawing; and / or generating a two-dimensional output color face image from a user-defined perspective based on the second color three-dimensional face.
[0182] According to an embodiment of the present invention, the method further includes: using the second three-dimensional face line drawing and the second colored three-dimensional face as a new first three-dimensional face line drawing and a new first colored three-dimensional face, providing the user with the function of continuously editing the three-dimensional face based on the line drawing.
[0183] To visually demonstrate the effects, the inventors have provided some schematic diagrams illustrating the face generation and editing processes. (See attached image.) Figures 6-12 To protect privacy, all color face images were stylized. For example, the training data used the FFHQ dataset, and AnimeGANv2 was used to generate stylized faces. The corresponding line art data was preprocessed using Photoshop image copying, and then generated using a line art simplification method. The illustrative example shows the effect of 3D face generation based on line art. Figure 6 As shown, the effect of 3D face editing based on line art is as follows: Figure 7 As shown, the effect of multi-view editing based on line art is as follows: Figure 8 As shown, the effect of first creating a line drawing and then performing detailed editing is as follows: Figure 9 As shown, the effect of local appearance editing based on the line art is as follows: Figure 10 As shown, the effect of 3D face synthesis based on line drawings of various styles is as follows: Figure 11 As shown, other face editing effects include... Figure 12 As shown.
[0184] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.
[0185] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0186] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.
[0187] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A system for synthesizing human faces based on line drawings, characterized in that, include: The three-plane-based implicit code synthesis module is configured to: acquire line art and appearance reference images in pairs; inject appearance information into the line art using the appearance reference images to synthesize the three-plane features corresponding to the line art; project the three-plane features into the latent space of the EG3D model to obtain the implicit code corresponding to the line art, wherein the line art is a two-dimensional line art containing a character's face, and the appearance reference image is a color image containing a character's face. The three-plane-based implicit code synthesis module includes a trained three-plane prediction network for the line art, which includes an appearance encoder, a transformation network, and a convolutional network; the three-plane-based implicit code synthesis module also includes a trained 2D encoder, which is the encoder in the trained pSp framework image autoencoder. The line art rendering module is configured to generate corresponding 3D line art of a face based on the hidden code. The line art rendering module is implemented by converting the output RGB three channels into a single channel EG3D model. The line art rendering module includes a StyleGAN backbone, a low-resolution line art decoder with a single channel output, and a super-resolution decoder with a single channel output. The image rendering module is configured to generate corresponding color 3D faces based on implicit codes, using a pre-trained EG3D model. The user operation module is configured as follows: The first line drawing and the first appearance reference image are obtained, and the first three-plane features, the first hidden code, the first three-dimensional line drawing of the face and the first color three-dimensional face corresponding to the first line drawing are obtained by using the hidden code synthesis module, the line drawing rendering module and the image rendering module. The system provides users with an interface to select different perspectives to observe the first three-dimensional line drawing of a face. Based on the user-selected perspective and the first three-dimensional line drawing of the face, it provides a two-dimensional editable custom perspective line drawing. The system also uses the image rendering module to obtain a second appearance reference image. Obtain the second line drawing obtained after the user edits the face on the custom view line drawing, and the mask used to distinguish the edited area from the non-edited area; Based on the second line drawing and the second appearance reference image, the second three-plane feature corresponding to the second line drawing is obtained using the hidden code synthesis module, and the fused hidden code is obtained by projecting the result of fusing the first three-plane feature and the second three-plane feature according to the mask. The mask is used to indicate the three-plane feature of the editing area and suppress the change of the three-plane feature of the non-editing area. Based on the fusion hidden code, the line drawing rendering module and the image rendering module are used to generate a second 3D line drawing of the face and a second color 3D face after the user edits the face.
2. The system according to claim 1, characterized in that, The trained line drawing three-plane prediction network was trained in the following manner: Obtain the first training set, which includes multiple first samples, each of which includes a set of training line drawings, appearance reference images, and three-plane feature ground values; The pre-defined three-plane prediction network for line art is trained iteratively using the multiple first samples to obtain the trained three-plane prediction network for line art, wherein: The appearance editor is used to extract appearance information based on the appearance reference image corresponding to the input line drawing. The appearance information includes the color and texture of each part of the face. The transformation network is used to extract color feature maps based on the line drawing and corresponding appearance information; The convolutional network is used to output the three-plane features corresponding to the line drawing based on the voxel features of the three planes constructed from the color feature map; The training loss is calculated using a preset loss function, and the gradient is obtained through backpropagation to update the trainable parameters of the line drawing three-plane prediction network.
3. The system according to claim 2, characterized in that, The loss of the training line drawing three-plane prediction network is calculated using the following loss function: in, Represents the three-plane features of the output and the corresponding three-plane feature truth values Between distance, This represents the image generated by the hidden code corresponding to the output three-plane features. and the corresponding image ground truth Between distance, Representing an image and the corresponding image ground truth Perceptual distance between them, image ground truth Use the corresponding appearance reference image from the line drawing. , , They are respectively for , , Preset weighting coefficients.
4. The system according to claim 2, characterized in that, The trained pSp framework image autoencoder is trained in the following manner: The image autoencoder under the pSp framework is trained iteratively using the multiple first samples to obtain the trained image autoencoder, wherein: The encoder in an image autoencoder is used to project the three-plane features of the input into a latent code in the latent space of EG3D. The decoder of an image autoencoder is used to reconstruct the three-plane features based on the implicit code projected by the encoder; The loss for training the image autoencoder is determined based on the loss function of the pSp framework. The gradient is calculated based on the loss of the image autoencoder, and the trainable parameters of the image autoencoder are updated by backpropagation.
5. The system according to claim 3, characterized in that, The line art rendering module was trained in the following manner: A second training set is obtained, which includes multiple second samples. Each second sample includes a training code, a line drawing ground value at a first resolution corresponding to the code, and a line drawing ground value at a second resolution obtained by downsampling the line drawing ground value at the first resolution. The low-resolution line art decoder and super-resolution decoder of the line art rendering module are trained using the second training set to generate line art of the corresponding resolution, wherein: The StyleGAN backbone is used to convert implicit codes into three-plane features; The low-resolution line art decoder is used to generate color features and density information of sampling points in space based on the converted three-plane features, and uses volume rendering combined with camera parameters to generate multi-channel feature maps and line art of the second resolution. The super-resolution line drawing decoder is used to generate a line drawing of the first resolution based on the multi-channel feature map and the line drawing of the second resolution; The sub-loss calculated from the reconstruction loss function based on the true value of the line drawing is added to the sub-loss calculated from the regularization loss function used to constrain the three-dimensional consistency of the line drawing to obtain the total loss. The gradient is then calculated and backpropagated to update the trainable parameters of the low-resolution line drawing decoder and the super-resolution line drawing decoder.
6. The system according to claim 5, characterized in that, The total loss during training of the line art rendering module is determined as follows: in, This represents the sub-loss calculated using the reconstruction loss function. The sub-loss calculated by the regularization loss function, This represents the generated line art at the first resolution. True value of line art at first resolution Between distance, This represents the generated second-resolution line drawing. True value of line art at second resolution Between distance, Line drawing And line art true value The perceived distance between them Line drawing And line art true value The perceived distance between them This represents the number of pixels in a set of pixels randomly sampled from the line art. Line drawing The median coordinate is pixel values, This represents the true value of the line drawing at second resolution using ray projection. Light obtained during volume rendering Rendered pixel values, express distance, This represents the weight coefficient of the corresponding item.
7. The system according to claim 1, characterized in that, The fusion hidden code is iteratively optimized multiple times using a preset optimization method to make the line drawing generated by the optimized fusion hidden code more compatible with the second line drawing; the second 3D line drawing of the face and the second color 3D face are generated based on the optimized fusion hidden code.
8. A method for generating and editing faces based on the system according to any one of claims 1-7, characterized in that, The method includes: Provides functionality for generating 3D faces based on line art, including: Obtain the first line drawing and the first appearance reference image. And by using the hidden code synthesis module, the line drawing rendering module and the image rendering module, the first three-plane features, the first hidden code, the first three-dimensional line drawing of the first face and the first color three-dimensional face corresponding to the first line drawing are obtained; Provides functionality for editing 3D faces based on line art, including: It provides users with an interface to select different perspectives to observe the first face 3D line drawing, and provides a 2D editable custom perspective line drawing based on the user-selected perspective and the first face 3D line drawing; Obtain the second line drawing obtained after the user edits the face on the custom view line drawing, and the mask used to distinguish the edited area from the non-edited area; Using the image rendering module, a second appearance reference image is obtained, which is obtained by observing the first color 3D face from the same perspective as the custom viewpoint line drawing; Based on the second line drawing and the second appearance reference image, the second three-plane feature corresponding to the second line drawing is obtained using the hidden code synthesis module, and the fused hidden code is obtained by projecting the result of fusing the first three-plane feature and the second three-plane feature according to the mask. The mask is used to indicate the three-plane feature of the editing area and suppress the change of the three-plane feature of the non-editing area. The fusion hidden code is iteratively optimized multiple times using a preset optimization method, and a second 3D line drawing of the face and a second color 3D face are generated based on the optimized fusion hidden code.
9. The method according to claim 8, characterized in that, The method further includes: Based on the second 3D facial line drawing, generate a 2D output line drawing from a user-defined perspective; and / or, Based on the second color 3D face, generate a 2D output color face image from a user-defined perspective.
10. The method according to claim 9, characterized in that, The method further includes: using the second 3D face line drawing and the second colored 3D face as a new first 3D face line drawing and a new first colored 3D face, providing users with the function of continuously editing 3D faces based on line drawings.
11. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method according to any one of claims 8-10.
12. An electronic device, characterized in that, include: One or more processors; as well as Memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 8-10 by executing the executable instructions.