Head portrait generation according to artistic styles

Through domain adaptation framework and art dataset training, combined with camera alignment, feature regularization and geometric deformation technology, the problem of arbitrary distribution of geometric shapes and textures of GAN models when generating art avatars is solved, achieving high-quality art style avatar generation.

CN120435731APending Publication Date: 2025-08-05SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380090011.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-29
Filing Date
2023-12-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing Generative Adversarial Networks (GAN) models are difficult to effectively generate art-style avatars, especially when domain adaptation is performed while maintaining facial features and art styles, there are challenges resulting from arbitrary distribution of geometric shapes and textures.

Method used

Using the domain adaptation framework, the 2D GAN of the target domain is trained using the art dataset, combined with the camera alignment module, feature regularization module, geometric deformation module and avatar generation module, the pre-trained 3D GAN is adjusted to generate artistic avatars, and the training is performed using StyleGAN-ADA and R1 regularization technologies, and the source synthesized 2D data of EG3D and StyleCariGAN and DualStyleGAN for domain adaptation.

Benefits of technology

It realizes the generation of realistic art avatars while maintaining facial identity, solves the problem of arbitrary distribution of geometric shapes and textures of the art dataset, and improves the artistic style adaptability and image quality of the generative model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120435731A_ABST
    Figure CN120435731A_ABST
Patent Text Reader

Abstract

A domain adaptation framework for generating a 3D avatar generative adversarial network (GAN) that is capable of generating an avatar based on a single photo image. A 3D avatar GAN is generated by training a target domain using an artistic data set. Each artistic data set includes a plurality of source images, each source image being associated with a style type, such as a cartoon, cartoon, and comic. In some embodiments, the domain adaptation framework begins from a source domain that has been trained according to 3D GAN and a target domain that has been trained with 2D GAN. The framework fine-tunes the 2D GAN by training it with an artistic data set. The obtained 3D head portrait GAN generates a 3D artistic head portrait; and the editing module is used for executing semantic and geometric editing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. application serial number 18 / 090,692, filed December 29, 2022, the contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] Examples presented in this disclosure relate to machine learning, generative models, and training datasets. More specifically, but not limited to, this disclosure describes an adaptation framework for training a target domain based on one or more art datasets to produce avatars rendered in a selected artistic style. Background Art

[0004] Machine learning refers to mathematical models or algorithms that gradually improve through experience. By processing a large number of diverse input data sets, machine learning algorithms can develop improved generalizations about a specific data set and then use these generalizations to produce accurate outputs or solutions when processing new data sets. Broadly speaking, a machine learning algorithm includes one or more parameters that are adjusted or changed in response to new experience, gradually improving the algorithm; the process is similar to learning.

[0005] Generative adversarial networks (GANs) are a class of machine learning frameworks in which two artificial neural networks (e.g., a generator and a discriminator) are trained together. Using a training dataset, the generator module is trained by generating new data (e.g., new synthetic images) that have the same or similar characteristics (e.g., statistically, mathematically, visually) as reference data in the training dataset (e.g., thousands of sample images). The generator module generates candidates (e.g., new images) based on the reference data. The discriminator module evaluates the candidates by determining how similar each candidate is to the reference data (e.g., by assigning a value between zero and one). If the discriminator concludes that the candidate is highly similar to the reference data, the candidate generated by the generator is classified as better (e.g., a value closer to one). If the discriminator concludes that the candidate is not very similar to the reference data (e.g., the candidate appears to be synthetic or forged), the candidate is classified as poor (e.g., a value closer to zero). Typically, the generator and discriminator are trained together. The generator learns and produces better and more excellent candidates, while the discriminator learns and becomes better at identifying poor candidates.

[0006] Generative frameworks like GANs can be used to generate realistic portraits based on photographic images of real faces. A framework called 3D-GAN is able to generate three-dimensional portraits from a single two-dimensional image.

[0007] An avatar is a graphical representation of a device user (e.g., a computer or mobile device user), or a persona or alter ego representing that user. An avatar can be personalized based on the preferences of the user or others. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The features of the various embodiments disclosed will be readily understood through the following detailed description, with reference to the accompanying drawings. Reference numerals are used for each element in the description and in the several views of the drawings. When multiple similar elements are present, a single reference numeral may be assigned to the same element, with added lowercase letters indicating the specific element. When referring to one or more non-specific elements, lowercase letters may be omitted.

[0009] Unless otherwise noted, the various elements shown in the figures are not drawn to scale. The dimensions of the various elements may be exaggerated or reduced for clarity. Several figures depict one or more embodiments and are presented by way of example only and should not be construed as limiting. The drawings include the following:

[0010] Figure 1 is a block diagram of the example adaptation framework;

[0011] Figure 2 is an illustration of the source image and several example artistic avatars generated according to the adaptation framework;

[0012] Figure 3 is a flowchart listing steps in an example method for generating an artistic headshot according to a headshot GAN generated and trained according to an adaptation framework;

[0013] Figure 4 is a flowchart listing steps in an example method for training a target domain according to the adaptation framework; and

[0014] Figure 5 is a block diagram of an example configuration of a machine suitable for implementing the method of generating a 3D representation of an object according to the systems and methods described herein. DETAILED DESCRIPTION

[0015] Generate artistic avatars ("artistic avatars") using a generative adversarial network (GAN). In an example embodiment, avatar generation includes selecting a source domain that has been trained with a source GAN (e.g., a 3D GAN) and selecting a target domain that has been trained with a target GAN (e.g., a 2D GAN). The target domain is trained with an artistic dataset (including artistic features, attributes, or a combination thereof) to produce an avatar GAN (e.g., a 3D avatar GAN). The trained avatar GAN generates artistic avatars based on real images and including features and attributes found in the artistic dataset.

[0016] The following detailed description includes systems, methods, techniques, instruction sequences, and computer program products that illustrate the examples set forth in this disclosure. In order to provide a thorough understanding of the disclosed subject matter and its related teachings, many details and examples are included. However, those skilled in the relevant art will understand how to apply the related teachings without these details. The aspects of the disclosed subject matter are not limited to the specific devices, systems, and methods described, as the related teachings can be applied or practiced in various ways. The terms and nomenclature used herein are only used to describe specific aspects and are not intended to be limiting. In general, well-known instruction instances, protocols, structures, and techniques are not necessarily shown in detail.

[0017] As used herein, the terms "connect," "connected," "coupled," and "coupled" refer to any logical, optical, physical, or electrical connection, including links that carry electrical or magnetic signals generated or supplied by one system element to another coupled or connected system element. Unless otherwise specified, coupled or connected elements or devices are not necessarily directly connected to each other and may be separated by intermediate components, elements, or communications media, one or more of which can modify, manipulate, or carry electrical signals. The term "on" means directly supported by an element or indirectly supported by the element through another element that is integrated into or supported by the element.

[0018] Additional objects, advantages, and novel features of the examples will be set forth in part in the following description and in part will become apparent to those skilled in the art upon examination of the following and accompanying drawings, or may be learned by production or operation of the examples. The objects and advantages of the subject matter may be realized and obtained by the methods, instrumentalities, and combinations particularly pointed out in the appended claims.

[0019] Reference will now be made in detail to the examples illustrated in the accompanying drawings and discussed below.

[0020] Photorealistic portrait face generation is a landmark application that demonstrates the power of generative models, especially GANs. Models called 3D GANs are able to learn 3D structures without 3D supervision. This unsupervised training is feasible when large-scale training datasets contain objects with relatively consistent geometric shapes (e.g., thousands of face images), allowing 3D GANs to learn from the distribution of shapes and textures. For example, Figure 1 As shown in the block diagram of , an example system for generating a realistic avatar 120 includes a source domain 100 (e.g., including a 3D GAN 150) and a target domain 110, which has been trained using a training dataset of real images (e.g., thousands of face images).

[0021] As used herein, the term "artistic" is used to refer to the aesthetic qualities of human or machine-generated content; and the term "artistic style" is used to refer to aspects or characteristics of these aesthetic qualities (e.g., purposeful or arbitrary exaggeration of geometry, texture, or both). In contrast to datasets containing real faces, art datasets 160 typically include artwork (e.g., sample images 180) with arbitrary exaggerations of facial geometry and texture. In some embodiments, art datasets 160 include a large number of sample art images 180 in artistic styles, such as comics, Pixar-style, cartoons, graphic novels, etc. Art datasets 160 may include unusual, highly variable geometry and arbitrary feature exaggerations—one or more of which may vary depending on the context, artist, prohibited style guides, and production requirements. For example, a nose, cheeks, and eyes may be arbitrarily drawn based on the artist's style and the subject's features. In some embodiments, the systems and methods described herein include a framework for adapting a pre-trained GAN model to work with an art dataset.

[0022] Figure 1 is a block diagram of an example domain adaptation framework 150. In the example shown, the source domain 100 is a pre-trained 3D GAN 105. The target domain 110 is a 2D GAN 115. In some embodiments, the 2D GAN 115 is trained on an art dataset 160, rather than using real facial images to train the 2D GAN 115.

[0023] Training 2D GAN 115 using art dataset 160 produces 3D avatar GAN 170, which is capable of generating an artistic avatar 190 (e.g., a caricature) based on an image 20 (e.g., a photograph) of an object 10 (such as a human face). Artistic avatar 190 includes features of object 10 and attributes established by art dataset 160. In this regard, adaptation framework 150 preserves the subject's identity. In other words, artistic avatar 190 is recognizable as an artistic representation of the person in image 20.

[0024] In some implementations, the domain adaptation framework 150 includes a camera alignment module 151 , a feature regularization module 152 , a geometric deformation module 153 , an avatar generation module 154 , and an editing module 175 .

[0025] In some embodiments, the camera alignment module 151 includes an optimization-based method for aligning the distribution of camera parameters across domains. Most available art datasets 160 do not include camera parameters. For example, if the sample image was drawn by an artist rather than captured using a camera, the data associated with the sample image will not include camera parameters.

[0026] In some embodiments, feature regularization module 152 includes regularization tools associated with texture, geometry, and depth established in learning art dataset 160. Regularization improves color and texture, avoids degenerate geometric solutions (e.g., flat shapes), and maintains the depth of the resulting portrait relative to the background.

[0027] Source domain (T S )100 and target domain (T t )110 has no paired data. Assume that the target domain (T t ) 110 are not annotated with camera parameters. These conditions influence the selection of the discriminator (D). In some embodiments, the feature regularization module 152 includes an unconditional version of the dual discriminator, as proposed in Efficient, Geometry-Aware 3D GAN (herein referred to as “EG3D”), where the discriminator is not conditioned on camera parameters. During training, the 3D avatar GAN (G t )170 uses M(θ′,φ′,c′,r′) to generate arbitrary images with poses. The discriminator (D) uses the images from the target domain (T t ) 110 to identify these images. In some embodiments, the feature regularization module 152 uses a training scheme called StyleGAN-ADA and R1 regularization to make the source domain (T S )100 Adapt to target domain (T t )110(T s →T t ).

[0028] In some embodiments, geometric deformation module 153 includes deformation-based techniques for modeling exaggerated geometric shapes found in some art datasets 160. In some embodiments, editing module 175 is based on techniques established for geometric deformation module 153.

[0029] The regularizer described in this paper adjusts T s →T t For larger geometric deformations (e.g., those found in the comic art dataset 160), in some embodiments, the adaptation framework 150 includes a geometric deformation module 153.

[0030] In some embodiments, the geometric warping module 153 includes a process for editing geometric shapes by utilizing the properties of three-plane features learned by EG3D. In this example, the geometric warping module 153 begins by analyzing the three planes in the source 3D GAN (GS) 105. Generally speaking, the frontal plane encodes the majority of the information needed to render the final image. To quantify this, in some embodiments, the geometric warping module 153 samples an image and depth map from the source 3D GAN (GS) 105 and then swaps the frontal plane with the other planes from two random images. The geometric warping module 153 then compares the difference in the RGB values of the images with the chamfer distances of the depth maps. When swapping the three frontal planes, the final image is completely swapped, and the chamfer distances change by approximately 80% to 90%, matching the depth map of the swapped image. In the case of the other two planes, the RGB image is not significantly affected, and in most cases the chamfer distance of the depth map is only reduced by approximately 20% to 30%.

[0031] In this regard, the geometric deformation module 153 manipulates the 2D front plane features in order to learn additional deformations and exaggerations. In some embodiments, the geometric deformation module 153 includes learning a TPS (Thin Plate Spline) network on top of the front plane.

[0032] In some embodiments, the avatar generation module 154 includes a process for linking latent spaces associated with the source domain and the target domain. The latent space associated with 3D GANs is more entangled than that of 2D GANs, which makes linking latent spaces between domains more challenging. A latent space is a compressed or simplified representation of the data associated with a particular feature. For example, the latent space associated with the tip of the nose can be reduced to a single data point in space that can be plotted on a graph, while the entire shape of the nose cannot. A simplified latent space is easier to analyze than analyzing the entire feature. In this regard, the latent space serves as an intermediate step during training.

[0033] GANs are generative models. One type of GAN, called StyleGAN, is particularly useful for smaller, high-quality datasets such as FFHQ, AFHQ, and LSUN objects. The disentangled latent space learned by StyleGAN has been shown to have semantic properties that are beneficial for semantic image editing. CLIP-based image editing and domain transfer are another group of works enabled by StyleGAN.

[0034] Algorithms that project existing images into the GAN latent space are part of most GAN-based image editing techniques. There are two main approaches to this projection: optimization-based and encoder-based. Within both streams, after obtaining the initial inversion results, the generator weights can be further modified.

[0035] Some existing systems attempt to extract 3D structures from pre-trained 2D GANs. Recently, inspired by Neural Radiance Field (NeRF) modeling, GAN architectures that combine implicit or explicit 3D representations with neural rendering techniques have been proposed.

[0036] In some embodiments, the systems and methods described herein build on EG3D, which includes state-of-the-art results for faces trained on the FFHQ dataset. In some embodiments, the systems and methods described herein employ 2D to 3D domain adaptation and refinement, leveraging synthetic 2D data from sources such as StyleCariGAN and DualStyleGAN.

[0037] Training a 3D GAN using an art dataset 160 presents challenges due to the arbitrary distribution of geometric shapes and textures found in many works of art. In some embodiments, each art dataset 60 includes a large number of 2D sample images 180, each image being associated with a style type 162 (e.g., comic, Pixar-style) and one or more attribute classifiers 164, as described herein. In this regard, the domain adaptation framework 150 involves fine-tuning the 3D GAN 105 using 2D artwork from the art dataset 160. In this context, the domain adaptation framework 150 involves training a 3D GAN 105 from a domain that is similar to the source domain (T s )100 related existing 3D-GAN (G s )105. One of the goals of the adaptation framework 150 is to generate t )110 associated 3D avatar GAN (G t )170, while maintaining 3D-GAN(G s )105, while preserving the semantic, stylistic, and geometric properties of the domain The identity of the subjects.

[0038] In some embodiments, to preserve the identity of the subject, the avatar generation module 154 uses one or more attribute classifiers 164 associated with the art dataset 160. The style type 162 and the attribute classifier 164 provide information about the coupling attributes between the image 20 and the artistic rendering (e.g., caricature). In some embodiments, the attribute classifier 164 is applied in an ex post facto manner. If applied during training, the attribute classifier 164 can affect the texture in the target domain and can sometimes degenerate into a relatively narrow style output. To avoid overfitting to the 3D-GAN (G s )105 and encourages easier transfer of optimized latent codes to 3D avatar GANs (G t) 170, in some embodiments, the avatar generation module 154 includes W space optimization. Finally, in some embodiments, the avatar generation module 154 initializes the 3D avatar GAN (G t )170 w code, and the target domain (T t ) 110 applies additional attribute classifier losses, as well as the depth regularization described herein (e.g., the equation for R(D)). In some embodiments, the attribute classifier 164 generalizes across all domains and applies W / W+ space optimization to improve the quality and diversity of the output.

[0039] In some aspects, the editing module 175 is based on techniques established for and implemented by the geometric warping module 153. For example, the editing module 175 is guided by the learned latent space. The domain adaptation framework 160 is designed to preserve the properties of the W and S latent spaces. In some embodiments, the editing module 175 includes processes for performing semantic editing using available tools such as InterFaceGAN, GANSpace, and StyleSpace. In some embodiments, the editing module 175 includes processes for performing geometric editing using the TPS module and Δs interpolation, as described herein. To perform video editing, in some embodiments, the editing module 175 includes an encoder for EG3D (which is based on e4e) to encode the video and extract the edits from the 3D-GAN (GAN) based on the w code. s )105 transferred to 3D avatar GAN (G t )170.

[0040] Figure 2 is an illustration of a source image 20 (e.g., a single photograph of a person) and several example artistic portraits 190 as a result of the domain adaptation framework 150 described herein. The results include example portraits 190 associated with various art datasets 160 and style types 162, including comics, Pixar-style, cartoons, and comic strips (e.g., graphic novels, classic comic books).

[0041] In the domain Selecting an appropriate range for camera parameters is one way to achieve high-fidelity geometry and texture details. t ) 110 does not contain camera parameter symbols, so the domain adaptation framework 150 will suppress undesirable artifacts such as low-quality textures and flat geometry in different views. Typically, camera parameters are estimated empirically, directly calculated from the dataset (e.g., using an off-the-shelf pose detector), or learned during training. Directly estimating camera parameters will be difficult because the state-of-the-art dataset 160 typically does not include any 3D information. Instead, in some embodiments, the camera alignment module 151 includes a method that ensures that the camera parameters are aligned based on the statistical distribution 401 ( Figure 4 ) in the domain There is consistency between them.

[0042] In some embodiments, for a target domain (T t ) 110, using StyleGAN2 trained on FFHQ, fine-tuned on the art dataset. In some embodiments, the camera alignment module 151 assumes that the intrinsic parameters (e.g., focal length, optical center, resolution) of all cameras are the same. Then, the camera alignment module 151 uses the 3D GAN (G s )105 external camera parameters (e.g., pose, orientation, position coordinates) and the statistical distribution 401 of 2D GAN (G 2D ) 115 for matching and then using the matching distribution 401 for training. In this regard, the camera alignment module 151 includes an optimization-based approach to matching the sought distribution 401. In some embodiments, one of the first steps is to identify a 2D GAN (G 2D )115. In some embodiments, the canonical pose image is an image with zero yaw, pitch, and roll parameters. The image corresponding to the average potential code satisfies this property. Let:

[0043] I s (w,θ,φ,c,r)=G s (w,M(θ,φ,c,r))

[0044] And make

[0045] I 2D (w) = G s (w)

[0046] Indicates that given the w code variable, 3D GAN (G s )105 and 2D GAN(G 2D )115 generates an arbitrary image.

[0047] Let k d By detector K d Detected facial key points; then

[0048]

[0049] Among them L kd (I1,I2)=||k d (I1)-k d (I2)||1

[0050] And where w avg It is similar to 2D GAN (G 2D )115 associated with the average w latent code, and wa ' vg It is similar to 3D GAN (G s ) 105 associated with the average w latent code. In our results, r′ is determined to be approximately 2.70, and c′ is approximately [0.0, 0.05, 0.17].

[0051] The next step is to determine the safe range for the θ and φ parameters. Following resources such as StyleFlow and FreeStyleGAN, we set these parameters to θ′∈[-0.45,0.45] and φ′∈[-0.35,0.35], in radians.

[0052] Figure 3 300 is a flowchart outlining steps in an example method for generating an artistic head portrait 190 according to a head portrait GAN 170 generated and trained according to the domain adaptation framework 150 described herein. Although these steps are described in the context of training a head portrait GAN 170, those skilled in the art will appreciate other uses and implementations of the steps for other types of systems based on the description herein. One or more of the steps shown and described may be performed simultaneously, in series, in an order different from that shown and described, or in combination with additional steps. Some steps may be omitted or repeated in some applications.

[0053] Block 302 depicts an example step of selecting a source domain 100 that has been trained on a source GAN 105. In some implementations, the source domain 100 is a pre-trained 3D GAN 105.

[0054] Block 304 recites an example step of selecting a target domain 110 that has been trained on a target GAN 115. In some implementations, the target domain 110 is a 2D GAN 115 trained using real facial images.

[0055] Block 306 describes example steps for training the target domain 110 based on the art dataset 160, such that the avatar GAN 170 is trained. Instead of using real facial images to train the target domain 110, block 306 describes example steps for training the target domain 100 using one or more art datasets 160 described herein.

[0056] Block 308 depicts an example step of capturing an image 20 of an object 10. In some embodiments, the object 10 is a person's face. In some embodiments, the source image 20 is a photograph of a face captured by a camera, retrieved from memory, or otherwise obtained.

[0057] Block 310 describes example steps for generating an artistic head portrait 190 based on the image 20 and according to the head portrait GAN 170 produced by training.

[0058] Block 312 recites the example step of projecting the image 20 onto the latent space 311 associated with the source GAN 105, where the latent space 311 is represented by the latent code 313. Block 314 recites the example step of transferring the latent code 313 to the avatar GAN 170. Block 316 recites the example step of rendering the artistic avatar 190 based on the latent code 313.

[0059] Figure 4 4 is a flowchart 400 outlining steps in an example method for training a target domain 110 according to the domain adaptation framework 150 described herein. Although these steps are described in the context of training an avatar GAN 170, those skilled in the art will appreciate, based on the description herein, other uses and implementations of the steps in other types of systems. One or more of the steps shown and described may be performed simultaneously, in series, in an order different from that shown and described, or in combination with additional steps. Some steps may be omitted or repeated in some applications.

[0060] Block 402 recites the example steps of computing a statistical distribution 401 based on a set of camera parameters 107 associated with the source domain 100 as described above.

[0061] Block 404 recites example steps of applying the statistical distribution 401 to estimate a set of camera parameters 117 for the target domain 110 .

[0062] Block 406 describes an example step of applying the loss function 405 and one or more feature regularizers 407. In some embodiments, the feature regularization module 152 includes one or more loss functions 405 and regularizers 407 to be applied to a 3D avatar GAN (G t ) 170. Because the art dataset 160 is typically not derived from a consistent 3D model or form (e.g., because the work is artistic), in some cases, the 3D avatar GAN (G t ) 170 associated generator modules tend to converge to simple degenerate solutions with flat geometry. In this regard, in some embodiments, the feature regularization module 152 seeks to extract the features from the 3D GAN (G s )105.

[0063] Block 408 describes example steps for controlling one or more geometric deformations 409 according to a thin plate spline network 411. In some embodiments, the TPS network 411 is conditioned on the front plane features and the W latent space features to achieve multiple transformations. The architecture of this module is similar in some respects to the standard SyleGAN2 layer, with an MLP added at the end to predict the control points of the transformed features. In some embodiments, this module is trained separately; in 3D avatar GAN (G t )170 has been trained. In some cases, joint training can produce instabilities, which may be related to the instability of the source domain (T S )100 and target domain (T t )110. Formally, we define this transformation as:

[0064] T(w,f):=Δc

[0065] where w is the latent code, f is the front plane, and c is the control point.

[0066] Let c1 be the initial control point that generates the identity transformation. Let (c1, c2) be the control points corresponding to the front plane (f1, f2) sampled using the W code (w1, w2). Let (c1′, c2′) be the point in the TPS network where the W code (w1, w2) is swapped. To standardize and encourage the TPS module to learn different deformations, in some embodiments, the geometric deformation module 153 includes:

[0067]

[0068] In some embodiments, the geometric deformation module 153 includes an additional loss term to learn the target domain (T t )110. Let S(I) be the soft argmax output of the face segmentation network, given an image I, and assuming that S generalizes to caricatures, then

[0069] R(T2):=||S(G t (w)),S(I t )||1

[0070] In some embodiments, the avatar generation module 154 includes a module for linking 3D-GAN (G s )105 latent space and 3D portrait GAN (G t )170. The process of linking the latent space is to use 3D avatar GAN (G t) 170 to generate a portion of a personalized 3D artistic avatar 190 based on a single image 20 (e.g., a reference photo) of an object 10 (e.g., a person's face). Generally speaking, when a 3D generator is used to process the projection of a real 2D photographic image, there are often differences between the coupled latent spaces.

[0071] In some embodiments, the avatar generation module 154 includes projecting the image 20 onto a 3D-GAN (G s )105, and then transfer the latent code 313 to the 3D avatar GAN (G t ) 170, and then further optimizes the image 20 to generate a 3D artistic head portrait 190. In some embodiments, the head portrait generation module 154 includes an optimization-based process to find a minimum between the generated head portrait 190 and the 3D-GAN (G s ) 105. The process includes aligning the camera, as described herein. In some embodiments, the avatar generation module 154 includes projecting the image 20 into a 3D-GAN (G s )105.

[0072] Block 410 depicts an example step of applying the adversarial loss 413 obtained during training as the main loss function 405A. In some embodiments, the feature regularization module 152 uses the adversarial loss 413 obtained during training as the main loss function 405A. In this example, the main loss function 405A is a standard non-saturating loss used to train the generator and discriminator networks (e.g., networks associated with the Efficient, Geometry-Aware 3D GAN referred to as “EG3D”). In some embodiments, the feature regularization module 152 also includes lazy density regularization to ensure that the final fine-tuned 3D Avatar GAN (G t )170 consistency of density values.

[0073] Block 412 describes an example step of applying texture regularization 407A as described herein. Texture data includes multiple layers and can be entangled with geometric information. In some embodiments, the feature regularization module 152 utilizes the fine-grained style information encoded in relatively later layers to update the tRGB layer parameters (outputting three-plane features) before the neural rendering stage. In addition, since the network needs to adapt to the target domain (T t ) 110, so in some embodiments, the feature regularization module 152 updates the decoder (MLP layer) of the neural rendering pipeline. Given the EG3D architecture, in some embodiments, the feature regularization module 152 updates the super-resolution layer parameters to improve the consistency between the low-resolution and high-resolution outputs seen by the discriminator.

[0074] Block 414 depicts the example step of applying geometric regularization 407B based on the derived parameters 413 associated with the S latent space 415. In some embodiments, the feature regularization module 152 updates relatively early layers with regularization to allow the network to learn the target domain (T t )110 and simultaneously improve the preservation of properties associated with the W and S latent spaces. Updating earlier layers encourages the source domain (T S )100 and target domain (T t ) 110. In this regard, the feature regularization module 152 updates the bias parameter (Δs) 413 based on the s activations of the S latent space 415. The s parameter is predicted by A(w), where A is the learned affine function in EG3D. In order to preserve the identity and geometry, the optimization of the bias parameter (Δs) 413 does not deviate from the original source domain (T s ) 100 is too far, in some embodiments, the feature regularization module 152 includes a regularizer given by:

[0075] R(Δs):=||Δs||1

[0076] In some embodiments, the regularizer R(Δs) is applied using density regularization. Surprisingly, after training, we can interpolate between s and (s+Δs) parameters to obtain the best fit in the source domain (T S )100 and target domain (T t )110 to interpolate between the geometric shapes of the samples.

[0077] Block 416 depicts the example step of applying depth regularization 407C based on the average background depth 417. Although the above-described geometric regularization 407B produces a target domain (T t )110 geometric improvements, but from 3D head GAN (G t ) 170 may still produce some samples with relatively flat geometry. Such cases are difficult to detect. In some embodiments, the feature regularization module 152 includes evaluating the depth of the background relative to the foreground. In this regard, the feature regularization module 152 includes an additional regularization called depth regularization 407C, which encourages the 3D avatar GAN (G t )170 associated with the average background depth 417 and the source 3DGAN (G S )105. For example, let S b Represents the face background segmentation network. The feature regularization module 152 calculates the feature regularization function generated by the source 3D GAN (G S )105 gives the average background depth 417 of the sample. The average depth 417 is given by the following formula:

[0078]

[0079] In this equation, D n From the source 3D GAN (G S )105 sampled image I n The symbol ⊙ represents the Hadamard product, M is the number of sampled images, and N n is image I n Finally, in some embodiments, depth regularization 407C is defined as:

[0080] R(D):=||a d ·J-(D t ⊙S b (I t ))|| F

[0081] Among them D t It is from 3D avatar GAN (G t )170 sampled image I t The depth map, and J is D t Matrices with the same spatial dimensions.

[0082] The techniques described herein can be used with one or more computing systems described herein or with one or more other systems. For example, the various processes described herein can be implemented using hardware or software, or a combination of both. For example, at least one of the processors, memories, storage devices, one or more output devices, one or more input devices, or communication connections discussed below can each be at least a portion of one or more hardware components. Dedicated hardware logic components can be constructed to implement at least a portion of one or more of the techniques described herein. For example, but not limited to, such hardware logic components can include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like. Applications of the various devices and systems can include a wide range of electronic and computing systems. The techniques can be implemented using two or more specific interconnected hardware modules or devices with associated control and data signals that can communicate between and through the modules, or as part of an application specific integrated circuit. Additionally, the techniques described herein can be implemented using a software program executable by a computing system. For example, implementations can include distributed processing, component / object distributed processing, and parallel processing. Furthermore, virtual computing system processing can be constructed to implement one or more of the techniques or functions described herein.

[0083] Figure 5 An example configuration of a machine 500 is shown that includes components that may be incorporated into a processor 502 adapted to manage 3D asset construction.

[0084] Specifically, Figure 5 A block diagram of an example machine 500 on which one or more configurations can be implemented is shown. In alternative configurations, the machine 500 can operate as a standalone device or can be connected (e.g., networked) to other machines. In a networked deployment, the machine 500 can operate as a server machine, a client machine, or both in a server-client network environment. In one example, the machine 500 can act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. In the example configuration, the machine 500 can be a personal computer (PC), a tablet computer, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smartphone, a network appliance, a server, a network router, a switch or a bridge, or any machine capable of executing instructions (sequentially or otherwise) specifying the actions to be taken by the machine. For example, the machine 500 can be used as a workstation, a front-end server, or a back-end server for a communication system. The machine 500 can implement the methods described herein by running software for implementing the features described herein. Furthermore, although only a single machine 500 is shown, the term "machine" should also be construed to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0085] As described herein, examples may include or may run on a processor, logic, or multiple components, modules, or mechanisms (referred to herein as "modules"). A module is a tangible entity (e.g., hardware) that is capable of performing a specified operation and may be configured or arranged in a particular manner. In one example, a circuit may be arranged in a specified manner as a module (e.g., internally or relative to an external entity, such as another circuit). In one example, all or part of one or more computing systems (e.g., stand-alone, client, or server computer systems) or one or more hardware processors may be configured by firmware or software (e.g., instructions, application portions, or applications) as a module that operates to perform a specified operation. In one example, the software may reside on a machine-readable medium. When executed by the underlying hardware of the module, the software causes the hardware to perform the specified operation.

[0086] Thus, the term "module" is understood to include at least one of a tangible hardware or software entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., temporarily) configured (e.g., programmed) to operate in a particular manner or perform part or all of any of the operations described herein. Given that modules are examples of temporary configurations, each module need not be instantiated at any time. For example, where a module includes a general-purpose hardware processor configured using software, the general-purpose hardware processor can be configured as corresponding different modules at different times. The software can configure the hardware processor accordingly, for example, to constitute a particular module at one time, and to constitute different modules at different times.

[0087] Machine (e.g., computing system or processor) 500 may include a hardware processor 502 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), main memory 504, and static memory 506, some or all of which may communicate with each other via an interconnect (e.g., a bus) 508. Machine 500 may also include a display unit 510 (shown as a video display), an alphanumeric input device 512 (e.g., a keyboard), and a user interface (UI) navigation device 514 (e.g., a mouse). In one example, display unit 510, input device 512, and UI navigation device 514 may be a touch screen display. Machine 500 may additionally include a mass storage device (e.g., a drive unit) 516, a signal generating device 518 (e.g., a speaker), a network interface device 520, and one or more sensors 522. Example sensors 522 include one or more of a global positioning system (GPS) sensor, a compass, an accelerometer, a temperature, light, camera, video camera, physical state or position sensor, pressure sensor, fingerprint sensor, retinal scanner, or other sensor. The machine 500 may include an output controller 524, such as a serial (e.g., universal serial bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, to communicate with or control one or more peripheral devices (e.g., a printer, a card reader, etc.).

[0088] The mass storage device 516 may include a machine-readable medium 526 having stored thereon one or more data structures or instructions 528 (e.g., software) embodying or utilized by any one or more of the techniques or functionality described herein. The instructions 528 may also reside, completely or at least partially, within the main memory 504, within the static memory 506, or within the hardware processor 502 during execution of the instructions 528 by the machine 500. In one example, one or any combination of the hardware processor 502, the main memory 504, the static memory 506, or the mass storage device 516 may constitute a machine-readable medium.

[0089] Although machine-readable medium 526 is shown as a single medium, the term "machine-readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database or associated cache and server) configured to store one or more instructions 528. The term "machine-readable medium" can include any medium capable of storing, encoding, or carrying instructions executed by machine 500 and causing machine 500 to perform any one or more of the techniques disclosed herein, or any medium capable of storing, encrypting, or carrying data structures used by or associated with these instructions. Non-limiting examples of machine-readable media can include solid-state memory and optical and magnetic media. Specific examples of machine-readable media can include non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; random access memory (RAM); solid-state drives (SSDs); and CD-ROM and DVD-ROM disks. In some examples, the machine-readable medium can include non-transitory machine-readable media. In some examples, the machine-readable medium can include machine-readable media that is not a transitory propagation signal.

[0090] The instructions 528 may also be sent or received over the communication network 532 via the network interface device 520 using a transmission medium. The machine 500 may communicate with one or more other machines using any of a variety of transmission protocols, such as Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc. Example communication networks may include a local area network (LAN), a wide area network (WAN), a packet data network (such as the Internet), a mobile telephone network (such as a cellular network), a plain old telephone (POTS) network, and a wireless data network (such as a wireless network). The network interface device 520 may include one or more physical jacks (e.g., Ethernet, coaxial cable, or telephone jacks) or one or more antennas 530 to connect to the communication network 532. In one example, the network interface device 520 may include multiple antennas 530 to communicate wirelessly using at least one of single-input multiple-output (SIMO), multiple-input multiple-input (MIMO), or multiple-input single-output (MISO) technology. In some examples, the network interface device 520 may use multi-user MIMO technology for wireless communication.

[0091] The features and flow charts described herein may be embodied in one or more methods as method steps, or in one or more applications as described above. According to some configurations, one or more "applications" are programs that perform the functions defined in the program. Various programming languages may be used to generate one or more of the applications structured in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, a third-party application (e.g., an application developed by an entity other than the vendor of a particular platform using an ANDROID TM or IOS TM Software Development Kit (SDK) can be used to develop applications on platforms such as IOS TM ANDROID TM 、 The present invention relates to mobile software running on the mobile operating system of the iPhone or other mobile operating systems. In this example, the third-party application can call API calls provided by the operating system to facilitate the functions described herein. The application can be stored in any type of computer-readable medium or computer storage device and executed by one or more general-purpose computers. In addition, the methods and processes disclosed herein can alternatively be embodied in dedicated computer hardware or an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a complex programmable logic device (CPLD).

[0092] The programmatic aspects of this technology can be considered a "product" or "article of manufacture," typically in the form of at least one of executable code or associated data, carried or embodied in a machine-readable medium. For example, the programming code may include code for a touch sensor or other functionality described herein. "Storage" media includes any or all tangible memory of a computer, processor, or the like, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., that can provide non-transitory storage for software programming at any time. All or part of the software may sometimes be communicated over the Internet or various other telecommunications networks. For example, such communication can load software from one computer or processor to another, such as from a service provider's server system or mainframe to the computer platform of a smartwatch or other portable electronic device. Thus, another type of media that may carry programs, media content, or metadata files includes optical, radio, and electromagnetic waves, such as those used at physical interfaces between local devices via wired and fiber optic landline networks and various airlinks. The physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that carry the software. As used herein, unless restricted to "non-transitory," "tangible," or "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions or data to a processor for execution.

[0093] Thus, machine-readable media can take the form of various forms of tangible storage media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer, etc., such as can be used to implement the client device, media gateway, transcoder, etc. shown in the accompanying drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that make up a bus within a computing system. Carrier-wave transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, diskettes, hard disks, magnetic tape, any other magnetic medium, CD-ROMs, DVDs or DVD-ROMs, any other optical medium, punched card tape, any other physical storage medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chip or cartridge, a carrier wave that transmits data or instructions, a cable or link that transmits such a carrier wave, or any other medium from which a computer can read at least one of programming code or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0094] The scope of protection is limited solely by the following claims. When interpreted in light of this specification and the subsequent prosecution history, the scope is intended and should be interpreted to be broad consistent with the ordinary meaning of the language used in the claims and to include all structural and functional equivalents. Notwithstanding the foregoing, none of the claims is intended to include subject matter that does not comply with the requirements of sections 101, 102, or 103 of the Patent Act, nor should it be interpreted in such a manner. Any unintended inclusion of such subject matter is hereby disclaimed.

[0095] Except as described above, nothing described or shown is intended or should be construed as conferring upon the public any component, step, feature, object, benefit, advantage, or equivalent, whether or not recited in the claims.

[0096] It should be understood that the terms and expressions used herein have the common meaning consistent with these terms and expressions in their corresponding respective investigation and research fields, unless otherwise specified herein. Relational terms such as first and second can be used only to distinguish one entity or action from another, without necessarily requiring or implying any actual such relationship or order between these entities or actions. The term "comprises," "comprising," "includes," "including," or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article, or device comprising or including a series of elements or steps not only includes these elements or steps, but can also include other elements or steps that are not explicitly listed or that are inherent to such process, method, article, or device. Without further limitation, an element starting with "a" or "an" does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0097] Furthermore, in the foregoing detailed description, it may be noted that various features are grouped together in various examples for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than expressly recited in each claim. Rather, as the following claims reflect, claimed subject matter encompasses less than all features of any single disclosed example. Accordingly, the following claims are hereby incorporated into the detailed description, with each claim standing on its own as separately claimed subject matter.

[0098] While the foregoing has described what is considered to be the best mode and other examples, it should be understood that various modifications may be made therein, and the subject matter disclosed herein may be embodied in a variety of forms and examples, and may be used in many applications, only some of which have been described herein. The following claims are intended to claim any and all modifications and variations that fall within the true scope of this concept.

Claims

1. A method for generating an avatar, comprising: Select the source domain based on the source generative adversarial network (GAN) training; Select the target domain to be trained on the target GAN; Training the target domain with an art dataset such that the training produces an avatar GAN; capturing an image of the subject; as well as The avatar is generated based on the image according to the avatar GAN.

2. The method according to claim 1, wherein Generating the avatar further includes: projecting the image onto a latent space associated with the source GAN, the latent space being represented by a latent code; transferring the latent code to the avatar GAN; and The avatar is rendered according to the latent code.

3. The method according to claim 1, wherein The source GAN includes a three-dimensional GAN, and The target GAN includes a two-dimensional GAN.

4. The method according to claim 1, wherein The art dataset includes: Multiple sample images, each sample image is associated with a style type and one or more attribute classifiers, The style type is an art style selected from the group consisting of comics, Pixar, cartoons, and comic strips.

5. The method according to claim 1, wherein Training the target domain further includes: calculating a statistical distribution based on a set of camera parameters associated with the source domain, and applying the statistical distribution to estimate a set of camera parameters for the target domain; applying a loss function and one or more feature regularizers; and Controls one or more geometric deformations based on a thin plate spline network.

6. The method according to claim 5, wherein: Applying the loss function and regularizer also includes: Applying the adversarial loss obtained during the training as a main loss function; Applying texture regularization to one or more layers; Applying geometric regularization according to the derived parameters associated with the latent space S; and Apply depth regularization based on the average background depth.

7. A domain adaptation framework, comprising: According to the source domain trained by the source GAN; According to the target domain of target GAN training; A head-generating generative adversarial network (GAN) generated by training the target domain on one or more art datasets; as well as An avatar generation module is configured to generate an avatar based on an image according to the avatar GAN.

8. The domain adaptation framework according to claim 7, wherein: The avatar generation module is further configured to: projecting the image onto a latent space associated with the source GAN, wherein the latent space is represented by a latent code; transferring the latent code to the avatar GAN; and The avatar is rendered according to the latent code.

9. The domain adaptation framework according to claim 7, wherein: The source GAN includes a three-dimensional GAN, Wherein, the target GAN includes a two-dimensional GAN, and The one or more art data sets include a plurality of sample images, each sample image being associated with two-dimensional data.

10. The domain adaptation framework according to claim 7, wherein: The one or more art datasets include a plurality of sample images, each sample image is associated with a style type and one or more attribute classifiers, The style type is an art style selected from the group consisting of comics, Pixar, cartoons, and comic strips.

11. The domain adaptation framework according to claim 7, further comprising: Camera alignment module; Feature regularization module; and Geometric deformation module.

12. The domain adaptation framework according to claim 11, wherein: The camera alignment module further includes: Based on a statistical distribution of a set of camera parameters associated with the source domain, wherein the statistical distribution is applied to estimate a set of camera parameters for the target domain.

13. The domain adaptation framework according to claim 11, wherein: The feature regularization module further includes: an adversarial loss obtained during said training; Texture regularization applied to one or more layers; a geometric regularization computed according to the derived parameters associated with the latent space S; and Depth regularization associated with the average background depth.

14. The domain adaptation framework according to claim 11, wherein: The geometric deformation module also includes: A thin plate spline network is configured to control one or more geometric deformations.

15. A non-transitory computer-readable medium comprising instructions for generating an avatar, the instructions, when executed by a processor, configuring the processor to perform functions comprising: Select the source domain based on the source GAN training; Select the target domain to be trained on the target GAN; Training a target domain with an art dataset such that the training produces an avatar GAN; capturing an image of the subject; as well as The avatar is generated based on the image according to the avatar GAN.

16. The medium according to claim 15, wherein Generating the avatar further includes: projecting the image onto a latent space associated with the source GAN, the latent space being represented by a latent code; transferring the latent code to the avatar GAN; and The avatar is rendered according to the latent code.

17. The medium according to claim 15, wherein The source GAN includes a three-dimensional GAN, and The target GAN includes a two-dimensional GAN.

18. The medium according to claim 15, wherein The art dataset includes: Multiple sample images, each sample image is associated with a style type and one or more attribute classifiers, The style type is an art style selected from the group consisting of comics, Pixar, cartoons, and comic strips.

19. The medium according to claim 15, wherein Training the target domain further includes: calculating a statistical distribution based on a set of camera parameters associated with the source domain, and applying the statistical distribution to estimate a set of camera parameters for the target domain; applying a loss function and one or more feature regularizers; and Controls one or more geometric deformations based on a thin plate spline network.

20. The medium according to claim 19, wherein Applying the loss function and regularizer also includes: Applying the adversarial loss obtained during the training as a main loss function; Applying texture regularization to one or more layers; Applying geometric regularization according to the derived parameters associated with the latent space S; and Apply depth regularization based on the average background depth.