Electronic device and control method thereof

By performing first and second learning on the neural network model and utilizing meta-learning and few-shot learning methods, the efficiency and realism issues of generating user conversation head video sequences in existing technologies have been solved, enabling the generation of highly realistic video sequences under conditions of limited images.

CN113544706BActive Publication Date: 2025-12-30SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080019713.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-30
Filing Date
2020-03-20
Publication Date
2025-12-30
Estimated Expiration
2040-03-20

AI Technical Summary

Technical Problem

Existing technologies have limitations in generating video sequences of user conversation heads, as they cannot realistically reflect head movement or rotation, and require a large amount of learning data and a long learning time.

Method used

By performing a first and second learning iteration on the neural network model, utilizing meta-learning and few-shot learning methods, and based on learning video sequences from multiple users and a small number of user images, the neural network model is fine-tuned to generate realistic video sequences of user conversation heads.

Benefits of technology

It enables the generation of highly realistic user conversation head video sequences with a limited number of images, improving generation efficiency and effectiveness while reducing the need for learning data and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113544706B_ABST
    Figure CN113544706B_ABST
Patent Text Reader

Abstract

An electronic device and a control method thereof are provided. A control method of an electronic device according to the disclosure includes performing first learning on a neural network model based on a plurality of learning video sequences including talking heads of a plurality of users to obtain a video sequence including a talking head of a random user, performing second learning to fine-tune the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image, and obtaining a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model on which the first learning and the second learning are performed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to an electronic device and a method for controlling the electronic device, and for example to an electronic device capable of acquiring a video sequence including the head of a user's conversation based on a relatively small number of images, and a method for controlling the device thereon. Background Technology

[0002] Recently, with the development of the field of artificial intelligence models, technologies related to conversation head models that can generate video sequences that express what it looks like when a user is talking have attracted attention.

[0003] As a conventional technique, there are methods to generate video sequences including conversational heads by warping static frames. This technique allows for the acquisition of video sequences using only a small number of images, similar to a single image. However, warping-based techniques have limitations, such as the inability to realistically reflect head movements or rotations.

[0004] As a conventional technique, Generative Adversarial Networks (GANs) exist to generate video sequences including conversational headlines, and this technique can produce highly realistic video sequences. However, so far, limitations have been identified in GAN-based techniques, such as the need for large amounts of training data and long training times.

[0005] Therefore, there is a growing demand for techniques that utilize a relatively small number of user images not included in the learning data to acquire video sequences including highly realistic conversational headshots. In particular, there is a significant need for neural network model structures and learning methods that efficiently and effectively train neural network models capable of generating video sequences including conversational headshots. Summary of the Invention

[0006] Technical issues

[0007] Embodiments of this disclosure provide an electronic device capable of acquiring realistic video sequences using a small number of images including the head of a particular user in conversation, and a method for controlling such a device.

[0008] Solution to the problem

[0009] According to an example embodiment of this disclosure, a method for controlling an electronic device includes: performing a first learning on a neural network model based on a plurality of learning video sequences including the conversation heads of a plurality of users to obtain a video sequence including the conversation heads of random users; performing a second learning on at least one image including the conversation head of a first user different from the plurality of users and first landmark information included in the at least one image to fine-tune the neural network model; and using the neural network model that has performed the first and second learning to obtain a first video sequence including the conversation head of the first user based on at least one image and pre-stored second landmark information.

[0010] According to one example embodiment, an electronic device includes: a memory storing at least one instruction; and a processor configured to execute the at least one instruction. By executing the at least one instruction, the processor is configured to: perform a first learning on a neural network model based on multiple learning video sequences including the heads of conversations of multiple users to obtain a video sequence including the heads of conversations of random users; perform a second learning on at least one image including the head of conversation of a first user different from the multiple users and first landmark information included in the at least one image to fine-tune the neural network model; and use the neural network model that has performed the first and second learning to obtain a first video sequence including the head of conversation of the first user based on the at least one image and pre-stored second landmark information.

[0011] According to one example embodiment, a non-transitory computer-readable recording medium has a program recorded thereon, which, when executed by a processor of an electronic device, causes the electronic device to perform operations including: performing a first learning on a neural network model based on multiple learning video sequences including the heads of conversations of multiple users to obtain a video sequence including the heads of conversations of random users; performing a second learning based on at least one image including the head of conversation of a first user different from the multiple users and first landmark information included in the at least one image to fine-tune the neural network model; and using the neural network model that has performed the first and second learning to obtain a first video sequence including the head of conversation of the first user based on at least one image and pre-stored second landmark information. Attached Figure Description

[0012] The above and other aspects, features, and advantages of certain embodiments of this disclosure will become clearer from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0013] Figure 1 This is a flowchart illustrating an example method of controlling an electronic device according to an embodiment of the present disclosure;

[0014] Figure 2This is a flowchart illustrating an example of a second learning process according to an embodiment of the present disclosure;

[0015] Figure 3 This is a flowchart illustrating an example first learning process according to an embodiment of the present disclosure;

[0016] Figure 4 This is a diagram illustrating an example architecture of an example neural network model according to an embodiment of the present disclosure, and example operations performed by an embedder, a generator, and a discriminator;

[0017] Figure 5 This is a block diagram illustrating an example configuration of an example electronic device according to an embodiment of the present disclosure; and

[0018] Figure 6 This is a block diagram illustrating an example configuration of an example electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0019] Various modifications can be made to the various exemplary embodiments of this disclosure, and various types of embodiments may exist. Therefore, specific embodiments will be shown in the accompanying drawings, and the embodiments will be described in detail in the disclosure. However, it should be noted that the various embodiments are not intended to limit the scope of this disclosure to the specific embodiments, but should be understood to include various modifications, equivalents, and / or alternatives to the embodiments of this disclosure. Furthermore, in relation to the description of the drawings, similar parts may be denoted by similar reference numerals.

[0020] In describing this disclosure, detailed descriptions may be omitted if it is determined that a detailed explanation of a relevant known function or component would unnecessarily obscure the main points of this disclosure.

[0021] Furthermore, the following embodiments can be modified in various different ways, and the scope of the technical spirit of this disclosure is not limited to the example embodiments described below.

[0022] The terminology used in this disclosure is for the purpose of explaining specific embodiments of the disclosure and is not intended to limit the scope of the disclosure. Furthermore, unless the context clearly indicates otherwise, singular expressions include plural expressions.

[0023] In this disclosure, expressions such as “have,” “may have,” “include,” and “may include” should be understood to indicate the presence of such characteristics (e.g., elements such as numerical values, functions, operations, and components), and such expressions are not intended to exclude the presence of additional characteristics.

[0024] In this disclosure, expressions such as “A or B”, “at least one of A and / or B”, or “one or more of A and / or B” can include all possible combinations of the listed items. For example, “A or B”, “at least one of A and B”, or “at least one of A or B” can refer to all of the following cases: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B.

[0025] Furthermore, expressions such as "first," "second," etc., used in this disclosure can be used to describe various elements, regardless of any order and / or degree of importance. Additionally, such expressions can be used to distinguish one element from another and are not intended to limit the elements.

[0026] In this disclosure, the description of one element (e.g., the first element) being "(operationally or in communication) coupled to" or "connected to" another element (e.g., the second element) should be understood to include both cases where one element is directly coupled to another element and cases where one element is coupled to another element through yet another element (e.g., the third element).

[0027] On the other hand, a description of an element (e.g., a first element) being "directly coupled" or "directly connected" to another element (e.g., a second element) does not include another element (e.g., a third element) between the two elements.

[0028] Furthermore, the expression "configured as" as used in this disclosure may be used interchangeably with other expressions such as "suitable for," "capable of," "designed to," "suitable for," "manufactured as," and "capable of doing," depending on the context. The term "configured as" may not necessarily mean, for example, that the device is "specifically designed for..." in terms of hardware.

[0029] In some cases, the expression "a device configured as..." can mean, for example, that the device is "capable" of performing operations together with another device or component. For example, the phrase "a processor configured to perform A, B, and C" can mean, for example, a dedicated processor (e.g., an embedded processor) for performing the corresponding operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform the corresponding operations by executing one or more software programs stored in a memory device. However, this disclosure is not limited thereto.

[0030] In embodiments of this disclosure, a 'module' or 'part' performs at least one function or operation, and these elements may be implemented as hardware or software, or as a combination of hardware and software. Additionally, besides 'modules' or 'parts' that need to be implemented as specific hardware, multiple 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor (not shown).

[0031] Various elements and areas may be schematically shown in the accompanying drawings. Therefore, the technical spirit of this disclosure is not limited to the relative sizes or spacing depicted in the drawings.

[0032] In the following, various exemplary embodiments according to the present disclosure will be described in more detail with reference to the accompanying drawings.

[0033] Figure 1 This is a flowchart illustrating an example method of controlling an electronic device according to an embodiment of the present disclosure.

[0034] The electronic device according to this disclosure can use a neural network model to acquire video sequences including the head of a user's conversation. The neural network model can, for example, refer to an artificial intelligence model including an artificial neural network, and therefore, the term neural network model can be used interchangeably with the term artificial intelligence model. For example, the neural network model can be a generative adversarial network (GAN) model including an embedder, a generator, and a discriminator, and can be configured to acquire video sequences including the head of a random user's conversation. The head of the conversation can, for example, refer to the portion of the head that represents the appearance of a user's conversation.

[0035] refer to Figure 1 The electronic device can perform a first learning operation on the neural network model according to this disclosure. For example, at operation S110, the electronic device can perform a first learning operation on the neural network model based on multiple learning video sequences including the conversation heads of multiple users to obtain a video sequence including the conversation heads of random users.

[0036] The initial learning can be performed using, for example, meta-learning methods. Meta-learning can refer, for example, to methods that automate the machine learning process and thereby enable the machine to learn learning rules (meta-knowledge) on its own. For example, meta-learning can refer to learning a learning method. For the initial learning, multiple learning video sequences are used, including conversational headlines from multiple users, and the learning video sequences refer to the video sequences used for the initial learning.

[0037] For example, the first learning according to this disclosure can be performed by: performing a training task in a manner similar to the second learning described below, and learning generalized rules to obtain video sequences including the heads of random users' conversations. During the first learning process, the input and output values ​​of each of the embedder, generator, and discriminator according to this disclosure, as well as multiple parameters, can be learned, and based on these, the neural network model can obtain video sequences including the heads of random users' conversations. References will be made below. Figure 3 and Figure 4 Let me describe the first learning experience in more detail.

[0038] The electronic device can perform a second learning operation on the neural network model according to this disclosure. For example, the electronic device can perform the second learning operation based on at least one image including the head of a user (hereinafter referred to as the first user) who is not included in the aforementioned plurality of learning video sequences. For example, at operation S120, the electronic device can perform the second learning operation to fine-tune the neural network model based on at least one image including the head of the first user (not included in the plurality of learning video sequences) and first landmark information included in the at least one image.

[0039] The second learning can be performed using, for example, a few-shot learning method. Few-shot learning can, for example, refer to a method of efficiently training a neural network model with a small amount of data. For the second learning, first landmark information can be used, and landmark information can, for example, refer to information about the main features of the user's face included in the image, and the first landmark information can refer to landmark information included in at least one image including the talking head of the first user.

[0040] For example, the second learning according to this disclosure may include the following process: after learning generalized rules for acquiring video sequences including the head of a random user's conversation through the aforementioned first learning, a neural network model is fine-tuned based on at least one image including the head of a first user's conversation to personalize the learning for the first user. During the second learning, the parameter set of the generator according to this disclosure may be fine-tuned to match at least one image including the head of the first user's conversation, and based on this, the neural network model can acquire video sequences including the head of the first user's conversation. Reference will be made below. Figure 2 and Figure 4 Let me describe the second learning session in more detail.

[0041] When performing the first learning and the second learning as described above, the electronic device can use the neural network model that has performed the first learning and the second learning at operation S130 to acquire a first video sequence including the head of the first user's conversation based on at least one image and pre-stored second landmark information.

[0042] For example, the electronic device can input at least one image and first landmark information into an embedder performing the first learning, and obtain a first embedding vector. The first embedding vector may be, for example, a vector of the Nth dimension obtained by the embedder, and may include information about the identity of the first user. Upon obtaining the first embedding vector, the electronic device can input the first embedding vector and pre-stored second landmark information into a generator, and obtain a first video sequence including the first user's head during conversation. The second landmark information may be, for example, pre-stored landmark information obtained from multiple images included in multiple learning video sequences.

[0043] Finally, the electronic device according to this disclosure can, for example, perform a first learning based on a plurality of learning video sequences according to a meta-learning method, thereby enabling the neural network model to acquire a video sequence including the head of a conversation of a random user, and perform a second learning based on at least one image of a first user who is a new user not included in the plurality of learning video sequences according to a few-shot learning method, thereby fine-tuning the neural network model to achieve personalization for the first user, and using the neural network model that has performed the first and second learning to acquire a video sequence including the head of a conversation of the first user.

[0044] According to various exemplary embodiments of this disclosure, an electronic device can efficiently and effectively train a neural network model capable of generating video sequences including conversational heads, and thereby acquire video sequences including highly realistic conversational heads using a relatively small number of images of the user not included in the learning data.

[0045] The first learning process, the second learning process, and the process of acquiring video sequences according to various example embodiments of the present disclosure will be described in more detail below.

[0046] Figure 2 This is a flowchart illustrating an example of a second learning process according to an embodiment of the present disclosure.

[0047] The above reference will be explained in more detail. Figure 1 The second learning step S120 discussed includes multiple operations. As described above, the second learning process according to this disclosure can be performed using a few-shot learning method, and can, for example, refer to a process for acquiring a video sequence of a user's conversational head not included in the learning video sequence of the first learning process. As described above, the first learning is performed before the second learning process. However, in the following, in order to explain the main features according to this disclosure in more detail, the second learning process will be explained first, followed by the first learning process.

[0048] like Figure 2As shown, the electronic device according to this disclosure can acquire at least one image including the head of a first user in conversation at operation S210. The at least one image can be, for example, 1 to 32 images. This merely indicates that the second learning according to this disclosure is performed using a small number of images, and the number of images used for the second learning according to this disclosure is not limited thereto. Even if the second learning is performed based on a single image, a video sequence including the head of the first user in conversation can be acquired. As the number of images used for the second learning increases, the degree of personalization for the first user may also increase.

[0049] Upon acquiring at least one image, the electronic device can acquire first landmark information based on the at least one image at operation S220. For example, the landmark information may include information about the head pose included in the image and information about the analog descriptor, and in addition, may include various information related to various characteristics included in the talking head. According to one embodiment of this disclosure, the landmarks can be rasterized into a 3-channel facial landmark image using, for example, a predefined color set, to connect specific landmarks with line segments.

[0050] Upon acquiring the first landmark information, the electronic device can, at operation S230, input at least one image and the first landmark information into the embedder that has performed the first learning, and acquire a first embedding vector including information about the identity of the first user. For example, the first embedding vector may include information about unique characteristics of the first user included in the first user's talking head, but may be independent of the first user's head posture. During the second learning process, the electronic device can acquire the first user's first embedding vector based on the parameters of the embedder acquired in the first learning.

[0051] Upon acquiring the first embedding vector, the electronic device can, at operation S240, fine-tune the generator's parameter set based on the first embedding vector to match at least one image including the head of a conversation of the first user. The features of fine-tuning the generator's parameter set may, for example, refer to fine-tuning the generator and the neural network model including the generator, and this may, for example, refer to optimizing the generator to correspond to the first user.

[0052] For example, the fine-tuning process may include instantiating the generator based on the generator's parameter set and a first embedding vector. Instantiation may be performed, for example, by a method such as Adaptive Instance Normalization (AdaIN).

[0053] For example, in the second learning, the neural network model according to this disclosure can learn not only individual general parameters related to characteristics generalized for multiple users, but also individual specific parameters.

[0054] The following will refer to Figure 4A more detailed description of the second learning method is provided, such as the detailed method for few-shot learning for fine-tuning neural network models according to this disclosure.

[0055] Figure 3 This is a flowchart illustrating an example first learning process according to an embodiment of the present disclosure.

[0056] The reference will be explained in more detail below. Figure 1 The explanation of the first learning step S110 includes multiple operations. As described above, the first learning process according to this disclosure can be performed by a meta-learning method, and can, for example, refer to a process for obtaining a video sequence of conversational heads including random users based on multiple learning video sequences including conversational heads of multiple users. As will be described in more detail below, in the first learning process, multiple parameters of the neural network model can be learned by an adversarial method.

[0057] like Figure 3 As shown, the electronic device according to this disclosure can acquire at least one learning image from a learning video sequence including the head of a conversation of a second user in a plurality of learning video sequences at operation S310. For example, the first learning can be performed by a K-sample learning method, which may refer, for example, to a method of performing learning by acquiring a number of K random image frames from one of the plurality of learning video sequences. The second user may refer, for example, to a specific user among a plurality of users included in the plurality of learning video sequences, and may be distinguished from the term first user, which is used to refer to a user not included in the plurality of learning video sequences.

[0058] When at least one learning image has been acquired from a learning video sequence including the head of a second user in conversation, the electronic device can, at operation S320, acquire third landmark information of the second user based on the acquired at least one learning image. The third landmark information may, for example, refer to landmark information included in at least one image including the head of the second user in conversation. For example, the third landmark information may include the same information as the first landmark information, since the third landmark information is information about the main features of the face of a specific user included in the image, but differs from the first landmark information in that the third landmark information is about the second user, not the first user.

[0059] Upon acquiring the third landmark information, the electronic device can, at operation S330, input at least one learning image acquired from the learning video sequence including the second user's conversational head, along with the third landmark information, into the embedder, and acquire a second embedding vector including information about the identity of the second user. The second embedding vector may, for example, refer to a vector of the Nth dimension acquired by the embedder, and may include information about the identity of the second user. For example, the second embedding vector may be specified by distinguishing the vector from a first embedding vector that includes information about the identity of the first user. For example, the second embedding vector may be the result of averaging embedding vectors acquired from randomly sampled images (e.g., frames) from the learning video sequence.

[0060] Upon acquiring the second embedding vector, the electronic device can instantiate the generator at operation S340 based on the generator's parameter set and the second embedding vector. Once the generator is instantiated, the electronic device can input the third landmark information and the second embedding vector into the generator at operation S350, and acquire a second video sequence including the second user's conversation header. The parameters of the embedder and the generator can be optimized to minimize and / or reduce an objective function, which includes, for example, a content loss term, an adversarial term, and an embedding matching term, as described in more detail below.

[0061] Upon acquiring the second video sequence, the electronic device can update the parameter sets of the generator and the embedder at operation S360 based on the degree of similarity between the second video sequence and the learning video sequence. For example, the electronic device can obtain a fidelity score for the second video sequence through a discriminator and update the parameter sets of the generator, the embedder, and the discriminator based on the obtained fidelity score. For example, the discriminator's parameter set can be updated to improve the fidelity score for the learning video sequence and decrease the fidelity score for the second video sequence.

[0062] The discriminator can be, for example, a projection discriminator that obtains a fidelity score based on a third embedding vector, which is different from the first and second embedding vectors. The third embedding vector can be, for example, a vector of the Nth dimension obtained by the discriminator, and can be distinguished from the first and second embedding vectors obtained by the embedder. For example, the third embedding vector can correspond to each of at least one image including the second user's head during conversation, and can include information related to the fidelity score for the second video sequence. During the first learning step, the difference between the second and third embedding vectors can be penalized, and the third embedding vector can be initialized based on the first embedding vector at the start of the second learning step. In other words, during the first learning step, the second and third embedding vectors can be learned to be similar to each other. Furthermore, at the start of the second learning step, the third embedding vector can be used when initializing the third embedding vector based on the first embedding vector, which includes information related to the identity of the first user that was not used in the learning of the first learning step.

[0063] The following will refer to Figure 4 A more detailed description of the first-learning method is provided, such as the detailed method for performing meta-learning on a neural network model according to this disclosure.

[0064] Figure 4 This is a diagram illustrating an example architecture of a neural network model according to an embodiment of the present disclosure, as well as example operations performed by an embedder, a generator, and a discriminator.

[0065] In the following text, before explaining the disclosure in more detail, several conventional techniques related to this disclosure will be explained, and a neural network model according to this disclosure for overcoming the limitations of conventional techniques will be explained. In explaining the neural network model, the architecture of the neural network model and methods for implementing various embodiments according to this disclosure will be explained in more detail.

[0066] This disclosure relates to a method for synthesizing highly realistic (photorealistic) and personalized conversational head models, such as highly realistic video sequences with high fidelity in their speech expression and simulation of a particular individual. For example, this disclosure relates to a method for synthesizing highly realistic and personalized head images given a set of facial landmarks that enable the animation of the model. This method can be practically applied to telepresence, including not only video conferencing and multiplayer games, but also the special effects industry.

[0067] It is known that synthesizing highly realistic conversational head sequences is difficult for two reasons. First, the human head has high photometric, geometric, and dynamic complexity. This complexity can occur not only in facial modeling, where numerous access modeling methods exist, but also in modeling the mouth, hair, and clothing. The second reason for this complexity is the sensitivity of the human visual system to minute errors that may occur in modeling the appearance of the human head (the so-called uncanny valley effect [Reference 24], in the following sections). Figure 6 The explanation will include a list of references cited in this disclosure. Because the tolerance for modeling errors is low, as stated above—that is, because the human visual system is very sensitive—users may still experience strong aversion even if there are slight differences between the conversational head sequence and the actual human face. In such cases, using an unrealistic avatar may actually create a better impression. Therefore, many current remote conferencing systems are using avatars that resemble unrealistic cartoon characters.

[0068] As a conventional technique for overcoming the aforementioned tasks, there are methods to synthesize connected head sequences by distorting single or multiple static frames. All distorted scenes synthesized using traditional distortion algorithms [References 5, 28] and machine learning (including deep learning) [References 11, 29, 40] can be used for this purpose. Distortion-based systems can generate conversational head sequences from a small number of images, such as a single image, but they are limited by factors such as de-occlusion, computational complexity, and head rotation without human intervention.

[0069] As a conventional technique, there are methods for directly (without distortion) synthesizing video frames using adversarially trained deep convolutional networks (ConvNets) [References 16, 20, 37]. However, for this approach to succeed, large-scale networks must be trained, with each of the generator and discriminator having tens of millions of parameters about the talking head. Therefore, generating a recently personalized talking head model from this system requires not only a massive dataset of videos or photos [Reference 16] several minutes long [References 20, 37], but also several hours of GPU training. This training data and time are arguably excessive for most real-world telepresence scenarios, where users need to generate personalized head models with minimal effort, but this requirement is lower than that of systems using complex physical and optical modeling to construct highly realistic head models [Reference 1].

[0070] As a conventional technique, there are methods for statistical modeling of the appearance of the human face [Reference 6], and specifically, there are methods using classical techniques [Reference 35], and more recently, there are methods using deep learning [References 22, 25]. Facial modeling and conversational head modeling are highly related, but conversational head modeling involves modeling the hair, neck, mouth, and often non-facial parts such as the shoulders / upper body, and therefore, facial modeling and conversational head modeling are not the same. These non-facial parts may not be handled by simply extending facial modeling methods, and this is because non-facial parts are less suitable for registration and tend to have higher variability and complexity than facial parts. In principle, the results of facial modeling [Reference 35] or lip modeling [Reference 31] can be stitched to a head video. However, in the case of this approach, the rotation of the head may not be fully controlled in the final video, and therefore, a true conversational head system may not be provided.

[0071] If Model-Independent Meta-Learning (MAML) [Reference 10], a conventional technique, is used, then the initial state of the image classifier is obtained, and based on this, the image classifier can be flexibly adapted to unseen classifications even with a small number of training samples. This approach can be used according to the method of this disclosure, but it differs in its implementation. Meanwhile, various methods exist for combining adversarial training and meta-learning. Data augmentation GANs [Reference 3], meta-GANs [Reference 43], and adversarial meta-learning [Reference 41] can use adversarially trained networks in the meta-learning step to generate additional images for unseen classifications. Such methods primarily focus on improving few-shot classification performance, but the method of this disclosure addresses training the image generation model using adversarial targets. In summary, adversarial fine-tuning can be introduced into the meta-learning framework in this disclosure. Fine-tuning can be applied after the initial state of the generator, and the discriminator network can be obtained through the meta-learning step.

[0072] As conventional techniques, there are two recent methods related to text-to-speech generation [References 4, 18]. The setup of these methods (few-shot learning of the generative model) and some components (independent embedder networks, fine-tuning generators) can also be used in this disclosure. However, this disclosure differs from conventional techniques in at least the following aspects: its application domain, the use of adversarial learning, the specific application of the meta-learning process, and the details of various implementation methods.

[0073] According to this disclosure, a method for generating conversational head models from a few photographs is provided (so-called few-shot learning). The method of this disclosure can generate reasonable results with a single photograph (single-shot learning), but the degree of personalization can be further improved by adding a few photographs. Similar to references 16, 20, and 37, the conversational heads generated by the neural network model according to this disclosure can be generated by a deep ConvNet, which synthesizes video frames directly through a series of convolutional operations rather than warping. Therefore, the conversational heads generated according to this disclosure can handle large pose variations beyond the capabilities of warp-based systems.

[0074] Few-shot learning ability can be acquired through extensive pre-training (meta-learning) on ​​a large corpus of conversational head-view videos corresponding to multiple users with diverse appearances who are different from each other. In the meta-learning process according to this disclosure, the method can learn to simulate a few-shot learning task and converge the landmark position to a highly realistic and personalized image based on a small training set. Subsequently, a small number of images of new users may introduce new adversarial learning problems to the discriminator, which has been pre-trained using a large-capacity generator and meta-learning. These new adversarial learning problems can converge to a state that generates highly realistic and personalized images after several training steps.

[0075] The architecture of the neural network model according to this disclosure can be implemented using at least some of the results from recent developments in image generation modeling. For example, adversarial training [Reference 12] and methods such as those used for conditional discriminators [Reference 23] including a projection discriminator [Reference 32] can be used according to the architecture of this disclosure. The meta-learning step can use, for example, the Adaptive Instance Normalization (AdaIN) mechanism [Reference 14], which is shown as applicable to large-scale conditional generation tasks [References 2, 34]. Therefore, according to this disclosure, the quality of synthesized images can be improved, and the uncanny valley effect can be eliminated and / or reduced from synthesized images.

[0076] According to the meta-learning steps of this disclosure, the availability of a number M video sequences (e.g., learning video sequences) comprising the conversational heads of multiple users distinct from each other can be assumed. xi indicates the i-th video sequence, and xi(t) indicates the t-th video frame of the video sequence. It can be assumed, not only during meta-learning but also during test time, that the positions of facial landmarks are available for all frames (e.g., standard facial alignment codes [Reference 7] can be used to obtain the positions of facial landmarks). Landmarks can be rasterized into 3-channel images (e.g., facial landmark images) using a predefined color set to connect specific landmarks with line segments. yi(t) indicates the resulting facial landmark image calculated relative to xi(t).

[0077] like Figure 4 As shown, the meta-learning architecture according to this disclosure may include: an embedder network that maps a head image (with estimated facial landmarks) to an embedding vector including pose-independent information; and a generator network that maps input facial landmarks to output frames via a set of convolutional layers modulated by the embedding vectors via adaptive instance normalization (AdaIN). Generally, during the meta-learning step, a set of frames obtained from the same video can be passed through the embedder network, and the resulting embeddings can be averaged and used to predict adaptive parameters of the generator network. Subsequently, an image generated after the landmarks of another frame are passed through the generator network can be compared with the ground truth. The objective function may include perceptual loss and adversarial loss. The adversarial loss can be implemented using a conditional projection discriminator network. The meta-learning architecture according to this disclosure and its corresponding operation will be described in more detail below.

[0078] In the meta-learning operation according to this disclosure, the following three networks (commonly referred to as adversarial networks or generative adversarial networks (GANs)) can be trained (see reference). Figure 4 ).

[0079] 1. Embedder E(x) i (S), y i (S); φ). The embedder can be configured to acquire video frames x i (S), and associated facial landmark images y i (S), and map these inputs to embedded N-dimensional vectors. Video frame x i (S) can be obtained from a sequence of training videos, i.e., multiple training video sequences, where the conversation head model includes conversation head images of multiple users, different from random users that will be synthesized in the future. φ indicates the embedder parameters learned during the meta-learning step. Generally, the goal of the meta-learning step for the embedder E is to learn φ such that the embedding is an N-dimensional vector. This includes video-specific information (such as human identity), which does not change with pose and simulated objects in a particular frame. The embedded N-dimensional vector calculated by the embedder is transcribed as...

[0080] 2. Generator The generator can be configured to work with the corresponding N-dimensional embedding vector computed by the embedder E. Get unseen video frames x i (t) and facial landmark image y i (t), and generate synthesized video frames. The generator G can be trained to maximize and / or increase the output (e.g., synthesize video frames). The similarity between the generator G and the corresponding real-world frame. The parameters of the generator G can be divided into two groups, for example, general parameters ψ and specific parameters. During the meta-learning step, individual-specific parameters Trainable projection matrices can be used during the fine-tuning steps of meta-learning (described in more detail below). From embedded N-dimensional vectors The prediction is performed in the model, while the individual general parameter ψ is trained directly.

[0081] 3. Discriminator D(x) i (t), y i (t); i; θ, W, w o b) The discriminator can be configured to acquire the input video frame x i (t), associated facial landmark image y i (t) and the index i of the learned video sequence, and compute the fidelity score r (a single scalar). θ, W, w o , b indicates the discriminator parameters learned during the meta-learning step. The discriminator may include the convolutional network (ConvNet) part V(x) i (t), y i (t); θ), the portion is configured to input video frame x i (t) and associated facial landmark image y i (t) is mapped to an N-dimensional vector. Then, the discriminator can calculate the fidelity score r based on the N-dimensional vector and the discriminator parameters W, w0, b. The fidelity score r indicates the input video frame x. i (t) Whether it is the actual (e.g., non-synthetic) video frame of the i-th learned video sequence, and the input video frame x i (t) Whether it is associated with the facial landmark image y i (t) Matching. Video frame x input to the discriminator. i (t) can be a composite video frame However, the discriminator does not know the input video frame. It is a fact that the video frames are synthesized.

[0082] During the meta-learning operation of the example method, the parameters of all three networks can be trained using an adversarial approach. This can be performed by simulating a K-sample learning phase. K can be, for example, 8, but is not limited to this, and can be chosen to be greater than or less than 8 depending on the performance of the hardware used in the meta-learning step, or the accuracy of the images generated by the meta-learned GAN and the purpose of the meta-learning of such GAN. In each phase, a learning video sequence i and a single ground truth video frame x from the sequence can be randomly extracted. i (t). Except for x i In addition to (t), K(s1, s2, ..., s) can be extracted from the same learning video sequence i. K The additional video frames are then used. Subsequently, at the embedder E, an N-dimensional embedding vector can be computed relative to the number of K additional video frames. By averaging, and therefore, an N-dimensional embedding vector can be computed relative to the learned video sequence i.

[0083]

[0084] At the generator G, the calculated embedded N-dimensional vector e can be computed. i And synthesized video frames (For example, the reconstruction of the t-th frame):

[0085]

[0086] The parameters of the embedder E and the generator G can be optimized to minimize and / or reduce the content loss term L. CNT , and the opposing item L ADV and embedded matching item L MCH Objective function:

[0087] L(φ,ψ,P,θ,W,w0,b)=L CNT (φ, ψ, P) + L ADV (φ,ψ,P,θ,W,w0,b)+L MCH (φ,W) (3)

[0088] In formula (3), the content loss term L CNT Perceptual similarity metrics can be used to measure real-world video frames x i (t) and the synthesized video frame The gap between them [Reference 19]. As an example, a perceptual similarity measure can be used, corresponding to the VGG19 [Reference 30] network trained relative to ILSVRC classification and the VGGFace [Reference 27] network trained to verify faces. However, any perceptual similarity measure known from conventional techniques can be used in this disclosure, and therefore, this disclosure is not limited to the foregoing examples. In the case where the VGG19 and VGGFace networks are used as perceptual similarity measures, the content loss term L CNT It can be calculated as a weighted sum of the L1 loss in the network's features.

[0089] The opposing term L in formula (3) ADV This can correspond to the following: the fidelity score r calculated by the discriminator D, which needs to be maximized and / or increased, and the feature matching term L, which was originally used as a perceptual similarity measure calculated using the discriminator. FM [Reference 38] (This can improve the stability of meta-learning):

[0090]

[0091] According to the projection discriminator access method [Reference 32], the columns of matrix w can include embedded N-dimensional vectors corresponding to individual videos. The discriminator D can first (e.g., input video frame x) i (t), associated facial landmark image y i (t) and the index i) of the learning video sequence are input into the N-dimensional vector V(x) i (t), y i (t); θ), and the fidelity score r is calculated as follows:

[0092]

[0093] Here, W i Indicates the i-th column of matrix W. Also, since W0 and b do not depend on the video index, such items can be compared with... The general degree of authenticity and the relationship with facial landmark images y i The compatibility of (t) corresponds to this.

[0094] Therefore, in the method according to this disclosure, there can be two types of embedded N-dimensional vectors: for example, the vector calculated by the embedder E and the vector corresponding to the column of matrix W at the discriminator D. The matching term L in formula (3) above... MCH (φ, W) can be used for The L1 difference between Wi is penalized, and the similarity between the two types of embedded N-dimensional vectors is increased.

[0095] As the parameters φ of the embedder E and ψ of the generator G are updated, the parameters θ, W, w0, b of the discriminator D can also be updated. The updates can be driven by minimizing and / or reducing the hinge loss objective function (6) as follows, and this can facilitate updates for real (e.g., non-spoof) video frames x. i The increase in the fidelity score r of (t) and for synthetic (e.g., fake) video frames The reduction in the realism score:

[0096]

[0097] Therefore, according to formula (6), the neural network model of this disclosure can identify fake instances. With actual example x i The degree of authenticity of (t) is compared, and the discriminator parameters are updated, thereby making the scores less than -1 and greater than +1 accordingly. Meta-learning can be performed by alternately updating the embedder E and the generator G to minimize the loss L. CNT L ADV and L MCH And update the discriminator D to minimize the loss L. DSC .

[0098] During meta-learning convergence, the neural network model according to this disclosure can be additionally trained during the meta-learning operation to synthesize a conversational head model for new, unseen users. Synthesis may be conditioned on facial landmark images. The neural network model according to this disclosure can be trained using a few-shot method, assuming facial landmark images, for which a number of training images of T (x(1), x(2), ..., x(T)) are provided (e.g., T frames from the same video), and y(1), y(2), ..., y(T) correspond to the training images. Here, the number of frames T need not be the same as K used in the meta-learning step. The neural network model according to this disclosure can generate reasonable results based on a single photograph (single-shot learning, T = 1), and the degree of personalization can be further improved if several additional photographs are added (few-shot learning, T > 1). For example, T can cover a range, for example, from 1 to 32. However, this disclosure is not limited thereto, and T can be selected in various ways depending on the performance of the hardware used for few-shot learning, the accuracy of the images generated by the GAN with few-shot learning after meta-learning, and the purpose of few-shot learning (e.g., fine-tuning) of the GAN after meta-learning.

[0099] The meta-learned embedding E can be used to compute the N-dimensional embedding vector of a new individual. In few-shot learning, the conversation head of the new individual will be synthesized. For example, The calculation can be performed as follows:

[0100]

[0101] In the meta-learning step, the parameters φ of the previously acquired embedder E can be reused. A simple way to generate new synthetic frames in response to new landmark images is to use not only the projection matrix P but also the computed embedding N-dimensional vector. The generator G is applied with meta-learned parameters ψ. However, in this case, it has been found that while the synthesized conversational head images look believable and realistic, there is an unacceptably large identity gap for most applications aimed at synthesizing personalized conversational head images.

[0102] Such identity gaps can be overcome according to this disclosure through a fine-tuning process. This fine-tuning process may appear to be a simplified version of meta-learning, which is performed based on a single video sequence and a small number of frames. For example, the fine-tuning process may include the following components:

[0103] 1. Generator This can now be replaced by a generator G'(yy(t); ψ, ψ′). Just as in the meta-learning step, the generator G' can be configured to acquire the facial landmark image y(t) and generate the synthesized video frames. Importantly, the individual-specific generator parameters, currently transcribed as ψ′, can be directly optimized along with the individual-general parameters ψ during the few-shot learning step. The N-dimensional embedding vector obtained in the meta-learning step... The projection matrix P can still be used to initialize the individual-specific generator parameters ψ′ (e.g., ).

[0104] 2. In the meta-learning operation, the discriminator D'(x(t), y(t); θ, w′, b) can be configured to compute the fidelity score r as before. The parameters θ and bias b of the ConvNet part v(x(r), y(t); θ) of the discriminator D' can be initialized with the same parameters θ and b obtained in the meta-learning step. The initialization of w′ will be explained below.

[0105] During the fine-tuning step, the fidelity score r of the discriminator D' can be obtained in a manner similar to that of the meta-learning step:

[0106]

[0107] As can be seen from the comparison of formulas (5) and (8), the role of vector w′ in the fine-tuning step may be related to that of vector W. i +w0 has the same role in the meta-learning step. During the initialization of w′ in the few-shot learning step, W...i Similar quantities may not be applicable to new individuals. This is because the video frames for new individuals were not used in the meta-learning training dataset. However, the matching term L in the meta-learning process... MCH This ensures the similarity between the discriminator's embedded N-dimensional vector and the embedded N-dimensional vector computed by the embedder. Therefore, w′ can be initialized to w0 and w′ in the few-shot learning step. The sum of .

[0108] When a new learning problem is defined, the loss function for the fine-tuning step can be directly derived from the meta-learning variables. Therefore, the individual-specific parameters ψ′ and the individual-general parameters ψ of the generator G' can be optimized to minimize the following simplified objective function:

[0109] L′(ψ,ψ′,θ,w′,b)=L′ CNT (ψ,ψ′)+L′ ADV (ψ,ψ′,θ,w′,b) (9)

[0110] Here, t∈{1...T} is the number of training samples.

[0111] The parameters θ and w of the discriminator NEW b can be optimized by minimizing the same hinge loss as in (6):

[0112]

[0113]

[0114] In most cases, a finely tuned generator can provide more suitable results for learning video sequences. Initializing all parameters through a meta-learning step is also crucial. As discovered through experiments, in this initialization, a highly realistic conversational head is first input, thereby enabling the neural network model according to this disclosure to extrapolate and predict highly realistic images relative to various head poses and facial expressions.

[0115] Generator Networks It could be based on the image-to-image transformation architecture proposed by Johnson et al. [Reference 19], but the downsampling and upsampling layers could be replaced by residual blocks through instance normalization [References 2, 15, 36]. Individual-specific parameters The adaptive instance normalization technique known in the relevant technical field [Reference 14] is used as the affine coefficients of the instance normalization layer, but it is still possible to use the facial landmark image y i (t) is a normalized layer for the regular (non-adaptive) instance of the downsampled block being encoded.

[0116] Embedder E(x) can be used i (s), y i (s); φ) and the ConvNet part V(x) of a discriminator-like network that includes the remaining downsampled blocks (the same as those used in the generator, but excluding the normalization layer). i (t), y i (t); θ). The discriminator network has an additional residual block at its end compared to the embedder, and they can operate at a 4x4 spatial resolution. To obtain vectorized outputs in both networks, global cumulative pooling for the spatial dimension can be performed before the rectified linear unit (ReLU).

[0117] Spectral normalization [Reference 33] can be used for all convolutional and fully connected layers in all networks. Alternatively, self-attention blocks [References 2, 42] can be used. These self-attention blocks can be inserted into all downsampled portions of the network at a 32x32 spatial resolution and into the upsampled portions of the generator at a 64x64 resolution.

[0118] To calculate L CNT The L1 loss can be evaluated between the activations of Conv1,6,11,20,29VGG19 layers and Conv1,6,11,18,25VGGFace layers used for real and fake images. In the case of VGG19, the loss with a certain weight is 1.10. -2 Furthermore, in the case of the VGGFace term, the summation yields 2.10. -3 For both networks, versions trained with Caffe can be used [Reference 17]. In L FM In this case, activation can be performed after the remaining blocks of each discriminator network, and is equivalent to 1.10. 1 The weights. Finally, in L MCH In this case, the weight can be set to 8.10. 1 .

[0119] The minimum number of channels in the convolutional layers can be set to, for example, 64, and both the size of the embedding vector N and the maximum number of channels can be set to, for example, 512. In total, the embedder can have 15 million parameters, and the generator can have 38 million parameters. The ConvNet part of the discriminator can have 20 million parameters. The network can be optimized using Adam [Reference 21]. In the case of the discriminator, the learning rate of both the embedder and generator networks can be set to 5 x 10^6. -5 and 2 x 10 -4 Furthermore, two update steps can be performed for the latter for each of the former [Reference 42].

[0120] This disclosure is not intended to be limited to the foregoing access methods, values, and details, and this is because modifications and alterations to the foregoing access methods, values, and details can be devised by those skilled in the art without any further effort. Therefore, it is assumed that such modifications and alterations are within the scope of the claims.

[0121] According to this disclosure, a method for synthesizing a conversational head sequence of a random individual using a generator network is provided, the generator network being configured to map a head pose and a mannequin descriptor to at least one image of a conversational head sequence of an individual in an electronic device. The method may include: performing few-shot learning of the generator network, the generator network undergoing meta-learning relative to a plurality of M video sequences, the plurality of M video sequences including conversational head images of people different from the random individual; and implementing a fine-tuned generator network, head pose, and mannequin descriptor y. NEW (t) previously used unseen sequences to synthesize the conversation head sequence of an individual.

[0122] Performing few-shot learning of a generator network (which has undergone meta-learning relative to a plurality of M video sequences, the plurality of M video sequences comprising talking head images of people different from random individuals) may include the following: receiving at least one video frame x′(t) from a single sequence of frames of a particular individual from which a talking head sequence will be synthesized; estimating a head pose and a simulated descriptor y′(t) for at least one video frame x′(t); and computing an embedded N-dimensional vector representing individual-specific information based on at least one video frame x′(t) using a meta-learned embedding network. The parameter set and embedded N-dimensional vectors based on the meta-learned generator network The generator network is instantiated; and its parameters are fine-tuned based on the head pose and analog object descriptor y′(t) provided to the generator network to match at least one video frame x′(t). The parameter set of the generator network after meta-learning can be the input of the few-shot learning step, and the parameter set of the fine-tuned generator network can be the output of the few-shot learning step.

[0123] Head pose and analog descriptors y′(t) and y NEW (t) may include, but is not limited to, facial landmarks. The head pose and analog object descriptor y′(t) can be used with at least one video frame x′(t) to compute the embedded N-dimensional vector.

[0124] Meta-learning of the generator network and the embedding network can be performed in K-sample learning stages, where K is a predefined integer, and each stage may include the following: receiving at least one video frame x(t) from one of a plurality of M video sequences, the plurality of M video sequences including conversational head images of people different from random individuals; estimating head pose and analog descriptor y(t) for at least one video frame x(t); and computing an embedded N-dimensional vector representing individual-specific information based on at least one video frame x(t). Based on the parameter set and embedded N-dimensional vector of the current generator network The generator network is instantiated; and the parameter sets of the generator network and the embedder network are updated based on the output of the generator network for the estimated head pose and the analog descriptor y(t) and the matching between the sequence of at least one video frame x(t).

[0125] The generator network and the embedder network can be, for example, convolutional networks. During the instantiation step, the normalization coefficients within the instantiated generator network can be calculated based on the embedded N-dimensional vector computed by the embedder network. The discriminator network can be meta-learned together with the generator network and the embedder network, and the method can further include: using the discriminator network to compute a fidelity score r for the output of the generator network, and updating the parameters of the generator network and the embedder network based on the fidelity score r; and updating the parameters of the discriminator network to increase the fidelity score r for a certain video frame in a plurality of M video sequences and decrease the fidelity score r for the output of the generator network (e.g., a synthesized image).

[0126] The discriminator network can be a projection discriminator network configured to compute a fidelity score r for the generator network's output using an embedded N-dimensional vector w, wherein the embedded N-dimensional vector w is different from an embedded N-dimensional vector trained relative to each of a plurality of M video sequences. It can embed N-dimensional vectors The difference between the embedded N-dimensional vector w and the discriminant is penalized, and a projection discriminator can be used during the fine-tuning step. The embedded N-dimensional vector w of the projection discriminator can be initialized to the embedded N-dimensional vector w at the start of fine-tuning.

[0127] The various embodiments described above can be executed by an electronic device. Reference will be made below. Figure 5 and Figure 6 The electronic device according to this disclosure will be described in more detail.

[0128] Figure 5 This is a block diagram illustrating an example configuration of an example electronic device according to an embodiment of the present disclosure.

[0129] like Figure 5 As shown, the electronic device 100 according to this disclosure may include a memory 110 and a processor (e.g., including processing circuitry) 120.

[0130] The memory 110 may store at least one instruction relating to the electronic device 100. Additionally, the memory 110 may store an operating system (O / S) to drive the electronic device 100. Furthermore, the memory 110 may store various software programs or applications to enable the electronic device 100 to operate according to various embodiments of this disclosure. Furthermore, the memory 110 may include semiconductor memory (such as flash memory) or magnetic storage media (such as a hard disk).

[0131] For example, memory 110 may store various types of software modules for operating electronic device 100 according to various embodiments of the present disclosure, and processor 120 may include various processing circuits and control the operation of electronic device 100 by executing the various types of software modules stored in memory 110. For example, memory 110 may be accessed by processor 120, and data reading / recording / correction / deletion / updating, etc., may be performed by processor 120.

[0132] In this disclosure, the term memory 110 may include memory 110, ROM (not shown) and RAM (not shown) inside processor 120, and / or a memory card (not shown) installed on electronic device 100 (e.g., micro SD card, memory stick).

[0133] For example, in various embodiments according to this disclosure, the memory 110 may store a neural network model according to this disclosure, as well as modules implemented to enable the use of the neural network model to implement various embodiments of this disclosure. Additionally, the memory 110 may store information related to algorithms used to perform first and second learning according to this disclosure. Furthermore, the memory 110 may store multiple video sequences, various images, landmark information, and information regarding parameters of the embedder, generator, and discriminator according to this disclosure.

[0134] In addition to the above, various necessary information within a certain range for achieving the purposes of this disclosure may also be stored in memory 110, and the information stored in memory 110 may be received from a server or external device, or updated when input by a user.

[0135] The processor 120 may include various processing circuits and control the overall operation of the electronic device 100. For example, the processor 120 may be connected to various components of the electronic device 100 including the aforementioned memory, and may control the overall operation of the electronic device 100 by executing at least one instruction stored in the aforementioned memory 110.

[0136] Processor 120 can be implemented in various ways. For example, processor 120 may include various processing circuits, such as, but not limited to, at least one of the following: application-specific integrated circuit (ASIC), embedded processor, microprocessor, hardware control logic, hardware finite state machine (FSM), digital signal processor (DSP), CPU, special-purpose processor, etc. In this disclosure, the term processor 120 may include, for example, but not limited to: central processing unit (CPU), graphics processing unit (GPU), main processing unit (MPU), etc.

[0137] For example, in various embodiments according to this disclosure, processor 120 may perform a first learning based on a plurality of learning video sequences according to a meta-learning method, and cause a neural network model to acquire a video sequence including the head of a conversation of a random user. Processor 120 may perform a second learning based on at least one image of a first user who is a new user not included in the plurality of learning video sequences according to a few-shot learning method, and fine-tune the neural network model to personalize for the first user, and use the neural network model that has performed the first and second learning to acquire a video sequence including the head of a conversation of the first user. (Note: The above references...) Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 Various embodiments according to this disclosure have been described, and therefore overlapping descriptions will not be repeated herein.

[0138] Figure 6 This is a block diagram illustrating an example configuration of an example electronic device according to an embodiment of the present disclosure.

[0139] like Figure 6 As shown, the electronic device 100 according to this disclosure may include not only a memory 110 and a processor 120, but also a communicator (e.g., including communication circuitry) 130, an image sensor (not shown), an output device (e.g., including output circuitry) 140, and an input device (e.g., including input circuitry) 150. However, such components are merely examples, and in implementing this disclosure, new components may be added in addition to such components, or some components may be omitted.

[0140] The communicator 130 may include various communication circuits and perform communication with external devices (e.g., including servers). For example, the processor 120 may receive various data or information from external devices connected via the communicator 130 and transmit various data or information to external devices. The communicator 130 may include various communication circuits included in various communication modules, such as, but not limited to, at least one of the following: a WiFi module, a Bluetooth module, a wireless communication module, an NFC module, etc.

[0141] For example, in various embodiments according to this disclosure, processor 120 may receive, via communicator 130, at least one image including the head of a conversation of a first user from an external device. Processor 120 may receive, via communicator 130, at least some of the following from the external device: information related to algorithms for performing first and second learning according to this disclosure, multiple video sequences according to this disclosure, various images, landmark information, and information regarding parameters of the embedder, the generator, and the discriminator. Processor 120 may control communicator 130 to transmit the first video sequence acquired according to this disclosure to the external device.

[0142] Output device 140 may include various output circuits, and processor 120 may output various functions that electronic device 100 can perform through output device 140. In addition, output device 140 may include, for example, but not limited to, at least one of the following: display, speaker, indicator, etc.

[0143] For example, in various embodiments of the present disclosure, the processor 120 may control the display to display video sequences or images. For example, when a first video sequence including the head of a conversation of a first user is acquired through the foregoing process, the processor 120 may control the display to display the acquired first video sequence.

[0144] The input device 150 may include various input circuits, and the processor 120 may receive user instructions through the input device 150 to control the operation of the electronic device 100. For example, the input device 150 may include various components containing input circuits, such as, but not limited to, a microphone, a camera, a signal receiver, etc. The input device 150 may be implemented as a touchscreen included in a display.

[0145] For example, in various embodiments according to this disclosure, the camera may include an image sensor and convert light entering through a lens into an electronic image signal. The processor 120 may acquire a raw image of an object via the camera. The image sensor may be a charge-coupled device (CCD) sensor or a complementary metal-oxide-semiconductor (CMOS) sensor, but is not limited thereto.

[0146] For example, according to one embodiment of this disclosure, the processor 120 can acquire at least one image including the head of a first user in conversation via a camera. Upon acquiring at least one image including the head of the first user in conversation, the acquired at least one image can be stored in the memory 110. The at least one image stored in the memory 110 can be used for at least one of the first or second learning processes performed by the aforementioned neural network model under the control of the processor 120, and can also be used to acquire a first video sequence.

[0147] The method for controlling the electronic device 100 according to the foregoing embodiments can be implemented as a program and provided to the electronic device 100. For example, a program including the control method of the electronic device 100 can be provided, and the program can be stored in a non-transitory computer-readable medium.

[0148] For example, in a computer-readable recording medium including a program for performing a control method of electronic device 100, the control method of electronic device 100 may include: performing a first learning on a neural network model based on multiple learning video sequences including the heads of conversations of multiple users to obtain a video sequence including the heads of conversations of random users; performing a second learning on a neural network model based on at least one image including the head of conversation of a first user different from the multiple users and first landmark information included in the at least one image to fine-tune the neural network model; and using the neural network model that has performed the first learning and the second learning to obtain a first video sequence including the head of conversation of the first user based on at least one image and pre-stored second landmark information.

[0149] In the foregoing, the electronic device 100 according to the present disclosure and the computer-readable recording medium including a program for executing the control method of the electronic device 100 have been explained schematically. However, this is only to omit overlapping explanations, and various embodiments of the control method of the electronic device 100 can be applied to the electronic device 100 according to the present disclosure and the computer-readable recording medium including a program for executing the control method of the electronic device 100.

[0150] According to the foregoing embodiments of this disclosure, an electronic device can efficiently and effectively train a neural network model capable of generating video sequences including conversational heads, and thereby, using a small number of user images not included in the learning data, acquire video sequences including highly realistic conversational heads with the user. Furthermore, according to this disclosure, the uncanny valley effect is substantially eliminated from the generated video sequences, and high-quality video sequences can be provided, including highly realistic conversational heads optimized for specific users.

[0151] The electronic device 100 according to this disclosure may be an electronic device, such as, but not limited to, a smartphone, tablet computer, PC, laptop computer, or, for example, AR glasses, VR glasses, smartwatches, etc. However, the electronic device 100 is not limited to these, and any electronic device capable of performing processes including first learning, second learning, and acquisition of video sequences according to this disclosure may be included in the electronic device 100 according to this disclosure.

[0152] Various methods, apparatuses, and systems for providing highly realistic avatars can be implemented based on the foregoing disclosure. Different methods, apparatuses, and systems using models / networks trained to provide highly realistic conversational head models and / or few-shot learning of highly realistic avatars can be generated based on the foregoing disclosure. Various exemplary embodiments of this disclosure can be implemented as a non-transitory machine-readable medium comprising computer-executable instructions that, when executed, cause an electronic device to perform the disclosed method of using an adversarial network to synthesize a random individual conversational head model when executed by a processing unit of a device.

[0153] This disclosure can be implemented as a system for synthesizing conversational head models of random individuals using adversarial networks. In such a system, the operation of the method can be implemented as different functional units, circuits, and / or processors 120. However, any suitable functional distribution can be used in the different functional units, circuits, and / or processors 120 without departing from the explained embodiments.

[0154] The various exemplary embodiments of this disclosure can be implemented in any suitable form, including hardware, software, firmware, or any combination thereof. Embodiments can be implemented, at least in part, as computer software selectively executed on at least one data processor 120 and / or digital signal processor 120. Elements and components of any embodiment can be implemented physically, functionally, and logically by any suitable method. In practice, functionality can be implemented as a single unit, multiple units, or part of another general unit.

[0155] The foregoing description of the embodiments of this disclosure is merely illustrative, and various modifications to the configuration and implementation should be considered within the scope of this disclosure, including the appended claims. For example, embodiments of this disclosure are generally explained relative to exemplary methods, but such explanations are provided as examples. Although this disclosure is described using language specific to structural features or method operations, it should be understood that the appended claims are not necessarily limited to the aforementioned specific features or operations. Rather, the aforementioned specific features and operations are disclosed as exemplary. This disclosure is not limited to the order of steps of the proposed method, and those skilled in the art can modify the order without much effort. Furthermore, some or all of the operations of the method may be performed sequentially or simultaneously.

[0156] Each of the components (e.g., modules or programs) in the foregoing various embodiments of this disclosure may include a single object or multiple objects. In the foregoing corresponding sub-components, some sub-components may be omitted, or other sub-components may be further included in the various embodiments. Typically or additionally, some components (e.g., modules or programs) may be integrated into an object and perform the functions performed by each component in the same or similar manner prior to integration.

[0157] Operations performed by modules, programs, or other components according to various embodiments may be performed sequentially, in parallel, repeatedly, or heuristically. At least some of the operations may be performed in a different order or omitted, or other operations may be added.

[0158] As used in this disclosure, the term "component" or "module" includes a unit comprising hardware, software, or firmware, and it may be used interchangeably with terms such as logic, logic block, component, or circuit. Additionally, a "component" or "module" can be a component comprising an integrated body or a minimum unit that performs one or more functions or a portion thereof. For example, a module may include an application-specific integrated circuit (ASIC).

[0159] Various embodiments of this disclosure can be implemented as software, which includes instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). A machine can, for example, refer to a means that invokes and can operate according to the instructions stored in the storage medium, and the means may include an electronic device (e.g., electronic device 100) according to embodiments described in this disclosure.

[0160] When instructions are executed by processor 120, processor 120 can perform the function corresponding to the instructions independently or using other components under the control of the processor. Instructions may include code generated by a compiler or code executable by an interpreter.

[0161] A machine-readable storage medium may be provided in the form of a non-transitory storage medium. A 'non-transitory storage medium' is a tangible device and may not include signals (e.g., electronic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where data is temporarily stored. For example, a 'non-transitory storage medium' may include a buffer for temporarily storing data.

[0162] According to one embodiment of this disclosure, the method may be provided when it is included in a computer program product comprising methods according to the various embodiments described herein. The computer program product is a product, and it can be traded between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., an optical disc read-only memory (CD-ROM)), or directly distributed between two user devices (e.g., a smartphone) and through an app store (e.g., the Play Store). TM Online distribution (e.g., download or upload). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be stored at least temporarily in a storage medium that can be read by a machine such as the memory 110 of the manufacturer's server, the app store's server, and the relay server, or may be temporarily generated.

[0163] Functions related to the neural network model according to this disclosure and functions related to artificial intelligence can be executed by memory 110 and processor 120.

[0164] Processor 120 may include one or more processors 120. The one or more processors 120 may be general-purpose processors, such as, but not limited to: CPU, AP, etc.; graphics-specific processors, such as GPU, VPU, etc.; artificial intelligence-specific processors, such as NPU, etc.

[0165] One or more processors 120 can perform control to process input data according to predefined operating rules or an artificial intelligence model stored in memory 110. The predefined operating rules or artificial intelligence models are characterized in that they are learned.

[0166] Features achieved through learning can refer, for example, to predefined operational rules or artificial intelligence models that achieve desired characteristics by applying a learning algorithm to multiple learning datasets. This learning can be performed independently by the device, where artificial intelligence according to this disclosure is executed, or via a separate server / system.

[0167] Artificial intelligence models may include multiple neural network layers. Each layer may include, for example, multiple weight values, and layer operations are performed by combining the results of the previous layer with the multiple weight values. As non-limiting examples of neural networks, there are convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), deep Q-networks, etc., and the neural networks in this disclosure are not limited to the foregoing examples, except where explicitly stated.

[0168] A learning algorithm can be, for example, a method of training a subject-specific machine (e.g., a robot) using multiple learning data and enabling the subject-specific machine to make decisions or predictions independently. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and the learning algorithms in this disclosure are not limited to the foregoing examples, except as expressly indicated.

[0169] While various exemplary embodiments of the present disclosure have been shown and described, the present disclosure is not limited to the foregoing embodiments, and it will be understood that various modifications may be made by those skilled in the art to which this disclosure pertains without departing from the spirit and scope of the present disclosure, including the appended claims.

[0170] The list of references mentioned above is as follows. It should be noted that these references are cited solely for the purpose of interpreting this disclosure and are not intended to be construed as limiting or restricting it. Furthermore, these references are incorporated herein by reference in their entirety.

[0171] [1] O. Alexander, M. Rogers, W. Lambeth, J.-Y. Chiang, W.-C. Ma, C.-C. Wang, and P. Debevec. The Digital Emily project: Achieving a photorealistic digital actor. IEEE Computer Graphics and Applications, 30(4): 20-31, 2010.

[0172] [2]KSAndrew Brock,Jeff Donahue.Large scale gan training for high fidelity natural image synthesis.arXiv:1809.11096,2018.

[0173] [3]A.Antoniou,A.J.Storkey,and H.Edwards.Augmenting image classifiersusing data augmentation generative adversarial networks.In Artificial NeuralNetworks and Machine Learning-ICANN,pages 594-603,2018.

[0174] [4]S.Arik,J.Chen,K.Peng,W.Ping,and Y.Zhou.Neural voice cloning with afew samples.In Proc.NIPS,pages 10040-10050,2018.

[0175] [5]H.Averbuch-Elor,D.Cohen-Or,J.Kopf,and M.F.Cohen.Bringing portraitsto life.ACM Transactions on Graphics(TOG),36(6):196,2017.

[0176] [6]V.Blanz,T.Vetter,et al.Amorphable model for the synthesis of 3dfaces.In Proc.SIGGRAPH,volume 99,pages 187-194,1999.

[0177] [7]A.Bulat and G.Tzimiropoulos.How far are we fromsolving the 2d&3dface alignment problem?(and a dataset of 230,000 3d facial landmarks).In IEEEInternational Conference on Computer Vision,ICCV 2017,Venice,Italy,October22-29,2017,pages 1021-1030,2017.

[0178] [8]JSChung,A.Nagrani,and A.Zisserman.Voxceleb2:Deep speaker recognition.In INTERSPEECH,2018.

[0179] [9]J.Deng,J.Guo,X.Niannan,and S.Zafeiriou.Arcface:Additive angular margin loss for deep face recognition.In CVPR,2019.

[0180]

[10] C.Finn,P.Abbeel,and S.Levine.Model-agnostic metal learning for fast adaptation of deep networks.In Proc.ICML,pages 1126-1135,2017.

[0181]

[11] Y.Ganin,D.Kononenko,D.Sungatullina,and V.Lempitsky.Deepwarp:Photorealistic image resynthesis for gaze manipulation.In European Conferenceon Computer Vision,pages 311-326.Springer,2016.

[0182]

[12] I.Goodfellow,J.Pouget-Abadie,M.Mirza,B.Xu,D.Warde-Farley,S.Ozair,A.Courville,and Y.Bengio.Generative adversarial nets.In Advances inneuralinformation processing systems,pages 2672-2680,2014.

[0183]

[13] M.Heusel,H.Ramsauer,T.Unterthiner,B.Nessler,and S.Hochreiter.Ganstrained by a two time-scale update rule converge to a local nashequilibrium.In I.Guyon,U.V.Luxburg,S.Bengio,H.Wallach,R.Fergus,S.Vishwanathan,and R.Garnett,editors,Advances inNeural Information ProcessingSystems 30,pages 6626-6637.Curran Associates,Inc.,2017.6

[0184]

[14] X.Huang and S.Belongie.Arbitrary style transfer inrealtime withadaptive instance normalization.In Proc.ICCV,2017.

[0185]

[15] S.Ioffe and C.Szegedy.Batch normalization:Accelerating deepnetwork training by reducing internal covariate shift.In Proceedings of the32Nd International Conference on International Conference on MachineLearning-Volume 37,ICML'15,pages 448-456.JMLR.org,2015.

[0186]

[16] P.Isola,J.Zhu,T.Zhou,and A.A.Efros.Image-to-image translationwith conditional adversarial networks.In Proc.CVPR,pages 5967-5976,2017.

[0187]

[17] Y.Jia,E.Shelhamer,J.Donahue,S.Karayev,J.Long,R.Girshick,S.Guadarrama,and T.Darrell.Caffe:Convolutional architecture for fast featureembedding.arXiv preprint arXiv:1408.5093,2014.

[0188]

[18] Y.Jia,Y.Zhang,R.Weiss,Q.Wang,J.Shen,F.Ren,P.Nguyen,R.Pang,ILMoreno,Y.Wu,et al.Transfer learning from speaker verification tomultispeaker text-tospeech synthesis.In Proc.NIPS,pages 4485-4495,2018.

[0189]

[19] J.Johnson,A.Alahi,and L.Fei-Fei.Perceptual losses for real-timestyle transfer and super-resolution.In Proc.ECCV,pages 694-711,2016.

[0190]

[20] H. Kim, P. Garrido, A. Tewari, W. Xu, J. Thies, M. Nieβner, P. Perez, C. Richardt, M. Zollh′ofer, and C. Theobalt.

[0191]

[21] DPKingma and J.Ba.Adam:A method for stochastic optimization.CoRR,abs / 1412.6980,2014.

[0192]

[22] S.Lombardi,J.Saragih,T.Simon,and Y.Sheikh.Deep appearance modelsfor face rendering.ACM Transactions on Graphics(TOG),37(4):68,2018.

[0193]

[23] S.O.Mehdi Mirza.Conditional generative adversarial nets.arXiv:1411.1784.

[0194]

[24] M.Mori.The uncanny valley.Energy,7(4):33-35,1970.

[0195]

[25] K.Nagano,J.Seo,J.Xing,L.Wei,Z.Li,S.Saito,A.Agarwal,J.Fursund,H.Li,R.Roberts,et al.paGAN:real-time avatars using dynamic textures.InSIGGRAPH Asia 2018 Technical Papers,page 258.ACM,2018.

[0196]

[26] A.Nagrani,J.S.Chung,and A.Zisserman.Voxceleb:a large-scalespeaker identification dataset.In INTERSPEECH,2017.

[0197]

[27] O.M.Parkhi,A.Vedaldi,and A.Zisserman.Deep face recognition.InProc.BMVC,2015.

[0198]

[28] S.M.Seitz and C.R.Dyer.View morphing.In Proceedings of the 23rdannual conference on Computer graphics and interactive techniques,pages 21-30.ACM,1996.

[0199]

[29] Z.Shu,M.Sahasrabudhe,R.Alp Guler,D.Samaras,N.Paragios,andI.Kokkinos.Deforming autoencoders:Unsupervised disentangling of shape andappearance.In The European Conference on Computer Vision(ECCV),September2018.

[0200]

[30] K.Simonyan and A.Zisserman.Very deep convolutional networks forlarge-scale image recognition.In Proc.ICLR,2015.

[0201]

[31] S.Suwajanakorn,S.M.Seitz,and I.KemelmacherShlizerman.SynthesizingObama:learning lip sync from audio.ACM Transactions on Graphics(TOG),36(4):95,2017.

[0202]

[32] M.K.Takeru Miyato.cgans with projection discriminator.arXiv:1802.05637,2018.

[0203]

[33] M.K.Y.Y.Takeru Miyato,Toshiki Kataoka.Spectral normalization forgenerative adversarial networks.arXiv:1802.05957,2018.

[0204]

[34] T.A.Tero Karras,Samuli Laine.A style-based generator architecturefor generative adversarial networks.arXiv:1812.04948.

[0205]

[35] J.Thies,M.Zollhofer,M.Stamminger,C.Theobalt,and M.Nieβner.Face2face:Real-time face capture and reenactment of RGB videos.InProceedings of the IEEE Conference on Computer Vision and PatternRecognition,pages 2387-2395,2016.

[0206]

[36] D.Ulyanov,A.Vedaldi,and V.S.Lempitsky.Instance normalization:Themissing ingredient for fast stylization.CoRR,abs / 1607.08022,2016.

[0207]

[37] T.-C.Wang,M.-Y.Liu,J.-Y.Zhu,G.Liu,A.Tao,J.Kautz,andB.Catanzaro.Video-to-video synthesis.arXiv preprint arXiv:1808.06601,2018.

[0208]

[38] T.-C.Wang,M.-Y.Liu,J.-Y.Zhu,A.Tao,J.Kautz,and B.Catanzaro.High-resolution image synthesis and semantic manipulation with conditional gans.InProceedings of the IEEE Conference on Computer Vision and PatternRecognition,2018.

[0209]

[39] Z.Wang,A.C.Bovik,H.R.Sheikh,and E.P.Simoncelli.Image qualityassessment:From error visibility to structural similarity.Trans.Img.Proc.,13(4):600-612,Apr.2004.

[0210]

[40] O.Wiles,A.Sophia Koepke,and A.Zisserman.X2face:A network forcontrolling face generation using images,audio,and pose codes.In The EuropeanConference on Computer Vision(ECCV),September 2018.

[0211]

[41] C.Yin,J.Tang,Z.Xu,and Y.Wang.Adversarial metalearning.CoRR,abs / 1806.03316,2018.2

[0212]

[42] H.Zhang,I.J.Goodfellow,D.N.Metaxas,and A.Odena.Self-attentiongenerative adversarial networks.arXiv:1805.08318,2018.

[0213]

[43] R.Zhang,T.Che,Z.Ghahramani,Y.Bengio,and Y.Song.Metagan:Anadversarial approach to few-shot learning.In NeurIPS,pages 2371-2380,2018。

Claims

1.A method of controlling an electronic device, the method comprising: performing first learning on a neural network model based on a plurality of learning video sequences including talking heads of a plurality of users to obtain a video sequence including a talking head of a random user, wherein the first learning is meta-learning for second learning; performing the second learning to fine-tune the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image, wherein the first landmark information is about major features of a face of the first user included in the at least one image; and generating a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model on which the first learning and the second learning are performed, wherein the second landmark information is obtained from a plurality of images included in the plurality of learning video sequences, the second landmark information being about major features of faces of a plurality of users included in the plurality of images, wherein the performing second learning comprises: obtaining at least one image including the talking head of the first user; obtaining the first landmark information based on the at least one image; obtaining a first embedding vector including information related to an identity of the first user by inputting the at least one image and the first landmark information into an embedder of the neural network model on which the first learning is performed; and fine-tuning a parameter set of a generator of the neural network model on which the first learning is performed to match the at least one image based on the first embedding vector. 2.The method of claim 1, wherein the generating a first video sequence comprises: obtaining the first video sequence by inputting the second landmark information and the first embedding vector into the generator. 3.The method of claim 1, wherein the first landmark information and the second landmark information include information about head poses and information about analog object descriptors. 4.The method of claim 1, wherein the embedder and the generator include convolutional networks, and based on the generator being instantiated, obtaining a normalization coefficient inside the instantiated generator based on a first embedding vector obtained by the embedder. 5.The method of claim 1, wherein the performing first learning comprises: obtaining at least one learning image from a learning video sequence including a talking head of a second user among the plurality of learning video sequences; obtaining third landmark information of the second user based on the at least one learning image; obtaining a second embedding vector including information related to an identity of the second user by inputting the at least one learning image and the third landmark information into the embedder; instantiating the generator based on a parameter set of the generator and the second embedding vector; obtaining a second video sequence including the talking head of the second user by inputting the second landmark information and the second embedding vector into the generator; and updating a parameter set of the neural network model based on a degree of similarity between the second video sequence and the learning video sequence. 6.The method of claim 5, wherein the performing the first learning further comprises: obtaining a realism score for the second video sequence by a discriminator of the neural network model; updating a parameter set of the generator and a parameter set of the embedder based on the realism score; and updating a parameter set of the discriminator. 7.The method of claim 6, wherein the discriminator comprises a projection discriminator that obtains the realism score based on a third embedding vector that is different from the first embedding vector and the second embedding vector. 8.The method of claim 7, wherein based on the first learning being performed, a difference between the second embedding vector and the third embedding vector is penalized, and the third embedding vector is initialized based on the first embedding vector at the beginning of the second learning. 9.The method of claim 1, wherein the at least one image comprises 1 to 32 images. 10.An electronic device comprising: a memory storing at least one instruction; and a processor configured to execute the at least one instruction, wherein by executing the at least one instruction, the processor is configured to: perform a first learning on a neural network model based on a plurality of learning video sequences comprising talking heads of a plurality of users to obtain a video sequence comprising a talking head of a random user, wherein the first learning is meta-learning for a second learning, perform the second learning to fine-tune the neural network model based on at least one image comprising a talking head of a first user that is different from the plurality of users and first landmark information included in the at least one image, wherein the first landmark information is about major features of a face of the first user included in the at least one image, and generate a first video sequence comprising the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model on which the first learning and the second learning are performed, wherein the second landmark information is about major features of faces of a plurality of users included in a plurality of images included in the plurality of learning video sequences, wherein the processor is configured to: obtain the at least one image comprising the talking head of the first user, obtain the first landmark information based on the at least one image, obtain a first embedding vector comprising information related to an identity of the first user by inputting the at least one image and the first landmark information to an embedder of the neural network model on which the first learning is performed, and fine-tune a parameter set of a generator of the neural network model on which the first learning is performed to match the at least one image based on the first embedding vector. 11.The electronic device of claim 10, wherein the processor is configured to: The first video sequence is acquired by inputting the second landmark information and the first embedding vector into the generator. 12.The electronic device of claim 10, wherein the first landmark information and the second landmark information include information about a head pose and information about a simulated object descriptor. 13.A non-transitory computer-readable recording medium having recorded thereon a program, the program, when executed by a processor of an electronic device, causing the electronic device to perform the method of claim 1.