Method for portrait animation, computing device and computer readable storage medium

By receiving scene videos and target images on mobile devices, and using deep neural networks for segmentation and background prediction, 2D deformation is applied to imitate facial expressions and head directions, the problem that realistic real-time portrait animation in the prior art is difficult to achieve on mobile devices, and a fast and realistic real-time portrait animation effect is achieved.

CN119963704APending Publication Date: 2025-05-09SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126743.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-01-18
Filing Date
2020-01-18
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to implement realistic real-time portrait animation on standard mobile devices. Although deep learning methods can generate realistic results, they take a long time and are not suitable for real-time applications.

Method used

Receive scene video and target images through computing devices, segmentation and background prediction based on deep neural networks, apply 2D deformation to mimic facial expressions and head orientation, and apply deformation to target images to generate output video.

Benefits of technology

It realizes the rapid generation of realistic real-time portrait animations on mobile devices, reducing computing time, suitable for real-time applications, while maintaining high realistic effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963704A_ABST
    Figure CN119963704A_ABST
Patent Text Reader

Abstract

The invention provides a portrait animation method, a computing device and a computer readable storage medium. An exemplary method includes receiving a scene video having at least one input frame. The input frame includes a first face. The method further includes receiving a target image having a second face. The method further includes determining a two-dimensional (2D) deformation based on the at least one input frame and the target image, where the 2D deformation, when applied to a second face, modifies the second face to mimic at least a facial expression and a head direction of the first face. The method further includes applying, by the computing device, the 2D deformation to the target image to obtain at least one output frame of the output video.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese national phase application of the PCT application with an application date of January 18, 2020, an international application number of PCT / US2020 / 014220, and an invention name of “System and method for realistic real-time portrait animation”. The entry date of this Chinese national phase application into the Chinese national phase is June 28, 2021, and the application number is 202080007510.5. Technical Field

[0002] The present disclosure relates generally to digital image processing and more particularly to methods and systems for realistic real-time portrait animation. Background Art

[0003] Portrait animation can be used in many applications, such as entertainment programs, computer games, video conversations, virtual reality, augmented reality, etc.

[0004] Some current techniques for portrait animation utilize deformable facial models to re-render faces with different facial expressions. Although generating faces using deformable facial models can be fast, the resulting faces are generally not realistic. Some other current techniques for portrait animation can be based on using deep learning methods to re-render faces with different facial expressions.

[0005] Deep learning methods can allow realistic results to be obtained. However, they are time-consuming and not suitable for performing real-time portrait animation on standard mobile devices. Summary of the invention

[0006] This section is provided to introduce some concepts in a simplified form, which are further described in the Detailed Description section below. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0007] According to one embodiment of the present disclosure, a method for realistic real-time portrait animation is provided. The method may include receiving a scene video through a computing device. The scene video may include at least one input frame. At least one input frame may include a first face. The method may further include receiving a target image through a computing device. The target image may include a second face. The method may further include: determining a two-dimensional (2D) deformation through a computing device and based on at least one input frame and a target image, wherein the 2D deformation, when applied to the second face, modifies the second face to mimic at least a facial expression and a head direction of the first face. The method may further include applying the 2D deformation to the target image through a computing device to obtain at least one output frame of an output video.

[0008] In some embodiments, the method may further include: before applying the 2D deformation, performing segmentation of the target image using a deep neural network (DNN) by a computing device to obtain an image of the second face and a background. The 2D deformation may be applied to the image of the second face to obtain a deformed face while keeping the background unchanged.

[0009] In some embodiments, the method may further include: inserting, by the computing device, the deformed face into the background after applying the 2D deformation. The method may further include: predicting, by the computing device and using the DNN, a portion of the background in a gap between the deformed face and the background. The method may further allow, by the computing device, to fill the gap with the predicted portion.

[0010] In some embodiments, determining the 2D deformation may include determining, by the computing device, a first control point on the first face and a second control point on the second face. The method may further include defining, by the computing device, a 2D deformation or affine transformation for aligning the first control point with the second control point.

[0011] In some embodiments, determining the 2D deformation may include establishing, by a computing device, a triangulation of the second control point. Determining the 2D deformation may further include determining, by a computing device, a displacement of the first control point in at least one input frame. Determining the 2D deformation may further include projecting, by the computing device and using an affine transformation, the displacement onto the target image to obtain a desired displacement of the second control point. Determining the 2D deformation may further include determining, by the computing device and based on the desired displacement, a warp field to be used as the 2D deformation.

[0012] In some embodiments, the warp field comprises a set of piecewise linear transformations defined by changes in triangles in the triangulation of the second control points.

[0013] In some embodiments, the method may further include generating, by a computing device, a mouth region and an eye region. The method may further include inserting, by a computing device, the mouth region and the eye region into at least one output frame.

[0014] In some embodiments, generating one of the mouth region and the eye region includes transferring, by the computing device, the mouth region and the eye region from the first face.

[0015] In some embodiments, generating the mouth region and the eye region may include fitting a three-dimensional (3D) facial model to a first control point by a computing device to obtain a first set of parameters. The first set of parameters may include at least a first facial expression. Generating one of the mouth region and the eye region may further include fitting the 3D facial model to a second control point by a computing device to obtain a second set of parameters. The second set of parameters may include at least a second facial expression. The first facial expression from the first set of parameters may be transferred to the second set of parameters. Generating the mouth region and the eye region may further include: synthesizing one of the mouth region and the eye region by a computing device and using the 3D facial model.

[0016] According to another embodiment, a system for realistic real-time portrait animation is provided. The system may include at least one processor and a memory storing processor executable code, wherein the at least one processor may be configured to implement the operations of the method for realistic real-time portrait animation when executing the processor executable code.

[0017] According to another aspect of the present disclosure, a non-transitory processor-readable medium is provided, which stores processor-readable instructions. When the processor-readable instructions are executed by a processor, the processor implements the method for realistic real-time portrait animation.

[0018] Other purposes, advantages and novel features of the embodiments will be partially described in the following description, and some of the contents will become obvious to those skilled in the art by reading the following description and the accompanying drawings, or can be learned through the production or operation of the examples. The purposes and advantages of the concept can be realized and obtained by the methods, means and combinations particularly pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Embodiments are illustrated by way of example, and not limitation, in the figures of the accompanying drawings in which like references indicate similar elements.

[0020] Figure 1 is a block diagram illustrating an exemplary environment in which a method for portrait animation may be implemented.

[0021] Figure 2 is a block diagram illustrating an exemplary embodiment of a computing device for implementing a method for portrait animation.

[0022] Figure 3 is a schematic diagram illustrating an exemplary process of portrait animation according to an exemplary embodiment.

[0023] Figure 4 A block diagram of a system for portrait animation according to an exemplary embodiment is shown.

[0024] Figure 5 A process flow diagram of a method for portrait animation according to some exemplary embodiments is shown.

[0025] Figure 6 A process flow diagram of a method for portrait animation according to some exemplary embodiments is shown.

[0026] Figure 7 An exemplary computer system that can be used to implement the method for portrait animation is shown.

[0027] Figure 8 is a block diagram of an exemplary deep neural network (DNN) for background prediction.

[0028] Fig. 9 is a block diagram of an exemplary compressed convolution block in a DNN.

[0029] Fig.10 is a block diagram of an exemplary decompressed convolution block in a DNN.

[0030] Fig.11 is a block diagram of an exemplary attention block in a DNN.

[0031] Fig.12 It is a block diagram of the DNN learning scheme.

[0032] Fig.13 is a block diagram of an exemplary discriminator network. DETAILED DESCRIPTION

[0033] The following detailed description of the embodiments includes reference to the accompanying drawings, which form a part of the detailed description. The methods described in this section are not prior art to the claims, and are not considered prior art by being included in this section. The accompanying drawings show diagrams according to exemplary embodiments. These exemplary embodiments are also referred to as "examples" in this article, which are described in sufficient detail to enable those skilled in the art to practice this subject matter. Without departing from the scope of the claimed protection, the embodiments can be combined, other embodiments can be utilized, or structural, logical and operational changes can be made. Therefore, the following specific embodiments should not be understood as restrictive, and the scope is limited by the attached claims and their equivalents.

[0034] Various techniques may be used to implement the present disclosure. For example, the methods described herein may be implemented by software running on a computer system and / or by hardware utilizing a combination of microprocessors or other specially designed application specific integrated circuits (ASICs), programmable logic devices, or any combination thereof. Specifically, the methods described herein may be implemented by a series of computer executable instructions residing on a non-transitory storage medium such as a disk drive or a computer-readable medium. It should be noted that the methods disclosed herein may be implemented by computing devices such as mobile devices, personal computers, servers, network nodes, and the like.

[0035] For the purposes of this patent document, the terms "or" and "and" shall mean "and / or" unless otherwise stated or the context in which they are used clearly indicates otherwise. The terms "a" and "an" shall mean "one or more" unless otherwise stated or where the use of "one or more" is clearly inappropriate. The terms "comprise," "comprising," "include," and "including" are interchangeable and are not intended to be limiting. For example, the term "comprising" shall be interpreted to mean "including, but not limited to."

[0036] The present disclosure relates to methods and systems for portrait animation. The present disclosure can be designed to work on mobile devices such as smart phones, tablet computers or mobile phones in real time and without being connected to the Internet or requiring the use of server-side computing resources, but embodiments can also be extended to methods involving web services or cloud-based resources.

[0037] Some embodiments of the present disclosure may allow animation of a target image with a target face. The target face may be manipulated in real time by the facial expressions of the source face. Some embodiments may be used to significantly reduce the computation time for realistic portrait animation. Embodiments of the present disclosure require only a single target image to achieve realistic results, whereas existing facial animation techniques typically use a video or a series of images of the target face.

[0038] Some embodiments of the present disclosure may allow the use of a 3D model of a source video to generate a 2D deformation field caused by changes in the 3D face, and apply the 2D deformation directly to the target image. Embodiments of the present disclosure may allow methods for realistic real-time portrait animation to be implemented on mobile devices and perform animation in real time. In contrast, other methods for editing 3D facial attributes require precise segmentation and texture mapping, and are therefore very time-consuming.

[0039] Embodiments of the present disclosure may allow users to create scenes so that users only need to indicate the expressions, actions, etc. that users want to see on the target face. For example, expressions and actions may be selected from the following list: frowning, smiling, looking down, etc.

[0040] According to one embodiment of the present disclosure, an exemplary method for portrait animation may include receiving a scene video through a computing device. The scene video may include at least one input frame. At least one input frame may include a first face. The method may further include receiving a target image through a computing device. The target image may include a second face. The method may further include: determining a two-dimensional (2D) deformation through a computing device and based on at least one input frame and a target image, wherein the 2D deformation, when applied to the second face, modifies the second face to mimic at least a facial expression and a head direction of the first face. The method may include applying the 2D deformation to the target image through a computing device to obtain at least one output frame of an output video.

[0041] With reference now to the accompanying drawings, exemplary embodiments are described. The accompanying drawings are schematic diagrams of idealized exemplary embodiments. Therefore, it should be apparent to those skilled in the art that the exemplary embodiments discussed herein should not be construed as being limited to the specific illustrations presented herein, but rather that these exemplary embodiments may include deviations and be different from the illustrations presented herein.

[0042] Figure 1 An exemplary environment 100 is shown in which a method for portrait animation can be practiced. The environment 100 may include a computing device 110 and a user 130. The computing device 110 may include a camera 115 and a graphics display system 120. The computing device 110 may refer to a mobile device such as a mobile phone, a smart phone, or a tablet computer. However, in other embodiments, the computing device 110 may refer to a personal computer, a laptop computer, a netbook, a set-top box, a television device, a multimedia device, a personal digital assistant, a game console, an entertainment system, an infotainment system, a vehicle computer, or any other computing device.

[0043] In some embodiments, computing device 110 may be configured to capture scene video via, for example, camera 115. The scene video may include at least the face of user 130 (also referred to as the source face). In some other embodiments, the scene video may be stored in a memory of computing device 110 or in a cloud-based computing resource that is communicatively connected to computing device 110. The scene video may include video of a person, such as user 130 or other person, who may speak, shake his head, and express various emotions.

[0044] In some embodiments of the present disclosure, computing device 110 may be configured to display target image 125. Target image 125 may include at least target face 140 and background 145. Target face 140 may belong to a person other than user 130 or other person depicted in the scene video. In some embodiments, target image 125 may be stored in a memory of computing device 110 or in a cloud-based computing resource that is communicatively connected to computing device 110.

[0045] In other embodiments, different scene videos and target images may be pre-recorded and stored in the memory of computing device 110 or in a cloud-based computing resource. User 130 may select a target image to be animated and one of the scene videos to be used to animate the target image.

[0046] According to various embodiments of the present disclosure, computing device 110 may be configured to analyze a scene video to extract facial expressions and actions of a person depicted in the scene video. Computing device 110 may be further configured to transfer the facial expressions and actions of the person to a target face in target image 125 so that target face 140 repeats the facial expressions and actions of the person in the scene video in real time and in a photo-realistic manner. In other embodiments, computing device 110 may be further configured to modify target image 125 so that target face 140 repeats the speech of the person depicted in the scene video.

[0047] exist Figure 2 In the example shown, the computing device 110 may include hardware components and software components. Specifically, the computing device 110 may include a camera 115 or any other image capture device or scanner to collect digital images. The computing device 110 may further include a processor module 210 and a storage module 215 for storing software components and processor-readable (machine-readable) instructions or codes that, when executed by the processor module 210, cause the computing device 200 to perform at least some steps of the method for portrait animation as described herein.

[0048] The computing device 110 may further include a portrait animation system 220 , which in turn may include hardware components (eg, separate processing modules and memory), software components, or a combination thereof.

[0049] like Figure 3As shown, the portrait animation system 220 can be configured to receive a target image 125 and a scene video 310 as input. The target image 125 can include a target face 140 and a background 145. The scene video 310 can depict at least the head and face of a person 320 who can speak, move his head, and express emotions. The portrait animation system 220 can be configured to analyze the frames of the scene video 310 to determine the facial expressions (emotions) and head movements of the person 320. The portrait animation system can be further configured to change the target image 125 by transferring the facial expressions and head movements of the person 320 to the target face 140, and thereby obtain the frames of the output video 330. The determination of the facial expressions and head movements of the person 320 and the transfer of the facial expressions and head movements to the target face 125 can be repeated for each frame of the scene video 310. The output video 330 can include the same number of frames as the scene video 310. Thus, the output video 330 can represent the animation of the target image 125. In some embodiments, animation can be achieved by performing 2D deformation of the target image 125, wherein the 2D deformation imitates facial expressions and head movements. In some embodiments, hidden areas and fine-scale details can be generated after the 2D deformation to achieve a photo-realistic result. The hidden area can include the mouth area of ​​the target face 140.

[0050] Figure 4 4 is a block diagram of a portrait animation system 220 according to an exemplary embodiment. The portrait animation system 220 may include a 3D face model 405, a sparse correspondence module 410, a scene video preprocessing module 415, a target image preprocessing module 420, an image segmentation and background prediction module 425, and an image animation and refinement module 430. Modules 405 to 430 may be implemented as software components for use with a hardware device such as a computing device 110, a server, or the like.

[0051] In some embodiments of the present disclosure, the 3D facial model 405 may be pre-generated based on a predefined number of images of individuals of different ages, genders, and ethnic backgrounds. For each individual, the images may include an image of the individual with a neutral facial expression and one or more images of the individual with different facial expressions. Facial expressions may include open mouth, smiling, angry, surprised, etc.

[0052] The 3D face model 405 may include a template mesh having a predetermined number of vertices. The template mesh may be represented as a 3D triangulation that defines the shape of the head. Each individual may be associated with an individual-specific blend shape. The individual-specific blend shapes may be adjusted according to the template network. The individual-specific blend shapes may correspond to specific coordinates of vertices in the template mesh. Thus, different images of an individual may correspond to a template mesh having the same structure; however, the coordinates of the vertices in the template mesh may be different for different images.

[0053] In some embodiments of the present disclosure, the 3D facial model 405 may include a bilinear facial model that depends on two parameters: facial identification and facial expression. The bilinear facial model may be established based on a blend shape corresponding to an image of the individual. Thus, the 3D facial model includes a template mesh having a predetermined structure, wherein the coordinates of the vertices depend on the facial identification and the facial expression. The facial identification may represent the geometric shape of the head.

[0054] In some embodiments, the sparse correspondence module 410 can be configured to determine sparse correspondences between frames of the scene video 310 and frames of the target image 125. The sparse correspondence module 410 can be configured to obtain a set of control points (facial landmarks) that can be robustly tracked through the scene video. The facial landmarks and other control points can be tracked using state-of-the-art tracking methods (e.g., optical flow). The sparse correspondence module 410 can be configured to determine an affine transformation that approximately aligns the facial landmarks in the first frame of the scene video 310 and in the target image 125. The affine transformation can be further used to predict the positions of other control points in the target image 125. The sparse correspondence module 410 can be further configured to establish a triangulation of the control points.

[0055] In some embodiments, the scene video preprocessing module 415 can be configured to detect 2D facial landmarks in each frame of the scene video 310. The scene video preprocessing module 415 can be configured to fit the 3D facial model 405 to the facial landmarks to find the parameters of the 3D facial model 405 for the person depicted in the scene video 310. The scene video preprocessing module 415 can be configured to determine the position of the 2D facial landmarks on the template grid of the 3D facial model. It can be assumed that the facial identification is the same for all frames of the scene video 310. The module 415 can be further configured to approximate the resulting changes in the 3D facial parameters for each frame of the scene video 310. The scene video preprocessing module 415 can be configured to receive manual annotations and add the annotations to the parameters of the frame. In some embodiments, third-party animation and modeling applications (such as Maya) can be used. TM ) is annotated. Module 415 can be further configured to select control points and track the positions of the control points in each frame of the scene video 310. In some embodiments, module 415 can be configured to perform segmentation of the inside of the mouth in each frame of the scene video 310.

[0056] In some embodiments, the target image preprocessing module 420 may be configured to detect 2D facial landmarks and visible parts of the head in the target image 125, and fit the 3D facial model to the 2D facial landmarks and visible parts of the head in the head of the target image 125. The target image may include a target face. The target face may not have a neutral facial expression, eyes closed or mouth open, and the age of the person depicted on the target image may be different from the age of the person depicted in the scene video 310. Module 430 may be configured to normalize the target face, such as rotating the head to a neutral state, closing the mouth or opening the eyes of the target face. Facial landmark detection and 3D facial model fitting can be performed using an iterative process. In some embodiments, the iterative process can be optimized for the central processing unit (CPU) and graphics processing unit (GPU) of the mobile device, which can allow the time required for preprocessing of the target image 125 and the scene video 310 to be significantly reduced.

[0057] In some embodiments, the target image pre-processing module 420 may be further configured to apply cosmetic effects and / or change the appearance of a person depicted on the target image 125. For example, the person's hair color or hairstyle may be changed, or the person may be made to look older or younger.

[0058] In some embodiments, the image segmentation and background separation module can be configured to perform segmentation of a person's head from an image of the person. Segmentation of the head can be performed on a target image to obtain an image of the head and / or target face 140. Animation can be further performed on the image of the head or target face 140. The animated head and / or target face 140 can be further inserted back into the background 145. Animating only the image of the head and / or face target 140 by applying a 2D deformation can help avoid unnecessary changes in the background 145 that may be caused by 2D deformation. Since the animation may include changes in head pose, some parts of the background that were previously invisible may become visible, resulting in gaps in the resulting image. In order to fill these gaps, the parts of the background covered by the head can be predicted. In some embodiments, a deep learning model can be trained to perform segmentation of a person's head from an image. Similarly, deep learning techniques can be used for prediction of the background. Refer to the following Figures 8 to 13 Describes the details of deep learning techniques for image segmentation and background prediction.

[0059] In some embodiments, the image animation and refinement module 430 can be configured to animate the target image frame by frame. For each frame of the scene video 310, the change in the position of the control point can be determined. The position change of the control point can be projected onto the target image 125. The module 430 can be further configured to establish a warp field. The warp field can include a set of piecewise linear transformations, which are caused by the change of each triangle in the triangulation of the control point. The module 430 can be further configured to apply the warp field to the target image 125, and thereby generate a frame of the output video 330. Applying the warp field to the image can be performed relatively quickly. It can allow animation to be performed in real time.

[0060] In some embodiments, the image animation and refinement module can be further configured to generate hidden areas, such as inner mouth areas. Several methods can be used to generate hidden areas. One method can include transferring the inside of a person's mouth in the scene video 310 to the inside of a person's mouth in the target image. Another method can include using a 3D mouth model to generate hidden areas. The 3D mouth model can match the geometry of the 3D face model.

[0061] In some embodiments, if the person in the scene video 310 closes his eyes or blinks, the module 430 can be configured to synthesize realistic eyelids in the target image by extrapolation. The skin color of the eyelids can be generated to match the color of the target image. To match the skin color, the module 430 can be configured to transfer the eye expression from the 3D facial model established for the frame of the scene video 310 to the 3D facial model established for the target image 125, and insert the generated eye region into the target image.

[0062] In some embodiments, module 430 can be configured to generate partially occluded areas (like mouth, iris or eyelids) and fine-scale details. Generative adversarial networks can be used to synthesize realistic textures and realistic eye images. Module 430 can be further configured to replace the eyes in hidden areas of the target image with realistic eye images generated using generative adversarial networks. Module 430 can be configured to generate photo-realistic textures and fine-scale details of the target image based on the target image and the original parameters and current parameters of the 3D face model. Module 430 can further refine the target image by replacing the hidden areas with the generated photo-realistic textures and applying the fine-scale details to the entire target image. Applying fine-scale details can include applying a shadow mask to each frame of the target image.

[0063] In some embodiments, module 430 may be further configured to apply other effects (eg, color correction and light correction) on the target image that are needed to make the animation look realistic.

[0064] Other embodiments of the present disclosure may allow for the transfer of not only facial expressions and head movements of objects in a scene video, but also body pose and orientation, gestures, etc. For example, a dedicated hair model may be used to improve hair representation during significant rotations of the head. Generative adversarial networks may be used to synthesize target body poses that mimic source body poses in a realistic manner. Figure 5 is a flow chart illustrating a method 500 for portrait animation according to an exemplary embodiment. The method 500 may be performed by the computing device 110 and the portrait animation system 220 .

[0065] Method 500 may include pre-processing the scene video in blocks 515 to 525. Method 500 may begin in block 505 by detecting, by computing device 110, control points (e.g., 2D facial landmarks) in a frame of the scene video. In block 520, method 500 may include generating, by computing device 110, displacements of the control points in the frame of the scene video. In block 525, method 500 may include fitting, by computing device 110, a 3D facial model to the control points in the frame of the scene video to obtain parameters of the 3D facial model for the frame of the scene video.

[0066] In blocks 530 to 540, method 500 may include preprocessing the target image. The target image may include a target face. In block 530, method 500 may include detecting, by computing device 110, control points (e.g., facial landmarks) in the target image. In block 535, method 500 may include fitting, by computing device 110, a 3D facial model to the control points in the target image to obtain parameters of the 3D facial model for the target image. In block 540, method 500 may include establishing a triangulation of the control points in the target image.

[0067] In block 545 , method 500 may include generating, by computing device 110 and based on parameters of the 3D facial model for the image and parameters of the 3D facial model for the frame of the scene video, deformations of the mouth and eye regions.

[0068] In block 550, method 500 may include generating, by computing device 110, a 2D facial deformation (warp field) based on the displacement of control points in the frame of the scene video and the triangulation of control points in the target image. The 2D facial deformation may include a set of affine transformations of some triangles of the 2D triangulation of the face in the target image and the background. The triangulation topology may be shared between the frame of the scene video and the target image.

[0069] In block 555, method 500 may include applying, by computing device 110, the 2D facial deformation to the target image to obtain a frame of the output video. Method 500 may further include generating, by computing device 110, mouth and eye regions in the target image based on the mouth and eye region deformations.

[0070] In block 560, method 500 may include performing refinement in frames of the output video. The refinement may include color and light correction.

[0071] Thus, the target image 125 can be animated by a series of 2D deformations of the facial transformations in the frames of the generated simulated scene video. This process can be very fast and appear as if the animation is being performed in real time. Some of the 2D deformations can be extracted from the frames of the source video and pre-stored. In addition, background restoration methods can be applied to achieve a photo-realistic effect that animates the target image.

[0072] Figure 6 6 is a flow chart illustrating a method 600 for portrait animation according to some exemplary embodiments. The method 600 may be performed by the computing device 110. The method 600 may begin in box 605 by receiving a scene video by the computing device. The scene video may include at least one input frame. The input frame may include a first face. In box 610, the method 600 may include receiving a target image by the computing device. The target image may include a second face. In box 615, the method 600 may include determining a two-dimensional (2D) deformation by the computing device and based on the at least one input frame and the target image, wherein the 2D deformation, when applied to the second face, modifies the second face to mimic at least a facial expression and a head direction of the first face. In box 620, the method 600 may include applying the 2D deformation to the target image by the computing device to obtain at least one output frame of the output video.

[0073] Figure 7 An exemplary computing system 700 that can be used to implement the methods described herein is shown. The computing system 700 can be implemented in the context of computing device 110, portrait animation system 220, 3D face model 405, sparse correspondence module 410, scene video preprocessing module 415, target image preprocessing module 420, and image animation and refinement module 430, among others.

[0074] like Figure 7As shown, the hardware components of the computing system 700 may include one or more processors 710 and a memory 720. The memory 720 stores instructions and data in part for execution by the processor 710. When the system 700 operates, the memory 720 may store executable code. The system 700 may further include an optional mass storage device 730, an optional portable storage medium drive 740, one or more optional output devices 750, one or more optional input devices 760, an optional network interface 770, and one or more optional peripheral devices 780. The computing system 700 may also include one or more software components 795 (e.g., a software component that can implement the method for portrait animation as described herein).

[0075] Figure 7 The components shown in are depicted as being connected via a single bus 790. The components may be connected via one or more data transmission devices or data networks. The processor 710 and memory 720 may be connected via a local microprocessor bus, and the mass storage device 730, peripheral device 780, portable storage device 740, and network interface 770 may be connected via one or more input / output (I / O) buses.

[0076] Mass storage device 730, which may be implemented as a magnetic disk drive, solid state disk drive, or optical disk drive, is a non-volatile storage device for storing data and instructions for use by processor 710. Mass storage device 730 may store system software (e.g., software component 795) used to implement the embodiments described herein.

[0077] The portable storage media drive 740 operates in conjunction with a portable non-volatile storage medium, such as a compact disk (CD) or digital video disk (DVD), to input and output data and code to and from the computing system 700. System software (e.g., software component 795) for implementing the embodiments described herein may be stored on such a portable medium and input to the computing system 600 via the portable storage media drive 740.

[0078] Optional input device 760 provides part of the user interface. Input device 760 may include an alphanumeric keypad (such as a keyboard) for entering alphanumeric and other information, or a pointing device such as a mouse, trackball, stylus, or cursor direction keys. Input device 760 may also include a camera or scanner. In addition, as Figure 7 The illustrated system 700 includes an optional output device 750. Suitable output devices include speakers, printers, network interfaces, and monitors.

[0079] The network interface 770 can be used to communicate with external devices, external computing devices, servers, and networked systems via one or more communication networks, such as one or more wired, wireless, or optical networks including, for example, the Internet, an intranet, a LAN, a WAN, a cellular telephone network, a Bluetooth radio, and a radio frequency network based on IEEE 802.11, as well as other networks. The network interface 770 can be a network interface card (such as an Ethernet card), an optical transceiver, a radio frequency transceiver, or any other type of device that can send and receive information. The optional peripheral devices 780 can include any type of computer support device to add additional functionality to the computer system.

[0080] The components included in computing system 700 are intended to represent a broad class of computer components. Thus, computing system 700 may be a server, a personal computer, a handheld computing device, a phone, a mobile computing device, a workstation, a minicomputer, a mainframe computer, a network node, or any other computing device. Computing system 700 may also include different bus configurations, networking platforms, multi-processor platforms, etc. Various operating systems (OS) may be used, including UNIX, Linux, Windows, Macintosh OS, Palm OS, and other suitable operating systems.

[0081] Some of the above functions may consist of instructions stored on a storage medium (e.g., a computer-readable medium or a processor-readable medium). The instructions may be retrieved and executed by a processor. Some examples of storage media are memory devices, tapes, disks, etc. The instructions are operable to direct the processor to operate in accordance with the present invention when executed by the processor. Those skilled in the art are familiar with instructions, processors, and storage media.

[0082] It is noteworthy that any hardware platform suitable for performing the processing described herein is applicable to the present invention. The terms "computer-readable storage medium" and "computer-readable storage medium" used herein refer to any one or more media that participate in providing instructions to a processor for execution. Such media can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks or disks, such as fixed disks. Volatile media include dynamic memory, such as system random access memory (RAM). Transmission media include coaxial cables, copper wires, and optical fibers, etc., including lines of an embodiment comprising a bus. Transmission media can also take the form of sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD read-only memory (ROM) disks, DVDs, any other optical media, any other physical media with patterns of marks or holes, RAM, PROM, EPROM, EEPROM, any other storage chip or cassette, carrier waves, or any other medium from which a computer can read.

[0083] Various forms of computer readable media may be involved in transmitting one or more sequences of one or more instructions to the processor for execution. The bus transfers the data to the system RAM, from which the processor retrieves and executes the instructions. The instructions received by the system processor may optionally be stored on a fixed disk before or after execution by the processor.

[0084] Figure 8 8 is a block diagram of a DNN 800 for background prediction according to an exemplary embodiment. DNN 800 may include convolutional layers 804, 828, and 830, compressed convolution blocks 806, 808, 810, 812, and 814, attention blocks 816, 818, 820, 822, 824, and 826, and decompressed convolution blocks 832, 834, 836, 838, and 840.

[0085] The compression convolution blocks 806, 808, 810, 812, and 814 may extract a semantic feature vector from the image 802. The decompression convolution blocks 832, 834, 836, 838, and 840 then transpose the semantic feature vector back to a resulting image 842 using information from the attention blocks 816, 818, 820, 822, 824, and 826. The image 802 may include the target image 125. The resulting image 842 may include a predicted background of a portion of the target image 125 covered by the target face 140.

[0086] Fig. 9 is a block diagram of an exemplary compressed convolution block 900. The compressed convolution block 900 may be used as Figure 8 The compressed convolution block 806, 808, 810, 812 or 814 in the DNN 800. The compressed convolution block 900 may include convolution layers 904 and 906 and a maximum pooling layer 908. The compressed convolution block 900 may generate outputs 910 and 920 based on the feature map 902.

[0087] Fig.10 is a block diagram of an exemplary decompressed convolution block 1000. The decompressed convolution block 1000 may be used as Figure 8 The decompressed convolution block 832, 834, 836, 838 or 840 in the DNN 800 of FIG. The decompressed convolution block 1000 may include convolution layers 1004 and 1008, a concatenated layer 1008, and a transposed convolution layer 1010. The decompressed convolution block 1000 may generate an output 1014 based on the feature map 1002 and the feature map 1012.

[0088] Fig.11 is a block diagram of an exemplary attention block 1100. The attention block 1100 may be used as Figure 8 The attention block 816, 818, 820, 822, 824, or 826 in the DNN 800. The attention block 1100 may include convolutional layers 1104 and 1106, a normalization layer 1108, an aggregation layer 1110, and a concatenation layer 1112. The attention block 1100 may generate a result map 1114 based on the feature map 1102.

[0089] Fig.12 is a block diagram of a learning scheme 1200 for training a DNN 800. The training scheme 1200 may include a loss calculator 1208, a DNN 800, discriminator networks 1212 and 1214, and a difference block 1222. The discriminator networks 1212 and 1214 may facilitate the DNN 800 to generate a photo-realistic background.

[0090] The DNN 800 may be trained based on a generated synthetic dataset. The synthetic dataset may include an image of a person in front of a background image. The background image may be used as a target image 1206. The image of the person in front of the background image may be used as input data 1202, and an insertion mask may be used as input data 1204.

[0091] The discriminator network 1212 can calculate the generator loss 1218 (g_loss) based on the output of the DNN 800 (predicted background). The discriminator network 1214 can calculate the predicted value based on the target image 1206. The difference block 1222 can calculate the discriminator loss 1220 (d_loss) based on the generator loss 1218 and the predicted value. The loss generator 1208 can calculate the training loss 1216 (im_loss) based on the output of the DNN 800 and the target image 1206.

[0092] The learning of DNN 800 may include a combination of the following steps:

[0093] 1. “Training step”. In the “training step”, the weights of the discriminator networks 1212 and 1214 remain unchanged, and im_loss and g_loss are used for back-propagation.

[0094] 2. “Pure training step”. In the “pure training step”, the weights of the discriminator networks 1212 and 1214 remain unchanged, and only im_loss is used for back-propagation.

[0095] 3. “Discriminator training step”. In the “Discriminator training step”, the weights of DNN 800 remain unchanged and d_loss is used for back-propagation.

[0096] The following pseudo code may describe the learning algorithm for DNN 800:

[0097] 1. Perform the “pure training step” 100 times;

[0098] 2. Repeat the following steps until the desired quality is achieved:

[0099] a. Perform the “Discriminator Training Step” 5 times

[0100] b. Execute the “training step”

[0101] Fig.13 is a block diagram of an exemplary discriminator network 1300. The discriminator network 1300 may be used as Fig.12 The discriminator network 1212 and 1214 in the learning scheme 1200 of FIG. The discriminator network 1300 may include a convolution layer 1304, a compressed convolution block 1306, a global average pooling layer 1308, and a dense layer 1310. The discriminator network 1300 may generate a prediction value 1312 based on the image 1302.

[0102] It should be noted that the architecture of the DNN for background prediction can be similar to Figures 8 to 13The architecture of the exemplary DNN 800 described in FIG. 8 is different. For example, the 3×3 convolutions 804-814 may be replaced by a combination of 3×1 convolutions and 1×3 convolutions. The DNN for background prediction may not include Figure 8 Some of the blocks shown. For example, attention blocks 816 to 820 may be excluded from the architecture of the DNN. Figure 8 The DNN for background prediction may also include a different number of hidden layers compared to the DNN 800 shown.

[0103] It should also be noted that similar Figures 8 to 13 The DNN 800 described in the above may be trained and used to predict other parts of the target image 125. For example, the DNN may be used to predict or generate hidden areas and fine-scale details of the target image 125 to achieve a photo-realistic result. The hidden areas may include the mouth area and the eye area of ​​the target face 140.

[0104] The present disclosure has the following implementation modes:

[0105] 1. A method for portrait animation, the method comprising:

[0106] Receiving, by a computing device, a scene video, the scene video comprising at least one input frame, the at least one input frame comprising a first face;

[0107] receiving, by the computing device, a target image, the target image including a second face;

[0108] determining, by the computing device and based on the at least one input frame and the target image, a two-dimensional (2D) deformation, wherein the 2D deformation, when applied to the second face, modifies the second face to mimic at least a facial expression and a head orientation of the first face; and

[0109] The 2D deformation is applied, by the computing device, to the target image to obtain at least one output frame of an output video.

[0110] 2. The method of clause 1, further comprising, before applying the 2D deformation:

[0111] performing segmentation of the target image by the computing device and using a deep neural network (DNN) to obtain an image of the second face and a background; and

[0112] Wherein applying the 2D deformation by the computing device comprises applying the 2D deformation to the image of the second face to obtain a deformed face while keeping the background unchanged.

[0113] 3. The method of clause 2, further comprising, after applying the 2D deformation:

[0114] inserting, by the computing device, the deformed face into the background; and

[0115] predicting, by the computing device and using the DNN, a portion of the background in a gap between the deformed face and the background; and

[0116] The gap is filled with the predicted portion by the computing device.

[0117] 4. The method of clause 1, wherein determining the 2D deformation comprises:

[0118] determining, by the computing device, a first control point on the first face;

[0119] determining, by the computing device, a second control point on the second face; and

[0120] A 2D deformation or affine transformation is defined by the computing device to align the first control point with the second control point.

[0121] 5. The method of clause 4, wherein determining the 2D deformation comprises establishing, by the computing device, a triangulation of the second control points.

[0122] 6. The method of clause 5, wherein determining the 2D deformation further comprises:

[0123] determining, by the computing device, a displacement of the first control point in the at least one input frame;

[0124] projecting the displacement onto the target image by the computing device and using the affine transformation to obtain a desired displacement of the second control point; and

[0125] A warp field to be used as the 2D deformation is determined by the computing device and based on the desired displacement.

[0126] 7. The method of clause 6, wherein the warp field comprises a set of piecewise linear transformations defined by changes in triangles in the triangulation of the second control points.

[0127] 8. The method according to clause 1, further comprising:

[0128] generating, by the computing device, one of a mouth region and an eye region; and

[0129] One of the mouth region and the eye region is inserted, by the computing device, into at least an output frame.

[0130] 9. The method of clause 8, wherein generating one of the mouth region and the eye region comprises transferring, by the computing device, one of the mouth region and the eye region from the first face.

[0131] 10. The method according to clause 8, wherein the mouth region and the eye region are generated

[0132] One of which includes:

[0133] fitting, by the computing device, a 3D facial model to the first control points to obtain a first set of parameters, the first set of parameters including at least a first facial expression;

[0134] fitting, by the computing device, the 3D facial model to the second control points to obtain a second set of parameters, the second set of parameters including at least a second facial expression;

[0135] transferring, by the computing device, the first facial expression from the first set of parameters to the second set of parameters; and

[0136] One of the mouth region and the eye region is synthesized, by the computing device and using the 3D facial model.

[0137] 11. A system for portrait animation, the system comprising at least one processor and a memory storing processor executable code, wherein the at least one processor is configured to

[0138] When the processor executable code is executed, the following operations are implemented:

[0139] Receiving a scene video, the scene video comprising at least one input frame, the at least one input frame comprising a first face;

[0140] receiving a target image, the target image including a second face;

[0141] determining a two-dimensional (2D) deformation based on the at least one input frame and the target image, wherein the 2D deformation, when applied to the second face, modifies the second face to mimic at least a facial expression and a head orientation of the first face; and

[0142] The 2D deformation is applied to the target image to obtain at least one output frame of an output video.

[0143] 12. The system of clause 11, further comprising, before applying the 2D deformation:

[0144] using a deep neural network (DNN) to perform segmentation of the target image to obtain the second facial image and background; and wherein,

[0145] Applying the 2D deformation includes applying the 2D deformation to the image of the second face to obtain a deformed face while keeping the background unchanged.

[0146] 13. The system of clause 12, further comprising after applying the 2D deformation:

[0147] inserting, by the computing device, the deformed face into the background; and

[0148] using the DNN to predict a portion of the background in a gap between the deformed face and the background;

[0149] The gap is filled with the predicted portion.

[0150] 14. The system of clause 11, wherein determining the 2D deformation comprises:

[0151] determining a first control point on the first face;

[0152] determining a second control point on the second face; and

[0153] A 2D deformation or affine transformation is defined for aligning the first control point with the second control point.

[0154] 15. The system of clause 14, wherein determining the 2D deformation comprises establishing a triangulation of the second control points.

[0155] 16. The method of clause 15, wherein determining the 2D deformation further comprises:

[0156] determining a displacement of the first control point in the at least one input frame;

[0157] The displacement is projected onto the target image using the affine transformation to obtain

[0158] The desired displacement of the second control point; and

[0159] Based on the desired displacement, a warp field is determined which will be used as the 2D deformation.

[0160] 17. The system of clause 16, wherein the warp field comprises a set of piecewise linear transformations defined by changes in triangles in the triangulation of the second control points.

[0161] 18. The system of clause 11, wherein the method further comprises:

[0162] generating one of a mouth region and an eye region; and

[0163] One of the mouth region and the eye region is inserted into at least an output frame.

[0164] 19. The system according to clause 18, wherein the mouth region and the eye region are generated

[0165] One of which includes:

[0166] fitting a 3D facial model to the first control points to obtain a first set of parameters, the first set of parameters comprising at least a first facial expression;

[0167] fitting the 3D facial model to the second control points to obtain a second set of parameters, the second set of parameters comprising at least a second facial expression;

[0168] transferring the first facial expression from the first set of parameters to the second set of parameters;

[0169] as well as

[0170] One of the mouth region and the eye region is synthesized using the 3D facial model.

[0171] 20. A non-transitory processor-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to implement

[0172] A method for portrait animation, the method comprising:

[0173] Receiving a scene video, the scene video comprising at least one input frame, the at least one input frame comprising a first face;

[0174] receiving a target image, the target image including a second face;

[0175] A two-dimensional (2D) deformation is determined based on the at least one input frame and the target image, wherein the 2D deformation is

[0176] when applied to the second face, modifying the second face to mimic at least the facial expression and head orientation of the first face; and

[0177] The 2D deformation is applied to the target image to obtain at least one output frame of an output video.

[0178] Thus, a method and system for photo-realistic real-time portrait animation has been described. Although the embodiments have been described with reference to specific exemplary embodiments, it will be apparent that various modifications and changes may be made to these exemplary embodiments without departing from the broader spirit and scope of the present application. Accordingly, the description and drawings should be regarded as illustrative rather than restrictive.

Claims

1. A method for portrait animation, the method comprising: receiving, by a computing device, scene data including information about a first face; receiving, by the computing device, a target image including a second face; determining, by the computing device and based on the target image and the information about the first face, a two-dimensional deformation of a second face in the target image; as well as The two-dimensional deformation is applied to the target image by the computing device to obtain at least one output frame of an output video.

2. The method according to claim 1, wherein: The scene data includes a scene video including at least one input frame, the at least one input frame including an image of a first face; and When the two-dimensional deformation is applied to the second face, the second face is modified to mimic at least a facial expression of the first face.

3. The method according to claim 1, wherein: The scene data includes a list of facial expressions and movements of the first face; and When the two-dimensional deformation is applied to the second face, the second face is modified to mimic the facial expression and the movement of the first face. 4 . The method of claim 3 , further comprising providing a user interface to a user to enable the user to generate a list of the facial expressions of the first face.

5. The method according to claim 1, further comprising: determining, based on the information about the first face, that at least one eye of the first face is closed by an eyelid; responsive to the determination, synthesizing an image of the closed eyelids in the second face; as well as The image of the closed eyelid is inserted into the target image.

6. The method according to claim 5, wherein: The color of the closed eyelids matches the color of the target face.

7. The method according to claim 1, further comprising: generating, by the computing device and based on the target image and the three-dimensional facial model, fine-scale details for the target image; as well as The fine scale of detail is applied to the target image by the computing device.

8. The method according to claim 7, wherein: Applying fine scale details includes applying one or more shadow masks to the target image.

9. The method according to claim 1, further comprising: Determining, by the computing device, that at least one area of ​​the second face is blocked; In response to the determination, the image of the at least one region is synthesized by the computing device and the image of the at least one region is inserted into the target image.

10. The method according to claim 9, wherein: The image of the at least one region is synthesized by a neural network.

11. A computing device comprising: processor; as well as a memory storing instructions that, when executed by the processor, configure the computing device to: receiving scene data including information about a first face; receiving a target image including a second face; determining a two-dimensional deformation of a second face in the target image based on the target image and the information about the first face; as well as The two-dimensional deformation is applied to the target image to obtain at least one output frame of an output video.

12. The computing device of claim 11, wherein: The scene data includes a scene video including at least one input frame, the at least one input frame including an image of a first face; and When the two-dimensional deformation is applied to the second face, the second face is modified to mimic at least a facial expression of the first face.

13. The computing device of claim 11, wherein: The scene data includes a list of facial expressions and movements of the first face; and When the two-dimensional deformation is applied to the second face, the second face is modified to mimic the facial expression and the movement of the first face.

14. The computing device of claim 13, wherein: The instructions further configure the computing device to provide a user interface to a user to enable the user to generate a list of the facial expressions of the first face.

15. The computing device of claim 11, wherein: The instructions further configure the computing device to: determining, based on the information about the first face, that at least one eye of the first face is closed by an eyelid; responsive to the determination, synthesizing an image of the closed eyelids in the second face; as well as The image of the closed eyelid is inserted into the target image.

16. The computing device of claim 15, wherein: The color of the closed eyelids matches the color of the target face.

17. The computing device of claim 11, wherein: The instructions further configure the computing device to: generating fine-scale details for the target image based on the target image and the three-dimensional facial model; and The fine scale details are applied to the target image.

18. The computing device of claim 17, wherein: Applying fine scale details includes applying one or more shadow masks to the target image.

19. The computing device of claim 11, wherein: The instructions further configure the computing device to: determining that at least one region of the second face is blocked; as well as In response to the determination, the image of the at least one region is synthesized and inserted into the target image.

20. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computing device, cause the computing device to: receiving scene data including information about a first face; receiving a target image including a second face; determining a two-dimensional deformation of a second face in the target image based on the target image and the information about the first face; as well as The two-dimensional deformation is applied to the target image to obtain at least one output frame of an output video.