System and method for template-based generation of personalized video

By receiving video configuration data and source facial images, the processor generates output video, solving the problem that existing messenger applications cannot realize complex video editing, and realizing the function of generating personalized videos based on templates.

CN120088372APending Publication Date: 2025-06-03SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510427483.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-10-23
Filing Date
2020-01-18
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing messenger applications cannot implement complex video editing, such as replacing one face with another, requiring the use of complex third-party video editing software.

Method used

Through a processor configuration, video configuration data and source facial images are received to generate output video. The video configuration data includes a sequence of frame images, facial area parameters, and facial landmark parameters, which are used to modify the source facial image and insert it into the frame image.

Benefits of technology

The function of generating personalized videos based on templates is realized, and users can perform complex video editing in simple ways, such as facial replacement, without the need to use third-party software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088372A_ABST
    Figure CN120088372A_ABST
Patent Text Reader

Abstract

Systems and methods for generating a personalized video based on a template are disclosed. An example method may begin with receiving video configuration data that includes a sequence of frame images, a sequence of face region parameters that define a location of a face region in the frame images, and a sequence of face landmark parameters that define a location of a face landmark in the frame images. The method may continue to receive an image of a source face. The method may also include generating an output video. Generation of the output video may include modifying a frame image of a sequence of frame images. Specifically, an image of a source face may be modified to obtain an additional image characterizing the source face that employs a facial expression corresponding to a facial landmark parameter. Additional images may be inserted into the frame image at positions determined by facial region parameters corresponding to the frame image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application, and the application number of its parent application is 202080009459.1, the application date is January 18, 2020, and the invention title is "System and Method for Generating Personalized Videos Based on Templates". Technical Field

[0002] The present disclosure generally relates to digital image processing. More specifically, the present disclosure relates to methods and systems for generating personalized videos based on templates. Background Art

[0003] Sharing media such as stickers and emojis has become a standard option in messaging applications (also referred to herein as messengers). Currently, some messengers provide users with options to generate images and short videos and send the images and short videos to other users via a communication chat. Certain existing messengers allow users to modify short videos before transmission. However, the modification of short videos provided by existing messengers is limited to visual effects, filters, and text. Users of current messengers cannot perform complex edits (e.g., replacing one face with another). Such video editing cannot be provided by current messengers and requires complex third-party video editing software. Summary of the Invention

[0004] The purpose of this section is to introduce selected concepts in a simplified form, and the specific content of these concepts is described in the detailed implementation section below. The summary of the invention is not intended to determine the key features or main features of the claimed subject matter, nor is it intended to assist in determining the scope of the claimed subject matter.

[0005] According to an embodiment of the present disclosure, a system for generating personalized videos based on templates is disclosed. The system may include at least one processor and a memory storing processor-executable code. The at least one processor may be configured to receive video configuration data by a computing device. The video configuration data may include: a sequence of frame images, a sequence of facial region parameters defining the positions of facial regions in the frame images, and a sequence of facial landmark parameters defining the positions of facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The at least one processor may be configured to receive an image of a source face by a computer device. The at least one processor may be configured to generate an output video by a computing device. The generation of the output video may include modifying the frame images of the sequence of frame images. Specifically, the image of the source face may be modified based on the facial landmark parameters corresponding to the frame images to obtain additional images, and the additional images represent the source face with the facial expressions corresponding to the facial landmark parameters. The additional images may be inserted into the frame images at positions determined by the facial region parameters corresponding to the frame images.

[0006] According to an exemplary embodiment of the present disclosure, a method for generating a personalized video based on a template is disclosed. The method may start with receiving video configuration data through a computing device. The video configuration data may include: a sequence of frame images, a sequence of facial region parameters defining the position of a facial region in the frame images, and a sequence of facial landmark parameters defining the position of facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The method may continue with receiving an image of a source face by a computer device. The method may also include generating an output video by the computing device. The generation of the output video may include modifying the frame images of the sequence of frame images. Specifically, the image of the source face may be modified to obtain another image that represents the source face adopting the facial expression corresponding to the facial landmark parameter. The modification of the image may be performed based on the facial landmark parameter corresponding to the frame image. The another image may be inserted into the frame image at a position determined by the facial region parameter corresponding to the frame image.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory processor-readable medium that stores processor-readable instructions. When the processor-readable instructions are executed by a processor, they cause the processor to implement the above-described method for generating a personalized video based on a template.

[0008] Additional objects, advantages, and novel features of the examples will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following description and the drawings, or may be learned by practice of the examples. The objectives and advantages of the concepts may be realized and attained by means of the methods, instrumentalities and combinations particularly pointed out in the appended claims. Description of the Drawings

[0009] Embodiments are illustrated in the drawings by way of example and not limitation, in which like reference numerals indicate similar elements.

[0010] Figure 1 is a block diagram showing an exemplary environment in which a system and method for generating a personalized video based on a template may be implemented.

[0011] Figure 2 is a block diagram showing an exemplary embodiment of a computing device for implementing a method for generating a personalized video based on a template.

[0012] Figure 3 is a flowchart showing a process for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure.

[0013] Figure 4 is a flowchart showing the functions of a system for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure.

[0014] Figure 5 is a flowchart showing a process for generating a live-action video for generating a video template according to some exemplary embodiments.

[0015] Figure 6 Frames of an example live-action video for generating a video template according to some exemplary embodiments are shown.

[0016] Figure 7 An original image of a face and an image of the face with normalized illumination according to an exemplary embodiment are shown.

[0017] Figure 8 A segmented head image, a head image with facial landmarks, and a facial mask according to an exemplary embodiment are shown.

[0018] Figure 9 Frames representing a user's face, a skin mask, and the result of recoloring the skin mask according to an exemplary embodiment are shown.

[0019] Figure 10 An image of the face of a facial synchronization actor, an image of the facial landmarks of the facial synchronization actor, an image of the facial landmarks of the user, and an image of the user's face with the facial expression of the facial synchronization actor according to an exemplary embodiment are shown.

[0020] Figure 11 A segmented facial image, a hair mask, a hair mask warped to a target image, and a hair mask applied to the target image according to an exemplary embodiment are shown.

[0021] Figure 12 An original image of an eye, an image with a reconstructed eye sclera, an image with a reconstructed iris, and an image with a moving reconstructed iris according to an exemplary embodiment are shown.

[0022] Figure 13 and Figure 14 Frames of an example personalized video generated based on a video template according to some exemplary embodiments are shown.

[0023] Figure 15 is a flowchart showing a method for generating a personalized video based on a template according to an exemplary embodiment of the present disclosure.

[0024] Figure 16 An example computer system that can be used to implement the method for generating a personalized video based on a template is shown. Detailed Description

[0025] The following detailed description of the embodiments includes reference to the accompanying drawings that form a part of the detailed description. The methods described in this section are not prior art to the claims, and are not admitted to be prior art by inclusion in this section. The drawings illustrate the description according to exemplary embodiments. The exemplary embodiments, also referred to herein as "examples," are described in sufficient detail to enable those skilled in the art to practice the subject matter. Embodiments may be combined, other embodiments may be utilized, or structural, logical, and operational changes may be made without departing from the scope claimed. Accordingly, the following detailed description should not be considered limiting, and the scope is defined by the appended claims and their equivalents.

[0026] For the purposes of this patent document, unless otherwise stated or clearly meant otherwise in the context in which it is used, the terms "or" and "and" shall mean "and / or". Unless otherwise stated or where the use of "one or more" is clearly inappropriate, the term "one" shall mean "one or more". The terms "comprise", "comprising", "include" and "including" are interchangeable and are not intended to be limiting. For example, the term "including" shall be construed to mean "including but not limited to".

[0027] The present disclosure relates to methods and systems for generating personalized videos based on templates. The embodiments provided by the present disclosure solve at least some problems of the prior art. The present disclosure may be designed to work in real time on a mobile device such as a smart phone, a tablet computer, or a telephone, but the embodiments may be extended to methods involving network services or cloud-based resources. The methods described herein may be implemented by software running on a computer system and / or by hardware utilizing a combination of microprocessors or other specially designed application specific integrated circuits (ASICs), programmable logic devices, or any combination thereof. Specifically, the methods described herein may be implemented by a series of computer-executable instructions residing on a non-transitory storage medium (such as a disk drive or a computer-readable medium).

[0028] Some embodiments of the present disclosure may allow for the real-time generation of personalized videos on a user computing device such as a smart phone. The personalized videos may be generated in the form of audiovisual media (e.g., video, animation, or any other type of media) that depicts the face of one user or the faces of multiple users. The personalized videos may be generated based on pre-generated video templates. The video templates may include video configuration data. The video configuration data may include a sequence of frame images, a sequence of facial region parameters that define the position of the facial region in the frame images, and a sequence of facial landmark parameters that define the position of the facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The frame images may be generated based on an animated video or a live-action video of a person. The facial landmark parameters may be generated based on another live-action video that depicts the face of an actor (also referred to as facesync as described in more detail below), an animated video, an audio file, text, or manually.

[0029] The video configuration file may further include a sequence of skin masks. The skin masks may define the skin region of the body of the actor depicted in the frame images or the skin region of a 2D / 3D animation of the body. In one exemplary embodiment, the skin masks and the facial landmark parameters may be generated based on two different live-action videos of different actors (referred to herein as the actor and the facesync actor, respectively). The video configuration data may further include a sequence of mouth region images and a sequence of eye parameters. The eye parameters may define the position of the iris in the sclera of the facesync actor depicted in the frame images. The video configuration data may include head parameters and a sequence of other parameters of the head that define the rotation, turn, position, and scale of the head. When taking an image and looking directly at the camera, the user may keep their head still, and thus, the scale and rotation of the head may be adjusted manually. The head parameters may be transferred from a different actor (also referred to herein as the facesync actor). As used herein, the facesync actor is the person whose facial landmark parameters are being used, and the actor is the other person whose body is being used in the video template and whose skin may be recolored, and the user is the person who takes an image of his / her face to generate the personalized video. Thus, in some embodiments, the personalized video includes the user's face modified to have the facial expression of the facesync actor and includes the body of the actor taken from the video template and recolored to match the user's facial color. The video configuration data includes a sequence of animated object images. Optionally, the video configuration data includes a soundtrack and / or voice.

[0030] Pre-generated video templates can be stored remotely in cloud-based computing resources and downloaded by users of computing devices such as smartphones. A user of a computing device can capture an image of a face through the computing device or select an image of a face from a camera roll, from a prepared set of images, or via a network link. In some embodiments, the image can include the face of an animal rather than a human, or can be in the form of a drawing. Based on one of the face-based image and the pre-generated video template, the computing device can also generate a personalized video. The user can send the personalized video to another user of another computing device via a communication chat, share it on social media, download it to the local storage device of the computing device, or upload it to a cloud storage device or a video sharing service.

[0031] According to one embodiment of the present disclosure, an example method for generating a personalized video based on a template can include receiving, by a computing device, video configuration data. The video configuration data can include a sequence of frame images, a sequence of face region parameters defining the position of a face region in the frame images, and a sequence of face landmark parameters defining the position of face landmarks in the frame images. Each face landmark parameter can correspond to a facial expression of a face-synchronized actor. The method can continue by the computing device receiving an image of a source face and generating an output video. The generation of the output video can include modifying the frame images of the sequence of frame images. The modification of the frame images can include modifying the image of the source face to obtain an additional image that represents the source face with a facial expression corresponding to the face landmark parameter, and inserting the additional image into the frame image at a position determined by the face region parameter corresponding to the frame image. Additionally, for example, the source face can be modified by changing the color, making the eyes larger, etc. The image of the source face can be modified based on the face landmark parameter corresponding to the frame image.

[0032] Now referring to the drawings, exemplary embodiments will be described. The drawings are schematic diagrams of idealized exemplary embodiments. Therefore, the exemplary embodiments discussed herein should not be construed as limited to the specific illustrations presented herein; rather, as will be apparent to those skilled in the art, these exemplary embodiments can include departures from and differences from the illustrations presented herein.

[0033] Figure 1Figure 100 shows an example environment in which systems and methods for generating personalized videos based on templates can be implemented. Environment 100 may include computing device 105, user 102, computing device 110, user 104, network 120, and messenger service system 130. Computing device 105 and computing device 110 may refer to mobile devices such as mobile phones, smartphones, or tablets. In other embodiments, computing device 110 may refer to a personal computer, laptop computer, netbook, set-top box, television device, multimedia device, personal digital assistant, gaming console, entertainment system, infotainment system, in-vehicle computer, or any other computing device.

[0034] Computing device 105 and computing device 110 may be communicatively coupled to messenger service system 130 via network 120. Messenger service system 130 may be implemented as cloud-based computing resources. Messenger service system 130 may include computing resources (hardware and software) available at a remote location and accessible via a network (e.g., the Internet). Cloud-based computing resources may be shared by multiple users and may be dynamically reallocated based on demand. Cloud-based computing resources may include one or more server farms / clusters that include a collection of computer servers that may be co-located with network switches and / or routers.

[0035] Network 120 may include any wired network, wireless network, or optical network (e.g., including the Internet, intranet, local area network (LAN), personal area network (PAN), wide area network (WAN), virtual private network (VPN), cellular telephone network (e.g., Global System for Mobile Communications (GSM)), etc.).

[0036] In some embodiments of the present disclosure, computing device 105 may be configured to initiate a communication chat between user 102 and user 104 of computing device 110. During the communication chat, user 102 and user 104 may exchange text messages and videos. The videos may include personalized videos. Personalized videos may be generated based on pre-generated video templates stored in computing device 105 or computing device 110. In some embodiments, the pre-generated video templates may be stored in messenger service system 130 and downloaded to computing device 105 or computing device 110 on demand.

[0037] Messenger service system 130 may include system 140 for preprocessing videos. System 140 may generate video templates based on animated videos or live-action videos of real people. Messenger service system 130 may include video template database 145 for storing video templates. The video templates may be downloaded to computing device 105 or computing device 110.

[0038] The messenger service system 130 may also be configured to store user profiles 135. The user profiles 135 may include images of the faces of user 102, user 104, and the faces of other people. Images of the faces may be downloaded to the computing device 105 or the computing device 110 on demand and based on a license. Additionally, an image of the face of user 102 may be generated using the computing device 105 and stored in the local memory of the computing device 105. An image of the face may be generated based on other images stored in the computing device 105. The computing device 105 may also use the image of the face to generate a personalized video based on a pre-generated video template. Similarly, the computing device 110 may be used to generate an image of the face of user 104. The image of the face of user 104 may be used to generate a personalized video on the computing device 110. In other embodiments, the image of the face of user 102 and the image of the face of user 104 may be used interchangeably to generate a personalized video on the computing device 105 or the computing device 110.

[0039] Figure 2 is a block diagram showing one exemplary embodiment of a computing device 105 (computing device 110) for implementing a method for generating a personalized video. In Figure 2 the example shown, the computing device 110 includes both hardware components and software components. Specifically, the computing device 110 includes a camera 205 or any other image capture device or scanner for acquiring digital images. The computing device 110 may also include a processor module 210 and a storage module 215 for storing software components and processor-readable (machine-readable) instructions or code that, when executed by the processor module 210, cause the computing device 105 to perform at least some of the steps of the method for generating a personalized video based on a template as described herein. The computing device 105 may include a graphics display system 230 and a communication module 240. In other embodiments, the computing device 105 may include additional or different components. Additionally, the computing device 105 may include fewer components that perform functions similar or equivalent to those depicted in Figure 2 here.

[0040] The computing device 110 may also include a messenger 220 for initiating a communication chat with another computing device (such as the computing device 110) and a system 250 for generating a personalized video based on a template. The system 250 will be described in more detail below with reference to Figure 4 here. The messenger 220 and the system 250 may be implemented as software components and processor-readable (machine-readable) instructions or code stored in the memory storage device 215 that, when executed by the processor module 210, cause the computing device 105 to perform at least some of the steps of the methods for providing a communication chat and generating a personalized video as described herein.

[0041] In some embodiments, the system 250 for generating personalized videos based on templates may be integrated in the messenger 220. The user interface of the messenger 220 and the system 250 for template-based personalized videos may be provided via the graphics display system 230. A communication chat may be initiated via the communication module 240 and the network 120. The communication module 240 may include a GSM module, a WIFI module, a Bluetooth TM module, etc.

[0042] Figure 3 FIG. is a flowchart showing steps of a process 300 for generating personalized videos based on templates according to some exemplary embodiments of the present disclosure. The process 300 may include production 305, post-production 310, resource preparation 315, skin recoloring 320, lip-syncing and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340. The resource preparation 315 may be performed by the system 140 for preprocessing videos in the messenger service system 130 (shown in Figure 1 ). The result of the resource preparation 315 is the generation of a video template that may include video configuration data.

[0043] The skin recoloring 320, lip-syncing and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340 may be performed by the system 250 for generating personalized videos based on templates in the computing device 105 (shown in Figure 2 ). The system 250 may receive an image of the user's face and the video configuration data and generate a personalized video representing the user's face.

[0044] The skin recoloring 320, lip-syncing and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340 may be performed by the system 140 for preprocessing videos in the messenger service system 130 (shown in Figure 1 ). The system 140 may receive a test image of the user's face and a video configuration profile. The system 140 may generate a test personalized video representing the user's face. The operator may check the test personalized video. Based on the result of the check, the video configuration profile may be stored in the video template database 145 and then downloaded to the computing device 105 or the computing device 110.

[0045] Production 305 may include idea and scene creation, pre-production (during which locations, props, actors, costumes, and effects are identified), and the production itself, which may require one or more recording sessions. In some exemplary embodiments, recording may be performed by recording scenes / actors against a chroma key background (also referred to herein as a green screen or chroma key screen). To allow for subsequent head tracking and resource cleanup, an actor may wear a chroma key face mask with tracking markers (e.g., balaclavas) that cover the actor's face but leave the neck and the bottom of the chin exposed. In Figure 5 Idea and scene creation are shown in detail in

[0046] In one exemplary embodiment, pre-production and subsequent production steps 305 are optional. Instead of recording actors, two-dimensional or three-dimensional animations may be created or third-party footage / images may be used. Additionally, the original background of the user's image may be used.

[0047] Figure 5 FIG. is a block diagram showing a process 500 for generating a live-action video. The live-action video may also be used to generate video templates for generating personalized videos. Process 500 may include generating an idea at step 505 and creating a scene at step 510. Process 500 may continue with pre-production at step 515, followed by production 305. Production 305 may include recording using a chroma key screen 525 or at a real-life location 530.

[0048] Figure 6 Frames of an example live-action video for generating a video template are shown. Frames of videos 605 and 615 are recorded at a real-life location 530. Frames of videos 610, 620, and 625 are recorded using a chroma key screen 525. An actor may wear a chroma key face mask 630 that has tracking markers covering the actor's face.

[0049] Post-production 310 may include video editing or animation, visual effects, cleanup, sound design, and voice recording.

[0050] During resource preparation 315, the resources to be further deployed may include the following components: background shots without the head of the actor (i.e., the cleaned-up background prepared to remove the head of the actor); shots of the actor on a black background (only for the recorded personalized video); the foreground sequence of the frames; example shots with a generic head and soundtrack; the coordinates of the head position, rotation, and scale; animated elements attached to the head (optional); soundtracks with and without narration; narration in a separate file (optional), etc. All these components are optional and can be presented in different formats. The number and configuration of the components depend on the format of the personalized video. For example, for a customized personalized video, no narration is required, and if the original background from the user's picture is used, etc., no background shots and head coordinates are required. In an exemplary embodiment, instead of preparing a file with coordinates, the area where the face needs to be located may be indicated (e.g., manually).

[0051] Skin recoloring 320 allows the color of the skin of the actor in the personalized video to be matched to the color of the face on the user's image. To implement this step, a skin mask may be prepared that specifically indicates which part of the background must be recolored. Preferably, there is a separate mask for each body part of the actor (neck, left hand, right hand, etc.).

[0052] Skin recoloring 320 may include facial image illumination normalization. Figure 7 Figure 705 shows the original image of the face and the image 710 of the face with normalized illumination according to an exemplary embodiment. Shadows or highlights caused by uneven illumination affect the color distribution and may result in the skin color being too dark or too bright after recoloring. To avoid this, the shadows and highlights in the user's face may be detected and removed. The facial image illumination normalization process includes the following steps. A deep convolutional neural network may be used to transform the image of the user's face. The network may receive the original image 705 in the form of a portrait image taken under any illumination and, while keeping the subject in the original image 705 the same, change the illumination of the original image 705 to make the original image 705 have uniform illumination. Therefore, the input of the facial image illumination normalization process includes the original image 705 in the form of an image of the user's face and facial landmarks. The output of the facial image illumination normalization process includes the image 710 of the face with normalized illumination.

[0053] Skin recoloring 320 may include mask creation and body statistics. There may be either a mask for the entire skin or separate masks for body parts. Additionally, different masks may be created for different scenes in the video (e.g., due to significant lighting changes). The masks may be created semi-automatically with some human guidance using techniques such as keying. The prepared masks may be merged into the video resource and then used in the recoloring. Also, to avoid unnecessary computations in real time, color statistics may be pre-computed for each mask. The statistics may include the mean, median, standard deviation, and some percentiles for each color channel. The statistics may be computed in the red, green, and blue (RGB) color space as well as other color spaces (hue, saturation, value (HSV) color space, CIELAB color space (also known as CIEL*a*b* or abbreviated as the "LAB" color space), etc.). The input to the mask creation process may include a grayscale mask for a body part of an actor with uncovered skin in the form of a video or image sequence. The output of the mask creation process may include the masks compressed and merged into the video and color statistics for each mask.

[0054] Skin recoloring 320 may also include facial statistics calculation. Figure 8 A segmented head image 805 according to one exemplary embodiment is shown, the segmented head image 805 having facial landmarks 810 and a facial mask 815. Based on the segmentation of the user's head image and facial landmarks, the user's facial mask 815 may be created. Regions such as the eyes, mouth, hair, or accessories (such as glasses) may not be included in the facial mask 815. The user's segmented head image 805 and facial mask may be used to calculate statistics for the user's facial skin. Thus, the input to the facial statistics calculation may include the user's segmented head image 805, facial landmarks 810, and face segmentation, and the output of the facial statistics calculation may include color statistics for the user's facial skin.

[0055] Skin recoloring 320 may further include skin color matching and recoloring. Figure 9Shows a frame 905 representing a user's face, a skin mask 910, and a result 915 of recoloring the skin mask 910 according to an exemplary embodiment. Skin color matching and recoloring can be performed using statistics describing the color distribution in the skin of an actor and the skin of the user, and the recoloring of the background frame can be performed in real time on a computing device. For each color channel, distribution matching can be performed and the values of the background pixels can be modified so that the distribution of the transformed values is close to the distribution of the face values. Distribution matching can be performed assuming that the color distribution is normal, or by applying techniques such as multi-dimensional probability density function transfer. Thus, the input to the skin color matching and recoloring process can include the background frame, the actor skin mask of the frame, the actor body skin color statistics for each mask, and the user face skin color statistics, and the output can include the background frame with the skin of all uncovered body parts recolored.

[0056] In some embodiments, to apply skin recoloring 320, several actors with different skin colors can be recorded, and then a version of the personalized video with the skin color closest to the skin color of the user's image can be used.

[0057] In an exemplary embodiment, instead of skin recoloring 320, a predetermined look-up table (LUT) can be used to adjust the color of the face for the lighting of the scene. The LUT can also be used to change the color of the face, for example, to make the face green.

[0058] Lip sync and facial reenactment 325 can produce realistic facial animations. Figure 10 Shows an example process of lip sync and facial reenactment 325. Figure 10 Shows an image 1005 of a face-synced actor's face, an image 1010 of face-synced actor face landmarks, an image 1015 of the user's face landmarks, and an image 1020 of the user's face with the facial expression of the face-synced actor according to an exemplary embodiment. The steps of lip sync and facial reenactment 325 can include recording the face-synced actor and preprocessing the source video / image to obtain the image 1005 of the face-synced actor's face. Then, as shown in the image 1010 of face-synced actor face landmarks, the face landmarks can be extracted. The steps can also include gaze tracking the face-synced actor. In some embodiments, instead of recording the face-synced actor, a pre-prepared animated 2D or 3D face and mouth region model can be used. The animated 2D or 3D face and mouth region model can be generated by machine learning techniques.

[0059] Optionally, fine-tuning of facial landmarks can be performed. In some exemplary embodiments, the fine-tuning of facial landmarks is performed manually. These steps can be performed in the cloud when preparing the video profile. In some exemplary embodiments, these steps can be performed during resource preparation 315. Then, as shown in the image 1015 of the user's facial landmarks, the user's facial landmarks can be extracted. The next step of synchronization and facial reenactment 325 can include animating the target image with the extracted landmarks to obtain an image 1020 of the user's face with the facial expression of the facial synchronization actor. The steps can be performed on a computing device based on an image of the user's face. The animation method is described in detail in U.S. Patent Application No. 16 / 251,472, the disclosure of which is incorporated herein by reference in its entirety. Lip synchronization and facial reenactment 325 can also be enriched with AI-generated head turns.

[0060] In some exemplary embodiments, after the user captures an image, a three-dimensional model of the user's head can be created. In this embodiment, the steps of lip synchronization and facial reenactment 325 can be omitted.

[0061] Hair animation 330 can be performed to animate the user's hair. For example, if the user has hair, the hair can be animated when the user moves or rotates his head. Hair animation 330 is shown in Figure 11 FIG. Figure 11 FIG. shows a segmented facial image 1105, a hair mask 1110, the hair mask 1110 moved to the facial image 1115, and the hair mask 1110 applied to the facial image 1120 according to an exemplary embodiment. Hair animation 330 can include one or more of the following steps: classifying the hair type, modifying the appearance of the hair, modifying the hairstyle, making the hair longer, changing the color of the hair, cutting the hair, and animating the hair, etc. As Figure 11 shown, a facial image can be obtained in the form of a segmented facial image 1105. Then, the hair mask 1110 can be applied to the segmented facial image 1105. Image 1115 shows the hair mask 1110 moved to the facial image. Image 1120 shows the hair mask 1110 applied to the facial image. Hair animation 330 is described in detail in U.S. Patent Application No. 16 / 551,756, the disclosure of which is incorporated herein by reference in its entirety.

[0062] Eye animation 335 can make the user's facial expression more realistic. In Figure 12Eye animation 335 is shown in detail. The processing of eye animation 335 may consist of the following steps: reconstruction of the eye region of the user's face, gaze movement step, and blink step. In the reconstruction process of the eye region, the eye region is segmented into the following parts: eyeball, iris, pupil, eyelashes, and eyelids. If some parts of the eye region (e.g., iris or eyelids) are not fully visible, the complete texture of that part can be synthesized. In some embodiments, a 3D deformable model of the eye can be fitted, and the 3D shape of the eye and the texture of the eye can be obtained. Figure 12 Shows the original image 1205 of the eye, the image 1210 with the reconstructed sclera of the eye, and the image 1215 with the reconstructed iris.

[0063] The gaze movement step includes tracking the gaze direction and pupil position in the video of the face-synchronized actor. If the eye movement of the face-synchronized actor is not rich enough, the data can be manually edited. Then, the gaze movement can be transferred to the user's eye region by synthesizing a new eye image with a transformed eye shape and the same iris position as that of the face-synchronized actor. Figure 12 Shows the image 1220 with the reconstructed moving iris.

[0064] During the blink step, the visible part of the user's eyes can be determined by tracking the eyes of the face-synchronized actor. The changed appearance of the eyelids and eyelashes can be generated based on the reconstruction of the eye region.

[0065] If a generative adversarial network (GAN) is used for face reenactment, the steps of eye animation 335 can be performed explicitly (as described above) or implicitly. In the latter case, the neural network can implicitly capture all the necessary information from the images of the user's face and the source video.

[0066] During deployment 340, the user's face can be realistically animated and automatically inserted into the shot template. The files from the previous steps (resource preparation 315, skin recoloring 320, lip sync and face reenactment 325, hair animation 330, and eye animation 335) can be used as the data of the configuration file. An example of a personalized video with a predetermined set of user faces can be generated for initial inspection. After eliminating the problems identified during the inspection, the personalized video can be deployed.

[0067] The configuration file may also include components that allow text parameters indicating customized personalized videos. A customized personalized video is a personalized video that allows a user to add any text desired by the user on top of the final video. The generation of personalized videos with customized text messages is described in more detail in U.S. Patent Application No. 16 / 661,122, entitled "SYSTEMS AND METHOD FOR GENERATING PERSONALIZED VIDEOS WITH CUSTOMIZED TEXT MESSAGES," filed on October 23, 2019, the disclosure of which is incorporated herein by reference in its entirety.

[0068] In one exemplary embodiment, the generation of the personalized video may further include steps of generating a distinct head turn of the user's head; body animation that changes clothing; facial enhancement (such as hair style change, beautification, adding accessories, etc.); changing scene lighting; synthesizing a voice that reads / sings the text input by the user or converting the voice to a voice that matches the user's voice; gender switching; constructing a background and foreground according to user input; and so on.

[0069] Figure 4It is a schematic diagram showing the functions 400 of a system 250 for generating personalized videos based on templates according to some exemplary embodiments. The system 250 can receive an image of a source face shown as a user face image 405 and a video template including video configuration data 410. The video configuration data 410 can include a data sequence 420. For example, the video configuration data 410 can include: a sequence of frame images, a sequence of facial region parameters defining the position of the facial region in the frame images, and a sequence of facial landmark parameters defining the position of the facial landmarks in the frame images. Each facial landmark parameter can correspond to a facial expression. The sequence of frame images can be generated based on an animated video or a live-action video of a real person. The sequence of facial landmark parameters can be generated based on a live-action video of a real person representing the face of a facial synchronization actor. The video configuration data 410 can also include a skin mask, eye parameters, mouth region images, head parameters, animated object images, preset text parameters, etc. The video configuration data can include a sequence of skin masks defining the skin regions of the bodies of at least one actor represented in the frame images. In one exemplary embodiment, the video configuration data 410 can also include a sequence of mouth region images. Each mouth region image can correspond to at least one frame image. In another exemplary embodiment, the video configuration data 410 can include a sequence of eye parameters defining the position of the iris in the sclera of the face of the facial synchronization actor represented in the frame images and / or a sequence of head parameters defining the rotation, turning, scale, and other parameters of the head. In another exemplary embodiment, the video configuration data 410 can also include a sequence of animated object images. Each animated object image can correspond to at least one frame image. The video configuration data 410 can also include a soundtrack 450.

[0070] The system 250 can determine user data 435 based on the user face image 405. The user data can include the user's facial landmarks, user facial mask, user color data, user hair mask, etc.

[0071] System 250 can generate frames 445 of an output video that is presented as personalized video 440 based on user data 435 and data sequence 420. System 250 can also add a soundtrack to personalized video 440. Personalized video 440 can be generated by modifying the frame images of a sequence of frame images. The modification of the frame images can include: modifying user face image 405 to obtain an additional image that represents a source face with a facial expression corresponding to facial landmark parameters. The modification can be performed based on the facial landmark parameters corresponding to the frame image. The additional image can be inserted into the frame image at a location determined by the facial region parameters corresponding to the frame image. In one exemplary embodiment, the generation of the output video can further include determining color data associated with the source face and recoloring the skin regions in the frame image based on the color data. Additionally, the generation of the output video includes inserting a mouth region corresponding to the frame image into the frame image. Other steps in generating the output video can include generating an image of an eye region based on eye parameters corresponding to the frame and inserting the image of the eye region into the frame image. In one exemplary embodiment, the generation of the output video can further include determining a hair mask based on the source face image, generating a hair image based on the hair mask and head parameters corresponding to the frame image, and inserting the hair image into the frame image. Additionally, the generation of the output video includes inserting an animated object image corresponding to the frame image into the frame image.

[0072] Figure 13 and Figure 14 shows a frame of an example personalized video generated based on a video template according to some exemplary embodiments. Figure 13 shows a captured personalized video 1305 with an actor, in which recoloring has been performed. Figure 13 Further shows a personalized video 1310 created based on stock video obtained from a third party. In personalized video 1310, user face 1320 is inserted into the stock video. Figure 13 Further shows a personalized video 1315 that is a 2D animation with a user head 1325 added on top of a two-dimensional animation.

[0073] Figure 14 shows a personalized video 1405 that is a 3D animation with a user face 1415 inserted into the 3D animation. Figure 14 Further shows a personalized video 1410 with effects, animated elements 1420, and optionally text added on top of an image of a user face.

[0074] Figure 15FIG. 1500 is a flowchart showing a method 1500 for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure. The method 1500 may be executed by a computing device 105. The method 1500 may start by receiving video configuration data at step 1505. The video configuration data may include a sequence of frame images, a sequence of facial region parameters defining the position of the facial region in the frame images, and a sequence of facial landmark parameters defining the position of the facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. In one exemplary embodiment, the sequence of frame images may be generated based on an animated video or a live-action video of a real person. The sequence of facial landmark parameters may be generated based on a live-action video of a real person representing the face of a facial synchronization actor. The video configuration data may include one or more of the following: a sequence of skin masks defining the skin regions of the bodies of at least one actor represented in the frame images; a sequence of mouth region images, where each mouth region image may correspond to at least one frame image; a sequence of eye parameters defining the position of the iris in the sclera of the facial synchronization actor represented in the frame images; a sequence of head parameters defining the rotation, scale, turn, and other parameters of the head; a sequence of animated object images, where each animated object image corresponds to at least one frame image; etc.

[0075] The method 1500 may continue to receive an image of the source face at step 1510. The method may also include generating an output video at step 1515. Specifically, the generation of the output video may include: modifying the frame images of the sequence of frame images. The frame images may be modified by modifying the image of the source face to obtain another image that represents the source face with a facial expression corresponding to the facial landmark parameters. The image of the source face may be modified based on the facial landmark parameters corresponding to the frame images. The other image may be inserted into the frame image at a position determined by the facial region parameters corresponding to the frame image. In one exemplary embodiment, the generation of the output video may also optionally include one or more of the following steps: determining color data associated with the source face and recoloring the skin regions in the frame images based on the color data; inserting the mouth region corresponding to the frame image into the frame image; generating an image of the eye region based on the eye parameters corresponding to the frame and inserting the image of the eye region into the frame image; determining a hair mask based on the source face image and generating a hair image based on the hair mask and the head parameters corresponding to the frame image, and inserting the hair image into the frame image; and inserting the animated object image corresponding to the frame image into the frame image.

[0076] Figure 16 FIG. 1600 shows an example computing system 1600 that may be used to implement the methods described herein. The computing system 1600 may be implemented in a similar environment to the computing devices 105 and 110, the messenger service system 130, the messenger 220, and the system 250 for generating a personalized video based on a template.

[0077] As Figure 16 shown, the hardware components of computing system 1600 can include one or more processors 1610 and a memory 1620. The memory 1620 stores in part instructions and data for execution by the processor 1610. The memory 1620 can store executable code when the system 1600 is running. The system 1600 can also include an optional mass storage device 1630, an optional portable storage media drive 1640, one or more optional output devices 1650, one or more optional input devices 1660, an optional network interface 1670, and one or more optional peripheral devices 1680. The computing system 1600 can also include one or more software components 1695 (e.g., software components that can implement the methods for generating personalized videos based on templates as described herein).

[0078] Figure 16 The components shown are depicted as being connected via a single bus 1690. The components can be connected via one or more data transfer devices or data networks. The processor 1610 and the memory 1620 can be connected via a local microprocessor bus, and the mass storage device 1630, the peripheral device 1680, the portable storage device 1640, and the network interface 1670 can be connected via one or more input / output (I / O) buses.

[0079] The mass storage device 1630, which can be implemented using a disk drive, a solid state disk drive, or an optical disk drive, is a non-volatile storage device for storing data and instructions for use by the processor 1610. The mass storage device 1630 can store system software (e.g., software component 1695) for implementing the embodiments described herein.

[0080] The portable storage media drive 1640 operates in conjunction with a portable non-volatile storage medium (such as a compact disk (CD) or a digital video disk (DVD)) to input data and code into the computing system 1600 and output data and code from the computing system 1600. System software (e.g., software component 1695) for implementing the embodiments described herein can be stored on such a portable medium and input into the computing system 1600 via the portable storage media drive 1640.

[0081] The optional input device 1660 provides a part of the user interface. The input device 1660 can include an alphanumeric keyboard (such as a keyboard) for inputting alphanumeric and other information or a pointing device (such as a mouse, a trackball, a stylus, or cursor direction keys). The input device 1660 can also include a camera or a scanner. In addition, Figure 16The system 1600 shown includes an optional output device 1650. Suitable output devices include speakers, printers, network interfaces, and monitors.

[0082] The network interface 1670 can be used to communicate with external devices, external computing devices, servers, and networking systems via one or more communication networks, such as one or more wired networks, wireless networks, or optical networks, including, for example, the Internet, intranet, local area network (LAN), wide area network (WAN), cellular telephone network, Bluetooth radio, and IEEE 802.11-based radio frequency networks, etc. The network interface 1670 can be a network interface card (such as an Ethernet card, optical transceiver, radio frequency transceiver) or any other type of device capable of sending and receiving information. The optional peripheral device 1680 can include any type of computer support device to add additional functionality to the computer system.

[0083] The components included in the computing system 1600 are intended to represent a large class of computer components. Thus, the computing system 1600 can be a server, personal computer, handheld computing device, telephone, mobile computing device, workstation, minicomputer, mainframe computer, network node, or any other computing device. The computing system 1600 can also include different bus configurations, networking platforms, multiprocessor platforms, etc. Various operating systems (OS) can be used, including UNIX, Linux, Windows, Macintosh OS, Palm OS, and other suitable operating systems.

[0084] Some of the above functions can consist of instructions stored on a storage medium (e.g., computer-readable medium or processor-readable medium). The instructions can be retrieved and executed by a processor. Some examples of storage media are storage devices, magnetic tapes, magnetic disks, etc. The instructions are operable when executed by the processor to direct the processor to operate in accordance with the present invention. Those skilled in the art are familiar with instructions, processors, and storage media.

[0085] It should be noted that any hardware platform suitable for performing the processes described herein is suitable for the present invention. The terms "computer-readable storage medium" and "computer-readable storage medium" as used herein refer to any medium that participates in providing instructions to a processor for execution. Such a medium can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media, for example, include optical discs or magnetic disks (such as fixed disks). Volatile media include dynamic memory (such as system random access memory (RAM)).

[0086] The transmission media include coaxial cables, copper wires, optical fibers, etc., and the transmission media include wires of an embodiment including a bus. The transmission media can also be in the form of acoustic waves or light waves (such as those generated during radio frequency (RF) and infrared (IR) data communications). Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD read-only memory (ROM) disks, DVDs, any other optical media, any other physical media with a pattern of marks or holes, RAM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), any other memory chip or cartridge, carrier waves, or any other medium from which a computer can read.

[0087] Various forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution. The bus carries data to the system RAM, and the processor retrieves and executes the instructions from the system RAM. The instructions received by the system processor can optionally be stored on a fixed disk before or after being executed by the processor.

[0088] Thus, methods and systems for generating personalized videos based on templates have been described. Although the embodiments have been described with reference to specific exemplary embodiments, it is apparent that various modifications and changes can be made to these exemplary embodiments without departing from the broader spirit and scope of the present application. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive.

Claims

1. A method for generating a personalized video based on a template, the method comprises: receiving, by a computing device, video configuration data, the video configuration data comprising: a sequence of frame images representing at least one body; a sequence of facial region parameters that define the position of a facial region in the frame images; and a sequence of skin masks that define the position of a skin region of a part of the at least one body in the frame images; receiving, by the computing device, an image of a source face; determining, by the computing device and based on the image of the source face, color data associated with the source face; and for a frame image in the sequence of frame images: recoloring, by the computing device and based on the color data, the skin region of the part of the at least one body in the frame image; and inserting, by the computing device, the image of the source face at a position determined by the facial region parameters corresponding to the frame image into the frame image to generate an output frame of an output video.

2. The method according to claim 1, wherein, the skin mask in the sequence of skin masks defines the position of a skin region of one of the following: the left hand of the at least one body, the neck of the at least one body, and the right hand of the at least one body.

3. The method according to claim 1, further comprising, before determining the color data associated with the source face, removing one or more of the following: shadows in the image of the source face and highlights in the image of the source face.

4. The method according to claim 1, further comprises: removing at least a part from the image of the source face before determining the color data associated with the source face.

5. The method according to claim 4, wherein, the at least a part comprises one of the following: the area of the eyes, the area of the mouth, the area of the hair, and glasses.

6. The method according to claim 1, wherein: determining the color data associated with the source face comprises determining a color distribution associated with the source face; and recoloring the skin region comprises modifying the values of pixels in the skin region based on the color distribution.

7. The method according to claim 6, wherein, modifying the values of pixels in the skin region to minimize the difference between the color distribution associated with the source face and the distribution of the modified values of the pixels in the skin region.

8. The method according to claim 1, wherein, the sequence of skin masks is generated based on a live action video representing at least one actor.

9. The method according to claim 1, wherein, the sequence of frame images is generated based on an animated video represented as one of the following: a two-dimensional animation of another body and a three-dimensional animation of another body.

10. The method according to claim 1, further comprising, before inserting the image of the source face into the frame image: Receive a sequence of facial landmark parameters that define the positions of facial landmarks in the frame images, where each of the sequence of facial landmark parameters corresponds to a facial expression; and Modify an image of the source face to adopt the facial expression based on the facial landmark parameters corresponding to the frame image.

11. A system for generating a personalized video based on a template, the system including at least one processor and a memory storing processor-executable code, wherein, the at least one processor is configured to perform the following operations when executing the processor-executable code: Receive video configuration data, the video configuration data including: A sequence of frame images representing at least one body; A sequence of facial region parameters that define the positions of facial regions in the frame images; and A sequence of skin masks that define the positions of skin regions of a part of the at least one body in the frame images; Receive an image of a source face; Determine color data associated with the source face based on the image of the source face; and For a frame image of the sequence of frame images: Recolor the skin region of the part of the at least one body in the frame image based on the color data; and Insert the image of the source face at a position determined by the facial region parameters corresponding to the frame image into the frame image to generate an output frame of an output video.

12. The system according to claim 11, wherein, The skin masks of the sequence of skin masks define the positions of skin regions of one of the following: the left hand of the at least one body, the neck of the at least one body, and the right hand of the at least one body.

13. The system according to claim 11, wherein, The at least one processor is configured to: before determining the color data associated with the source face, remove one or more of the following: shadows in the image of the source face and highlights in the image of the source face.

14. The system according to claim 11, wherein, The at least one processor is configured to: before determining the color data associated with the source face, remove at least a part from the image of the source face.

15. The system according to claim 14, wherein, The at least a part includes one of the following: the area of the eyes, the area of the mouth, the area of the hair, and glasses.

16. The system according to claim 11, wherein: Determining the color data associated with the source face includes determining the color distribution associated with the source face; and Recoloring the skin region includes modifying the values of the pixels in the skin region based on the color distribution.

17. The system according to claim 16, wherein, Modify the values of the pixels in the skin region to minimize the difference between the color distribution associated with the source face and the distribution of the modified values of the pixels in the skin region.

18. The system according to claim 11, wherein, The sequence of skin masks is generated based on a live-action video representing at least one actor.

19. The system according to claim 11, wherein, a sequence of the frame images is generated based on an animated video characterized as one of the following: a two-dimensional animation of another body and a three-dimensional animation of the other body.

20. A non-transitory processor-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to implement a method for generating a personalized video based on a template, the method comprising: receiving video configuration data, the video configuration data including: a sequence of frame images representing at least one body; a sequence of facial region parameters that define positions of a facial region in the frame images; and a sequence of skin masks that define positions of skin regions of a part of the at least one body in the frame images; receiving an image of a source face; determining color data associated with the source face based on the image of the source face; and for a frame image of the sequence of frame images: re-coloring the skin regions of the part of the at least one body in the frame image based on the color data; and inserting the image of the source face at a position determined by the facial region parameter corresponding to the frame image into the frame image to generate an output frame of an output video.