System and method for template-based generation of personalized video
By receiving and processing video configuration data and generating personalized videos, the problem of face replacement cannot be realized in the prior art is solved, and real-time personalized video editing on mobile devices is realized.
Patent Information
- Application Number
- CN202510427468.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-23
- Filing Date
- 2020-01-18
- Publication Date
- 2025-07-04
AI Technical Summary
Existing messaging applications cannot implement complex video editing. If one face is replaced with another, it needs to rely on third-party video editing software.
By receiving video configuration data, including frame images, facial areas and landmark parameters, personalized video is generated, and facial expression images are modified and inserted using computing devices to realize facial replacement.
Generate personalized videos in real time on mobile devices, support facial replacement and expression editing, simplifying the video editing process and avoiding dependence on third-party software.
Smart Images

Figure CN120259497A_ABST
Abstract
Description
[0001] This application is a divisional application, and the application number of its parent application is 202080009459.1, the application date is January 18, 2020, and the invention title is "System and Method for Generating Personalized Videos Based on Templates". Technical Field
[0002] The present disclosure generally relates to digital image processing. More specifically, the present disclosure relates to methods and systems for generating personalized videos based on templates. Background Art
[0003] Sharing media such as stickers and emojis has become a standard option in messaging applications (also referred to herein as messengers). Currently, some messengers provide users with options to generate images and short videos and send the images and short videos to other users via a communication chat. Certain existing messengers allow users to modify short videos before transmission. However, the modification of short videos provided by existing messengers is limited to visual effects, filters, and text. Users of current messengers cannot perform complex edits (e.g., replacing one face with another). Such video editing cannot be provided by current messengers and requires complex third-party video editing software. Summary of the Invention
[0004] The purpose of this section is to introduce selected concepts in a simplified form, and the specific content of these concepts is described in the detailed implementation section below. The summary of the invention is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to assist in determining the scope of the claimed subject matter.
[0005] According to an embodiment of the present disclosure, a system for generating personalized videos based on templates is disclosed. The system may include at least one processor and a memory storing processor-executable code. The at least one processor may be configured to receive video configuration data by a computing device. The video configuration data may include: a sequence of frame images, a sequence of facial region parameters defining the position of a facial region in the frame images, and a sequence of facial landmark parameters defining the position of facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The at least one processor may be configured to receive an image of a source face by a computer device. The at least one processor may be configured to generate an output video by the computing device. The generation of the output video may include modifying the frame images of the sequence of frame images. Specifically, the image of the source face may be modified based on the facial landmark parameters corresponding to the frame images to obtain additional images, and the additional images represent the source face with the facial expression corresponding to the facial landmark parameters. The additional images may be inserted into the frame images at positions determined by the facial region parameters corresponding to the frame images.
[0006] According to an exemplary embodiment of the present disclosure, a method for generating a personalized video based on a template is disclosed. The method may begin with receiving video configuration data through a computing device. The video configuration data may include: a sequence of frame images, a sequence of facial region parameters defining the position of a facial region in the frame images, and a sequence of facial landmark parameters defining the position of facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The method may continue with receiving an image of a source face by a computer device. The method may further include generating an output video by the computing device. The generation of the output video may include modifying the frame images of the sequence of frame images. Specifically, the image of the source face may be modified to obtain another image that represents the source face adopting the facial expression corresponding to the facial landmark parameter. The modification of the image may be performed based on the facial landmark parameter corresponding to the frame image. The another image may be inserted into the frame image at a position determined by the facial region parameter corresponding to the frame image.
[0007] According to another aspect of the present disclosure, a non-transitory processor-readable medium is provided that stores processor-readable instructions. When the processor-readable instructions are executed by a processor, they cause the processor to implement the above-described method for generating a personalized video based on a template.
[0008] Additional objects, advantages, and novel features of the examples will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following description and the drawings, or may be learned by practice of the examples. The objectives and advantages of the concepts may be realized and attained by means of the instrumentalities, methods, and combinations particularly pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Embodiments are illustrated in the drawings by way of example and not limitation, in which like reference numerals indicate similar elements.
[0010] Figure 1 is a block diagram showing an exemplary environment in which a system and method for generating a personalized video based on a template may be implemented.
[0011] Figure 2 is a block diagram showing an exemplary embodiment of a computing device for implementing a method for generating a personalized video based on a template.
[0012] Figure 3 is a flowchart showing a process for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure.
[0013] Figure 4 is a flowchart showing the functions of a system for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure.
[0014] Figure 5 is a flowchart showing a process for generating a live-action human video for generating a video template according to some exemplary embodiments.
[0015] Figure 6 Frames of an example live-action human video for generating a video template according to some exemplary embodiments are shown.
[0016] Figure 7 An original image of a face and an image of the face with normalized illumination according to an exemplary embodiment are shown.
[0017] Figure 8 A segmented head image, a head image with facial landmarks, and a facial mask according to an exemplary embodiment are shown.
[0018] Figure 9 Frames representing a user's face, a skin mask, and the result of recoloring the skin mask according to an exemplary embodiment are shown.
[0019] Figure 10 An image of a face of a facial-synchronized actor, an image of facial landmarks of the facial-synchronized actor, an image of facial landmarks of the user, and an image of the user's face with the facial expression of the facial-synchronized actor according to an exemplary embodiment are shown.
[0020] Figure 11 A segmented facial image, a hair mask, the hair mask warped to a target image, and the hair mask applied to the target image according to an exemplary embodiment are shown.
[0021] Figure 12 An original image of an eye, an image with a reconstructed eye sclera, an image with a reconstructed iris, and an image with a moving reconstructed iris according to an exemplary embodiment are shown.
[0022] Figure 13 and Figure 14 Frames of an example personalized video generated based on a video template according to some exemplary embodiments are shown.
[0023] Figure 15 is a flowchart showing a method for generating a personalized video based on a template according to an exemplary embodiment of the present disclosure.
[0024] Figure 16 An example computer system that can be used to implement a method for generating a personalized video based on a template is shown. Detailed Description
[0025] The following detailed description of the embodiments includes reference to the accompanying drawings that form a part of the detailed description. The methods described in this section are not prior art to the claims, and are not admitted to be prior art by inclusion in this section. The drawings illustrate the description according to exemplary embodiments. The exemplary embodiments, also referred to herein as "examples," are described in sufficient detail herein to enable those skilled in the art to practice the subject matter. Embodiments may be combined, other embodiments may be utilized, or structural, logical, and operational changes may be made without departing from the scope claimed. Accordingly, the following detailed description should not be considered limiting, and the scope is defined by the appended claims and their equivalents.
[0026] For the purposes of this patent document, unless otherwise stated or clearly meant otherwise in the context in which it is used, the terms "or" and "and" shall mean "and / or". Unless otherwise stated or where the use of "one or more" is clearly inappropriate, the term "one" shall mean "one or more". The terms "comprise", "comprising", "include" and "including" are interchangeable and are not intended to be limiting. For example, the term "including" shall be construed to mean "including but not limited to".
[0027] This disclosure relates to methods and systems for generating personalized videos based on templates. The embodiments provided by this disclosure solve at least some problems of the prior art. This disclosure can be designed to work in real time on mobile devices such as smart phones, tablets or telephones, but the embodiments can be extended to methods involving network services or cloud-based resources. The methods described herein can be implemented by software running on a computer system and / or by hardware using a combination of microprocessors or other specially designed application specific integrated circuits (ASICs), programmable logic devices, or any combination thereof. Specifically, the methods described herein can be implemented by a series of computer-executable instructions residing on a non-transitory storage medium (such as a disk drive or computer-readable medium).
[0028] Some embodiments of the present disclosure may allow for the real-time generation of personalized videos on a user computing device such as a smart phone. The personalized videos may be generated in the form of audiovisual media (e.g., video, animation, or any other type of media) that depicts the face of one user or the faces of multiple users. The personalized videos may be generated based on pre-generated video templates. The video templates may include video configuration data. The video configuration data may include a sequence of frame images, a sequence of facial region parameters that define the position of the facial region in the frame images, and a sequence of facial landmark parameters that define the position of the facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The frame images may be generated based on an animated video or a live-action video of a person. The facial landmark parameters may be generated based on another live-action video that depicts the face of an actor (also referred to as facesync as described in more detail below), an animated video, an audio file, text, or manually.
[0029] The video configuration file may also include a sequence of skin masks. The skin masks may define the skin region of the body of the actor depicted in the frame images or the skin region of a 2D / 3D animation of the body. In one exemplary embodiment, the skin masks and the facial landmark parameters may be generated based on two different live-action videos of different actors (referred to herein as the actor and the facesync actor, respectively). The video configuration data may also include a sequence of mouth region images and a sequence of eye parameters. The eye parameters may define the position of the iris in the sclera of the facesync actor depicted in the frame images. The video configuration data may include head parameters and a sequence of other parameters of the head that define the rotation, turning, position, and scale of the head. When taking an image and looking directly at the camera, the user may keep their head stationary, and thus, the scale and rotation of the head may be adjusted manually. The head parameters may be transferred from a different actor (also referred to herein as the facesync actor). As used herein, the facesync actor is the person whose facial landmark parameters are being used, and the actor is the other person whose body is being used in the video template and whose skin may be recolored, and the user is the person who takes an image of his / her face to generate the personalized video. Thus, in some embodiments, the personalized video includes the user's face modified to have the facial expression of the facesync actor and includes the body of the actor taken from the video template and recolored to match the user's facial color. The video configuration data includes a sequence of animated object images. Optionally, the video configuration data includes a soundtrack and / or voice.
[0030] Pre-generated video templates can be remotely stored in cloud-based computing resources and downloaded by users of computing devices such as smartphones. A user of a computing device can capture an image of a face through the computing device or select an image of a face from the camera roll, from a prepared set of images, or via a network link. In some embodiments, the image can include the face of an animal rather than a human, or can be in the form of a drawing. Based on one of the face-based image and the pre-generated video template, the computing device can also generate a personalized video. The user can send the personalized video to another user of another computing device via a communication chat, share it on social media, download it to the local storage device of the computing device, or upload it to a cloud storage device or a video sharing service.
[0031] According to one embodiment of the present disclosure, an example method for generating a personalized video based on a template can include receiving video configuration data by a computing device. The video configuration data can include a sequence of frame images, a sequence of face region parameters defining the position of a face region in the frame images, and a sequence of face landmark parameters defining the position of face landmarks in the frame images. Each face landmark parameter can correspond to a facial expression of a face-sync actor. The method can continue by the computing device receiving an image of a source face and generating an output video. The generation of the output video can include modifying the frame images of the sequence of frame images. The modification of the frame images can include modifying the image of the source face to obtain another image that represents the source face with a facial expression corresponding to the face landmark parameter, and inserting the other image into the frame image at a position determined by the face region parameter corresponding to the frame image. Additionally, for example, the source face can be modified by changing the color, making the eyes larger, etc. The image of the source face can be modified based on the face landmark parameter corresponding to the frame image.
[0032] Referring now to the drawings, exemplary embodiments will be described. The drawings are schematic diagrams of idealized exemplary embodiments. Accordingly, the exemplary embodiments discussed herein should not be construed as limited to the specific illustrations presented herein; rather, as will be apparent to those skilled in the art, these exemplary embodiments can include departures from and differences from the illustrations presented herein.
[0033] Figure 1Figure 100 shows an example environment in which a system and method for generating personalized videos based on templates can be implemented. Environment 100 may include computing device 105, user 102, computing device 110, user 104, network 120, and messenger service system 130. Computing device 105 and computing device 110 may refer to mobile devices such as mobile phones, smartphones, or tablets. In other embodiments, computing device 110 may refer to a personal computer, laptop, netbook, set-top box, television device, multimedia device, personal digital assistant, gaming console, entertainment system, infotainment system, in-vehicle computer, or any other computing device.
[0034] Computing device 105 and computing device 110 may be communicatively connected to messenger service system 130 via network 120. Messenger service system 130 may be implemented as cloud-based computing resources. Messenger service system 130 may include computing resources (hardware and software) that are available at a remote location and accessible via a network (e.g., the Internet). Cloud-based computing resources may be shared by multiple users and may be dynamically reallocated based on demand. Cloud-based computing resources may include one or more server farms / clusters that include a collection of computer servers that may be co-located with a network switch and / or router.
[0035] Network 120 may include any wired network, wireless network, or optical network (e.g., including the Internet, intranet, local area network (LAN), personal area network (PAN), wide area network (WAN), virtual private network (VPN), cellular phone network (e.g., Global System for Mobile Communications (GSM)), etc.).
[0036] In some embodiments of the present disclosure, computing device 105 may be configured to initiate a communication chat between user 102 and user 104 of computing device 110. During the communication chat, user 102 and user 104 may exchange text messages and videos. The videos may include personalized videos. Personalized videos may be generated based on pre-generated video templates stored in computing device 105 or computing device 110. In some embodiments, the pre-generated video templates may be stored in messenger service system 130 and downloaded to computing device 105 or computing device 110 on demand.
[0037] Messenger service system 130 may include system 140 for preprocessing videos. System 140 may generate video templates based on animated videos or live-action videos of real people. Messenger service system 130 may include video template database 145 for storing video templates. The video templates may be downloaded to computing device 105 or computing device 110.
[0038] The messenger service system 130 may also be configured to store user profiles 135. The user profiles 135 may include images of the faces of user 102, user 104, and the faces of other persons. Images of the faces may be downloaded to the computing device 105 or the computing device 110 on demand and based on a license. Additionally, an image of the face of user 102 may be generated using the computing device 105 and stored in the local memory of the computing device 105. An image of the face may be generated based on other images stored in the computing device 105. The computing device 105 may also use the image of the face to generate a personalized video based on a pre-generated video template. Similarly, the computing device 110 may be used to generate an image of the face of user 104. The image of the face of user 104 may be used to generate a personalized video on the computing device 110. In other embodiments, the image of the face of user 102 and the image of the face of user 104 may be used interchangeably to generate a personalized video on the computing device 105 or the computing device 110.
[0039] Figure 2 is a block diagram showing an exemplary embodiment of the computing device 105 (computing device 110) for implementing a method for generating a personalized video. In Figure 2 the example shown, the computing device 110 includes both hardware components and software components. Specifically, the computing device 110 includes a camera 205 or any other image capture device or scanner for acquiring digital images. The computing device 110 may also include a processor module 210 and a storage module 215 for storing software components and processor-readable (machine-readable) instructions or code that, when executed by the processor module 210, cause the computing device 105 to perform at least some of the steps of the method for generating a personalized video based on a template as described herein. The computing device 105 may include a graphics display system 230 and a communication module 240. In other embodiments, the computing device 105 may include additional or different components. Additionally, the computing device 105 may include fewer components that perform functions similar or equivalent to those depicted in Figure 2 this.
[0040] The computing device 110 may also include a messenger 220 for initiating a communication chat with another computing device (such as the computing device 110) and a system 250 for generating a personalized video based on a template. The system 250 will be described in more detail below with reference to Figure 4 this. The messenger 220 and the system 250 may be implemented as software components and processor-readable (machine-readable) instructions or code stored in the memory storage device 215 that, when executed by the processor module 210, cause the computing device 105 to perform at least some of the steps of the methods for providing a communication chat and generating a personalized video as described herein.
[0041] In some embodiments, the system 250 for generating personalized videos based on templates may be integrated in the messenger 220. The user interface of the messenger 220 and the system 250 for template-based personalized videos may be provided via the graphics display system 230. A communication chat may be initiated via the communication module 240 and the network 120. The communication module 240 may include a GSM module, a WIFI module, a Bluetooth TM module, etc.
[0042] Figure 3 is a flowchart showing the steps of a process 300 for generating personalized videos based on templates according to some exemplary embodiments of the present disclosure. The process 300 may include production 305, post-production 310, resource preparation 315, skin recoloring 320, lip sync and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340. The resource preparation 315 may be performed by the system 140 for preprocessing videos in the messenger service system 130 (shown in Figure 1 ). The result of the resource preparation 315 is the generation of a video template that may include video configuration data.
[0043] Skin recoloring 320, lip sync and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340 may be performed by the system 250 for generating personalized videos based on templates in the computing device 105 (shown in Figure 2 ). The system 250 may receive an image of the user's face and video configuration data and generate a personalized video representing the user's face.
[0044] Skin recoloring 320, lip sync and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340 may be performed by the system 140 for preprocessing videos in the messenger service system 130 (shown in Figure 1 ). The system 140 may receive a test image of the user's face and a video configuration file. The system 140 may generate a test personalized video representing the user's face. The operator may inspect the test personalized video. Based on the result of the inspection, the video configuration file may be stored in the video template database 145 and then downloaded to the computing device 105 or the computing device 110.
[0045] Production 305 may include idea and scene creation, pre-production (during which locations, props, actors, costumes, and effects are identified), and the production itself, which may require one or more recording sessions. In some exemplary embodiments, recording may be performed by recording scenes / actors against a chroma key background (also referred to herein as a green screen or chroma key screen). To allow for subsequent head tracking and resource cleanup, actors may wear a chroma key face mask with tracking markers (e.g., balaclavas) that cover the actor's face but leave the neck and the bottom of the chin exposed. In Figure 5 Idea and scene creation are shown in detail in
[0046] In one exemplary embodiment, pre-production and the subsequent production steps 305 are optional. Instead of recording actors, two-dimensional or three-dimensional animations may be created or third-party footage / images may be used. Additionally, the original background of the user's image may be used.
[0047] Figure 5 FIG. is a block diagram showing a process 500 for generating a live-action human video. The live-action human video may also be used to generate video templates for generating personalized videos. Process 500 may include generating an idea at step 505 and creating a scene at step 510. Process 500 may continue with pre-production at step 515, followed by production 305. Production 305 may include recording using a chroma key screen 525 or at a real-life location 530.
[0048] Figure 6 Frames of an example live-action human video for generating a video template are shown. Frames of videos 605 and 615 are recorded at a real-life location 530. Frames of videos 610, 620, and 625 are recorded using a chroma key screen 525. Actors may wear a chroma key face mask 630 that has tracking markers covering the actor's face.
[0049] Post-production 310 may include video editing or animation, visual effects, cleanup, sound design, and voice recording.
[0050] During resource preparation 315, the resources to be further deployed may include the following components: background shots without the head of the actor (i.e., the cleaned-up background ready to remove the head of the actor); shots of the actor on a black background (only for the recorded personalized video); the foreground sequence of the frames; example shots with a generic head and soundtrack; the coordinates of the head position, rotation, and scale; animated elements attached to the head (optional); soundtracks with and without narration; narration in a separate file (optional), etc. All these components are optional and can be presented in different formats. The number and configuration of the components depend on the format of the personalized video. For example, for a customized personalized video, no narration is required, and if the original background from the user's picture is used, etc., no background shots and head coordinates are required. In an exemplary embodiment, instead of preparing a file with coordinates, the area where the face needs to be located may be indicated (e.g., manually).
[0051] Skin recoloring 320 allows the color of the skin of the actor in the personalized video to be matched to the color of the face on the user's image. To implement this step, a skin mask may be prepared that specifically indicates which part of the background must be recolored. Preferably, there is a separate mask for each body part of the actor (neck, left hand, right hand, etc.).
[0052] Skin recoloring 320 may include facial image illumination normalization. Figure 7 Figure 705 shows the original image of the face and the image 710 of the face with normalized illumination according to an exemplary embodiment. Shadows or highlights caused by uneven illumination affect the color distribution and may result in the skin color being too dark or too bright after recoloring. To avoid this, shadows and highlights in the user's face may be detected and removed. The facial image illumination normalization process includes the following steps. A deep convolutional neural network may be used to transform the image of the user's face. The network may receive the original image 705 in the form of a portrait image taken under arbitrary illumination and, while keeping the subject in the original image 705 the same, change the illumination of the original image 705 to make the original image 705 have uniform illumination. Therefore, the input of the facial image illumination normalization process includes the original image 705 in the form of an image of the user's face and facial landmarks. The output of the facial image illumination normalization process includes the image 710 of the face with normalized illumination.
[0053] Skin recoloring 320 may include mask creation and body statistics. There may be only a mask for the entire skin or separate masks for body parts. Additionally, different masks may be created for different scenes in the video (e.g., due to significant lighting changes). The masks may be created semi-automatically with some human guidance using techniques such as keying. The prepared masks may be merged into the video resource and then used in the recoloring. Also, to avoid unnecessary computations in real time, color statistics may be pre-computed for each mask. The statistics may include the mean, median, standard deviation, and some percentiles for each color channel. The statistics may be computed in the Red, Green, and Blue (RGB) color space as well as other color spaces (Hue, Saturation, Value (HSV) color space, CIELAB color space (also known as CIEL*a*b* or abbreviated as the "LAB" color space), etc.). The input to the mask creation process may include a grayscale mask for a body part of an actor with uncovered skin in the form of a video or image sequence. The output of the mask creation process may include the masks compressed and merged into the video and the color statistics for each mask.
[0054] Skin recoloring 320 may also include facial statistics calculation. Figure 8 A segmented head image 805 according to one exemplary embodiment is shown, the segmented head image 805 having facial landmarks 810 and a facial mask 815. Based on the segmentation of the user's head image and facial landmarks, a facial mask 815 of the user may be created. Regions such as the eyes, mouth, hair, or accessories (such as glasses) may not be included in the facial mask 815. The user's segmented head image 805 and facial mask may be used to calculate the statistics of the user's facial skin. Thus, the input to the facial statistics calculation may include the user's segmented head image 805, facial landmarks 810, and face segmentation, and the output of the facial statistics calculation may include the color statistics of the user's facial skin.
[0055] Skin recoloring 320 may further include skin color matching and recoloring. Figure 9Shows a frame 905 representing a user's face, a skin mask 910, and the result 915 of recoloring the skin mask 910 according to an exemplary embodiment. Skin color matching and recoloring can be performed using statistics describing the color distribution in the skin of the actor and the user, and the recoloring of the background frame can be performed in real time on a computing device. For each color channel, distribution matching can be performed and the values of the background pixels can be modified so that the distribution of the transformed values is close to the distribution of the face values. Distribution matching can be performed assuming that the color distribution is normal, or by applying techniques such as multi-dimensional probability density function transfer. Thus, the inputs to the skin color matching and recoloring process can include the background frame, the actor skin mask of the frame, the actor body skin color statistics for each mask, and the user face skin color statistics, and the output can include the background frame with the skin of all uncovered body parts recolored.
[0056] In some embodiments, to apply skin recoloring 320, several actors with different skin colors can be recorded, and then a version of the personalized video with the skin color closest to that of the user's image can be used.
[0057] In an exemplary embodiment, instead of skin recoloring 320, a predetermined look-up table (LUT) can be used to adjust the color of the face for the lighting of the scene. The LUT can also be used to change the color of the face, for example, to make the face green.
[0058] Lip sync and facial reenactment 325 can produce realistic facial animations. Figure 10 Shows an example process of lip sync and facial reenactment 325. Figure 10 Shows an image 1005 of a face-synced actor's face, an image 1010 of face-synced actor face landmarks, an image 1015 of the user's face landmarks, and an image 1020 of the user's face with the facial expression of the face-synced actor according to an exemplary embodiment. The steps of lip sync and facial reenactment 325 can include recording the face-synced actor and preprocessing the source video / image to obtain the image 1005 of the face-synced actor's face. Then, as shown in the image 1010 of face-synced actor face landmarks, the face landmarks can be extracted. The steps can also include gaze tracking the face-synced actor. In some embodiments, instead of recording the face-synced actor, a pre-prepared animated 2D or 3D face and mouth region model can be used. The animated 2D or 3D face and mouth region model can be generated by machine learning techniques.
[0059] Optionally, fine-tuning of facial landmarks can be performed. In some exemplary embodiments, the fine-tuning of facial landmarks is performed manually. These steps can be performed in the cloud when preparing the video profile. In some exemplary embodiments, these steps can be performed during resource preparation 315. Then, as shown in the image 1015 of the user's facial landmarks, the user's facial landmarks can be extracted. The next step of synchronization and facial reenactment 325 can include animating the target image with the extracted landmarks to obtain an image 1020 of the user's face with the facial expression of the facial synchronized actor. The steps can be performed on a computing device based on an image of the user's face. The animation method is described in detail in U.S. Patent Application No. 16 / 251,472, the disclosure of which is incorporated herein by reference in its entirety. Lip synchronization and facial reenactment 325 can also be enriched with AI-generated head turns.
[0060] In some exemplary embodiments, after the user captures an image, a three-dimensional model of the user's head can be created. In this embodiment, the steps of lip synchronization and facial reenactment 325 can be omitted.
[0061] Hair animation 330 can be performed to animate the user's hair. For example, if the user has hair, the hair can be animated when the user moves or rotates his head. Hair animation 330 is shown in Figure 11 shown. Figure 11 Shown are a segmented facial image 1105, a hair mask 1110, the hair mask 1110 moved to the facial image, and the hair mask 1110 applied to the facial image according to one exemplary embodiment. Hair animation 330 can include one or more of the following steps: classifying the hair type, modifying the appearance of the hair, modifying the hairstyle, making the hair longer, changing the color of the hair, cutting the hair, and animating the hair, etc. As Figure 11 shown, a facial image can be obtained in the form of a segmented facial image 1105. Then, the hair mask 1110 can be applied to the segmented facial image 1105. Image 1115 shows the hair mask 1110 moved to the facial image. Image 1120 shows the hair mask 1110 applied to the facial image. Hair animation 330 is described in detail in U.S. Patent Application No. 16 / 551,756, the disclosure of which is incorporated herein by reference in its entirety.
[0062] Eye animation 335 can make the user's facial expression more realistic. In Figure 12Eye animation 335 is shown in detail. The processing of eye animation 335 may consist of the following steps: reconstruction of the eye region of the user's face, gaze movement step, and blink step. In the reconstruction process of the eye region, the eye region is segmented into the following parts: eyeball, iris, pupil, eyelashes, and eyelids. If some parts of the eye region (e.g., iris or eyelids) are not fully visible, the complete texture of that part can be synthesized. In some embodiments, a 3D deformable model of the eye can be fitted, and the 3D shape of the eye and the texture of the eye can be obtained. Figure 12 An original image 1205 of the eye is shown, an image 1210 with the reconstructed sclera of the eye, and an image 1215 with the reconstructed iris.
[0063] The gaze movement step includes tracking the gaze direction and pupil position in the video of the face-synchronized actor. If the eye movement of the face-synchronized actor is not rich enough, the data can be manually edited. Then, the gaze movement can be transferred to the user's eye region by synthesizing a new eye image with a transformed eye shape and the same iris position as that of the face-synchronized actor. Figure 12 An image 1220 with the reconstructed moving iris is shown.
[0064] During the blink step, the visible part of the user's eyes can be determined by tracking the eyes of the face-synchronized actor. The changed appearance of the eyelids and eyelashes can be generated based on the reconstruction of the eye region.
[0065] If a generative adversarial network (GAN) is used for face reenactment, the steps of eye animation 335 can be performed explicitly (as described above) or implicitly. In the latter case, the neural network can implicitly capture all the necessary information from the images of the user's face and the source video.
[0066] During deployment 340, the user's face can be realistically animated and automatically inserted into the shot template. The files from the previous steps (resource preparation 315, skin recoloring 320, lip sync and face reenactment 325, hair animation 330, and eye animation 335) can be used as the data of the configuration file. An example of a personalized video with a predetermined set of user faces can be generated for initial inspection. After eliminating the problems identified during the inspection, the personalized video can be deployed.
[0067] The configuration file may also include components that allow text parameters indicating customized personalized videos. A customized personalized video is a personalized video that allows a user to add any text the user desires on top of the final video. The generation of personalized videos with customized text messages is described in more detail in U.S. Patent Application No. 16 / 661,122, filed on October 23, 2019, entitled "SYSTEMS AND METHOD FOR GENERATING PERSONALIZED VIDEOS WITH CUSTOMIZED TEXT MESSAGES", the disclosure of which is incorporated herein by reference in its entirety.
[0068] In one exemplary embodiment, the generation of a personalized video may also include steps of generating a distinct head turn of the user's head; body animation that changes clothing; facial enhancements (such as hair style changes, beautification, adding accessories, etc.); changing scene lighting; synthesizing a voice that reads / sings the text input by the user or converting speech to a voice that matches the user's voice; gender switching; constructing a background and foreground based on user input; etc.
[0069] Figure 4It is a schematic diagram showing the functions 400 of a system 250 for generating personalized videos based on templates. The system 250 can receive an image of a source face shown as a user face image 405 and a video template including video configuration data 410. The video configuration data 410 can include a data sequence 420. For example, the video configuration data 410 can include: a sequence of frame images, a sequence of facial region parameters defining the position of the facial region in the frame images, and a sequence of facial landmark parameters defining the position of the facial landmarks in the frame images. Each facial landmark parameter can correspond to a facial expression. The sequence of frame images can be generated based on an animated video or a live-action video of a real person. The sequence of facial landmark parameters can be generated based on a live-action video of a real person representing the face of a facial synchronization actor. The video configuration data 410 can also include a skin mask, eye parameters, mouth region images, head parameters, animated object images, preset text parameters, etc. The video configuration data can include a sequence of skin masks defining the skin regions of the bodies of at least one actor represented in the frame images. In one exemplary embodiment, the video configuration data 410 can also include a sequence of mouth region images. Each mouth region image can correspond to at least one frame image. In another exemplary embodiment, the video configuration data 410 can include a sequence of eye parameters defining the position of the iris in the sclera of the face of the facial synchronization actor represented in the frame images and / or a sequence of head parameters defining the rotation, turning, scale, and other parameters of the head. In another exemplary embodiment, the video configuration data 410 can also include a sequence of animated object images. Each animated object image can correspond to at least one frame image. The video configuration data 410 can also include a background music 450.
[0070] The system 250 can determine user data 435 based on the user face image 405. The user data can include the user's facial landmarks, user face mask, user color data, user hair mask, etc.
[0071] System 250 can generate frames 445 of an output video that is presented as personalized video 440 based on user data 435 and data sequence 420. System 250 can also add a soundtrack to personalized video 440. Personalized video 440 can be generated by modifying the frame images of a sequence of frame images. The modification of the frame images can include: modifying user face image 405 to obtain an additional image that represents a source face with a facial expression corresponding to facial landmark parameters. The modification can be performed based on the facial landmark parameters corresponding to the frame image. The additional image can be inserted into the frame image at a location determined by facial region parameters corresponding to the frame image. In one exemplary embodiment, the generation of the output video can further include determining color data associated with the source face and recoloring skin regions in the frame image based on the color data. Additionally, the generation of the output video includes inserting a mouth region corresponding to the frame image into the frame image. Other steps of generating the output video can include generating an image of an eye region based on eye parameters corresponding to the frame and inserting the image of the eye region into the frame image. In one exemplary embodiment, the generation of the output video can further include determining a hair mask based on the source face image, generating a hair image based on the hair mask and head parameters corresponding to the frame image, and inserting the hair image into the frame image. Additionally, the generation of the output video includes inserting an animated object image corresponding to the frame image into the frame image.
[0072] Figure 13 and Figure 14 shows a frame of an example personalized video generated based on a video template according to some exemplary embodiments. Figure 13 shows a captured personalized video 1305 with an actor, in which recoloring has been performed. Figure 13 Further shows a personalized video 1310 created based on stock video obtained from a third party. In personalized video 1310, user face 1320 is inserted into the stock video. Figure 13 Further shows a personalized video 1315 that is a 2D animation with a user head 1325 added on top of a two-dimensional animation.
[0073] Figure 14 shows a personalized video 1405 that is a 3D animation with a user face 1415 inserted into the 3D animation. Figure 14 Further shows a personalized video 1410 with effects, animated elements 1420, and optionally text added on top of an image of a user face.
[0074] Figure 15is a flowchart showing a method 1500 for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure. The method 1500 may be executed by a computing device 105. The method 1500 may start by receiving video configuration data at step 1505. The video configuration data may include a sequence of frame images, a sequence of facial region parameters defining the positions of facial regions in the frame images, and a sequence of facial landmark parameters defining the positions of facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. In one exemplary embodiment, the sequence of frame images may be generated based on an animated video or a live-action video of a real person. The sequence of facial landmark parameters may be generated based on a live-action video of a real person representing the face of a facial synchronization actor. The video configuration data may include one or more of the following: a sequence of skin masks defining skin regions of the bodies of at least one actor represented in the frame images; a sequence of mouth region images, where each mouth region image may correspond to at least one frame image; a sequence of eye parameters defining the positions of irises in the scleras of the facial synchronization actors represented in the frame images; a sequence of head parameters defining the rotation, scale, orientation, and other parameters of the head; a sequence of animated object images, where each animated object image corresponds to at least one frame image; etc.
[0075] The method 1500 may continue to receive an image of a source face at step 1510. The method may also include generating an output video at step 1515. Specifically, the generation of the output video may include: modifying the frame images of the sequence of frame images. The frame images may be modified by modifying the image of the source face to obtain another image that represents the source face with a facial expression corresponding to the facial landmark parameters. The image of the source face may be modified based on the facial landmark parameters corresponding to the frame images. The another image may be inserted into the frame image at a position determined by the facial region parameters corresponding to the frame image. In one exemplary embodiment, the generation of the output video may also optionally include one or more of the following steps: determining color data associated with the source face and recoloring the skin regions in the frame images based on the color data; inserting the mouth region corresponding to the frame image into the frame image; generating an image of an eye region based on the eye parameters corresponding to the frame and inserting the image of the eye region into the frame image; determining a hair mask based on the source face image and generating a hair image based on the hair mask and the head parameters corresponding to the frame image and inserting the hair image into the frame image; and inserting the animated object image corresponding to the frame image into the frame image.
[0076] Figure 16 Shows an example computing system 1600 that may be used to implement the methods described herein. The computing system 1600 may be implemented in a similar environment to the computing devices 105 and 110, the messenger service system 130, the messenger 220, and the system 250 for generating a personalized video based on a template.
[0077] As Figure 16 shown, the hardware components of computing system 1600 may include one or more processors 1610 and a memory 1620. The memory 1620 stores, in part, instructions and data for execution by the processor 1610. The memory 1620 may store executable code while the system 1600 is running. The system 1600 may also include an optional mass storage device 1630, an optional portable storage media drive 1640, one or more optional output devices 1650, one or more optional input devices 1660, an optional network interface 1670, and one or more optional peripheral devices 1680. Computing system 1600 may also include one or more software components 1695 (e.g., software components that may implement the methods for generating personalized videos based on templates as described herein).
[0078] Figure 16 The components shown are depicted as being connected via a single bus 1690. The components may be connected via one or more data transfer devices or data networks. The processor 1610 and the memory 1620 may be connected via a local microprocessor bus, and the mass storage device 1630, the peripheral device 1680, the portable storage device 1640, and the network interface 1670 may be connected via one or more input / output (I / O) buses.
[0079] The mass storage device 1630, which may be implemented with a disk drive, a solid state disk drive, or an optical disk drive, is a non-volatile storage device for storing data and instructions for use by the processor 1610. The mass storage device 1630 may store system software (e.g., software components 1695) for implementing the embodiments described herein.
[0080] The portable storage media drive 1640 operates in conjunction with a portable non-volatile storage medium (such as a compact disk (CD) or a digital video disk (DVD)) to input data and code into the computing system 1600 and to output data and code from the computing system 1600. System software (e.g., software components 1695) for implementing the embodiments described herein may be stored on such a portable medium and input into the computing system 1600 via the portable storage media drive 1640.
[0081] The optional input device 1660 provides a part of the user interface. The input device 1660 may include an alphanumeric keyboard (such as a keyboard) for inputting alphanumeric and other information or a pointing device (such as a mouse, a trackball, a stylus, or cursor direction keys). The input device 1660 may also include a camera or a scanner. In addition, Figure 16The illustrated system 1600 includes an optional output device 1650. Suitable output devices include speakers, printers, network interfaces, and monitors.
[0082] The network interface 1670 can be used to communicate with external devices, external computing devices, servers, and networked systems via one or more communication networks, such as one or more wired networks, wireless networks, or optical networks, including, for example, the Internet, intranet, local area network (LAN), wide area network (WAN), cellular telephone network, Bluetooth radio, and IEEE 802.11-based radio frequency network, etc. The network interface 1670 can be a network interface card (such as an Ethernet card, optical transceiver, radio frequency transceiver) or any other type of device capable of sending and receiving information. The optional peripheral device 1680 can include any type of computer support device to add additional functionality to the computer system.
[0083] The components included in the computing system 1600 are intended to represent a large class of computer components. Thus, the computing system 1600 can be a server, personal computer, handheld computing device, telephone, mobile computing device, workstation, minicomputer, mainframe computer, network node, or any other computing device. The computing system 1600 can also include different bus configurations, networking platforms, multiprocessor platforms, etc. Various operating systems (OS) can be used, including UNIX, Linux, Windows, Macintosh OS, Palm OS, and other suitable operating systems.
[0084] Some of the above functions can consist of instructions stored on a storage medium (e.g., computer-readable medium or processor-readable medium). The instructions can be retrieved and executed by the processor. Some examples of storage media are storage devices, magnetic tapes, magnetic disks, etc. The instructions are operable when executed by the processor to direct the processor to operate in accordance with the present invention. Those skilled in the art are familiar with instructions, processors, and storage media.
[0085] It should be noted that any hardware platform suitable for performing the processes described herein is suitable for the present invention. The terms "computer-readable storage medium" and "computer-readable storage medium" as used herein refer to any medium that participates in providing instructions to a processor for execution. Such a medium can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media, for example, include optical discs or magnetic disks (such as fixed disks). Volatile media include dynamic memory (such as system random access memory (RAM)).
[0086] The transmission medium includes coaxial cables, copper wires, optical fibers, etc. The transmission medium includes a wire of one embodiment that includes a bus. The transmission medium can also be in the form of acoustic waves or light waves (such as those generated during radio frequency (RF) and infrared (IR) data communications). Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD read-only memory (ROM) disks, DVDs, any other optical media, any other physical media with a pattern of marks or holes, RAM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), any other memory chip or cartridge, carrier waves, or any other medium from which a computer can read.
[0087] Various forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution. The bus carries data to the system RAM, and the processor retrieves and executes instructions from the system RAM. The instructions received by the system processor can optionally be stored on a fixed disk before or after being executed by the processor.
[0088] Accordingly, methods and systems for generating personalized videos based on templates have been described. Although the embodiments have been described with reference to specific exemplary embodiments, it is apparent that various modifications and changes can be made to these exemplary embodiments without departing from the broader spirit and scope of the present application. Therefore, the specification and the drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method, comprising: Receiving, by a computing device: A sequence of frame images; Facial region parameters corresponding to the position of a facial region in the frame images of the sequence of frame images; And Facial landmark parameters corresponding to the frame images of the sequence of frame images, wherein the facial landmark parameters do not exist in the frame images; Receiving, by the computing device, an image of a source face; Modifying, by the computing device, the image of the source face based on the facial landmark parameters corresponding to the frame images to obtain another facial image, the another facial image representing the source face with a facial expression associated with the facial landmark parameters; and Inserting, by the computing device, the another facial image at a position determined by the facial region parameters associated with the facial image into the frame images, thereby generating an output frame of an output video.
2. The method according to claim 1, wherein The facial landmark parameters are generated based on a video representing a person.
3. The method according to claim 1, wherein The facial landmark parameters are generated based on user input.
4. The method according to claim 1, wherein, The facial landmark parameters are generated based on an animated video.
5. The method according to claim 1, wherein The facial landmark parameters are generated based on an audio file.
6. The method according to claim 1, wherein The facial landmark parameters are generated based on text.
7. The method according to claim 1, further comprising: Receiving, by the computing device, head parameters associated with the size of the facial region in the frame images of the sequence of frame images; And Modifying, by the computing device, the image of the source face based on the head parameters to fit the size of the facial region.
8. The method according to claim 1, further comprising: Receiving, by the computing device, head parameters associated with the rotation of a head to be inserted in the frame images of the sequence of frame images; And Modifying, by the computing device, the image of the source face based on the head parameters to adopt the rotation of the head.
9. The method according to claim 1, wherein The frame images include animals.
10. The method according to claim 1, wherein The frame images include drawn pictures.
11. A computing device, comprising: A processor; And A memory storing instructions that, when executed by the processor, configure the computing device to: Receive: A sequence of frame images; Facial region parameters corresponding to the position of a facial region in the frame images of the sequence of frame images; And Facial landmark parameters corresponding to the frame images of the sequence of frame images, wherein the facial landmark parameters do not exist in the frame images; Receive an image of a source face; Modify the image of the source face based on the facial landmark parameters corresponding to the frame images to obtain another facial image, the another facial image representing the source face with a facial expression associated with the facial landmark parameters; and Insert the another facial image at a position determined by the facial region parameters associated with the facial image into the frame images, thereby generating an output frame of an output video.
12. The computing device according to claim 11, wherein, The facial landmark parameters are generated based on a video representing a person.
13. The computing device according to claim 11, wherein, The facial landmark parameters are generated based on user input.
14. The computing device according to claim 11, wherein, The facial landmark parameters are generated based on an animated video.
15. The computing device according to claim 11, wherein, The facial landmark parameters are generated based on an audio file.
16. The computing device according to claim 11, wherein, The facial landmark parameters are generated based on text.
17. The computing device according to claim 11, wherein, The instructions further configure the computing device to: Receive, by the computing device, a head parameter associated with a size of a facial region in a frame image of a sequence of the frame images; and Modify, based on the head parameter, an image of the source face to fit the size of the facial region.
18. The computing device according to claim 11, wherein, The instructions further configure the computing device to: Receive, by the computing device, a head parameter associated with a rotation of a head to be inserted in a frame image of a sequence of the frame images; And Modify, by the computing device, the image of the source face based on the head parameter to adopt the rotation of the head.
19. The computing device according to claim 11, wherein, The frame images include animals.
20. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that, when executed by a computing device, cause the computing device to: Receive: A sequence of frame images; Facial region parameters corresponding to the position of the facial region in the frame images of the sequence of frame images; And Facial landmark parameters of the frame images corresponding to the sequence of the frame images, the facial landmark parameters not being present in the frame images; Receive an image of a source face; Modify, based on the facial landmark parameters corresponding to the frame images, the image of the source face to obtain another facial image, the another facial image characterizing the source face adopting a facial expression associated with the facial landmark parameters; and Insert the another facial image at a position determined by a facial region parameter associated with the facial image into the frame image, thereby generating an output frame of an output video.