System and method for template-based generation of personalized video

By configuring the processor system in the messenger application and receiving video configuration data and source facial images, complex video editing, such as facial replacement, solves the problem that existing messenger applications cannot achieve such editing and simplifies the editing process.

CN120070678APending Publication Date: 2025-05-30SNAP INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510427473.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-10-23
Filing Date
2020-01-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing messenger applications cannot implement complex video editing, such as replacing one face with another, requiring the use of complex third-party video editing software.

Method used

Through a processor configuration system, video configuration data and source facial images are received to generate output video. The system modifies the sequence of frame images, modifies the source facial image based on the facial landmark parameters, and inserts the modified image at the determined location.

Benefits of technology

This enables complex video editing, such as face replacement in messenger applications, and simplifies the video editing process without the use of third-party software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070678A_ABST
    Figure CN120070678A_ABST
Patent Text Reader

Abstract

Systems and methods for generating a personalized video based on a template are disclosed. An example method may begin with receiving video configuration data that includes a sequence of frame images, a sequence of face region parameters that define a location of a face region in the frame images, and a sequence of face landmark parameters that define a location of a face landmark in the frame images. The method may continue to receive an image of a source face. The method may also include generating an output video. Generation of the output video may include modifying a frame image of a sequence of frame images. Specifically, an image of a source face may be modified to obtain an additional image characterizing the source face that employs a facial expression corresponding to a facial landmark parameter. Additional images may be inserted into the frame image at positions determined by facial region parameters corresponding to the frame image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application, and the application number of its parent application is 202080009459.1, the application date is January 18, 2020, and the invention title is "Systems and Methods for Generating Personalized Videos Based on Templates". Technical Field

[0002] The present disclosure generally relates to digital image processing. More specifically, the present disclosure relates to methods and systems for generating personalized videos based on templates. Background Art

[0003] Sharing media such as stickers and emojis has become a standard option in messaging applications (also referred to herein as messengers). Currently, some messengers provide users with the option to generate images and short videos and send the images and short videos to other users via a communication chat. Certain existing messengers allow users to modify short videos before transmission. However, the modification of short videos provided by existing messengers is limited to visual effects, filters, and text. Users of current messengers cannot perform complex edits (e.g., replacing one face with another). Such video editing cannot be provided by current messengers and requires complex third-party video editing software. Summary of the Invention

[0004] The purpose of this section is to introduce selected concepts in a simplified form, and the specific content of these concepts is described in the detailed implementation section below. The summary of the invention is not intended to identify the key features or main features of the claimed subject matter, nor is it intended to assist in determining the scope of the claimed subject matter.

[0005] According to an embodiment of the present disclosure, a system for generating personalized videos based on templates is disclosed. The system may include at least one processor and a memory storing processor-executable code. The at least one processor may be configured to receive video configuration data by a computing device. The video configuration data may include: a sequence of frame images, a sequence of facial region parameters defining the position of a facial region in the frame images, and a sequence of facial landmark parameters defining the position of facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The at least one processor may be configured to receive an image of a source face by a computer device. The at least one processor may be configured to generate an output video by the computing device. The generation of the output video may include modifying the frame images of the sequence of frame images. Specifically, the image of the source face may be modified based on the facial landmark parameters corresponding to the frame images to obtain additional images, and the additional images represent the source face with the facial expression corresponding to the facial landmark parameters. The additional images may be inserted into the frame images at positions determined by the facial region parameters corresponding to the frame images.

[0006] According to an exemplary embodiment of the present disclosure, a method for generating a personalized video based on a template is disclosed. The method may start with receiving video configuration data via a computing device. The video configuration data may include: a sequence of frame images, a sequence of facial region parameters defining the position of a facial region in the frame images, and a sequence of facial landmark parameters defining the position of facial landmarks in the frame images. Each facial landmark parameter may correspond to a facial expression. The method may continue with receiving an image of a source face by a computer device. The method may further include generating an output video by the computing device. The generation of the output video may include modifying the frame images of the sequence of frame images. Specifically, the image of the source face may be modified to obtain another image that depicts the source face adopting the facial expression corresponding to the facial landmark parameter. The modification of the image may be performed based on the facial landmark parameter corresponding to the frame image. The another image may be inserted into the frame image at a position determined by the facial region parameter corresponding to the frame image.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory processor-readable medium storing processor-readable instructions. When the processor-readable instructions are executed by a processor, they cause the processor to implement the above-described method for generating a personalized video based on a template.

[0008] Additional objectives, advantages, and novel features of the examples will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following description and the drawings, or may be learned by production or operation of the examples. The objectives and advantages of the concepts may be realized and attained by means of the methods, instrumentalities and combinations particularly pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Embodiments are illustrated in the drawings by way of example and not limitation, in which like reference numerals indicate similar elements.

[0010] Figure 1 is a block diagram showing an exemplary environment in which a system and method for generating a personalized video based on a template may be implemented.

[0011] Figure 2 is a block diagram showing an exemplary embodiment of a computing device for implementing a method for generating a personalized video based on a template.

[0012] Figure 3 is a flowchart showing a process for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure.

[0013] Figure 4 is a flowchart showing the functions of a system for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure.

[0014] Figure 5 is a flowchart showing a process for generating a live-action human video for generating a video template according to some exemplary embodiments.

[0015] Figure 6 shows a frame of an example live-action human video for generating a video template according to some exemplary embodiments.

[0016] Figure 7 shows an original image of a face and an image of the face with normalized illumination according to one exemplary embodiment.

[0017] Figure 8 shows a segmented head image, a head image with facial landmarks, and a facial mask according to one exemplary embodiment.

[0018] Figure 9 shows a frame representing a user's face, a skin mask, and the result of recoloring the skin mask according to one exemplary embodiment.

[0019] Figure 10 shows an image of a face of a facial synchronization actor, an image of facial landmarks of the facial synchronization actor, an image of facial landmarks of the user, and an image of the user's face with the facial expression of the facial synchronization actor according to one exemplary embodiment.

[0020] Figure 11 shows a segmented facial image, a hair mask, a hair mask warped to a target image, and a hair mask applied to the target image according to one exemplary embodiment.

[0021] Figure 12 shows an original image of an eye, an image with a reconstructed eye sclera, an image with a reconstructed iris, and an image with a moving reconstructed iris according to one exemplary embodiment.

[0022] Figure 13 and Figure 14 shows frames of an example personalized video generated based on a video template according to some exemplary embodiments.

[0023] Figure 15 is a flowchart showing a method for generating a personalized video based on a template according to an exemplary embodiment of the present disclosure.

[0024] Figure 16 shows an example computer system that can be used to implement a method for generating a personalized video based on a template. DETAILED DESCRIPTION

[0025] The following detailed description of the embodiments includes reference to the accompanying drawings that form a part of the detailed description. The methods described in this section are not prior art to the claims, and are not admitted to be prior art by inclusion in this section. The drawings illustrate the description according to exemplary embodiments. The exemplary embodiments, also referred to herein as "examples", are described in sufficient detail herein to enable those skilled in the art to practice the subject matter. Embodiments may be combined, other embodiments may be utilized, or structural, logical, and operational changes may be made without departing from the scope claimed. Accordingly, the following detailed description should not be considered limiting, and the scope is defined by the appended claims and their equivalents.

[0026] For the purposes of this patent document, unless otherwise stated or clearly meant otherwise in the context in which it is used, the terms "or" and "and" shall mean "and / or". Unless otherwise stated or where the use of "one or more" is clearly inappropriate, the term "one" shall mean "one or more". The terms "comprise", "comprising", "include" and "including" are interchangeable and are not intended to be limiting. For example, the term "including" shall be interpreted to mean "including but not limited to".

[0027] This disclosure relates to methods and systems for generating personalized videos based on templates. The embodiments provided by this disclosure solve at least some problems of the prior art. This disclosure can be designed to work in real time on mobile devices such as smart phones, tablets or telephones, but the embodiments can be extended to methods involving network services or cloud-based resources. The methods described herein can be implemented by software running on a computer system and / or by hardware using a combination of microprocessors or other specially designed application specific integrated circuits (ASICs), programmable logic devices or any combination thereof. Specifically, the methods described herein can be implemented by a series of computer-executable instructions residing on a non-transitory storage medium (such as a disk drive or computer-readable medium).

[0028] Some embodiments of the present disclosure may allow for the real-time generation of personalized videos on a user computing device such as a smart phone. The personalized videos may be generated in the form of audiovisual media (e.g., video, animation, or any other type of media) that depicts the face of one user or the faces of multiple users. The personalized videos may be generated based on pre-generated video templates. The video templates may include video configuration data. The video configuration data may include a sequence of frame images, a sequence of facial region parameters that define the position of a facial region within the frame images, and a sequence of facial landmark parameters that define the position of facial landmarks within the frame images. Each facial landmark parameter may correspond to a facial expression. The frame images may be generated based on an animated video or a live-action video of a person. The facial landmark parameters may be generated based on another live-action video that depicts the face of an actor (also referred to as facesync as described in more detail below), an animated video, an audio file, text, or manually.

[0029] The video configuration file may further include a sequence of skin masks. The skin masks may define the skin region of the body of the actor depicted in the frame images or the skin region of a 2D / 3D animation of the body. In one exemplary embodiment, the skin masks and the facial landmark parameters may be generated based on two different live-action videos of different actors (referred to herein as the actor and the facesync actor, respectively). The video configuration data may further include a sequence of mouth region images and a sequence of eye parameters. The eye parameters may define the position of the iris within the sclera of the facesync actor depicted in the frame images. The video configuration data may include head parameters and a sequence of other parameters of the head that define the rotation, yaw, position, and scale of the head. When taking an image and looking directly at the camera, the user may keep their head stationary, and thus, the scale and rotation of the head may be adjusted manually. The head parameters may be transferred from a different actor (also referred to herein as the facesync actor). As used herein, the facesync actor is the person whose facial landmark parameters are being used, and the actor is the other person whose body is being used in the video template and whose skin may be recolored, and the user is the person who takes an image of his / her face to generate the personalized video. Thus, in some embodiments, the personalized video includes the user's face that is modified to have the facial expression of the facesync actor and includes the body of the actor taken from the video template and recolored to match the user's facial color. The video configuration data includes a sequence of animated object images. Optionally, the video configuration data includes a soundtrack and / or voice.

[0030] Pre-generated video templates can be remotely stored in cloud-based computing resources and downloaded by users of computing devices such as smartphones. A user of a computing device can capture an image of a face through the computing device or select an image of a face from a camera roll, from a prepared set of images, or via a network link. In some embodiments, the image can include the face of an animal rather than a human, or can be in the form of a drawing. Based on one of the face-based image and the pre-generated video template, the computing device can also generate a personalized video. The user can send the personalized video to another user of another computing device via a communication chat, share it on social media, download it to a local storage device of the computing device, or upload it to a cloud storage device or a video sharing service.

[0031] According to one embodiment of the present disclosure, an example method for generating a personalized video based on a template can include receiving, by a computing device, video configuration data. The video configuration data can include a sequence of frame images, a sequence of face region parameters defining the position of a face region in the frame images, and a sequence of face landmark parameters defining the position of face landmarks in the frame images. Each face landmark parameter can correspond to a facial expression of a face synchronization actor. The method can continue by the computing device receiving an image of a source face and generating an output video. The generation of the output video can include modifying the frame images of the sequence of frame images. The modification of the frame images can include modifying the image of the source face to obtain an additional image that represents the source face with a facial expression corresponding to the face landmark parameter and inserting the additional image into the frame image at a position determined by the face region parameter corresponding to the frame image. Additionally, for example, the source face can be modified by changing the color, making the eyes larger, etc. The image of the source face can be modified based on the face landmark parameter corresponding to the frame image.

[0032] Referring now to the drawings, exemplary embodiments are described. The drawings are schematic diagrams of idealized exemplary embodiments. Accordingly, the exemplary embodiments discussed herein should not be construed as limited to the specific illustrations presented herein; rather, as will be apparent to those skilled in the art, these exemplary embodiments can include departures from and differences from the illustrations presented herein.

[0033] Figure 1An example environment 100 is shown, in which a system and method for generating personalized videos based on templates can be implemented. Environment 100 may include a computing device 105, a user 102, a computing device 110, a user 104, a network 120, and a messenger service system 130. The computing device 105 and the computing device 110 may refer to mobile devices such as mobile phones, smartphones, or tablets. In other embodiments, the computing device 110 may refer to a personal computer, laptop, netbook, set-top box, television device, multimedia device, personal digital assistant, gaming console, entertainment system, infotainment system, in-vehicle computer, or any other computing device.

[0034] The computing device 105 and the computing device 110 may be communicatively connected to the messenger service system 130 via the network 120. The messenger service system 130 may be implemented as cloud-based computing resources. The messenger service system 130 may include computing resources (hardware and software) available at a remote location and accessible via a network (e.g., the Internet). The cloud-based computing resources may be shared by multiple users and may be dynamically reallocated based on demand. The cloud-based computing resources may include one or more server farms / clusters that include a collection of computer servers that may be co-located with a network switch and / or router.

[0035] The network 120 may include any wired network, wireless network, or optical network (e.g., including the Internet, intranet, local area network (LAN), personal area network (PAN), wide area network (WAN), virtual private network (VPN), cellular phone network (e.g., Global System for Mobile Communications (GSM)), etc.).

[0036] In some embodiments of the present disclosure, the computing device 105 may be configured to initiate a communication chat between the user 102 and the user 104 of the computing device 110. During the communication chat, the user 102 and the user 104 may exchange text messages and videos. The videos may include personalized videos. The personalized videos may be generated based on pre-generated video templates stored in the computing device 105 or the computing device 110. In some embodiments, the pre-generated video templates may be stored in the messenger service system 130 and downloaded to the computing device 105 or the computing device 110 on demand.

[0037] The messenger service system 130 may include a system 140 for preprocessing videos. The system 140 may generate video templates based on animated videos or live-action videos of real people. The messenger service system 130 may include a video template database 145 for storing video templates. The video templates may be downloaded to the computing device 105 or the computing device 110.

[0038] The messenger service system 130 may also be configured to store user profiles 135. The user profiles 135 may include images of the faces of user 102, user 104, and the faces of other persons. The images of the faces may be downloaded to the computing device 105 or the computing device 110 on demand and based on a license. Additionally, an image of the face of user 102 may be generated using the computing device 105 and stored in the local memory of the computing device 105. The image of the face may be generated based on other images stored in the computing device 105. The computing device 105 may also use the image of the face to generate a personalized video based on a pre-generated video template. Similarly, the computing device 110 may be used to generate an image of the face of user 104. The image of the face of user 104 may be used to generate a personalized video on the computing device 110. In other embodiments, the image of the face of user 102 and the image of the face of user 104 may be used interchangeably to generate a personalized video on the computing device 105 or the computing device 110.

[0039] Figure 2 is a block diagram showing an exemplary embodiment of a computing device 105 (computing device 110) for implementing a method of generating a personalized video. In Figure 2 the example shown, the computing device 110 includes both hardware components and software components. Specifically, the computing device 110 includes a camera 205 or any other image capture device or scanner for acquiring digital images. The computing device 110 may also include a processor module 210 and a storage module 215 for storing software components and processor-readable (machine-readable) instructions or code that, when executed by the processor module 210, cause the computing device 105 to perform at least some of the steps of the method for generating a personalized video based on a template as described herein. The computing device 105 may include a graphics display system 230 and a communication module 240. In other embodiments, the computing device 105 may include additional or different components. Additionally, the computing device 105 may include fewer components that perform functions similar or equivalent to those depicted in Figure 2 herein.

[0040] The computing device 110 may also include a messenger 220 for initiating a communication chat with another computing device (such as the computing device 110) and a system 250 for generating a personalized video based on a template. The system 250 will be described in more detail below with reference to Figure 4 herein. The messenger 220 and the system 250 may be implemented as software components and processor-readable (machine-readable) instructions or code stored in the memory storage device 215 that, when executed by the processor module 210, cause the computing device 105 to perform at least some of the steps of the methods for providing a communication chat and generating a personalized video as described herein.

[0041] In some embodiments, the system 250 for generating personalized videos based on templates may be integrated in the messenger 220. The user interface of the messenger 220 and the system 250 for template-based personalized videos may be provided via the graphics display system 230. A communication chat may be initiated via the communication module 240 and the network 120. The communication module 240 may include a GSM module, a WIFI module, a Bluetooth TM module, etc.

[0042] Figure 3 FIG. is a flowchart showing the steps of a process 300 for generating personalized videos based on templates according to some exemplary embodiments of the present disclosure. The process 300 may include production 305, post-production 310, resource preparation 315, skin recoloring 320, lip sync and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340. The resource preparation 315 may be performed by the system 140 for preprocessing videos in the messenger service system 130 (shown in Figure 1 ). The result of the resource preparation 315 is the generation of a video template that may include video configuration data.

[0043] Skin recoloring 320, lip sync and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340 may be performed by the system 250 for generating personalized videos based on templates in the computing device 105 (shown in Figure 2 ). The system 250 may receive an image of the user's face and video configuration data and generate a personalized video representing the user's face.

[0044] Skin recoloring 320, lip sync and facial reenactment 325, hair animation 330, eye animation 335, and deployment 340 may be performed by the system 140 for preprocessing videos in the messenger service system 130 (shown in Figure 1 ). The system 140 may receive a test image of the user's face and a video configuration file. The system 140 may generate a test personalized video representing the user's face. The operator may inspect the test personalized video. Based on the result of the inspection, the video configuration file may be stored in the video template database 145 and then downloaded to the computing device 105 or the computing device 110.

[0045] Production 305 may include idea and scene creation, pre-production (during which locations, props, actors, costumes, and effects are identified), and the production itself, which may require one or more recording sessions. In some exemplary embodiments, recording may be performed by recording scenes / actors against a chroma key background (also referred to herein as a green screen or chroma key screen). To allow for subsequent head tracking and resource cleanup, actors may wear a chroma key face mask with tracking markers (e.g., balaclavas) that cover the actor's face but leave the neck and the bottom of the chin exposed. In Figure 5 Idea and scene creation are shown in detail in

[0046] In one exemplary embodiment, pre-production and the subsequent production steps 305 are optional. Instead of recording actors, 2D or 3D animations may be created or third-party footage / images may be used. Additionally, the original background of the user's image may be used.

[0047] Figure 5 FIG. is a block diagram showing a process 500 for generating a live-action human video. The live-action human video may also be used to generate video templates for generating personalized videos. Process 500 may include generating an idea at step 505 and creating a scene at step 510. Process 500 may continue with pre-production at step 515, followed by production 305. Production 305 may include recording using a chroma key screen 525 or at a real-life location 530.

[0048] Figure 6 FIG. shows frames of an example live-action human video for generating a video template. Frames of videos 605 and 615 are recorded at a real-life location 530. Frames of videos 610, 620, and 625 are recorded using a chroma key screen 525. Actors may wear a chroma key face mask 630 that has tracking markers covering the actor's face.

[0049] Post-production 310 may include video editing or animation, visual effects, cleanup, sound design, and voice recording.

[0050] During resource preparation 315, the resources to be further deployed may include the following components: background shots without the head of the actor (i.e., the cleaned-up background ready to remove the head of the actor); shots of the actor on a black background (only for the recorded personalized video); the foreground sequence of the frames; example shots with a generic head and soundtrack; the coordinates of the head position, rotation, and scale; animated elements attached to the head (optional); soundtracks with and without narration; the narration in a separate file (optional), etc. All these components are optional and can be presented in different formats. The number and configuration of the components depend on the format of the personalized video. For example, for a customized personalized video, no narration is required, and if the original background from the user's picture is used, etc., no background shots and head coordinates are required. In one exemplary embodiment, instead of preparing a file with coordinates, the area where the face needs to be located may be indicated (e.g., manually).

[0051] Skin recoloring 320 allows the color of the skin of the actor in the personalized video to be matched to the color of the face on the user's image. To implement this step, a skin mask specifically indicating which part of the background must be recolored may be prepared. Preferably, there is a separate mask for each body part of the actor (neck, left hand, right hand, etc.).

[0052] Skin recoloring 320 may include facial image illumination normalization. Figure 7 FIG. 705 shows the original image of the face and FIG. 710 shows the image of the face with normalized illumination according to one exemplary embodiment. Shadows or highlights caused by uneven illumination affect the color distribution and may result in the skin color being too dark or too bright after recoloring. To avoid this, the shadows and highlights in the user's face may be detected and removed. The facial image illumination normalization process includes the following steps. A deep convolutional neural network may be used to transform the image of the user's face. The network may receive the original image 705 in the form of a portrait image taken under arbitrary illumination and, while keeping the subject in the original image 705 the same, change the illumination of the original image 705 to make the original image 705 have uniform illumination. Therefore, the input of the facial image illumination normalization process includes the original image 705 in the form of an image of the user's face and the facial landmarks. The output of the facial image illumination normalization process includes the image 710 of the face with normalized illumination.

[0053] Skin recoloring 320 may include mask creation and body statistics. There may be either a mask for the entire skin or separate masks for body parts. Additionally, different masks may be created for different scenes in the video (e.g., due to significant lighting changes). The masks may be created semi-automatically with some human guidance by techniques such as keying. The prepared masks may be incorporated into the video resource and then used in the recoloring. And, to avoid unnecessary computations in real time, color statistics may be pre-computed for each mask. The statistics may include the mean, median, standard deviation, and some percentiles for each color channel. The statistics may be computed in the red, green, and blue (RGB) color space as well as other color spaces (hue, saturation, value (HSV) color space, CIELAB color space (also known as CIEL*a*b* or abbreviated as the "LAB" color space), etc.). The input to the mask creation process may include a grayscale mask for a body part of an actor with uncovered skin in the form of a video or image sequence. The output of the mask creation process may include the masks compressed and incorporated into the video and the color statistics for each mask.

[0054] Skin recoloring 320 may also include facial statistics calculation. Figure 8 A segmented head image 805 according to an exemplary embodiment is shown, the segmented head image 805 having facial landmarks 810 and a facial mask 815. Based on the segmentation of the user's head image and facial landmarks, a facial mask 815 of the user may be created. Regions such as the eyes, mouth, hair, or accessories (such as glasses) may not be included in the facial mask 815. The user's segmented head image 805 and facial mask may be used to calculate the statistics of the user's facial skin. Thus, the input to the facial statistics calculation may include the user's segmented head image 805, facial landmarks 810, and face segmentation, and the output of the facial statistics calculation may include the color statistics of the user's facial skin.

[0055] Skin recoloring 320 may further include skin color matching and recoloring. Figure 9Shows a frame 905 representing a user's face, a skin mask 910, and the result 915 of recoloring the skin mask 910 according to an exemplary embodiment. Skin color matching and recoloring can be performed using statistics describing the color distribution in the skin of the actor and the user's skin, and the recoloring of the background frame can be performed in real time on a computing device. For each color channel, distribution matching can be performed and the values of the background pixels can be modified so that the distribution of the transformed values is close to the distribution of the face values. Distribution matching can be performed assuming that the color distribution is normal, or by applying techniques such as multi-dimensional probability density function transfer. Thus, the input to the skin color matching and recoloring process can include the background frame, the actor skin mask of the frame, the actor body skin color statistics of each mask, and the user face skin color statistics, and the output can include the background frame with the skin of all uncovered body parts recolored.

[0056] In some embodiments, to apply skin recoloring 320, several actors with different skin colors can be recorded, and then a version of the personalized video with the skin color closest to that of the user's image can be used.

[0057] In an exemplary embodiment, instead of skin recoloring 320, a predetermined look-up table (LUT) can be used to adjust the color of the face for the lighting of the scene. The LUT can also be used to change the color of the face, for example, to make the face green.

[0058] Lip sync and facial reenactment 325 can produce realistic facial animations. Figure 10 Shows an example process of lip sync and facial reenactment 325. Figure 10 Shows an image 1005 of a face-synced actor's face, an image 1010 of face-synced actor face landmarks, an image 1015 of the user's face landmarks, and an image 1020 of the user's face with the facial expression of the face-synced actor according to an exemplary embodiment. The steps of lip sync and facial reenactment 325 can include recording the face-synced actor and preprocessing the source video / image to obtain the image 1005 of the face-synced actor's face. Then, as shown in the image 1010 of the face-synced actor face landmarks, the face landmarks can be extracted. The steps can also include gaze tracking the face-synced actor. In some embodiments, instead of recording the face-synced actor, a pre-prepared animated 2D or 3D face and mouth region model can be used. The animated 2D or 3D face and mouth region model can be generated by machine learning techniques.

[0059] Optionally, fine-tuning of facial landmarks can be performed. In some exemplary embodiments, the fine-tuning of facial landmarks is performed manually. These steps can be performed in the cloud when preparing the video profile. In some exemplary embodiments, these steps can be performed during resource preparation 315. Then, as shown in the image 1015 of the user's facial landmarks, the user's facial landmarks can be extracted. The next step of synchronization and facial reenactment 325 can include animating the target image with the extracted landmarks to obtain an image 1020 of the user's face with the facial expression of the facial-synchronized actor. The steps can be performed on a computing device based on an image of the user's face. The animation method is described in detail in U.S. Patent Application No. 16 / 251,472, the disclosure of which is incorporated herein by reference in its entirety. Lip synchronization and facial reenactment 325 can also be enriched with AI-generated head turns.

[0060] In some exemplary embodiments, after the user captures an image, a three-dimensional model of the user's head can be created. In this embodiment, the steps of lip synchronization and facial reenactment 325 can be omitted.

[0061] Hair animation 330 can be performed to animate the user's hair. For example, if the user has hair, the hair can be animated when the user moves or rotates his head. Hair animation 330 is shown in Figure 11 shown. Figure 11 FIG. 11 shows a segmented facial image 1105, a hair mask 1110, the hair mask 1110 moved to the facial image, and the hair mask 1110 applied to the facial image according to an exemplary embodiment. Hair animation 330 can include one or more of the following steps: classifying the hair type, modifying the appearance of the hair, modifying the hairstyle, making the hair longer, changing the color of the hair, cutting the hair, and animating the hair, etc. As Figure 11 shown, a facial image in the form of a segmented facial image 1105 can be obtained. Then, the hair mask 1110 can be applied to the segmented facial image 1105. Image 1115 shows the hair mask 1110 moved to the facial image. Image 1120 shows the hair mask 1110 applied to the facial image. Hair animation 330 is described in detail in U.S. Patent Application No. 16 / 551,756, the disclosure of which is incorporated herein by reference in its entirety.

[0062] Eye animation 335 can make the user's facial expression more realistic. In Figure 12Eye animation 335 is shown in detail. The processing of eye animation 335 may consist of the following steps: reconstruction of the eye region of the user's face, gaze movement step, and blink step. In the reconstruction process of the eye region, the eye region is segmented into the following parts: eyeball, iris, pupil, eyelashes, and eyelids. If some parts of the eye region (e.g., iris or eyelids) are not fully visible, the complete texture of that part can be synthesized. In some embodiments, a 3D deformable model of the eye can be fitted, and the 3D shape of the eye and the texture of the eye can be obtained. Figure 12 Shows the original image 1205 of the eye, the image 1210 with the reconstructed sclera of the eye, and the image 1215 with the reconstructed iris.

[0063] The gaze movement step includes tracking the gaze direction and pupil position in the video of the face-synchronized actor. If the eye movement of the face-synchronized actor is not rich enough, the data can be manually edited. Then, the gaze movement can be transferred to the user's eye region by synthesizing a new eye image with a transformed eye shape and the same iris position as that of the face-synchronized actor. Figure 12 Shows the image 1220 with the reconstructed moving iris.

[0064] During the blink step, the visible part of the user's eye can be determined by tracking the eyes of the face-synchronized actor. The changed appearance of the eyelids and eyelashes can be generated based on the reconstruction of the eye region.

[0065] If facial reenactment is performed using a generative adversarial network (GAN), the steps of eye animation 335 can be performed explicitly (as described above) or implicitly. In the latter case, the neural network can implicitly capture all the necessary information from the images of the user's face and the source video.

[0066] During deployment 340, the user's face can be realistically animated and automatically inserted into the shot template. The files from the previous steps (resource preparation 315, skin recoloring 320, lip sync and facial reenactment 325, hair animation 330, and eye animation 335) can be used as the data for the configuration file. An example of a personalized video with a predetermined set of user faces can be generated for initial inspection. After eliminating the problems identified during the inspection, the personalized video can be deployed.

[0067] The configuration file may also include components that allow text parameters indicating customized personalized videos. A customized personalized video is a personalized video that allows a user to add any text desired by the user on top of the final video. The generation of personalized videos with customized text messages is described in more detail in U.S. Patent Application No. 16 / 661,122, entitled "SYSTEMS AND METHOD FOR GENERATING PERSONALIZED VIDEOS WITH CUSTOMIZED TEXT MESSAGES," filed on October 23, 2019, the disclosure of which is incorporated herein by reference in its entirety.

[0068] In one exemplary embodiment, the generation of a personalized video may further include steps of generating a distinct head turn of the user's head; body animation that changes clothing; facial enhancements (such as a hairstyle change, beautification, adding accessories, etc.); changing scene lighting; synthesizing a voice that reads / sings the text input by the user or converting speech to a voice that matches the user's voice; gender switching; constructing a background and foreground based on user input; etc.

[0069] Figure 4FIG. 400 is a schematic diagram showing the functions of a system 250 for generating a personalized video based on a template according to some exemplary embodiments. The system 250 can receive an image of a source face shown as a user face image 405 and a video template including video configuration data 410. The video configuration data 410 can include a data sequence 420. For example, the video configuration data 410 can include: a sequence of frame images, a sequence of face region parameters defining the position of the face region in the frame images, and a sequence of face landmark parameters defining the position of the face landmarks in the frame images. Each face landmark parameter can correspond to a facial expression. The sequence of frame images can be generated based on an animated video or a live-action video of a real person. The sequence of face landmark parameters can be generated based on a live-action video of a real person representing the face of a synchronized actor. The video configuration data 410 can also include a skin mask, eye parameters, mouth region images, head parameters, animated object images, preset text parameters, etc. The video configuration data can include a sequence of skin masks defining the skin regions of the bodies of at least one actor represented in the frame images. In one exemplary embodiment, the video configuration data 410 can also include a sequence of mouth region images. Each mouth region image can correspond to at least one frame image. In another exemplary embodiment, the video configuration data 410 can include a sequence of eye parameters defining the position of the iris in the sclera of the face of the synchronized actor represented in the frame images and / or a sequence of head parameters defining the rotation, turning, scale, and other parameters of the head. In another exemplary embodiment, the video configuration data 410 can also include a sequence of animated object images. Each animated object image can correspond to at least one frame image. The video configuration data 410 can also include a soundtrack 450.

[0070] The system 250 can determine user data 435 based on the user face image 405. The user data can include the user's face landmarks, user face mask, user color data, user hair mask, etc.

[0071] System 250 can generate frames 445 of an output video that is displayed as a personalized video 440 based on user data 435 and data sequence 420. System 250 can also add a soundtrack to the personalized video 440. The personalized video 440 can be generated by modifying the frame images of a sequence of frame images. The modification of the frame images can include: modifying the user face image 405 to obtain an additional image that represents a source face with a facial expression corresponding to facial landmark parameters. The modification can be performed based on the facial landmark parameters corresponding to the frame image. The additional image can be inserted into the frame image at a position determined by the facial region parameters corresponding to the frame image. In one exemplary embodiment, the generation of the output video can further include determining color data associated with the source face and recoloring the skin regions in the frame image based on the color data. Additionally, the generation of the output video includes inserting a mouth region corresponding to the frame image into the frame image. Other steps in generating the output video can include generating an image of an eye region based on eye parameters corresponding to the frame and inserting the image of the eye region into the frame image. In one exemplary embodiment, the generation of the output video can further include determining a hair mask based on the source face image, generating a hair image based on the hair mask and head parameters corresponding to the frame image, and inserting the hair image into the frame image. Additionally, the generation of the output video includes inserting an animated object image corresponding to the frame image into the frame image.

[0072] Figure 13 and Figure 14 show a frame of an example personalized video generated based on a video template according to some exemplary embodiments. Figure 13 show a captured personalized video 1305 with an actor, in which recoloring has been performed. Figure 13 Further show a personalized video 1310 created based on stock video obtained from a third party. In the personalized video 1310, a user face 1320 is inserted into the stock video. Figure 13 Further show a personalized video 1315 that is a 2D animation with a user head 1325 added on top of a two-dimensional animation.

[0073] Figure 14 show a personalized video 1405 that is a 3D animation with a user face 1415 inserted into the 3D animation. Figure 14 Further show a personalized video 1410 with effects, animated elements 1420, and optionally text added on top of an image of the user face.

[0074] Figure 15is a flowchart showing a method 1500 for generating a personalized video based on a template according to some exemplary embodiments of the present disclosure. The method 1500 can be executed by a computing device 105. The method 1500 can start by receiving video configuration data at step 1505. The video configuration data can include a sequence of frame images, a sequence of facial region parameters defining the positions of facial regions in the frame images, and a sequence of facial landmark parameters defining the positions of facial landmarks in the frame images. Each facial landmark parameter can correspond to a facial expression. In one exemplary embodiment, the sequence of frame images can be generated based on an animated video or a live-action video of a real person. The sequence of facial landmark parameters can be generated based on a live-action video representing the face of a facial synchronization actor. The video configuration data can include one or more of the following: a sequence of skin masks defining skin regions of the bodies of at least one actor represented in the frame images; a sequence of mouth region images, where each mouth region image can correspond to at least one frame image; a sequence of eye parameters defining the positions of irises in the scleras of the facial synchronization actors represented in the frame images; a sequence of head parameters defining the rotation, scale, turn, and other parameters of the head; a sequence of animated object images, where each animated object image corresponds to at least one frame image; etc.

[0075] The method 1500 can continue to receive an image of a source face at step 1510. The method can also include generating an output video at step 1515. Specifically, the generation of the output video can include: modifying the frame images of the sequence of frame images. The frame images can be modified by modifying the image of the source face to obtain another image that represents the source face adopting the facial expression corresponding to the facial landmark parameters. The image of the source face can be modified based on the facial landmark parameters corresponding to the frame images. The another image can be inserted into the frame image at the position determined by the facial region parameters corresponding to the frame image. In one exemplary embodiment, the generation of the output video can also optionally include one or more of the following steps: determining color data associated with the source face and recoloring the skin regions in the frame images based on the color data; inserting the mouth region corresponding to the frame image into the frame image; generating an image of an eye region based on the eye parameters corresponding to the frame and inserting the image of the eye region into the frame image; determining a hair mask based on the source face image and generating a hair image based on the hair mask and the head parameters corresponding to the frame image and inserting the hair image into the frame image; and inserting the animated object image corresponding to the frame image into the frame image.

[0076] Figure 16 Shows an example computing system 1600 that can be used to implement the methods described herein. The computing system 1600 can be implemented in a similar environment to the computing devices 105 and 110, the messenger service system 130, the messenger 220, and the system 250 for generating a personalized video based on a template.

[0077] As Figure 16 shown, the hardware components of computing system 1600 may include one or more processors 1610 and a memory 1620. The memory 1620 stores, in part, instructions and data for execution by the processor 1610. The memory 1620 may store executable code while the system 1600 is running. The system 1600 may also include an optional mass storage device 1630, an optional portable storage media drive 1640, one or more optional output devices 1650, one or more optional input devices 1660, an optional network interface 1670, and one or more optional peripheral devices 1680. The computing system 1600 may also include one or more software components 1695 (e.g., software components that implement the methods for generating personalized videos based on templates as described herein).

[0078] Figure 16 The components shown are depicted as being connected via a single bus 1690. The components may be connected by one or more data transfer devices or data networks. The processor 1610 and the memory 1620 may be connected via a local microprocessor bus, and the mass storage device 1630, the peripheral device 1680, the portable storage device 1640, and the network interface 1670 may be connected via one or more input / output (I / O) buses.

[0079] The mass storage device 1630, which may be implemented with a disk drive, a solid state disk drive, or an optical disk drive, is a non-volatile storage device for storing data and instructions for use by the processor 1610. The mass storage device 1630 may store system software (e.g., software component 1695) for implementing the embodiments described herein.

[0080] The portable storage media drive 1640 operates in conjunction with a portable non-volatile storage medium (such as a compact disk (CD) or a digital video disk (DVD)) to input data and code into the computing system 1600 and to output data and code from the computing system 1600. System software (e.g., software component 1695) for implementing the embodiments described herein may be stored on such a portable medium and input into the computing system 1600 via the portable storage media drive 1640.

[0081] The optional input device 1660 provides a portion of the user interface. The input device 1660 may include an alphanumeric keyboard (such as a keyboard) for inputting alphanumeric and other information or a pointing device (such as a mouse, a trackball, a stylus, or cursor direction keys). The input device 1660 may also include a camera or a scanner. In addition, Figure 16The system 1600 shown includes an optional output device 1650. Suitable output devices include speakers, printers, network interfaces, and monitors.

[0082] The network interface 1670 can be used to communicate with external devices, external computing devices, servers, and networking systems via one or more communication networks, such as one or more wired networks, wireless networks, or optical networks, including, for example, the Internet, intranet, local area network (LAN), wide area network (WAN), cellular telephone network, Bluetooth radio, and IEEE 802.11-based radio frequency networks, etc. The network interface 1670 can be a network interface card (such as an Ethernet card, optical transceiver, radio frequency transceiver) or any other type of device capable of sending and receiving information. The optional peripheral device 1680 can include any type of computer support device to add additional functionality to the computer system.

[0083] The components included in the computing system 1600 are intended to represent a large class of computer components. Thus, the computing system 1600 can be a server, personal computer, handheld computing device, telephone, mobile computing device, workstation, minicomputer, mainframe computer, network node, or any other computing device. The computing system 1600 can also include different bus configurations, networking platforms, multiprocessor platforms, etc. Various operating systems (OS) can be used, including UNIX, Linux, Windows, Macintosh OS, Palm OS, and other suitable operating systems.

[0084] Some of the above functions can consist of instructions stored on a storage medium (e.g., computer-readable medium or processor-readable medium). The instructions can be retrieved and executed by a processor. Some examples of storage media are storage devices, magnetic tapes, magnetic disks, etc. The instructions are operable when executed by the processor to direct the processor to operate in accordance with the present invention. Those skilled in the art are familiar with instructions, processors, and storage media.

[0085] It should be noted that any hardware platform suitable for performing the processes described herein is suitable for the present invention. The terms "computer-readable storage medium" and "computer-readable storage medium" as used herein refer to any medium that participates in providing instructions to a processor for execution. Such a medium can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media, for example, include optical discs or magnetic disks (such as fixed disks). Volatile media include dynamic memory (such as system random access memory (RAM)).

[0086] The transmission media include coaxial cables, copper wires, optical fibers, etc., and the transmission media include wires in one embodiment that include a bus. The transmission media can also be in the form of acoustic or optical waves (such as those generated during radio frequency (RF) and infrared (IR) data communications). Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD read-only memory (ROM) disks, DVDs, any other optical media, any other physical media with patterns of marks or holes, RAM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), any other memory chip or cartridge, carrier waves, or any other medium from which a computer can read.

[0087] Various forms of computer-readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution. The bus carries data to the system RAM, and the processor retrieves and executes the instructions from the system RAM. The instructions received by the system processor can optionally be stored on a fixed disk before or after being executed by the processor.

[0088] Thus, methods and systems for generating personalized videos based on templates have been described. Although the embodiments have been described with reference to specific exemplary embodiments, it is apparent that various modifications and changes can be made to these exemplary embodiments without departing from the broader spirit and scope of the present application. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method, comprising: receiving, by a computing device: a sequence of frame images; facial region parameters corresponding to the positions of facial regions in the frame images of the sequence of frame images; and facial landmark parameters corresponding to the frame images of the sequence of frame images; receiving, by the computing device, an image of a source face; modifying, by the computing device and based on the facial landmark parameters corresponding to the frame images, the image of the source face to obtain another facial image, the another facial image characterizing the source face with a facial expression corresponding to the facial landmark parameters; and inserting, by the computing device, the another facial image at a position determined by the facial region parameters corresponding to the frame images into the frame images to generate an output frame of an output video.

2. The method according to claim 1, wherein the facial landmark parameters are generated based on another frame image different from the frame images in the sequence of frame images.

3. The method according to claim 1, wherein the sequence of frame images is generated based on one of: an animated video and a real-time action video.

4. The method according to claim 1, wherein the facial landmark parameters are generated based on one of: a live action video, an animated video, an audio file, text, a manual indication.

5. The method according to claim 1, further comprising: receiving, by the computing device, a sequence of skin masks defining skin regions of a body represented in the frame images or skin regions of a 2D / 3D animation of another body; and wherein generating the output frame of the output video comprises: determining color data associated with the source face; and re-coloring the skin regions of the body in the frame images or the skin regions of the 2D / 3D animation of the another body based on the color data.

6. The method according to claim 1, further comprising: receiving, by the computing device, a sequence of mouth region images, each of the mouth region images corresponding to at least one of the frame images; and wherein generating the output frame of the output video comprises inserting the mouth region corresponding to the frame image into the frame image.

7. The method according to claim 1, further comprising: receiving, by the computing device, a sequence of eye parameters defining the positions of irises in the scleras of the faces represented in the frame images; and wherein generating the output frame of the output video comprises: generating an image of an eye region based on the eye parameters corresponding to the frame images; and inserting the image of the eye region into the frame image.

8. The method according to claim 1, further comprising: receiving, by the computing device, a sequence of head parameters defining one or more of rotation, rotation, position, and scale of a head, wherein generating the output frame of the output video is based on the sequence of head parameters.

9. The method according to claim 1, wherein generating the output frame of the output video comprises: determining a hair mask based on the image of the source face; generating a hair image based on the hair mask; and inserting the hair image into the frame image.

10. The method according to claim 1, further comprising: Receive a sequence of animated object images by the computing device, where each of the animated object images corresponds to at least one of the frame images; and wherein, generating an output frame of the output video includes inserting an animated object image corresponding to the frame image into the frame image.

11. A computing device, comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the computing device to: Receive: a sequence of frame images; facial region parameters corresponding to the position of the facial region in the frame images of the sequence of frame images; and facial landmark parameters corresponding to the frame images of the sequence of frame images; Receive an image of a source face; Modify the image of the source face based on the facial landmark parameters corresponding to the frame image to obtain another facial image, the another facial image representing the source face with a facial expression corresponding to the facial landmark parameters; and Insert the another facial image at a position determined by the facial region parameters corresponding to the frame image into the frame image to generate an output frame of the output video.

12. The computing device according to claim 11, wherein, the facial landmark parameters are generated based on another frame image different from the frame images in the sequence of frame images.

13. The computing device according to claim 11, wherein, the sequence of frame images is generated based on one of: an animated video and a real-time action video.

14. The computing device according to claim 11, wherein, the facial landmark parameters are generated based on one of: a live action video, an animated video, an audio file, text, a manual indication.

15. The computing device according to claim 11, wherein, the instructions further configure the computing device to receive a sequence of skin masks defining a skin region of a body represented in the frame image or a skin region of a 2D / 3D animation of another body; and wherein, generating an output frame of the output video includes: Determining color data associated with the source face; and Recoloring the skin region of the body in the frame image or the skin region of the 2D / 3D animation of the another body based on the color data.

16. The computing device according to claim 11, wherein, the instructions further configure the computing device to receive a sequence of mouth region images, each of the mouth region images corresponding to at least one of the frame images; and wherein, generating an output frame of the output video includes inserting a mouth region corresponding to the frame image into the frame image.

17. The computing device according to claim 11, wherein, the instructions further cause the computing device to receive a sequence of eye parameters defining the position of an iris in the sclera of a face represented in the frame image; and wherein, generating an output frame of the output video includes: Generating an image of an eye region based on the eye parameters corresponding to the frame image; and Inserting the image of the eye region into the frame image.

18. The computing device according to claim 11, further comprising receiving, by the computing device, a sequence of head parameters that define one or more of rotation, turning, position, and scale of the head, wherein, the output frames of the output video are generated based on the sequence of head parameters.

19. The computing device according to claim 11, wherein, generating the output frames of the output video includes: determining a hair mask based on the image of the source face; generating a hair image based on the hair mask; and inserting the hair image into the frame image.

20. A non-transitory computer-readable storage medium, the computer-readable storage medium comprising instructions that, when executed by a computing device, cause the computing device to: receive: a sequence of frame images; facial region parameters corresponding to the position of the facial region in the frame images of the sequence of frame images ; and facial landmark parameters corresponding to the frame images of the sequence of frame images; receive an image of a source face; modifying the image of the source face based on the facial landmark parameters corresponding to the frame images to obtain another facial image, the another facial image representing the source face with a facial expression corresponding to the facial landmark parameters; and inserting the another facial image at a position determined by the facial region parameters corresponding to the frame images into the frame images to generate output frames of an output video.

Citation Information

Cited By

  • Video generation method and system based on environmental perception, server and medium

    CN120711257A