Method and computing device for providing personalized video
By providing systems and methods in instant messaging software, receiving and modifying images of the source face to generate personalized videos, the problem of complex video editing in the prior art is solved, and the functions of facial replacement and video generation in instant messaging software are realized.
Patent Information
- Application Number
- CN202510118794.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-07
- Filing Date
- 2020-01-18
- Publication Date
- 2025-05-06
AI Technical Summary
Existing instant messaging software cannot perform complex video editing, such as replacing one face with another, requiring the use of third-party video editing software.
A system and method are provided to receive preprocessed videos through a computing device, receive images of the source faces, and generate personalized videos by modifying images of the source faces to adopt facial expressions of the target faces. The system includes a processor and memory, which can store preprocessed videos in the memory of the computing device, and display images through a graphic display system to guide the user to locate the face.
This implements complex video editing in instant messaging software, allowing users to replace one face with another, generate personalized videos, and select and send these videos in communication chat.
Smart Images

Figure CN119942418A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese national phase application with an international application date of January 18, 2020, an international application number of PCT / US2020 / 014223, and an invention name of “System and method for providing personalized video”. The national phase entry date of the Chinese national phase application is July 16, 2021, the application number is 202080009764.0, and the invention name is “System and method for providing personalized video”. Technical Field
[0002] The present invention relates generally to digital image processing and more particularly to methods and systems for providing personalized video. Background Art
[0003] Sharing media such as stickers and emoticons has become a standard option in messaging applications (also referred to herein as instant messaging software (messenger)). Currently, some instant messaging software provides users with the option of generating and sending images and short videos to other users through communication chat. Some existing instant messaging software allows users to modify short videos before transmission. However, the modification of short videos provided by existing instant messaging software is limited to visualization effects, filters and text. Users of current instant messaging software cannot perform complex editing, such as replacing one face with another face. Current instant messaging software does not provide such video editing, and such complex video editing requires the use of third-party video editing software. Summary of the invention
[0004] This section is provided to introduce a simplified form of the technical solution, which will be further described in the detailed description section below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.
[0005] According to one embodiment of the present invention, a system for providing personalized videos is disclosed. The system may include at least one processor and a memory, the memory storing processor executable code. The processor may be configured to store one or more pre-processed videos in the memory of a computing device. The one or more pre-processed videos may include at least one frame having at least one target face. The processor may be configured to receive an image of a source face. The image of the source face may be received as an additional image selected by a user from a set of images stored in the memory. The additional image may be segmented into a portion including the source face and a background. In another exemplary embodiment, the image of the source face may be received by capturing the additional image through a camera of the computing device, and segmenting the additional image into a portion including the source face and a background. Before capturing the additional image, the processor may display the additional image via a graphics display system of the computing device, and guide the user to position the face in the additional image within a predetermined area of the screen.
[0006] The processor may be configured to modify one or more pre-processed videos to generate one or more personalized videos. The modification of the one or more pre-processed videos may be performed by modifying an image of a source face to adopt a facial expression of a target face. Modifying the one or more pre-processed videos may also include replacing at least one target face with an image of the modified source face.
[0007] Before modifying the image of the source face, the processor may determine target facial expression parameters associated with the parameterized facial model based on at least one frame. In this embodiment, modifying the image of the source face may include determining source parameters associated with the parameterized facial model based on the image of the source face. The source parameters may include source facial expression parameters, source facial identity parameters, and source facial texture parameters. Modifying the image of the source face may also include synthesizing the modified image of the source face based on the parameterized facial model and the target facial expression parameters, the source facial identity parameters, and the source facial texture parameters.
[0008] The processor may also be configured to receive additional images from additional sources and modify one or more pre-processed videos based on the additional images to generate one or more additional personalized videos. The processor may also be configured to enable a communication chat between the user and at least one other user of at least one remote computing device, receive a video selected by the user from the one or more personalized videos, and send the selected video to the at least one other user via the communication chat.
[0009] The processor may also be configured to display the selected video in the window of the communication chat. The selected video may be displayed in a collapsed mode. Upon receiving an indication that the user has clicked on the selected video in the communication chat window, the processor may display the selected video in a full screen mode. The processor may also be configured to mute the sound associated with the selected video when displaying the selected video in the collapsed mode, and to replay the sound associated with the selected video when displaying the selected video in the full screen mode.
[0010] According to an exemplary embodiment, a method for providing personalized videos is disclosed. The method includes: receiving a preprocessed video including a target face by a computing device; providing a first user interface by the computing device to enable a user to generate an image of a source face; modifying the preprocessed video by the computing device to generate one or more personalized videos by replacing the target face with the source face, the source face being modified to adopt the facial expression of the target face; and providing a second user interface by the computing device to select one or more personalized videos.
[0011] According to an exemplary embodiment, a computing device is disclosed, including: a processor; and a memory storing instructions, which, when executed by the processor, configure the computing device to: receive a preprocessed video including a target face; provide a first user interface to enable a user to generate an image of a source face; modify the preprocessed video to generate one or more personalized videos by replacing the target face with the source face, the source face being modified to adopt the facial expressions of the target face; and provide a second user interface to select one or more personalized videos.
[0012] According to an exemplary embodiment, a non-transitory computer-readable storage medium is disclosed, the computer-readable storage medium including instructions that, when executed by a computing device, cause the computing device to: receive a pre-processed video including a target face; provide a first user interface that enables a user to generate an image of a source face; modify the pre-processed video to generate one or more personalized videos by replacing the target face with the source face, the source face being modified to adopt the facial expressions of the target face; and provide a second user interface to select one or more personalized videos.
[0013] According to an exemplary embodiment, a method for providing personalized video is disclosed. The method may include storing one or more pre-processed videos by a computing device. The one or more pre-processed videos may include at least one frame having at least one target face. The method may then continue to receive an image of a source face by a computing device. The image of the source face may be received as a user selects an additional image from a set of images stored in a memory of the computing device, and the additional image is segmented into a portion including the source face and a background. In another exemplary embodiment, the image of the source face may be received by capturing the additional image by a camera of the computing device, and the additional image is segmented into a portion including the source face and a background. Before capturing the additional image, the additional image may be displayed via a graphics display system of the computing device, and the user may be guided to position the facial image with the additional image within a predetermined area of the graphics display system.
[0014] The method may also include modifying, by the computing device, one or more pre-processed videos to generate one or more personalized videos. The modification may include modifying an image of a source face to generate an image of a modified source face. The modified source face may adopt a facial expression of a target face. The modification may also include replacing at least one target face with the image of the modified source face. The method may also include receiving, by the computing device, additional images of additional source faces, and modifying, by the computing device, one or more pre-processed videos based on the additional images to generate one or more additional personalized videos.
[0015] The method may also include enabling, by the computing device, a communication chat between a user of the computing device and at least one other user of at least one other computing device, receiving, by the computing device, a video selected by the user from one or more personalized videos, and sending, by the computing device, the selected video to at least one other user via the communication chat. The method may continue to display, by the computing device, the selected video in a window of the communication chat in a collapsed mode. Upon receiving, by the computing device, an indication that the user has clicked on the selected video in the window of the communication chat, the selected video may be displayed in a full-screen mode. The method may include muting the sound associated with the selected video when displaying the selected video in the collapsed mode, and playing back the sound associated with the selected video when displaying the selected video in the full-screen mode.
[0016] The method may also include determining target facial expression parameters associated with the parameterized facial model based on at least one frame before modifying the image of the source face. The at least one frame may include metadata, such as the target facial expression parameters. In this embodiment, modifying the source facial image may include determining source parameters associated with the parameterized facial model based on the source facial image. The source parameters may include source facial expression parameters, source facial identity parameters, and source facial texture parameters. Modifying the source facial image may also include synthesizing the modified image of the source face based on the parameterized facial model and the target facial expression parameters, the source facial identity parameters, and the source facial texture parameters.
[0017] According to another aspect of the present invention, a non-transitory processor-readable medium is provided, storing processor-readable instructions. When a processor executes the processor-readable instructions, the processor-readable instructions enable the processor to implement the above method for providing personalized video.
[0018] Other purposes, advantages and novel features of the examples will be described in part in the following description, and in part will become apparent to those skilled in the art after knowing the following description and the accompanying drawings, or can be learned through the making or operation of the examples. The purposes and advantages of these technical solutions can be achieved and obtained through the methods, means and combinations specifically pointed out in the attached claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The embodiments are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like reference numerals refer to similar elements.
[0020] Figure 1 is a block diagram illustrating an exemplary environment in which systems and methods for providing personalized video may be implemented.
[0021] Figure 2 is a block diagram illustrating an exemplary embodiment of a computing device for implementing a method for providing personalized video.
[0022] Figure 3 is a block diagram illustrating a system for providing personalized video according to some exemplary embodiments of the present invention.
[0023] Figure 4 is a schematic diagram illustrating a process of generating a personalized video according to an exemplary embodiment.
[0024] Figure 5 is a block diagram of a personalized video generation module according to some exemplary embodiments of the present invention.
[0025] Figures 6 to 11 Screens showing a user interface of a system for providing personalized video in instant messaging software according to some exemplary embodiments are shown.
[0026] Fig.12 is a flowchart illustrating a method for providing personalized video according to an exemplary embodiment.
[0027] Fig.13 is a flow chart illustrating a method for sharing a personalized video according to an exemplary embodiment.
[0028] Fig.14 An exemplary computer system is shown, which can be used to implement a method for providing personalized video. DETAILED DESCRIPTION
[0029] The following detailed description of the embodiments includes reference to the accompanying drawings, which form a part of the detailed description. The methods described in this section are not prior art to the claims and are not considered prior art by being included in this section. The accompanying drawings show diagrams according to exemplary embodiments. These exemplary embodiments, also referred to as "examples" herein, are described as detailed as possible to enable those skilled in the art to implement the present technical solution. Without departing from the scope of the claimed protection, the embodiments may be combined, other embodiments may be utilized, or structural, logical and operational changes may be made thereto. Therefore, the following detailed description should not be considered restrictive, and the scope is limited by the attached claims and their equivalents.
[0030] For the purposes of this patent document, the terms "or" and "and" shall mean "and / or" unless otherwise specified or clearly intended in the context of their use. The term "a" shall mean "one or more" unless otherwise specified or the use of "one or more" is clearly inappropriate. The terms "include" and "including" are interchangeable and are not intended to be limiting. For example, the term "include" shall be interpreted as "including but not limited to."
[0031] The present invention relates to methods and systems for providing personalized videos. The embodiments provided by the present invention solve at least some of the problems of the prior art. The present invention can be designed to work in real time on a mobile device such as a smart phone, tablet computer or mobile phone, although the embodiments can be extended to methods involving network services or cloud-based resources. The methods described herein can be implemented by software running on a computer system and / or by hardware utilizing a combination of microprocessors or other specially designed application-specific integrated circuits (ASICs), programmable logic devices or any combination thereof. In particular, the methods described herein can be implemented by a series of computer executable instructions residing on a non-transitory storage medium such as a disk drive or a computer-readable medium.
[0032] Some embodiments of the present invention may allow for the generation of personalized videos in real time on a user computing device such as a smart phone. Personalized videos may be generated based on pre-generated videos, such as videos that feature close-ups of actors. Certain embodiments of the present invention may allow for the replacement of an actor's face in a pre-generated video with the face of a user or another person to generate a personalized video. When replacing the actor's face with the face of the user or another person, the face of the user or another person is modified to adopt the facial expressions of the actor. Personalized videos may be generated in a communication chat between a user and another user of another computing device. A user may select one or more of the personalized videos and send them to another user through the communication chat. Personalized videos may be indexed and searched based on pre-defined keywords associated with a template of a pre-processed video, which are used to insert an image of the user's face to generate a personalized video. Personalized videos may be ranked and classified based on the emotions and actions in the video.
[0033] According to one embodiment of the present invention, an exemplary method for providing personalized videos may include storing one or more pre-processed videos by a computing device. The one or more pre-processed videos may include at least one frame having at least a target face. The method may also include enabling a communication chat between a user and at least one other user of at least one remote computing device. The computing device may receive an image of a source face and modify one or more pre-processed videos to generate one or more personalized videos. The modification may include modifying the image of the source face to generate an image of a modified source face. The modified source face may adopt the facial expression of the target face. The modification may also include replacing at least one target face with an image of a modified source face. During the modification, a video selected by a user from one or more personalized videos may be received, and the selected video may be sent to at least one other user through the communication chat.
[0034] With reference now to the accompanying drawings, exemplary embodiments are described. The accompanying drawings are schematic diagrams of idealized exemplary embodiments. Therefore, the exemplary embodiments discussed herein should not be construed as being limited to the specific illustrations presented herein, but rather, these exemplary embodiments may include deviations and different illustrations than those presented herein, which will be apparent to those skilled in the art.
[0035] Figure 1An exemplary environment 100 is shown in which a method for providing personalized video can be practiced. Environment 100 may include computing device 105, user 102, computing device 110, user 104, network 120, and instant messaging software service system 130. Computing device 105 and computing device 110 may refer to mobile devices such as mobile phones, smart phones, or tablet computers. However, in other embodiments, computing device 105 or computing device 110 may refer to a personal computer, a laptop computer, a netbook, a set-top box, a television device, a multimedia device, a personal digital assistant, a game console, an entertainment system, an infotainment system, a vehicle computer, or any other computing device.
[0036] Computing device 105 and computing device 110 can be communicatively connected to instant messaging software service system 130 via network 120. Instant messaging software service system 130 can be implemented as a cloud-based computing resource. The instant messaging software service system may include computing resources (hardware and software) that are available at a remote location and can be accessed via a network (e.g., the Internet). Cloud-based computing resources can be shared by multiple users and can be dynamically reallocated as needed. Cloud-based computing resources may include one or more server farms / clusters, which include a group of computer servers that can be co-located with network switches and / or routers.
[0037] Network 120 may include any wired, wireless, or optical network, including, for example, the Internet, an intranet, a local area network (LAN), a personal area network (PAN), a wide area network (WAN), a virtual private network (VPN), a cellular telephone network (e.g., a Global System for Mobile (GSM) communication network, etc.).
[0038] In some embodiments of the present invention, computing device 105 may be configured to enable a communication chat between user 102 and user 104 of computing device 110. During the communication chat, user 102 and user 104 may exchange text messages and videos. The video may include a personalized video. The personalized video may be generated based on a pre-generated video stored in computing device 105 or computing device 110. In some embodiments, the pre-generated video may be stored in instant messaging software service system 130 and downloaded to computing device 105 or computing device 110 as needed.
[0039] The instant messaging software service system 130 may also be configured to store user profiles. The user profile may include a facial image of user 102, a facial image of user 104, and facial images of other people. The facial image may be downloaded to computing device 105 or computing device 110 on demand and based on a license. In addition, a facial image of user 102 may be generated using computing device 105 and stored in a local memory of computing device 105. A facial image may be generated based on other images stored in computing device 105. Computing device 105 may further use the facial image to generate a personalized video based on a pre-generated video. Similarly, computing device 110 may be used to generate a facial image of user 104. The facial image of user 104 may be used to generate a personalized video on computing device 110. In a further embodiment, the facial image of user 102 and the facial image of user 104 may be used to generate a personalized video on computing device 105 or computing device 110.
[0040] Figure 2 is a block diagram illustrating an exemplary embodiment of a computing device 105 (or computing device 110) for implementing a method for personalizing a video. Figure 2 In the example shown, computing device 110 includes hardware components and software components. In particular, computing device 110 includes camera 205 or any other image acquisition device or scanner to acquire digital images. Computing device 110 may also include processor module 210 and storage module 215 for storing software components and processor-readable (machine-readable) instructions or codes that, when executed by processor module 210, cause computing device 105 to perform at least some steps of the method for providing personalized video as described herein. Computing device 105 may include graphics display system 230 and communication module 240. In other embodiments, computing device 105 may include additional or different components. In addition, computing device 105 may include executing similar or equivalent Figure 2 Fewer components than those depicted in the .
[0041] The computing device 110 may also include: instant messaging software 220 for implementing communication and chatting with another computing device (such as the computing device 110); and system 300 for providing personalized video. Figure 3 The system 300 is described in more detail. The instant messaging software 220 and the system 300 may be implemented as software components and processor-readable (machine-readable) instructions or codes stored in the memory 215, which, when executed by the processor module 210, cause the computing device 105 to perform at least some steps of the method for providing communication chat and personalized video as described herein.
[0042] In some embodiments, the system 300 for providing personalized video can be integrated into the instant messaging software 220. The user interface of the instant messaging software 220 and the system 300 for providing personalized video can be provided by the graphic display system 230. The communication chat can be realized through the communication module 240 and the network 120. The communication module 240 can include a GSM module, a WiFi module, a Bluetooth module, and a TM Modules, etc.
[0043] Figure 3 3 is a block diagram of a system 300 for providing personalized video according to some exemplary embodiments of the present invention. The system 300 may include a user interface 305, a facial image acquisition module 310, a video database 320, and a personalized video generation module 330.
[0044] The video database 320 may store one or more videos. The video may include previously recorded videos of an actor or actors. The video may include a 2D video or a 3D scene. The video may be pre-processed to segment the actor's face (also referred to as the target face) and the background in each frame, and identify a set of parameters that may be used to further insert a source face instead of the actor's face (target face). The set of parameters may include facial texture, facial expression parameters, facial color, facial identity parameters, the position and angle of the face, etc. The set of parameters may also include a list of manipulations and operations that may be performed on the actor's face, such as replacement of the actor's face in a photo-realistic manner.
[0045] The facial image acquisition module 320 may receive an image of a person and generate an image of the person's face. The image of the person's face may be used as a source face to replace a target face in a video stored in the video database 320. The image of the person may be acquired by the camera 205 of the computing device 105. The image of the person may include an image stored in the memory 215 of the computing device 105. Figure 7 Details of the facial image acquisition module 320 are provided in.
[0046] Based on the image of the source face, the personalized video generation module 330 can generate a personalized video from one or more pre-generated videos stored in the database 320. The module 330 can replace the face of the actor in the pre-generated video with the source face while maintaining the facial expression of the actor's face. The module 330 can replace the facial texture, facial color, and facial identity of the actor with the facial texture, facial color, and facial identity of the source face. The module 330 can also add an image of glasses on the eye area of the source face in the personalized video. Similarly, the module 330 can add an image of headwear (e.g., a beanie, a brimmed hat, a helmet, etc.) on the head of the source face in the personalized video. The images of glasses and headwear can be pre-stored in the user's computing device 105, or the images of glasses and headwear can be generated. The images of glasses and headwear can be generated using a DNN. The module 330 can also apply shadows or colors to the source face in the personalized video. For example, the module 330 can add a tan to the face of the source face.
[0047] Figure 4 4 is a schematic diagram illustrating the functionality 400 of the personalized video generation module 330 according to some exemplary embodiments. The personalized video generation module 330 may receive an image of a source face 405 and a pre-generated video 410. The pre-generated video 410 may include one or more frames 420. The frame 420 may include a target face 415. The facial expression of the source face 405 may be different from the facial expression of the target face 415.
[0048] In some embodiments of the present invention, the personalized video generation module 330 may be configured to analyze the image of the source face 405 to extract the source face parameters 430. The source face parameters 430 may be extracted by fitting a parameterized face model to the image of the source face 405. The parameterized face model may include a template mesh. The coordinates of the vertices in the template mesh may depend on two parameters: facial identity and facial expression. Therefore, the source parameters 430 may include the facial identity and facial expression corresponding to the source face 405. The source parameters 405 may also include the texture of the source face 405. The texture may include the color at the vertices in the template mesh. In some embodiments, a texture model associated with the template mesh may be used to determine the texture of the source face 405.
[0049] In some embodiments of the present invention, the personalized video generation module 330 may be configured to analyze the frames 420 of the target video 410 to extract the target facial parameters 335 of each frame 420. The target facial parameters 435 may be extracted by fitting a parameterized facial model to the target face 415. The target parameters 435 may include a facial identity and facial expression corresponding to the target face 415. The target facial parameters 430 may also include a texture of the target face 415. The texture model may be used to obtain the texture of the target face 415. In some embodiments of the present invention, each frame 420 may include metadata. The metadata may include parameters determined for the frame. For example, the parameters may be provided by the instant messaging software service system 130 (e.g., Figure 1 The parameters may be stored in metadata of a frame of the pre-generated video 410. The pre-generated video may be further downloaded to the computing device 105 and stored in the video database 320. Alternatively, the personalized video generation module 330 may pre-process the pre-generated video 410 to determine target facial parameters 435 and position parameters of the target face 415 in the frame 420. The personalized video generation module 330 may also store the target facial parameters 435 and position parameters of the target face in the metadata of the corresponding frame 420. In this way, the target facial parameters 435 are not recalculated each time the pre-generated video 410 is selected to be personalized using a different source face.
[0050] In some embodiments of the present invention, the personalized video generation module 330 may also be configured to replace the facial expressions in the source facial parameters 430 with the facial expressions from the target parameters 435. The personalized video generation module 330 may be further configured to synthesize the output face 445 using the parameterized facial model, the texture module and the target parameters 430 and the replaced facial expressions. The output face 435 may be used to replace the target face 415 in the frame of the target video 410 to obtain the frame 445 of the output video shown as the personalized video 440. The output face 435 is the source face 405 adopting the facial expressions of the target face 415. The output video is a personalized video 440 generated based on the predetermined video 410 and the image of the source face 405.
[0051] Figure 5 is a block diagram of a personalized video generation module 330 according to an exemplary embodiment. The personalized video generation module 330 may include a parameterized face model 505, a texture model 510, a DNN 515, a preprocessing module 520, a parameter extraction module 525, a face synthesis module 525, and a mouth and eye generation module 530. Modules 505 to 530 may be implemented as software components for use by hardware devices, such as computing device 105, computing device 110, instant messaging software service system 130, etc.
[0052] In some embodiments of the present invention, the parameterized facial model 505 may be pre-generated based on images of a predetermined number of individuals of different ages and genders. For each individual, the images may include an image of the individual with a neutral facial expression and one or more images of the individual with different facial expressions. Facial expressions may include open mouth, smile, anger, surprise, etc.
[0053] The parameterized face model 505 may include a template mesh having a predetermined number of vertices. The template mesh may be represented as a 3D triangulation that defines the shape of the head. Each individual may be associated with an individual-specific blend shape. The individual-specific blend shapes may be adjusted based on the template mesh. The individual-specific blend shapes may correspond to specific coordinates of vertices in the template mesh. Thus, different individual images may correspond to a template mesh of the same structure; however, the coordinates of the vertices in the template mesh may be different for different images.
[0054] In some embodiments of the present invention, the parameterized facial model may include a bilinear facial model that depends on two parameters: facial identity and facial expression. The bilinear facial model may be constructed based on blend shapes corresponding to individual images. Thus, the parameterized facial model includes a template mesh of a predetermined structure, in which the coordinates of the vertices depend on the facial identity and the facial expression.
[0055] In some embodiments of the invention, texture model 510 may include a linear space of texture vectors corresponding to individual images. The texture vectors may be determined as the colors at the vertices of the template mesh.
[0056] Parameterized face model 505 and texture model 510 can be used to synthesize faces based on known parameters of face identity, facial expression, and texture. Parameterized face model 505 and texture model 510 can also be used to determine unknown parameters of face identity, facial expression, and texture based on new images of new faces.
[0057] Synthesizing a face using the parameterized facial model 505 and the texture model 510 is not time-consuming; however, the synthesized face may not be realistic, especially in the mouth and eye areas. In some embodiments of the present invention, the DNN 515 can be trained to generate realistic images of the mouth and eye areas of a face. The DNN 515 can be trained using a set of videos of talking individuals. The mouth and eye areas of the talking individuals can be collected from the frames of the video. The DNN 515 can be trained using a generative adversarial network (GAN) to predict the mouth and eye areas of a face based on a predetermined number of previous frames of the mouth and eye areas and the desired facial expression of the current frame. The previous frames of the mouth and eye areas can be extracted in the specific moment parameters of the facial expression. The DNN 515 can allow the synthesis of mouth and eye areas with the parameters required for facial expressions. The DNN 515 can also allow the use of previous frames to obtain spatial coherence.
[0058] The GAN performs adjustments on the mouth and eye regions rendered from the facial model, current expression parameters, and embedded features from previously generated images, and produces the same but more realistic regions. The mouth and eye regions generated using the DNN 515 can be used to replace the mouth and eye regions synthesized by the parameterized facial model 505. It should be noted that synthesizing the mouth and eye regions by the DNN may be less time-consuming than synthesizing the entire face by the DNN. Therefore, the DNN can be used to generate the mouth and eye regions in real time by one or more processors of a mobile device such as a smartphone or tablet.
[0059] In some embodiments, the pre-processing module 520 may be configured to receive the pre-generated video 410 and the image of the source face 405. The target video 410 may include the target face. The pre-processing module 520 may also be configured to segment at least one frame of the target video to obtain an image of the target face 415 and the target background. The segmentation may be performed using a neural network, matting, and smoothing.
[0060] In some embodiments, the pre-processing module 520 may also be configured to use the parameterized facial model 505 and the texture model 510 to determine a set of target facial parameters based on at least one frame of the target video 410. In some embodiments, the target parameters may include a target facial identity, a target facial expression, and a target texture. In some embodiments, the pre-processing module 520 may also be configured to use the parameterized facial model 505 and the texture model 510 to determine a set of source facial parameters based on an image of the source face 405. The set of source facial parameters may include a source facial identity, a source facial expression, and a source texture.
[0061] In some embodiments, the face synthesis module 525 can be configured to replace the source facial expression in a set of source facial parameters with the target facial expression to obtain a set of output parameters. The face synthesis module 525 can also be configured to use the set of output parameters and the parameterized face model 505 and the texture model 510 to synthesize the output face.
[0062] In some embodiments, a two-dimensional (2D) deformation may be applied to a target face to obtain a realistic image of an output facial region hidden in the target face. Parameters of the 2D deformation may be determined based on a set of source parameters of a parameterized facial model.
[0063] In some embodiments, the mouth and eye generation module 530 can be configured to generate mouth and eye regions using the DNN 515 based on the source facial expression and at least one previous frame of the target video 410. The mouth and eye generation module 530 can also be configured to replace the mouth and eye regions in the output face synthesized using the parameterized facial model 505 and the texture model 510 with the mouth and eye regions synthesized using the DNN 515.
[0064] Figure 6 An exemplary screen of a user interface of a system for providing personalized video in a messaging application (messenger) according to some exemplary embodiments is shown. The user interface 600 may include a chat window 610 and a portion containing a video 640. The video 640 may include a pre-rendered video with a face portion 650 instead of a face. The pre-rendered video may include a preview video that is intended to show the user an exemplary representation of how the personalized video may look. The face portion 650 may be shown in the form of a white oval. In some embodiments, the video 640 may include multiple face portions 650 to be able to create a multi-person video, i.e., a video with multiple faces. The user may click on any of the videos 640 to select one of the videos 640 to be modified and sent to the chat window 610. The modification may include receiving a selfie from the user (i.e., an image of the user's face taken by the front camera of the computing device), obtaining a source face from the selfie, and modifying the selected video 640 using the source face to create a personalized video, also referred to herein as a "Reel". Therefore, as used herein, a Reel is a personalized video generated by modifying a video template (a video without a user's face) to a video with a user's face inserted. Thus, a personalized video may be generated in the form of an audiovisual medium (e.g., video, animation, or any other type of media) that features a close-up of the user's face. The modified video may be sent to the chat window 610. The user interface 600 may also have a button 630, upon clicking which the user may switch from the messaging application to the system for providing personalized video according to the present invention and use the functionality of the system.
[0065] Figure 7Exemplary screens of user interfaces 710 and 720 of a system for providing personalized video in instant messaging software according to some exemplary embodiments are shown. User interfaces 710 and 720 show a selfie acquisition mode in which a user can take an image of the user's face and then use it as a source face. When the user intends to capture a selfie image, the user interface 710 displays a real-time view of the camera of the computing device. The real-time view can display the user's face 705. The user interface 710 can display a selfie oval 730 and a camera button 740. In an exemplary embodiment, the camera button 740 can slide up from the bottom of the screen in the selfie acquisition mode. The user may need to change the position of the camera so that the user's face 705 is positioned within the boundary of the selfie oval 730. When the user's face 705 is not in the center of the selfie oval 730, the selfie oval 730 can be designed in the form of a dotted line, and the camera button 740 is translucent and inoperable to indicate that the camera button 740 is inactive. In order to notify the user that his face is not centered, text 760 can be displayed below the selfie oval 730. The text 760 may include instructions to the user, such as “center your face,” “find good lighting,” etc.
[0066] User interface 720 shows a real-time view of the camera of the computing device after the user changes the position of the camera to capture a selfie image and the user's face 705 becomes centered in the selfie oval 730. In particular, when the user's face 705 becomes centered in the selfie oval 730, the selfie oval 730 becomes a thick solid line and the camera button 740 becomes opaque and operable to indicate that the camera button 740 is now active. To notify the user, text 760 can be displayed below the selfie oval 730. The text 760 can instruct the user to take a selfie, such as "Take a selfie", "Try not to smile", etc. In some embodiments, the user can select an existing selfie from the photo library by pressing the camera album button 750.
[0067] Figure 8Exemplary screens of user interfaces 810 and 820 of a system for providing personalized video in instant messaging software according to some exemplary embodiments are shown. After a user takes a selfie, user interfaces 810 and 820 are displayed on the screen. User interface 810 may display a background 800, an illustration 805 of a currently created scroll, and text 815. Text 815 may include, for example, "Create my scroll". User interface 820 may display a created scroll 825 and text portions 830 and 835. Scroll 825 may be displayed in full screen mode. Text 815 may include, for example, "Your scroll is ready". A dark gradient may be provided behind scroll 825 so that text 830 is visible. Text portion 835 may display, for example, "Use this selfie to send a scroll in a chat or retake to try again" to inform the user that the selfie photo that the user has already taken may be used or another selfie photo may be taken. In addition, two buttons may be displayed on user interface 820. Button 840 may be displayed with a blue and filled background and may instruct the user to "use this selfie". When the user clicks button 840, two-person scrolling can be enabled. Button 845 can be displayed with a white, outlined, and transparent background and can instruct the user to "retake selfie". When the user clicks button 845, Figure 7 The user interface 710 shown in FIG. 10 can start the steps of creating a scroll, as shown in FIG. Figure 7 The user interface 820 may also display a lower text 850 below the buttons 840 and 845. The lower text 850 may inform the user how to delete the scroll, for example, "You can delete your scroll selfie in settings."
[0068] Fig. 9 An exemplary screen of a user interface 900 of a system for providing personalized videos in instant messaging software according to some exemplary embodiments is shown. The user interface 900 may be displayed after the user selects and confirms the user's selfie picture. The user interface 900 may display a chat window 610 and a scroll portion with personalized videos 910. In an exemplary embodiment, the personalized videos 910 may be displayed in a vertically scrolling tile list, with four personalized video 910 tiles per row. All personalized videos 910 may be automatically played (automatically replayed) and looped (continuously played). Regardless of the sound settings of the computing device or the user clicking the volume button, the sound in all personalized videos 910 may be turned off. Similar to tags commonly used in instant messaging software, personalized videos 910 may be indexed and searchable.
[0069] Fig.10An exemplary screen of a user interface 1000 of a system for providing personalized video in instant messaging software according to some exemplary embodiments is shown. The user interface 1000 may include a chat window 610, a personalized video list with a selected video 1010, and an action bar 1020. The action bar 1020 may slide upward from the bottom of the screen to enable the user to take action on the selected video 1010. The user may take certain actions on the selected video 1010 from the action bar 1020 through buttons 1030 to 1060. Button 1030 is a "view" button that enables the user to view the selected video 1010 in full screen mode. Button 1040 is an "export" button that enables the user to export the selected video 1010 using another application or save the selected video 1010 to the memory of the computing device. Button 1050 is a "new selfie" button that enables the user to take a new selfie. Button 1060 is a "send" button that enables the user to send the selected video 1010 to the chat window 610.
[0070] The user can click button 1030 to view the selected video 1010 in full screen mode. When clicking button 1030, button 1060 ("Send" button) can remain in place on action bar 1020, enabling the user to insert the selected video 1010 into chat window 610. When the selected video 1010 is reproduced in full screen mode, other buttons can fade out. The user can click the right side of the screen or slide left / right to browse between videos in full screen mode. The user can move to the next video until the user completes a row and then moves to the first video in the next row. When the selected video 1010 is displayed in full screen mode, the volume of the selected video 1010 corresponds to the volume setting of the computing device. If the volume is on, the selected video 1010 can be played at volume. If the volume is off, the selected video 1010 can be played without volume. If the volume is off but the user clicks the volume button, the selected video 1010 can be played at volume. If the user selects a different video, the same settings are applied, i.e., the selected video can be played at volume. If the user leaves the chat conversation view displayed on user interface 1000, the volume setting of the video can be reset to correspond to the volume setting of the computing device.
[0071] Once the selected video 1010 is sent, both the sender and the receiver can view the selected video 1010 in the same manner in the chat window 610. When the selected video 1010 is in a collapsed view, the sound of the selected video 1010 can be turned off. The sound of the selected video 1010 can only be played when the selected video 1010 is viewed in full screen mode.
[0072] The user can view the selected video 1010 in full screen mode and slide the selected video 1010 downward to exit from full screen mode and return to the chat conversation view. The user can also click the downward arrow in the upper left corner to close it.
[0073] Clicking button 1040, the "Export" button, can trigger the presentation of a share sheet. The user can directly share the selected video 1010 through any other platform or save it to the photo gallery on the computing device. Some platforms may automatically play the video in the chat, while other platforms may not automatically play the video. If the platform does not automatically play the video, the selected video 1010 can be exported in a graphics interchange format (GIF) format. The operating system of some computing devices may have a share menu that allows you to select which file to share to which platform, so it may not be necessary to add a custom action sheet. Some platforms may not be able to play GIF files and display them as static images, and the selected video 1010 may be exported to these platforms as a video.
[0074] Fig.11 Exemplary screens of user interfaces 1110 and 1120 of a system for providing personalized videos in instant messaging software according to some exemplary embodiments are shown. The user interface 1110 may include a chat window 610, a personalized video list with a selected video 1115, and an action bar 1020. When the new selfie button 1050 is clicked, the user interface 1120 may be displayed. Specifically, when the state shown on the user interface 1110 is selected in the video, the user can click the new selfie 1050 to view an action table to allow the user to select whether to select a selfie from the gallery (via the "Select from Camera Album" button 1125) or take a new selfie using the camera of the computing device (via the "Take Selfie" button 1130). Clicking the "Take Selfie" button 1130 may guide the user to complete the following steps: Figure 7 The process shown.
[0075] After the user takes a selfie using the camera or selects a selfie from the camera roll, they can start Figure 8 Clicking the "Select Face from Camera Album" button 1125 takes the user to the selfie on the Camera Album page, which can be slid up from the bottom of the screen at the top of the chat window 610. The selfie can then be positioned in the reference Figure 7 Described in the selfie ellipse.
[0076] When a user receives a scroll for the first time and has not yet created his own scroll, the system can encourage the user to create his own scroll. For example, when the user is viewing a scroll that the user has received from another user in full screen mode, a "Create My Scroll" button can be displayed at the bottom of the scroll. The user can click the button or swipe up on the scroll to bring the camera button to the screen and enter the reference Figure 7 A selfie mode described in detail.
[0077] In an exemplary embodiment, the scroll can be categorized to allow the user to easily find the general emotion the user wants to convey. A predetermined number of categories for various emotions can be provided, such as close-up, greeting, love, happiness, frustration, celebration, etc. In some exemplary embodiments, search tags can be used instead of categories.
[0078] Fig.12 1 is a flow chart illustrating a method 1200 for providing personalized videos according to an exemplary embodiment. The method 1200 may be performed by a computing device 105. The method 1200 may begin in block 1205, storing one or more pre-processed videos by the computing device. The one or more pre-processed videos may include at least one frame. The at least one frame may include at least one target face. The method 1200 may continue to receive an image of a source face by the computing device, as shown in block 1210. The method 1200 may further continue at block 1215, where the one or more pre-processed videos may be modified to generate one or more personalized videos. The modification may include modifying the image of the source face to generate an image of a modified source face. The modified source face may adopt a facial expression of the target face. The modification may also include replacing at least one target face with the image of the modified source face.
[0079] Fig.13 1 is a flowchart illustrating a method 1300 for sharing a personalized video according to some exemplary embodiments of the present invention. The method 1300 may be performed by the computing device 105. The method 1300 may provide Fig.12 The method 1300 may begin in block 1305 by enabling, by the computing device, a communication chat between a user of the computing device and at least one other user of at least one other computing device. The method 1300 may continue in block 1310 by receiving, by the computing device, a video selected by the user from one or more personalized videos. The method 1300 may also include sending, by the computing device, the selected video to the at least one other user via the communication chat, as shown in block 1315.
[0080] Fig.14An exemplary computing system 1400 that can be used to implement the methods described herein is illustrated. The computing system 1400 can be implemented in contexts such as computing devices 105 and 110, instant messaging software service system 130, instant messaging software 220, and system 300 for providing personalized video.
[0081] like Fig.14 As shown, the hardware components of the computing system 1400 may include one or more processors 1410 and memory 1420. The memory 1420 stores instructions and data for execution by the processor 1410 in part. When the system 1400 is running, the memory 1420 may store the code of the executable file. The system 1400 may also include an optional mass storage device 1430, an optional portable storage medium drive 1440, one or more optional output devices 1450, one or more optional input devices 1460, an optional network interface 1470, and one or more optional peripheral devices 1480. The computing system 1400 may also include one or more software components 1495 (e.g., those software components that may implement the method for providing personalized video as described herein).
[0082] Fig.14 The components shown in are depicted as being connected via a single bus 1490. These components may be connected via one or more data transmission devices or data networks. The processor 1410 and memory 1420 may be connected via a local microprocessor bus, and the mass storage device 1430, peripheral device 1480, portable storage device 1440, and network interface 1470 may be connected via one or more input / output (I / O) buses.
[0083] Mass storage device 1430, which may be implemented with a magnetic disk drive, solid state disk drive, or optical disk drive, is a non-volatile storage device for storing data and instructions for use by processor 1410. Mass storage device 1430 may store system software (e.g., software component 1495) for implementing the embodiments described herein.
[0084] The portable storage media drive 1440 operates with portable nonvolatile storage media, such as compact disks (CDs) or digital video disks (DVDs), to input and output data and code to and from the computing system 1400. System software (e.g., software component 1495) for implementing the embodiments described herein may be stored on such portable media and input to the computing system 1400 via the portable storage media drive 1440.
[0085] Optional input device 1460 provides part of the user interface. Input device 1460 may include an alphanumeric keypad (such as a keyboard) for entering alphanumeric and other information, or a pointing device such as a mouse, trackball, stylus, or cursor direction keys. Input device 1460 may also include a camera or scanner. In addition, as Fig.14 The illustrated system 1400 includes an optional output device 1450. Suitable output devices include speakers, printers, network interfaces, and monitors.
[0086] The network interface 1470 can be used to communicate with external devices, external computing devices, servers, and networked systems via one or more communication networks (such as one or more wired, wireless, or optical networks, including, for example, the Internet, an intranet, a local area network, a wide area network, a cellular telephone network, a Bluetooth radio, and a radio frequency network based on IEEE802.14, etc.). The network interface 1470 can be a network interface card, such as an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device that can send and receive information. The optional peripherals 1480 can include any type of computer support device to add additional functionality to the computer system.
[0087] The components contained in computing system 1400 are intended to represent a broad class of computer components. Thus, computing system 1400 may be a server, a personal computer, a handheld computing device, a phone, a mobile computing device, a workstation, a minicomputer, a mainframe computer, a network node, or any other computing device. Computing system 1400 may also include different bus configurations, networking platforms, multi-processor platforms, etc. Various operating systems (OS) may be used, including UNIX, Linux, Windows, Macintosh OS, Palm OS, and other suitable operating systems.
[0088] Some of the above functions may consist of instructions stored on a storage medium (e.g., a computer-readable medium or a processor-readable medium). The instructions may be retrieved and executed by a processor. Some examples of storage media are storage devices, tapes, disks, etc. The instructions are operable when executed by a processor to direct the processor to operate in accordance with the present invention. Those skilled in the art are familiar with instructions, processors, and storage media.
[0089] It is worth noting that any hardware platform suitable for performing the processing described herein is suitable for the present invention. The term "computer-readable storage medium" used herein refers to any medium that participates in providing instructions to a processor for execution. Such media can take a variety of forms, including but not limited to non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical disks or disks, such as fixed disks. Volatile media include dynamic memory, such as system random access memory (RAM). Transmission media include coaxial cables, copper wires, and optical fibers, etc., including wires, which include an embodiment of a bus. Transmission media can also take the form of sound waves or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, floppy disks, floppy disks, hard disks, tapes, any other magnetic media, CD read-only memory (ROM) disks, DVDs, any other optical media, any other physical media with markings or hole patterns, RAM, PROM, EPROM, EEPROM, any other storage chip or cassette, carrier waves, or any other computer-readable medium.
[0090] Various forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to the processor for execution. The bus transfers the data to the system RAM, from which the processor retrieves and executes the instructions. The instructions received by the system processor may be selectively stored on a fixed disk before or after execution by the processor.
[0091] Thus, methods and systems for providing personalized video have been described. Although the embodiments have been described with reference to specific exemplary embodiments, it is apparent that various modifications and changes may be made to these exemplary embodiments without departing from the broader spirit and scope of the present application. Therefore, the description and drawings are to be regarded as illustrative rather than restrictive.
Claims
1. A method for providing personalized video, the method comprising: receiving, by a computing device, a preprocessed video including a target face; providing, by the computing device, a first user interface to enable a user to generate an image of a source face; modifying, by the computing device, the pre-processed video to generate one or more personalized videos by replacing the target face with the source face, the source face being modified to adopt a facial expression of the target face; as well as A second user interface is provided by the computing device to select the one or more personalized videos.
2. The method according to claim 1, further comprising: determining that the user has selected the video from the pre-processed videos; as well as In response to the determination, a third user interface is provided to enable the user to select an action to be applied to the selected video from a list of actions.
3. The method according to claim 2, wherein: The action list includes: changing a first action for modifying an image of the source face of the pre-processed video; Second action to view the selected video in full screen mode; A third action of sending the selected video to another computing device; and The fourth action exports the selected video to a file in a predetermined video format.
4. The method according to claim 3, further comprising: an act of determining that the user has selected to view the selected video in the full screen mode; and In response to the determination: displaying the selected video in the full screen mode; as well as Selection of the first action, the third action, and the fourth action is disabled.
5. The method according to claim 1, wherein: Providing the first user interface includes prompting the user to enter a selfie capture mode to generate an image of the source face using a camera of the computing device.
6. The method according to claim 5, further comprising: Determining that the user has entered the selfie collection mode; In response to the determination, displaying a selfie oval and a camera button; as well as The user is prompted to center their face within the selfie oval.
7. The method according to claim 6, further comprising: Determining that the user's face is not centered in the selfie ellipse; as well as In response to the determination, the camera button is disabled.
8. The method according to claim 6, further comprising: Determining that the user's face is located at the center of the selfie ellipse; as well as In response to the determination, the camera button is enabled.
9. The method of claim 8, further comprising, in response to determining that the user's face is centered within the selfie ellipse, changing a type of contour associated with the selfie ellipse.
10. The method according to claim 6, further comprising: determining that the user has pressed the camera button; as well as In response to the determination, the user is prompted to confirm that the image of the source face will be used to modify the pre-processed video.
11. A computing device comprising: processor; as well as A memory storing instructions, which, when executed by the processor, configure the computing device to: receiving a preprocessed video including a target face; providing a first user interface to enable a user to generate an image of a source face; modifying the pre-processed video to generate one or more personalized videos by replacing the target face with the source face, the source face being modified to adopt the facial expressions of the target face; as well as A second user interface is provided for selecting the one or more personalized videos.
12. The computing device of claim 11, wherein: The instructions further configure the computing device to: determining that the user has selected the video from the pre-processed videos; and In response to the determination, a third user interface is provided to enable the user to select an action to be applied to the selected video from a list of actions.
13. The computing device of claim 12, wherein: The action list includes: changing a first action for modifying an image of the source face of the pre-processed video; Second action to view the selected video in full screen mode; A third action of sending the selected video to another computing device; and The fourth action exports the selected video to a file in a predetermined video format.
14. The computing device of claim 13, wherein: The instructions further configure the computing device to: determining that the user has selected an action to view the selected video in the full screen mode; and In response to the determination: displaying the selected video in the full screen mode; as well as Selection of the first action, the third action, and the fourth action is disabled.
15. The computing device of claim 11, wherein: Providing the first user interface includes prompting the user to enter a selfie capture mode to generate an image of the source face using a camera of the computing device.
16. The computing device of claim 15, wherein: The instructions further configure the computing device to: Determining that the user has entered the selfie collection mode; In response to the determination, displaying a selfie oval and a camera button; and The user is prompted to center their face within the selfie oval.
17. The computing device of claim 16, wherein: The instructions further configure the computing device to: Determining that the user's face is not centered within the selfie ellipse; and In response to the determination, the camera button is disabled.
18. The computing device of claim 16, wherein: The instructions further configure the computing device to: Determining that the user's face is centered within the selfie ellipse; and In response to the determination, the camera button is enabled.
19. The computing device of claim 18, wherein: The instructions further configure the computing device to, in response to determining that the user's face is centered within the selfie ellipse, change a type of contour associated with the selfie ellipse.
20. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computing device, cause the computing device to: receiving a preprocessed video including a target face; providing a first user interface to enable a user to generate an image of a source face; modifying the pre-processed video to generate one or more personalized videos by replacing the target face with the source face, the source face being modified to adopt the facial expressions of the target face; as well as A second user interface is provided for selecting the one or more personalized videos.