Video generation method and related apparatus
By acquiring facial images and video frame sequences, extracting facial and head pose feature information, and rendering to generate target synthetic videos, the problem of low-quality videos generated from single images is solved, and higher-quality video generation is achieved.
Patent Information
- Application Number
- PCT/CN2025/090294
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-27
- Filing Date
- 2025-04-22
- Publication Date
- 2026-01-02
AI Technical Summary
Existing technologies generate low-quality videos based on single face images, lacking detailed information, resulting in videos that are not realistic or vivid enough.
By acquiring facial images and video frame sequences, facial features, head pose features, and facial expression features are extracted, and combined with rendering technology to generate target synthetic videos, supplementing dynamic information to improve video quality.
It enhances the realism and appeal of generated videos, making them more vivid and interesting, with facial features consistent with the original images, and natural movements and expressions.
Smart Images

Figure CN2025090294_02012026_PF_FP_ABST
Abstract
Description
Video generation method and related device
[0001] The present application claims priority to the Chinese patent application No. 202410854816.4, filed on June 27, 2024, and entitled "Video generation method and related device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present disclosure relates to the field of artificial intelligence, and more particularly, to a video generation method. BACKGROUND
[0003] Talking face generation technology refers to generating a video of a person talking based on a single image including the person's face. Talking face generation technology has a wide range of applications. For example, it can be used to create realistic talking face animations for movies, TV series, and animations, thereby improving visual effects and audience experience. It can also be used to create virtual teachers or characters for online education and training, thereby enhancing teaching effectiveness and interactivity. It can also be used to generate realistic virtual characters for virtual reality and augmented reality applications, thereby providing a more immersive experience.
[0004] However, the quality of the generated video is low based on the way of generating a video of a person talking based on a single image including the person's face. SUMMARY
[0005] Embodiments of the present application provide a video generation method and related device, which solve the problem of low quality of the generated video of a person talking due to the lack of facial detail information in a single reference image in the prior art.
[0006] In an aspect, the present application provides a video generation method, comprising:
[0007] obtaining a face image and a video frame sequence, wherein the video frame sequence includes facial movement information when talking, and the video frame sequence includes M video frame images, M being an integer greater than 1;
[0008] extracting facial feature information from the face image;
[0009] extracting head pose features and facial expression features of the face when talking from the M video frame images, to obtain M head pose feature information and M facial expression feature information;
[0010] rendering according to the facial feature information, the M head pose feature information, and the M facial expression feature information, to obtain M target frame images;
[0011] generate a target synthesis video according to the M target frame images.
[0012] Another aspect of the present application provides a video generation apparatus, comprising:
[0013] a data acquisition module configured to acquire a face image and a video frame sequence, wherein the video frame sequence comprises face motion information when speaking, and the video frame sequence comprises M video frame images, M being an integer greater than 1;
[0014] a face information extraction module configured to extract face feature information from the face image;
[0015] a pose expression extraction module configured to extract head pose features and face expression features of the face when speaking from the M video frame images, to obtain M head pose feature information and M face expression feature information;
[0016] a feature rendering module configured to render according to the face feature information, the M head pose feature information, and the M face expression feature information, to obtain M target frame images;
[0017] a video synthesis module configured to generate a target synthesis video according to the M target frame images.
[0018] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0019] This application provides a video generation method and related apparatus. First, a face image and a video frame sequence including facial motion information during speech are acquired. This video frame sequence reflects subtle movements and facial expressions of the face over time during speech, providing richer raw data and supplementing missing facial details in a single face image. Next, facial feature information is extracted from the face image to provide the face's identity information. Then, head pose features and facial expression features of the face during speech are extracted from each video frame image in the video frame sequence, providing dynamic information about the face during speech. Then, the facial feature information, head pose features, and facial expression features are rendered to obtain a target frame image. Specifically, the facial feature information is adjusted and optimized using the head pose and facial expression features, ensuring that the identity features of the face in the target synthetic video generated from the target frame image are identical to those in the face image. This maintains consistency in identity features while allowing the face to perform the same actions and expressions as in the video frame sequence, improving the quality of the target synthetic video. The method provided in this application supplements static facial images with missing dynamic information during speech, such as head posture features and facial expression features, by introducing a video frame sequence containing facial motion information during speech. This increases the richness of the data required to generate the target synthetic video, thereby improving the quality of the target synthetic video. Consequently, it enhances the realism and appeal of the person speaking in the target synthetic video, making the generated video more vivid and interesting. Attached Figure Description
[0020] Figure 1 is a schematic diagram of one application mode of the video generation method provided in a certain embodiment of this application;
[0021] Figure 2 is a second schematic diagram of the application mode of the video generation method provided in a certain embodiment of this application;
[0022] Figure 3 is a schematic diagram of the structure of a terminal device for applying the video generation method provided in a certain embodiment of this application;
[0023] Figure 4 is an application environment diagram of a video generation method provided in a certain embodiment of this application;
[0024] Figure 5 is a flowchart of one embodiment of the video generation method provided in this application;
[0025] Figure 6 is a schematic diagram of a video generation method provided in one embodiment of this application;
[0026] Figure 7 is a second flowchart of a video generation method provided in one embodiment of this application;
[0027] FIG. 8 is a structural diagram of a masked multi-reference map identity guide according to an embodiment of the present application;
[0028] FIG. 9 is a flowchart of a video generation method according to a third embodiment of the present application;
[0029] FIG. 10 is a structural diagram of a masked multi-reference map identity guide according to a second embodiment of the present application;
[0030] FIG. 11 is a flowchart of a video generation method according to a fourth embodiment of the present application;
[0031] FIG. 12 is a structural diagram of a masked multi-reference map identity guide according to a third embodiment of the present application;
[0032] FIG. 13 is a flowchart of a video generation method according to a fifth embodiment of the present application;
[0033] FIG. 14 is a structural diagram of a masked perception motion guide according to an embodiment of the present application;
[0034] FIG. 15 is a flowchart of a video generation method according to a sixth embodiment of the present application;
[0035] FIG. 16 is a structural diagram of a denoising network according to a first embodiment of the present application;
[0036] FIG. 17 is a flowchart of a video generation method according to a seventh embodiment of the present application;
[0037] FIG. 18 is a structural diagram of a denoising network according to a second embodiment of the present application;
[0038] FIG. 19 is a flowchart of a video generation method according to an eighth embodiment of the present application;
[0039] FIG. 20 is a structural diagram of a denoising network according to a third embodiment of the present application;
[0040] FIG. 21 is a framework diagram of an application model of a video generation method according to an embodiment of the present application;
[0041] FIG. 22 is a flowchart of a video generation method according to a ninth embodiment of the present application;
[0042] FIG. 23 is a schematic diagram of face shape alignment according to an embodiment of the present application;
[0043] FIG. 24 is a schematic diagram of a video generation method according to a second embodiment of the present application;
[0044] FIG. 25 is a flowchart of a video generation method according to a tenth embodiment of the present application;
[0045] FIG. 26 is a third schematic diagram of a video generation method according to an embodiment of the present application;
[0046] FIG. 27 is an eleventh flowchart of a video generation method according to an embodiment of the present application;
[0047] FIG. 28 is a fourth schematic diagram of a video generation method according to an embodiment of the present application;
[0048] FIG. 29 is a first test diagram of a video generation result according to an embodiment of the present application;
[0049] FIG. 30 is a second test diagram of a video generation result according to an embodiment of the present application;
[0050] FIG. 31 is a twelfth flowchart of a video generation method according to an embodiment of the present application;
[0051] FIG. 32 is a fifth schematic diagram of a video generation method according to an embodiment of the present application;
[0052] FIG. 33 is a schematic diagram of a video frame interpolation inference of a speaking face according to an embodiment of the present application;
[0053] FIG. 34 is a thirteenth flowchart of a video generation method according to an embodiment of the present application;
[0054] FIG. 35 is a fourteenth flowchart of a video generation method according to an embodiment of the present application;
[0055] FIG. 36 is a fifteenth flowchart of a video generation method according to an embodiment of the present application;
[0056] FIG. 37 is a generation result on a face video frame interpolation task according to an embodiment of the present application;
[0057] FIG. 38 is a structural schematic diagram of a video generation apparatus according to an embodiment of the present application;
[0058] FIG. 39 is a structural schematic diagram of a server according to an embodiment of the present application. DETAILED DESCRIPTION
[0059] The embodiments of the present application provide a video generation method. By introducing a video frame sequence with face motion information in speaking, the dynamic information missing in speaking, such as head posture feature information and facial expression feature information, is supplemented for a static face image. Thus, the richness of data required for generating a target synthetic video is improved, and the quality of the target synthetic video is improved. Furthermore, the realism and attractiveness of a person in speaking in the target synthetic video are improved, so that the generated video is more vivid and interesting.
[0060] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of these terms herein is to be construed to cover a generalised use of these terms to describe elements distinguishable from other elements. It is to be understood that the data so used in these embodiments can be interchanged, under suitable circumstances, without departing from the scope of the embodiments of the present application described herein. Furthermore, the terms "comprising" and "including" and any of their derivatives, as used herein, are intended to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense unless otherwise indicated. Also, the use of "including" and "comprising" and variations thereof, does not imply that more, less or only so many steps are necessary or essential to the embodiments of the present application described herein.
[0061] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0062] In order to facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used by the embodiments of the present application are explained here:
[0063] Diffusion models: Diffusion models are a type of latent variable generative model that is trained using Markov chains applied through variational estimation. The goal is to learn the underlying structure of a dataset by modeling the way data points diffuse in latent space. Specifically, the model corrupts the original data by adding continuous Gaussian noise and gradually recovers the original data through an inverse process. Diffusion models can be applied to tasks such as image denoising, image inpainting, super-resolution imaging, and image generation.
[0064] Talking face generation: This technology aims to generate a talking face video from a single image that includes a face. This process can be guided by text or audio used for driving, and the face image in the video is driven and synthesized through the audio signal, so that the mouth shape and expression of the face are synchronized with the sound in the audio. This technology is mainly applied in the fields of video production, virtual reality, animated movies, etc., and can improve the naturalness and immersion of audiovisual media.
[0065] Existing talking face generation techniques are mostly based on GAN (Generative adversarial network) schemes, which can be further divided into 2D (Two Dimensional) based methods and 3D (Three Dimensional) based methods.
[0066] 2D based methods mainly use sparse key points, motion field estimation and first-order local affine transformation techniques to generate talking face videos through analysis and processing of a single image. This method has the advantage of high computational efficiency and can quickly generate results. However, since it is based on 2D images, it may not fully capture the depth and realism of the face.
[0067] 3D based methods use face parameter models such as 3DMM (3D Morphable Face Model) to learn prior information from a large amount of face data, providing a more convenient tool for driving facial expressions. These methods can usually generate more realistic and stereoscopic faces, but have higher computational costs and require more computing resources and time.
[0068] That is, existing talking face generation techniques can only generate videos based on a single reference image, and it is difficult for a single reference image to include all the information of the person in the generated video, such as the person in the reference image not showing teeth, while the generated video requires the person to open their mouth and show teeth. It is difficult to generate high-quality and detailed teeth. In other words, the method of generating a video based on a single image including a face (hereinafter referred to as single image generating video) based on the face talking, due to the less information included in the single image, resulting in the quality of the generated video is low. Based on this, the present application introduces a video frame sequence including face motion information when talking, which can reflect the subtle movements and expression changes of the face over time during talking, thereby providing more abundant raw data, which can supplement the missing face details of a single face image to improve the quality of the subsequently generated target synthesized video.
[0069] In addition, in the manner of generating a video only through a single image, different users have different requirements for the generated video, but the single image provides limited facial details. Therefore, in order to better meet the needs of users and make the information of the face in the generated video more abundant, that is, the face information is expandable, the present application designs a multi-reference image strategy combined with an image mask mechanism, so that users can select the form of reference image according to actual needs, such as selecting only a single picture, or selecting a single picture and an additional picture only including a face area (without background), or selecting two complete pictures including the same face as a reference. The mask multi-reference image identity guide can adaptively extract the information required for generating a video from the additional reference image, so that users can improve the details of the generated result as needed.
[0070] Furthermore, in the manner of generating a video only through a single image, the video can only be generated in a frame-by-frame driving manner, which is not the most required by users. On the one hand, the motion sequence is not necessarily accurate, and on the other hand, even if the motion sequence is accurate, the frame-by-frame driving may cause the motion of the face and the characteristics of the person in the generated video to be inconsistent. Based on this, the present application designs a mask-aware motion guide. This strategy discards part of the driving frames at random, so that the model better models the natural face motion manner, and the generated result is more natural. When users finally use the product provided by the video generation method of the embodiments of the present application, they can select a more accurate driving manner (frame-by-frame driving) or a more natural and more consistent driving manner with the characteristics of the person (keyframe driving) according to needs.
[0071] In summary, the manner of generating a video only through a single image is not flexible enough, which not only affects the upper limit of the generation quality of the visually driven speaker face, but also hinders the actual landing of the task. The method provided by the embodiments of the present application generates a video with a natural speaker face in a more controllable manner of visually driven speaker face, such as expanding the appearance and customizing the motion, thereby improving the realism and attractiveness of the person when speaking in the target synthesized video, and improving the quality of the target synthesized video, such as the identity features of the face in the target synthesized video (hereinafter referred to as face features) being consistent with the face features in the face image, improving the authenticity of the video, and introducing more information such as video frame sequence to enrich the details required for face generation, making the generated face more complete, and thus making the generated video more lively and interesting.
[0072] In one implementation scenario, as shown in FIG. 1, which is one of the application modes of the video generation method provided by the embodiments of the present application, the application mode is suitable for some related data calculation that can be completed by the terminal device 400 relying on the calculation capability of the graphic processing hardware of the terminal device 400. The output of the target synthetic video is completed by various terminal devices 400 such as smart phones, tablet computers, and virtual reality / augmented reality devices.
[0073] As an example, the types of graphic processing hardware include a central processing unit (CPU) and a graphic processing unit (GPU).
[0074] When forming the visual perception of the target synthetic video, the terminal device 400 calculates the data required for display by graphic calculation hardware, and completes the loading, parsing, and rendering of the display data. The video frame capable of forming the visual perception of the target synthetic video is output by the graphic output hardware, for example, the two-dimensional video frame is presented on the display screen of the smart phone, or the video frame achieving the three-dimensional display effect is projected on the lens of the augmented reality / virtual reality glasses; in addition, in order to enrich the perception effect, the terminal device 400 can also form one or more of auditory perception, tactile perception, motion perception, and gustatory perception by means of different hardware.
[0075] As an example, please refer to (a) in FIG. 1, the client 410 (for example, an application for video generation) is running on the terminal device 400. The user can upload the face image 100 through the client 410, select the style 200 of the video to be generated, or select the video template such as the facial expression in the video to be generated, or input the information such as the content of the speech to be generated. Please refer to (b) in FIG. 1, the client 410 acquires the corresponding video frame sequence according to the video template, and generates the target synthetic video 300 according to the face image and the video frame sequence. The terminal device 400 displays or plays the target synthetic video 300.
[0076] For example, when the client 410 receives the face image 100 uploaded by the user, and acquires the corresponding video frame sequence according to the video style 200 selected by the user, then extracts the facial feature information from the face image 100, and extracts the head posture feature and the facial expression feature of the face when speaking from the video frame sequence. Then, the target frame image is obtained by rendering according to the facial feature information, the head posture feature information, and the facial expression feature information. Finally, the target synthetic video 300 is generated according to the target frame image. The face feature of the target synthetic video 300 is the same as the face feature in the face image. The terminal device 400 displays or plays the target synthetic video 300.
[0077] In one implementation scenario, as shown in FIG. 2, which is a schematic diagram of an application mode of the video generation method provided in the embodiments of the present application, the terminal device 400 runs a client 410 (for example, an application of video generation), and the user can upload a face image 100 through the client 410 and select a video style 200 to be generated, or select a video template such as a facial expression in the video to be generated, or input information such as the content of speech to be generated. The client 410 acquires a corresponding video frame sequence according to the video template, and sends the face image and the video frame sequence to the server 200, and the server 200 generates a target synthesis video according to the face image and the video frame sequence, and sends the target synthesis video to the client 410, and the terminal device 400 displays or plays the target synthesis video.
[0078] For example, in the case of forming a visual perception of the target synthesis video, the server 200 performs calculation on the face image and the video frame sequence including the facial motion information when speaking, and sends the calculation result to the terminal device 400 through the network 300, and the terminal device 400 relies on the graphic calculation hardware to complete loading, parsing and rendering of the display data, and relies on the graphic output hardware to output the target synthesis video to form a visual perception, for example, the two-dimensional video frame can be presented on the display screen of a smart phone, or the video frame achieving a three-dimensional display effect can be projected on the lens of augmented reality / virtual reality glasses; for the perception of the form of the target synthesis video, it can be understood that the corresponding hardware output of the terminal device 400 can be used, for example, a microphone is used to form an auditory perception, a vibrator is used to form a tactile perception, and the like.
[0079] For example, when the client 410 receives the face image uploaded by the user and the video style selected by the user, the client uploads the face image uploaded by the user and the video style selected by the user to the server 200, the server 200 receives the face image and the video style, acquires a corresponding video frame sequence according to the video style selected by the user, then extracts facial feature information from the face image, and extracts head posture features and facial expression features of the face when speaking from the video frame sequence, then renders according to the facial feature information, the head posture feature information and the facial expression feature information to obtain a target frame image, and finally generates a target synthesis video according to the target frame image, the facial features of the target synthesis video being the same as the facial features in the face image. The server 200 sends the target synthesis video to the client 410, and the terminal device 400 displays or plays the target synthesis video.
[0080] In some embodiments, the application can also be implemented by means of cloud technology. Cloud technology refers to a kind of hosting technology that unifies a series of resources such as hardware, software and network to realize data calculation, storage, processing and sharing in a wide area network or a local area network. The structure of the terminal device 400 shown in FIG. 3 will be described below. Referring to FIG. 3, FIG. 3 is a structural schematic diagram of the terminal device 400 provided by the application. The terminal device 400 shown in FIG. 3 includes at least one processor 420, a memory 460, at least one network interface 430 and a user interface 440. The various components in the user interface 400 are coupled together through a bus system 450. It can be understood that the user interface 440 is used to realize the connection communication between the components. The bus system 450 includes not only a data bus, but also a power bus, a control bus and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 450 in FIG. 3.
[0081] The processor 420 can be an integrated circuit chip with a signal processing capability, such as a general purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general purpose processor can be a microprocessor or any conventional processor.
[0082] The user interface 440 includes one or more output devices 441 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 440 also includes one or more input devices 442 that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0083] The memory 460 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. The memory 460 optionally includes one or more storage devices physically located in proximity to the processor 420.
[0084] The memory 460 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 460 described in the embodiments of the application is intended to include any suitable type of memory.
[0085] In some embodiments, the memory 460 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, examples of which are illustratively described below.
[0086] The operating system 461 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks;
[0087] The network communication module 462 is configured to communicate with other computing devices via one or more (wired or wireless) network interfaces 430, examples of which include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), and the like;
[0088] The presentation module 463 is configured to enable presentation of information via one or more output devices 441 (e.g., a display screen, a speaker, and the like) associated with the user interface 440 (e.g., a user interface for operating a peripheral device and displaying content and information);
[0089] The input processing module 464 is configured to detect and interpret one or more user inputs or interactions from one or more input devices 442.
[0090] In some embodiments, the video generation apparatus provided by the embodiments of the present application can be implemented in a software manner, and FIG. 3 shows a video generation apparatus 465 stored in the memory 460, which can be software in the form of programs and plug-ins, and includes the following software modules: a data acquisition module 4651, a face information extraction module 4652, a posture expression extraction module 4653, a feature rendering module 4654, and a video synthesis module 4655. These modules are logical, and thus can be combined or further split according to the implemented functions. It should be noted that, in FIG. 3, all the above modules are shown at one time for the convenience of expression, but should not be regarded as excluding the implementation that the video generation apparatus 465 can only include the data acquisition module 4651, the face information extraction module 4652, the posture expression extraction module 4653, the feature rendering module 4654, and the video synthesis module 4655. The functions of each module will be described below.
[0091] In some embodiments, the video generation apparatus provided in the embodiments of the present application can be implemented in a hardware manner. For example, the video generation apparatus provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to perform the video generation method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can be implemented by using one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic elements.
[0092] The video generation method provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The video generation method provided in the embodiments of the present application can be performed by the terminal device 400 in FIG. 1 alone, or can be performed by the terminal device 400 and the server 200 in FIG. 2 in cooperation.
[0093] For ease of understanding, please refer to FIG. 4, which is an application environment diagram of the video generation method in the embodiments of the present application. As shown in FIG. 4, the video generation method in the embodiments of the present application is applied to a video generation system. The video generation system includes a server and a terminal device. The server can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a palm computer, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and the embodiments of the present application do not limit this.
[0094] The video generation method provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The video generation method provided in the embodiments of the present application can be performed by the terminal device 400 in FIG. 1 alone, or can be performed by the terminal device 400 and the server 200 in FIG. 2 in cooperation.
[0095] S110, obtaining a face image and a video frame sequence.
[0096] The video frame sequence includes facial motion information of a person speaking, and the video frame sequence includes M video frame images, where M is an integer greater than 1.
[0097] It can be understood that the user can upload a face image through the client and select a style of a video to be generated or select a video template of a facial expression in the video to be generated. The face image includes a face to be generated for a video of speaking, i.e., a target synthesized video. For example, if a video of a face A speaking is to be generated, a face image including the face A needs to be provided.
[0098] The server receives the face image sent by the client, and the face image includes feature information of the face. For example, from a visual perspective, the face image is a two-dimensional or three-dimensional representation of a person's face, including shapes, positions, proportions of features such as eyes, nose, mouth, ears, and intuitive features such as facial contours and skin. From an information perspective, the face image contains rich personal identity information (i.e., identity features), which has a high degree of uniqueness. Each person's face is unique and can be used for identity recognition and verification.
[0099] The server receives selection information of the user for the video template sent by the client, and obtains a video frame sequence corresponding to the video template according to the video template selected by the user. The video frame sequence includes facial motion information of a person speaking, i.e., information of changes in the face during speaking. The video frame sequence includes M video frame images, where M is an integer greater than 1. The M video frame images refer to a series of ordered video image frames. These video image frames are continuous in time and can present smooth dynamic pictures when played in sequence.
[0100] The video frame sequence has temporal continuity, representing continuous sampling over a period of time. Adjacent frames usually have a small time interval, so that subtle movements and expression changes of the face over time during speaking can be captured. The video frame sequence also has content relevance. Each frame includes information associated with the previous and next frames, which together build a complete facial motion process during speaking. The video frame sequence is an organized collection, and each video frame has a specific position and order, which is crucial for accurately understanding and reproducing facial motion. The video frame sequence carries various feature information of the face during speaking, such as changes in head position and evolution of facial expressions, which plays a key role in subsequent feature extraction and video generation. The video frame sequence can be regarded as a digital record of dynamic behavior of the face during speaking. By analyzing this sequence, the rules and characteristics of facial motion can be understood in depth.
[0101] The video frame image is a single image unit in the video frame sequence. In terms of content, it specifically presents the state of the face at a certain time, including the appearance of the facial features, the contour of the face, and the details of the skin. From the functional point of view, it is one of the basic elements constituting the entire video content, and the numerous video frame images played in sequence form a dynamic video effect. From the technical characteristics, it has a specific resolution and pixel information, which accurately depicts the visual performance of the face at that moment. It also contains time-related information, which can be analyzed through comparison with other video frame images to analyze the motion trend and change of the face in this period of time, such as the rotation of the head and the change of the expression. Moreover, each video frame image is an indispensable part of the entire video, and they are combined together to fully show the continuous action and expression of the face when speaking. It can be said that the video frame image is a freeze frame of the face at a certain time, and through the analysis and processing of a series of such images, various operations and applications of the face video can be realized.
[0102] As a possible implementation manner, the face in the face image and the face in the image frame sequence can be not only different but also the same, so that the image frame sequence representing the same face provides more accurate and less confusing information about the face in the face image, thereby improving the accuracy of the subsequent target synthesis video.
[0103] As a possible implementation manner, the number of face images is not specifically limited to one or more, and it should be noted that the multiple face images include the same face. In addition, after obtaining a face image, the required face in the face image can be cropped to obtain a partial face image including only the required face, so that the subsequent processing is performed according to the face image and the partial face image, so as to improve the detail information of the required face and improve the accuracy of the subsequent target synthesis video.
[0104] It can be understood that after obtaining the face image, the face feature information is extracted by using a specific algorithm and technology. The face feature information can include key features such as facial contour, position and shape of facial features, and the face feature information will serve as a reference in the subsequent rendering process.
[0105] The facial feature information is information used to describe the features of a face, so as to provide the identity information of the face based on the facial feature information. The facial feature information usually includes but is not limited to the following categories: (1) facial feature: such as the shape, size and position of the eyes, the shape and direction of the eyebrows, the contour of the nose, the shape and size of the mouth, etc. These features are very important for recognizing and distinguishing different faces. (2) facial contour features: including the overall shape of the face, such as whether the face shape is round, square or other shapes, and the width of the cheeks, the shape of the chin, etc. (3) facial texture features: such as skin texture, wrinkles, etc., which can provide information about age, expression, etc.
[0106] The process of extracting facial feature information may use some models trained based on machine learning or deep learning. These models are trained on a large amount of data and can accurately identify and extract various facial features. By quantifying and representing these features, a face image can be converted into a set of feature vectors or feature descriptors, which can be used for subsequent analysis, comparison or synthesis. The accuracy and completeness of this step are crucial to the quality and effect of the entire video generation process, and it provides the basis data for subsequent rendering and generating target video based on these features.
[0107] In S130, head pose features and facial expression features of the face when speaking are extracted from the M video frame images, obtaining M head pose feature information and M facial expression feature information.
[0108] It can be understood that first, computer vision and image processing techniques are needed to analyze the face in the video frame images. This can involve steps such as face detection, face alignment, etc. to ensure accurate acquisition of the position and pose of the face. For the extraction of head pose features, the direction and angle of the face in the image can be analyzed. This can be achieved by calculating the positions and relative relationships of the key points of the face, such as the eyes, nose, mouth, etc. Head pose features refer to a series of attributes or parameters that can describe the position, direction, shape, etc. of the head in space. Head pose features can include information such as the rotation angle and tilt degree of the head. Facial expression features refer to a series of attributes or parameters that can describe the changes in facial expressions. The extraction of facial expression features requires the analysis of muscle movements and expression changes of the face. This can be achieved by using expression recognition algorithms, which are usually based on feature points or texture patterns of facial expressions. Common facial expression features include smiling, frowning, blinking, etc. To extract these features, deep learning models such as convolutional neural networks (CNN) can be used to train a large amount of facial image data. These models can learn the feature representation of the face and automatically extract head pose features and facial expression features. After extracting the features, each video frame image can obtain corresponding head pose feature information and facial expression feature information. These information can be represented in the form of numerical values or vectors for subsequent processing and analysis. By extracting head pose features and facial expression features from M video frame images, dynamic information about the face during speaking can be obtained. These information is of great significance for understanding the emotional state, intention, and coordination with speech of the speaker.
[0109] In addition, further analysis and processing of these features can be performed, such as calculating the statistics of the features, performing feature fusion, or correlating with other modalities of information (such as speech) for more comprehensive and in-depth understanding. It should be noted that the accuracy and reliability of feature extraction are crucial for subsequent applications and analysis. Therefore, in practical applications, more complex and accurate algorithms may be used, and multiple features may be combined to improve performance. At the same time, for different application scenarios and requirements, appropriate feature representation and analysis methods may need to be selected.
[0110] The head pose feature information refers to information used to describe the head pose features, and the facial expression feature information refers to information used to describe the facial expression features.
[0111] Taking one of the M video frame images as an example, the head posture feature of the face when speaking can be extracted from the video frame image to obtain the head posture feature information of the face, and the facial expression feature of the face when speaking can be extracted from the video frame image to obtain the facial expression feature information of the face. Thus, extraction is performed for each video frame image to obtain M head posture feature information and M facial expression feature information. As a possible implementation manner, if the M video frame images include multiple faces, extraction can be performed for the same face in the M video frame images, so that the obtained M head posture feature information and M facial expression feature information are from the same face, so as to improve the accuracy of the subsequent target synthesis video.
[0112] In S140, rendering is performed according to the face feature information, the M head posture feature information, and the M facial expression feature information to obtain M target frame images.
[0113] The target frame image is an image obtained according to the face feature information, the head posture feature information, and the facial expression feature information. In the target frame image, the head posture feature information is used to control the face represented by the face feature information to make a corresponding head action, and the facial expression feature information is used to control the face represented by the face feature information to make a corresponding expression action.
[0114] As a possible implementation manner, one target frame image can be obtained based on the face feature information, one head posture feature information, and one facial expression feature information, and the head posture feature information and the facial expression feature information are from the same video frame image. Thus, M target frame images are obtained through the head posture feature information and the facial expression feature information included in the M video frame images. In addition, the M target frame images can have a time sequence relationship, that is, the arrangement order of the M target frame images is the same as that of the M video frame images, so as to ensure that the motion track of the face in the subsequent target synthesis video is more natural.
[0115] It can be understood that first, the facial feature information, head pose feature information and facial expression feature information need to be integrated. This can be achieved by associating or fusing these information with the video frame images. Next, the integrated information is processed using rendering techniques. Rendering can include adjusting the light, color, texture, etc. of the image to make it more in line with specific requirements or to present specific effects. During the rendering process, the pose and expression of the face can be adjusted according to the head pose feature and facial expression feature. For example, if the head pose feature indicates that the face is tilted to the left, the image can be adjusted accordingly during rendering so that the face in the target frame image also presents a left tilt. Similarly, according to the facial expression feature, the expression of the face in the image can be simulated and rendered. In addition, other post-processing of the target frame image can be performed as needed, such as adding background, special effects, etc. to further enhance the effect of the image.
[0116] Through rendering processing, M target frame images are finally obtained. These images, based on preserving the original video frame content, are adjusted and optimized according to the extracted feature information to better show the pose and expression of the face when speaking.
[0117] It should be noted that the specific implementation of rendering can vary depending on the technology and tools used. Common image processing and computer graphics techniques can be used to achieve the rendering effect. At the same time, the quality and effect of rendering will be affected by factors such as the accuracy of feature extraction and the performance of rendering algorithms. In actual application, multiple tests and adjustments may be needed to obtain satisfactory target frame images.
[0118] S150, generating a target synthesis video according to the M target frame images.
[0119] The target synthesis video is a video obtained based on the M target frame images. It can be obtained by arranging the M target frame images in a time sequence, or by further processing a video obtained by arranging the M target frame images in a time sequence. The present application does not make specific limitations on this.
[0120] It can be understood that the parameters of the video are determined, including the frame rate, resolution, duration, etc. of the video. These parameters will determine the playback speed and picture quality of the synthesis video. According to the set frame rate, the M target frame images are written into the video one by one. Image processing libraries or video editing software can be used to achieve this step. When writing the images, attention should be paid to the order and time interval of the images to ensure the smoothness and coherence of the synthesis video. The synthesis video can be further edited and processed as needed, such as adding audio, adjusting picture brightness, contrast, etc. Finally, the target synthesis video is obtained.
[0121] The face features of the target synthesized video are the same as those in the face image, and from the appearance features, it means that in the synthesized video, the basic appearance characteristics of the face, such as the shape, proportion, relative position of the features, and the facial contour, skin, etc., are consistent with the initial provided face image. From the perspective of identity recognition, the same face features ensure that the same person can be recognized through these features, that is, uniqueness and distinguishability. In terms of expression and posture, although the face in the video may have various dynamic expressions and head posture changes, the unique morphology of the basic features of the face, such as the eyes, nose, and mouth, does not change. In the entire video generation process, the algorithm and model successfully preserve the key feature information in the original face image and accurately reproduce these features in each frame of the generated video.
[0122] For example, if there is a unique mole or special eyebrow shape in the original face image, then in each corresponding frame of the generated target synthesized video, the mole and eyebrow shape should appear in the same position, shape, and size. During the conversion from the face image to the target synthesized video, the core features of the face remain unchanged, ensuring the identity consistency and feature stability of the face in the video.
[0123] For ease of understanding, please refer to FIG. 6, which is one of the schematic diagrams of the video generation method provided in the embodiments of the present application. First, the face image and the video frame sequence including the face motion information when speaking are obtained, then the face feature information is extracted from the face image, and the head posture feature information and the facial expression feature information of the face when speaking are extracted from the video frame sequence, then the target frame image is obtained by rendering according to the face feature information, the head posture feature information, and the facial expression feature information, and finally, the target synthesized video is generated according to the target frame image. The face features of the target synthesized video are the same as those in the face image, and the expressions, head movements, etc. of the face in the video frame sequence are synthesized into the face in the target frame image. For example, the face in the face image is smiling, the face in the video frame image is pursing the mouth, and the face features in the target frame image are the same as those in the face image, but the face in the target frame image is pursing the mouth. The method provided in the embodiments of the present application can generate a video with a natural speaker's face, improve the realism and appeal of the person speaking in the target synthesized video, improve the quality and integrity of the target synthesized video, and make the generated video more lively and interesting.
[0124] The video generation method provided in the application ensures that rich original data, including static features of the face and dynamic facial movement information, is obtained through the face image and the video frame sequence including facial movement information when speaking, lays a foundation for subsequent generation of high-quality and diversified videos, and the continuous video frame sequence can capture subtle changes in the face when speaking, so that the generated video is more realistic and natural. By extracting facial feature information from the face image, the facial feature information is more accurate, which provides a key reference for subsequent rendering and synthesis, ensures the basic feature consistency of the generated video and the face image, helps to improve the accuracy and efficiency of video generation, and avoids unnecessary deviation and errors. By extracting the head posture feature and facial expression feature of the face when speaking from the video frame image, dynamic information about the face when speaking is obtained, so that the generated video can accurately reflect the emotional state and intention of the speaker, enhance the expressiveness and appeal of the video, and provide key dynamic feature data for realizing more realistic and natural face video generation. According to the extracted features, the video frames are adjusted and optimized, so that the generated video is more consistent with the user's expectations and needs, improves the quality and visual effect of the video, and shows more lively and natural posture and expression of the face when speaking. The generated video maintains the consistency of the facial features, which is helpful for identity recognition and ensures the credibility of the video. Through the video generation method provided in the application, the generation of high-quality, personalized, realistic and natural target synthesized videos from original face images and video frame sequences can be realized, which meets the needs of users in various application scenarios, such as entertainment, education, communication, etc.
[0125] Thus, the present application provides a video generation method. First, a face image and a video frame sequence including face motion information when speaking are obtained. The video frame sequence can reflect the subtle movements and expression changes of the face over time during speaking, thereby providing more abundant raw data and enabling the missing face details of a single face image to be supplemented. Next, facial feature information is extracted from the face image to provide identity information of the face, and head pose features and facial expression features of the face when speaking are extracted from each video frame image in the video frame sequence to obtain head pose feature information of each video frame image and facial expression feature information of each video frame image, thereby providing dynamic information about the face when speaking. Then, rendering is performed according to the facial feature information, the head pose feature information, and the facial expression feature information to obtain a target frame image. That is, the facial feature information is adjusted and optimized through the head pose feature information and the facial expression feature information, so that the face features in the target synthetic video generated according to the target frame image are the same as the face features in the face image, thereby maintaining identity consistency while enabling the face to make the same movements and expressions as the video frame sequence, thereby improving the quality of the target synthetic video. The method provided in the embodiments of the present application supplements the missing dynamic information such as head pose feature information and facial expression feature information in the speaking process for the static face image by introducing the video frame sequence having face motion information when speaking, thereby improving the richness of the data required for generating the target synthetic video and improving the quality of the target synthetic video. Furthermore, the realism and attractiveness of the person when speaking in the target synthetic video are improved, so that the generated video is more lively and interesting.
[0126] In one optional embodiment of the video generation method provided in the embodiment corresponding to FIG. 5 of the present application, referring to FIG. 7, step S120 further includes sub-step S121 to sub-step S123. Sub-step S121 to sub-step S123 can be implemented by a mask multi-reference figure identity guide. Specifically:
[0127] S121, performing feature sampling on the face image to obtain a sampling image.
[0128] It can be understood that specific points or regions are selected from the original face image for feature extraction. Through feature sampling, the dimension of the data can be reduced, the amount of calculation in subsequent processing can be reduced, and the key parts or representative feature regions of the face can be focused on. The sampling method can be uniform sampling, key point-based sampling, or adaptive sampling based on image content, etc.
[0129] Preferably, during the sampling process, multiple samplings can be performed, and the output of each sampling is used as the input of the next sampling, and the image obtained by each sampling is used as a sampling image, so as to obtain multiple sampling images, thereby improving the degree of focusing on the key parts or representative feature areas of the face. In step S122, the image features of each sampling image need to be spliced with the image features of the face image.
[0130] S122, the image features of the face image and the image features of the sampling image are spliced to obtain face image splicing features.
[0131] It can be understood that first, the image features of the face image and the image features of the sampling image are extracted respectively to obtain the image features of the face image and the image features of the sampling image. The image features are a set of information used to describe and distinguish the content of the image, such as the image features of the face image used to identify the face image, and the image features of the sampling image used to identify the sampling image. Then, the image features of the face image and the image features of the sampling image are spliced to obtain face image splicing features, which combines the overall features of the original face image and the local features obtained by sampling. In this way, the global information and local key information of the face can be considered comprehensively, so that the subsequent processing can fully utilize features of different levels and ranges, thereby more comprehensively and accurately describing the features of the face.
[0132] As a possible implementation, the embodiment corresponding to FIG. 7 can also include S123, which encodes the face image splicing features to obtain encoded features.
[0133] It can be understood that the purpose of encoding the spliced features is to convert the complex feature representation into a more compact and representative form. The encoding process can be implemented by using a CLIP (Contrastive Language-Image Pre-Training, a multi-modal learning model using language and image) encoder, an encoder in a VAE (Variational Autoencoder, Variational Autoencoder) (hereinafter referred to as VAE encoder) or other specific encoding algorithms, in order to better capture the internal relationship and pattern between the features. By encoding the face image splicing features, the complex feature representation can be converted into a more compact form, which is beneficial for storage and transmission, and also facilitates subsequent processing and analysis.
[0134] S124, the face image splicing features are processed by a first attention mechanism to obtain face feature information.
[0135] The first attention mechanism is an attention mechanism, which is a technique in artificial neural networks that simulates cognitive attention. The core idea is to dynamically focus on key parts of the input data, thereby improving the model's ability to capture important information. The embodiments of the present application do not specifically limit the first attention mechanism, and those skilled in the art can set it according to actual needs, such as self-attention mechanism (Self-Attention), multi-head attention mechanism (Multi-Head Attention), etc.
[0136] It can be understood that the introduction of the first attention mechanism can selectively focus on important parts of the encoded features, giving higher weights to key features and suppressing less important features. This can highlight features that have a significant impact on generating facial feature information, improve the effectiveness and accuracy of the features, and thus obtain more accurate and targeted facial feature information, providing high-quality input for subsequent video generation steps.
[0137] As a possible implementation, if the face image stitching features are encoded to obtain the encoded features, the encoder features can be processed by the first attention mechanism to obtain the facial feature information.
[0138] For ease of understanding, please refer to FIG. 8, which is one of the structural diagrams of the mask multi-reference figure identity guide provided by the embodiments of the present application. The mask multi-reference figure identity guide includes a sampling module, an image feature extraction module, a feature stitching module, an encoding module, and a first attention mechanism module. First, the face image is sampled at least once by the sampling module to obtain at least one sample image. Then, the image features of the face image and the sample image are extracted by the image feature extraction module to obtain the image features of the face image and the sample image. Then, the image features of the face image and the sample image are stitched by the feature stitching module to obtain the face image stitching features. Then, the face image stitching features are encoded by the encoding module to obtain the encoded features. Finally, the encoded features are processed by the first attention mechanism module to obtain the facial feature information.
[0139] The video generation method provided in the application can focus attention on the key parts or representative feature areas of the face by sampling the face image, which helps to better capture important features of the face. The face image splicing features obtained by splicing the overall features of the face image and the local features of the sampled image can comprehensively consider global and local information, making the generated facial feature information more comprehensive and accurate. The first attention mechanism can selectively focus on important features, improving the effectiveness and accuracy of the features, thereby generating facial feature information that is more targeted and of higher quality. The high-quality facial feature information provides good input for the subsequent video generation step, which helps to generate more natural, realistic and lively speaker face videos.
[0140] In one optional embodiment of the video generation method provided in the embodiment corresponding to FIG. 7 of the application, referring to FIG. 9, the sub-step S123 further includes a sub-step S1231 and a sub-step S1232. It should be noted that the sub-step S1231 and the sub-step S1232 can be parallel steps or sub-steps with execution sequence, which are not limited in the embodiments of the application. The embodiments of the application take the sub-step S1231 and the sub-step S1232 as parallel steps for illustration. Specifically:
[0141] S1231, encoding the face image splicing features through a semantic feature encoder to obtain a semantic feature encoding vector.
[0142] It can be understood that the semantic feature encoder is used to extract semantic information of the face image, such as facial expression, emotion and other high-level semantic features. The semantic feature encoding vector obtained by encoding the face image splicing features through the semantic feature encoder can capture the semantic content of the face.
[0143] Illustratively, the semantic feature encoder can be a CLIP encoder. The CLIP encoder is a neural network model that has a significant impact in the field of multi-modal. CLIP encoding is an image encoding method based on contrastive learning, which learns the semantic information of the image by contrastive learning between the image and the text. The CLIP encoder is usually composed of an image encoder and a text encoder, and the two encoders can be neural networks based on the Transformer architecture. The CLIP encoder adopts a contrastive learning training method. In the process of training the CLIP encoder, the model learns to distinguish matching and non-matching image-text pairs, thereby capturing the semantic similarity between the image and the text. The training data of CLIP encoding is a large number of image-text pairs. The scale and diversity of these data help the model to learn universal image and text representations.
[0144] S1232, encode the face image splicing feature through an image feature encoder to obtain an image feature encoding vector.
[0145] It can be understood that the image feature encoder is mainly used to extract the image features of the face image, such as the texture, shape and other underlying features of the image. The image feature encoding vector can reflect the visual features of the face image. The image feature encoding vector and the semantic feature encoding vector jointly constitute the encoding feature. Illustratively, the image feature encoder can be a VAE encoder. The VAE encoder is an image encoding method based on variational autoencoder. It learns the latent distribution of the image to realize the encoding and decoding of the image. VAE is a generative model composed of an encoder and a decoder. The encoder maps the input data to the latent space, and the decoder decodes the vector in the latent space into the output data. Unlike traditional autoencoders, the latent space of VAE is continuous, and the outputs of the encoder and the decoder are probability density distributions of parameter-constrained variables.
[0146] When training the VAE, the model learns the probability distribution of the latent variable (feature) Z. In order to make this probability distribution easy to handle, VAE introduces a variational inference model (encoder) to approximate this distribution by optimizing the parameters through a deep neural network. The loss function of VAE consists of two parts. One part is the reconstruction loss, which measures the difference between the output generated by the decoder and the original input. The other part is the KL divergence, which measures the difference between the distribution learned by the encoder and the prior distribution. By minimizing this loss function, VAE can learn the latent representation of the data and generate new data.
[0147] For ease of understanding, please refer to FIG. 10, which is a second structural diagram of the mask multi-reference image identity guide provided by the embodiments of the present application. The mask multi-reference image identity guide includes a sampling module, an image feature extraction module, a feature splicing module, an encoding module, and a first attention mechanism module, wherein the encoding module includes a semantic feature encoder and an image feature encoder. First, at least one sampling image is obtained by sampling the face image at least once through the sampling module. Then, the image features of the face image and the sampling image are obtained by performing image feature extraction on the face image and the sampling image through the image feature extraction module. Then, the semantic feature encoding vector is obtained by encoding the face image splicing feature through the semantic feature encoder (CLIP encoder), and the image feature encoding vector is obtained by encoding the face image splicing feature through the image feature encoder (VAE encoder). The encoding vector includes the semantic feature encoding vector and the image feature encoding vector. Finally, the encoding feature composed of the image feature encoding vector and the semantic feature encoding vector is processed through the first attention mechanism module to obtain the face feature information.
[0148] The video generation method provided in the present application can obtain multi-dimensional feature representations of face images by using different encoders (semantic feature encoder and image feature encoder). This can more comprehensively describe the features of the face, including semantic information and image features, and improve the accuracy and effect of subsequent processing. The semantic feature encoder can extract semantic information of the face image, such as expression, emotion, etc. This helps better understand the meaning and intention of the face, which is very helpful for applications involving emotion analysis, human-computer interaction, etc. The image feature encoder focuses on capturing image features of the face image, such as texture, shape, etc. This helps to retain the detailed information of the image, making the generated video more realistic and natural. The semantic feature encoding vector and the image feature encoding vector are fused, which can comprehensively utilize the advantages of the two kinds of features. This helps better control and generate facial expressions, postures, etc. in the subsequent video generation process, improving the quality of the generated video. Using different encoders can be flexibly selected and adjusted according to specific needs and application scenarios. For example, in some cases, semantic features may be more important, while in other cases, image features may be more critical. This flexibility enables the method to adapt to different tasks and requirements. By introducing multi-dimensional feature representation and fusion mechanism, the performance and generalization ability of the entire model (i.e. the carrier of the video generation method provided in the present application) can be improved. The model can better handle different face image and video generation tasks, thereby producing more accurate and natural results.
[0149] In one optional embodiment of the video generation method provided in the embodiment corresponding to FIG. 9 of the present application, referring to FIG. 11, the first attention mechanism is composed of K self-attention layers and K cross-attention layers, K being an integer greater than 1; the sub-step S124 further includes a sub-step S1241. Specifically:
[0150] S1241, the semantic feature encoding vector is taken as the input of the K cross-attention layers, and the image feature encoding vector is taken as the input of the first attention mechanism, and the semantic feature encoding vector and the image feature encoding vector are processed by the first attention mechanism to obtain the facial feature information.
[0151] It can be understood that the first attention mechanism includes 2K layers, which are composed of K self-attention layers and K cross-attention layers to form the first attention mechanism. For example, the first cross-attention layer is after the first self-attention layer, the second self-attention layer is after the first cross-attention layer, and so on.
[0152] Self-Attention Layer is a mechanism widely used in deep learning to compute attention weights for each element in an input sequence and then perform a weighted sum of the elements based on these weights to get a representation of the sequence. The main idea of Self-Attention Layer is to determine the importance of elements by computing their relevance to each other, thus capturing long-range dependencies in the sequence. Unlike traditional Recurrent Neural Networks (RNN) or Convolutional Neural Networks (CNN), Self-Attention Layer can operate on the entire sequence directly without processing each element sequentially. In Self-Attention Layer, a Query vector, a Key vector, and a Value vector are usually used to compute attention weights. The Query and Key vectors compute their relevance through dot product or other similarity measures, and then use a softmax function to convert the relevance into a probability distribution, i.e., attention weights. Finally, the attention weights are multiplied with the Value vector to get the weighted sum result, i.e., the representation of the sequence.
[0153] Cross-Attention Layer is a mechanism used in deep learning, often for handling multi-modal data or interacting between different feature spaces. It can help models (such as first attention mechanisms) better understand and fuse information from different sources or representations. In Cross-Attention Layer, the model considers two or more input sequences simultaneously and computes attention weights between them. These attention weights represent the importance or relevance of each input element to others. By performing a weighted sum of the attention weights, a fused feature representation can be obtained. Unlike Self-Attention Layer, which only focuses on the internal relationships within a single input sequence, Cross-Attention Layer focuses on the relationships between different input sequences. This makes Cross-Attention Layer useful for handling multi-modal data, fusing different features, or interacting across domains. The specific implementation of Cross-Attention Layer can vary depending on the application scenario and model architecture. One common implementation is to use Multi-Head Attention, where multiple attention heads compute different attention weights and combine them to get the final output.
[0154] The first attention mechanism is composed of K self-attention layers and K cross-attention layers. The self-attention layer can capture the relationship between different positions in the input sequence (such as image feature encoding vectors), and the cross-attention layer can capture the relationship between different feature spaces. By combining these two attention mechanisms, the feature information in the input data can be more comprehensively extracted, and the representation ability of the model can be improved. The self-attention mechanism can establish a dependency relationship between any positions in the input sequence, which is helpful for processing long sequence data. By combining self-attention layers and cross-attention layers, the modeling ability of the model for long-distance dependencies can be further enhanced, and the performance of the model can be improved. This cross structure can flexibly adjust the number and combination of self-attention layers and cross-attention layers according to the characteristics of specific tasks and data to adapt to different application scenarios. When processing multi-modal data, the cross-attention layer can be used to fuse the features of different modalities, and the self-attention layer can be used to capture the relationship within each modality. This structure helps the model better understand and process multi-modal information. By introducing multiple attention mechanisms, the model can have stronger robustness to changes and noise in input data, reducing the risk of overfitting.
[0155] For ease of understanding, please refer to FIG. 12, which is a third structure diagram of the mask multi-reference figure identity guide provided by the embodiments of the present application. The mask multi-reference figure identity guide includes a sampling module, an image feature extraction module, a feature splicing module, an encoding module, and a first attention mechanism module, wherein the encoding module includes a semantic feature encoder and an image feature encoder, and the first attention mechanism module includes K self-attention layers and K cross-attention layers. First, at least one sampling image is obtained by sampling the face image at least once through the sampling module, then the image features of the face image and the sampling image are obtained by performing image feature extraction on the face image and the sampling image through the image feature extraction module, then the semantic feature encoding vector is obtained by performing encoding processing on the face image splicing feature through the semantic feature encoder (CLIP encoder), and the image feature encoding vector is obtained by performing encoding processing on the face image splicing feature through the image feature encoder, finally, the semantic feature encoding vector is taken as the input of the K cross-attention layers in the first attention mechanism module, and the image feature encoding vector is taken as the input of the first attention mechanism, and the face feature information is obtained by performing feature processing on the semantic feature encoding vector and the image feature encoding vector through the first attention mechanism.
[0156] The video generation method provided in the application can realize the fusion of semantic features and image features by taking the semantic feature encoding vector as the input of the cross-attention layer and taking the image feature encoding vector as the input of the first attention mechanism. This helps to comprehensively consider the semantic features (such as expressions, emotions, etc.) and image features (such as textures, shapes, etc.) of the face, thereby more comprehensively describing the features of the face. The attention mechanism can weight the features according to the correlation between the features, thereby highlighting important feature information and suppressing unimportant information. This helps to improve the accuracy of feature representation, making the subsequent processing (such as video generation) more accurate and natural. The cross-attention layer can capture the dependency between different positions in the input sequence, thereby better processing long-distance information in the face image. This is very important for accurately simulating the expression and posture changes of the face. By fusing semantic and image features and using the attention mechanism for feature processing, the model can learn more general and representative feature representations. This helps to improve the generalization ability of the model, enabling it to better process new, unseen face images and videos. Accurate facial feature information is crucial for generating high-quality target synthesis videos.
[0157] In one optional embodiment of the video generation method provided in any of the above embodiments of the application, referring to FIG. 13, step S130 further includes sub-step S131 to sub-step S133. Sub-step S131 to sub-step S133 can be implemented by a mask-aware motion guide.
[0158] S131, obtaining a mask feature sequence.
[0159] The mask feature sequence includes M mask features.
[0160] It can be understood that each mask feature is a binary array or matrix, wherein each mask feature includes a plurality of elements, and each element represents the state of the corresponding position, such as whether the position is covered or not covered. The mask feature sequence is composed of M mask features.
[0161] S132, performing feature extraction on the M video frame images to obtain M video frame image features.
[0162] For example, feature extraction is performed on each video frame image respectively to obtain video frame image features corresponding to each video frame image respectively, thereby obtaining M video frame image features. The video frame image feature is used to describe the content or attributes in the video frame image.
[0163] It can be understood that by feature extraction on the M video frame images, M video frame image features corresponding to the M video frame images are obtained. As a possible implementation manner, when performing feature extraction, the head posture feature and the facial expression feature of the person in the video frame image need to be extracted. First, computer vision and image processing techniques are needed to analyze the human face in the video frame image. This can involve steps such as face detection and face alignment to ensure accurate acquisition of the position and posture of the face. For the extraction of the head posture feature, the direction and angle of the face in the image can be analyzed to determine it. This can be achieved by calculating the positions and relative relationships of the key points of the face, such as the eyes, nose, mouth, etc. The head posture feature can include the rotation angle of the head, the degree of tilt, etc. The extraction of the facial expression feature requires the analysis of the muscle movement and expression change of the face. This can be achieved by using expression recognition algorithms, which are usually based on feature points or texture patterns of facial expressions for recognition. Common facial expression features include smiling, frowning, blinking, etc.
[0164] S133, generating a mask motion sequence according to the mask feature sequence and the M video frame image features.
[0165] The mask motion sequence includes M mask motion vectors, and each mask motion vector includes head posture feature information and facial expression feature information.
[0166] For example, the i-th mask feature in the mask feature sequence and the i-th video frame image feature in the M video frame image features are obtained to obtain the i-th mask motion vector in the mask motion sequence, and the i-th mask motion vector is obtained based on the head posture feature information and the facial expression feature information in the i-th video frame image. i is a positive integer less than or equal to M.
[0167] It can be understood that the M mask features in the mask feature sequence and the M video frame image features are subjected to matrix operation to adjust the dimension of the video frame image features to the same dimension as the mask features, thereby obtaining the mask motion sequence. The binary mask sequence can selectively retain or suppress information in the video frame image features. The binary mask sequence captures the motion information of the face in the video frame image features, including the head posture feature information and the facial expression feature information.
[0168] For ease of understanding, please refer to FIG. 14, which is a structural diagram of the mask-aware motion guide provided by the embodiment of the present application. First, a mask feature sequence composed of M mask features is obtained, and feature extraction is performed on the M video frame images to obtain M video frame image features. Then, matrix operation is performed on the mask feature sequence and the M video frame image features to generate a mask motion sequence including head posture feature information and facial expression feature information.
[0169] The video generation method provided in the application can process irregular parts or irrelevant information in the data set through the mask feature sequence, thereby improving the attention of the model (i.e., the mask-aware motion guide) to key information. For video frame sequences of different lengths, the mask feature sequence can help the model process variable-length data, ensuring that the model can effectively process various sequence lengths. Using the mask feature sequence can introduce regularization, reducing the over-reliance of the model on specific data, thereby improving the generalization performance of the model. By focusing on extracting head posture and facial expression features in the video frame images, the mask motion sequence can better capture the motion information of the face, thereby improving the quality of the target synthesized video.
[0170] In an optional embodiment of the video generation method provided in any of the above embodiments of the application, referring to FIG. 15, step S140 further includes sub-step S141 to sub-step S143. Sub-step S141 to sub-step S143 can be implemented by a denoising network. Specifically:
[0171] S141, performing noise adding processing on the M video frame images to obtain video frame noise-added feature information.
[0172] It can be understood that the noise adding processing can increase the diversity of data and prevent overfitting of the model. By adding noise to the original video frame images, the model can learn more robust feature representations, thereby improving the generalization ability of the model. Common noise types include Gaussian noise, salt and pepper noise, etc. These noises can simulate various disturbances and uncertainties in actual data. During the training process, the model needs to learn how to extract useful information from the noise-added feature information while ignoring the interference of noise. This helps to improve the robustness and anti-interference ability of the model.
[0173] The video frame noise-added feature information is obtained by performing noise adding processing on the M video frame images. The application embodiments do not specifically limit the noise adding processing process, which is described below by way of example.
[0174] Specifically, first, image feature extraction is performed on the M video frame images to obtain M video frame image features. Then, the M video frame image features are encoded by an image feature encoder (VAE encoder) to obtain video frame image encoded feature vectors. Finally, the video frame image encoded feature vectors are spliced with a noise vector to obtain video frame noise-added feature information. The noise vector is a vector used to describe noise, and the video frame image encoded feature vectors are spliced with the noise vector to implement the noise adding processing.
[0175] S142, performing feature processing on the video frame noise-added feature information, the facial feature information, the M head posture feature information, and the M facial expression feature information through a second attention mechanism to obtain M attention feature information.
[0176] The second attention mechanism is an attention mechanism. Embodiments of the present application do not specifically limit the second attention mechanism, and a person skilled in the art can set it according to actual needs, which can be the same as or different from the first attention mechanism.
[0177] It can be understood that introducing the second attention mechanism can help the model focus on important feature information and improve the expression ability of the features. By weighting and fusing different feature information, the key information in the video frame, such as changes in facial expression and head pose, can be better captured. The second attention mechanism can be regarded as a feature selection and weighting method, which assigns different weights according to the importance of the input features. In this step, the second attention mechanism processes the video frame noise feature information, facial feature information, head pose feature information and facial expression feature information to determine which features are more critical for generating the target frame image. Through the second attention mechanism, the model can learn the relationship and importance between different features. For features with high relevance to the generated target frame image, they are given a larger weight; while for irrelevant or secondary features, the weight is smaller. In this way, the model can pay more attention to important feature information, improving the quality and accuracy of the generated attention feature information.
[0178] For example, by processing the video frame noise feature information, facial feature information, head pose feature information and facial expression feature information through the second attention mechanism, an attention feature information is obtained, so that after the processing of the second attention mechanism, M attention feature information is finally obtained. These attention feature information is a weighted sum and selection of the original features, highlighting the important features related to the generation of the target frame image. These feature information will be used as input for the subsequent decoding step to generate M target frame images.
[0179] S143, decoding the M attention feature information to obtain M target frame images.
[0180] It can be understood that the decoding process converts the attention feature information into the target frame image, realizing the generation or conversion of the video frame. Through decoding, the target frame image with specific facial expression and head pose can be generated for subsequent video synthesis or processing. Through the decoding operation, the attention feature information is mapped back to the pixel space of the image to obtain the corresponding target frame image. This is a key step from feature representation to image generation. The specific decoding method can be determined according to the architecture and design of the model. A common method is to use deconvolution or upsampling operation to gradually increase the resolution of the attention feature information, and finally generate a target frame image with the same size as the input video frame image. In the decoding process, other information such as facial feature information, head pose feature information, etc. can be combined to further refine and optimize the generation of the target frame image.
[0181] For ease of understanding, please refer to FIG. 16, which is one of the structure diagrams of the denoising network provided in the embodiments of the present application. The denoising network includes a noise adding module, a second attention mechanism module and a decoding module. First, the noise adding module is used to perform noise adding processing on the video frame sequence including M video frame images to obtain video frame noise feature information. Then, the facial feature information output by the mask multi-reference image identity guide, the mask motion sequence including M head pose feature information and M facial expression feature information output by the mask perception motion guide, and the video frame noise feature information obtained by adding noise to the M video frame images are subjected to feature processing through the second attention mechanism in the second attention mechanism module to obtain M attention feature information. Finally, the M attention feature information is decoded through the decoding module to obtain M target frame images.
[0182] The video generation method provided in the present application adds noise in the M video frame images, so that the model learns more robust feature representation and improves the generalization ability of the model. By using the second attention mechanism for feature processing, the attention and extraction ability of the model to key features can be improved, so that more accurate and representative target frame images can be generated. By decoding the attention feature information, the conversion from features to images can be realized, and target frame images with specific features and contents can be generated.
[0183] In one optional embodiment of the video generation method provided in the embodiment corresponding to FIG. 15 of the present application, the second attention mechanism is composed of L self-attention layers, L reference attention layers and L time sequence attention layers, and L is an integer greater than 1; please refer to FIG. 17, step S142 further includes sub-step S1421. Specifically:
[0184] S1421. Facial feature information is used as input to L reference attention layers. Video frame noise feature information, M head pose feature information and M facial expression feature information are used as input to the second attention mechanism. The second attention mechanism is used to process the video frame noise feature information, facial feature information, M head pose feature information and M facial expression feature information to obtain M attention feature information.
[0185] Understandably, the second attention mechanism comprises 3L layers, consisting of L self-attention layers, L reference attention layers, and L temporal attention layers intersecting each other. For example, the first self-attention layer is followed by the first reference attention layer, the first reference attention layer by the first temporal attention layer, the first temporal attention layer by the second self-attention layer, and so on.
[0186] For example, facial feature information is used as input to L reference attention layers. Video frame noise feature information, M head pose feature information and M facial expression feature information are concatenated. The concatenated feature information is used as input to the second attention mechanism, and the second attention mechanism outputs M attention feature information.
[0187] Self-attention layers can capture the internal relationships between input features, reference attention layers can use facial feature information as a reference to guide attention to other features, and temporal attention layers can consider information in the time dimension. By combining these three attention mechanisms, the features of video frame images can be processed more comprehensively, thereby generating more accurate and meaningful attention feature information.
[0188] A reference attention layer is a layer used in attention mechanisms to guide attention to other features using reference information. In the steps described above, facial feature information is used as a reference, interacting with video frame noise features, head pose features, and facial expression features through a second attention mechanism to obtain attention feature information. Specifically, the reference attention layer works by comparing and relating the reference information to other features to determine which features should receive more attention. In this way, the model can better understand the relationships between different features and allocate attention weights based on the importance of the reference information. In the steps described above, facial feature information serves as input to the reference attention layer and is processed along with other feature information through the second attention mechanism. The benefit of this is that it allows the model to pay more attention to information related to facial features, thereby improving its understanding and processing capabilities of video frame images. For example, if facial feature information indicates that a person is smiling, the model may pay more attention to facial expression features related to smiling to better understand the person's emotional state.
[0189] The temporal attention layer is a layer used in the attention mechanism, which plays a role in allocating attention weights according to the time sequence of data when processing time-series data. Unlike traditional attention mechanisms, the temporal attention layer considers the time dimension of the data and can better capture dynamic information in time-series data. In the above steps, the working principle of the temporal attention layer is to process the video frame noise feature information, facial feature information, head pose feature information, and facial expression feature information to obtain attention feature information. Specifically, the temporal attention layer will allocate different attention weights to each feature according to the time sequence of these feature information, thereby highlighting important feature information and improving the performance and effect of the model.
[0190] For ease of understanding, please refer to FIG. 18, which is a structure diagram two of the denoising network provided in the embodiments of the present application. The denoising network includes a noise adding module, a second attention mechanism module, and a decoding module. First, the noise adding module is used to add noise to the video frame sequence including M video frame images to obtain video frame noise feature information. Then, the facial feature information is taken as the input of the L reference attention layers, the video frame noise feature information, the mask motion sequence including M head pose feature information and M facial expression feature information are spliced, and the spliced feature information is taken as the input of the second attention mechanism module. M attention feature information is output through the second attention mechanism in the second attention mechanism module. Finally, the M attention feature information is decoded through the decoding module to obtain M target frame images.
[0191] The video generation method provided in the present application can capture various feature relationships in video frame images by combining different types of attention mechanisms, thereby obtaining more rich and accurate feature representations. The introduction of the temporal attention layer can consider information in the time dimension and better capture dynamic changes in the video. The reference attention layer uses facial feature information as a reference to guide the model to pay attention to other related features, thereby improving the accuracy of feature extraction. The design and use of the second attention mechanism can improve the understanding and processing ability of the model for video frame images, thereby generating more accurate and meaningful target frame images.
[0192] In one optional embodiment of the video generation method provided in the embodiment corresponding to FIG. 15 of the present application, please refer to FIG. 19, step S141 further includes sub-step S1411 to sub-step S1413. Specifically:
[0193] S1411, encode the M video frame images through the image feature encoder to obtain video frame image encoded feature vectors.
[0194] It can be understood that the M video frame images are encoded by the image feature encoder, which converts the images from the original pixel space to the feature vector representation. This helps to extract key features in the image, reduce data dimension, and make subsequent processing more efficient.
[0195] S1412, obtaining a noise vector.
[0196] It can be understood that the noise vector is obtained to increase the robustness of the model or provide certain randomness.
[0197] S1413, concatenating the video frame image encoding feature vector and the noise vector to obtain video frame noise feature information.
[0198] It can be understood that the video frame image encoding feature vector and the noise vector are concatenated. This is done to combine image features and noise information so that the model can consider both aspects of information.
[0199] For ease of understanding, please refer to FIG. 20, which is a third structure diagram of the denoising network provided by the embodiments of the present application. The denoising network includes an encoding module, a noise adding module, a second attention mechanism module, and a decoding module. First, the M video frame images are encoded by the encoding module (i.e., the image feature encoder) to obtain video frame image encoding feature vectors. Then, the noise vector is obtained by the noise adding module, so that the video frame image encoding feature vector and the noise vector are concatenated to obtain the video frame noise feature information, realizing the noise processing of the video frame image encoding feature vector. Then, the face feature information is taken as the input of the L reference attention layers, the video frame noise feature information, and the mask motion sequence including the M head pose feature information and the M face expression feature information are concatenated, and the concatenated feature information is taken as the input of the second attention mechanism module. The second attention mechanism outputs M attention feature information through the second attention mechanism in the second attention mechanism module. Finally, the attention feature information is decoded by the decoding module to obtain M target frame images.
[0200] The video generation method provided by the present application can represent video frame images in a more concise and representative way through encoding processing, reduce redundant information, facilitate model learning and processing, and also lay a foundation for subsequent fusion and analysis with other features. By adding noise, the model is prevented from relying too much on specific patterns, and the generalization ability of the model is enhanced, so that it can better cope with various situations.
[0201] For ease of understanding, please refer to FIG. 21, which is a framework diagram of an application model of the video generation method provided in the embodiments of the present application. The application model of the video generation method includes a mask multi-reference image identity guide, a mask perception motion guide, and a denoising network. The mask multi-reference image identity guide is used to extract facial feature information from a face image, the mask perception motion guide is used to extract a mask motion sequence from a video frame sequence, and the denoising network is used to generate a target frame image according to a noisy video frame sequence, facial feature information, and a mask motion sequence. Specifically:
[0202] The mask multi-reference image identity guide includes a sampling module, an image feature extraction module, a feature splicing module, an encoding module, and a first attention mechanism module. First, the sampling module samples the face image at least once to obtain at least one sample image. Next, the image feature extraction module extracts image features of the face image and the sample image, respectively, to obtain image features of the face image and the sample image, splices the image features of the face image and the sample image to obtain face image splicing features. Then, the semantic feature encoder (CLIP encoder) in the encoding module encodes the face image splicing features to obtain a semantic feature encoding vector, and the image feature encoder in the encoding module encodes the face image splicing features to obtain an image feature encoding vector. Finally, the semantic feature encoding vector is input into K cross-attention layers in the first attention mechanism module, and the image feature encoding vector is input into the first attention mechanism, and the semantic feature encoding vector and the image feature encoding vector are processed by the first attention mechanism to obtain facial feature information.
[0203] The mask perception motion guide is used to obtain a mask feature sequence composed of M mask features, and to extract M video frame image features from M video frame images. Then, matrix operations are performed on the mask feature sequence and the M video frame image features to generate a mask motion sequence including head pose feature information and facial expression feature information.
[0204] The denoising network comprises a VAE encoding module, a noise adding module, a second attention mechanism module, and a VAE decoding module. First, the M video frame images are encoded by the encoding module (i.e., an image feature encoder) to obtain video frame image encoding feature vectors. Then, a noise vector is obtained by the noise adding module, so as to splice the video frame image encoding feature vectors and the noise vector to obtain video frame noise-added feature information, thereby realizing noise addition processing on the M video frame images. Next, the face feature information output by the mask multi-reference image identity guide, the mask motion sequence comprising the M head posture feature information and the M facial expression feature information output by the mask perception motion guide, and the video frame noise-added feature information obtained by adding noise to the M video frame images are subjected to feature processing by the second attention mechanism in the second attention mechanism module to obtain M attention feature information. Finally, the M attention feature information is decoded by the VAE decoding module to obtain M target frame images.
[0205] The video generation method provided by the embodiments of the present application realizes expansion of appearance details and customization of driving modes according to the needs of users through the mask multi-reference image identity guide and the mask perception motion guide, and finally obtains a high-quality visual driving speaker face generation result.
[0206] In one optional embodiment of the video generation method provided by any of the above embodiments of the present application, referring to FIG. 22, before step S140, steps S210 to S240 are further included. Specifically:
[0207] S210, extracting identity features from the face image to obtain reference identity feature information.
[0208] The identity features are features that uniquely identify a person's identity, and the reference identity feature information is information used to describe the identity features in the face image.
[0209] It can be understood that a specific algorithm or model is used to process the face image to extract features that can uniquely identify a person's identity. These features are usually based on various physiological structures, contours, textures, and other information of the face.
[0210] S220, extracting identity features from any video frame image to obtain driving identity feature information.
[0211] The driving identity feature information is information used to describe the identity features in the video frame image.
[0212] It can be understood that any video frame image is selected from the M video frame images in the video frame sequence, and the identity feature information of the face is extracted from the selected video frame image.
[0213] S230, in the case that the identity feature indicated by the reference identity feature information is same as the identity feature indicated by the driving identity feature information, performing step S140 and step S150.
[0214] It can be understood that, in the case that the identity feature indicated by the reference identity feature information is same as the identity feature indicated by the driving identity feature information, i.e. when the reference identity feature information corresponding to the face image is same as the identity indicated by the driving identity feature information corresponding to the video frame sequence, the head pose feature and the facial expression feature of the face during speaking can be extracted from the M video frame images to obtain M head pose feature information and M facial expression feature information.
[0215] S240, in the case that the identity feature indicated by the reference identity feature information is different from the identity feature indicated by the driving identity feature information, performing face shape alignment on the face feature information and the feature information combination to obtain the M video frame images after face shape alignment, and then performing step S140 and step S150.
[0216] The feature information combination refers to the M head pose feature information and the M facial expression feature information.
[0217] It can be understood that, in the case that the identity feature indicated by the reference identity feature information is different from the identity feature indicated by the driving identity feature information, i.e. when the reference identity feature information corresponding to the face image is different from the identity indicated by the driving identity feature information corresponding to the video frame sequence, face alignment is needed, and the head pose feature and the facial expression feature of the face during speaking are extracted from the M video frame images after obtaining the M video frame images after face shape alignment to obtain M head pose feature information and M facial expression feature information.
[0218] For ease of understanding, please refer to FIG. 23, which is a schematic diagram of face shape alignment provided by an embodiment of the present application. First, identity features are extracted from the face image to obtain reference identity feature information, and identity features are extracted from any video frame image to obtain driving identity feature information.
[0219] In the case that the identity feature indicated by the reference identity feature information is same as the identity feature indicated by the driving identity feature information, rendering is performed according to the face feature information and the M head pose feature information and the M facial expression feature information to obtain M target frame images.
[0220] In the case that the identity feature indicated by the reference identity feature information is different from the identity feature indicated by the driving identity feature information, the face shape of the face feature information and the M head posture feature information and the M face expression feature information are aligned to obtain M video frame images after face shape alignment. The face feature information is extracted from the face image, and the head posture feature and the face expression feature of the face during speaking are extracted from the M video frame images after face shape alignment. Then, the target frame image is obtained by rendering according to the face feature information, the head posture feature information and the face expression feature information. Finally, the target synthesized video is generated according to the target frame image, and the face feature of the target synthesized video is the same as the face feature in the face image.
[0221] For ease of understanding, please refer to FIG. 24, which is a second schematic diagram of the video generation method provided by the embodiment of the present application. First, a face image and a video frame sequence including face motion information during speaking are obtained. The identity feature is extracted from the face image to obtain reference identity feature information, and the identity feature is extracted from any video frame image to obtain driving identity feature information.
[0222] If the identity feature indicated by the reference identity feature information is the same as the identity feature indicated by the driving identity feature information, the face feature information is extracted from the face image, and the head posture feature and the face expression feature of the face during speaking are extracted from the video frame sequence. Then, the target frame image is obtained by rendering according to the face feature information, the head posture feature information and the face expression feature information. Finally, the target synthesized video is generated according to the target frame image, and the face feature of the target synthesized video is the same as the face feature in the face image.
[0223] If the identity indicated by the reference identity feature information is different from the identity indicated by the driving identity feature information, the face shape of the face feature information and the M head posture feature information and the M face expression feature information are aligned to obtain M video frame images after face shape alignment. The face feature information is extracted from the face image, and the head posture feature and the face expression feature of the face during speaking are extracted from the M video frame images after face shape alignment. Then, the target frame image is obtained by rendering according to the face feature information, the head posture feature information and the face expression feature information. Finally, the target synthesized video is generated according to the target frame image, and the face feature of the target synthesized video is the same as the face feature in the face image.
[0224] The video generation method provided by the present application directly performs subsequent processing if the identity feature of the person in the face image is the same as the identity feature of the person in the video frame sequence, thereby improving the processing efficiency, enabling the extraction and utilization of head posture and face expression features of the same person, and making the generated content more coherent and natural. If the identity feature of the person in the face image is different from the identity feature of the face in the video frame sequence, the face shape is aligned to ensure that the features are as reasonably presented as possible even if the persons are different, thereby increasing the adaptability and flexibility of the method, enabling the processing of multiple different identities, and expanding the application range.
[0225] In an optional embodiment of the video generation method provided by the embodiment corresponding to FIG. 5 of the present application, referring to FIG. 25, step S120 further includes a sub-step S1201, and step S140 further includes a sub-step S1401. Specifically:
[0226] S1201, extracting facial feature information and background feature information from the face image.
[0227] It can be understood that after obtaining the face image, specific algorithms and techniques are used to extract the facial feature information and the background feature information.
[0228] The background feature information is used to describe the background features in the face image, so as to represent the background through the background feature information and represent the face through the facial feature information, thereby realizing the distinction between the face and the background in the face image. The background feature information is a description about the background of the face image, such as the color, texture, and lighting condition of the background. The extraction of the background feature information helps to better restore and process the background part in the subsequent rendering process. The background features can reflect the environment in which the image is taken, such as indoor, outdoor, daytime, or night, etc. These information are very important for understanding the scene and context of the image. The features such as the color, brightness, and shadow of the background can indicate the lighting condition in the image. The lighting condition will affect the appearance and features of the face, so understanding the lighting information of the background helps to better process and analyze the face image. The background features can provide context information about the face. For example, other faces, objects, or scene elements in the background may be related to the identity, emotion, or behavior of the face.
[0229] S1401, rendering according to the background feature information, the facial feature information, the M head pose feature information, and the M facial expression feature information to obtain M target frame images.
[0230] As a possible implementation manner, one target frame image can be obtained based on the background feature information, the facial feature information, one head pose feature information, and one facial expression feature information, and the head pose feature information and the facial expression feature information come from the same video frame image. Thus, M target frame images are obtained through the head pose feature information and the facial expression feature information included in the M video frame images respectively.
[0231] It can be understood that first, the background feature information, the facial feature information, the head pose feature information, and the facial expression feature information need to be integrated. This can be realized by associating or fusing these information with the video frame image. Next, the integrated information is processed using rendering technology.
[0232] Through the rendering processing, finally, M target frame images are obtained, which are adjusted and optimized according to the extracted feature information on the basis of preserving the original video frame content, so as to better show the posture and expression of the face when speaking.
[0233] For ease of understanding, please refer to Fig. 26, which is the third schematic diagram of the video generation method provided by the embodiment of the application. First, the face image and the video frame sequence including the face motion information when speaking are obtained. Then, the face feature information and the background feature information are extracted from the face image, and the head posture feature information and the facial expression feature information of the face when speaking are extracted from the video frame sequence. Then, the M target frame images are obtained by rendering according to the background feature information, the face feature information, and the M head posture feature information and the M facial expression feature information. Finally, the target synthesis video is generated according to the target frame images. The face feature of the target synthesis video is the same as the face feature in the face image, and the background feature of the target synthesis video is the same as the background feature in the face image.
[0234] The video generation method provided by the application accurately extracts the face feature information and the background feature information, provides accurate data basis for subsequent processing, and helps to more accurately perform identification, analysis and rendering and the like. At the same time, the face and the background are considered, the image content can be more comprehensively understood, the subsequent processing is more in line with the actual scene, and the deviation caused by one-sided processing is avoided. The rendering is performed by comprehensively considering various feature information, the target frame image generated is more real and natural, whether the posture and expression of the face or the background environment is more in line with the actual situation. The background is processed in combination with the background feature information, the face and the environment are better integrated, and the coordination and the situational feeling of the overall picture are improved.
[0235] In one optional embodiment of the video generation method provided by the embodiment corresponding to Fig. 5 of the application, please refer to Fig. 27, step S150 further includes sub-step S151. Specifically:
[0236] S151, generating a target synthesis video according to the face image and the M target frame images.
[0237] The first frame of the target synthesis video is the face image, that is, the face image is taken as the first frame of the target synthesis video.
[0238] The obtained face image is taken as the first frame of the target synthesized video, which provides a starting point and a reference for the entire video. In this way, it is ensured that the synthesized video can accurately present the specific face from the beginning. Then, the M target frame images generated in the previous step are combined, which have been rendered and processed, including the adjustment and optimization of the posture and expression of the talking face according to the facial feature information, head posture feature information and facial expression feature information. By combining the face image and the M target frame images in sequence, a continuous video sequence can be formed.
[0239] For ease of understanding, please refer to FIG. 28, which is the fourth schematic diagram of the video generation method provided by the embodiment of the present application. First, the face image and the video frame sequence including the facial motion information of the talking face are obtained. Then, the facial feature information is extracted from the face image, and the head posture feature and the facial expression feature of the talking face are extracted from the video frame sequence. Then, the target frame image is obtained by rendering according to the facial feature information, the head posture feature information and the facial expression feature information. Finally, the target synthesized video is generated according to the face image and the M target frame images, wherein the face image is taken as the first frame of the target synthesized video.
[0240] The video generation method provided by the present application first takes the face image as the first frame of the target synthesized video, which can ensure that the target synthesized video presents the real face specified by the user at the beginning, providing a clear starting point and personalized identification for the entire target synthesized video, and increasing the recognition and uniqueness of the target synthesized video. Secondly, the synthesized video is generated by combining the M target frame images, which realizes the dynamic presentation of the talking face. This makes the target synthesized video have a coherent and smooth visual effect, which can vividly show the various postures and expression changes of the face during talking, and improves the authenticity and appeal of the target synthesized video. Thirdly, the video generation method provided by the present application can meet the personalized needs of different users for video style and content. Users can choose face images and video templates according to their own preferences, thereby creating unique videos that meet their imagination, providing great freedom of creation and personalized experience.
[0241] The backbone network of the algorithm model used in the video generation method provided by the embodiment of the present application accepts multiple frames of Gaussian noise (the result after adding Gaussian noise to the training video frame during training) as input, and generates a talking face video through iterative denoising. The process is jointly guided by the identity information provided by the mask multi-reference figure identity guide and the motion information provided by the mask perception motion guide.
[0242] To ensure the consistency of the generated video with the face in the face image, the video generation method provided by the embodiments of the present application uses a double-flow UNet (U-Net, U-shaped network) structure. One of the UNets is used as a mask multi-reference image identity guide for extracting facial feature information and background feature information from the face image. The other UNet is a denoising UNet, and the cross-attention layer in the denoising UNet is modified into a reference attention layer that accepts the feature information (i.e., the facial feature information and the background feature information) from the mask multi-reference image identity guide and then iteratively denoises. To make the face image have a specific motion sequence, a mask-aware motion guide is used to encode the mask motion sequence. Then the encoded features are connected with the noise (such as the noise vector described above). The inter-frame consistency of the generated target synthetic video is ensured, and a timing attention layer is added after the reference attention layer in the denoising UNet. This addition helps to maintain the temporal consistency of the video frames.
[0243] Mask multi-reference image identity guide: the face image is input into the mask multi-reference image identity guide, and then the output features of each self-attention layer in the multi-reference image identity guide are saved. In the denoising process of the denoising UNet, the reference attention layer receives the features of each corresponding self-attention layer (hereinafter referred to as the corresponding layer) in the mask multi-reference image identity guide. The reference attention layer calculates the cross-attention between the feature map (such as obtained based on the noise-added feature information of the video frame and the mask motion sequence) in the denoising UNet and the feature map (such as obtained based on the facial feature information) provided by the corresponding layer of the mask multi-reference image identity guide. It should be noted that a single image, a single image plus another image containing only facial information, or two images containing the same facial information can be input into the mask multi-reference image identity guide as needed. In this way, the mask multi-reference image identity guide can adaptively provide image information related to the required face for various different tasks such as single-shot visual-driven talking face video generation, few-shot visual-driven talking face video generation, or low-frame-rate video generation.
[0244] Mask-aware motion guide: It can be a lightweight motion guide that encodes the motion sequence (such as a sequence of video frames) into a feature map, which has the same dimension as the noise input to the denoising UNet (hereinafter referred to as input noise). Then we add the feature map to the input noise. In order to make the model compatible with both frame-by-frame and keyframe driving modes, a random mask condition layer is combined with the motion guide to form a mask-aware motion guide. The frame-by-frame driving method can produce more accurate motion transfer effects, while the keyframe driving method can model more natural facial motion effects, which are closely related to individual characteristics. The random mask condition layer changes the binary mask feature sequence into the same dimension as the encoded motion sequence, thereby generating a mask motion sequence. The binary mask feature sequence acts as a mechanism for selectively retaining or suppressing information in the motion sequence. As the mask ratio increases, it can also transition the network to low-frame-rate video interpolation, high-frame-rate video compression transmission reconstruction, and other tasks.
[0245] Temporal attention layer: For feature map x, we first adjust its shape to: Then apply temporal attention, which involves self-attention along the f dimension. Here, b represents the batch size, f is the number of generated frames, h and w represent the spatial dimensions of the feature map, and c is the feature dimension. The result is integrated into the original feature using a residual connection from the inter-frame smoothing layer. This layer is inserted at the end of each basic block and only used within the denoising UNet basic blocks, not within the UNet basic blocks of the mask multi-reference map identity guide.
[0246] Vision-driven speaker face generation inference process: In the same identity scenario, there is no face shape difference between the face image and the face in the video frame sequence, so we directly extract the key points from the video frame sequence for driving. However, in the cross-identity scenario, different identities have different facial shapes, and we use Deep3DFaceReconstruction (a deep learning-based face 3D reconstruction technology) to extract 3DMM parameters from the face image and the video frame sequence. Then, the reference identity feature information from the face image is combined with the head pose feature information and the facial expression feature information from the video frame sequence. Subsequently, the parameters are rendered into an RGB image (an image including a red channel, a blue channel, and a green channel), and key points are extracted for cross-identity driving.
[0247] The embodiments of the present application aim to make the visual-driven speaker face generation more controllable and customized. In view of the problem of missing face detail information in a single reference image, the present application designs a multi-reference image strategy combined with an image mask mechanism, so that users can select the form of reference image according to actual needs, such as selecting only a single picture, or selecting a single picture and an additional picture containing only the face area (without background), or selecting two complete pictures containing the same face and the same background as the reference. The mask multi-reference image identity guide can adaptively extract the information needed for video generation from the additional reference image, so that users can improve the insufficient details in the generated results as needed.
[0248] In addition, in view of the problem of inflexible driving mode, the present application designs a mask-aware motion guide. This strategy discards part of the driving frames randomly, so that the model can better model the natural face movement mode, and the generated result is more natural. When finally using the present application, users can select a more accurate driving mode (frame-by-frame driving) or a more natural driving mode (keyframe driving) according to needs.
[0249] In summary, the embodiments of the present application make the visual-driven speaker face generation method more controllable, specifically making the appearance expandable and the action customizable.
[0250] The embodiments of the present application accept one or two reference images (i.e. the aforementioned face images) containing the same face and a sequence of video frames containing the face as input, drive the face in the reference image according to the face action and expression in the video frame sequence, and generate a video of the speaker face. Since this framework supports appearance expansion and action customization, and uses a diffusion model as the backbone network, it guarantees the quality of the generated video while improving the user's freedom in actual use, allowing users to enrich the details of the face in the generated video and freely switch between more accurate frame-by-frame driving and more natural keyframe driving according to needs.
[0251] Figures 29 and 30 respectively show the results of the qualitative experiments of the embodiments of the present application on the HDTF (High-resolution Talking Face Dataset) dataset. The method of the embodiments of the present application outperforms other methods in terms of the visual quality of the generated videos. Although StyleHEAT (Style-Based High-Resolution Editable Talking Face, a pre-trained StyleGAN-based high-resolution editable talking face generation framework) can generate 1024*1024 resolution videos, it lacks rich details and has many distortions. In Figure 29, the case where the reference face and the driven face identity are the same is shown, and most methods can accurately complete the transfer of facial expressions and head poses. As shown in Figure 30, when the reference face and the driven face identity are different, only the method of the embodiments of the present application can still accurately complete the transfer of facial expressions and head poses, and other methods can only generate satisfactory results in a small number of cases. Among them, Reference is the reference image, DaGAN (Deep Attention Generative Adversarial Networks), StyleHEAT, MCNET (Multi-Scale Context-Attention Network), and DPE (Disentanglement of Pose and Expression for General Video Portrait Editing) are comparative algorithms of the embodiments of the present application.
[0252] Tables 1 and 2 respectively show the quantitative experimental results of the embodiments of the present application on the HDTF dataset, in which the best results are shown in bold. At present, seven evaluation indicators commonly used in the field of visually driven talking face generation: FID, PSNR, SSIM, LPIPS, CSIM, APD and AED. Among the seven quantitative indicators, FID, PSNR, SSIM and LPIPS are four indicators reflecting the quality of the generated video, CSIM reflects the similarity of the face identity in the generated video and the face identity in the reference image, APD reflects the accuracy of head pose transfer, and AED reflects the accuracy of facial expression transfer. The results in Tables 1 and 2 show that the embodiments of the present application outperform all methods in APD and AED to achieve the best performance, and our method achieves optimal or competitive results in the four indicators that measure the quality of the generated video.
[0253] Table 1
[0254] Table 2
[0255] In another implementation scenario provided by the embodiments of the present application, the generated synthesized video may have a problem of unsmooth pictures or visual defects due to insufficient frame rate. Therefore, the embodiments of the present application provide a video generation method, which generates an inserted frame from an original video and inserts the inserted frame into the original video to solve the problem of lag of the synthesized video, improve the continuity of the video, make the action of the face in the video more coherent and smooth, eliminate the possible lag, and compensate for the unnatural and delicate performance of the face when the frame rate of the original video is low.
[0256] In one optional embodiment of the video generation method provided by the embodiments of the present application corresponding to FIG. 5, referring to FIG. 31, the video generation method further includes steps S310 to S350. Specifically:
[0257] S310, obtaining a face video.
[0258] The face video includes N face images of video frames, and N is an integer greater than 2.
[0259] It can be understood that the user can upload the face video through the client, and the server receives the face video sent by the client, which includes the feature information of the face. Exemplarily, from the visual point of view, the face video is a two-dimensional or three-dimensional presentation of a person's face, which includes the shape, position, proportion of the features such as the five organs (eyes, nose, mouth, ears, etc.), and the intuitive features such as the facial contour and skin. From the information point of view, the face image contains rich personal identity information, which has high uniqueness, and the face of each person is unique, which can be used for identity recognition and verification. The server receives the face video sent by the client, which includes N face images of video frames.
[0260] S320, extracting video face feature information from the face video.
[0261] For example, the video face feature information is extracted from any video frame face image in the N face images of video frames of the face video. The video face feature information is used to describe the facial features of the face included in the face video.
[0262] After obtaining any video frame face image, a specific algorithm and technology are used to extract the facial feature information. These feature information may include facial contour, position and shape of the five organs, and other key features, which will serve as reference information and play an important role in the subsequent rendering process.
[0263] Video face facial feature information generally includes but is not limited to the following categories: (1) facial features: such as the shape, size and position of the eyes, the shape and direction of the eyebrows, the outline of the nose, the shape and size of the mouth, etc. These features are very important for identifying and distinguishing different faces. (2) Facial contour features: including the overall shape of the face, such as round, square or other shapes, and the width of the cheeks, the shape of the chin, etc. (3) Facial texture features: such as skin texture, wrinkles, etc., which can provide information about age, expression, etc.
[0264] The process of extracting video face facial feature information may use some machine learning or deep learning models that have been trained on a large amount of data and can accurately identify and extract various facial features. By quantifying and representing these features, a face image can be converted into a set of feature vectors or feature descriptors for subsequent analysis, comparison or synthesis, etc. The accuracy and completeness of this step are crucial to the quality and effectiveness of the entire video generation process, and it provides the basis for subsequent rendering and generating target videos based on these features.
[0265] S330, extracting key point features from N video frame face images to obtain N key point feature information.
[0266] Among them, the N key point feature information is used to represent the facial movement information when speaking. The key point feature information describes the information of the key points in the face.
[0267] It can be understood that extracting key point feature information from multiple video frame face images is a very important step. These key point feature information usually represents the positions that have significant changes in the movement of the face, such as the corners of the eyes, the corners of the mouth, the tip of the nose, etc. By extracting these key point feature information, we can obtain the specific movement information of the face when speaking or performing other facial actions. These key point feature information can provide accurate guidance for subsequent rendering, helping to determine how each part should move and change in the newly generated predicted frame.
[0268] As a possible implementation, key point features can be extracted from a video frame face image to obtain key point feature information for the video frame face image, and key point features can be extracted from each of the N video frame face images to obtain key point feature information corresponding to each video frame face image, and thus N key point feature information is obtained.
[0269] S340, rendering according to video face facial feature information and N key point feature information to obtain N predicted frame images.
[0270] The predicted frame image is an image rendered based on the video face feature information and the key point feature information.
[0271] It can be understood that the video face feature information and the key point feature information are used for rendering to construct a new predicted frame image based on the existing feature data. By analyzing the face feature information, the basic morphology and features of the face can be determined, and the key point feature information indicates the direction and amplitude of the face movement. The rendering process is to generate a new image that conforms to the movement rules of the face according to these information. This process may involve complex algorithms and models to ensure that the generated predicted frame image conforms to the features of the original face and accurately reflects the changes caused by movement.
[0272] As a possible implementation, the video face feature information and one key point feature information can be used for rendering to obtain a predicted frame image, so that the video face feature information and N key point feature information can be used for rendering to obtain N predicted frame images.
[0273] S350, fuse the N predicted frame images with the face video to obtain a predicted video.
[0274] As a possible implementation, the N predicted frame images have a time sequence relationship, and the time sequence relationship is determined based on the N video frame face images, so that the jth predicted frame image and the jth video frame face image in the face video can be fused to obtain the jth video frame image in the predicted video, j is a positive integer less than or equal to N.
[0275] It can be understood that the generated N predicted frame images are fused with the face video, so that the newly generated predicted frame can be naturally integrated into the original video. In this way, the frame rate and smoothness of the video can be increased without damaging the overall structure and content of the original video. The fusion process needs to consider factors such as color, lighting, transparency, etc. to ensure that the predicted video has no visual sense of strangeness and realizes natural transition. In this way, the possible stuttering problem of the synthesized video can be effectively solved, the coherence of the video is improved, the movement of the face is more coherent and smooth, and the unnatural and delicate performance of the face when the frame rate of the original video is low can be compensated. For example, in a video of a person speaking, after these steps, the mouth movement of the person speaking will be more natural and smooth, and the expression will be more lively.
[0276] For ease of understanding, please refer to FIG. 32, which is the fifth schematic diagram of the video generation method provided by the embodiments of the present application. First, a face video including N face frame images is obtained. Then, video face feature information is extracted from the face video, and key point features representing face motion information when speaking are extracted from the N face frame images to obtain N key point feature information. Then, the video face feature information and the N key point feature information are rendered to obtain N predicted frame images. Finally, the N predicted frame images are fused with the face video to obtain a predicted video. The method provided by the embodiments of the present application enables the audience to obtain a more comfortable and natural visual experience when watching face-related videos, making the facial expressions and movements more realistic and credible, and helping to improve the quality and appeal of video content. For example, in virtual reality and augmented reality applications, smoother face videos can bring better immersion.
[0277] For ease of understanding, please refer to FIG. 33, which is a schematic diagram of video interpolation reasoning of a speaker's face provided by the embodiments of the present application. When performing video interpolation on a face video, first, key points are extracted from the face frame images of the face video, and the obtained key point feature information is used as motion conditions. Then, the video face feature information is paired with its corresponding key point feature information. These face frame images are simultaneously input into the mask multi-reference image identity guide as reference images. At the same time, the key point feature information is input into the mask perception motion guide as motion condition information of two frames. Then, the denoising network predicts frame images based on the output of the mask multi-reference image identity guide and the mask perception motion guide. Finally, the predicted frame images are fused with the face video to obtain a predicted video.
[0278] The video generation method provided by the application first captures and analyzes the unique features of the face by obtaining a face video and extracting rich facial feature information from it, which lays the foundation for subsequent generation of high-quality and personalized video content. Second, the key point feature information is extracted and used to represent the facial movement information when speaking, so that the generated video can more realistically reflect the natural state of the character when speaking or performing other facial actions, enhancing the realism and credibility of the video. Third, the extracted video face feature information and key point feature information are used to render predicted frame images. This approach can flexibly generate new video content according to actual needs and specific scenarios, with strong adaptability and creativity. Then, the predicted frame images are fused with the face video, which not only improves the frame rate and smoothness of the video, solves the possible stuttering problem, but also makes the video more natural and coherent visually, providing a better viewing experience for the audience. Finally, the comprehensive application of these steps can make video generation more intelligent and automated, reducing the need for human intervention, improving work efficiency, and providing strong technical support for various face video processing application scenarios such as virtual avatar generation and video special effect production, expanding the possibilities and application range of video processing.
[0279] In one optional embodiment of the video generation method provided by the embodiment corresponding to Figure 31 of the present application, please refer to Figure 34, step S350 further includes sub-step S351 to sub-step S352. Specifically:
[0280] S351, determining a first video frame face image and a second video frame face image from the N video frame face images.
[0281] Among them, the first video frame face image is the first frame in the face video, and the second video frame face image is the last frame in the face video.
[0282] It can be understood that the first video frame face image (starting frame) and the second video frame face image (ending frame) with special position significance are determined from the numerous video frame face images of the face video. This step provides a key starting and ending point reference for subsequent video synthesis. By clearly defining these two boundary frames, it can ensure that the synthesized video has a clear definition in time sequence, avoiding confusion.
[0283] S352, video synthesis according to the first video frame face image, the second video frame face image and the N predicted frame images, to obtain a predicted video.
[0284] The first video frame face image is taken as the first frame of the predicted video, and the second video frame face image is taken as the last frame of the predicted video, so that the first frame of the predicted video is the first video frame face image, and the last frame of the predicted video is the second video frame face image.
[0285] It can be understood that the video synthesis is performed based on the first video frame face image, the second video frame face image and the generated N predicted frame images determined in the foregoing. The first video frame face image is taken as the first frame of the predicted video, so that the starting state of the newly synthesized video is consistent with the starting state of the original face video, and the consistency is ensured. Similarly, the second video frame face image is taken as the last frame, so that the ending state is consistent. In this way, the predicted video after synthesis has good connection with the original video at the beginning and the end, and the entire video looks more natural and smooth.
[0286] The video generation method provided in the application ensures the integrity and consistency of video synthesis, so that the newly generated predicted video has a clear starting and ending point on the time axis, and is better integrated with the original video. The characteristics of the beginning and the end of the video are maintained, so that the performance of the synthesized video at the beginning and the end is consistent with the original video, and there is no abrupt change. This is conducive to improving the quality and effect of video synthesis, providing a better visual experience for the audience, and avoiding unnatural situations at the beginning or end of the video due to synthesis. A standardized and orderly method is provided for video processing, facilitating efficient video synthesis operations in various application scenarios.
[0287] In one optional embodiment of the video generation method provided in the embodiment corresponding to FIG. 31 of the application, referring to FIG. 35, step S350 further includes sub-step S3501 to sub-step S3502. Specifically:
[0288] S3501, determining N frame insertion positions according to the face video.
[0289] The frame insertion position is a position in the face video for inserting a predicted frame image.
[0290] It can be understood that the N frame insertion positions are determined according to the face video. This means that the time sequence of the original face video needs to be analyzed and planned to find suitable points to insert the newly generated predicted frame images. This accurate determination of the frame insertion position is to better adapt to the video content and rhythm, and to ensure that the inserted predicted frame can be naturally integrated into the video stream without appearing discordant or unreasonable.
[0291] S3502, inserting N predicted frame images into the face video according to the N frame insertion positions to obtain a predicted video.
[0292] It can be understood that according to the determined N interpolation positions, the corresponding N prediction frame images are accurately inserted into the face video. Wherein, each interpolation position is inserted with one of the prediction frame images, in this way, the details and fluency of the video can be increased without destroying the overall structure and logic of the original video. These inserted prediction frames can supplement the missing intermediate states in the original video, making the facial movements and other performances of the characters more delicate and coherent.
[0293] The video generation method provided by the present application improves the fluency and naturalness of the video, allowing the audience to experience more delicate and realistic facial dynamics. The interpolation positions and number can be flexibly adjusted according to specific needs, adapting to different video styles and effect requirements. It helps to enrich the content of the video, increases more detailed performance by inserting prediction frames, and improves the quality of the video. It provides more possibilities for video post-processing and optimization, and can be personalized according to actual conditions. The synthesized prediction video is more in line with people's expectations for high-quality videos, enhancing the visual experience.
[0294] In one optional embodiment of the video generation method provided by the embodiment corresponding to Figure 31 of the present application, please refer to Figure 36, step S350 further includes sub-step S3511 to sub-step S3512. Specifically:
[0295] S3511, obtaining a target frame rate.
[0296] The target frame rate is the frame rate set for the prediction video.
[0297] It can be understood that the target frame rate is obtained in order to determine the desired frame rate when the prediction frame image is fused with the face video. The selection of the target frame rate usually depends on the specific application requirements and scenarios. Higher frame rate can bring more fluent video effect, but it will also increase the requirements of data processing and storage.
[0298] S3512, according to the target frame rate, fuse the N prediction frame images with the face video to obtain a prediction video.
[0299] It can be understood that the N prediction frame images are fused with the face video according to the target frame rate. This process is to combine the prediction frame images with the original face video according to the requirements of the target frame rate to generate the final prediction video. By fusing the prediction frame images, the frame rate of the video can be increased, making it more fluent and reducing the sense of lag.
[0300] The video generation method provided in the application can increase the frame rate to make the video appear smoother, reduce the sense of jumping and discontinuity of the picture, and improve the viewing experience. For some applications with high requirements for video quality, such as film and television production, virtual reality, etc., high frame rate can provide more realistic visual effects. The appropriate target frame rate can be selected according to the specific application scenario and device performance to balance the video quality and processing efficiency. Smooth video can better attract the attention of the audience, improve the participation and satisfaction of users to the video content.
[0301] FIG. 37 shows the results of the video interpolation test of the embodiment of the application on the HDTF dataset. The first and last frames of the video are input, the middle video frames are predicted, and the longest missing 10 frames of the video are supported. FIG. 37 shows that the embodiment of the application has good video interpolation generation effect.
[0302] The video generation device in the application will be described in detail below. Please refer to FIG. 38. FIG. 38 is a schematic diagram of one embodiment of the video generation device 10 in the embodiment of the application. The video generation device 10 includes a data acquisition module 110, a facial information extraction module 120, a posture and expression extraction module 130, a feature rendering module 140, and a video synthesis module 150. Specifically:
[0303] The data acquisition module 110 is configured to acquire a face image and a video frame sequence, wherein the video frame sequence includes face motion information when speaking, and the video frame sequence includes M video frame images, M being an integer greater than 1;
[0304] The facial information extraction module 120 is configured to extract facial feature information from the face image;
[0305] The posture and expression extraction module 130 is configured to extract head posture features and facial expression features of the face when speaking from the M video frame images, to obtain M head posture feature information and M facial expression feature information;
[0306] The feature rendering module 140 is configured to render according to the facial feature information, the M head posture feature information, and the M facial expression feature information, to obtain M target frame images;
[0307] The video synthesis module 150 is configured to generate a target synthesis video according to the M target frame images.
[0308] The video generation apparatus provided in the application ensures that rich original data is obtained, including static features of the face and dynamic facial motion information, lays a foundation for subsequent generation of high-quality and diversified videos, and the continuous video frame sequence can capture subtle changes in the face when speaking, making the generated video more realistic and natural. By extracting facial feature information from the face image, the facial feature information is more accurate, providing a key reference for subsequent rendering and synthesis, ensuring that the generated video is consistent with the basic features of the face image, helping to improve the accuracy and efficiency of video generation, and avoiding unnecessary deviations and errors. By extracting the head posture features and facial expression features of the face when speaking from the video frame image, dynamic information about the face when speaking is obtained, so that the generated video can accurately reflect the emotional state and intention of the speaker, enhance the expressiveness and appeal of the video, and provide key dynamic feature data for generating more realistic and natural face videos. According to the extracted features, the video frames are adjusted and optimized, so that the generated video is more consistent with the user's expectations and needs, improves the quality and visual effect of the video, and shows more lively and natural posture and expression of the face when speaking. The generated video maintains the consistency of the face features, which is helpful for identity recognition and ensures the credibility of the video. Through the video generation apparatus provided in the application, the generation of high-quality, personalized, realistic and natural target synthesized videos from original face images and video frame sequences can be realized, meeting the needs of users in various application scenarios, such as entertainment, education, communication, etc.
[0309] Thus, the application provides a video generation apparatus. First, a face image and a video frame sequence including face motion information when speaking are acquired. The video frame sequence can reflect subtle movements and expression changes of the face over time during speaking, thereby providing more abundant original data and being capable of supplementing missing face details of the single face image. Then, face feature information is extracted from the face image to provide identity information of the face, and head pose features and facial expression features of the face when speaking are extracted from each video frame image in the video frame sequence to obtain head pose feature information of each video frame image and facial expression feature information of each video frame image, so as to provide dynamic information about the face when speaking. Then, rendering is performed according to the face feature information, the head pose feature information and the facial expression feature information to obtain a target frame image. That is, the face feature information is adjusted and optimized through the head pose feature information and the facial expression feature information, so that the face features in the target synthetic video generated according to the target frame image are the same as the face features in the face image, thereby maintaining identity consistency and being capable of controlling the face to make the same movements and expressions as the video frame sequence, and improving the quality of the target synthetic video. The method provided in the embodiments of the application supplements the missing dynamic information such as the head pose feature information and the facial expression feature information in the speaking process for the static face image by introducing the video frame sequence having the face motion information when speaking, thereby improving the richness of the data required for generating the target synthetic video and improving the quality of the target synthetic video. Furthermore, the realness and attractiveness of the person when speaking in the target synthetic video are improved, and the generated video is more vivid and interesting.
[0310] In one optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the application, referring to FIG. 38, the face information extraction module 120 is further configured to:
[0311] perform feature sampling on the face image to obtain a sampling image;
[0312] perform feature splicing on the image features of the face image and the image features of the sampling image to obtain face image splicing features;
[0313] perform feature processing on the face image splicing features through a first attention mechanism to obtain the face feature information.
[0314] The video generation apparatus provided in the application can focus attention on the key parts or representative feature areas of the face by sampling the face image, which helps to better capture important features of the face. The overall features of the face image and the local features of the sampled image are spliced to obtain face image splicing features, which can comprehensively consider global and local information, making the generated facial feature information more comprehensive and accurate. The first attention mechanism can selectively focus on important features to improve the effectiveness and accuracy of the features, thereby generating facial feature information that is more targeted and of higher quality. The high-quality facial feature information provides a good input for the subsequent video generation step, which helps to generate more natural, realistic and lively speaker face videos.
[0315] In one optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the application, referring to FIG. 38, the face information extraction module 120 is further used for:
[0316] encoding the face image splicing features through a semantic feature encoder to obtain a semantic feature encoding vector;
[0317] and encoding the face image splicing features through an image feature encoder to obtain an image feature encoding vector;
[0318] The face information extraction module 120 is further used for performing feature processing on the encoding features through the first attention mechanism to obtain the facial feature information, wherein the encoding features include the semantic feature encoding vector and the image feature encoding vector.
[0319] The video generation apparatus provided in the application can obtain multi-dimensional feature representation of the face image by using different encoders (semantic feature encoder and image feature encoder). In this way, the features of the face can be described more comprehensively, including semantic information and image features, improving the accuracy and effect of subsequent processing. The semantic feature encoder can extract semantic information of the face image, such as expression, emotion, etc. This helps better understand the meaning and intention of the face, which is very helpful for applications involving emotion analysis, human-computer interaction, etc. The image feature encoder focuses on capturing image features of the face image, such as texture, shape, etc. This helps to retain the detailed information of the image, making the generated video more realistic and natural. The semantic feature encoding vector and the image feature encoding vector are fused, which can comprehensively utilize the advantages of the two kinds of features. This helps better control and generate the expression, posture, etc. of the face in the subsequent video generation process, improving the quality of the generated video. Using different encoders can be flexibly selected and adjusted according to specific needs and application scenarios. For example, in some cases, semantic features may be more important, while in other cases, image features may be more critical. This flexibility enables the apparatus to adapt to different tasks and requirements. By introducing multi-dimensional feature representation and fusion mechanism, the performance and generalization ability of the whole model can be improved. The model can better handle different face image and video generation tasks, thus producing more accurate and natural results.
[0320] In one optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the application, the first attention mechanism is composed of K self-attention layers and K cross-attention layers in cross combination, K being an integer greater than 1.
[0321] Referring to FIG. 38, the face information extraction module 120 is further configured to: input the semantic feature encoding vector as input of the K cross-attention layers, and input the image feature encoding vector as input of the first attention mechanism; and perform feature processing on the semantic feature encoding vector and the image feature encoding vector through the first attention mechanism to obtain the face feature information.
[0322] The video generation apparatus provided in the application can realize the fusion of semantic features and image features by taking the semantic feature encoding vector as the input of the cross-attention layer and taking the image feature encoding vector as the input of the first attention mechanism. This helps to comprehensively consider the semantic features (such as expressions, emotions, etc.) and image features (such as texture, shape, etc.) of the face, so as to more comprehensively describe the features of the face. The attention mechanism can weight the features according to the correlation between the features, so as to highlight important feature information and suppress unimportant information. This helps to improve the accuracy of feature representation, so that the subsequent processing (such as video generation) is more accurate and natural. The cross-attention layer can capture the dependency between different positions in the input sequence, so as to better process the long-distance information in the face image. This is very important for accurately simulating the expression and posture changes of the face. By fusing semantic and image features and using the attention mechanism for feature processing, the model can learn more general and representative feature representations. This helps to improve the generalization ability of the model, so that it can better process new and unseen face images and videos. Accurate facial feature information is crucial for generating high-quality target synthesis videos.
[0323] In one optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the application, referring to FIG. 38, the pose-expression extraction module 130 is further configured to:
[0324] obtain a mask feature sequence, wherein the mask feature sequence includes M mask features;
[0325] perform feature extraction on the M video frame images to obtain M video frame image features;
[0326] generate a mask motion sequence according to the mask feature sequence and the M video frame image features, wherein the mask motion sequence includes M mask motion vectors, and each mask motion vector includes head pose feature information and facial expression feature information.
[0327] The video generation apparatus provided in the application can process irregular parts or irrelevant information in the data set through the mask feature sequence, and improve the attention of the model (i.e., the mask-aware motion guide) to key information. For video frame sequences of different lengths, the mask feature sequence can help the model to process variable-length data, so as to ensure that the model can effectively process various sequence lengths. Using the mask feature sequence can introduce regularization to reduce the over-reliance of the model on specific data, thereby improving the generalization performance of the model. By focusing on extracting head pose and facial expression features in the video frame images, the mask motion sequence can better capture the motion information of the face and improve the quality of the target synthesis video.
[0328] In an optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the present application, referring to FIG. 38, the feature rendering module 140 is further configured to:
[0329] perform noise adding processing on the M video frame images to obtain video frame noise-added feature information;
[0330] perform feature processing on the video frame noise-added feature information, the face feature information, the M head pose feature information, and the M facial expression feature information through the second attention mechanism to obtain M attention feature information;
[0331] decode the M attention feature information to obtain M target frame images.
[0332] The video generation apparatus provided in the present application adds noise to the M video frame images, so that the model learns more robust feature representation and improves the generalization ability of the model. By using the second attention mechanism for feature processing, the model's attention and extraction ability for key features can be improved, so as to generate more accurate and representative target frame images. By decoding the attention feature information, the conversion from features to images can be realized, and target frame images with specific features and contents can be generated.
[0333] In an optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the present application, the second attention mechanism is composed of L self-attention layers, L reference attention layers, and L time-series attention layers, and L is an integer greater than 1;
[0334] Referring to FIG. 38, the feature rendering module 140 is further configured to: input the face feature information as input of the L reference attention layers, input the video frame noise-added feature information, the M head pose feature information, and the M facial expression feature information as input of the second attention mechanism, and perform feature processing on the video frame noise-added feature information, the face feature information, the M head pose feature information, and the M facial expression feature information through the second attention mechanism to obtain M attention feature information.
[0335] The video generation apparatus provided in the present application combines different types of attention mechanisms, which can capture various feature relationships in the video frame images, so as to obtain more rich and accurate feature representation. The introduction of the time-series attention layer can consider the information in the time dimension and better capture the dynamic changes in the video. The reference attention layer uses the face feature information as a reference, which can guide the model to pay attention to other related features and improve the accuracy of feature extraction. The design and use of the second attention mechanism can improve the understanding and processing ability of the model for the video frame images, so as to generate more accurate and meaningful target frame images.
[0336] In an optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the present application, referring to FIG. 38, the feature rendering module 140 is further configured to:
[0337] encode the M video frame images through the image feature encoder to obtain video frame image encoded feature vectors;
[0338] obtain a noise vector;
[0339] splice the video frame image encoded feature vectors and the noise vector to obtain video frame noise-added feature information.
[0340] The video generation apparatus provided in the present application can represent video frame images in a more concise and representative manner through encoding processing, reduce redundant information, facilitate model learning and processing, and lay a foundation for subsequent fusion and analysis with other features. Through noise addition, the model is prevented from over-relying on specific patterns, the generalization ability of the model is enhanced, and the model can better cope with various situations.
[0341] In an optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the present application, referring to FIG. 38, the pose expression extraction module 130 is further configured to:
[0342] extract identity features from the face image to obtain reference identity feature information;
[0343] and extract identity features from any of the video frame images to obtain driving identity feature information;
[0344] in a case where the identity features indicated by the reference identity feature information are the same as the identity features indicated by the driving identity feature information, perform the steps of extracting head pose features and facial expression features of the face when speaking from the M video frame images to obtain M head pose feature information and M facial expression feature information, and subsequent steps;
[0345] in a case where the identity features indicated by the reference identity feature information are different from the identity features indicated by the driving identity feature information, perform the steps of aligning the face shape of the face feature information and the feature information combination to obtain M video frame images after face shape alignment, extracting head pose features and facial expression features of the face when speaking from the M video frame images after face shape alignment to obtain the M head pose feature information and the M facial expression feature information, and the feature information combination includes the M head pose feature information and the M facial expression feature information.
[0346] The video generation apparatus provided in the application directly performs subsequent processing if the identity feature of the person in the face image is the same as the identity feature of the person in the video frame sequence, improves processing efficiency, can quickly extract and use the head posture and facial expression features of the same person, and makes the generated content more coherent and natural. If the identity feature of the person in the face image is different from the identity feature of the face in the video frame sequence, the face shape is aligned, so that even if the person is different, the fusion and reasonable presentation of the features can be realized as much as possible, the adaptability and flexibility of the method are increased, various different identity situations can be processed, and the application range is expanded.
[0347] In one optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the application, referring to FIG. 38, the feature rendering module 140 is further used for:
[0348] extracting background feature information from the face image;
[0349] rendering according to the background feature information, the facial feature information, the M head posture feature information, and the M facial expression feature information to obtain M target frame images.
[0350] The video generation apparatus provided in the application accurately extracts the facial feature information and the background feature information, provides accurate data basis for subsequent processing, and is helpful for more accurate identification, analysis and rendering operations. At the same time, the face and the background are considered, the image content can be more comprehensively understood, the subsequent processing is more in line with the actual scene, and the deviation caused by one-sided processing is avoided. The rendering is performed by comprehensively considering various feature information, so that the generated target frame image is more real and natural, whether the facial posture and expression or the background environment is more in line with the actual situation. The background is processed in combination with the background feature information, the face and the environment are better integrated, and the coordination and situational feeling of the overall picture are improved.
[0351] In one optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the application, referring to FIG. 38, the video synthesis module 150 is further used for generating a target synthesis video according to the face image and the M target frame images, wherein the first frame of the target synthesis video is the face image.
[0352] The video generation device provided by the application first takes the face image as the first frame of the target synthesized video, which can ensure that the target synthesized video presents the real face specified by the user at the beginning, provides a clear starting point and personalized identification for the entire target synthesized video, and increases the recognition and uniqueness of the target synthesized video. Secondly, the synthesized video is generated in combination with the M target frame images, which realizes the dynamic presentation of the face when speaking. This makes the target synthesized video have a coherent and smooth visual effect, can vividly show various postures and expression changes of the face during speaking, and improves the authenticity and attractiveness of the target synthesized video. Thirdly, the video generation method provided by the application can meet the personalized needs of different users for video style and content. Users can select face images and video templates according to their own preferences, thereby creating unique videos that meet their imagination, and providing great creative freedom and personalization.
[0353] In one optional embodiment of the video generation device provided by the embodiment corresponding to FIG. 38 of the application, please refer to FIG. 38,
[0354] The data acquisition module 110 is also configured to acquire a face video, wherein the face video includes N face images of video frames, and N is an integer greater than 2.
[0355] The face information extraction module 120 is also configured to extract video face feature information from the face video.
[0356] The posture and expression extraction module 130 is also configured to extract key point features from the N face images of video frames to obtain N key point feature information, wherein the N key point feature information is used to represent the face motion information when speaking.
[0357] The feature rendering module 140 is also configured to render the video face feature information and the N key point feature information to obtain N predicted frame images.
[0358] The video synthesis module 150 is also configured to fuse the N predicted frame images and the face video to obtain a predicted video.
[0359] The video generation device provided by the application first captures and analyzes the unique features of the face by obtaining a face video and extracting rich facial feature information from it, which lays the foundation for subsequent generation of high-quality and personalized video content. Second, the key point feature information is extracted and used to represent the facial movement information when speaking, so that the generated video can more realistically reflect the natural state of the character when speaking or performing other facial actions, enhancing the realism and credibility of the video. Third, the predicted frame image is obtained by rendering based on the extracted video face feature information and key point feature information. This approach can flexibly generate new video content according to actual needs and specific scenarios, with strong adaptability and creativity. Then, the predicted frame image is fused with the face video, which not only improves the frame rate and smoothness of the video, solves the possible stuttering problem, but also makes the video more natural and coherent in vision, providing a better viewing experience for the audience. Finally, the comprehensive application of these steps can make the video generation more intelligent and automated, reducing the need for manual intervention, improving work efficiency, and providing strong technical support for various face video processing application scenarios, such as virtual avatar generation, video special effect production, etc., expanding the possibilities and application range of video processing.
[0360] In one optional embodiment of the video generation device provided in the embodiment corresponding to Figure 38 of the present application, referring to Figure 38, the video synthesis module 150 is also used for:
[0361] determining a first video frame face image and a second video frame face image from the N video frame face images, wherein the first video frame face image is the first frame in the face video, and the second video frame face image is the last frame in the face video;
[0362] performing video synthesis according to the first video frame face image, the second video frame face image, and the N predicted frame images to obtain a predicted video, wherein the first frame of the predicted video is the first video frame face image, and the last frame of the predicted video is the second video frame face image.
[0363] The video generation device provided by the application ensures the integrity and coherence of video synthesis, so that the newly generated predicted video has a clear starting and ending point on the time axis, and is better integrated with the original video. The characteristics of the beginning and end of the video are maintained, so that the synthesized video at the beginning and end is consistent with the original video, and there is no abrupt change. This is conducive to improving the quality and effect of video synthesis, providing a better visual experience for the audience, and avoiding unnatural situations at the beginning or end of the video due to synthesis problems. It provides a standardized and orderly device for video processing, facilitating efficient video synthesis operations in various application scenarios.
[0364] In an optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the present application, referring to FIG. 38, the video synthesis module 150 is further configured to:
[0365] determine N frame insertion positions according to the face video;
[0366] insert N predicted frame images into the face video according to the N frame insertion positions to obtain a predicted video, one predicted frame image being inserted into each frame insertion position.
[0367] The video generation apparatus provided in the present application improves the smoothness and naturalness of the video, allowing the audience to experience more delicate and realistic facial dynamics. The frame insertion positions and the number of frames can be flexibly adjusted according to specific requirements, adapting to different video styles and effect requirements. This helps to enrich the content of the video, increases more detailed performance by inserting predicted frames, and improves the quality of the video. It provides more possibilities for post-processing and optimization of the video, and can be personalized according to actual conditions. The synthesized predicted video is more in line with people's expectations for high-quality videos, enhancing the visual experience.
[0368] In an optional embodiment of the video generation apparatus provided in the embodiment corresponding to FIG. 38 of the present application, referring to FIG. 38, the video synthesis module 150 is further configured to:
[0369] obtain a target frame rate;
[0370] fuse the N predicted frame images with the face video according to the target frame rate to obtain a predicted video.
[0371] The video generation apparatus provided in the present application increases the frame rate to make the video appear smoother, reduces the sense of jumping and discontinuity of the picture, and improves the viewing experience. For some applications with high requirements for video quality, such as film and television production, virtual reality, etc., high frame rate can provide more realistic visual effects. The appropriate target frame rate can be selected according to the specific application scenario and device performance to balance the video quality and processing efficiency. Smooth video can better attract the audience's attention, improve the user's participation and satisfaction with the video content.
[0372] FIG. 39 is a schematic diagram of a server structure according to an embodiment of the present application. The server 300 can have a great difference in configuration or performance, and can include one or more central processing units (CPUs) 322 (e.g., one or more processors) and a memory 332, and one or more storage media 330 (e.g., one or more mass storage devices) storing applications 342 or data 344. The memory 332 and the storage media 330 can be temporary or persistent storage. The programs stored in the storage media 330 can include one or more modules (not shown in the figure), each of which can include a series of instructions operated in the server. Further, the central processing unit 322 can be configured to communicate with the storage media 330 and execute the series of instructions operated in the storage media 330 on the server 300.
[0373] The server 300 can also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0374] The steps performed by the server in the above embodiments can be based on the server structure shown in FIG. 39.
[0375] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0376] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0377] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0378] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0379] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the essential part of the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0380] The above, the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A video generation method, comprising: Acquire a face image and a video frame sequence, wherein the video frame sequence includes facial motion information of the face during speech, and the video frame sequence includes M video frame images, where M is an integer greater than 1; Extract facial feature information from the face image; The head posture features and facial expression features of the speaking face are extracted from the M video frame images to obtain M head posture feature information and M facial expression feature information. Rendering is performed based on the facial feature information, the M head pose feature information, and the M facial expression feature information to obtain M target frame images; Generate a target synthetic video based on the M target frame images.
2. The video generation method as described in claim 1, wherein extracting facial feature information from the face image includes: Feature sampling is performed on the face image to obtain a sampled image; The image features of the face image are combined with the image features of the sampled image to obtain the face image stitched features; The facial image splicing features are processed through a first attention mechanism to obtain the facial feature information.
3. The video generation method as described in claim 2, characterized in that, The method further includes: The facial image stitching features are encoded using a semantic feature encoder to obtain a semantic feature encoding vector; Furthermore, the facial image stitching features are encoded using an image feature encoder to obtain an image feature encoding vector; The step of processing the facial image stitching features through a first attention mechanism to obtain the facial feature information includes: The encoded features are processed through the first attention mechanism to obtain the facial feature information, wherein the encoded features include the semantic feature encoding vector and the image feature encoding vector.
4. The video generation method as described in claim 3, characterized in that, The first attention mechanism is composed of K self-attention layers and K cross-attention layers, where K is an integer greater than 1; The step of processing the encoded features through the first attention mechanism to obtain the facial feature information includes: The semantic feature encoding vector is used as the input to the K cross-attention layers, and the image feature encoding vector is used as the input to the first attention mechanism. The first attention mechanism is used to process the semantic feature encoding vector and the image feature encoding vector to obtain facial feature information.
5. The video generation method according to any one of claims 1-4, wherein extracting the head pose features and facial expression features of the speaking face from the M video frame images to obtain M head pose feature information and M facial expression feature information includes: Obtain a mask feature sequence, wherein the mask feature sequence includes M mask features; Feature extraction is performed on the M video frame images to obtain the M video frame image features; A mask motion sequence is generated based on the mask feature sequence and the M video frame image features, wherein the mask motion sequence includes M mask motion vectors, and each mask motion vector includes the head pose feature information and the facial expression feature information.
6. The video generation method according to any one of claims 1-4, wherein rendering based on the facial feature information, the M head pose feature information, and the M facial expression feature information to obtain M target frame images includes: The M video frame images are subjected to noise addition processing to obtain video frame noise addition feature information; The video frame noise feature information, the facial feature information, the M head pose feature information and the M facial expression feature information are processed through a second attention mechanism to obtain M attention feature information. The M attention feature information is decoded to obtain the M target frame images.
7. The video generation method as described in claim 6, wherein the second attention mechanism is composed of L self-attention layers, L reference attention layers and L temporal attention layers, where L is an integer greater than 1; The step involves processing the video frame noise feature information, facial feature information, M head pose feature information, and M facial expression feature information through an attention mechanism to obtain M attention feature information, including: The facial feature information is used as the input to the L reference attention layers. The video frame noise feature information, the M head pose feature information, and the M facial expression feature information are used as the input to the second attention mechanism. The second attention mechanism is used to perform feature processing on the video frame noise feature information, the facial feature information, the M head pose feature information, and the M facial expression feature information to obtain the M attention feature information.
8. The video generation method as described in claim 6, wherein the step of adding noise to the M video frame images to obtain video frame noise feature information includes: The M video frame images are encoded using an image feature encoder to obtain video frame image encoded feature vectors; Obtain the noise vector; The video frame image encoding feature vector is concatenated with the noise vector to obtain the video frame noise-added feature information.
9. The video generation method according to any one of claims 1-8, further comprising, before extracting the head pose features and facial expression features of the speaking face from the M video frame images: Identification features are extracted from the facial image to obtain reference identity feature information; Furthermore, identity features are extracted from any of the video frame images to obtain driving identity feature information; If the identity features indicated by the reference identity feature information are the same as the identity features indicated by the driving identity feature information, the steps of extracting the head pose features and facial expression features of the speaking face from the M video frame images to obtain M head pose feature information and M facial expression feature information, and subsequent steps are performed. When the identity features indicated by the reference identity feature information are different from the identity features indicated by the driving identity feature information, the facial feature information and the feature information combination are aligned to obtain M video frame images after face shape alignment. The head posture features and facial expression features of the face when speaking are extracted from the M video frame images after face shape alignment to obtain the M head posture feature information and the M facial expression feature information. The feature information combination includes the M head posture feature information and the M facial expression feature information.
10. The video generation method according to any one of claims 1-9, wherein extracting facial feature information from the face image includes: Extract the facial feature information and background feature information from the face image; The process of rendering based on the facial feature information, the M head pose feature information, and the M facial expression feature information to obtain M target frame images includes: The M target frame images are obtained by rendering based on the background feature information, the facial feature information, the M head pose feature information, and the M facial expression feature information.
11. The video generation method according to any one of claims 1-10, wherein generating a target composite video based on the M target frame images comprises: The target synthetic video is generated based on the face image and the M target frame images, wherein the first frame of the target synthetic video is the face image.
12. The video generation method according to any one of claims 1-11, the method further comprising: Acquire a face video, wherein the face video includes N video frames of face images, where N is an integer greater than 2; Extract facial feature information of the face from the face video; Key point features are extracted from the face images of the N video frames to obtain N key point feature information, wherein the N key point feature information is used to characterize the facial motion information of the face when speaking; Render the video face feature information and the N key point feature information to obtain N prediction frame images; The N predicted frame images are fused with the face video to obtain the predicted video.
13. The video generation method as described in claim 12, wherein fusing the N predicted frame images with the face video to obtain a predicted video comprises: A first video frame face image and a second video frame face image are determined from the N video frame face images, wherein the first video frame face image is the first frame in the face video, and the second video frame face image is the last frame in the face video; A predicted video is obtained by synthesizing the face image of the first video frame, the face image of the second video frame, and the N predicted frame images, wherein the first frame of the predicted video is the face image of the first video frame, and the last frame of the predicted video is the face image of the second video frame.
14. The video generation method as described in claim 12, wherein fusing the N predicted frame images with the face video to obtain a predicted video comprises: Based on the facial video, determine N frame interpolation positions; Based on the N interpolation positions, the N predicted frame images are inserted into the face video to obtain the predicted video, with one predicted frame image inserted at each interpolation position.
15. The video generation method according to any one of claims 12-14, wherein fusing the N predicted frame images with the face video to obtain a predicted video comprises: Obtain the target frame rate; Based on the target frame rate, the N predicted frame images are fused with the face video to obtain the predicted video.
16. A video generation apparatus, comprising: The data acquisition module is used to acquire face images and video frame sequences, wherein the video frame sequence includes facial motion information of the face when speaking, and the video frame sequence includes M video frame images, where M is an integer greater than 1; A facial information extraction module is used to extract facial feature information from the facial image; The posture and expression extraction module is used to extract the head posture features and facial expression features of the speaking face from the M video frame images, and obtain M head posture feature information and M facial expression feature information. The feature rendering module is used to render M target frame images based on the facial feature information, the M head pose feature information, and the M facial expression feature information. The video synthesis module is used to generate a target synthesized video based on the M target frame images.
17. A computer device, comprising: Memory, processor, and bus system; The memory is used to store programs; The processor is configured to execute a program in the memory, including executing the video generation method as described in any one of claims 1 to 15; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
18. A computer-readable storage medium comprising instructions that, when executed on a computer device, cause a computer to perform the video generation method as claimed in any one of claims 1 to 15.
19. A computer program product comprising a computer program that is executed by a processor using the video generation method as described in any one of claims 1 to 15.
Citation Information
Patent Citations
Synthetic video generation method based on three-dimensional face reconstruction and video key frame optimization
CN113269872A
Face forgery detection system and method based on convolutional neural network
CN116824708A
Speaking face video generation method and device, equipment and storage medium
CN116980697A
Speaking face video generation method and device based on multi-modal information control
CN117456587A
Speaking face video generation method, computer equipment and storage medium
CN117789751A