Video generation method and related apparatus

By dynamically adjusting the mesh model of the reference face and fusing facial motion features using mesh parameter difference sequences, the problem of poor target video quality is solved, and high-quality videos are generated.

WO2025232361A1PCT designated stage Publication Date: 2025-11-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084277
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-07
Filing Date
2025-03-24
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

The target video generated by existing technology is of poor quality, and the fusion of the reference face and the original facial expression is unnatural.

Method used

By acquiring a reference facial image and a sequence of content to be converted, a sequence of mesh parameter differences is determined, and the mesh model of the reference face is adjusted to dynamically regulate facial motion features, thereby generating the target video.

Benefits of technology

It improves the fusion quality of the shape features of the reference face in the target video and the facial motion features of the original face, generating videos with accurate lip movements and smooth visuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025084277_13112025_PF_FP_ABST
    Figure CN2025084277_13112025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are a video generation method and a related apparatus, which may be applied in scenarios such as artificial intelligence, digital humans, virtual humans, games, virtual reality and extended reality. The method comprises: acquiring a reference facial image and a content sequence to be converted; on the basis of the reference facial image, determining a mesh model of a reference face, so as to explicitly embody external features of the reference face; on the basis of the content sequence to be converted, determining a mesh parameter difference value sequence, so as to identify the differences between facial motion features indicated by a plurality of pieces of content to be converted; then, adjusting the mesh model of the reference face by means of the mesh parameter difference value sequence, that is, no longer using fixed facial contour parameters, but performing dynamic adjustment by means of time-varying facial motion features, such that the external features of the reference face in a facial mesh model sequence obtained by means of adjustment are better fused with the facial motion features of an original face; and then, on the basis of the facial mesh model sequence and the reference facial image, generating a target video. Therefore, the quality of the target video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A video generation method and related apparatus

[0001] This application claims priority to Chinese Patent Application No. 2024105572778, filed on May 7, 2024, entitled "A Video Generation Method and Related Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence technology, and in particular to video generation. Background Technology

[0003] With the development of artificial intelligence technology, a video of the reference face speaking (i.e., the target video) can be generated from a reference face image that includes a reference face, thereby enabling live video streaming, video production, and other functions.

[0004] In related technologies, models are generally used to generate target videos. Taking diffusion models as an example, motion coefficients related to the original facial expressions are extracted from the sequence of content to be converted (such as audio or video). The reference facial image and motion coefficients are then input into the diffusion model, which generates a series of image frames to obtain the target video.

[0005] However, the target video generated by the above method is of poor quality, such as unnatural fusion of the expressions of the reference face and the original face. Summary of the Invention

[0006] To address the aforementioned technical problems, this application provides a video generation method and related apparatus for improving the quality of the generated target video.

[0007] The embodiments of this application disclose the following technical solutions:

[0008] On one hand, embodiments of this application provide a video generation method, the method comprising:

[0009] A reference facial image and a sequence of content to be converted are obtained. The reference facial image includes the shape features of a reference face, and the sequence of content to be converted includes multiple content items that change over time. The content items are used to indicate the facial motion features of the original face.

[0010] Based on the reference facial image, determine the mesh model of the reference face;

[0011] Based on the sequence of content to be converted, a grid parameter difference sequence is determined, which is used to identify the differences between facial motion features indicated by multiple contents to be converted;

[0012] The mesh model of the reference face is adjusted by the mesh parameter difference sequence to obtain a facial mesh model sequence;

[0013] A target video is generated based on the facial mesh model sequence and the reference facial image.

[0014] On the other hand, embodiments of this application provide a video generation apparatus, the apparatus comprising: an acquisition unit, a first determination unit, a second determination unit, an adjustment unit, and a generation unit;

[0015] The acquisition unit is used to acquire a reference facial image and a sequence of content to be converted. The reference facial image includes the shape features of a reference face, and the sequence of content to be converted includes multiple content to be converted that changes over time. The content to be converted is used to indicate the facial motion features of the original face.

[0016] The first determining unit is configured to determine the mesh model of the reference face based on the reference face image;

[0017] The second determining unit is configured to determine a grid parameter difference sequence based on the sequence of content to be converted, wherein the grid parameter difference sequence is used to identify the differences between facial motion features indicated by multiple contents to be converted;

[0018] The adjustment unit is used to adjust the mesh model of the reference face through the mesh parameter difference sequence to obtain a facial mesh model sequence;

[0019] The generation unit is used to generate a target video based on the facial mesh model sequence and the reference facial image.

[0020] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory:

[0021] The memory is used to store computer programs and to transfer the computer programs to the processor;

[0022] The processor is configured to execute the methods described above according to instructions in the computer program.

[0023] On the other hand, embodiments of this application provide a computer-readable storage medium for storing a computer program for performing the methods described above.

[0024] On the other hand, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods described above.

[0025] As can be seen from the above technical solution, a reference facial image and a sequence of content to be converted are obtained. The reference facial image includes the shape features of the reference face, and the sequence of content to be converted includes multiple content items that change over time. Each content item is used to indicate the facial motion features of the original face. To generate a target video of the reference face exhibiting the aforementioned facial motion features, a mesh model identifying the reference face can be determined based on the reference facial image. A mesh parameter difference sequence is then determined based on the sequence of content to be converted. Since this mesh parameter difference sequence is used to identify the differences between the facial motion features indicated by the multiple content items, and does not reflect the facial contour parameters corresponding to each content item, the differences identified by this mesh parameter difference sequence are equivalent to removing the facial features of the original face while retaining the facial motion features of different content items. Therefore, after adjusting the mesh model of the reference face using this mesh parameter difference sequence, the resulting facial mesh model sequence not only accurately reflects the facial motion features involved in the sequence of content items to be converted, but also has a higher degree of conformity between the identified facial contours and the reference face, avoiding situations where the facial contours resemble the original face. By dynamically adjusting the mesh model of the reference face based on the facial motion features that change over time, the shape features of the reference face in the adjusted facial mesh model sequence are better integrated with the facial motion features of the original face, thereby improving the quality of the target video. Attached Figure Description

[0026] Figure 1 is a schematic diagram of an application scenario of a video generation method provided in an embodiment of this application;

[0027] Figure 2 is a flowchart illustrating the video generation method provided in an embodiment of this application;

[0028] Figure 3 is a schematic diagram of an application scenario for video generation provided in an embodiment of this application;

[0029] Figure 4 is a schematic diagram of an application scenario for video generation provided in an embodiment of this application;

[0030] Figure 5 is a schematic diagram of a speech conversion module provided in an embodiment of this application;

[0031] Figure 6 is a schematic diagram of a video conversion module provided in an embodiment of this application;

[0032] Figure 7 is a schematic diagram of a video generation module provided in an embodiment of this application;

[0033] Figure 8 is a schematic diagram of a video generation device provided in an embodiment of this application;

[0034] Figure 9 is a schematic diagram of the structure of a server provided in an embodiment of this application;

[0035] Figure 10 is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0036] The embodiments of this application will now be described with reference to the accompanying drawings.

[0037] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “corresponding,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0038] In related technologies, the quality of target videos generated based on models is generally poor. Analysis reveals that regardless of whether the content sequence to be converted is video or audio, it generally carries facial motion features of the original face. For example, videos directly include the original face, while audio, although not directly including the original face, implicitly includes the facial motion features of the face making the sound. For instance, the volume and content of the voice will lead to differences in facial expressions and head movements; if the voice in the audio remains calm, the original face may remain still. However, the models in related technologies use fixed facial contour parameters to generate target videos. If the reference face in the reference image differs significantly from the original face in the content sequence to be converted—that is, if the similarity between the reference face and the original face is low—it may result in an unnatural fusion of the target face generated based on the reference face and the original face, leading to poor quality of the generated target video.

[0039] Based on this, embodiments of this application provide a video generation method and related apparatus, which no longer use fixed facial contour parameters to generate target videos, but dynamically adjusts them through time-varying facial motion features, so that the shape features of the reference face in the adjusted facial mesh model sequence are better integrated with the facial motion features of the original face, thereby improving the quality of the target video generated based on the facial mesh model.

[0040] The video generation method provided in this application can be applied to computer devices with video generation capabilities, such as terminal devices and servers.

[0041] Specifically, terminal devices can be desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Smart in-vehicle devices can be in-vehicle navigation terminals and in-vehicle computers, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc., but are not limited to these.

[0042] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal devices and servers can be connected directly or indirectly via wired or wireless communication; this application does not impose any restrictions on this.

[0043] To facilitate understanding of the video generation method provided in this application embodiment, the following example uses a server as the execution subject of the video generation method to illustrate its application scenarios.

[0044] Referring to Figure 1, this figure is a schematic diagram of an application scenario of a video generation method provided in an embodiment of this application. As shown in Figure 1, this application scenario includes a terminal device 110 and a server 120, and the terminal device 110 and the server 120 can communicate through a communication network. The communication network uses standard communication technologies and / or protocols, typically the Internet, but can also be any network, including but not limited to Bluetooth, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), mobile, private networks, or any combination of virtual private networks. In some embodiments, customized or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.

[0045] For example, a video generation client is installed in terminal device 110. The user uploads a sequence of content to be converted and a reference facial image through the client. The reference facial image includes the shape features of a reference face. The sequence of content to be converted includes multiple pieces of content that change over time, and each piece of content is used to indicate the facial motion features of the original face. As shown in Figure 1, the sequence of content to be converted can be a video. Terminal device 110 sends the sequence of content to be converted and the reference facial image to server 120.

[0046] Server 120 is equipped with a video generation server to provide video generation services. After acquiring a reference facial image and a sequence of content to be converted, server 120 determines a mesh model of the reference face based on the reference facial image. The mesh model of the reference face can explicitly reflect the external features of the reference face. Based on the sequence of content to be converted, a mesh parameter difference sequence is determined. This mesh parameter difference sequence is used to identify the differences between facial motion features indicated by multiple contents to be converted, that is, to reflect the changes in facial motion features of the original face over time during the process of emitting sound.

[0047] Server 120 adjusts the mesh model of the reference face by using a mesh parameter difference sequence. This means adjusting the reference face's mesh model based on the facial motion features of the original face, rather than using fixed facial contour parameters. Instead, it dynamically adjusts the model based on time-varying facial motion features, allowing for better integration of the reference face's shape features with the original face's motion features in the adjusted facial mesh model sequence. Based on the facial mesh model sequence and the reference face image, server 120 generates the target video, thereby improving the quality of the target video.

[0048] The server 120 sends the generated target video to the terminal device 110 so that the target video can be displayed on the terminal device 110.

[0049] The video generation method provided in this application embodiment can be executed by a server. However, in other embodiments of this application, the terminal device may also have similar functions to the server to execute the video generation method provided in this application embodiment, or the terminal device and the server may jointly execute the video generation method provided in this application embodiment. This embodiment does not limit this.

[0050] Furthermore, the video generation method provided in this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, and the Internet of Things. Four specific scenarios are given below as examples.

[0051] Scenario 1: Personalized video generation scenario.

[0052] When a user needs personalized video generation, they can provide a facial image (such as an image of another person's face, a virtual object's face, etc.) and an audio clip or video containing audio through terminal device 110. The given facial image is used as a reference facial image, and the audio or video clip is used as a sequence of content to be converted. Both the reference facial image and the sequence of content to be converted are transmitted to server 120 for processing. Server 120 generates a target video based on the video generation method provided in this application embodiment. In this target video, the user-provided facial image can be used to generate audio, with accurate lip movements and smooth visuals; that is, the target video has high quality and meets the user's need for personalized video.

[0053] Scenario 2: Virtual object live streaming scenario.

[0054] Virtual objects are generally non-real objects. When a virtual object is live-streamed, a facial image of the virtual object and the audio required for the live stream can be provided. The terminal device 110 uses the virtual object's facial image as a reference facial image and the audio as a sequence of content to be converted. Both the reference facial image and the sequence of content to be converted are transmitted to the server 120 for processing. The server 120 generates a target video for live-streaming based on the video generation method provided in this embodiment. In this target video, the virtual object can speak based on the audio, with accurate lip movements, smooth visuals, and even head movements that match the rhythm of the audio. This results in a high-quality target video, improving the live-streaming effect and attracting viewers.

[0055] Scene 3: Film and television production.

[0056] To replace object A with object B in a pre-made video, a facial image of object B and the video segment containing object A can be provided. Terminal device 110 uses the facial image of object B as a reference facial image and the video segment containing object A as the sequence of content to be converted. Both the reference facial image and the sequence of content to be converted are transmitted to server 120 for processing. Server 120 generates a target video based on the video generation method provided in this embodiment. Compared to the video segment containing object A, the target video replaces object A with object B. That is, in this target video, object B replaces object A based on audio pronunciation, with accurate lip movements and smooth visuals. Therefore, the target video has higher quality and meets the user's needs for video face swapping.

[0057] Scene 4: The Game.

[0058] Multiple virtual objects can exist in the game. Audio control can manipulate the lip movements and head motions of these virtual objects during vocalization, making them more lifelike and enhancing the user experience. Taking a singing game as an example, when a user sings, a virtual object is added to their party. The user can specify the virtual object's facial image to match the audio during the singing. The terminal device 110 uses the virtual object's facial image as a reference image and the audio as a sequence to be converted. Both the reference image and the sequence are then transmitted to the server 120 for processing. The server 120 generates a target video of the virtual object singing based on the video generation method provided in this embodiment. In this target video, the virtual object can vocalize based on the audio, with accurate lip movements, smooth visuals, and even head motion matching the audio rhythm. This indicates a high-quality target video, improving game interactivity and attracting user participation.

[0059] The video generation method provided in this application can be applied not only to generating personalized videos, but also to other scenarios requiring video compositing, such as movies, games, and advertisements, and has broad application prospects. By using the video generation method provided in this application, high-quality target videos can be generated quickly, improving the user experience.

[0060] The following describes in detail a video generation method provided in this application through method embodiments.

[0061] Referring to Figure 2, this figure is a schematic flowchart of the video generation method provided in an embodiment of this application. For ease of description, the following embodiments will still use a server as the execution subject of the video generation method as an example. As shown in Figure 2, the video generation method includes S201-S205.

[0062] S201: Obtain the reference facial image and the sequence of content to be converted.

[0063] The reference facial image includes the shape features of the reference face. These shape features refer to characteristics used to describe the shape of the reference face, such as facial contour features and hairstyle features, so that a target video with the reference face's voice can be generated based on the reference face. It should be noted that the face can be a human face, but can also refer to the face of an animal, robot, virtual human, etc.; this application does not specifically limit this.

[0064] The sequence of content to be converted includes multiple pieces of content that change over time; that is, the multiple pieces of content to be converted are arranged in chronological order, such as video and audio. Video includes multiple video frames that change over time, and audio includes multiple sound waveforms that change over time.

[0065] The content to be converted not only provides the audio required in the target video, but also indicates the facial motion features of the original face. Facial motion features refer to the changes that occur on the original face during the process of producing sound, such as facial expressions and head movement trajectories. This allows the content to be converted in the sequence to be merged one by one into the reference face image, thereby obtaining the target video of the reference face producing sound. This ensures that the lip movements of the reference face are consistent with the audio and the head movement trajectories are consistent with the audio rhythm during the sound production process, thus improving the quality of the target video.

[0066] It is understood that all data obtained in this application (such as reference facial images, sequences of content to be converted, etc.) are obtained with the separate consent and authorization of the data subject (such as users, institutions or enterprises), and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of relevant countries and regions, and facial images and other data must be processed in strict accordance with legal requirements and personal information processing rules, and technical measures must be taken to ensure the security of related data.

[0067] S202: Determine the mesh model of the reference face based on the reference face image.

[0068] A reference face mesh model refers to a model of a reference face built based on a mesh, which can describe the shape and features of the reference face. The reference face mesh model is generally a three-dimensional model, which is composed of polygons. Taking a triangle as an example, a complex polygon can be composed of multiple triangles. Therefore, the surface of a three-dimensional model is composed of multiple interconnected triangles. In three-dimensional space, the set of points that make up these triangles and the sets of edges of the triangles constitute the mesh.

[0069] This application does not specifically limit the method of obtaining the mesh model of the reference face; those skilled in the art can set it according to actual needs. The following description uses one method as an example.

[0070] For example, key features such as edges, corners, and textures are extracted from a reference facial image. Then, based on these extracted features, the 3D shape of the reference face is constructed using methods such as stereo vision, structured light, and depth cameras, resulting in a 3D model of the reference face. Finally, after obtaining the 3D model of the reference face, its surface is discretized into a series of triangular faces to obtain a mesh model of the reference face.

[0071] S203: Determine the grid parameter difference sequence based on the sequence of content to be converted.

[0072] The sequence of content to be converted includes multiple pieces of content that change over time, meaning there may be differences between them. Each piece of content is used to indicate the facial motion features of the original face. Therefore, based on the multiple pieces of content included in the sequence, the differences between the facial motion features indicated by two pieces of content can be determined, i.e., a mesh parameter difference sequence. Compared to using a mesh model of the original face included in each piece of content, the mesh parameter difference sequence can more specifically reflect the changes that occur in the original face during vocalization.

[0073] Suppose there are N pieces of content to be converted. For any one of these N pieces of content, such as the i-th piece of content, it can have one difference in the network parameter difference sequence, such as reflecting the difference between its facial motion features and those of the j-th piece of content. Alternatively, it can have at most N-1 differences in the network parameter difference sequence. For example, if there are two differences, the first difference reflects the difference between its facial motion features and those of the j-th piece of content, and the second difference reflects the difference between its facial motion features and those of the k-th piece of content.

[0074] Therefore, the difference for each content to be converted in the grid parameter difference sequence can reflect the difference between the facial motion features of that content and the other N content to be converted. Thus, the entire network parameter difference sequence effectively shows where the differences lie between the facial motion features of different content to be converted.

[0075] Furthermore, since the facial motion features of each content to be converted are represented through interpolation, the common feature of the content to be converted, the "facial contour of the original face," is weakened or even eliminated in the mesh parameter interpolation sequence. This allows the mesh parameter interpolation sequence to accurately identify the time-varying facial motion features carried by the content sequence to be converted, but without highlighting the facial contour features of the original face. In this way, when the mesh model of the reference face is subsequently adjusted using the mesh parameter interpolation sequence, the facial contour features of the original face will not be incorporated into the reference face's mesh model. Consequently, in the final generated target video, the facial contour reflecting the facial motion features identified by the content sequence to be converted will better match the reference face, rather than exhibiting a similar contour to the original face.

[0076] In one possible implementation, the mesh parameter difference sequence includes multiple mesh parameter differences, the number of which matches the number of contents to be converted, meaning there is a one-to-one correspondence between the contents to be converted and the mesh parameter differences. Each mesh parameter difference is used to identify the difference between the mesh model corresponding to the original face in the current contents to be converted and the standard mesh model. The standard mesh model is the mesh model corresponding to the original face in one of the contents to be converted within the sequence. For example, from multiple original faces included in the sequence, the original face with the expression most similar to a reference face is selected, and its mesh model is used as the standard mesh model. Alternatively, the mesh model corresponding to the original face in the previous contents to be converted can be used as the standard mesh model for the original face in the next contents to be converted. This application does not impose specific limitations on this, and those skilled in the art can set it according to actual needs.

[0077] This application does not specifically limit the method of obtaining the grid parameter difference sequence. Three methods are described below as examples.

[0078] Method 1: Based on the sequence of content to be converted, a speech-to-grid sequence model is used to convert it to obtain a grid parameter difference sequence.

[0079] In other words, a speech-to-grid sequence model is trained to establish a mapping relationship between the content sequence to be converted and the grid parameter difference sequence. Taking audio as an example, the audio includes the prosody of the original face emitting the audio. For instance, if the sound in the audio is consistently calm, the original face may remain still. Or, if the sound in the audio gradually decreases, the facial movements of the original face may gradually decrease. Thus, based on the content sequence to be converted, a corresponding grid parameter difference sequence is generated.

[0080] The training process of the speech-to-grid sequence model is explained below.

[0081] Obtain facial vocalization video samples with a first label, which is used to identify the mesh parameter difference sequence corresponding to the original facial samples in the facial vocalization video samples. For example, a video frame can be selected from the facial vocalization video samples, and the mesh model of the original facial samples included in that video frame can be used as the standard mesh model. Then, the difference between the mesh model of the original facial samples in each video frame and the standard mesh model can be calculated to obtain the mesh parameter difference of each video frame, thereby obtaining the mesh parameter difference sequence corresponding to the original facial samples, i.e., the first label.

[0082] Based on facial vocalization video samples, an initial speech-to-grid sequence model is used to convert the speech into a predicted grid parameter difference sequence. The model parameters of the initial speech-to-grid sequence model are adjusted according to the difference between the predicted grid parameter difference sequence and the first label, so that the predicted grid parameter difference sequence gets closer and closer to the first label, thus obtaining the speech-to-grid sequence model.

[0083] The embodiments of this application do not specifically limit the speech-to-grid sequence model, but can be based on the Wav2vec model (an unsupervised pre-training model for speech recognition), etc.

[0084] Method 2: Based on the sequence of content to be converted, the speech-to-grid model is used to convert the content to be converted, and the grid model of each content to be converted in the sequence is obtained. A standard grid model is determined from the multiple grid models, and the difference between the multiple grid models and the standard grid model is calculated to obtain the grid parameters of each content to be converted, thus obtaining the grid parameter difference sequence.

[0085] The training process of the speech-to-mesh model is explained below.

[0086] A facial speech image sample with a second label is obtained, which includes multiple mesh models used to identify the facial speech image sample. Based on the facial speech image sample, it is converted using an initial speech-to-mesh model to obtain multiple predicted mesh models included in the facial speech image sample. Based on the differences between the multiple predicted mesh models and the second label, the model parameters of the initial speech-to-mesh model are adjusted so that the predicted mesh models become closer and closer to the second label, thus obtaining the speech-to-mesh model.

[0087] Method 3: Based on the sequence of content to be converted, determine the mesh model of the original face in each piece of content to be converted. From the mesh models of multiple original faces, determine a standard mesh model, calculate the difference between the mesh models of multiple original faces and the standard mesh, obtain the mesh parameter difference value of each piece of content to be converted, and thus obtain the mesh parameter difference value sequence.

[0088] It should be noted that if the content to be converted is video, and the video includes images of the original face, a mesh model can be directly obtained from the video. In this case, method three, which obtains the mesh parameter difference sequence, is more suitable. If the content to be converted is audio, and the audio does not include images of the original face, a mesh model cannot be directly obtained. In this case, method one or method two is more suitable.

[0089] Furthermore, the embodiments of this application do not specifically limit the method of determining the standard mesh model. It can be selected arbitrarily, or it can be selected based on the similarity between it and the reference facial image, or it can be selected based on the default method. This will be explained in detail later, and will not be repeated here.

[0090] S204: Adjust the mesh model of the reference face by using the mesh parameter difference sequence to obtain the facial mesh model sequence.

[0091] In related technologies, a mesh model for generating the target video is directly extracted from the sequence of content to be converted. However, the facial contours of the faces in the target video generated based on this mesh model are quite similar to the original faces in the content to be converted. If the reference face differs significantly from the original face, the faces in the target video may appear to be identical in appearance to the reference face image, but their facial contours may be identical to the original face. This results in the problem that the faces in the target video are not similar to the reference face, meaning that the quality of the target video is poor.

[0092] Furthermore, in a two-dimensional image, an object only has two dimensions: length and width. In three-dimensional space, an object gains a depth dimension. Therefore, when attempting to map a two-dimensional object to three-dimensional space, there are multiple possible mapping methods, each corresponding to a different position and shape in three-dimensional space. For example, a circle in a two-dimensional image can be represented in many ways when mapped to three-dimensional space. It can be a circle on a plane, a three-dimensional sphere or cylinder, or even a distorted, irregular three-dimensional shape.

[0093] The content to be converted is generally audio or video, and the original face is usually displayed as a two-dimensional image. The mesh parameter difference sequence is generally obtained based on a mesh model, which is a three-dimensional model. This involves mapping two-dimensional objects to three-dimensional space, resulting in multiple mapping methods. It's understandable that while audio doesn't directly display the original face as a two-dimensional image, it is trained based on a two-dimensional image including the audio to discover the mapping relationship between the audio and the original face in the two-dimensional image; that is, audio can implicitly contain a two-dimensional image of the original face.

[0094] In other words, if a grid model is directly obtained from the sequence of content to be converted, and then a target video is generated based on the grid model, the faces in the target video are greatly affected by the original faces. Moreover, due to various mapping methods, there may be problems with the faces in the target video being sometimes large and sometimes small, resulting in low quality of the target video.

[0095] Therefore, this application no longer directly uses the models extracted from the content sequence to be converted. Instead, it fuses the models based on the differences between them with the mesh model of the reference face. That is, it adjusts the mesh model of the reference face based on the mesh parameter difference sequence to obtain a facial mesh model sequence. The facial mesh model sequence includes multiple facial mesh models, each of which is obtained by adjusting the mesh model of the reference face based on the mesh parameter differences. The number of facial mesh models in the facial mesh model sequence is consistent with the number of contents to be converted included in the content sequence to be converted.

[0096] Among them, the mesh parameter difference sequence can specifically reflect the changes of the original face during the speech process, such as changes in expression and head movement trajectory. The mesh model of the reference face is used to represent the shape of the reference face, such as facial contour and hairstyle. Based on the mesh parameter difference sequence, the mesh model of the reference face is adjusted. For example, by directly performing mesh parameter superposition calculation, the changes of the original face during the speech process can be integrated into the reference face while keeping the shape of the reference face unchanged, resulting in a facial mesh model sequence. This is equivalent to obtaining a fusion result based on the changes in facial motion features of the original face during the speech process and the shape features of the reference face. This fusion result is not fixed, thereby improving the quality of the subsequently generated target video.

[0097] S205: Generate the target video based on the facial mesh model sequence and the reference facial image.

[0098] A facial mesh model sequence can reflect facial feature changes, such as shape and expression. A reference facial image can provide texture features, age features, color features, etc. Therefore, based on the facial mesh model sequence and the reference facial image, multiple video frames can be obtained. Each video frame is obtained based on the facial mesh model in the facial mesh model sequence and the reference facial image. Finally, the target video is obtained based on multiple video frames.

[0099] The faces appearing in the target video are reference faces identified by the reference face image. In the target video, the reference face will make facial motion features of the content to be converted based on the characteristics of the time changes identified in the sequence of content to be converted.

[0100] Understandably, the target video can also include audio, meaning it is synthesized from multiple video frames and audio. Thus, in the target video, speech is produced based on the shape of a reference facial image, and the facial movements of the reference image during speech production remain consistent with the sequence of content to be converted, such as lip movements matching the audio, facial expressions matching the audio, and head movement trajectories matching the audio rhythm, thereby improving the quality of the target video.

[0101] As can be seen from the above technical solution, a reference facial image and a sequence of content to be converted are acquired. The reference facial image includes the shape features of the reference face, and the sequence of content to be converted includes multiple content items that change over time. Each content item indicates the facial motion features of the original face. A mesh model of the reference face is determined based on the reference facial image, which explicitly reflects the external features of the reference face. A mesh parameter difference sequence is determined based on the sequence of content to be converted. This sequence identifies the differences between the facial motion features indicated by the multiple content items, reflecting the changes in facial motion features of the original face over time during sound production. The mesh model of the reference face is then adjusted using the mesh parameter difference sequence, i.e., adjusted based on the facial motion features of the original face. This avoids using fixed facial contour parameters and instead dynamically adjusts the mesh model based on the time-varying facial motion features, allowing for better integration of the shape features of the reference face with the facial motion features of the original face in the adjusted mesh model sequence. Finally, a target video is generated based on the facial mesh model sequence and the reference facial image, thereby improving the quality of the target video.

[0102] As can be seen from the foregoing, the embodiments of this application do not specifically limit the selection of standard mesh models. The following will illustrate three methods as examples.

[0103] Method 1: Select a mesh model with similar facial expressions. See A1-A4 for details.

[0104] A1: Determine the original facial mesh model sequence based on the sequence of content to be converted.

[0105] The sequence of original facial mesh models includes multiple original facial mesh models, the number of which corresponds to the number of content items to be converted in the sequence. For example, if the content sequence is video, the mesh models for each original face can be obtained in a similar manner to obtaining the reference facial mesh model; this will not be elaborated further here. Similarly, if the content to be converted is audio, it can be converted using the aforementioned speech-to-mesh model to obtain the mesh models for each piece of content to be converted.

[0106] A2: Determine the similarity between mesh models of multiple original faces and mesh models of reference faces.

[0107] Similarity refers to the degree of similarity between two mesh models. This application does not specifically limit the method for determining similarity. For example, the mesh model can be converted into a point cloud, and then the distance metric between the two point clouds can be calculated to determine the similarity. Alternatively, features of the mesh model, such as volume, surface area, and curvature, can be extracted, and the similarity can be determined based on the differences between the features of the two mesh models.

[0108] As one possible approach, since the goal is to fuse the facial motion features of the original face into the reference face, the expression of the original face's mesh model and the expressions of the reference mesh models can be determined, thus establishing the similarity between the two mesh models. Expressions can be determined through features, expression weight coefficients, etc., thereby calculating only the expression of the mesh model. This not only reduces computational load but also allows for targeted calculation of expressions, improving the accuracy of subsequent fusion and enhancing the quality of the target video.

[0109] A3: Among multiple original facial mesh models, the mesh model whose similarity meets the similarity criteria is determined as the standard mesh model.

[0110] The mesh models in the original facial mesh model sequence that meet the similarity criteria are determined as standard mesh models, meaning the standard mesh model is relatively similar to the reference facial mesh model. This application does not specifically limit the similarity criteria, such as the highest similarity.

[0111] A4: Based on the differences in facial motion features between each mesh model in the original facial mesh model sequence and the standard mesh model, a mesh parameter difference sequence is obtained.

[0112] Therefore, using a standard mesh model with high similarity to the reference face mesh model as a benchmark, the differences between the mesh models of each original face and the standard mesh model in the dimension of facial motion features are calculated. This yields a mesh parameter difference sequence including multiple differences, which is equivalent to obtaining the differences between the mesh models of each original face and the mesh model of the reference face. This allows for more accurate, convenient, and faster subsequent adjustment of the reference face mesh model based on the mesh parameter difference sequence. Furthermore, this method has no restrictions on the reference face image, meaning users can freely choose any reference face image, improving the user experience.

[0113] Method 2: Limit the facial expression of the reference face to the first state, and select a mesh model that matches the first state. See B1-B3 for details.

[0114] B1: Determine the original facial mesh model sequence based on the sequence of content to be converted.

[0115] The sequence of original facial mesh models includes multiple original facial mesh models.

[0116] B2: Among multiple original facial mesh models, the mesh model whose expression conforms to the first state is determined as the standard mesh model.

[0117] It should be noted that in this embodiment, the facial expression of the reference face in the reference facial image is in the first state. As one possible implementation, the first state is an expressionless state, so that the original facial expression is superimposed on the reference facial image in the expressionless state, which has a better effect and results in a higher quality target video.

[0118] This allows us to identify the mesh model in the original facial mesh model sequence that matches the first state as the standard mesh model. This is equivalent to selecting a mesh model from the original facial mesh model sequence that is similar to the reference face as the standard mesh model.

[0119] B3: Based on the differences in facial motion features between each mesh model in the original facial mesh model sequence and the standard mesh model, a mesh parameter difference sequence is obtained.

[0120] If the first state is an expressionless state, then each difference in the mesh parameter difference sequence represents the difference between the mesh model of each original face and the standard mesh model in the expressionless state. This makes it easier to adjust the mesh model of the reference face image in the same expressionless state based on the mesh parameter difference sequence, thereby improving the accuracy of the subsequent target video.

[0121] Therefore, by defining the expression of the reference face as the first state, a mesh model that conforms to the first state can be directly selected from the mesh model sequence of the original face as the standard mesh model. This is equivalent to selecting a mesh model that is more similar to the mesh model of the reference face from multiple mesh models. This not only has the advantages of the aforementioned method one (A1-A4), but also improves the quality of the subsequent target video without the need for similarity calculation, reducing the amount of computation, shortening the generation time of the target video, and improving the user experience.

[0122] Method 3: Do not limit the reference facial expression to the first state; select a mesh model that conforms to the first state. See C1-C5 for details.

[0123] C1: Obtain the pending reference facial image.

[0124] The pending reference facial image is an image specified by the user. It is desired to generate a video of facial sounds included in the pending reference facial image. Compared with the reference facial image in Method 2 (B1-B3), this embodiment does not specifically limit the facial expression state included in the pending reference facial image. For example, if the expression of the reference face in the pending reference facial image is in the second state, that is, the expression of the reference face specified by the user is not limited.

[0125] C2: Based on the first state, adjust the facial motion features of the reference face in the undetermined reference face image to obtain the reference face image.

[0126] The facial motion features of the reference face in the desired reference face image are adjusted to change it from the second state to the first state, thereby obtaining the reference face image.

[0127] The embodiments of this application do not specifically limit the adjustment method. For example, diffusion models can be used. Diffusion models can generate multiple states of the face, such as generating a smiling state based on a blank expression state. Thus, based on a reference face image with a second expression, the diffusion model can be used to obtain an image with a first expression, i.e., a reference face image.

[0128] C3: Determine the original facial mesh model sequence based on the sequence of content to be converted.

[0129] C4: The mesh model whose expression matches the first state in the original facial mesh model sequence is identified as the standard mesh model.

[0130] C5: Based on the differences between each mesh model in the original facial mesh model sequence and the standard mesh model, obtain the mesh parameter difference sequence.

[0131] C3-C5 can be found in the aforementioned B1-B3, and will not be repeated here.

[0132] Therefore, without limiting the expression to the first state, the expression can be adjusted from the second state to the first state. That is, a reference facial image is obtained based on the undetermined reference facial image. Then, a mesh model that conforms to the first state can be directly selected from the mesh model sequence of the original face as the standard mesh model. This is equivalent to selecting a mesh model that is more similar to the mesh model of the reference face from multiple mesh models. It not only has the advantages of the aforementioned method two (B1-B3), but also does not limit the state of the expression, expands the application scenarios, and improves the user experience.

[0133] As one possible implementation, this application embodiment also provides a specific implementation of S205, namely, a specific implementation of generating a target video based on a facial mesh model sequence and a reference facial image. Specifically, the facial mesh model sequence can first be reprojected to obtain a facial pose sequence. Then, the target video is generated based on the facial pose sequence and the reference facial image.

[0134] A facial pose sequence comprises multiple facial poses, each obtained by reprojecting a facial mesh model. Taking a facial mesh model from the sequence as an example, the corresponding facial pose is obtained by reprojecting that model, thus generating the facial pose sequence.

[0135] Reprojection refers to generating a new image, or facial pose image, by projecting a facial mesh model from any viewpoint. Reprojection ensures the accurate mapping of the three-dimensional structure of the face onto a two-dimensional image, thereby enabling more precise acquisition of key facial feature points and contour information, and thus a more accurate description of facial pose (such as rotation angle, offset, etc.).

[0136] Furthermore, facial mesh models are 3D models, while facial poses are 2D images. Converting a 3D model to a 2D image makes it easier to generate corresponding video frames based on the 2D facial pose and reference facial image. Thus, the target video is generated from the facial pose sequence and reference facial image, resulting in a higher quality target video. In addition, facial pose data is more efficient than mesh model data in its description method; that is, facial pose has fewer parameters, resulting in relatively lower computational costs. This helps to accelerate processing speed and reduce hardware requirements.

[0137] Therefore, the three-dimensional facial mesh model is first converted into a two-dimensional facial pose through reprojection, thereby ensuring the geometric accuracy of the face and reducing the amount of computation. Then, based on the facial pose in the facial pose sequence and the shape features of the reference facial image, the target video is obtained, thus improving the generation efficiency and quality of the target video.

[0138] It should be noted that during the process of speaking, not only will facial expressions change, but the head will also change, such as nodding or turning the head to follow the sound. In other words, the head will change with the rhythm of the audio. If you want to improve the quality of the target video, you also need to consider the head movement trajectory.

[0139] In the first scenario, if the content sequence to be converted is a video, since the video directly includes facial images, the grid parameter difference sequence obtained from the video can not only reflect changes in facial expressions but also changes in head movement trajectories. As a result, in the target video, the head movement trajectory will change based on the audio rhythm.

[0140] Scenario 2: If the content sequence to be converted is audio, since audio does not directly include facial images, prediction needs to be made based on the audio. That is, to improve prediction accuracy, facial expressions and head movement trajectories need to be predicted based on the audio. As mentioned above, a speech-to-grid sequence model or speech-to-grid model can be used to obtain a grid parameter difference sequence from the content sequence to be converted.

[0141] Taking a speech-to-grid sequence model as an example, the model can convert the sequence of content to be converted into a grid parameter difference sequence. However, if the grid parameter difference sequence includes not only changes in facial expressions but also changes in head movement trajectories, it may not converge during the training of the speech-to-grid sequence model, thus failing to complete the training.

[0142] Research has shown that predicting facial expression changes and predicting head movement trajectory changes are equivalent to two prediction tasks. If the same model is used to perform two prediction tasks, it may cause the model to lose focus on learning, resulting in non-convergence.

[0143] Therefore, this embodiment separates the two prediction tasks. First, based on the sequence of content to be converted, a grid parameter difference sequence is determined. This sequence identifies the differences between expressions indicated by multiple contents to be converted, thereby achieving the expression prediction task. Second, based on the sequence of content to be converted, a head pose vector sequence is determined. This sequence includes multiple pose vectors, each describing the head's position and orientation, thus reflecting changes in the head's movement trajectory and achieving the head movement trajectory prediction task. Finally, a facial grid model sequence is obtained based on the grid parameter difference sequence. The facial grid model sequence and the head pose vector sequence are then reprojected to obtain the facial pose sequence.

[0144] This application does not specifically limit the method of obtaining the head pose vector sequence. Two methods are described below as examples.

[0145] Method 1: Based on the sequence of content to be converted, a speech-to-gesture sequence model is used to convert it to obtain a head gesture vector sequence.

[0146] In other words, a speech-to-pose sequence model is trained, which can establish a mapping relationship between the content sequence to be converted and the head pose vector sequence, thereby generating a corresponding grid parameter difference sequence based on the content sequence to be converted.

[0147] The training process of the speech-to-gesture sequence model is explained below.

[0148] A facial speech video sample with a third label is obtained. This third label identifies the head pose vector sequence corresponding to the original facial sample in the facial speech video sample. Based on the facial speech video sample, an initial speech-to-pose sequence model is used to convert the speech to pose sequence to obtain a predicted head pose vector sequence. According to the difference between the predicted head pose vector sequence and the third label, the model parameters of the initial speech-to-pose sequence model are adjusted so that the predicted head pose vector sequence gets closer and closer to the third label, thus obtaining the speech-to-pose sequence model.

[0149] The embodiments of this application do not specifically limit the speech-to-gesture sequence model, and can be based on the Wav2vec model, etc.

[0150] Method 2: Based on the sequence of content to be converted, the speech-to-gesture model is used to convert the content to obtain the head gesture vector of each content to be converted in the sequence, and thus a head gesture vector sequence is obtained based on multiple head gesture vectors.

[0151] The training process of the speech-to-gesture model will be explained below.

[0152] A facial vocalization image sample with a fourth label is obtained, which is used to identify the head pose vectors of multiple facial vocalization image samples. Based on the facial vocalization image sample, an initial speech-to-pose model is used for conversion to obtain multiple predicted head pose vectors corresponding to the facial vocalization image sample. According to the differences between the multiple predicted head pose vectors and the fourth label, the model parameters of the initial speech-to-pose model are adjusted so that the predicted head pose vectors are closer and closer to the fourth label, thus obtaining the speech-to-pose model.

[0153] Therefore, if the content sequence to be converted is audio, the facial expression prediction task and the head motion trajectory prediction task can be separated. For example, based on the content sequence to be converted, a grid parameter difference sequence can be obtained through a speech-to-grid sequence model, and a head pose vector sequence can be obtained through a speech-to-pose sequence model. This allows each prediction task to be more targeted, improving prediction accuracy and thus enhancing the quality of the target video.

[0154] As one possible implementation, this application embodiment also provides a specific implementation of S205, namely, a specific implementation of generating a target video based on a facial mesh model sequence and a reference facial image, see S2051-S2053 for details.

[0155] S2051: Extract features from the reference facial image to obtain the image features of the reference facial image.

[0156] Image features are used to describe the characteristics of the reference facial image, such as texture features, identity features of the reference face, color features, etc., so as to learn the characteristics of the reference facial image again so that the target video generated later is more similar to the reference facial image.

[0157] As one possible implementation, since the reference facial image is large in size and the generated target video needs to be calculated on multi-frame superimposed data, the amount of data can be reduced by encoding in order to reduce the amount of computation. Specifically, the reference facial image is encoded to obtain the reference facial image code, wherein the data dimension of the reference facial image code is smaller than that of the reference facial image to reduce the feature dimension. For example, a reference facial image of size 3×512×512 is mapped to a low-dimensional feature of 4×64×64. Subsequently, feature extraction can be performed on the reference facial image code to obtain the image features of the reference facial image.

[0158] The embodiments of this application do not specifically limit the encoding method, such as using the VAE encoder in Variational Autoencoders (VAE) for encoding.

[0159] S2052: Extract features from the facial mesh model sequence to obtain the facial feature sequence.

[0160] The facial feature sequence includes facial features required for multiple video frames. These facial features are obtained by adjusting the mesh model of the reference face based on the mesh parameter difference sequence. For example, it integrates the facial contour of the reference face and the expression of the original face, which can reflect the characteristics of the fused face.

[0161] As one possible approach, a pose guide can be used to extract various facial features from the facial feature sequence, such as the location of key points on the face, changes in facial feature movement, and expression weight coefficients, thereby more accurately describing the characteristics of the face.

[0162] S2053: Generate the target video based on image features and facial feature sequences.

[0163] The facial feature sequence is fused with each facial feature and the image features to obtain multiple video frames, thus obtaining the target video.

[0164] As one possible implementation, if the image features are based on the encoding of a reference facial image, video can be generated based on the image features and the facial feature sequence to obtain the target video encoding, and then the target video encoding can be decoded to obtain the target video. For example, if a VAE encoder is used for encoding, a corresponding VAE decoder can be used for decoding.

[0165] Therefore, image features are extracted from the reference facial image, and facial feature sequences are extracted from the facial mesh model. The image features and facial feature sequences are then fused to obtain the target video. This is equivalent to first learning the shape features of the reference face in the reference facial image, and then selectively learning the image features of the reference facial image. That is, learning from the reference facial image a small number of times makes the style of the target video more similar to the reference facial image, thereby improving the quality of the target video.

[0166] As one possible implementation, this application embodiment also provides a specific implementation of S2051, namely, a specific implementation of extracting features from a reference facial image to obtain the image features of the reference facial image, see D1-D5 for details.

[0167] D1: Obtain the noise vector for facial features in the facial feature sequence.

[0168] For ease of explanation, the following example uses a target facial feature from among multiple facial features in a facial feature sequence. During the process of obtaining the corresponding video frame for the target face, i.e., the target video frame, a noise vector, such as the vector corresponding to Gaussian noise, can be acquired.

[0169] D2: Based on the noise vector, shallow features are obtained by performing convolution calculations through the convolutional layers included in the diffusion model.

[0170] In this embodiment, the diffusion model may include a stack of convolutional layers, cross-attention layers, and stereo cross-attention layers. This embodiment does not specifically limit the convolutional layers; for example, a residual neural network (ResNet) may be used. Based on the noise vector, convolutional calculations are performed through the convolutional layers included in the diffusion model to obtain shallow features, which are used to describe some detailed information.

[0171] D3: Based on shallow features and facial features, the features are fused through the cross-attention layer included in the diffusion model to obtain deep features.

[0172] By overlaying target facial features onto shallow features, pose guidance information is injected to obtain deep features. Deep features can capture the complex structure and patterns of data, thereby improving recognition accuracy and generalization ability.

[0173] D4: Based on deep features and image features, the features are fused through the stereo cross-attention layer included in the diffusion model to obtain the video frames in the target video.

[0174] By overlaying the image features of the reference facial image on top of the deep features, information injection from the reference facial image is achieved. Thus, by combining the noise vector, the target facial features, and the image features, the target video frame corresponding to the target facial features is obtained.

[0175] It is understandable that after introducing a noise vector, the desired output target video frame can be obtained through multiple cyclic denoising processes.

[0176] D5: Generate the target video from the video frames obtained by using the facial features included in the facial feature sequence.

[0177] The facial features included in the facial feature sequence are used as target facial features to obtain video frames corresponding to each facial feature, thereby obtaining the target video based on multiple video frames.

[0178] Therefore, by introducing a noise vector, the subsequent diffusion model can introduce subtle variations when generating target video frames, making each generated video frame different and enriching the details of the target video. This avoids repetitive or overly regular images during the generation process, improving the quality of the target video. Moreover, the noise vector also helps the diffusion model better capture details and textures in the data, making the generated video frames more realistic and natural, and the picture more coherent.

[0179] To facilitate a further understanding of the technical solutions provided in the embodiments of this application, the following description takes the execution subject of the video generation method provided in the embodiments of this application as a server as an example, and provides an overall exemplary introduction to the video generation method.

[0180] The following section explains the process of using the video generation method.

[0181] Users can upload sequences of content to be converted, and users can specify reference images, such as uploading a reference image or selecting one from multiple candidate reference images, etc. This application does not impose specific limitations on this. Depending on the sequence of content to be converted uploaded by the user, there are two implementation methods, as shown in Figures 3 and 4 respectively, which are described below.

[0182] If the content sequence to be converted is audio, it is processed according to the method shown in Figure 3. Specifically, after acquiring the audio and a reference facial image, the user-uploaded audio is used as the content to be converted. Multiple video frames are obtained using the video generation method provided in this application embodiment, and the number of these video frames matches the length of the audio. Finally, the multiple video frames are combined with the audio to obtain the target video.

[0183] If the content sequence to be converted is a video, it is processed according to the method shown in Figure 4. Specifically, the user-uploaded video is used as the content to be converted, and multiple video frames are obtained through the video generation method provided in this application embodiment. Finally, audio is extracted from the video, and the multiple video frames are combined with the audio to obtain the target video. It can be understood that the process of extracting audio from the video and then combining the audio with multiple video frames to obtain the target video is also within the video generation method provided in this application embodiment.

[0184] Then, we will explain from the perspective of the video generation system.

[0185] This video generation system comprises three modules: a speech conversion module, a video conversion module, and a video generation module. These will be described in detail below.

[0186] (1) Voice conversion module.

[0187] When the content to be converted is identified as audio, a speech conversion module is used to obtain a facial pose sequence. See Figure 5, which is a schematic diagram of a speech conversion module provided in an embodiment of this application.

[0188] A mesh model of a reference face is obtained based on a reference facial image. This mesh model can include the facial contour information of the reference face. Furthermore, in scenarios where the audio is a sequence of content to be converted, the expression in the reference facial image can be an expressionless image. If the expression in the reference facial image is another expression, an expressionless reference facial image can be generated using a diffusion model to facilitate the subsequent generation of a facial mesh model sequence.

[0189] The audio is segmented into multiple segments based on its duration, such as dividing it into segments at 30fps, thus obtaining a facial pose for each segment. These multiple audio segments are then input into a speech-to-pose sequence model, which converts them into a head pose vector sequence, representing the head movement trajectory matched to the audio. Finally, the multiple audio segments are input into a speech-to-mesh sequence model, which converts them into a mesh parameter difference sequence, representing changes in lip shape and blinking during speech.

[0190] The reference face mesh model is adjusted by the mesh parameter difference sequence to obtain the face mesh model sequence. The face mesh model sequence and the head pose vector sequence are then reprojected to obtain the face pose sequence, which can be a face pose sequence of a plane that emits the audio.

[0191] (2) Video conversion module.

[0192] When the content to be converted is identified as video, a facial pose sequence is obtained using a video conversion module. See Figure 6, which is a schematic diagram of a video conversion module provided in an embodiment of this application.

[0193] A mesh model of the reference face is obtained based on the reference facial image. Furthermore, in scenarios where the video is a sequence of content to be converted, the facial expression in the reference image can be arbitrary. Based on each video frame, a mesh model of the original face in each frame is obtained, thus producing a sequence of mesh models of the original face.

[0194] The similarity between multiple original facial mesh models and the reference facial mesh model is determined. Based on the similarity between the multiple original facial mesh models and the reference facial mesh model, the mesh models in the original facial mesh model sequence that meet the similarity conditions are determined as standard mesh models. Based on the difference between each mesh model in the original facial mesh model sequence and the standard mesh model, a mesh parameter difference sequence is obtained.

[0195] The facial mesh model is obtained by adjusting the reference face mesh model through a mesh parameter difference sequence, thus realizing the transformation of the facial contour. This facial mesh model sequence includes a head pose vector sequence, so it can be directly reprojected based on the facial mesh model sequence to obtain the facial pose sequence.

[0196] (3) Video generation module.

[0197] Referring to Figure 7, this figure is a schematic diagram of a video generation module provided in an embodiment of this application. As mentioned above, both the video conversion module and the speech conversion module will obtain corresponding facial pose sequences. The facial pose sequences are input into the pose guide, which extracts the facial poses to obtain a facial feature sequence, that is, the facial pose sequences are evolved into a coherent image sequence of corresponding actions, and the facial feature sequences are input into the diffusion model.

[0198] The reference facial image is input into the VAE encoder and encoded to obtain the reference facial image encoding. Then, feature extraction is performed on the reference facial image encoding to obtain the image features of the reference facial image, and the image features of the reference facial image are input into the diffusion model.

[0199] The diffusion model obtains video encoding based on facial feature sequences, image features of a reference facial image, and noise vectors. The video encoding is then decoded by a VAE decoder to obtain multiple video frames.

[0200] Finally, after obtaining multiple video frames, the target video is obtained based on these multiple video frames.

[0201] Therefore, the embodiments of this application can utilize audio or video to drive any reference facial image to obtain high-quality lip-sync driven video, i.e., the target video. Moreover, it does not require any training data for the reference facial image, exhibiting strong generalization ability. The resulting video is rich in detail, has coherent visuals, accurate lip movements, and the trajectory of the human head matches the rhythm of the audio, making the overall target video more natural and realistic.

[0202] In relation to the video generation method described above, this application also provides a corresponding video generation apparatus so that the above video generation method can be applied and implemented in practice.

[0203] Referring to Figure 8, this figure is a schematic diagram of the structure of a video generation device provided in an embodiment of this application. As shown in Figure 8, the video generation device 800 includes: an acquisition unit 801, a first determination unit 802, a second determination unit 803, an adjustment unit 804, and a generation unit 805;

[0204] The acquisition unit 801 is used to acquire a reference facial image and a sequence of content to be converted. The reference facial image includes the shape features of a reference face, and the sequence of content to be converted includes multiple content to be converted that changes over time. The content to be converted is used to indicate the facial motion features of the original face.

[0205] The first determining unit 802 is used to determine the mesh model of the reference face based on the reference face image;

[0206] The second determining unit 803 is used to determine a grid parameter difference sequence based on the sequence of content to be converted, wherein the grid parameter difference sequence is used to identify the differences between facial motion features indicated by multiple contents to be converted;

[0207] The adjustment unit 804 is used to adjust the mesh model of the reference face through the mesh parameter difference sequence to obtain a facial mesh model sequence;

[0208] The generation unit 805 is used to generate a target video based on the facial mesh model sequence and the reference facial image.

[0209] As can be seen from the above technical solutions, the video generation apparatus provided in this application includes: an acquisition unit, a first determination unit, a second determination unit, an adjustment unit, and a generation unit. The acquisition unit acquires a reference facial image and a sequence of content to be converted. The reference facial image includes the shape features of a reference face, and the sequence of content to be converted includes multiple pieces of content to be converted, which change over time. Each piece of content to be converted is used to indicate the facial motion features of the original face. The first determination unit determines a mesh model of the reference face based on the reference facial image. The mesh model of the reference face can explicitly reflect the external features of the reference face. The second determination unit determines a mesh parameter difference sequence based on the sequence of content to be converted. This mesh parameter difference sequence is used to identify the differences between the facial motion features indicated by the multiple pieces of content to be converted, i.e., to reflect the changes in the facial motion features of the original face over time during the sound emission process. Therefore, by utilizing the adjustment unit, the mesh model of the reference face is adjusted through a mesh parameter difference sequence. That is, the reference face's mesh model is adjusted based on the facial motion features of the original face. This eliminates the use of fixed facial contour parameters and instead dynamically adjusts the model based on time-varying facial motion features, resulting in a better integration of the reference face's shape features with the original face's facial motion features in the adjusted facial mesh model sequence. Then, through the generation unit, the target video is generated based on the facial mesh model sequence and the reference face image, thereby improving the quality of the target video.

[0210] As one possible implementation, the second determining unit 803 is specifically used for:

[0211] Based on the sequence of content to be converted, determine the mesh model sequence of the original face, which includes multiple mesh models of the original face;

[0212] Determine the similarity between the mesh models of the multiple original faces and the mesh model of the reference face;

[0213] Among the multiple original facial mesh models, the mesh model whose similarity meets the similarity condition is determined as the standard mesh model;

[0214] The mesh parameter difference sequence is obtained based on the differences in facial motion features between each mesh model in the original facial mesh model sequence and the standard mesh model.

[0215] As one possible implementation, if the expression of the reference face is in a first state, then the second determining unit 803 is specifically used for:

[0216] Based on the sequence of content to be converted, determine the mesh model sequence of the original face, which includes multiple mesh models of the original face;

[0217] Among the multiple original facial mesh models, the mesh model whose expression matches the first state is determined as the standard mesh model;

[0218] The mesh parameter difference sequence is obtained based on the differences in facial motion features between each mesh model in the original facial mesh model sequence and the standard mesh model.

[0219] As one possible implementation, the acquisition unit 801 is further configured to acquire a reference face image to be determined, wherein the expression of the reference face in the reference face image to be determined is in a second state;

[0220] The device further includes a preprocessing unit for adjusting the facial motion features of the reference face in the undetermined reference face image based on the first state, thereby obtaining the reference face image.

[0221] As one possible implementation, the generation unit 805 is specifically used for:

[0222] The facial mesh model sequence is reprojected to obtain a facial pose sequence;

[0223] The target video is generated based on the facial pose sequence and the reference facial image.

[0224] As one possible implementation, if the content sequence to be converted is audio, the device further includes a third determining unit, used to obtain a head posture vector sequence based on the content sequence to be converted, the head posture vector sequence including multiple posture vectors, the posture vectors being used to describe the position of the head and the orientation of the head;

[0225] The generation unit 805 is specifically used to reproject the facial mesh model sequence and the head pose vector sequence to obtain a facial pose sequence.

[0226] As one possible implementation, if the content sequence to be converted is audio, the device further includes a third determining unit, used to convert the content sequence to be converted by a speech-to-gesture sequence model to obtain the head gesture vector sequence.

[0227] The training method for the speech-to-gesture sequence model is as follows:

[0228] Obtain facial vocalization video samples with a third label, the third label being used to identify the head pose vector sequence of the facial vocalization video samples;

[0229] Based on the facial vocalization video samples, the initial speech-to-pose sequence model is used to convert the samples to obtain a predicted head pose vector sequence.

[0230] Based on the difference between the predicted head pose vector sequence and the third label, the model parameters of the initial speech-to-pose sequence model are adjusted to obtain the speech-to-pose sequence model.

[0231] The speech-to-gesture sequence model can be trained by the training unit included in the device, or it can be trained by other devices. This application does not make any specific limitation on this.

[0232] As one possible implementation, the generation unit 805 is specifically used for:

[0233] Feature extraction is performed on the reference facial image to obtain the image features of the reference facial image;

[0234] Feature extraction is performed on the facial mesh model sequence to obtain a facial feature sequence;

[0235] The target video is obtained by generating a video based on the image features and the facial feature sequence.

[0236] As one possible implementation, the generation unit 805 is specifically used for:

[0237] For the facial features in the facial feature sequence, obtain the noise vector;

[0238] Based on the noise vector, shallow features are obtained by performing convolution calculations through the convolutional layers included in the diffusion model.

[0239] Based on the shallow features and the facial features, the features are fused through the cross-attention layer included in the diffusion model to obtain deep features;

[0240] Based on the deep features and the image features, the features are fused through the stereo cross-attention layer included in the diffusion model to obtain the video frames in the target video;

[0241] The target video is generated by obtaining video frames from the facial features included in the facial feature sequence.

[0242] As one possible implementation, the generation unit 805 is specifically used for:

[0243] The reference facial image is encoded to obtain a reference facial image code, wherein the data dimension of the reference facial image code is smaller than the data dimension of the reference facial image;

[0244] Feature extraction is performed on the encoding of the reference facial image to obtain the image features of the reference facial image;

[0245] The target video code is generated based on the image features and the facial feature sequence.

[0246] The target video is decoded to obtain the target video.

[0247] As one possible implementation, if the content to be converted is audio, then the second determining unit 803 is specifically used for:

[0248] Based on the sequence of content to be converted, the speech-to-grid sequence model is used to convert the content to obtain the grid parameter difference sequence.

[0249] The training method for the speech-to-grid sequence model is as follows:

[0250] Obtain facial vocalization video samples with a first label, wherein the first label is used to identify the grid parameter difference sequence of the facial vocalization video samples;

[0251] Based on the facial vocalization video samples, the initial speech-to-grid sequence model is used to convert the data to obtain a predicted grid parameter difference sequence.

[0252] Based on the difference between the predicted grid parameter difference sequence and the first label, the model parameters of the initial speech-to-grid sequence model are adjusted to obtain the speech-to-grid sequence model.

[0253] The speech-to-grid sequence model can be trained by the training unit included in the device, or it can be trained by other devices. This application does not make any specific limitation on this.

[0254] This application also provides a computer device, which can be a server or a terminal device. The computer device provided in this application will be described below from the perspective of hardware implementation. Figure 9 shows a schematic diagram of the server structure, and Figure 10 shows a schematic diagram of the terminal device structure.

[0255] Referring to Figure 9, which is a schematic diagram of a server structure provided in an embodiment of this application, the server 1400 can vary considerably due to different configurations or performance. It may include one or more processors 1422, such as a Central Processing Unit (CPU), memory 1432, and one or more storage media 1430 (e.g., one or more mass storage devices) for application programs 1442 or data 1444. The memory 1432 and storage media 1430 can be temporary or persistent storage. The program stored in the storage media 1430 may include one or more modules (not shown in the figure), each module including a series of instruction operations on the server. Furthermore, the processor 1422 may be configured to communicate with the storage media 1430 and execute the series of instruction operations in the storage media 1430 on the server 1400.

[0256] Server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input / output interfaces 1458, and / or one or more operating systems 1441, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM etc.

[0257] The steps performed by the server in the above embodiments can be based on the server structure shown in Figure 9.

[0258] The processor 1422 is used to perform the following steps:

[0259] A reference facial image and a sequence of content to be converted are obtained. The reference facial image includes the shape features of a reference face, and the sequence of content to be converted includes multiple content items that change over time. The content items are used to indicate the facial motion features of the original face.

[0260] Based on the reference facial image, determine the mesh model of the reference face;

[0261] Based on the sequence of content to be converted, a grid parameter difference sequence is determined, which is used to identify the differences between facial motion features indicated by multiple contents to be converted;

[0262] The mesh model of the reference face is adjusted by the mesh parameter difference sequence to obtain a facial mesh model sequence;

[0263] A target video is generated based on the facial mesh model sequence and the reference facial image.

[0264] Optionally, the processor 1422 may also execute method steps of any specific implementation of the video generation method in the embodiments of this application.

[0265] Referring to Figure 10, this figure is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Taking a smartphone as an example, Figure 10 shows a block diagram of part of the structure of the smartphone, which includes: a radio frequency (RF) circuit 1510, a memory 1520, an input unit 1530, a display unit 1540, a sensor 1550, an audio circuit 1560, a Wi-Fi module 1570, a processor 1580, and a power supply 1590, etc. Those skilled in the art will understand that the smartphone structure shown in Figure 10 does not constitute a limitation on the smartphone, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0266] The following section, with reference to Figure 10, provides a detailed introduction to the various components of a smartphone:

[0267] The RF circuit 1510 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1580; in addition, it transmits uplink data to the base station.

[0268] The memory 1520 can be used to store software programs and modules, and the processor 1580 runs the software programs and modules stored in the memory 1520 to realize various functions and data processing of the smartphone.

[0269] Input unit 1530 can be used to receive input numeric or character information and generate key signal inputs related to user settings and function control of the smartphone. Specifically, input unit 1530 may include touch panel 1531 and other input devices 1532. Touch panel 1531, also known as a touch screen, can collect touch operations on or near the user and drive corresponding connected devices according to a pre-set program. In addition to touch panel 1531, input unit 1530 may also include other input devices 1532. Specifically, other input devices 1532 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0270] The display unit 1540 can be used to display information input by the user or information provided to the user, as well as various menus of the smartphone. The display unit 1540 may include a display panel 1541, which may optionally be configured as a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0271] Smartphones may also include at least one sensor 1550, such as a light sensor, a motion sensor, and other sensors. Other sensors that smartphones may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be detailed here.

[0272] Audio circuit 1560, speaker 1561, and microphone 1562 provide an audio interface between the user and the smartphone. Audio circuit 1560 converts received audio data into electrical signals and transmits them to speaker 1561, where speaker 1561 converts them into sound signals for output. On the other hand, microphone 1562 converts collected sound signals into electrical signals, which are received by audio circuit 1560, converted into audio data, and then processed by processor 1580 before being transmitted via RF circuit 1510 to, for example, another smartphone, or the audio data can be output to memory 1520 for further processing.

[0273] The processor 1580 is the control center of the smartphone, connecting various parts of the smartphone through various interfaces and lines. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 1520, and by calling data stored in the memory 1520. Optionally, the processor 1580 may include one or more processing units.

[0274] The smartphone also includes a power supply 1590 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 1580 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0275] Although not shown, smartphones may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0276] In this embodiment of the application, the memory 1520 included in the smartphone can store computer programs and transmit the computer programs to the processor.

[0277] The processor 1580 included in the smartphone can execute the video generation method provided in the above embodiments according to the instructions in the computer program.

[0278] This application also provides a computer-readable storage medium for storing a computer program for executing the video generation method provided in the above embodiments.

[0279] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video generation method provided in the various optional implementations of the above aspects.

[0280] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk or optical disk, and other media that can store computer programs.

[0281] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0282] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0283] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video generation method, the method being executed by a computer device, the method comprising: A reference facial image and a sequence of content to be converted are obtained. The reference facial image includes the shape features of a reference face, and the sequence of content to be converted includes multiple content items that change over time. The content items are used to indicate the facial motion features of the original face. Based on the reference facial image, determine the mesh model of the reference face; Based on the sequence of content to be converted, a grid parameter difference sequence is determined, which is used to identify the differences between facial motion features indicated by multiple contents to be converted; The mesh model of the reference face is adjusted by the mesh parameter difference sequence to obtain a facial mesh model sequence; A target video is generated based on the facial mesh model sequence and the reference facial image.

2. The method according to claim 1, wherein determining the grid parameter difference sequence based on the sequence of content to be converted comprises: Based on the sequence of content to be converted, determine the mesh model sequence of the original face, which includes multiple mesh models of the original face; Determine the similarity between the mesh models of the multiple original faces and the mesh model of the reference face; Among the multiple original facial mesh models, the mesh model whose similarity meets the similarity condition is determined as the standard mesh model; The mesh parameter difference sequence is obtained based on the differences in facial motion features between each mesh model in the original facial mesh model sequence and the standard mesh model.

3. The method according to claim 1, wherein if the expression of the reference face is in a first state, then determining the mesh parameter difference sequence based on the sequence of content to be converted includes: Based on the sequence of content to be converted, determine the mesh model sequence of the original face, which includes multiple mesh models of the original face; Among the multiple original facial mesh models, the mesh model whose expression matches the first state is determined as the standard mesh model; The mesh parameter difference sequence is obtained based on the differences in facial motion features between each mesh model in the original facial mesh model sequence and the standard mesh model.

4. The method according to claim 3, further comprising: Obtain a reference facial image to be determined, wherein the expression of the reference face in the reference facial image is in a second state; Based on the first state, the facial motion features of the reference face in the undetermined reference face image are adjusted to obtain the reference face image.

5. The method according to claim 1, wherein generating the target video based on the facial mesh model sequence and the reference facial image comprises: The facial mesh model sequence is reprojected to obtain a facial pose sequence; The target video is generated based on the facial pose sequence and the reference facial image.

6. The method according to claim 5, wherein if the sequence of content to be converted is audio, the method further comprises: Based on the sequence of content to be converted, a head pose vector sequence is obtained. The head pose vector sequence includes multiple pose vectors, which are used to describe the position of the head and the orientation of the head. The reprojection of the facial mesh model sequence to obtain a facial pose sequence includes: The facial mesh model sequence and the head pose vector sequence are reprojected to obtain the facial pose sequence.

7. The method according to claim 6, wherein obtaining the head pose vector sequence based on the content sequence to be converted comprises: Based on the sequence of content to be converted, the head posture vector sequence is obtained by converting it using a speech-to-gesture sequence model. The training method for the speech-to-gesture sequence model is as follows: Obtain facial vocalization video samples with a third label, the third label being used to identify the head pose vector sequence of the facial vocalization video samples; Based on the facial vocalization video samples, the initial speech-to-pose sequence model is used to convert the samples to obtain a predicted head pose vector sequence. Based on the difference between the predicted head pose vector sequence and the third label, the model parameters of the initial speech-to-pose sequence model are adjusted to obtain the speech-to-pose sequence model.

8. The method according to claim 1, wherein generating the target video based on the facial mesh model sequence and the reference facial image comprises: Feature extraction is performed on the reference facial image to obtain the image features of the reference facial image; Feature extraction is performed on the facial mesh model sequence to obtain a facial feature sequence; The target video is obtained by generating a video based on the image features and the facial feature sequence.

9. The method according to claim 8, wherein generating the target video based on the image features and the facial feature sequence comprises: For the facial features in the facial feature sequence, obtain the noise vector; Based on the noise vector, shallow features are obtained by performing convolution calculations through the convolutional layers included in the diffusion model. Based on the shallow features and the facial features, the features are fused through the cross-attention layer included in the diffusion model to obtain deep features; Based on the deep features and the image features, the features are fused through the stereo cross-attention layer included in the diffusion model to obtain the video frames in the target video; The target video is generated by obtaining video frames from the facial features included in the facial feature sequence.

10. The method according to claim 8, wherein extracting features from the reference facial image to obtain image features of the reference facial image comprises: The reference facial image is encoded to obtain a reference facial image code, wherein the data dimension of the reference facial image code is smaller than the data dimension of the reference facial image; Feature extraction is performed on the encoding of the reference facial image to obtain the image features of the reference facial image; The step of generating the target video based on the image features and the facial feature sequence includes: The target video code is generated based on the image features and the facial feature sequence. The target video is decoded to obtain the target video.

11. The method according to any one of claims 1-10, wherein if the content to be converted is audio, the step of determining the grid parameter difference sequence based on the sequence of the content to be converted includes: Based on the sequence of content to be converted, the speech-to-grid sequence model is used to convert the content to obtain the grid parameter difference sequence. The training method for the speech-to-grid sequence model is as follows: Obtain facial vocalization video samples with a first label, wherein the first label is used to identify the grid parameter difference sequence of the facial vocalization video samples; Based on the facial vocalization video samples, the initial speech-to-grid sequence model is used to convert the data to obtain a predicted grid parameter difference sequence. Based on the difference between the predicted grid parameter difference sequence and the first label, the model parameters of the initial speech-to-grid sequence model are adjusted to obtain the speech-to-grid sequence model.

12. A video generation apparatus, the apparatus comprising: The system comprises an acquisition unit, a first determination unit, a second determination unit, an adjustment unit, and a generation unit. The acquisition unit is used to acquire a reference facial image and a sequence of content to be converted. The reference facial image includes the shape features of a reference face, and the sequence of content to be converted includes multiple content to be converted that changes over time. The content to be converted is used to indicate the facial motion features of the original face. The first determining unit is configured to determine the mesh model of the reference face based on the reference face image; The second determining unit is configured to determine a grid parameter difference sequence based on the sequence of content to be converted, wherein the grid parameter difference sequence is used to identify the differences between facial motion features indicated by multiple contents to be converted; The adjustment unit is used to adjust the mesh model of the reference face through the mesh parameter difference sequence to obtain a facial mesh model sequence; The generation unit is used to generate a target video based on the facial mesh model sequence and the reference facial image.

13. A computer device, the computer device comprising a processor and a memory: The memory is used to store computer programs and to transfer the computer programs to the processor; The processor is configured to perform the method according to any one of claims 1-11 according to the computer program.

14. A computer-readable storage medium for storing a computer program for performing the method according to any one of claims 1-11.

15. A computer program product comprising a computer program, which, when run on a computer device, causes the computer device to perform the method of any one of claims 1-11.

Citation Information

Patent Citations

  • Image-based control method and device

    CN106331572A

  • Face adjustment method and device, live broadcast method and device, electronic equipment and storage medium

    CN111652794A

  • Three-dimensional head model construction method, device and system and storage medium

    CN113470162A

  • Video generation method and device, electronic equipment and storage medium

    CN113689538A

  • Method, device and equipment for generating dynamic image based on audio and storage medium

    CN117523051A