Digital human-driven model generation method and device, electronic equipment and storage medium

By adopting multimodal feature fusion and two-stage sub-model training methods in digital human video generation, the problem of generating high-quality digital human videos when there is a lack of a driving signal source is solved, and natural and efficient video generation is achieved.

CN119992416APending Publication Date: 2025-05-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510065975.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the absence of a driving signal source, it is difficult to generate natural and high-quality digital human videos when only voice input is available. In the case of a lack of a driving signal source, the prior art is prone to introduce errors when voice prediction of human body motion signals, resulting in a decline in video quality.

Method used

Through the multimodal feature fusion based on reference video and driver speech, a two-stage sub-model is trained to generate a digital human-driven model. The specific steps include: determining the synthetic sequence and feature mark based on the reference video, training the first sub-model to generate the first driver video; training the second sub-model to generate the second driver video based on the second synthetic sequence and feature mark; integrating the trained sub-model to generate the digital human-driven model.

Benefits of technology

In the absence of a driving signal source, high-quality digital human videos are generated through voice input, which improves the naturalness and production efficiency of the video, and ensures the consistency of driving human body movement and audio rhythm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992416A_ABST
    Figure CN119992416A_ABST
Patent Text Reader

Abstract

The invention provides a digital human driven model generation method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models, augmented reality and the like, and can be applied to scenes such as digital human. According to the specific implementation scheme, a first synthesis sequence, a second synthesis sequence, a voice feature mark sequence, a reference feature mark and a facial feature mark are determined based on a reference video; training a first sub-model based on the first synthesis sequence, the voice feature mark sequence, the reference feature mark and the facial feature mark, so that the first sub-model outputs a first driving video; training a second sub-model based on the second synthetic sequence, the reference feature mark and the facial feature mark, so that the second sub-model outputs a second driving video; and generating the digital human-driven model based on the trained first sub-model and the second sub-model. According to the scheme, the quality of the digital human video generated by the digital human driving model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, especially to technical fields such as computer vision, deep learning, large models, and augmented reality, and can be applied to scenarios such as digital humans, and specifically to methods, devices, electronic devices, and storage media for generating digital human-driven models. Background Art

[0002] Digital humans have been widely used in many industries, such as entertainment, education, and e-commerce. As industry demand and user experience gradually increase, the demand for digital human videos in various industries is also increasing. Summary of the invention

[0003] The present disclosure provides a method, device, electronic device and storage medium for generating a digital human driving model.

[0004] According to a first aspect of the present disclosure, a method for generating a digital human driving model is provided, comprising: determining a first synthesis sequence, a second synthesis sequence, a speech feature marker sequence, a reference feature marker and a facial feature marker based on a reference video; training a first sub-model based on the first synthesis sequence, the speech feature marker sequence, the reference feature marker and the facial feature marker so that the first sub-model outputs a first driving video; training a second sub-model based on the second synthesis sequence, the reference feature marker and the facial feature marker so that the second sub-model outputs a second driving video; and generating a digital human driving model based on the trained first sub-model and the second sub-model.

[0005] According to a second aspect of the present disclosure, a digital human driving method is provided, including: determining a digital human reference feature marker and a digital human facial feature marker based on a reference image; extracting a driving voice feature marker sequence based on a driving voice; determining a third synthetic sequence for generating each video clip based on the reference image and the driving voice; inputting the third synthetic sequence, the driving voice feature marker sequence, the digital human reference feature marker and the digital human facial feature marker into a first sub-model of a digital human driving model to obtain a third driving video; determining a fourth synthetic sequence for generating each video clip based on the third driving video; inputting the fourth synthetic sequence, the digital human reference feature marker and the digital human facial feature marker into a second sub-model of the digital human driving model to obtain a fourth driving video; and generating a digital human video based on the fourth driving video corresponding to all video clips.

[0006] According to a third aspect of the present disclosure, a digital human driving model generation device is provided, comprising: a feature extraction module, used to determine a first synthesis sequence, a second synthesis sequence, a speech feature marker sequence, a reference feature marker and a facial feature marker based on a reference video; a first training module, used to train a first sub-model based on the first synthesis sequence, the speech feature marker sequence, the reference feature marker and the facial feature marker, so that the first sub-model outputs a first driving video; a second training module, used to train a second sub-model based on the second synthesis sequence, the reference feature marker and the facial feature marker, so that the second sub-model outputs a second driving video; and a model generation module, used to generate a digital human driving model based on the trained first sub-model and the second sub-model.

[0007] According to a fourth aspect of the present disclosure, a digital human driving device is provided, comprising: a feature marking module, used to determine a digital human reference feature marking and a digital human facial feature marking based on a reference image; a speech extraction module, used to extract a driving voice feature marking sequence based on a driving voice; a first sequence module, used to determine a third synthetic sequence for generating each video clip based on the reference image and the driving voice; a first generation module, used to input the third synthetic sequence, the driving voice feature marking sequence, the digital human reference feature marking and the digital human facial feature marking into a first sub-model of a digital human driving model to obtain a third driving video; a second sequence module, used to determine a fourth synthetic sequence for generating each video clip based on the third driving video; a second generation module, used to input the fourth synthetic sequence, the digital human reference feature marking and the digital human facial feature marking into a second sub-model of the digital human driving model to obtain a fourth driving video; and a video generation module, used to generate a digital human video based on the fourth driving video corresponding to all video clips.

[0008] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0009] at least one processor; and

[0010] a memory communicatively connected to the at least one processor; wherein,

[0011] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0012] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute any method according to the embodiments of the present disclosure.

[0013] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.

[0014] The solution disclosed in the present invention can improve the quality of digital human videos generated by digital human driving models.

[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0017] Figure 1 is a flow chart of a method for generating a digital human driving model according to an embodiment of the present disclosure;

[0018] Figure 2 is an architectural diagram of a first-stage neural network structure according to an embodiment of the present disclosure;

[0019] Figure 3 is an architectural diagram of a second-stage neural network structure according to an embodiment of the present disclosure;

[0020] Figure 4 is a flowchart of a digital human driving method according to an embodiment of the present disclosure;

[0021] Figure 5 is a structural diagram of a digital human driving model generating device according to an embodiment of the present disclosure;

[0022] Figure 6 is a structural diagram of a digital human driving device according to an embodiment of the present disclosure;

[0023] Figure 7 is a scene schematic diagram of a method for generating a digital human driven model according to an embodiment of the present disclosure;

[0024] Figure 8 is a schematic diagram of a scene of a digital human driving method according to an embodiment of the present disclosure;

[0025] Fig. 9 It is a structural diagram of an electronic device used to implement the digital human driving model generation method and / or the digital human driving method of the embodiments of the present disclosure. DETAILED DESCRIPTION

[0026] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0027] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them. They do not mean to limit the order or to limit them to only two. For example, the first feature and the second feature refer to two types / two features. The first feature can be one or more, and the second feature can also be one or more.

[0028] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.

[0029] Before introducing the technical solutions of the embodiments of the present disclosure, the technical terms that may be used in the present disclosure are further explained:

[0030] Digital human: a character image generated by digital technology based on a video recorded by a real person. After collecting and analyzing the appearance, expression, and movements of the real person, these data are processed using technical means, and lip movements and body movements are arranged so that they can express and display actions according to the preset content. The digital human generated in this way can simulate the performance of a real person to a certain extent and can be used in various scenarios, such as virtual anchors, virtual customer service, etc. Its advantage is that it can retain certain real-life characteristics, making the audience feel more real and intimate.

[0031] Token: In natural language processing, a token is the basic unit in a text, usually a word, punctuation mark, subword, or character.

[0032] In the related technologies, most of them use a human body replay-based method to realize the driving of digital human. Specifically, on the one hand, the human body driving signal can be obtained from an existing human body motion video through three-dimensional (3D) human body modeling and rendering technology, or directly using two-dimensional (2D) detection technology. Then, based on the obtained human body driving signal and a selected target person reference image, the human body motion video is generated. In this process, the pre-processed driving signal, such as the 3D rendered human body model or the 2D key points obtained by detection, is directly input into the generator of the convolutional neural network or diffusion model. However, when there is no driving signal source and only voice input, this type of technology cannot work. On the other hand, some existing technologies realize the driving of human body motion based on voice input, and the process is roughly as follows: first predict the human body driving signal from the voice input, and then complete the video generation based on the human body driving signal. However, this type of technology will introduce errors in the link of voice prediction of human body motion signals, and the accumulation of these errors will lead to a decrease in the quality of the final generated video, and the movement of the person does not appear natural enough. In addition, there are some warping-based methods in the prior art. These methods first implicitly represent and operate dynamic information, and then use the generator of the adversarial generative network to generate digital human videos. However, these methods have certain limitations. They can only achieve good driving effects on limited data sets, are difficult to extend to driving arbitrary images, and the quality of the generated videos is relatively low.

[0033] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, the present disclosure proposes a digital human driving solution, which can realize digital human driving according to reference images and driving voices, simplify the digital human driving process, and ensure that the driving human body movement and audio rhythm are consistent, thereby greatly improving the production efficiency and quality of digital human videos.

[0034] The present disclosure provides a method for generating a digital human driving model. Figure 1 It is a flow chart of a method for generating a digital human-driven model according to an embodiment of the present disclosure. The method for generating a digital human-driven model can be applied to a digital human-driven model generating device. The digital human-driven model generating device is located in an electronic device. The electronic device includes but is not limited to fixed devices and / or mobile devices. For example, fixed devices include but are not limited to servers, and the servers can be cloud servers or ordinary servers. For example, mobile devices include but are not limited to: mobile phones, tablet computers. In some possible implementations, the method for generating a digital human-driven model can also be implemented by a processor calling computer-readable instructions stored in a memory. For example Figure 1As shown, the digital human driving model generation method includes:

[0035] S101, determining a first synthesis sequence, a second synthesis sequence, a speech feature marker sequence, a reference feature marker, and a facial feature marker based on a reference video;

[0036] S102, training a first sub-model based on the first synthesis sequence, the speech feature marker sequence, the reference feature marker, and the facial feature marker, so that the first sub-model outputs a first driving video;

[0037] S103, training a second sub-model based on the second synthetic sequence, the reference feature markers, and the facial feature markers, so that the second sub-model outputs a second driving video;

[0038] S104: Generate a digital human driving model based on the trained first sub-model and the second sub-model.

[0039] In the disclosed embodiment, the digital human driving model is used to output a digital human video based on a reference image.

[0040] In the disclosed embodiments, the reference video refers to an existing video, which is used as the basic data source for generating the video. The reference video provides the basic features of the target person, such as body movements, facial expressions, etc., which are used to analyze and extract relevant information during the driving process, and can be used as a basic template for generating a digital human. For example, the reference video can be a video recorded by a real person.

[0041] In the disclosed embodiment, the first synthetic sequence refers to a reduced-dimensional feature sequence generated by processing the reference video, mainly including a low-dimensional representation of the overall video content. The first synthetic sequence can be used for training the first sub-model.

[0042] In the disclosed embodiment, the second synthetic sequence refers to a reduced-dimensional feature sequence generated after further processing based on the reference video or the first driving video. The second synthetic sequence can be used for training the second sub-model.

[0043] In the disclosed embodiment, the speech feature token sequence (referred to as Audio Tokens) refers to the feature token sequence generated after the speech is extracted from the audio of the reference video and processed. The Token sequence contains the timing rhythm of the audio and the information of the speech content, which can drive the overall action rhythm and facial expression of the target person.

[0044] In the disclosed embodiment, the reference feature tokens (denoted as Appearance Tokens) refer to the overall picture feature information extracted from the reference video. The reference feature tokens can ensure that the characters in the generated video are consistent with the characters in the reference video, including information such as facial structure and head position, and can also ensure that the background in the generated video is consistent with the background in the reference video.

[0045] In the disclosed embodiment, facial feature tags (referred to as Identity Tokens) refer to feature tags generated after processing the target person's facial area, including detailed information of facial features. Facial feature tags can ensure that the facial features in the generated video match the facial features in the reference frame.

[0046] In the disclosed embodiment, after constructing the first sub-model, first, the first synthesis sequence, the speech feature marker sequence, the reference feature marker and the facial feature marker are input into the first sub-model. Then, the first sub-model can be trained using supervised learning or adversarial learning. After the training is completed, the first sub-model can generate a first driving model.

[0047] In the disclosed embodiment, after the second sub-model is constructed, first, the second synthetic sequence, the reference feature markers, and the facial feature markers are input into the second sub-model. Then, the second sub-model can be trained using supervised learning or adversarial learning. After the training is completed, the second sub-model can generate a second driving video.

[0048] In the disclosed embodiment, a cascaded architecture is used to integrate the first sub-model and the second sub-model to achieve the generation of a digital human video from a reference video.

[0049] The technical solution of the disclosed embodiment effectively separates the complexity of the video generation task by constructing a two-stage generation mechanism of the first sub-model and the second sub-model, avoiding the distortion or missing details that may be caused by one-step generation, thereby significantly improving the quality of the generated video. By constructing sub-models with different structures, the two sub-models are each focused on different generation tasks, the optimization goals are clear, and the task adaptation ability of the model is improved. By integrating the trained first sub-model and the second sub-model to generate a digital human driving model, direct generation from reference images to driving videos is achieved, the process of digital human driving is simplified, and the production efficiency of digital human videos is improved. By combining the synthetic sequence and speech feature markers, it can ensure that the driving human body movement and audio rhythm are consistent, thereby improving the generation quality of digital human videos.

[0050] In some embodiments, determining a first synthetic sequence based on a reference video includes: determining a feature marker sequence of the ith video segment and a feature marker sequence of its preceding video segment based on the reference video, where i is an integer not less than 1; determining a speech feature marker sequence corresponding to the ith video segment based on the reference video; generating a first synthetic sequence for the ith video segment by combining the feature marker sequence of the ith video segment and a feature marker sequence of its preceding video segment, as well as a speech feature marker sequence corresponding to the ith video segment; the ith video segment is used to predict the ith video segment in the first driving video.

[0051] In the disclosed embodiment, the preceding video segment of the ith video segment refers to the (i-1)th video segment. In particular, the preceding video segment of the ith video segment has the same length as the ith video segment.

[0052] In the disclosed embodiment, the process of determining the feature marker sequence of the ith video segment can first use a video processing tool to divide the reference video into multiple video segments according to a fixed time length, and at the same time, divide the voice of the reference video into multiple voice segments according to the same time length. Then, starting from the second video segment, each video segment is selected in turn as the ith video segment. Next, a pre-trained deep learning model can be used to extract features from each frame of the ith video segment, and then the visual features of all frames in the ith video segment are aggregated to generate a feature marker sequence for the ith video segment. The above is only an exemplary description and is not intended to limit all possible situations for determining the feature marker sequence of the ith video segment, but it is not exhaustive here.

[0053] In the disclosed embodiment, the process of determining the feature tag sequence of the preceding video segment of the i-th video segment may first use a pre-trained deep learning model to extract features from each frame of the preceding video segment, and then aggregate the visual features of all frames in the preceding video segment to generate the feature tag sequence of the preceding video segment. The above is only an exemplary description and is not intended to limit all possible situations for determining the feature tag sequence of the preceding video segment of the i-th video segment, but it is not exhaustive here.

[0054] In the disclosed embodiment, the process of determining the speech feature marker sequence corresponding to the ith video segment based on the reference video can first determine the corresponding time period in the reference video according to the time range of the ith video segment in the driving video, and then use the separation tool to separate the audio and video of the reference video to extract the corresponding speech segment. Subsequently, a deep learning model can be used to extract high-level speech features, including speech content, semantics, and emotional information. Next, the extracted speech features are output as vectors, organized in sequence form, aligned with the time segments, and a speech feature marker sequence is generated. The above is only an exemplary description and is not intended to limit all possible situations for determining the speech feature marker sequence, but it is not exhaustive here.

[0055] In the disclosed embodiment, the process of generating the first synthetic sequence of the ith video segment can splice the feature marker sequence of the ith video segment and the feature marker sequence of its preceding video segment and the speech feature marker sequence in the temporal dimension, thereby generating the first synthetic sequence of the ith video segment. In particular, in order to fully learn the correlation and timing information between multimodal features, a feature fusion network can also be used to process the spliced ​​first synthetic sequence of the previous step. Exemplarily, a sequence model based on a convolutional neural network or a model based on an attention mechanism can be used to perform feature fusion modeling. The above is only an exemplary description and is not intended to limit all possible situations for generating the first synthetic sequence of the ith video segment, except that they are not exhaustive here.

[0056] In this way, by combining the feature marker sequence of the previous segment, the generated video is ensured to be natural and smooth in the time dimension. By combining the speech feature marker sequence, the action of the generated video is ensured to be highly matched with the speech rhythm and semantics. By using the segment-level feature marker sequence, the appearance and dynamic features of the target person in the generated video are guaranteed to be consistent to avoid distortion. The segmented generation strategy is suitable for video generation tasks of different lengths and complexities, and at the same time enhances the diversity of the generated results. The introduction of the feature marker sequence and segmented generation simplifies the calculation process and improves the generation efficiency.

[0057] In some embodiments, determining a feature marker sequence of the ith video clip based on a reference video includes: determining a first video feature marker sequence of the ith video clip based on the reference video; wherein the first video feature marker sequence is generated based on a plurality of consecutive images in the reference video; generating a first noise feature marker sequence for the ith video clip; generating a first head feature marker sequence for the ith video clip; and generating a feature marker sequence of the ith video clip based on the first video feature marker sequence, the first noise feature marker sequence and the first head feature marker sequence.

[0058] In the disclosed embodiment, the process of determining the first video feature marker sequence of the i-th video clip based on the reference video can first divide the reference video into continuous frames at a fixed frame rate to form a frame sequence, and then set the frame length of the video clip to divide the frame sequence into multiple video clips. In particular, each frame in the video clip can be normalized, cropped, and other preprocessing to ensure the same size. Then, starting from the second video clip, each video clip is selected in turn as the i-th video clip. Then, a three-dimensional variational autoencoder (Three Dimensional Variational Autoencoder, 3D-VAE) can be used to encode the video clip, compress the spatiotemporal information of the video into a low-dimensional latent space, and generate a low-dimensional continuous latent vector. Then, a convolutional network can be used to divide the latent vector into blocks, each block can be regarded as a local feature representation of the video clip. Finally, all blocks can be vector quantized, and the continuous latent vector can be mapped to a discrete Token, thereby determining the first video feature marker sequence of the i-th video clip.

[0059] In the disclosed embodiment, the process of generating the first noise signature sequence for the i-th video segment may first determine the shape of the first noise signature sequence based on the first video signature sequence of the i-th video segment, and the shape may include a time step and a feature dimension, etc. Subsequently, the first noise signature sequence may be generated based on a preset noise distribution. Exemplarily, the noise distribution of the first noise signature sequence may be a Gaussian distribution.

[0060] In the disclosed embodiment, the process of generating the first head feature marker sequence for the i-th video clip can first perform face detection on each frame in the i-th video clip to extract the position information of the head, which may include the bounding box and key point coordinates of the face. Subsequently, the head position is represented as a binary image consistent with the video resolution, with the pixel value of the head area being 1 and the remaining pixel values ​​being 0, to generate a head position 01 mask. Next, the above-mentioned head position 01 mask is input into a convolutional neural network for encoding to obtain a high-dimensional feature representation of the head position. Then, the high-dimensional feature representation of the head position can be divided into blocks of fixed size, with each block as an independent feature unit. Finally, all blocks can be vector quantized to map the continuous potential vectors to discrete Tokens, thereby generating the first head feature marker sequence.

[0061] In the disclosed embodiment, the process of generating the characteristic marker sequence of the i-th video segment may first superimpose the first noise characteristic marker sequence on the first video characteristic marker sequence, and then perform channel splicing with the first head characteristic marker sequence to generate the characteristic marker sequence of the i-th video segment.

[0062] In this way, by introducing the head feature marker sequence, explicit head position and dynamic information is provided, so that the model can control the generation of the head area more accurately, effectively reducing the phenomenon of head shaking, incoherence or unnaturalness in the generated video, ensuring the stability of head movement, and helping to improve the realism of the overall video. By introducing the noise feature marker sequence, randomness is introduced into the generation process. This randomness helps the model generate more diverse results. Without changing the video generation logic, by adjusting or controlling the intensity of the noise feature, the generated video is allowed to have different details or slight changes, thereby improving the diversity and flexibility of the generated video. The fusion of the first video feature marker sequence, head features and noise features provides more comprehensive generation conditions and avoids the problems of local incoherence or unnatural details that may be caused by a single feature generation method. By separating the video features, head features and noise features, the control of the generation model is more convenient.

[0063] In some embodiments, determining a speech feature marker sequence corresponding to the ith video segment based on a reference video includes: determining a reference speech of the ith video segment based on the reference video; extracting and encoding features of the reference speech to generate a speech feature marker at a single time point; and arranging the speech feature markers in chronological order to generate a speech feature marker sequence for the ith video segment.

[0064] In the disclosed embodiment, the process of determining the reference speech of the i-th video segment based on the reference video can first determine the corresponding time period in the reference video according to the time range of the i-th video segment in the driving video, and then use a separation tool to separate the audio and video of the reference video, extract the corresponding speech segment, and use the speech segment as the reference speech of the i-th video segment.

[0065] In the disclosed embodiment, the process of extracting and encoding features of a reference speech and generating a speech feature marker for a single time point can first use a processor of a speech encoder to convert waveform data of the reference speech into a network input format, which is then input into a speech coding network to extract feature vectors for each time point. The feature vector corresponding to each time point is the speech feature marker for the single time point.

[0066] In the disclosed embodiment, the process of arranging the speech feature markers in chronological order to generate the speech feature marker sequence of the i-th video clip can traverse each time point and arrange the feature vectors of all time points in chronological order to form a speech feature marker sequence.

[0067] In this way, the reference video is used to determine the reference speech corresponding to the video clip, ensuring the consistency of the video and speech in the time dimension, thereby improving the coherence and realism of the video generation. By extracting the speech feature markers at a single time point, detailed speech features can be extracted, so that the generated video can accurately correspond to the speech signal.

[0068] In some embodiments, determining reference feature markers and facial feature markers based on a reference video includes: determining any frame image from the reference video; encoding any frame image to generate a reference feature marker; extracting facial information of a target object from any frame image, and encoding the facial information of the target object to generate a facial feature marker.

[0069] In the disclosed embodiment, the process of determining any frame image from the reference video can be performed by randomly selecting a frame to determine whether the clarity of the currently selected frame meets the requirements, and if it does not meet the requirements, a new random selection is made. The selection can also be made by specifying a time point. The above is only an exemplary description and is not intended to limit all possible situations for determining any frame image, but it is not exhaustive here.

[0070] In the disclosed embodiment, when encoding any frame of image and generating reference feature markers, the image may first be preprocessed by resizing, normalizing, tensorizing, etc. Subsequently, a convolutional neural network or an image coding model is selected to encode the image and generate a high-dimensional feature vector. Finally, the generated high-dimensional feature vector is used as a reference feature marker.

[0071] In the disclosed embodiment, the facial information of the target object is extracted from any frame image, and the facial information of the target object is encoded to generate a facial feature tag. The facial region of the target object can be first located using a facial detection algorithm, and then the detected face can be cropped to a standard size and normalized. Finally, a pre-trained facial feature extraction model can be used to encode the facial information to generate high-dimensional facial features, and the generated high-dimensional facial features can be used as facial feature tags.

[0072] In this way, encoding video frames to generate reference feature markers can capture the global information of the frame image and avoid missing key background or overall features. By focusing on feature extraction in the facial area, the accurate modeling of facial details is enhanced, which helps to generate more realistic and detailed facial dynamics.

[0073] In some embodiments, determining a second synthetic sequence based on a reference video includes: determining a feature marker sequence of the kth video segment and a feature marker sequence of its preceding video segment based on the reference video or the first driving video, where k is an integer not less than 1; determining a second synthetic sequence of the kth video segment according to the feature marker sequence of the kth video segment and a feature marker sequence of its preceding video segment; the kth video segment is used to predict the kth video segment in the second driving video.

[0074] In the embodiment of the present disclosure, the preceding video segment of the kth video segment refers to the (k-1)th video segment. In particular, the preceding video segment of the kth video segment has the same length as the kth video segment.

[0075] In the disclosed embodiment, the process of determining the feature marker sequence of the kth video segment can first use a video processing tool to divide the first drive video into multiple video segments according to a fixed time length, and at the same time, divide the voice of the first drive video into multiple voice segments according to the same time length. Then, starting from the second video segment, each video segment is selected in turn as the kth video segment. Next, a pre-trained deep learning model can be used to extract features from each frame of the kth video segment, and then the visual features of all frames in the kth video segment are aggregated to generate a feature marker sequence for the kth video segment. The above is only an exemplary description and is not intended to limit all possible situations for determining the feature marker sequence of the kth video segment, but it is not exhaustive here.

[0076] In the disclosed embodiment, the process of determining the feature tag sequence of the preceding video segment of the kth video segment may first use a pre-trained deep learning model to extract features from each frame of the preceding video segment, and then aggregate the visual features of all frames in the preceding video segment to generate the feature tag sequence of the preceding video segment. The above is only an exemplary description and is not intended to limit all possible situations for determining the feature tag sequence of the preceding video segment of the kth video segment, but it is not exhaustive here.

[0077] In the disclosed embodiment, the process of generating the second synthetic sequence of the kth video segment can splice the feature marker sequence of the kth video segment and the feature marker sequence of the preceding video segment in chronological order in the temporal dimension, thereby generating the second synthetic sequence of the kth video segment. In particular, in order to fully learn the correlation and timing information between multimodal features, a feature fusion network can also be used to process the spliced ​​second synthetic sequence of the previous step. Exemplarily, a sequence model based on a convolutional neural network or a model based on an attention mechanism can be used to perform feature fusion modeling. The above is only an exemplary description and is not intended to limit all possible situations for generating the second synthetic sequence of the kth video segment, but it is not exhaustive here.

[0078] In this way, by combining the feature markers of the first drive video or reference video and the preceding clip, the coherence, naturalness and style consistency of the generated video are achieved, and the action and dynamic characteristics of the video generation are guaranteed by utilizing the action drive information provided by the first drive video.

[0079] In some embodiments, determining a feature marker sequence of the kth video clip and a feature marker sequence of its preceding video clip based on a reference video or a first drive video includes: determining a second video feature marker sequence of the kth video clip based on the reference video or the first drive video; wherein the second video feature marker sequence is a video feature sequence after masking the face and hand regions of multiple consecutive images in the reference video or the first drive video; generating a three-dimensional grid feature marker for the face and hand regions for the kth video clip based on the reference video or the first drive video; generating a second noise feature marker sequence for the kth video clip; and generating a feature marker sequence for the kth video clip based on the second video feature marker sequence, the second noise feature marker sequence and the three-dimensional grid feature marker.

[0080] In the disclosed embodiment, the process of determining the second video feature marker sequence of the kth video segment based on the reference video or the first drive video can first divide the reference video or the first drive video into continuous frames at a fixed frame rate to form a frame sequence, and then set the frame length of the video segment to divide the frame sequence into multiple video segments. In particular, each frame in the video segment can be normalized, cropped, and other preprocessing to ensure that the size is consistent. Then, starting from the second video segment, each video segment is selected in turn as the kth video segment. Afterwards, a deep learning detector can be used to locate the face and hand areas, and for each frame image, a mask operation is applied to the detected face and hand areas, and black filling can be used. Then, feature markers can be extracted for each frame image after masking to generate a video feature sequence, and then the second video feature marker sequence of the kth video segment can be determined.

[0081] In the disclosed embodiment, the process of generating a 3D mesh feature tag for the face and hand regions for the kth video clip based on the reference video or the first driving video can first use the existing face and hand 3D estimation model to estimate the 3D mesh of the face and hand from each frame of the image. Subsequently, the generated 3D mesh representation is encoded and converted into a Token. Finally, each frame of the video clip is processed separately, and the 3D mesh Token of each frame is extracted, thereby generating a 3D mesh Token sequence of the entire video clip.

[0082] In the disclosed embodiment, the process of generating the second noise feature marker sequence for the kth video segment may first determine the shape of the second noise Token sequence based on the second video Token sequence of the kth video segment, and the shape may include a time step and a feature dimension, etc. Subsequently, the second noise Token sequence may be generated based on a preset noise distribution. Exemplarily, the noise distribution of the second noise Token sequence may be a Gaussian distribution.

[0083] In the disclosed embodiment, the process of generating the feature marker sequence of the kth video clip based on the second video feature marker sequence, the second noise feature marker sequence and the three-dimensional grid feature marker can first superimpose the second noise Token sequence on the second video Token sequence, and then perform channel splicing with the three-dimensional grid Token sequence to generate the feature marker sequence of the kth video clip.

[0084] In this way, by integrating the video feature sequence, the three-dimensional grid feature marker and the noise feature marker sequence, the feature marker sequence of the video clip is generated, ensuring the comprehensiveness and diversity of the generated results. Among them, the video feature sequence provides the dynamic background information of the reference video, the three-dimensional grid feature marker provides the geometric details of the local face and hands, and the noise feature marker introduces randomness, so that the generated results can not only retain the core information of the input, but also have certain controllable changes. By masking the face and hand areas, the interference of irrelevant areas on the generation is avoided, and the focus on the background and action changes is strengthened. At the same time, the fine modeling of the three-dimensional grid features ensures the coherence of dynamic details. The face and hand areas are finely modeled by the three-dimensional grid feature markers, capturing the geometric shape, action transformation and three-dimensional structure information of these local areas. Compared with the traditional generation method that relies only on planar features, this method greatly improves the modeling ability of complex actions, ensuring that the generated results can reflect the detailed features in the reference video or driving video, and have a high degree of restoration ability.

[0085] In some embodiments, the first sub-model includes N transformer layers, each transformer layer includes a first self-attention layer and a speech cross-attention layer; the first self-attention layer includes a first feature adapter and a second feature adapter, the first feature adapter is used to adjust the features of the corresponding object in the first target image according to the reference feature marker; the second feature adapter is used to adjust the facial features in the first target image according to the facial feature marker; the speech cross-attention layer is used to adjust the facial expressions in the target image according to the speech feature marker sequence so that the facial expressions match the speech content; wherein N is an integer not less than 2.

[0086] In this way, the reference features and facial features are processed by the first feature adapter and the second feature adapter respectively to ensure the content consistency and detail restoration of the generated video. The speech cross-attention layer is introduced to associate the speech features with the target image features to achieve speech-driven dynamic expression generation. The layer-by-layer processing mechanism helps to accurately model the complex relationship between speech, reference features and facial features, making the generated actions, expressions and backgrounds more natural and smooth.

[0087] In some embodiments, a first sub-model is trained based on a first synthesis sequence, a speech feature marker sequence, a reference feature marker and a facial feature marker, including: taking the first synthesis sequence, the reference feature marker and the facial feature marker as inputs of a first self-attention layer, and taking the speech feature marker sequence and the output of the first self-attention layer as inputs of a speech cross-attention layer, and training the first sub-model so that the difference between a first driving video output by the first sub-model and the reference video is less than a first preset range.

[0088] In the disclosed embodiment, the process of training the first sub-model can first use the first synthesis sequence, reference feature markers, facial feature markers, and speech feature marker sequences as input, forward propagate to calculate the first driving video, and simultaneously calculate the values ​​of different loss functions, and back propagate to update the model parameters. Afterwards, as the training progresses, the global features, facial details, and voice-driven dynamic expressions of the generated video are gradually optimized through a multi-layer Transformer, so that the generated video gradually approaches the reference video, and the difference is less than the first preset range. Here, the first preset range can be set according to the training accuracy or efficiency.

[0089] In this way, by training the first sub-model, the difference between the generated first driving video and the reference video in terms of global features, detail features and voice expression matching can be less than a first preset range, thus meeting the requirements for high-quality digital human driving video generation.

[0090] In some embodiments, the first synthesized sequence, reference feature markers, and facial feature markers are used as inputs to the first self-attention layer, and the speech feature marker sequence and the output of the first self-attention layer are used as inputs to the speech cross-attention layer, and the first sub-model is trained, including: inputting the first synthesized sequence into the first self-attention layer in the first Transformer layer; inputting the output of the first self-attention layer and the reference feature markers into the first feature adapter; inputting the output of the first feature adapter and the facial feature markers into the second feature adapter; inputting the output of the second feature adapter and the speech feature marker sequence into the speech cross-attention layer to obtain an updated first synthesized sequence output by the speech cross-attention layer, and the updated first synthesized sequence is used as the first synthesized sequence of the input to the next Transformer layer, until the last Transformer layer outputs the first driving video.

[0091] In the disclosed embodiment, the process of inputting the first synthetic sequence into the first self-attention layer of the first Transformer layer can calculate the relationship between the features in the first synthetic sequence through the self-attention mechanism and capture their global dependencies.

[0092] In the disclosed embodiment, the process of inputting the output of the first self-attention layer and the reference feature marker into the first feature adapter can adjust the global features such as the background and posture of the generated sequence to align them with the reference feature marker.

[0093] In the disclosed embodiment, the process of inputting the output of the first feature adapter and the facial feature markers to the second feature adapter can optimize the features of the facial area such as the geometry of the facial features and facial details, so that the facial features are aligned with the facial feature markers.

[0094] In the disclosed embodiment, the process of inputting the output of the second feature adapter and the speech feature label sequence into the speech cross-attention layer can use the cross-attention mechanism to associate the speech features with facial and global features, and dynamically adjust the expression details in the generated sequence to match the speech content.

[0095] In the disclosed embodiment, the process of obtaining the first synthetic sequence output by the speech cross-attention layer and used as the input of the next Transformer layer can iteratively train the first sub-model and optimize the generated sequence layer by layer until the last Transformer layer outputs the feature representation of the first driving video. Finally, the generated Tokens are converted into actual video frames through the decoder to obtain the first driving video.

[0096] In this way, by iteratively training the first sub-model, a fully optimized first driving video can be output to ensure the global consistency of the generated video, the accuracy of facial details, and the naturalness of voice driving.

[0097] In some embodiments, the second sub-model includes N Transformer layers, each Transformer layer includes a second self-attention layer; the second self-attention layer includes a third feature adapter and a fourth feature adapter, the third feature adapter is used to adjust the features of the corresponding object in the second target image according to the reference feature markers; the fourth feature adapter is used to adjust the facial features in the second target image according to the facial feature markers, where N is an integer not less than 2.

[0098] In this way, the third feature adapter and the fourth feature adapter process the reference features and facial features respectively to ensure the content consistency and detail restoration of the generated video. The layer-by-layer processing mechanism helps to accurately model the complex relationship between the reference features and facial features, making the generated actions, expressions and backgrounds more natural and smooth.

[0099] In some embodiments, training the second sub-model based on the second synthetic sequence, the reference feature markers, and the facial feature markers includes: using the second synthetic sequence, the reference feature markers, and the facial feature markers as inputs of the second self-attention layer, and training the second sub-model so that the difference between the second driving video output by the second sub-model and the reference video is less than a second preset range. Here, the second preset range can be set according to training accuracy or efficiency.

[0100] In the disclosed embodiment, the process of inputting the second synthetic sequence, the reference feature markers and the facial feature markers into the second self-attention layer of the first Transformer layer can calculate the relationship between the features through the self-attention mechanism and capture their global dependencies.

[0101] In this way, by iteratively training the second sub-model, the first driving video can be improved, and then a fully optimized second driving video can be output to ensure the global consistency of the generated video, the accuracy of facial details, and the naturalness of voice driving.

[0102] Figure 2 The architecture diagram of the first stage neural network structure is shown as Figure 2 As shown, the first stage neural network structure includes a first sub-model, specifically including a first feature encoding part, a second feature encoding part and a video generation part.

[0103] Among them, the first feature encoding part uses the VAE encoder to encode the preamble segment to generate the preamble token; uses the VAE encoder to encode the video segment to generate the video token; generates the head position based on the video segment, and then uses the head position encoder to encode it to generate the head token; superimposes the generated noise token with the video token, and then performs channel splicing with the head token. Finally, the first feature encoding part outputs the preamble token and the spliced ​​video token.

[0104] The second feature encoding part uses the input audio to generate an audio token; after determining the reference image from the reference video, it is encoded using the VAE encoder to generate a reference token; after extracting the face from the reference image, it is encoded using the facial encoder to generate a facial token. Finally, the second feature encoding part outputs an audio token, a reference token, and a facial token.

[0105] Among them, the video generation part integrates the preamble Token, the spliced ​​video Token, the audio Token, the reference Token and the face Token output by the first feature encoding part and the second feature encoding part to obtain an input Token sequence; the input Token sequence is sent to the self-attention layer and the audio cross-attention layer in the N Transformer layers; the self-attention layer includes a reference adapter and a face adapter, the reference adapter takes the reference Token as a parameter, and the face adapter takes the face Token as a parameter; after passing through the Transformer layer, a denoised video Token is output; finally, after passing through the VAE decoder, a denoised video, i.e., the first driving video, is finally generated.

[0106] Figure 3 The architecture diagram of the second stage neural network structure is shown as Figure 3 As shown, the second stage neural network structure includes a feature encoding part, a second sub-model part and a video generation part.

[0107] Among them, after the feature encoding part generates the face and hand masks, it performs channel splicing with the denoised video generated by the first stage diffusion transformer, and generates a video token after encoding with the VAE encoder; based on the audio drive results, the hand and face estimation results are used to generate a 3D grid, which is then encoded using the grid encoder to generate a grid token; the generated noise token is superimposed with the video token and channel spliced ​​with the grid token. Finally, the feature encoding part outputs the spliced ​​video token.

[0108] Among them, the second sub-model part Figure 3Not shown in detail, it includes N Transformer layers, each Transformer layer includes a self-attention layer; wherein the self-attention layer includes a reference adapter and a face adapter, the reference adapter takes the reference Token as a parameter, and the face adapter takes the face Token as a parameter; after passing through the Transformer layer, the denoised video Token is output.

[0109] The video generation part Figure 3 Not shown in detail, it includes a VAE decoder; wherein the denoised video Token output by the second sub-model part finally passes through the VAE decoder to finally generate a denoised video, that is, the second driving video.

[0110] It should be understood that Figure 2 and Figure 3 The schematic diagram shown is only exemplary and not restrictive, and it is scalable, and those skilled in the art can Figure 2 and Figure 3 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0111] The present disclosure provides a method for driving a digital human. Figure 4 : is a flow chart of a digital human driving method according to an embodiment of the present disclosure, and the digital human driving method can be applied to a digital human driving device. The digital human driving device is located in an electronic device. The electronic device includes but is not limited to a fixed device and / or a mobile device. For example, a fixed device includes but is not limited to a server, and the server can be a cloud server or an ordinary server. For example, a mobile device includes but is not limited to: a mobile phone, a tablet computer. In some possible implementations, the digital human driving method can also be implemented by a processor calling a computer-readable instruction stored in a memory. Figure 4 As shown, the digital human driving method includes:

[0112] S401, determining a digital human reference feature mark and a digital human facial feature mark based on a reference image;

[0113] S402, extracting a driving speech feature tag sequence based on the driving speech;

[0114] S403, determining a third synthesis sequence for generating each video segment based on the reference image and the driving voice;

[0115] S404, inputting the third synthesis sequence, the driving voice feature mark sequence, the digital human reference feature mark and the digital human facial feature mark into the first sub-model of the digital human driving model to obtain a third driving video;

[0116] S405, determining a fourth synthesis sequence for generating each video segment based on the third driving video;

[0117] S406, inputting the fourth synthesis sequence, the digital human reference feature mark and the digital human facial feature mark into the second sub-model of the digital human driving model to obtain a fourth driving video;

[0118] S407: Generate a digital human video based on the fourth driving video corresponding to all the video clips.

[0119] In the disclosed embodiment, the reference image refers to an externally input reference image, which at least includes a person for the digital human video to be generated, and may also include a background for the digital human video to be generated.

[0120] In the disclosed embodiment, the digital human reference feature tag refers to the overall picture feature information extracted from the reference image. The digital human reference token includes at least a digital human body reference token and may also include a digital human background reference token. The digital human reference token can ensure that the person in the generated video is consistent with the person in the reference image, including information such as facial structure and head position, and can also ensure that the background in the generated video is consistent with the background in the reference image.

[0121] In the disclosed embodiment, the digital human facial feature token is a token generated by processing the facial area of ​​the person in the reference image, which contains detailed information of the facial features. The digital human facial token can ensure that the facial features in the generated video match the facial features in the audio reference image.

[0122] In the embodiments of the present disclosure, the driving voice refers to the voice used to drive or control the movements and expressions of the digital human. The driving voice can be used as input, and through a series of technical processing, the digital human can generate matching facial expressions, lip shapes, and body movements according to the content, rhythm, and intonation of the voice.

[0123] In the disclosed embodiment, the driving voice feature tag sequence refers to the speech feature sequence extracted from the driving voice and the token sequence generated after processing. The driving voice token sequence contains the timing rhythm and voice content information of the driving voice, which can drive the overall movement rhythm and facial expression of the digital human.

[0124] In the disclosed embodiment, the third synthesis sequence can be input into the first sub-model to generate a third driving video.

[0125] In the disclosed embodiment, the fourth synthesized sequence can be input into the training of the second sub-model to generate a fourth driving video.

[0126] In the disclosed embodiment, the digital human video refers to a video in which a person in a reference image can make corresponding lip movements and body movements according to the content of the driving voice.

[0127] In the disclosed embodiment, a neural network can be first used to extract global features of a reference image, and the extracted features can be represented as a token in the form of a vector as a digital human reference token. Afterwards, a neural network can be used to extract the facial area in the reference image, and the dynamic features of the facial area can be extracted, and the extracted facial dynamic features can be represented as a token as a digital human facial token. Furthermore, a neural network can be used to extract voice features in a driving voice, and then the extracted voice features can be divided into blocks in chronological order to generate a corresponding voice token sequence. Furthermore, a neural network can be used to extract character features of a reference image, and the driving voice can be used to determine character actions. After applying the character actions to the character features, a third synthetic sequence can be determined.

[0128] In the disclosed embodiment, a key point detection model can be used to detect the face and hand regions from the frame sequence of the third driving video and mask the region, and then a neural network can be used to perform dimensionality reduction processing on the frame sequence of the third driving video, and the video features after dimensionality reduction are arranged in chronological order. In addition, a 3D estimation model is used to generate 3D mesh tokens of the face and hand. A fourth synthetic sequence is generated based on the dimensionality reduced video features and the 3D mesh tokens.

[0129] In the disclosed embodiment, all fourth driving videos generated based on the reference image and the driving voice are combined to obtain a digital human video.

[0130] The technical solution of the disclosed embodiment enables the generated digital human video to naturally present voice-driven expressions and movements by fusing reference images, facial features, and driving voice features. The first sub-model and the second sub-model are generated in stages to gradually optimize the details and consistency of the generated video, ensuring higher generation quality. By generating and synthesizing each segment, the coherence of movements and expressions between video segments is ensured, enhancing the fluency of the generated video. The reference image provides global and facial features, so that the generated video can maintain visual and identity features consistent with the reference image.

[0131] In some embodiments, determining a digital human reference feature marker and a digital human facial feature marker based on a reference image includes: feature encoding the reference image to generate a digital human reference feature marker; extracting the digital human's facial information from the reference image, and encoding the facial information to generate a digital human facial feature marker.

[0132] In the disclosed embodiment, the process of encoding the reference image and generating the digital human reference feature tag may first pre-process the reference image by resizing, normalizing, tensorizing, etc. Subsequently, a convolutional neural network or image coding model is selected to encode the image and generate a high-dimensional feature vector. Finally, the generated high-dimensional feature vector is used as the digital human reference token.

[0133] In the disclosed embodiment, the facial information of the digital person is extracted from the reference image, and the facial information is encoded to generate the facial feature tag of the digital person. The facial area of ​​the target object can be located by using a facial detection algorithm, and then the detected face can be cropped to a standard size and normalized. Finally, the facial information can be encoded using a pre-trained facial feature extraction model to generate high-dimensional facial features, and the generated high-dimensional facial features can be used as the digital person's facial token.

[0134] In this way, by encoding global information with reference feature markers and local information with facial feature markers, complete modeling of the overall and detailed features of the digital human can be achieved. The independent extraction and encoding of facial feature markers enhances the capture of facial geometry and expression details, ensuring accurate restoration of the facial characteristics of the generated digital human. The separate processing of reference feature markers and facial feature markers helps to optimize the consistency and naturalness of global and local features when generating videos.

[0135] In some embodiments, extracting a driving speech feature marker sequence based on the driving speech includes: extracting and encoding features of the driving speech to generate a driving speech feature marker at a single time point; arranging the driving speech feature markers in chronological order to generate a driving speech feature marker sequence.

[0136] In the disclosed embodiment, the process of extracting and encoding the features of the driving speech and generating the feature marker of the driving speech at a single time point can first use the processor of the speech encoder to convert the waveform data of the driving speech into a network input format, and then input it into the speech coding network to extract the feature vector of each time point. The feature vector corresponding to each time point is the driving speech Token at the single time point.

[0137] In the disclosed embodiment, the process of arranging the driving speech feature markers in chronological order to generate a driving speech feature marker sequence can traverse each time point and arrange the feature vectors of all time points in chronological order to form a driving speech Token sequence.

[0138] In this way, by extracting speech feature markers at each time point, the dynamic changes of speech can be accurately captured. The speech feature marker sequence is generated in chronological order to ensure the coherence of the speech-driven process and the synchronization with video generation. Feature extraction and encoding compress the speech information, which is convenient for subsequent model processing and improves generation efficiency. The speech feature sequence can better guide the digital human's movements and expressions, making the generated video more natural and realistic.

[0139] In some embodiments, a third synthetic sequence for generating each video segment is determined based on a reference image and a driving voice, including: when predicting the j-th video segment of the third driving video, a feature marker sequence of the j-th video segment and a feature marker sequence of its preceding video segment are determined based on the reference image and the driving voice, where j is an integer not less than 1; and a third synthetic sequence of the j-th video segment is generated by combining the feature marker sequence of the j-th video segment and a feature marker sequence of its preceding video segment, as well as a driving voice feature marker sequence corresponding to the j-th video segment.

[0140] In the embodiment of the present disclosure, the preceding video segment of the jth video segment refers to the (j-1)th video segment. In particular, the preceding video segment of the jth video segment has the same length as the jth video segment.

[0141] In the disclosed embodiment, the process of determining the feature tag sequence of the jth video segment can first predict the character action using the driving voice, then use the pre-trained deep learning model to extract features from the reference image, combine the extracted features with the predicted character action, and then generate the feature tag sequence of the jth video segment. The above is only an exemplary description and is not intended to limit all possible situations for determining the feature tag sequence of the jth video segment, but it is not exhaustive here.

[0142] In the disclosed embodiment, the process of determining the feature tag sequence of the preceding video segment of the jth video segment can first use a pre-trained deep learning model to extract features from each frame of the generated preceding video segment, and then aggregate the visual features of all frames in the preceding video segment to generate the feature tag sequence of the preceding video segment. The above is only an exemplary description and is not intended to limit all possible situations for determining the feature tag sequence of the preceding video segment of the jth video segment, but it is not exhaustive here.

[0143] In the disclosed embodiment, the process of generating the third synthetic sequence of the jth video segment can splice the feature marker sequence of the jth video segment, the feature marker sequence of the preceding video segment, and the speech feature marker sequence in chronological order in the temporal dimension, thereby generating the third synthetic sequence of the jth video segment. In order to fully learn the correlation and timing information between multimodal features, a feature fusion network can also be used to process the spliced ​​third synthetic sequence of the previous step. Exemplarily, a sequence model based on a convolutional neural network or a model based on an attention mechanism can be used to perform feature fusion modeling. The above is only an exemplary description and is not intended to limit all possible situations for generating the third synthetic sequence of the jth video segment, but it is not exhaustive here.

[0144] In this way, by combining the feature tag sequence of the current video clip with the previous video clip, the coherence of the actions and expressions between the generated video clips is ensured. By combining the driving voice feature tag sequence, the generated video clip can more accurately reflect the dynamic changes of voice driving. The previous clip features are used to provide contextual information to improve the accuracy and naturalness of the current clip generation.

[0145] In some embodiments, a feature marker sequence of the jth video segment and a feature marker sequence of its preceding video segment are determined based on a reference image and a driving voice, including: when j=1, determining the feature marker sequence of the 1st video segment based on the reference image; setting the feature marker sequence of the preceding video segment of the 1st video segment to a default value; when j≥2, determining the feature marker sequence of the jth video segment based on the reference image and a feature sequence of the driving voice corresponding to the jth video segment; and using the determined feature marker sequence of the j-1th video segment as the feature marker sequence of the preceding video segment of the jth video segment.

[0146] In the disclosed embodiment, when j=1, the feature tag sequence of the first video clip is determined based on the reference image; in the process of setting the feature tag sequence of the preceding video clip of the first video clip to the default value, since j=1 is equivalent to the first reasoning, there is no preceding frame, so the preceding Token can be set to the default value. Exemplarily, the preceding Token can be set to zero, or to other values ​​according to actual conditions. The above is only an exemplary description and is not intended to limit all possible situations of the preceding Token default value, but it is not exhaustive here.

[0147] In the embodiment of the present disclosure, when j≥2, the process of determining the feature marker sequence of the jth video clip based on the reference image and the driving voice feature sequence corresponding to the jth video clip can first use the driving voice to predict the character's movements, and then use the pre-trained deep learning model to extract features from the reference image, combine the extracted features with the predicted character movements, and then generate the feature marker sequence of the jth video clip.

[0148] In the disclosed embodiment, in the process of using the feature marker sequence of the j-1th video segment that has been determined as the feature marker sequence of the preceding video segment of the j-th video segment, since j≥2 is equivalent to not being the first inference and there are video segments that have been inferred, the preceding Token can be set as the feature marker sequence of the j-1th video segment.

[0149] In this way, by setting a default value for the leading feature tag sequence of the first video clip, the processing flow is simplified and the stability of the generation is ensured. For non-first video clips, the determined feature tag sequence of the leading clip is used as context information to enhance the continuity and temporal consistency between video clips. By referencing the leading clip features frame by frame, the generated video has a more natural and coherent dynamic performance.

[0150] In some embodiments, determining a feature marker sequence of the jth video segment based on a reference image and a driving voice feature sequence corresponding to the jth video segment includes: determining a third video feature marker sequence of the jth video segment based on the reference image and the driving voice feature sequence corresponding to the jth video segment; generating a third noise feature marker sequence for the jth video segment; generating a second head feature marker sequence for the jth video segment; and generating a feature marker sequence of the jth video segment based on the third video feature marker sequence, the third noise feature marker sequence, and the second head feature marker sequence.

[0151] In the disclosed embodiment, the process of determining the third video feature tag sequence of the jth video segment based on the reference image and the driving voice feature sequence corresponding to the jth video segment can first combine the character action predicted according to the driving voice with the character features extracted according to the reference image to generate a low-dimensional continuous latent vector. Then, a convolutional network can be used to divide the latent vector into blocks, each of which can be regarded as a local feature representation. Finally, all blocks can be vector quantized, and the continuous latent vector can be mapped to discrete Tokens, thereby determining the first video Token sequence of the jth video segment.

[0152] In the embodiment of the present disclosure, the process of generating the third noise feature marker sequence for the jth video segment may first determine the shape of the third noise Token sequence based on the third video Token sequence of the jth video segment, and the shape may include a time step and a feature dimension, etc. Subsequently, the third noise Token sequence may be generated based on a preset noise distribution. Exemplarily, the noise distribution of the third noise Token sequence may be a Gaussian distribution.

[0153] In the disclosed embodiment, the process of generating the second head feature marker sequence for the jth video clip can first perform facial detection on each frame in the jth video clip to extract the position information of the head, which may include the bounding box and key point coordinates of the face. Subsequently, the head position is represented as a binary image consistent with the video resolution, with the pixel value of the head area being 1 and the remaining pixel values ​​being 0, to generate a head position 01 mask. Next, the above-mentioned head position 01 mask is input into a convolutional neural network for encoding to obtain a high-dimensional feature representation of the head position. Then, the high-dimensional feature representation of the head position can be divided into blocks of fixed size, with each block as an independent feature unit. Finally, all blocks can be vector quantized to map continuous potential vectors to discrete Tokens, thereby generating a second head Token sequence.

[0154] In the embodiment of the present disclosure, the process of generating the characteristic marker sequence of the jth video clip based on the third video characteristic marker sequence, the third noise characteristic marker sequence and the second head characteristic marker sequence can first superimpose the third noise Token sequence on the third video Token sequence, and then perform channel splicing with the second head Token sequence to generate the characteristic marker sequence of the jth video clip.

[0155] In this way, by combining video features, noise features and head features, global and local features are captured comprehensively, and the generated feature marker sequence is more comprehensive and accurate. The introduction of noise feature marker sequence enhances the robustness of the model, reduces errors in the generation process, and improves the detail quality of the results. By generating a head feature marker sequence separately, the capture of head posture and dynamic changes is enhanced, making the generated video more natural and realistic. Combined with the reference image and driving speech features, it ensures that the generated feature marker sequence can accurately reflect the dynamic changes of the driving speech and the corresponding visual performance.

[0156] In some embodiments, a fourth synthetic sequence for generating each video clip is determined based on the third driving video, including: determining a fourth video feature marker sequence of the jth video clip based on the third driving video, the fourth video feature marker sequence being a video feature sequence after masking the face and hand regions in the jth video clip; determining a three-dimensional grid feature marker for the face and hand regions of the jth video clip based on the third driving video; determining a fourth noise feature marker sequence for the jth video clip; and determining a fourth synthetic sequence for the jth video clip based on the three-dimensional grid feature marker, the fourth noise feature marker sequence and the fourth video feature marker sequence of the jth video clip.

[0157] In the disclosed embodiment, the fourth video feature marker sequence of the jth video segment is determined based on the third driving video. The face and hand regions can be located using a deep learning detector first, and a mask operation can be applied to the detected face and hand regions for each frame image. Black filling can be used. Then, feature markers can be extracted from each masked frame image to generate a video feature sequence, thereby determining the fourth video Token sequence of the jth video segment.

[0158] In the disclosed embodiment, the process of determining the 3D mesh feature labels for the face and hand regions of the jth video clip based on the third driving video can use the existing 3D estimation model for the face and hand to estimate the 3D mesh of the face and hand from each frame of the image. Subsequently, the generated 3D mesh representation is encoded and converted into a Token. Finally, each frame of the video clip is processed separately, and the 3D mesh Token of each frame is extracted, thereby generating a 3D mesh Token sequence of the entire video clip.

[0159] In the embodiment of the present disclosure, the process of determining the fourth noise feature marker sequence for the jth video segment may first determine the shape of the fourth noise Token sequence based on the fourth video Token sequence of the jth video segment, and the shape may include a time step and a feature dimension, etc. Subsequently, the fourth noise Token sequence may be generated based on a preset noise distribution. Exemplarily, the noise distribution of the fourth noise Token sequence may be a Gaussian distribution.

[0160] In the disclosed embodiment, the process of determining the fourth synthetic sequence of the jth video segment based on the three-dimensional grid feature marker, the fourth noise feature marker sequence and the fourth video feature marker sequence of the jth video segment can firstly superimpose the fourth noise Token sequence on the fourth video Token sequence, and then perform channel splicing with the three-dimensional grid Token sequence to generate the feature marker sequence of the jth video segment.

[0161] In this way, by marking the 3D mesh features of the face and hand areas, the dynamic performance and detail accuracy of these key areas are improved. By masking the face and hand areas, the global and local features are separated to ensure that the generated features are clearer and more accurate. By combining the 3D mesh features and noise feature markers, the authenticity and naturalness of the generated video clips are enhanced. By introducing noise feature markers, the generalization ability of the model is enhanced, noise interference is reduced, and the stability and diversity of video generation are improved.

[0162] The present disclosure provides a digital human driving model generation device, such as Figure 5As shown, the device may include: a feature extraction module 501, used to determine a first synthesis sequence, a second synthesis sequence, a speech feature marker sequence, a reference feature marker and a facial feature marker based on a reference video; a first training module 502, used to train a first sub-model based on the first synthesis sequence, the speech feature marker sequence, the reference feature marker and the facial feature marker, so that the first sub-model outputs a first driving video; a second training module 503, used to train a second sub-model based on the second synthesis sequence, the reference feature marker and the facial feature marker, so that the second sub-model outputs a second driving video; a model generation module 504, used to generate a digital human driving model based on the trained first sub-model and the second sub-model.

[0163] In some embodiments, the feature extraction module 501 includes: a first video feature submodule, used to determine the feature marker sequence of the ith video clip and the feature marker sequence of its preceding video clip based on a reference video, where i is an integer not less than 1; a speech feature determination submodule, used to determine the speech feature marker sequence corresponding to the ith video clip based on the reference video; a first sequence generation submodule, used to combine the feature marker sequence of the ith video clip and the feature marker sequence of its preceding video clip, as well as the speech feature marker sequence corresponding to the ith video clip, to generate a first synthetic sequence for the ith video clip, where the ith video clip is used to predict the ith video clip in the first driving video.

[0164] In some embodiments, the first video feature submodule is used to: determine a first video feature marker sequence of the i-th video clip based on a reference video; wherein the first video feature marker sequence is generated based on multiple consecutive images in the reference video; generate a first noise feature marker sequence for the i-th video clip; generate a first head feature marker sequence for the i-th video clip; and generate a feature marker sequence of the i-th video clip based on the first video feature marker sequence, the first noise feature marker sequence and the first head feature marker sequence.

[0165] In some embodiments, the first video feature submodule is further used to: determine the reference speech of the i-th video segment based on the reference video; extract and encode the reference speech to generate a speech feature marker at a single time point; arrange the speech feature markers in chronological order to generate a speech feature marker sequence for the i-th video segment.

[0166] In some embodiments, the feature extraction module 501 also includes: an image determination submodule, used to determine any frame image from a reference video; an image encoding submodule, used to encode any frame image to generate a reference feature marker; a facial feature submodule, used to extract facial information of a target object from any frame image, and encode the facial information of the target object to generate a facial feature marker.

[0167] In some embodiments, the feature extraction module 501 also includes: a second video feature submodule, used to determine the feature marker sequence of the kth video clip and the feature marker sequence of its preceding video clip based on the reference video or the first driving video, k is an integer not less than 1; a second sequence generation submodule, used to determine the second synthetic sequence of the kth video clip according to the feature marker sequence of the kth video clip and the feature marker sequence of its preceding video clip; the kth video clip is used to predict the kth video clip in the second driving video.

[0168] In some embodiments, the second video feature submodule is used to: determine a second video feature marker sequence of the kth video clip based on a reference video or a first drive video; wherein the second video feature sequence is a video feature sequence after masking the face and hand areas of multiple consecutive images in the reference video or the first drive video; generate a three-dimensional grid feature marker for the face and hand areas for the kth video clip based on the reference video or the first drive video; generate a second noise feature marker sequence for the kth video clip; and generate a feature marker sequence for the kth video clip based on the second video feature marker sequence, the second noise feature marker sequence and the three-dimensional grid feature marker.

[0169] In some embodiments, the first sub-model includes N Transformer layers, each Transformer layer includes a first self-attention layer and a speech cross-attention layer; the first self-attention layer includes a first feature adapter and a second feature adapter, the first feature adapter is used to adjust the features of the corresponding object in the first target image according to the reference feature marker; the second feature adapter is used to adjust the facial features in the first target image according to the facial feature marker; the speech cross-attention layer is used to adjust the facial expressions in the target image according to the speech feature marker sequence so that the facial expressions match the speech content; wherein N is an integer not less than 2.

[0170] In some embodiments, the first training module 502 includes: a first training sub-module, used to take the first synthesis sequence, reference feature markers and facial feature markers as inputs of a first self-attention layer, and take the speech feature marker sequence and the output of the first self-attention layer as inputs of a speech cross-attention layer, to train the first sub-model so that the difference between the first driving video output by the first sub-model and the reference video is less than a first preset range.

[0171] In some embodiments, a first training submodule is used to: input the first synthesized sequence into the first self-attention layer in the first Transformer layer; input the output of the first self-attention layer and the reference feature label into the first feature adapter; input the output of the first feature adapter and the facial feature label into the second feature adapter; input the output of the second feature adapter and the speech feature label sequence into the speech cross-attention layer to obtain an updated first synthesized sequence output by the speech cross-attention layer, and the updated first synthesized sequence is used as the first synthesized sequence of the input of the next Transformer layer until the last Transformer layer outputs the first driving video.

[0172] In some embodiments, the second sub-model includes N Transformer layers, each Transformer layer includes a second self-attention layer; the second self-attention layer includes a third feature adapter and a fourth feature adapter, the third feature adapter is used to adjust the features of the corresponding object in the second target image according to the reference feature markers; the fourth feature adapter is used to adjust the facial features in the second target image according to the facial feature markers, where N is an integer not less than 2.

[0173] In some embodiments, the second training module 503 includes: a second training sub-module, used to take the second synthetic sequence, reference feature markers and facial feature markers as inputs of the second self-attention layer, and train the second sub-model so that the difference between the second driving video output by the second sub-model and the reference video is less than a second preset range.

[0174] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0175] The digital human driving model generation device in the embodiment of the present disclosure can improve the quality and efficiency of the digital human driving model in generating digital human videos.

[0176] The embodiment of the present disclosure provides a digital human driving device, such as Figure 6As shown, the device may include: a feature marking module 601, which is used to determine a digital human reference feature marking and a digital human facial feature marking based on a reference image; a speech extraction module 602, which is used to extract a driving voice feature marking sequence based on a driving voice; a first sequence module 603, which is used to determine a third synthetic sequence for generating each video clip based on the reference image and the driving voice; a first generation module 604, which is used to input the third synthetic sequence, the driving voice feature marking sequence, the digital human reference feature marking and the digital human facial feature marking into a first sub-model of a digital human driving model to obtain a third driving video; a second sequence module 605, which is used to determine a fourth synthetic sequence for generating each video clip based on the third driving video; a second generation module 606, which is used to input the fourth synthetic sequence, the digital human reference feature marking and the digital human facial feature marking into a second sub-model of the digital human driving model to obtain a fourth driving video; and a video generation module 607, which is used to generate a digital human video based on the fourth driving video corresponding to all video clips.

[0177] In some embodiments, the feature marking module 601 includes: a reference feature marking submodule, which is used to perform feature encoding on a reference image to generate a reference feature mark of a digital person; and a facial feature marking submodule, which is used to extract facial information of a digital person from a reference image and encode the facial information to generate a facial feature mark of a digital person.

[0178] In some embodiments, the speech extraction module 602 includes: a speech feature extraction submodule, which is used to extract and encode the driving speech features and generate a driving speech feature marker at a single time point; and a speech feature sorting submodule, which is used to arrange the driving speech feature markers in chronological order and generate a driving speech feature marker sequence.

[0179] In some embodiments, the first sequence module 603 includes: a third video feature submodule, which is used to determine the feature marker sequence of the jth video segment and the feature marker sequence of its preceding video segment based on the reference image and the driving voice when predicting the jth video segment of the third driving video, where j is an integer not less than 1; and a third sequence generation submodule, which is used to combine the feature marker sequence of the jth video segment and the feature marker sequence of its preceding video segment, as well as the driving voice feature marker sequence corresponding to the jth video segment, to generate a third synthetic sequence of the jth video segment.

[0180] In some embodiments, the third video feature submodule is used to: when j=1, determine the feature marker sequence of the first video segment based on the reference image; set the feature marker sequence of the preceding video segment of the first video segment to a default value; when j≥2, determine the feature marker sequence of the jth video segment based on the reference image and the driving speech feature sequence corresponding to the jth video segment; and use the determined feature marker sequence of the j-1th video segment as the feature marker sequence of the preceding video segment of the jth video segment.

[0181] In some embodiments, the third video feature submodule is further used to: determine a third video feature marker sequence of the jth video segment based on a reference image and a driving speech feature sequence corresponding to the jth video segment; generate a third noise feature marker sequence for the jth video segment; generate a second head feature marker sequence for the jth video segment; and generate a feature marker sequence of the jth video segment based on the third video feature marker sequence, the third noise feature marker sequence and the second head feature marker sequence.

[0182] In some embodiments, the second sequence module 605 includes: a fourth video feature submodule, used to determine a fourth video feature marker sequence of the jth video clip based on the third driving video, the fourth video feature sequence being a video feature sequence after masking the face and hand areas in the jth video clip; a grid feature marker submodule, used to determine a three-dimensional grid feature marker for the face and hand areas of the jth video clip based on the third driving video; a noise feature marker submodule, used to determine a fourth noise feature marker sequence for the jth video clip; and a fourth sequence generation submodule, used to determine a fourth synthetic sequence of the jth video clip based on the three-dimensional grid feature marker, the fourth noise feature marker sequence and the fourth video feature marker sequence of the jth video clip.

[0183] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0184] The digital human driving device in the disclosed embodiment can improve the quality and efficiency of generating digital human videos.

[0185] The present disclosure provides a scenario diagram of a method for generating a digital human driving model. Figure 7 shown.

[0186] As mentioned above, the digital human-driven model generation method provided by the embodiment of the present disclosure is applied to electronic devices. The electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0187] Specifically, the electronic device can perform the following operations:

[0188] Determine a first synthesis sequence, a second synthesis sequence, a speech feature marker sequence, a reference feature marker, and a facial feature marker based on the reference video;

[0189] Training a first sub-model based on the first synthesis sequence, the speech feature marker sequence, the reference feature markers, and the facial feature markers so that the first sub-model outputs a first driving video;

[0190] training a second sub-model based on the second synthesized sequence, the reference feature markers, and the facial feature markers so that the second sub-model outputs a second driving video;

[0191] A digital human driving model is generated based on the trained first sub-model and the second sub-model.

[0192] It should be understood that Figure 7 The scene diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 7 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0193] The present disclosure provides a scene schematic diagram of a digital human driving method, such as Figure 8 shown.

[0194] As mentioned above, the digital human driving method provided by the embodiment of the present disclosure is applied to electronic devices. The electronic devices are intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0195] Specifically, the electronic device can perform the following operations:

[0196] Determine digital human reference feature markers and digital human facial feature markers based on the reference image;

[0197] Extracting a driving speech feature tag sequence based on the driving speech;

[0198] determining a third synthetic sequence for generating each video segment based on the reference image and the driving speech;

[0199] Inputting the third synthesis sequence, the driving speech feature mark sequence, the digital human reference feature mark and the digital human facial feature mark into the first sub-model of the digital human driving model to obtain a third driving video;

[0200] determining a fourth synthesis sequence for generating each video segment based on the third driving video;

[0201] Inputting the fourth synthetic sequence, the digital human reference feature mark and the digital human facial feature mark into the second sub-model of the digital human driving model to obtain a fourth driving video;

[0202] A digital human video is generated based on the fourth driving video corresponding to all the video clips.

[0203] It should be understood that Figure 8 The scene diagram shown is only illustrative and not restrictive. Those skilled in the art can Figure 8 Various obvious changes and / or substitutions can be made to the examples, and the resulting technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0204] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0205] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0206] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0207] like Fig. 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 to a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0208] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0209] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as a digital human driving model generation method and / or a digital human driving method. For example, in some embodiments, the digital human driving model generation method and / or the digital human driving method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the digital human driving model generation method and / or the digital human driving method described above may be executed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the digital human driving model generation method and / or the digital human driving method in any other appropriate manner (e.g., by means of firmware).

[0210] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0211] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0212] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0213] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0214] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), and the Internet.

[0215] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0216] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in this application can be achieved, and this document is not limited here.

[0217] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for generating a digital human driving model, comprising: Determine a first synthesis sequence, a second synthesis sequence, a speech feature marker sequence, a reference feature marker, and a facial feature marker based on the reference video; Training a first sub-model based on the first synthesis sequence, the speech feature marker sequence, the reference feature marker, and the facial feature marker, so that the first sub-model outputs a first driving video; training a second sub-model based on the second synthesized sequence, the reference feature markers, and the facial feature markers so that the second sub-model outputs a second driving video; A digital human driving model is generated based on the trained first sub-model and the second sub-model.

2. The method according to claim 1, wherein: The determining of the first synthesis sequence based on the reference video comprises: Determining a feature marker sequence of an i-th video segment and a feature marker sequence of a preceding video segment based on the reference video, where i is an integer not less than 1; the i-th video segment is used to predict the i-th video segment in the first driving video; Determine the speech feature marker sequence corresponding to the i-th video segment based on the reference video; The first synthetic sequence of the i-th video segment is generated by combining the feature marker sequence of the i-th video segment and the feature marker sequence of its preceding video segment, and the speech feature marker sequence corresponding to the i-th video segment.

3. The method according to claim 2, wherein: The determining the feature marker sequence of the i-th video segment based on the reference video comprises: Determine a first video feature marker sequence of the i-th video segment based on the reference video; wherein the first video feature marker sequence is generated based on a plurality of continuous images in the reference video; generating a first noise signature sequence and a first head signature sequence for the i-th video segment; A signature sequence of the i-th video segment is generated based on the first video signature sequence, the first noise signature sequence, and the first head signature sequence.

4. The method according to claim 2, wherein: The determining, based on the reference video, the speech feature marker sequence corresponding to the i-th video segment comprises: Determining a reference speech of the i-th video segment based on the reference video; Extracting and encoding the reference speech to generate speech feature markers at a single time point; The speech feature markers are arranged in chronological order to generate the speech feature marker sequence of the i-th video segment.

5. The method according to claim 1, wherein: Determining reference feature markers and facial feature markers based on the reference video includes: Determine any frame image from the reference video; Encoding any one of the frame images to generate the reference feature mark; The facial information of the target object is extracted from any frame image, and the facial information of the target object is encoded to generate the facial feature mark.

6. The method according to claim 1, wherein: Determining a second synthesized sequence based on a reference video includes: Determine a feature marker sequence of a k-th video segment and a feature marker sequence of a preceding video segment based on the reference video or the first driving video, where k is an integer not less than 1; the k-th video segment is used to predict a k-th video segment in the second driving video; The second synthesized sequence of the kth video segment is determined according to the feature marker sequence of the kth video segment and the feature marker sequence of its preceding video segment.

7. The method according to claim 6, wherein: The determining, based on the reference video or the first driving video, a feature marker sequence of the kth video segment and a feature marker sequence of a preceding video segment thereof comprises: Determine a second video feature tag sequence of the kth video clip based on the reference video or the first driving video; wherein the second video feature sequence is a video feature sequence after masking of face and hand regions of a plurality of consecutive images in the reference video or the first driving video; Generating a three-dimensional mesh feature tag for the face and hand regions for the kth video segment based on the reference video or the first driving video; generating a second noise signature sequence for the kth video segment; A signature sequence of the kth video segment is generated based on the second video signature sequence, the second noise signature sequence, and the three-dimensional grid signature.

8. The method according to claim 1, wherein: The first sub-model includes N transformer layers, each of which includes a first self-attention layer and a speech cross-attention layer; the first self-attention layer includes a first feature adapter and a second feature adapter, the first feature adapter is used to adjust the features of the corresponding object in the first target image according to the reference feature marker; the second feature adapter is used to adjust the facial features in the first target image according to the facial feature marker; the speech cross-attention layer is used to adjust the facial expressions in the target image according to the speech feature marker sequence so that the facial expressions match the speech content; wherein N is an integer not less than 2.

9. The method according to claim 8, wherein: The training of the first sub-model based on the first synthesis sequence, the speech feature mark sequence, the reference feature mark and the facial feature mark comprises: The first synthetic sequence, the reference feature markers and the facial feature markers are used as inputs of the first self-attention layer, and the speech feature marker sequence and the output of the first self-attention layer are used as inputs of the speech cross-attention layer, and the first sub-model is trained so that the difference between the first driving video output by the first sub-model and the reference video is less than a first preset range.

10. The method according to claim 9, wherein: The first synthesis sequence, the reference feature markers and the facial feature markers are used as inputs of the first self-attention layer, and the speech feature marker sequence and the output of the first self-attention layer are used as inputs of the speech cross-attention layer to train the first sub-model, comprising: The first synthesized sequence is input into the first self-attention layer in the first Transformer layer; the output of the first self-attention layer and the reference feature label are input into the first feature adapter; the output of the first feature adapter and the facial feature label are input into the second feature adapter; the output of the second feature adapter and the speech feature label sequence are input into the speech cross-attention layer to obtain an updated first synthesized sequence output by the speech cross-attention layer, and the updated first synthesized sequence is used as the input of the next Transformer layer until the last Transformer layer outputs the first driving video.

11. The method according to claim 1, wherein: The second sub-model includes N Transformer layers, each Transformer layer includes a second self-attention layer; the second self-attention layer includes a third feature adapter and a fourth feature adapter, the third feature adapter is used to adjust the features of the corresponding object in the second target image according to the reference feature marker; The fourth feature adapter is used to adjust the facial features in the second target image according to the facial feature markers, where N is an integer not less than 2.

12. The method according to claim 11, wherein: The training of the second sub-model based on the second synthesized sequence, the reference feature marker and the facial feature marker comprises: The second synthesized sequence, the reference feature markers and the facial feature markers are used as inputs of the second self-attention layer, and the second sub-model is trained so that the difference between the second driving video output by the second sub-model and the reference video is less than a second preset range.

13. A digital human driving method, comprising: Determine digital human reference feature markers and digital human facial feature markers based on the reference image; Extracting a driving speech feature tag sequence based on the driving speech; Determine a third synthesis sequence for generating each video segment based on the reference image and the driving voice; Inputting the third synthesis sequence, the driving speech feature mark sequence, the digital human reference feature mark and the digital human facial feature mark into the first sub-model of the digital human driving model to obtain a third driving video; Determining a fourth synthesis sequence for generating each video segment based on the third driving video; Inputting the fourth synthesized sequence, the digital human reference feature mark and the digital human facial feature mark into the second sub-model of the digital human driving model to obtain a fourth driving video; A digital human video is generated based on the fourth driving video corresponding to all video clips.

14. The method according to claim 13, wherein: The step of determining a digital human reference feature mark and a digital human facial feature mark based on a reference image comprises: Performing feature encoding on the reference image to generate a reference feature mark of the digital human; The facial information of the digital person is extracted from the reference image, and the facial information is coded to generate the facial feature mark of the digital person.

15. The method according to claim 13, wherein: The step of extracting a driving speech feature tag sequence based on the driving speech comprises: Extracting and encoding the driving speech features to generate a driving speech feature marker at a single time point; The driving speech feature markers are arranged in time sequence to generate the driving speech feature marker sequence.

16. The method according to claim 13, wherein: The determining, based on the reference image and the driving voice, a third synthetic sequence for generating each video segment comprises: When predicting the jth video segment of the third driving video, determining a feature marker sequence of the jth video segment and a feature marker sequence of a preceding video segment based on the reference image and the driving voice, where j is an integer not less than 1; The third synthetic sequence of the jth video segment is generated by combining the feature marker sequence of the jth video segment and the feature marker sequence of its preceding video segment, as well as the driving speech feature marker sequence corresponding to the jth video segment.

17. The method according to claim 16, wherein: The step of determining the feature marker sequence of the jth video segment and the feature marker sequence of its preceding video segment based on the reference image and the driving voice comprises: When j=1, determining a feature marker sequence of the first video segment based on the reference image; setting a feature marker sequence of a preceding video segment of the first video segment to a default value; When j≥2, determine the feature marker sequence of the jth video segment based on the reference image and the driving speech feature sequence corresponding to the jth video segment; and use the determined feature marker sequence of the j-1th video segment as the feature marker sequence of the preceding video segment of the jth video segment.

18. The method according to claim 17, wherein: The determining the feature marker sequence of the jth video segment based on the reference image and the driving speech feature sequence corresponding to the jth video segment comprises: Determine a third video feature marker sequence of the jth video segment based on the reference image and the driving speech feature sequence corresponding to the jth video segment; Generating a third noise signature sequence and a second head signature sequence for the j-th video segment; A signature sequence of the j-th video segment is generated based on the third video signature sequence, the third noise signature sequence and the second head signature sequence.

19. The method according to claim 13, wherein: The determining, based on the third driving video, a fourth synthesis sequence for generating each video segment comprises: Determine a fourth video feature tag sequence of the jth video segment based on the third driving video, wherein the fourth video feature sequence is a video feature sequence after masking of the face and hand regions in the jth video segment; Determine a three-dimensional grid feature marker for the face and hand regions of the j-th video segment based on the third driving video; Determining a fourth noise feature marker sequence for the j-th video segment; The fourth synthesis sequence of the jth video segment is determined according to the three-dimensional grid feature mark of the jth video segment, the fourth noise feature mark sequence and the fourth video feature mark sequence.

20. A digital human driving model generation device, comprising: A feature extraction module, used to determine a first synthesis sequence, a second synthesis sequence, a speech feature marker sequence, a reference feature marker, and a facial feature marker based on a reference video; A first training module, configured to train a first sub-model based on the first synthesis sequence, the speech feature marker sequence, the reference feature marker and the facial feature marker, so that the first sub-model outputs a first driving video; A second training module, configured to train a second sub-model based on the second synthesized sequence, the reference feature markers, and the facial feature markers, so that the second sub-model outputs a second driving video; A model generation module is used to generate a digital human driving model based on the trained first sub-model and the second sub-model.

21. A digital human driving device, comprising: A feature marking module, used to determine a digital human reference feature marking and a digital human facial feature marking based on a reference image; A speech extraction module, used for extracting a driving speech feature tag sequence based on the driving speech; A first sequence module, configured to determine a third synthesis sequence for generating each video segment based on the reference image and the driving speech; A first generating module is used to input the third synthesis sequence, the driving speech feature mark sequence, the digital human reference feature mark and the digital human facial feature mark into the first sub-model of the digital human driving model to obtain a third driving video; A second sequence module, configured to determine a fourth synthesis sequence for generating each video segment based on the third driving video; A second generating module, configured to input the fourth synthesized sequence, the digital human reference feature mark and the digital human facial feature mark into a second sub-model of the digital human driving model to obtain a fourth driving video; The video generation module is used to generate a digital human video based on the fourth driving video corresponding to all video clips.

22. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to perform the method of any one of claims 1-19.

23. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are for causing a computer to perform a method according to any one of claims 1-19.

24. A computer program product comprising a computer program stored on a storage medium, the computer program implementing the method according to any one of claims 1 to 19 when executed by a processor.

Citation Information

Cited By

  • Training method of interactive digital human generation model, interactive digital human generation method and device, storage medium and program product

    CN120146139A