Digital human video generation method, electronic equipment, storage medium and program product
By performing feature decoupling processing of face images and combining the generation of adversarial networks, the feature dimension of the diffusion model is reduced, and the problem of long inference time of video generation in the existing technology is solved, and efficient digital video generation is achieved to achieve real-time output effect.
Patent Information
- Application Number
- CN202510560203.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-25
AI Technical Summary
The existing digital human video generation technology based on diffusion model has the problem of long video generation inference time and low generation efficiency.
By decoupling the face image, face features and background features are obtained, and input them from the audio feature sequence to the diffusion model to generate face feature sequences. Finally, digital human videos are generated based on background features and face feature sequences, reducing the feature dimensions of the diffusion model, combining the generation speed of the generative adversarial network and the diversity of the diffusion model.
It improves the efficiency of digital human video generation, achieves real-time generation effect, and can achieve real-time output of 30fps on a 3090Ti GPU.
Smart Images

Figure CN120375448A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital human generation technology, and in particular, to a method for generating digital human videos, an electronic device, a storage medium, and a program product. Background Art
[0002] In related technologies, a diffusion model is used to generate digital human videos. The diffusion model edits and generates an output video on a single face image or a real speaker video based on the text or speech content input by the user. The above diffusion model is a generative model that maps the noise distribution to the target distribution by learning the step-by-step denoising process of the image, and is used to generate high-quality images. Audio-driven lip movements and head movement conditions are introduced during the denoising process, and it has the ability to generate realistic details and can restore high-resolution dynamic facial expressions.
[0003] However, the biggest defect of the above digital human video generation technology based on the diffusion model is that the video generation inference time is relatively long. The reason is that the above digital human video generation technology based on the diffusion model directly generates images using the diffusion model, which takes a long time. Specifically, the denoising process of the diffusion model requires multiple steps of inference, resulting in low generation efficiency and making it difficult to achieve inference. Summary of the Invention
[0004] Embodiments of this application provide a method for generating digital human videos, an electronic device, a storage medium, and a program product, which are used to solve at least one of the above technical problems.
[0005] In a first aspect, embodiments of this application provide a method for generating digital human videos, including: Receiving a face image and an audio file input by a user; Processing the face image to obtain face features and background features; Extracting features from the audio file to obtain an audio feature sequence; Using a diffusion model to generate a face feature sequence according to the face features and the audio feature sequence; Generating a digital human video at least according to the background features and the face feature sequence.
[0006] In some embodiments, generating a digital human video at least according to the background features and the face feature sequence includes: Generating a face image sequence according to the background features and the face feature sequence; Generating a digital human video based on the face image sequence and the audio file.
[0007] In some embodiments, extracting features from the audio file to obtain an audio feature sequence includes: Use Whisper to extract features from the audio file to obtain an initial audio feature sequence; Use a preset neural network to filter the initial audio feature sequence to obtain an audio feature sequence, which includes a feature sequence corresponding to lip movements.
[0008] In some embodiments, the face features include expression features and head pose features; Use a diffusion model to generate a face feature sequence based on the face features and the audio feature sequence, including: Use a diffusion model to process the expression features and head pose features according to the feature sequence corresponding to the lip movements to obtain a face feature sequence including lip information containing corresponding audio features and natural and continuous head movement information.
[0009] In some embodiments, the face features are 69-dimensional features and the audio features are 512-dimensional features.
[0010] In some embodiments, generating a digital human video based on the face image sequence and the audio file includes: According to the audio feature sequence, perform temporal synchronization processing on the face image sequence and the audio file to generate a digital human video file; or, According to the audio feature sequence, perform temporal synchronization processing on the face image sequence and the audio file and stream out a digital human video stream.
[0011] In some embodiments, generating a face image sequence based on the background features and the face feature sequence includes: Use the generator of a pre-trained generative adversarial network to perform streaming rendering based on the background features and the face feature sequence to generate a face image sequence.
[0012] In a second aspect, an embodiment of the present application provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the digital human video generation method according to any one of the embodiments of the present application.
[0013] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the digital human video generation method according to any one of the embodiments of the present application are implemented.
[0014] Fourthly, an embodiment of the present application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the digital human video generation method described in any one of the embodiments of the present application.
[0015] In the embodiment of the present application, face features and background features are obtained by performing feature decoupling processing on a face image, and the face features and the audio feature sequence are input into a diffusion model to obtain a face feature sequence. Finally, a digital human video is generated based on the face feature sequence and the background features. This enables the diffusion model to only process low-dimensional face features and audio features and output a face feature sequence, without directly processing high-dimensional face images, not only making use of the generation diversity characteristic of the diffusion model, but also improving the efficiency of generating digital human videos. Description of the Drawings
[0016] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flowchart of an embodiment of the digital human video generation method of the present application; Figure 2 It is a flowchart of another embodiment of the digital human video generation method of the present application; Figure 3 It is a flowchart of another embodiment of the digital human video generation method of the present application; Figure 4 It is a flowchart of another embodiment of the digital human video generation method of the present application; Figure 5 It is a schematic structural diagram of an embodiment of the electronic device of the present application. Detailed Embodiments
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0019] It should also be noted that in this text, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or elements inherent to such a process, method, article, or device. Without further limitations, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article, or device that includes the said elements.
[0020] In the process of implementing this application, the inventor tried to reduce the feature dimension of the diffusion model to solve the problem of relatively long inference time for video generation. However, due to these defects being technical limitations of the diffusion model itself, because the diffusion model itself needs to gradually denoise to complete feature mapping, even when using Stable Diffusion and reducing the feature dimension based on VAE, the denoising work is usually carried out in the dimension of 64*64*4, with a relatively high feature dimension, and ultimately still directly generates images, so the generation efficiency cannot be substantially improved. The inventor further found that the generative adversarial network can efficiently complete the generation of face images, but the diversity of face images generated only based on the generative adversarial network is relatively low, and it is also difficult to control the generation of detailed expressions.
[0021] Furthermore, the inventor proposed a digital human video generation method that combines the generative adversarial network and the diffusion model. That is, it ensures the generation speed of the generative adversarial network while retaining the generation diversity of the diffusion model.
[0022] As Figure 1 shown, an embodiment of this application provides a digital human video generation method, including: S10. Receive a face image and an audio file input by a user.
[0023] Exemplarily, in the actual application process, the user can input any face image and audio file. The face image (for example, the user's own face) is used as the reference face, and the audio file (for example, the user's own recorded audio) is used as the reference audio, so as to generate the user's own digital human according to the user's own face and the user's own recorded audio.
[0024] S20. Process the face image to obtain face features and background features.
[0025] Exemplarily, a pre-trained feature decoupling network can be used to decouple the face image to obtain face features and background features. The specific training method of the feature decoupling network is not limited in this application.
[0026] S30. Extract feature sequences from the audio file to obtain audio feature sequences.
[0027] Exemplarily, feature extraction is performed frame by frame on the audio file to obtain an audio feature sequence, which is used to drive the face image to generate a face feature sequence that conforms to the audio feature sequence.
[0028] S40. Use a diffusion model to generate a face feature sequence based on the face feature and the audio feature sequence.
[0029] Exemplarily, a pre-trained diffusion model is used to generate a face feature sequence based on the face feature and the audio feature sequence. Among them, the goal of training the diffusion model is to use the single face feature and the audio feature sequence decoupled by the decoupling network as input conditions to generate natural and continuous head movement features and expression features highly corresponding to the audio. At this time, the input audio feature is the audio feature file extracted from the corresponding audio file, and the target generated feature is the face feature extracted and decoupled from the corresponding video frame.
[0030] S50. Generate a digital human video based on at least the background feature and the face feature sequence.
[0031] In the embodiment of the present application, face features and background features are obtained by performing feature decoupling processing on a face image, and the face features and the audio feature sequence are input into a diffusion model, thereby obtaining a face feature sequence. Finally, a digital human video is generated based on the face feature sequence and the background feature. The diffusion model only needs to process the low-dimensional face features and audio features and output the face feature sequence, without directly processing the high-dimensional face image, which not only utilizes the generation diversity characteristic of the diffusion model, but also improves the efficiency of generating a digital human video.
[0032] As Figure 2 shown is a flowchart of another embodiment of the digital human video generation method of the present application. In this embodiment, generating a digital human video based on at least the background feature and the face feature sequence includes: S51. Generate a face image sequence based on the background feature and the face feature sequence.
[0033] Exemplarily, a generator of a pre-trained generative adversarial network is used to perform streaming rendering based on the background feature and the face feature sequence to generate a face image sequence. The generative adversarial network includes a discriminator and a generator, but only the discriminator is used during training, and only the generator is used during the inference phase, and the generator is used to perform image rendering generation. It should be noted that the present application does not limit the specific method for training the generative adversarial network.
[0034] S52. Generate a digital human video based on the face image sequence and the audio file.
[0035] Exemplarily, generating a digital human video based on the face image sequence and the audio file includes: performing temporal synchronization processing on the face image sequence and the audio file according to the audio feature sequence to generate a digital human video file; or, performing temporal synchronization processing on the face image sequence and the audio file according to the audio feature sequence and streaming outputting a digital human video stream. Among them, the face image sequence is generated based on a piece of audio, and the mouth shape of each face in this face image sequence corresponds to the audio. The image sequence and the audio file are combined to output a video, and the mouth shape in the video matches the audio.
[0036] The solution of this embodiment combines the two major advantages of the generative adversarial network and the diffusion model, that is, ensuring the generation speed of the generative adversarial network while retaining the generation diversity of the diffusion model.
[0037] Such as Figure 3 shown is a flowchart of another embodiment of the digital human video generation method of the present application. In this embodiment, extracting features from the audio file to obtain an audio feature sequence includes: S31. Using whisper to extract features from the audio file to obtain an initial audio feature sequence. When an input audio file is received, audio feature extraction is immediately performed (the extracted audio features are used to drive mouth movements and represent the features of the audio content), and an audio feature sequence is obtained. Here, we first use whisper-base as the basic audio feature (the specific acquisition method can refer to relevant prior arts, and the present application does not make limitations).
[0038] S32. Using a preset neural network to filter the initial audio feature sequence to obtain an audio feature sequence, and the audio feature sequence includes a feature sequence corresponding to mouth movement actions.
[0039] Exemplarily, after obtaining the initial audio feature sequence, a series of network processes are performed to further extract useful information in the audio (exemplarily, after using whisper to extract features, several neural networks are constructed to further extract useful features) to prevent feature leakage.
[0040] Among them, the useful features refer to the features specifically corresponding to mouth movement actions. Because the features extracted based on whisper are multi-scale features, and also include other features such as voiceprints, redundant features need to be filtered out. If the features directly extracted by whisper are input into the diffusion model, then there are many redundant other features, which are likely to cause feature leakage. Therefore, a preset neural network is used after whisper to further extract useful features to prevent feature leakage.
[0041] In some embodiments, the facial features include expression features and head pose features; using a diffusion model to generate a facial feature sequence based on the facial features and the audio feature sequence, including: using a diffusion model to process the expression features and head pose features according to the feature sequence of the corresponding mouth shape movement to obtain a facial feature sequence including lip information containing the corresponding audio features and natural and continuous head movement information. In some embodiments, the facial features (including expression features and head pose features) are 69-dimensional features, and the audio features are 512-dimensional features.
[0042] In this embodiment, the feature dimensions input into the diffusion model include: a total of 133 dimensions of all facial features, 69 dimensions to be generated (expression features, head pose features), and the remaining 64 dimensions remain fixed (person features). The input audio feature dimension is 512. That is, the feature dimensions input into the diffusion model are much lower than the commonly used dimensions in the industry described above (64*64*4). Therefore, although this embodiment applies a diffusion model, it will not cause excessive time loss, and at the same time retains a major advantage of the diversity of the diffusion model.
[0043] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, each embodiment is described with emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0044] In some embodiments, the actual application of this application is to generate a target person's speaking face video from a single face image driven by audio, including natural and continuous head movement postures and lip movements highly corresponding to the audio content. The application content includes the training and generation methods of this solution, which can achieve real-time interactive generation in actual applications. It can output a real speaking face video for any audio content, including situations such as speaking and singing.
[0045] Specifically, there are two stages in the training process, namely, the training of the adversarial generation network and the feature decoupling network (i.e., the encoder) of the reference face image, and the training of the face feature generation network (i.e., the diffusion model) using audio features. Similarly, in the inference and generation stage, the feature decoupling network (i.e., the encoder) decouples the reference face image into face features and background features; the diffusion model generates a sequence of face features based on the input audio and face features; and the generator of the adversarial generation network renders and generates a sequence of face images according to the background features and the sequence of face features, and finally combines the continuous sequence of face images with the audio to output the talking face video of the target person.
[0046] Before training, a large amount of face audio-visual data is pre-collected, and after processes such as cropping, cleaning, and screening, high-quality talking face video data is retained for subsequent training.
[0047] The training objective of the first-stage network is that the feature decoupling network can accurately decouple the features of the input face image, that is, separate the face features (including expression features, head pose features, identity features, etc.) and background features (meanwhile, the diffusion model can arbitrarily modify the target features, change the expression actions or head poses of the target person, and obtain a sequence of face features), and the generative adversarial network combines the sequence of face features and background features to generate a series of real face images. In this stage, we use the feature decoupling network and the generative adversarial network for training, which not only ensures high fidelity but also provides the possibility of real-time inference in actual applications.
[0048] In the second stage, the goal is to generate natural and continuous head movement features and expression features highly corresponding to the audio based on the diffusion model, using the single face features (e.g., expression features, head pose features) decoupled by the first-stage network and the audio feature sequence as input conditions. At this time, the input audio features are the audio features extracted from the corresponding audio file, and the target generated features are the face features extracted and decoupled from the corresponding video frames.
[0049] At this time, the feature dimensions input into the diffusion model include: a total of 133 dimensions of all face features, 69 dimensions to be generated (head pose features and expression features), and the remaining 64 dimensions remain fixed (person features). The dimension of the input audio features is 512. That is, the feature dimensions input into the diffusion model are much lower than the commonly used dimensions (64*64*4) in the industry described above. Therefore, although this solution applies the diffusion model, it will not cause excessive time loss, and at the same time retains a major advantage of the diversity of the diffusion model. At the same time, the generative adversarial network can be used to efficiently complete the generation of face images. Therefore, this application not only ensures the generation speed of the generative adversarial network but also retains the generation diversity of the diffusion model.
[0050] Such as Figure 4The following is a flowchart of another embodiment of the digital human generation method of the present application. In the actual application process of this embodiment, the user can input any face image and audio file. The inference process is as follows Figure 4 shown. First, when the user inputs a face image, the background will immediately respond and perform feature decoupling and encoding work; after the audio file is input, a continuous face feature sequence is generated through the above-mentioned second-stage diffusion model network; then, the generator of the generative adversarial network is used to stream-render the image and output a sequence of face images, so that real-time output at 30fps can be achieved. It can also be connected to an interactive large model to interact with customers in real time through tts voice and a pre-set character image. The process of each step will be introduced in detail below.
[0051] 1. Face feature decoupling: For the input reference face image, feature decoupling processing will be immediately performed to output face features and background features in different dimensions. The face features (including different features such as expression features and action features, with different feature dimensions) will be used as reference features and input into the neural network (i.e., the diffusion model) during the subsequent generation of the face feature sequence, while the background features will be used as inputs in the final face image rendering network (i.e., the generator), so as to ensure that the background remains unchanged and only the face actions and expressions are changed.
[0052] 2. Audio feature extraction: When the model receives the input audio file, audio feature extraction is immediately performed (the extracted audio features are features representing the audio content and are used to drive the mouth shape movements), and an audio feature sequence is obtained. Here, we first use whisper-base as the basic audio feature, and then further extract useful information in the audio through a series of network processes (exemplarily, after using whisper to extract features, several neural networks are constructed for a series of subsequent network processes to further extract useful features) to prevent feature leakage.
[0053] Among them, the useful features refer to the features specifically corresponding to the mouth shape movements. Because the features extracted based on whisper are multi-scale features, including other features such as voiceprints, redundant features need to be filtered out. If the features directly extracted by whisper are input into the diffusion model, then there are many redundant other features, which are likely to cause feature leakage. Therefore, several preset neural networks are used after whisper to further extract useful features to prevent feature leakage.
[0054] 3. Generation of face feature sequence: After obtaining the target face features and audio features, they are jointly input into the diffusion model. At this time, the audio feature is a sequence feature, and a face feature sequence can be generated based on the single-face expression feature. This face feature sequence contains lip information corresponding to the audio feature and natural and continuous head movement information.
[0055] 4. Facial Image Rendering: After generating the facial feature sequence, the facial features of each frame can be combined with the background features obtained above and input into the renderer to generate a rendered target facial image (finally obtaining a sequence of facial images corresponding to the facial feature sequence).
[0056] 5. Video Synthesis: According to the actual application scenario, after all facial images (sequence of facial images) are rendered, a video file (e.g., MP4 file) can be output by combining with an audio file; or a video stream can be output in a streaming manner by combining with an audio file for real-time interaction. Exemplarily, the sequence of facial images is generated based on a piece of audio, and the mouth shape of each face in this sequence of facial images corresponds to the audio. The image sequence and the audio file are combined to output a video, and the mouth shape in the video matches the audio.
[0057] In the embodiments of the present application, generating facial images through the renderer can achieve real-time generation while maintaining high fidelity. Compared with the common diffusion model-based solutions in the industry (which cannot achieve real-time generation), this solution can achieve real-time generation at 30fps on a 3090Ti GPU.
[0058] In the embodiments of the present application, facial features are generated based on the diffusion model. At this time, the feature dimension is relatively low, and problems such as slow generation time will not occur. At the same time, the advantages of the diffusion model such as generation diversity and high consistency are retained.
[0059] The key innovation point of the embodiments of the present application is to propose a new digital human generation technology. Only any one facial image and any audio file are required as inputs, and a natural and realistic target person's speaking facial video can be generated. The technical solution of the present application can not only be applied in the scenarios described above, but also in application scenarios such as live broadcasts and real-time MOOC systems. In addition, the aforementioned first stage can be applied to application scenarios related to video reproduction.
[0060] In some embodiments, the embodiments of the present application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the digital human video generation methods described above in the present application.
[0061] In some embodiments, the embodiments of the present application further provide a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute any one of the digital human video generation methods described above.
[0062] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the digital human video generation method.
[0063] In some embodiments, the present application further provides a digital human video generation device, including: A data receiving module, configured to receive a face image and an audio file input by a user; A feature decoupling network, configured to process the face image to obtain face features and background features; An audio feature extraction network, configured to extract features from the audio file to obtain an audio feature sequence; A diffusion model, configured to generate a face feature sequence according to the face features and the audio feature sequence; A video generation module, configured to generate a digital human video at least according to the background features and the face feature sequence.
[0064] In some embodiments, generating a digital human video at least according to the background features and the face feature sequence includes: Generating a face image sequence according to the background features and the face feature sequence; Generating a digital human video based on the face image sequence and the audio file.
[0065] In some embodiments, extracting features from the audio file to obtain an audio feature sequence includes: Using whisper to extract features from the audio file to obtain an initial audio feature sequence; Using a preset neural network to filter the initial audio feature sequence to obtain an audio feature sequence, where the audio feature sequence includes a feature sequence corresponding to lip movements.
[0066] In some embodiments, the face features include expression features and head pose features; Using a diffusion model to generate a face feature sequence according to the face features and the audio feature sequence includes: Using a diffusion model to process the expression features and head pose features according to the feature sequence corresponding to the lip movements to obtain a face feature sequence including lip information corresponding to the audio features and natural and continuous head movement information.
[0067] In some embodiments, the face features are 69-dimensional features, and the audio features are 512-dimensional features.
[0068] In some embodiments, generating a digital human video based on the sequence of face images and the audio file includes: Performing temporal synchronization processing on the sequence of face images and the audio file according to the sequence of audio features to generate a digital human video file; or, Performing temporal synchronization processing on the sequence of face images and the audio file according to the sequence of audio features and streaming out a digital human video stream.
[0069] In some embodiments, generating a sequence of face images according to the background features and the sequence of face features includes: Using the generator of a pre-trained generative adversarial network to perform streaming rendering based on the background features and the sequence of face features to generate a sequence of face images.
[0070] The digital human video generation device in the embodiments of the present application can be used to execute the digital human video generation method in the embodiments of the present application, and correspondingly achieve the technical effects achieved by the digital human video generation method in the embodiments of the present application, which will not be elaborated here. In the embodiments of the present application, relevant functional modules can be implemented by a hardware processor.
[0071] Figure 5 FIG. is a schematic hardware structure diagram of an electronic device for executing the digital human video generation method provided in another embodiment of the present application. As Figure 5 shown, the device includes: One or more processors 510 and a memory 520. Figure 5 Taking one processor 510 as an example.
[0072] The device for executing the digital human video generation method may further include: an input device 530 and an output device 540.
[0073] The processor 510, the memory 520, the input device 530, and the output device 540 may be connected through a bus or other means. Figure 5 Taking the connection through a bus as an example.
[0074] The memory 520, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the digital human video generation method in the embodiments of the present application. The processor 510 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 520, that is, implementing the digital human video generation method in the above method embodiments.
[0075] The memory 520 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the digital human video generation device, etc. In addition, the memory 520 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 520 may optionally include a memory remotely disposed relative to the processor 510, and these remote memories may be connected to the digital human video generation device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0076] The input device 530 may receive input digital or character information and generate signals related to user settings and function controls of the digital human video generation device. The output device 540 may include a display device such as a display screen.
[0077] The one or more modules are stored in the memory 520 and, when executed by the one or more processors 510, execute the digital human video generation method in any of the above method embodiments.
[0078] The above product may execute the method provided in the embodiments of the present application, and has function modules and beneficial effects corresponding to the execution of the method. Technical details not described in detail in this embodiment may be referred to the method provided in the embodiments of the present application.
[0079] The electronic device in the embodiments of the present application exists in various forms, including but not limited to: (1) Mobile communication devices: Such devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc.
[0080] (2) Ultra-mobile personal computer devices: Such devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as iPad.
[0081] (3) Portable entertainment devices: Such devices can display and play multimedia content. Such devices include: audio and video players (such as iPod), handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.
[0082] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but due to the need to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, manageability, etc.
[0083] (5) Other electronic devices with data interaction functions.
[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0085] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the parts that contribute to the related technologies, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for generating a digital human video, comprising: Receiving a face image and an audio file input by a user; Processing the face image to obtain face features and background features; Extracting features from the audio file to obtain an audio feature sequence; Using a diffusion model to generate a face feature sequence according to the face features and the audio feature sequence; Generating a digital human video based on at least the background features and the face feature sequence.
2. The method according to claim 1, wherein Generating a digital human video based on at least the background features and the face feature sequence, including: Generating a face image sequence according to the background features and the face feature sequence; Generating a digital human video based on the face image sequence and the audio file.
3. The method according to claim 1, characterized in that, The extracting features from the audio file to obtain an audio feature sequence includes: Using whisper to extract features from the audio file to obtain an initial audio feature sequence; Using a preset neural network to filter the initial audio feature sequence to obtain an audio feature sequence, and the audio feature sequence includes a feature sequence corresponding to lip movements.
4. The method according to claim 3, wherein The face features include expression features and head pose features; Using a diffusion model to generate a face feature sequence according to the face features and the audio feature sequence includes: Using a diffusion model to process the expression features and head pose features according to the feature sequence corresponding to the lip movements to obtain a face feature sequence including lip information containing corresponding audio features and natural and continuous head movement information.
5. The method according to claim 4, characterized in that, The face features are 69-dimensional features, and the audio features are 512-dimensional features.
6. The method according to claim 2, wherein The generating a digital human video based on the face image sequence and the audio file includes: Performing temporal synchronization processing on the face image sequence and the audio file according to the audio feature sequence to generate a digital human video file; or, Performing temporal synchronization processing on the face image sequence and the audio file according to the audio feature sequence and streaming out a digital human video stream.
7. The method according to claim 2 or 6, characterized in that, Generating a face image sequence according to the background features and the face feature sequence includes: Using the generator of a pre-trained generative adversarial network to perform streaming rendering based on the background features and the face feature sequence to generate a face image sequence.
8. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.