Stylized Digital Human Video Generation Method, Electronic Device, and Storage Medium

By receiving user photos and dubbing files, and using a pre-trained lip-driven model to generate stylized digital human videos, the problem of insufficient integration of stylized elements in the existing technology is solved, high flexibility and real-time performance is achieved, and user operation threshold is lowered.

CN119211659BActive Publication Date: 2025-06-20HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411699073.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-20
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

The prior art is difficult to effectively incorporate stylized elements into digital verb-driven videos, resulting in insufficient flexibility and real-timeness and high user operation thresholds.

Method used

A stylized digital video generation method is proposed. By receiving user photos, target stylized types and dubbing files, using a pre-trained lip-driven model, the target stylized images and audio features are combined to generate synchronized stylized digital videos.

Benefits of technology

It realizes the generation of targeted stylized digital human images based on user photos, improves flexibility and user satisfaction, improves the real-time video generation, and simplifies user operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119211659B_ABST
    Figure CN119211659B_ABST
Patent Text Reader

Abstract

The present application provides a method for generating a stylized digital human video, an electronic device, and a storage medium. Belonging to the technical field of image processing, the method includes: receiving a stylized digital human video generation instruction, where the stylized digital human video generation instruction includes a user photo, a target stylized type, and a voice-over file; converting the user photo into a target stylized image according to the target stylized type; inputting the target stylized image and the voice-over file into a pre-trained lip-sync driving model, where the pre-trained lip-sync driving model extracts the identity features of the target stylized image and the audio features of the voice-over file, and generates a stylized digital human video according to the identity features and the audio features; obtaining the stylized digital human video output by the pre-trained lip-sync driving model, and the lip-sync of the stylized digital human video is synchronized with the voice-over file. The present application can also provide a more personalized, real-time, and high-quality stylized digital human video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and in particular, to a method for generating a stylized digital human video, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of artificial intelligence technologies, image processing and video generation technologies have made remarkable progress. This progress has given rise to the digital human lip-sync driving video technology, which converts an audio signal into a corresponding lip animation and applies it to a digital human image to generate highly realistic and interactive video content. However, with the diversification of user needs and the complexity of application scenarios, the traditional digital human lip-sync driving technology that pursues accurate restoration rather than stylized performance is difficult to apply to specific scenarios that require giving digital humans a unique style.

[0003] To integrate stylized elements into digital human lip-sync driving videos, the existing technologies are mainly divided into two categories of methods. The first category of methods is to pre-produce stylized digital human images for users to choose from, and then use non-real-time lip driving technology to convert the user's voice or text input into the lip animation of the digital human. Since the digital human images in this way are pre-produced, they have poor flexibility, cannot fully meet the personalized needs of each user, and the real-time performance of video generation is poor. The second category of methods requires users to dress themselves in the corresponding style, then record a video, and then convert the recorded video into a digital human video through image processing technology. Although this method can provide highly personalized results, its implementation process is relatively cumbersome, and users need to invest a lot of time and energy in dressing up and recording. This high-threshold operation process limits its universality and cannot meet the needs of all users.

[0004] Therefore, in order to integrate stylized elements into digital human lip-sync driving videos, it is urgent to further develop and improve relevant algorithms and technologies to overcome the limitations of the existing technologies. Summary of the Invention

[0005] The purpose of the present application is to provide a method for generating a stylized digital human video, an electronic device, and a storage medium to solve the above problems.

[0006] To achieve the above purpose, in a first aspect, the present application proposes a method for generating a stylized digital human video, and the method includes:

[0007] Receiving a stylized digital human video generation instruction, where the stylized digital human video generation instruction includes a user photo, a target stylized type, and a voiceover file;

[0008] Converting the user photo into a target stylized image according to the target stylized type;

[0009] Input the target stylized image and the voice-over file into a pre-trained lip-sync driving model. The pre-trained lip-sync driving model extracts the identity features of the target stylized image and the audio features of the voice-over file, and generates a stylized digital human video based on the identity features and the audio features.

[0010] Obtain the stylized digital human video output by the pre-trained lip-sync driving model, and the lip-sync of the stylized digital human video is synchronized with the voice-over file.

[0011] In some embodiments, the lip-sync driving model includes a generator network and a discriminator network. The discriminator network includes a lip-sync discriminator. Before inputting the target stylized image and the voice-over file into the pre-trained lip-sync driving model, it further includes:

[0012] Obtain a training set, which includes digital human lip-sync driving videos of various stylized types.

[0013] Through the digital human lip-sync driving videos of various stylized types, train the evaluation ability of the lip-sync discriminator for the synchronization of lip-sync and audio.

[0014] After training the lip-sync discriminator, extract reconstructed target frames from the training set. By masking the lip-sync part of the reconstructed target frames, train the generator network to generate reconstructed image frames based on the reconstructed target frames with the masked lip-sync part, and calculate the reconstruction loss based on the reconstructed target frames and the reconstructed image frames.

[0015] Based on the trained lip-sync discriminator, calculate the synchronization loss between the lip-sync of the reconstructed image frames and the lip-sync of the reconstructed target frames.

[0016] Perform feedback update on the generator network based on the reconstruction loss and the synchronization loss until the training is completed.

[0017] In some embodiments, the discriminator network further includes a visual quality discriminator. After training the generator network to generate reconstructed image frames based on the reconstructed target frames with the masked lip-sync part by masking the lip-sync part of the reconstructed target frames, it further includes:

[0018] Supervise the image quality of the reconstructed image frames through the visual quality discriminator, and calculate the adversarial loss. The adversarial loss is used to measure the adversarial performance of the reconstructed image frames against the discriminator network in adversarial training.

[0019] The performing feedback update on the generator network based on the reconstruction loss and the synchronization loss until the training is completed includes:

[0020] The generator network is updated by feedback by minimizing the weighted sum of the reconstruction loss, the synchronization loss, and the adversarial loss until the training is completed.

[0021] In some embodiments, training the lip-sync discriminator's ability to evaluate the synchronization of lip movements with audio by driving a video with digital mouth shapes of each stylized type includes:

[0022] Extract the audio information of the digital mouth shape-driven videos of each stylized type and cut the audio information into audio chunks;

[0023] Pair the audio chunks with the image frames in the digital mouth shape-driven videos of each stylized type to form matching pairs of audio chunks and corresponding image frames and non-matching pairs of audio chunks and non-corresponding image frames;

[0024] Train the lip-sync discriminator's ability to evaluate the synchronization of lip movements with audio through the matching pairs and the non-matching pairs.

[0025] In some embodiments, training the lip-sync discriminator's ability to evaluate the synchronization of lip movements with audio through the matching pairs includes:

[0026] Extract features from the audio chunks and image frames in the matching pairs and the non-matching pairs to generate audio feature vectors and image feature vectors;

[0027] Calculate the cosine similarity between the audio feature vectors and the image feature vectors of the matching pairs and the non-matching pairs respectively;

[0028] Based on the cosine similarities of the matching pairs and the non-matching pairs, update the lip-sync discriminator by feedback by minimizing the matching loss until the training is completed.

[0029] In some embodiments, extracting a reconstruction target frame from the training set, training the generator network to generate a reconstructed image frame based on the reconstruction target frame with the lip part masked, and calculating a reconstruction loss based on the reconstruction target frame and the reconstructed image frame includes:

[0030] Based on the training set, obtain a reference frame, a reconstruction target frame, and an audio segment corresponding to the time dimension of the reconstruction target frame, and use the reference frame, the reconstruction target frame, and the audio segment as inputs to the generator network, where the reference frame contains complete face features and the lip part of the reconstruction target frame is masked;

[0031] Obtain the reconstructed image frame generated by the generator network based on the reference frame, the reconstructed target frame, and the audio segment;

[0032] Calculate a reconstruction loss based on the reconstructed target frame and the reconstructed image frame where the lip part is not masked;

[0033] Update the generator network by minimizing the reconstruction loss until the training is completed.

[0034] In some embodiments, the generator network includes an identity encoder, a voice encoder, and a face decoder. Obtaining a reference frame, a reconstructed target frame, and an audio segment corresponding to the time dimension of the reconstructed target frame based on the training set, and using the reference frame, the reconstructed target frame, and the audio segment as inputs to the generator network includes:

[0035] Based on the training set, obtain a reference frame, a reconstructed target frame, and an audio segment corresponding to the time dimension of the reconstructed target frame;

[0036] Concatenate the reference frame and the reconstructed target frame along the channel dimension as the input to the identity encoder to obtain the identity features output by the identity encoder; and

[0037] Use the audio segment as the input to the voice encoder to obtain the voice features output by the voice encoder;

[0038] Concatenate the identity features and the voice features as the input to the face decoder to obtain the reconstructed image frame generated by the face decoder.

[0039] In some embodiments, before obtaining the training set, where the training set includes digital human lip-sync driving videos of various stylized types, it includes:

[0040] Obtain a set of lip-sync driving videos containing digital humans of various stylized types;

[0041] Screen out the lip-sync driving videos that meet the preset audio-visual synchronization standard from the set of lip-sync driving videos to construct an initial training set;

[0042] Based on the stylized labels of the lip-sync driving videos in the initial training set, classify the lip-sync driving videos with the same stylized label into the same data set;

[0043] Based on the classified data sets, form a training set.

[0044] In a second aspect, the present application also proposes an electronic device, including:

[0045] One or more processors;

[0046] A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the stylized digital human video generation method as described above.

[0047] In a third aspect, the present application also provides a storage medium storing executable instructions, which when executed by a processor cause the processor to execute the stylized digital human video generation method as described above.

[0048] Compared with the prior art, the beneficial effects of the present application include:

[0049] In a first aspect, the present application can generate a digital human image of a target stylized type according to the received user photo, avoiding the limitations of pre-made digital human images, and improving flexibility and user satisfaction. In a second aspect, the present application uses a pre-trained lip movement driving model to be able to generate a synchronized lip movement driving video in real time according to a dubbing file. Compared with traditional non-real-time lip movement driving technologies, the real-time performance of video generation is greatly improved. In a third aspect, to generate a stylized digital human video based on the technical solution of the present application, the user only needs to provide a photo, a dubbing file, and select a target stylized type. This greatly simplifies the user operation process and reduces the participation threshold. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope of the present application.

[0051] Figure 1 It is a schematic flowchart of an embodiment of the stylized digital human video generation method of the present application;

[0052] Figure 2 It is a partial schematic flowchart of an embodiment of the stylized digital human video generation method of the present application;

[0053] Figure 3 It is a partial schematic flowchart of an embodiment of the stylized digital human video generation method of the present application;

[0054] Figure 4 It is a schematic flowchart of an embodiment of the stylized digital human video generation method of the present application;

[0055] Figure 5 It is a partial schematic flowchart of an embodiment of the stylized digital human video generation method of the present application;

[0056] Figure 6Schematic flowchart of an embodiment of the method for generating a stylized digital human video in this application;

[0057] Figure 7 Schematic structural diagram of an electronic device involved in the method for generating a stylized digital human video in any embodiment of this application. Detailed implementation manners

[0058] In order to make the objectives, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0059] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0060] For example, terms such as "first" and "second" used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.

[0061] For another example, terms such as "include" and "comprise" used in this application indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0062] As described above, in order to integrate stylized elements into a digital human lip-sync driving video, the existing technologies are mainly divided into two categories of methods. The first category of methods is to pre-produce stylized digital human images for users to choose from, and then use non-real-time lip-sync driving technology to convert the user's voice or text input into the lip-sync animation of the digital human. Since the digital human images in this method are pre-produced, they have poor flexibility, cannot fully meet the personalized needs of each user, and the real-time performance of video generation is poor. The second category of methods requires the user to dress up in the corresponding style, then record a video, and then convert the recorded video into a digital human video through image processing technology. Although this method can provide highly personalized results, its implementation process is more cumbersome and requires the user to invest a lot of time and energy in dressing up and recording. This high-threshold operation process limits its universality and cannot meet the needs of all users. Therefore, in order to integrate stylized elements into a digital human lip-sync driving video, it is urgent to further develop and improve relevant algorithms and technologies to overcome the limitations of the existing technologies. For this reason, this application proposes a method, an electronic device and a storage medium for generating a stylized digital human video, which not only overcomes the limitations of the existing technologies, but also can provide more personalized, real-time and high-quality stylized digital human videos.

[0063] Figure 1 It is a schematic flowchart of an embodiment of the generation of a stylized digital human video in this application. As Figure 1 shown, the generation of the stylized digital human video in this embodiment includes the following steps:

[0064] Step S10, receiving a stylized digital human video generation instruction, where the stylized digital human video generation instruction includes a user photo, a target stylized type, and a voice-over file.

[0065] In this embodiment, the application scenarios of the stylized digital human video generation method cover a variety of intelligent terminals, including but not limited to smart phones, personal computers, robots, and other electronic devices with data processing, network communication, and program running functions. The specific online implementation forms can be websites, mobile application programs (APPs), WeChat mini-programs, H5 pages, and API interfaces, etc. Users access the Internet through these terminals and interact with the interactive interface that provides the stylized digital human video generation service. This may involve opening a website, starting an APP, entering a mini-program, or loading an H5 page in a browser.

[0066] In some embodiments, the electronic device receives a stylized digital human video generation instruction triggered by the user based on the interactive interface, and the stylized digital human video generation instruction includes a photo upload instruction, a voice-over file upload instruction, and a target stylized type selection instruction.

[0067] The electronic device obtains the user photo uploaded by the user terminal associated with the photo upload instruction. The user photo refers to the portrait picture uploaded by the user and is used to generate a stylized digital human image that matches the user's portrait features.

[0068] The electronic device obtains the voiceover file uploaded by the user terminal associated with the voiceover file upload instruction. The voiceover file includes text content or voice content. If the voiceover file obtained by the electronic device is text content, a voice style selection interface is published through the interaction interface, a voice style selection instruction is received based on the voice style selection interface, and a voice file is generated according to the voice parameters associated with the voice style selection instruction and the text content. Among them, the voice parameters include speech rate, intonation, gender, timbre, language, etc.

[0069] The electronic device determines that the stylized type selected by the user is the target stylized type based on the target stylized type selection instruction. In this embodiment, the stylized types include cartoon style, realistic style, retro style, etc. Among them, the realistic style includes but is not limited to professional style, casual style, dress style, denim style, ladylike style, sports style, etc.

[0070] Step S20: Convert the user photo into a target stylized image according to the target stylized type.

[0071] In this embodiment, the electronic device obtains the style image associated with the target stylized type according to the target stylized type. The content features of the user photo and the style features of the style image are extracted through a pre-trained convolutional neural network. Among them, the content features are used to retain the structure and content of the user photo, and the style features are used to capture the texture and color information of the style image. A target stylized image is generated according to the content features and the style image.

[0072] In some embodiments, an initial stylized image is randomly generated according to the content features and the style image. The content loss between the initial stylized image and the user photo is calculated. The content loss is used to measure the difference between the content features of the generated initial stylized image and the content features of the user photo. The style loss between the initial stylized image and the style image is calculated. The style loss is used to measure the difference between the style features of the generated initial stylized image and the style features of the style image.

[0073] Using optimization algorithms such as gradient descent, gradually adjust the pixel values of the generated stylized image by minimizing the combination of content loss and style loss. In each iteration, calculate the content features and style features of the generated stylized image, and compare them with the content features of the user photo and the style features of the style image, so as to update the generated stylized image.

[0074] When it is detected that the combination of content loss and style loss tends to be stable, confirm the end of the iteration, and obtain the finally generated target stylized image, which is a high-quality image that retains both the content features of the user photo and the style features of the style image.

[0075] It can be understood that before feature extraction of the user photo, the user photo can be preprocessed, and the preprocessing includes operations such as resizing and normalization, so as to facilitate subsequent stylization processing.

[0076] Step S30, input the target stylized image and the voiceover file into a pre-trained lip-sync driving model. The pre-trained lip-sync driving model extracts the identity features of the target stylized image and the audio features of the voiceover file, and generates a stylized digital human video according to the identity features and the audio features.

[0077] In this embodiment, the electronic device inputs the received speech content, or the speech file generated based on the received text content, together with the target stylized image into the pre-trained lip-sync driving model. After receiving the speech content or the speech file, the pre-trained lip-sync driving model extracts the audio features of the speech content or the speech file, and the audio features are features related to pronunciation, such as syllables, tones, etc. After receiving the target stylized image, the pre-trained lip-sync driving model extracts the identity features of the target stylized image, and the identity features refer to the facial features of the digital human in the target stylized image, and the facial features include the features of key parts such as facial contour, eyes, nose, mouth, etc. According to the audio features and the facial features, generate a lip animation. According to the lip animation and the target stylized image, generate a stylized digital human video.

[0078] Step S40, obtain the stylized digital human video output by the pre-trained lip-sync driving model, and the lip-sync of the stylized digital human video is synchronized with the voiceover file.

[0079] In this embodiment, obtain the stylized digital human video output by the pre-trained lip motion driving model, where the lip motion driving of the stylized digital human video is synchronized with the voice content or voice file. And store the stylized digital human video in a server or cloud storage. Create a download path and / or a preview path for the stylized digital human video, and send the download path and / or the preview path of the stylized digital human video to an interactive platform for access by a user terminal.

[0080] In this embodiment, on the one hand, the present application can generate a digital human image of a target stylized type according to a received user photo, avoiding the limitations of pre-made digital human images, and improving flexibility and user satisfaction. On the other hand, the present application uses a pre-trained lip motion driving model to be able to generate a synchronized lip motion driving video in real time according to a dubbing file. Compared with traditional non-real-time lip motion driving technologies, the real-time performance of video generation is greatly improved. On the third hand, based on the technical solution of the present application to generate a stylized digital human video, the user only needs to provide a photo, a dubbing file, and select a target stylized type. This greatly simplifies the user operation process and reduces the participation threshold.

[0081] In one embodiment, the lip motion driving model includes a generator network and a discriminator network, and the discriminator network includes a lip synchronization discriminator, as Figure 2 shown, before the step S30, it further includes:

[0082] Step A10, obtain a training set, where the training set includes digital human lip motion driving videos of various stylized types.

[0083] In this embodiment, the digital human lip motion driving videos of various stylized types included in the training set are videos obtained by performing lip motion driving on the stylized digital human in the stylized digital human image through a non-real-time lip motion driving technology (such as sadtalker, etc.) based on the pre-made stylized digital human image.

[0084] In some embodiments, perform facial feature point detection on the pre-made stylized digital human image; according to the distribution of the feature points, determine whether there are regions with abnormal distribution of feature points in each pre-made stylized digital human image, where the regions with abnormal distribution of feature points include abnormal mouth opening and closing (too large), abnormal eye closure (incomplete closure), abnormal jaw (distorted), etc.; remove the stylized digital human images with regions with abnormal distribution of feature points to obtain a stylized digital human image sample set, and the stylized digital human image sample set is used for subsequent production of a digital human lip motion driving video set.

[0085] In some embodiments, as Figure 3 shown, before the step A10, it includes:

[0086] Step A00: Obtain a set of lip-sync driving videos of digital humans with various stylized types;

[0087] It should be noted that compared with real humans or highly realistic digital humans, the proportions and sizes of the facial features of stylized digital humans will be beautified accordingly. If the existing digital human lip-sync driving technology is directly used for real-time lip-sync driving, there will be a problem that due to the adjustment of the facial features of the stylized digital human, the lip-sync driving cannot be correctly performed.

[0088] In this embodiment, the existing digital human lip-sync driving technology is used to produce a set of lip-sync driving videos of digital humans with various stylized types. Since the videos in the lip-sync driving video set are pre-produced and the production process can be artificially intervened, there are no videos with abnormal lip-sync driving in the lip-sync driving video set of this embodiment.

[0089] Step A01: Screen out the lip-sync driving videos that meet the preset audio-video synchronization standard from the lip-sync driving video set to construct an initial training set;

[0090] The preset audio-video synchronization standard in this embodiment includes at least one of the following: delay threshold, which refers to the maximum allowable delay between audio and video; synchronization accuracy, which refers to the degree of synchronization between audio and video and can be measured by calculating the synchronization rate or cross-correlation of audio and video; lip-sync accuracy, which refers to the degree of synchronization between the lip movement changes of the digital human and the pronunciation of the audio.

[0091] Step A02: Based on the stylized labels of the lip-sync driving videos in the initial training set, classify the lip-sync driving videos with the same stylized label into the same data set;

[0092] In this embodiment, traverse the initial training set, read the stylized label of each lip-sync driving video; create a dictionary with the stylized label as the key and the video list of the corresponding stylized type as the value; according to the stylized labels of each lip-sync driving video, add each lip-sync driving video to the data set of the corresponding stylized type (i.e., under the corresponding key of the dictionary).

[0093] Step A03: Based on the classified data sets, form a training set.

[0094] In this embodiment, the structured training set formed based on the classified data sets can not only be uniformly used for the subsequent training of the lip-sync driving model, playing the role of improving the training efficiency, enhancing the generalization ability of the model and simplifying the evaluation process, but also the individual data sets in the training set can be respectively used for the training of the lip-sync driving model, so as to achieve the purpose of targeted training, multi-style support and avoiding style entanglement.

[0095] Step A20: Train the lip-sync discriminator's ability to evaluate the synchronization between lip movements and audio by driving the video with digital lip shapes of each stylized type.

[0096] In this embodiment, by driving the video with digital lip shapes of each stylized type, the lip-sync discriminator's ability to evaluate the synchronization between lip movements and audio is trained. This training method can help the lip-sync discriminator better understand and recognize the synchronization between digital lip shapes of different styles and audio, thereby improving its evaluation accuracy and generalization ability.

[0097] In some embodiments, as Figure 4 shown, step A20 includes:

[0098] Step A21: Extract the audio information of the digital lip shape-driven video of each stylized type and cut the audio information into audio chunks.

[0099] In this embodiment, the extracted audio information is cut into audio chunks of a fixed length, and each audio chunk corresponds to one frame of the digital lip shape-driven video.

[0100] In some embodiments, obtain the frame rate of the digital lip shape-driven video and the audio sampling rate, and determine the cutting length according to the frame rate and the audio sampling rate. Cut the audio information into several audio chunks of the cutting length.

[0101] Exemplarily, assuming the frame rate of the digital lip shape-driven video is 30fps and the audio sampling rate is 44.1kHz, the length of the audio chunk corresponding to one frame of the digital lip shape-driven video is 33ms (i.e., 33 * 44100 / 1000 = 1455 sampling points).

[0102] Exemplarily, assuming the frame rate of the digital lip shape-driven video is 25fps and the audio sampling rate is 16kHz, the length of the audio chunk corresponding to one frame of the digital lip shape-driven video is 16ms (i.e., 16 * 1000 / 16000 = 256 sampling points).

[0103] Step A22: Pair the audio chunks with the image frames in the digital lip shape-driven video of each stylized type to form matching pairs of the audio chunks and the corresponding image frames and non-matching pairs of the audio chunks and non-corresponding image frames.

[0104] Step A23: Train the lip-sync discriminator's ability to evaluate the synchronization between lip movements and audio through the matching pairs and the non-matching pairs.

[0105] In this embodiment, a matching pair refers to an audio block and the corresponding image frame, which are synchronized in content. Through the matching pairs, the lip-sync discriminator can learn the correct correspondence between the audio blocks and the image frames. A non-matching pair refers to an audio block and a non-corresponding image frame, which are not synchronized in content. Through the non-matching pairs, the model can learn the incorrect correspondence between the audio blocks and the image frames.

[0106] In this embodiment, the matching pairs are used as positive samples, and the non-matching pairs are used as negative samples. By training the positive samples and the negative samples simultaneously, the lip-sync discriminator can learn to distinguish between synchronous and asynchronous situations, thereby improving its discrimination ability.

[0107] In some embodiments, as Figure 5 shown, step A23 further includes:

[0108] Step A24, extracting features from the audio blocks and the image frames in the matching pairs and the non-matching pairs to generate audio feature vectors and image feature vectors.

[0109] In this embodiment, features are extracted from the audio blocks in the matching pairs and the non-matching pairs to generate audio feature vectors. The audio feature extraction methods that can be used include MFCC (Mel Frequency Cepstral Coefficients), Spectrogram, Chroma features, etc. Also, features are extracted from the image frames in the matching pairs and the non-matching pairs to generate image feature vectors. The image feature extraction methods that can be used include using pre-trained convolutional neural networks (CNNs) such as VGG, ResNet, etc., or combining facial key point detection techniques.

[0110] Step A25, respectively calculating the cosine similarities between the audio feature vectors and the image feature vectors of the matching pairs and the non-matching pairs.

[0111] The cosine similarities between the audio feature vectors and the image feature vectors of the matching pairs and the non-matching pairs are calculated respectively. The cosine similarity is a method for measuring the similarity between two vectors, and the closer the value is to 1, the higher the similarity.

[0112] Step A26, based on the cosine similarities of the matching pairs and the non-matching pairs, feedback-update the lip-sync discriminator by minimizing the matching loss until the training is completed.

[0113] In this embodiment, the cosine similarity between the matching pairs and the non-matching pairs is mapped to the probability interval of [0, 1] and converted into a matching probability value. The matching loss is calculated based on the matching probability values of the matching pairs and the non-matching pairs. The matching loss is a loss function used to measure the errors made by the lip-sync discriminator when determining whether an audio block and an image frame match. By minimizing the matching loss, feedback updates can be made to the lip-sync discriminator, thereby optimizing the performance of the lip-sync discriminator so that it can more accurately distinguish between matching pairs and non-matching pairs.

[0114] Step A30: After training the lip-sync discriminator, extract the reconstructed target frames from the training set. By masking the lip part of the reconstructed target frames, train the generator network to generate reconstructed image frames based on the reconstructed target frames with the masked lip parts, and calculate the reconstruction loss based on the reconstructed target frames and the reconstructed image frames.

[0115] In this embodiment, the reconstructed target frames with the masked lip parts are input into the generator network to train the generator network to generate a reconstructed image frame that is as similar as possible to the reconstructed target frame, especially the lip part is as similar as possible. To evaluate the performance of the generator network, the reconstruction loss is calculated based on the reconstructed target frames and the reconstructed image frames. The reconstruction loss is used to measure the difference between the reconstructed target frames and the generated reconstructed image frames.

[0116] In some embodiments, as Figure 6 shown, the step A30 includes:

[0117] Step A31: Based on the training set, obtain the reference frames, the reconstructed target frames, and the audio segments corresponding to the time dimension of the reconstructed target frames, and use the reference frames, the reconstructed target frames, and the audio segments as the inputs of the generator network. Among them, the reference frames contain complete face features, and the lip parts of the reconstructed target frames are masked.

[0118] Step A32: Obtain the reconstructed image frames generated by the generator network based on the reference frames, the reconstructed target frames, and the audio segments.

[0119] In this embodiment, the reference frame is an image frame containing complete face features, which is used to provide the structure and features of the lip part. The lip parts of the reconstructed target frames have been masked, which is used to provide the image structure and features other than the lip part. The audio segments corresponding to the time dimension of the reconstructed target frames include lip movement information.

[0120] Input the reference frame, the reconstructed target frame, and the audio segment into the generator network together. The generator network generates a reconstructed image frame based on the structure and features of the lip part provided by the reference frame, the image structure and features other than the lip part provided by the reconstructed target frame, and the lip movement information.

[0121] Step A33: Calculate the reconstruction loss based on the reconstructed target frame and the reconstructed image frame where the lip part is not masked.

[0122] Step A34: Update the generator network by minimizing the reconstruction loss until the training is completed.

[0123] Calculate the reconstruction loss based on the pixel value difference between the reconstructed target frame and the reconstructed image frame where the lip part is not masked. The reconstruction loss is used to adjust the parameters of the generator network to obtain better performance in the next iteration.

[0124] Step A40: Calculate the synchronization loss between the lip of the reconstructed image frame and the lip of the reconstructed target frame based on the trained lip synchronization discriminator.

[0125] In this embodiment, the reconstructed image frames generated by the generator network are consecutive frames. Calculate the synchronization loss based on the training samples associated with the reconstructed target frame for the consecutive frames. The synchronization loss takes into account the context information of the reconstructed image frame and the reconstructed target frame and can more accurately measure the lip synchronization quality of the generated reconstructed image frames. The training samples are video segments with a fixed number of frames cut from the digital human lip-driven videos in the training set. The number of frames in the fixed number is the same as the number of consecutive frames generated.

[0126] Step A50: Update the generator network based on the reconstruction loss and the synchronization loss until the training is completed.

[0127] Calculate the comprehensive loss function based on the reconstruction loss and the synchronization loss. The comprehensive loss function includes hyperparameters for balancing the importance of the reconstruction loss and the synchronization loss. Adjust the parameters of the generator network according to the comprehensive loss function to obtain better performance in the next iteration until the comprehensive loss function maintains at a minimum value and the training is completed.

[0128] In this embodiment, by creating a training set including digital human lip-driven videos of various stylized types, extracting features, training the generator network, calculating the synchronization loss, and performing feedback updates, the lip-driven model can learn the characteristics of different styles and generate digital human videos with corresponding styles. At the same time, by calculating the synchronization loss and performing feedback updates, it can ensure good lip synchronization between the generated digital human video and the original input data, thereby improving the realism and naturalness of the generated digital human video.

[0129] In one embodiment, the generator network includes an identity encoder, a voice encoder, and a face decoder. The step A31 includes:

[0130] Based on the training set, obtain a reference frame, a reconstructed target frame, and an audio segment corresponding to the reconstructed target frame in the time dimension; concatenate the reference frame and the reconstructed target frame along the channel dimension as the input of the identity encoder to obtain the identity features output by the identity encoder; and use the audio segment as the input of the voice encoder to obtain the voice features output by the voice encoder; concatenate the identity features and the voice features as the input of the face decoder to obtain the reconstructed image frame generated by the face decoder.

[0131] In this embodiment, the identity encoder, the voice encoder, and the face decoder are all neural network modules within the generator network. By separating the identity features and the voice features and processing them separately by the identity encoder and the voice encoder, the feature information in the reconstructed target frame and the audio segment can be better captured. This helps to improve the ability of the face decoder to generate realistic and synchronized reconstructed image frames.

[0132] In one embodiment, the discriminator network further includes a visual quality discriminator. After the step A30, it further includes:

[0133] Supervise the image quality of the reconstructed image frame through the visual quality discriminator and calculate the adversarial loss. The adversarial loss is used to measure the adversarial performance of the reconstructed image frame against the discriminator network in the adversarial training.

[0134] In this embodiment, compared with the lip discriminator for discriminating the lip synchronization with the audio, the visual quality discriminator is used to discriminate the image quality of the generated reconstructed image frame to avoid blurring or artifacts in the reconstructed image frame. The adversarial loss is used to measure the adversarial performance of the reconstructed image frame against the discriminator network in the adversarial training, helping the generator network learn a more accurate image reconstruction method and improve the visual quality of the generated reconstructed image frame.

[0135] Among them, the larger the adversarial loss, the easier it is for the visual quality discriminator to identify the difference between the reconstructed image frame and the reconstructed target frame, the worse the performance of the generator, and the lower the quality of the generated reconstructed image frame. Among them, the smaller the adversarial loss, the less likely it is for the visual quality discriminator to identify the difference between the reconstructed image frame and the reconstructed target frame, the better the performance of the generator, and the higher the quality of the generated reconstructed image frame. That is, the adversarial loss is inversely proportional to the image quality of the reconstructed image frame.

[0136] The step A50 includes:

[0137] The generator network is updated by feedback by minimizing the weighted sum of the reconstruction loss, the synchronization loss, and the adversarial loss until the training is completed.

[0138] In this embodiment, in each iteration of the generator network, the parameters of the generator network are updated according to the sum of the gradients of the reconstruction loss, the synchronization loss, and the adversarial loss. This process continues until a certain stopping condition is met, such as reaching a preset number of training epochs or the convergence of the loss function. By balancing different types of losses, the generator network can generate high-quality and realistic high-resolution images while retaining image details.

[0139] In one embodiment, a computer-readable storage medium is provided, on which executable instructions are stored, and when the instructions are executed by a processor, the processor executes the steps in the above method embodiments.

[0140] In one embodiment, an electronic device is further provided, including one or more processors; a memory, and one or more programs are stored in the memory, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the steps in the above method embodiments.

[0141] In one embodiment, as Figure 7 shown, it shows a schematic structural diagram of an electronic device for implementing the embodiments of the present application. The electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0142] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.

[0143] In particular, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer-readable medium carrying instructions. In such an embodiment, the instructions can be downloaded and installed from a network through a communication section 709, and / or installed from a removable medium 711. When the instructions are executed by a central processing unit (CPU) 701, the various method steps described in the present application are executed.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0145] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the above claims, any one of the claimed embodiments can be used in any combination. The information disclosed in this background section is only intended to deepen the understanding of the overall background of the present application, and should not be regarded as an admission or any form of implication that this information constitutes prior art known to those skilled in the art.

Claims

1. A method for generating stylized digital human videos, characterized in that: The method comprises: Receiving a stylized digital human video generation instruction, wherein the stylized digital human video generation instruction includes a user photo, a target stylization type, and a dubbing file; According to the target stylization type, a style image associated with the target stylization type is obtained, content features of the user photo and style features of the style image are extracted through a pre-trained convolutional neural network, the content features are used to retain the structure and content of the user photo, and the style features are used to capture the texture and color information of the style image, an initial stylized image is randomly generated according to the content features and the style image, content loss of the initial stylized image and the user photo is calculated, the content loss is used to measure the difference between the content features of the initial stylized image and the content features of the user photo, style loss of the initial stylized image and the style image is calculated, the style loss is used to measure the difference between the style features of the initial stylized image and the style features of the style image, and pixel values ​​of the generated stylized image are adjusted in each iteration by minimizing a combination of the content loss and the style loss; when it is detected that the combination of the content loss and the style loss tends to be stable, the iteration is confirmed to be finished, and the target stylized image finally generated is obtained; Inputting the target stylized image and the dubbing file into a pre-trained lip-sync driven model, wherein the pre-trained lip-sync driven model extracts identity features of the target stylized image and audio features of the dubbing file, and generates a stylized digital human video according to the identity features and the audio features; Acquire the stylized digital human video output by the pre-trained lip-driven model, wherein the lip-driven stylized digital human video is synchronized with the dubbing file; The lip-sync driven model includes a generator network and a discriminator network, the discriminator network includes a lip-sync discriminator, and before the target stylized image and the dubbing file are input into the pre-trained lip-sync driven model, the method further includes: A training set is obtained, wherein the training set includes digital population-driven videos of various stylized types, audio information of the digital population-driven videos of various stylized types is extracted, and the audio information is cut into audio blocks, and the audio blocks are paired with image frames in the digital population-driven videos of various stylized types to form matching pairs of the audio blocks and corresponding image frames and non-matching pairs of the audio blocks and non-corresponding image frames, and feature extraction is performed on the audio blocks and image frames in the matching pairs and the non-matching pairs to generate audio feature vectors and image feature vectors, and cosine similarities between the audio feature vectors and image feature vectors of the matching pairs and the non-matching pairs are calculated respectively, and the cosine similarities are mapped to [0, 1], convert it into a matching probability value, calculate the matching loss based on the matching probability value, and feedback update the lip synchronization discriminator by minimizing the matching loss until the training of the lip synchronization discriminator is completed. After the lip synchronization discriminator is trained, a reconstructed target frame is extracted from the training set, and the generator network is trained to generate a reconstructed image frame based on the reconstructed target frame with the masked lip shape part by masking the lip shape part of the reconstructed target frame, and the reconstruction loss is calculated based on the reconstructed target frame and the reconstructed image frame, and the synchronization loss between the lip shape of the reconstructed image frame and the lip shape of the reconstructed target frame is calculated based on the trained lip synchronization discriminator, and the generator network is feedback updated based on the reconstruction loss and the synchronization loss until the training of the generator network is completed.

2. The method for generating stylized digital human videos according to claim 1, characterized in that: The discriminator network further includes a visual quality discriminator, and after masking the lip shape part of the reconstructed target frame and training the generator network to generate a reconstructed image frame based on the reconstructed target frame of the masked lip shape part, further includes: The image quality of the reconstructed image frame is supervised by the visual quality discriminator, and an adversarial loss is calculated, where the adversarial loss is used to measure the adversarial performance of the reconstructed image frame against the discriminator network in adversarial training; The step of performing feedback updating on the generator network based on the reconstruction loss and the synchronization loss until the training is completed comprises: The generator network is feedback updated by minimizing the weighted sum of the reconstruction loss, the synchronization loss and the adversarial loss until the training is completed.

3. The method for generating stylized digital human videos according to claim 1, characterized in that: The step of extracting a reconstructed target frame from the training set, masking the lip shape part of the reconstructed target frame, training the generator network to generate a reconstructed image frame based on the reconstructed target frame of the masked lip shape part, and calculating the reconstruction loss based on the reconstructed target frame and the reconstructed image frame includes: Based on the training set, a reference frame, a reconstructed target frame, and an audio clip of a time dimension corresponding to the reconstructed target frame are obtained, and the reference frame, the reconstructed target frame, and the audio clip are used as inputs of the generator network, wherein the reference frame contains complete facial features, and the mouth shape part of the reconstructed target frame is masked; Obtain a reconstructed image frame generated by the generator network based on the reference frame, the reconstructed target frame and the audio clip; Calculating a reconstruction loss based on the reconstructed target frame and the reconstructed image frame in which the lip shape part is not masked; By minimizing the reconstruction loss, the generator network is fed back and updated until the training is completed.

4. The method for generating stylized digital human video according to claim 1, characterized in that: The generator network includes an identity encoder, a speech encoder and a face decoder. Based on the training set, a reference frame, a reconstructed target frame and an audio clip of a time dimension corresponding to the reconstructed target frame are obtained, and the reference frame, the reconstructed target frame and the audio clip are used as inputs of the generator network, including: Based on the training set, a reference frame, a reconstructed target frame, and an audio segment of a time dimension corresponding to the reconstructed target frame are obtained; splicing the reference frame and the reconstructed target frame together according to the channel dimension as input of the identity encoder to obtain the identity feature output by the identity encoder; and Using the audio segment as input of the speech encoder to obtain speech features output by the speech encoder; The identity feature and the voice feature are concatenated as input to the facial decoder to obtain a reconstructed image frame generated by the facial decoder.

5. The method for generating stylized digital human video according to claim 1, characterized in that: Before obtaining a training set, wherein the training set includes digital population-driven videos of various stylized types, the method includes: Obtain a collection of lip-sync-driven videos containing various stylized types of digital humans; Selecting lip-sync-driven videos that meet a preset audio and video synchronization standard from the lip-sync-driven video set to construct an initial training set; Based on the stylized labels of the lip-sync-driven videos in the initial training set, classifying the lip-sync-driven videos with the same stylized labels into the same dataset; Based on the classified data set, a training set is formed.

6. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the stylized digital human video generation method according to any one of claims 1 to 5.

7. A storage medium, characterized in that: The storage medium stores executable instructions, and when the instructions are executed by the processor, the processor executes the method for generating a stylized digital human video according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Virtual digital human lip shape synchronization method and system

    CN116668611A

  • Digital human driving and diffusion model training method and device

    CN118522303A