Method and device for generating digital human video, electronic equipment and storage medium

By using a keypoint localization model, an affine transformation offset model, and a generative adversarial network in the digital human video generation process, combined with sentiment analysis and interpolation algorithms, the problems of feature offset and distortion in digital human videos are solved, achieving more accurate and natural video generation.

CN116071223BActive Publication Date: 2026-04-28AVATAR WORKS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AVATAR WORKS INC
Filing Date
2022-12-13
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In the existing technology, there are feature offsets and distortions in the process of generating digital human videos. The lack of consideration for emotions leads to deformation of the mouth and face areas, resulting in uncontrolled deviations in parameters such as the average value of the mouth, which undermines the authenticity of the broadcast video.

Method used

By acquiring target facial images and text, facial key points are identified using a preset key point localization model. Combined with a preset affine transformation offset model and a generative adversarial network, digital human videos are generated. The influence of emotions and text content on the position of facial key points is considered, and a cubic interpolation algorithm and a dual discriminator are used to optimize video smoothness.

Benefits of technology

More accurate and realistic digital human videos were generated, improving the accuracy of facial deformation and video smoothness, thus enhancing the realism and naturalness of digital human videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116071223B_ABST
    Figure CN116071223B_ABST
Patent Text Reader

Abstract

The application discloses a digital person video generation method and device, electronic equipment and storage medium, and the method comprises the steps of obtaining a target image comprising a target face and target text to be broadcast; inputting the target image into a preset key point positioning model to perform face key point recognition, and obtaining multiple face key points in the target image; inputting the target text and the face key points into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text, wherein each set of affine transformation offset parameters represents an affine transformation offset value of each face key point when the target face broadcasts the character; and generating a digital person video of the target face broadcasting the target text according to the target image, the face key points in the target image and the sets of affine transformation offset parameters. Based on the face key points, an affine transformation corresponding to the target text is performed to obtain a more accurate face region deformation result, thereby generating a more accurate digital person video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a method, apparatus, electronic device, and storage medium for generating digital human videos. Background Technology

[0002] A digital human is a virtual simulation of the human body at different levels of form and function, using information science methods. Digital human video broadcasts have been widely used in fields such as virtual avatars and virtual anchors.

[0003] In existing technologies, audio is used as input. The audio is fed into a generator to obtain facial parameters of the target face. These parameters, along with an image of the target face with the mouth area obscured, are then input into a second generator. This two-step generation process yields a digital human broadcasting video. However, this method results in distortion due to feature shifts and averaging. Furthermore, it fails to consider emotional distortions of the mouth and facial areas, leading to uncontrolled shifts in parameters such as the average mouth size and unguided spontaneous shaking, thus compromising the realism of the broadcasting video.

[0004] Therefore, how to generate digital human videos more accurately is a technical problem that needs to be solved.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] This application provides a method, apparatus, electronic device, and storage medium for generating digital human videos, so as to generate digital human videos more accurately.

[0007] In a first aspect, a method for generating a digital human video is provided. The method includes: acquiring a target image including a target face and target text to be read; inputting the target image into a preset key point localization model to perform facial key point recognition, obtaining multiple facial key points in the target image; inputting the target text and each of the facial key points into a preset affine transformation offset model, obtaining multiple sets of affine transformation offset parameters corresponding to each character in the target text, each set of affine transformation offset parameters representing the affine transformation offset value generated by each of the facial key points when the target face reads the character; and generating a digital human video of the target face reading the target text based on the target image, each facial key point in the target image, and each set of affine transformation offset parameters.

[0008] Secondly, a digital human video generation apparatus is provided, the apparatus comprising: an acquisition module for acquiring a target image including a target face and target text to be read; a positioning module for inputting the target image into a preset key point positioning model for facial key point recognition to obtain multiple facial key points in the target image; a parsing module for inputting the target text and each of the facial key points into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text, each set of affine transformation offset parameters representing the affine transformation offset value generated by each of the facial key points when the target face reads the character; and a generation module for generating a digital human video of the target face reading the target text based on the target image, each facial key point in the target image, and each set of affine transformation offset parameters.

[0009] Thirdly, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the digital human video generation method of the first aspect by executing the executable instructions.

[0010] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method for generating digital human videos as described in the first aspect.

[0011] By applying the above technical solutions, a target image including the target face and the target text to be broadcast are obtained; the target image is input into a preset key point localization model for facial key point recognition to obtain multiple facial key points in the target image; the target text and each of the facial key points are input into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text, each set of affine transformation offset parameters representing the affine transformation offset value generated by each of the facial key points when the target face broadcasts the character; a digital human video of the target face broadcasting the target text is generated based on the target image, each facial key point in the target image, and each set of affine transformation offset parameters. Based on the facial key points, an affine transformation corresponding to the target text is performed to obtain a more accurate facial region deformation result, thereby generating a more accurate digital human video. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating a method for generating digital human videos according to an embodiment of the present invention is shown.

[0014] Figure 2 A flowchart illustrating a method for generating digital human videos according to another embodiment of the present invention is shown.

[0015] Figure 3 A flowchart illustrating a method for generating digital human videos according to another embodiment of the present invention is shown.

[0016] Figure 4 A schematic diagram of the structure of a digital human video generation device according to an embodiment of the present invention is shown;

[0017] Figure 5 A schematic diagram of the structure of an electronic device according to an embodiment of the present invention is shown. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] It should be noted that other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this application are indicated in the claims section.

[0020] It should be understood that this application is not limited to the precise structure described below and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

[0021] It should be noted that the following application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the implementation of this application is not limited in any way. On the contrary, the implementation of this application can be applied to any applicable scenario.

[0022] This application provides a method for generating digital human videos, such as... Figure 1 As shown, the method includes the following steps:

[0023] Step S101: Obtain the target image including the target face and the target text to be broadcast.

[0024] In this embodiment, the digital human video to be generated is a digital human video in which a target face broadcasts target text. Therefore, it is necessary to first obtain a target image including the target face and the target text to be broadcast. The target face can be a real face or a virtual avatar face. The target image can be an image obtained after capturing the target face, or an image including the target face extracted from a preset video. Optionally, the target image format can be any of the formats including JPEG, TIF, EPS, BMP, PCX, etc. The target text can be a piece of text, and the target text format can be any of the formats including txt, doc, etc. The target image and target text can be uploaded by the user or obtained from an external terminal or server.

[0025] Step S102: Input the target image into a preset key point localization model to perform facial key point recognition, and obtain multiple facial key points in the target image.

[0026] In this embodiment, a preset keypoint localization model is pre-built and trained, and facial keypoint recognition is performed based on this preset keypoint localization model. Facial keypoint recognition refers to locating the key regions of a face from a facial image. These key regions include eyebrows, eyes, nose, mouth, and facial contours. After completing facial keypoint recognition, the preset keypoint localization model outputs multiple facial keypoints, each corresponding to a pixel coordinate value. Those skilled in the art can identify multiple facial keypoints of different numbers and locations according to actual needs. Optionally, the preset keypoint localization model can be any of the following models: Cascaded Regression CNN Facial Keypoint Detection Model, Dlib Facial Keypoint Detection Model, libfacedetect Facial Keypoint Detection Model, Seetaface Facial Keypoint Detection Model, etc.

[0027] Step S103: Input the target text and each of the facial key points into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text.

[0028] In this embodiment, the affine transformation is a linear transformation between two-dimensional coordinates, which can maintain the flatness and parallelism of two-dimensional graphics. During the process of the target face pronouncing each character in the target text, the positions of each face key point will change accordingly, and this change is reflected in the different affine transformation offset values of each face key point under different characters. A preset affine transformation offset model is established and trained in advance. By inputting the target text and each face key point into the preset affine transformation offset model, multiple groups of affine transformation offset parameters can be obtained. Each group of affine transformation offset parameters represents the affine transformation offset value generated by each face key point when the target face pronounces each character, that is, each group of affine transformation offset parameters can cause a position change of each face key point.

[0029] In some embodiments of the present application, the target text is composed of multiple statement units, and each statement unit has an emotion label representing emotion. Before inputting the target text and each face key point into the preset affine transformation offset model, the method further includes:

[0030] Input each character of the sample text with the emotion label into the preset recurrent neural network model in sequence, and use the affine transformation offset value of the face key points of the preset standard face as the output to train the preset recurrent neural network model. After the training is completed, the preset affine transformation offset model is obtained.

[0031] Considering that the target face may show different facial area changes when reading the same text content under different emotions, in this embodiment, emotion labels representing emotions are set for each statement unit in the target text in advance. For example, the emotion label of the statement unit "Hello" is "excited", and the emotion label of the statement unit "Nice to meet you" is "gentle", etc. In the specific application scenario of the present application, the format example of the target text with emotion labels is: {{text: Hello, emotion: excited}, {text: Nice to meet you, emotion: gentle}}. Training the preset recurrent neural network model with the sample text with emotion labels can obtain the preset affine transformation offset model. Specifically, input each character in the sample text into the preset recurrent neural network model in sequence, so that the preset recurrent neural network model outputs the affine transformation offset value of the face key points of the preset standard face, so as to perform affine transformation learning on the relative positions of the face key points of the preset standard face. After the training is completed, the preset affine transformation offset model is obtained. Since the influence of emotion and text content on the position of face key points is considered, the preset affine transformation offset model outputs more accurate affine transformation offset parameters, thereby enhancing the authenticity of the digital human video.

[0032] Optionally, sentiment labels can be manually assigned to each sentence unit in the target text, or a trained sentiment analysis model can be used to analyze the target text and obtain sentiment labels for each sentence unit.

[0033] Optionally, the preset recurrent neural network model can be replaced with other deep learning models, such as convolutional neural network models, generative adversarial network models, long short-term memory network models, etc.

[0034] In some embodiments of this application, after obtaining multiple sets of affine transformation offset parameters corresponding to each character in the target text, the method further includes:

[0035] The cubic interpolation algorithm is used to interpolate every two adjacent affine transformation offset parameters to obtain multiple sets of intermediate affine transformation offset parameters.

[0036] The intermediate affine transformation offset parameters are inserted as a new set of affine transformation offset parameters between the two adjacent sets of affine transformation offset parameters corresponding to themselves.

[0037] In this embodiment, every two sets of adjacent affine transformation offset parameters cause a positional change in each facial keypoint. Significant positional changes can affect video smoothness; therefore, interpolation is required for every two sets of adjacent affine transformation offset parameters. Specifically, a cubic interpolation function is fitted based on a cubic interpolation algorithm. This function is then used to interpolate every two sets of adjacent affine transformation offset parameters to obtain multiple intermediate affine transformation offset parameters. These intermediate parameters are then inserted between their corresponding two sets of adjacent affine transformation offset parameters. After interpolation, each intermediate affine transformation offset parameter serves as a transition parameter between its corresponding two sets of adjacent affine transformation offset parameters, improving video smoothness while maintaining the basic frame rate.

[0038] Optionally, the cubic interpolation algorithm can be replaced with other types of interpolation algorithms, such as nearest neighbor interpolation, bilinear interpolation, etc.

[0039] Step S104: Generate a digital human video of the target face broadcasting the target text based on the target image, each facial key point in the target image, and each set of affine transformation offset parameters.

[0040] In this embodiment, each set of affine transformation offset parameters can cause corresponding positional changes in each facial key point in the target image, thereby obtaining multiple frames of images that have changed relative to the target image. Each set of affine transformation offset parameters corresponds to each character of the target text. Therefore, by combining each frame of images in sequence, a digital human video of the target face broadcasting the target text can be generated.

[0041] By applying the above technical solutions, a target image including the target face and the target text to be broadcast are obtained; the target image is input into a preset key point localization model for facial key point recognition to obtain multiple facial key points in the target image; the target text and each of the facial key points are input into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text, each set of affine transformation offset parameters representing the affine transformation offset value generated by each of the facial key points when the target face broadcasts the character; a digital human video of the target face broadcasting the target text is generated based on the target image, each facial key point in the target image, and each set of affine transformation offset parameters. Based on the facial key points, an affine transformation corresponding to the target text is performed to obtain a more accurate facial region deformation result, thereby generating a more accurate digital human video.

[0042] This application also proposes a method for generating digital human videos, such as... Figure 2 As shown, it includes the following steps:

[0043] Step S201: Obtain the target image including the target face and the target text to be broadcast.

[0044] In this embodiment, the digital human video to be generated is a digital human video in which a target face broadcasts target text. Therefore, it is necessary to first obtain a target image including the target face and the target text to be broadcast. The target face can be a real face or a virtual avatar face. The target image can be an image obtained after capturing the target face, or an image including the target face extracted from a preset video. The target image and target text can be uploaded by the user or obtained from an external terminal or server.

[0045] Step S202: Input the target image into a preset key point localization model to perform facial key point recognition, and obtain multiple facial key points in the target image.

[0046] In this embodiment, a preset keypoint localization model is pre-built and trained, and facial keypoint recognition is performed based on this model. Facial keypoint recognition refers to locating the key regions of a face from a facial image. These key regions include eyebrows, eyes, nose, mouth, and facial contours. After completing facial keypoint recognition, the preset keypoint localization model outputs multiple facial keypoints, each corresponding to a pixel coordinate value. Those skilled in the art can identify different numbers and locations of multiple facial keypoints according to actual needs.

[0047] Step S203: Input the target text and each of the facial key points into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text.

[0048] In this embodiment, as the target face reads each character in the target text, the positions of each facial keypoint will change accordingly. This change is reflected in the different affine transformation offset values ​​of each facial keypoint under different characters. A preset affine transformation offset model is established and trained in advance. By inputting the target text and each facial keypoint into the preset affine transformation offset model, multiple sets of affine transformation offset parameters can be obtained. Each set of affine transformation offset parameters represents the affine transformation offset value generated by each facial keypoint when the target face reads each character. That is, each set of affine transformation offset parameters causes each facial keypoint to undergo a positional change.

[0049] Step S204: Input the target image and each facial key point in the target image into the generator in the preset generative adversarial network, and input a set of affine transformation offset parameters into the generator in sequence each time, and generate a digital human video of the target face broadcasting the target text based on the multiple frames of images output by the generator in sequence.

[0050] Generative Adversarial Networks (GANs) are a neural network paradigm that includes a generator and a discriminator. The generator produces predictions, while the discriminator evaluates these predictions. During training, the two components engage in a zero-sum game. In this embodiment, the generator in a pre-trained GAN predicts each frame of the image. The target image and its facial landmarks serve as the basic data. The generator's basic input is the current frame image and its facial landmarks, with the first current frame image being the target image. The generator generates one frame based on each set of affine transformation offset parameters. Therefore, the generator's conditional input is the affine transformation offset parameters for the next frame image, and its output is the next frame image.

[0051] The target image and the facial key points in the target image are used as the first frame data input to the generator. A set of affine transformation offset parameters are input to the generator in sequence each time. The generator outputs multiple frames of images corresponding to each set of affine transformation offset parameters in sequence. After the images of each frame are synthesized in sequence, a digital human video in which the target face broadcasts the target text can be generated.

[0052] The generator in the pre-defined generative adversarial network processes the target image, the key facial points in the target image, and the affine transformation offset parameters of each set to obtain multiple frames of images, thereby achieving more accurate generation of digital human videos.

[0053] Those skilled in the art can use different network structures to train the generator according to actual needs. In some embodiments of this application, the ResNet-101 fast training residual network is used as the basic network architecture of the generator. Preset guidance conditions are added during the generator training process. These preset guidance conditions are used to regenerate preset key regions in the face. The preset key regions may include the mouth, teeth, and other regions. This makes the output frames of the generator more consistent with the target face and avoids distortion and jitter in local key regions (mouth, teeth, etc.).

[0054] Optionally, the generator can also employ a recurrent neural network to more accurately capture the relationship between consecutive frames, resulting in a more natural and smooth flow of micro-movements on the target face.

[0055] In some embodiments of this application, the preset generative adversarial network further includes a frame discriminator and a flow discriminator. The frame discriminator is used to determine whether a preset key region in the next frame image is consistent with the real image of the preset key region. The flow discriminator is used to determine whether the optical flow data change value between the frames output by the generator is consistent with the preset change value.

[0056] In this embodiment, the pre-defined generative adversarial network includes two discriminators: a frame discriminator and a stream discriminator. The frame discriminator determines whether a preset key region in the next frame matches the real image of that preset key region. This allows for adjustments to the preset key region in the next frame based on the real image of the face, improving the realism of local facial details. The stream discriminator determines whether the optical flow data changes between frames output by the generator match preset values. Optical flow represents changes in image brightness patterns. By determining whether the optical flow data changes match preset values, better transitions between video frames are guided, ensuring the smoothness of pixel flow between frames and resulting in a smoother and more natural final video.

[0057] In some embodiments of this application, before inputting the target image and each facial key point in the target image into a generator in a preset generative adversarial network, and sequentially inputting a set of affine transformation offset parameters into the generator each time, the method further includes:

[0058] The preset initial generative adversarial network is trained based on preset sample data to obtain the first loss value generated by the frame discriminator and the second loss value generated by the stream discriminator.

[0059] The first loss value and the second loss value are weighted and summed to obtain the third loss value;

[0060] The parameters of the generator are adjusted based on the third loss value;

[0061] When the preset training stopping condition is met, the preset generative adversarial network is obtained.

[0062] In this embodiment, a preset initial generative adversarial network is trained using preset sample data. During training, the frame discriminator generates a first loss value based on its own discrimination results. This first loss value represents the difference between a preset key region and the real image of the preset key region in the next frame. The flow discriminator generates a second loss value based on its own discrimination results. This second loss value represents the difference between the optical flow data change value and a preset change value. The first and second loss values ​​are weighted and summed to obtain a third loss value. Then, the generator parameters (such as weight parameters) are adjusted based on this third loss value, and training continues until a preset training stopping condition is met, resulting in the preset generative adversarial network. By weighting and summing the first and second loss values ​​and using the obtained third loss value to optimize the generator parameters, the realism of individual frames and the smoothness between frames are considered simultaneously, making the video more natural and smooth overall.

[0063] Optionally, a third loss value can be obtained from the sum of the first loss value and the second loss value.

[0064] Optionally, the preset training stopping condition can be reaching a preset number of iterations or the third loss value being less than a preset threshold.

[0065] Optionally, the preset sample data is a dataset of multiple videos featuring human voices broadcasting, with facial landmarks manually annotated.

[0066] In some embodiments of this application, after obtaining the preset generative adversarial network, the method further includes:

[0067] The individual feature data of the target face are obtained based on the real broadcast video of the target face;

[0068] The individual feature data is used as the generation guiding condition for the generator;

[0069] The individual feature data includes dental features and facial texture motion and distribution parameters during expression changes.

[0070] In this embodiment, the target face exhibits unique individual features during real-world broadcasting, such as tooth features, facial texture movement and distribution parameters when facial expressions change. Therefore, corresponding individual feature data can be obtained from the real-world broadcasting video of the target face. This individual feature data can be used as a generation guide condition for the generator, allowing the generator to consider these individual feature data when generating images, thus making the next frame image output by the generator more consistent with the target face.

[0071] By applying the above technical solutions, a target image including the target face and the target text to be broadcast are obtained; the target image is input into a preset key point localization model for face key point recognition, resulting in multiple face key points in the target image; the target text and each face key point are input into a preset affine transformation offset model, resulting in multiple sets of affine transformation offset parameters corresponding to each character in the target text, each set of affine transformation offset parameters representing the affine transformation offset value generated by each face key point when the target face broadcasts the character; the target image and each face key point in the target image are input into a generator in a preset generative adversarial network, and a set of affine transformation offset parameters are input into the generator sequentially each time, generating a digital human video of the target face broadcasting the target text based on the multiple frames of images output sequentially by the generator; wherein, the basic input of the generator is the current frame image and each face key point in the current frame image, the conditional input of the generator is the affine transformation offset parameters of the next frame image, the output of the generator is the next frame image, and the first current frame image is the target image. Based on facial key points, an affine transformation corresponding to the target text is performed to obtain more accurate facial region deformation results. Then, based on the generator in the preset generative adversarial network, multiple frames of images are generated to achieve the generation of more accurate digital human videos.

[0072] To further illustrate the technical concept of this invention, the technical solution of this invention will now be described in conjunction with specific application scenarios.

[0073] This application proposes a method for generating digital human videos, such as... Figure 3 As shown, it includes the following steps:

[0074] Step S1: Input the target image into the preset key point localization model to perform facial key point recognition and obtain multiple facial key points in the target image.

[0075] In this embodiment, a preset key point localization model is trained using a training set of accurately labeled facial key points. Based on this preset key point localization model, facial key point recognition is performed, and multiple facial key points are output, each corresponding to a pixel coordinate value.

[0076] Step S2: Input the target text and each facial key point into the preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text.

[0077] In this embodiment, the target text consists of multiple sentence units, each with an emotion tag representing emotion. By inputting the target text and each facial key point into a preset affine transformation offset model, multiple sets of affine transformation offset parameters can be obtained. Each set of affine transformation offset parameters represents the affine transformation offset value generated by each facial key point when the target face pronounces each character.

[0078] To improve smoothness and ensure a basic frame rate, after obtaining multiple sets of affine transformation offset parameters, a cubic interpolation algorithm is used to interpolate every two adjacent sets of affine transformation offset parameters to obtain multiple sets of intermediate affine transformation offset parameters. The intermediate affine transformation offset parameters are then inserted as a new set of affine transformation offset parameters between the two sets of adjacent affine transformation offset parameters corresponding to them.

[0079] Step S3: Input the target image and the facial key points in the target image into the generator in the preset generative adversarial network, and input a set of affine transformation offset parameters into the generator in sequence. The generator outputs multiple frames of images in sequence.

[0080] In this embodiment, the basic input of the generator is the current frame image and the key facial points in the current frame image, the conditional input of the generator is the affine transformation offset parameters of the next frame image, the output of the generator is the next frame image, and the first current frame image is the target image.

[0081] The pre-defined generative adversarial network also includes a frame discriminator and a flow discriminator. The frame discriminator is used to determine whether the preset key region in the next frame image is consistent with the real image of the preset key region. The flow discriminator is used to determine whether the optical flow data change value between the frames output by the generator is consistent with the preset change value.

[0082] As shown by the dashed arrow in the figure, during the training of the preset initial generative adversarial network based on preset sample data, the first loss value generated by the frame discriminator and the second loss value generated by the flow discriminator are obtained. The first loss value and the second loss value are weighted and summed to obtain the third loss value. The parameters of the generator are adjusted based on the third loss value. Since the realism of the single frame image and the smoothness between each frame image are considered at the same time, the video is more natural and smooth as a whole.

[0083] Step S4: Combine multiple frames of images in sequence to obtain a digital human video in which the target face broadcasts the target text.

[0084] By applying the above technical solutions, a broadcast video can be generated using only a single target image containing the target face, reducing data acquisition costs and lowering the barrier to entry. Considering the combined influence of emotion and text content on the position of facial key points, the preset affine transformation offset model outputs more accurate affine transformation offset parameters, thereby enhancing the realism of the digital human video. Based on dual discriminator processing, the preset key regions in the next frame are adjusted using real images of the preset key regions of the face, improving the realism of local facial details while ensuring the smoothness of pixel flow between frames, resulting in a more fluid and natural final video.

[0085] This application also proposes an apparatus for generating digital human videos, such as... Figure 4 As shown, the device includes:

[0086] The acquisition module 401 is used to acquire a target image including the target face and the target text to be broadcast;

[0087] The positioning module 402 is used to input the target image into a preset key point positioning model to perform facial key point recognition and obtain multiple facial key points in the target image;

[0088] The parsing module 403 is used to input the target text and each of the facial key points into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text. Each set of affine transformation offset parameters represents the affine transformation offset value generated by each of the facial key points when the target face broadcasts the character.

[0089] The generation module 404 is used to generate a digital human video in which the target face broadcasts the target text based on the target image, each facial key point in the target image and each set of affine transformation offset parameters.

[0090] In specific application scenarios, the 404 generation module is used specifically for:

[0091] The target image and the facial key points in the target image are input into the generator in the preset generative adversarial network, and a set of affine transformation offset parameters are input into the generator in sequence each time. The digital human video is generated based on the multiple frames of images output by the generator in sequence.

[0092] The generator's basic inputs are the current frame image and the facial key points in the current frame image, the generator's conditional inputs are the affine transformation offset parameters of the next frame image, the generator's output is the next frame image, and the first current frame image is the target image.

[0093] In specific application scenarios, the target text consists of multiple sentence units, each sentence unit carrying an emotion tag representing emotion. The device also includes a first training module for:

[0094] Each character of the sample text with the emotion tag is input into a preset recurrent neural network model in sequence. The affine transformation offset value of the facial key points of a preset standard face is used as the output to train the preset recurrent neural network model. After training, the preset affine transformation offset model is obtained.

[0095] In specific application scenarios, the device further includes an interpolation module, used for:

[0096] The cubic interpolation algorithm is used to interpolate every two adjacent affine transformation offset parameters to obtain multiple sets of intermediate affine transformation offset parameters.

[0097] The intermediate affine transformation offset parameters are inserted as a new set of affine transformation offset parameters between the two adjacent sets of affine transformation offset parameters corresponding to themselves.

[0098] In specific application scenarios, the preset generative adversarial network also includes a frame discriminator and a flow discriminator. The frame discriminator is used to determine whether the preset key region in the next frame image is consistent with the real image of the preset key region. The flow discriminator is used to determine whether the optical flow data change value between the frames output by the generator is consistent with the preset change value.

[0099] In specific application scenarios, the device further includes a second training module, used for:

[0100] The preset initial generative adversarial network is trained based on preset sample data to obtain the first loss value generated by the frame discriminator and the second loss value generated by the stream discriminator.

[0101] The first loss value and the second loss value are weighted and summed to obtain the third loss value;

[0102] The parameters of the generator are adjusted based on the third loss value;

[0103] When the preset training stopping condition is met, the preset generative adversarial network is obtained.

[0104] In specific application scenarios, the device further includes a feature extraction module, used for:

[0105] The individual feature data of the target face are obtained based on the real broadcast video of the target face;

[0106] The individual feature data is used as the generation guiding condition for the generator;

[0107] The individual feature data includes dental features and facial texture motion and distribution parameters during expression changes.

[0108] By applying the above technical solutions, the digital human video generation device includes: an acquisition module for acquiring a target image including a target face and target text to be broadcast; a positioning module for inputting the target image into a preset key point positioning model to perform facial key point recognition, thereby obtaining multiple facial key points in the target image; a parsing module for inputting the target text and each of the facial key points into a preset affine transformation offset model, thereby obtaining multiple sets of affine transformation offset parameters corresponding to each character in the target text, each set of affine transformation offset parameters representing the affine transformation offset value generated by each of the facial key points when the target face broadcasts the character; and a generation module for generating a digital human video of the target face broadcasting the target text based on the target image, each facial key point in the target image, and each set of affine transformation offset parameters, by performing an affine transformation corresponding to the target text based on the facial key points to obtain a more accurate facial region deformation result, thereby generating a more accurate digital human video.

[0109] This invention also provides an electronic device, such as... Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0110] Memory 503 is used to store the processor's executable instructions;

[0111] Processor 501 is configured to execute the following via executing the executable instructions:

[0112] Acquire the target image, including the target face, and the target text to be broadcast;

[0113] The target image is input into a preset key point localization model for facial key point recognition to obtain multiple facial key points in the target image;

[0114] The target text and each of the facial key points are input into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text. Each set of affine transformation offset parameters represents the affine transformation offset value generated by each of the facial key points when the target face broadcasts the character.

[0115] A digital human video is generated based on the target image, the facial key points in the target image, and the affine transformation offset parameters of each group, so that the target face broadcasts the target text.

[0116] The aforementioned communication bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0117] The communication interface is used for communication between the aforementioned terminal and other devices.

[0118] The memory may include RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0119] The processors mentioned above can be general-purpose processors, including CPUs (Central Processing Units), NPs (Network Processors), etc.; they can also be DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0120] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the method for generating digital human videos as described above.

[0121] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the method for generating digital human videos as described above.

[0122] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0123] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0124] The various embodiments in this specification are described in a related manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0125] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A method for generating digital human videos, characterized in that, The method includes: Acquire the target image, including the target face, and the target text to be broadcast; The target image is input into a preset key point localization model for facial key point recognition to obtain multiple facial key points in the target image; The target text and each of the facial key points are input into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text. Each set of affine transformation offset parameters represents the affine transformation offset value generated by each of the facial key points when the target face broadcasts the character. Generating a digital human video of the target face reading the target text based on the target image, facial key points in the target image, and each set of affine transformation offset parameters, includes: inputting the target image and facial key points in the target image into a generator in a preset generative adversarial network, and sequentially inputting a set of affine transformation offset parameters into the generator each time, and generating the digital human video based on multiple frames of images sequentially output by the generator; wherein, the basic input of the generator is the current frame image and facial key points in the current frame image, the conditional input of the generator is the affine transformation offset parameters of the next frame image, the output of the generator is the next frame image, and the first current frame image is the target image.

2. The method as described in claim 1, characterized in that, The target text consists of multiple sentence units, each sentence unit carrying an emotion label representing emotion. Before inputting the target text and each of the facial key points into a preset affine transformation offset model, the method further includes: Each character of the sample text with the emotion tag is input into a preset recurrent neural network model in sequence. The affine transformation offset value of the facial key points of a preset standard face is used as the output to train the preset recurrent neural network model. After training, the preset affine transformation offset model is obtained.

3. The method as described in claim 1, characterized in that, After obtaining multiple sets of affine transformation offset parameters corresponding to each character in the target text, the method further includes: The cubic interpolation algorithm is used to interpolate every two adjacent affine transformation offset parameters to obtain multiple sets of intermediate affine transformation offset parameters. The intermediate affine transformation offset parameters are inserted as a new set of affine transformation offset parameters between the two adjacent sets of affine transformation offset parameters corresponding to themselves.

4. The method as described in claim 1, characterized in that, The preset generative adversarial network further includes a frame discriminator and a flow discriminator. The frame discriminator is used to determine whether the preset key region in the next frame image is consistent with the real image of the preset key region. The flow discriminator is used to determine whether the optical flow data change value between the frames output by the generator is consistent with the preset change value.

5. The method as described in claim 4, characterized in that, Before inputting the target image and the facial key points in the target image into a generator in a preset generative adversarial network, and sequentially inputting a set of affine transformation offset parameters into the generator, the method further includes: The preset initial generative adversarial network is trained based on preset sample data to obtain the first loss value generated by the frame discriminator and the second loss value generated by the stream discriminator. The first loss value and the second loss value are weighted and summed to obtain the third loss value; The parameters of the generator are adjusted based on the third loss value; When the preset training stopping condition is met, the preset generative adversarial network is obtained.

6. The method as described in claim 5, characterized in that, After obtaining the preset generative adversarial network, the method further includes: The individual feature data of the target face are obtained based on the real broadcast video of the target face; The individual feature data is used as the generation guiding condition for the generator; The individual feature data includes dental features and facial texture motion and distribution parameters during expression changes.

7. A device for generating digital human videos, characterized in that, The device includes: The acquisition module is used to acquire the target image, including the target face, and the target text to be broadcast; The positioning module is used to input the target image into a preset key point positioning model to perform facial key point recognition and obtain multiple facial key points in the target image; The parsing module is used to input the target text and each of the facial key points into a preset affine transformation offset model to obtain multiple sets of affine transformation offset parameters corresponding to each character in the target text. Each set of affine transformation offset parameters represents the affine transformation offset value generated by each of the facial key points when the target face broadcasts the character. A generation module is used to generate a digital human video in which the target face reads the target text based on the target image, facial key points in the target image, and each set of affine transformation offset parameters. The module includes: inputting the target image and facial key points in the target image into a generator in a preset generative adversarial network, and sequentially inputting a set of affine transformation offset parameters into the generator each time; generating the digital human video based on multiple frames of images sequentially output by the generator; wherein the basic input of the generator is the current frame image and facial key points in the current frame image, the conditional input of the generator is the affine transformation offset parameters of the next frame image, the output of the generator is the next frame image, and the first current frame image is the target image.

8. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the digital human video generation method according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating digital human videos according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image synthesis method and device, equipment and storage medium

    CN112785670A

  • Video generation method, electronic equipment, storage medium and digital human server

    CN114173188A