Real-time digital human video generation method and device, electronic equipment and storage medium

By using predefined facial motion templates and a material database, the problem of high computational costs in existing technologies is solved, enabling low-latency real-time digital human video generation and reducing hardware requirements and long-term operating costs.

CN121309905BActive Publication Date: 2026-04-14BEIJING HONGMIAN XIAOBING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HONGMIAN XIAOBING TECH CO LTD
Filing Date
2025-12-11
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing methods for generating digital human videos rely on massive amounts of training data and complex models, resulting in high computational costs and difficulty in meeting the low latency requirements of real-time interactive applications.

Method used

By using predefined facial motion templates and a database of materials, digital human videos are generated in real time through audio feature extraction and matching, reducing the computational burden of the real-time generation stage and completing computationally intensive tasks in the pre-production stage.

Benefits of technology

It enables real-time digital human video generation with low computational cost, reduces hardware requirements, improves resource utilization efficiency, and meets the low latency requirements of real-time interactive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121309905B_ABST
    Figure CN121309905B_ABST
Patent Text Reader

Abstract

The application provides a real-time digital human video generation method and device, electronic equipment and storage medium, and relates to the technical field of digital human video generation. The method can realize real-time and low-computing-cost digital human video generation by using a predefined face motion template and a material database. The computationally intensive tasks such as design and optimization of the face motion template, customization of the digital human template, and generation of the face motion material are completed in the pre-production stage, which can greatly reduce the computing burden of the real-time digital human video generation stage, simplify the process of the real-time generation stage, and accordingly reduce the performance requirements of the hardware device. The method does not need to rely on expensive computing resources such as high-end GPUs, thereby significantly reducing the overall computing cost and hardware investment of the real-time digital human video generation. Meanwhile, the face motion material obtained in the pre-production stage can be reused, which further improves the resource utilization efficiency and reduces the long-term operating cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital human video generation technology, and in particular to a real-time digital human video generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the continuous development of technologies such as artificial intelligence, digital human video generation technology has gradually become a popular research field and has been widely used in many industries such as film and television production, virtual anchors, and online education.

[0003] Currently, most common methods for generating digital human videos are based on deep learning models. These methods typically rely on large amounts of training data and complex model architectures to achieve precise control over the digital human's movements and expressions. Common deep learning models include Wav2Lip and SadTalker. Wav2Lip overlays synthesized lip movements onto existing video content and uses a discriminative mechanism to ensure lip-sync. SadTalker, on the other hand, generates 3D motion coefficients from audio and produces a close-up of the speaker's head (Talking Head).

[0004] Existing deep learning models such as Wav2Lip and SadTalker rely on massive amounts of training data to learn and generate realistic digital human patterns and features. The collection, labeling, and organization of training data are time-consuming and labor-intensive, requiring specialized data labeling teams and substantial storage resources. Furthermore, the complex structure of deep learning models typically demands powerful computing capabilities for training and inference, necessitating the use of multiple high-end graphics processing units (GPUs) and consuming significant amounts of electricity and computation time. Moreover, for applications requiring real-time interaction, such as real-time virtual customer service and interactive online education live streaming, the current deep learning models often struggle to meet low-latency requirements for digital human video generation, resulting in weak real-time interactive capabilities. Summary of the Invention

[0005] This invention provides a method, apparatus, electronic device, and storage medium for generating real-time digital human videos, in order to overcome the deficiencies existing in related technologies.

[0006] This invention provides a method for generating real-time digital human videos, comprising:

[0007] Obtain the target audio and digital human template;

[0008] For the current audio frame in the target audio, feature extraction and encoding are performed on the current audio frame to obtain the current audio feature, and the current audio feature is matched with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio feature; the facial motion template includes multiple facial motion patterns;

[0009] Based on the current facial motion pattern identifier, the current facial motion material is retrieved from the material database of the digital human template; the material database includes multiple pre-produced facial motion materials, and each facial motion material corresponds one-to-one with multiple facial motion patterns;

[0010] The current facial motion data is fused with the digital human template to generate a digital human video in real time.

[0011] According to a real-time digital human video generation method provided by the present invention, the step of determining the material database includes:

[0012] Face detection is performed on the digital human template to obtain a face image and parameter information of the face image in the digital human image of the digital human template;

[0013] Based on the facial motion template, facial driving is performed on the face image to generate the facial motion material corresponding to each facial motion mode;

[0014] Each facial motion data and its corresponding parameter information are stored in the data database.

[0015] According to a real-time digital human video generation method provided by the present invention, the step of performing facial driving on the face image based on the facial motion template to generate facial motion material corresponding to each facial motion mode includes:

[0016] The face image and the facial motion template are input into the facial motion driving model to obtain the facial motion materials output by the facial motion driving model; wherein, the facial motion driving model is based on a pre-trained diffusion model and is constructed using a control network framework and a local control module.

[0017] According to the present invention, a real-time digital human video generation method is provided, wherein the digital human template is a template video;

[0018] The step of performing face detection on the digital human template to obtain a face image includes:

[0019] Face detection is performed on each frame of the template video to obtain the face image in each frame.

[0020] According to the present invention, a real-time digital human video generation method is provided, wherein fusing the current facial motion data with the digital human template to generate a digital human video in real time includes:

[0021] Based on the parameter information corresponding to the current facial motion material, the current facial motion material is fused with the digital human template to generate the digital human video in real time.

[0022] According to a real-time digital human video generation method provided by the present invention, the step of matching the current audio features with a predefined facial motion template to obtain a current facial motion pattern identifier corresponding to the current audio features includes:

[0023] The current audio features are input into the feature mapping model to obtain the current facial motion pattern identifier output by the feature mapping model;

[0024] The feature mapping model is used to calculate the association between the current audio feature and the feature information corresponding to each facial motion pattern in the facial motion template, and to determine the current facial motion pattern identifier based on the association.

[0025] According to a real-time digital human video generation method provided by the present invention, the step of retrieving current facial motion material from the material database of the digital human template based on the current facial motion pattern identifier includes:

[0026] Based on the current facial motion pattern identifier, determine the access information of the current facial motion material;

[0027] Based on the access information, the current facial motion footage is retrieved from the media database.

[0028] The present invention also provides a real-time digital human video generation device, comprising:

[0029] The acquisition module is used to acquire the target audio and the digital human template;

[0030] The motion pattern determination module is used to perform feature extraction and encoding on the current audio frame in the target audio to obtain the current audio feature, and to match the current audio feature with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio feature; the facial motion template includes multiple facial motion patterns;

[0031] The motion material determination module is used to retrieve the current facial motion material from the material database of the digital human template based on the current facial motion pattern identifier; each facial motion material in the material database corresponds one-to-one with multiple facial motion patterns;

[0032] The video generation module is used to fuse the current facial motion footage with the digital human template to generate a digital human video in real time.

[0033] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the real-time digital human video generation method as described above.

[0034] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the real-time digital human video generation method as described above.

[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the real-time digital human video generation method as described above.

[0036] This invention provides a real-time digital human video generation method, apparatus, electronic device, and storage medium. This method utilizes predefined facial motion templates and a material database to achieve real-time, low-computational-cost digital human video generation. By concentrating computationally intensive tasks such as facial motion template design and optimization, digital human template customization, and facial motion material generation in the pre-production stage, the computational burden of the real-time digital human video generation stage can be significantly reduced, simplifying the real-time generation process and correspondingly lowering the performance requirements of hardware devices. It eliminates the need to rely on expensive computing resources such as high-end GPUs, thereby significantly reducing the overall computational cost and hardware investment for real-time digital human video generation. Simultaneously, the facial motion materials obtained in the pre-production stage can be reused, further improving resource utilization efficiency and reducing long-term operating costs. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is one of the flowcharts illustrating the real-time digital human video generation method provided by the present invention.

[0039] Figure 2 This is the second flowchart of the real-time digital human video generation method provided by the present invention.

[0040] Figure 3This is a schematic diagram of the material database determination method in the real-time digital human video generation method provided by the present invention.

[0041] Figure 4 This is a schematic diagram of the structure of the real-time digital human video generation device provided by the present invention.

[0042] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0044] Figure 1 This is a flowchart illustrating a real-time digital human video generation method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0045] S1, acquire the target audio and digital human template;

[0046] S2, for the current audio frame in the target audio, perform feature extraction and encoding processing on the current audio frame to obtain the current audio feature, and match the current audio feature with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio feature; the facial motion template includes multiple facial motion patterns;

[0047] S3, based on the current facial motion pattern identifier, retrieve the current facial motion material from the material database of the digital human template; the material database includes multiple pre-produced facial motion materials, and each facial motion material corresponds one-to-one with multiple facial motion patterns;

[0048] S4, the current facial motion material is fused with the digital human template to generate a digital human video in real time.

[0049] Specifically, the real-time digital human video generation method provided in this embodiment of the invention is executed by a real-time digital human video generation device. This device can be configured in an electronic device such as a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, tablet, etc., and no specific limitation is made here.

[0050] First, step S1 is executed to obtain the target audio and the digital human template. The target audio can be input by the user, and the digital human template can be selected by the user. The digital human video can be generated by driving the digital human template with the target audio. Here, the user can select the digital human image on the display interface to select the digital human template. Each digital human template corresponds to a digital human image, which is the digital human displayed in the digital human video.

[0051] Each digital human template defines the overall appearance of a digital human, including facial structure, half-body or full-body motion videos, etc., and is the basic framework for generating the digital human image in digital human videos.

[0052] Each digital human template can be an image or a frame from a video; no specific limitations are made here.

[0053] Then, step S2 is executed. The target audio may include multiple audio frames. The same operation is performed on each audio frame. That is, for the current audio frame in the target audio, feature extraction and encoding processing can be performed on the current audio frame through an audio coding model to obtain the current audio features. Here, the audio coding model can be a trained deep learning audio coding model, such as Wav2Vec.

[0054] During training, audio coding models can automatically learn important features in audio data, such as spectral features, melodic features, and rhythmic features, and convert them into discrete feature vectors. These discrete feature vectors can, to a certain extent, represent key information such as the semantics and emotion of the audio content.

[0055] Therefore, the current audio features can be in the form of discrete feature vectors, including the spectral features, melodic features, rhythmic features, and other features of the target audio, representing key information such as the semantics and emotion of the target audio.

[0056] In this embodiment of the invention, the number of values ​​for the current audio feature is limited, which is consistent with the number of facial motion patterns in the predefined facial motion template. This provides a concise and effective audio feature for subsequent facial motion template matching, and provides a basis for determining the corresponding current facial motion material, thereby achieving a precise association between audio and facial motion.

[0057] The current audio feature value can be one of a predetermined number of values, which can be set as needed, for example, 1024.

[0058] Subsequently, the current audio features are matched with predefined facial motion templates to quickly and accurately obtain the current facial motion pattern identifier corresponding to the current audio features. The facial motion template serves as the basic reference for the digital human's facial motion and includes multiple facial motion patterns used to drive subsequent facial movements. Each facial motion pattern in the facial motion template corresponds to a facial expression, lip shape, etc., and can be considered a sub-motion template within the facial motion template.

[0059] Each facial movement pattern has a corresponding facial movement pattern identifier (ID) to uniquely identify the facial movement pattern.

[0060] The current audio features are matched with predefined facial motion templates. Specifically, the current audio features are matched with each facial motion pattern in the facial motion template. The one that matches successfully from multiple facial motion patterns in the facial motion template is selected to obtain the current facial motion pattern identifier. This establishes a direct correlation between the target audio and the digital human's facial motion, providing a key intermediate result for enabling the digital human to perform corresponding facial expressions and lip movements based on the target audio.

[0061] Next, step S3 is executed, whereby the current facial motion pattern identifier is used to retrieve the current facial motion pattern from the digital human template's material database. The material database may include multiple pre-produced facial motion patterns, each consisting of a facial image corresponding to a specific facial motion pattern. These patterns are typically stored as image sequences or video clips. The corresponding current facial motion pattern can be identified using the current facial motion pattern identifier.

[0062] When extracting current facial motion data, decoding and format conversion are required to ensure that the obtained current facial motion data can be integrated with the digital human template.

[0063] Finally, step S4 is executed to fuse the current facial motion footage with the digital human template, obtaining the current frame in the digital human video, thus achieving real-time generation of the digital human video. The fusion of the current facial motion footage with the digital human template can be achieved using image synthesis techniques, such as alpha channel fusion, to naturally and accurately blend the current facial motion footage with the digital human template, ultimately generating a complete digital human video. This allows the user to expect the digital human to display real-time, natural facial expressions and lip movements based on the target audio, meeting the needs of various real-time interactive application scenarios.

[0064] During the fusion process, it is necessary to consider the static or dynamic display data of the digital human image of the digital human template as the basic background for fusion with the current facial motion material. It is also necessary to deal with the problem of continuous fusion of multiple facial motion materials with the digital human template in time sequence to ensure the continuity and smoothness of the generated digital human video in time sequence.

[0065] The digital human videos generated in real time can be displayed on terminal devices, such as computer screens and mobile phone screens, for users to watch and interact with.

[0066] like Figure 2 As shown, the target audio can be processed by an audio coding model to obtain discrete audio features, which include the audio features of each frame of the target audio.

[0067] By matching discrete audio features with predefined facial motion templates, a set of facial motion pattern identifiers can be obtained, which includes facial motion pattern identifiers corresponding to each frame of audio.

[0068] Users select a digital avatar, determine the digital avatar template and material database, and retrieve the corresponding set of facial motion materials from the material database through the set of facial motion pattern identifiers. The set of facial motion materials includes the facial motion materials corresponding to each facial motion pattern identifier in the set of facial motion pattern identifiers.

[0069] By fusing each facial motion data element from the facial motion data set with its corresponding digital human template, a digital human video can be generated in real time. Specifically, when the digital human template is a video, it is necessary to fuse each facial motion data element from the facial motion data set with its corresponding frame from the video of the digital human template.

[0070] The real-time digital human video generation method provided in this embodiment of the invention first acquires the target audio and a digital human template; then, for the current audio frame in the target audio, feature extraction and encoding are performed on the current audio frame to obtain the current audio features, and the current audio features are matched with a predefined facial motion template to obtain the current facial motion mode identifier corresponding to the current audio features; subsequently, using the current facial motion mode identifier, the current facial motion material is retrieved from the digital human template's material database; finally, the current facial motion material is fused with the digital human template to generate a digital human video in real time. This method utilizes a predefined facial motion template and a material database to achieve real-time, low-computational-cost digital human video generation. By concentrating computationally intensive tasks such as the design and optimization of facial motion templates, the customization of digital human templates, and the generation of facial motion materials in the pre-production stage, the computational burden of the real-time digital human video generation stage can be greatly reduced, simplifying the real-time generation process and correspondingly reducing the performance requirements of hardware devices. It eliminates the need to rely on expensive computing resources such as high-end GPUs, thereby significantly reducing the overall computational cost and hardware investment for real-time digital human video generation. Simultaneously, the facial motion materials obtained in the pre-production stage can be reused, further improving resource utilization efficiency and reducing long-term operating costs.

[0071] Furthermore, this method leverages efficient audio feature processing technology, enabling the audio coding model to quickly extract current audio features and then rapidly match them with the corresponding current facial motion pattern identifier. This efficient processing flow significantly improves the speed of digital human generation, meeting the stringent low-latency requirements of real-time interactive applications, such as real-time virtual customer service and online education live interactive scenarios, ensuring users receive a smooth and natural interactive experience.

[0072] By matching current audio features with predefined facial motion templates, efficient association between audio and facial motion can be achieved, avoiding the complexity of processing large amounts of audio data in real time. Current facial motion material is extracted from the material database based on the current facial motion pattern identifier and fused with the digital human template. This fusion method is simple and efficient, significantly improving the generation speed of digital human videos while ensuring generation quality.

[0073] Based on the above embodiments, the steps for determining the material database include:

[0074] Face detection is performed on the digital human template to obtain a face image and parameter information of the face image in the digital human image of the digital human template;

[0075] Based on the facial motion template, facial driving is performed on the face image to generate the facial motion material corresponding to each facial motion mode;

[0076] Each facial motion data and its corresponding parameter information are stored in the data database.

[0077] Specifically, when building a material database, such as Figure 3 As shown, a deep learning-based face detection algorithm can be used to detect faces in the digital human template, accurately locate the face position, and extract the face image. Here, the parameter information of the face image within the digital human image of the template can include the face image's position, size, angle, etc. The face detection algorithm can be RetinaFace, etc.

[0078] Subsequently, facial motion templates can be used to drive the face image, causing the face image to perform corresponding facial expressions and mouth movements according to the various facial motion patterns in the facial motion template, thereby generating facial motion material corresponding to each facial motion pattern.

[0079] Finally, the generated facial motion data and corresponding parameter information are stored in the data database to provide rich data resources for the real-time generation stage.

[0080] Based on the above embodiments, the step of performing facial driving on the face image based on the facial motion template to generate the facial motion material corresponding to each facial motion mode includes:

[0081] The face image and the facial motion template are input into the facial motion driving model to obtain the facial motion materials output by the facial motion driving model; wherein, the facial motion driving model is based on a pre-trained diffusion model and is constructed using a control network framework.

[0082] Specifically, when performing facial motion driving, a facial motion driving model can be introduced. This model can be a conditional diffusion model, such as X-Portrait, built using a pre-trained diffusion model and a ControlNet framework with local control modules. Through implicit cross-identity motion control via the ControlNet, dynamic information is directly interpreted from the driving video, and local control modules focus on local movements such as the eyes and mouth. Random scaling is also used to enhance training. In terms of performance, it can accurately capture and transfer subtle and extreme facial expressions, maintain identity features, generate high-quality animations, and outperform existing technologies in multiple benchmark tests, demonstrating superior perceptual quality, motion accuracy, and identity similarity.

[0083] A face image and a facial motion template are input into a facial motion-driven model, which then generates various facial motion data. Here, each facial motion data point can be a frame from the video output by the facial motion-driven model.

[0084] In this embodiment of the invention, facial motion data can be determined efficiently and accurately through a facial motion-driven model.

[0085] Based on the above embodiments, the digital human template is a template video;

[0086] The step of performing face detection on the digital human template to obtain a face image includes:

[0087] Face detection is performed on each frame of the template video to obtain the face image in each frame.

[0088] Specifically, the digital human template can be a template video. Then, when performing face detection on the digital human template, face detection can be performed on each frame of the template video to obtain the face image in each frame and the parameter information of the face image in the digital human image of each frame.

[0089] In this embodiment of the invention, template videos can be used to provide the overall movements of the digital human image in digital human videos, making the overall movements of the digital human image richer.

[0090] Based on the above embodiments, the step of fusing the current facial motion data with the digital human template to generate a digital human video in real time includes:

[0091] Based on the parameter information corresponding to the current facial motion material, the current facial motion material is fused with the digital human template to generate the digital human video in real time.

[0092] Specifically, when merging the current facial motion data with the digital human template, it is necessary to consider the position, size, angle and other parameters of the current facial motion data in the digital human image of the digital human template, as well as the consistency of the lighting and shadow effects, so as to ensure that the merged digital human video is visually natural, realistic and without obvious flaws.

[0093] Based on the above embodiments, the step of matching the current audio feature with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio feature includes:

[0094] The current audio features are input into the feature mapping model to obtain the current facial motion pattern identifier output by the feature mapping model;

[0095] The feature mapping model is used to calculate the association between the current audio feature and the feature information corresponding to each facial motion pattern in the facial motion template, and to determine the current facial motion pattern identifier based on the association.

[0096] Specifically, in this embodiment of the invention, the current audio features can be matched with a predefined facial motion template through a feature mapping model, and the current audio features can be input into the feature mapping model.

[0097] Each facial motion pattern in the facial motion template corresponds to feature information, which may include feature labels, feature vectors, etc. The feature mapping model can calculate the association between the current audio features and the feature information corresponding to each facial motion pattern in the facial motion template. This association can be represented by similarity; the higher the similarity, the stronger the association.

[0098] By establishing relationships, the facial motion pattern with the highest similarity can be selected from the facial motion templates and used as the current facial motion pattern, which is then identified as the current facial motion pattern identifier.

[0099] In this embodiment of the invention, the accuracy and reliability of the identified current facial motion pattern can be improved by using a feature mapping model.

[0100] Based on the above embodiments, the step of retrieving current facial motion data from the digital human template's data database based on the current facial motion pattern identifier includes:

[0101] Based on the current facial motion pattern identifier, determine the access information of the current facial motion material;

[0102] Based on the access information, the current facial motion footage is retrieved from the media database.

[0103] Specifically, when retrieving current facial motion data from the digital human template's data database, the access information for the current facial motion data can be determined first using the current facial motion mode identifier. This access information may include storage location, access method, and access permissions.

[0104] Subsequently, by utilizing the access information, the current facial motion data can be quickly and accurately retrieved from the data database, ensuring the efficiency and consistency of the digital human generation process.

[0105] like Figure 4 As shown, based on the above embodiments, this embodiment of the invention provides a real-time digital human video generation device, comprising:

[0106] Module 41 is used to acquire the target audio and the digital human template;

[0107] The motion pattern determination module 42 is used to perform feature extraction and encoding processing on the current audio frame in the target audio to obtain the current audio feature, and to match the current audio feature with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio feature; the facial motion template includes multiple facial motion patterns;

[0108] The motion material determination module 43 is used to retrieve the current facial motion material from the material database of the digital human template based on the current facial motion mode identifier; each facial motion material in the material database corresponds one-to-one with multiple facial motion modes;

[0109] The video generation module 44 is used to fuse the current facial motion material with the digital human template to generate a digital human video in real time.

[0110] Based on the above embodiments, the real-time digital human video generation device provided in this embodiment of the invention includes a material database determination module, used for:

[0111] Face detection is performed on the digital human template to obtain a face image and parameter information of the face image in the digital human image of the digital human template;

[0112] Based on the facial motion template, facial driving is performed on the face image to generate the facial motion material corresponding to each facial motion mode;

[0113] Each facial motion data and its corresponding parameter information are stored in the data database.

[0114] Based on the above embodiments, the real-time digital human video generation apparatus provided in this embodiment of the invention, wherein the material database determination module is specifically used for:

[0115] The face image and the facial motion template are input into the facial motion driving model to obtain the facial motion materials output by the facial motion driving model; wherein, the facial motion driving model is based on a pre-trained diffusion model and is constructed using a control network framework and a local control module.

[0116] Based on the above embodiments, the real-time digital human video generation device provided in this embodiment of the invention uses a template video as the digital human template.

[0117] The material database determination module is specifically used for:

[0118] Face detection is performed on each frame of the template video to obtain the face image in each frame.

[0119] Based on the above embodiments, the real-time digital human video generation device provided in this embodiment of the invention, wherein the video generation module is specifically used for:

[0120] Based on the parameter information corresponding to the current facial motion material, the current facial motion material is fused with the digital human template to generate the digital human video in real time.

[0121] Based on the above embodiments, the motion mode determination module in the real-time digital human video generation apparatus provided in this embodiment of the invention is specifically used for:

[0122] The current audio features are input into the feature mapping model to obtain the current facial motion pattern identifier output by the feature mapping model;

[0123] The feature mapping model is used to calculate the association between the current audio feature and the feature information corresponding to each facial motion pattern in the facial motion template, and to determine the current facial motion pattern identifier based on the association.

[0124] Based on the above embodiments, the real-time digital human video generation apparatus provided in this embodiment of the invention, wherein the motion material determination module is specifically used for:

[0125] Based on the current facial motion pattern identifier, determine the access information of the current facial motion material;

[0126] Based on the access information, the current facial motion footage is retrieved from the media database.

[0127] Specifically, the functions of each module in the real-time digital human video generation device provided in this embodiment correspond one-to-one with the operation flow of each step in the above method embodiment, and the achieved effect is also the same. Please refer to the above embodiments for details, and this will not be repeated in this embodiment.

[0128] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the real-time digital human video generation method provided in the above embodiments.

[0129] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0130] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the real-time digital human video generation method provided in the above embodiments.

[0131] In another aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the real-time digital human video generation method provided in the above embodiments. This computer-readable storage medium can be either a non-transitory computer-readable storage medium or a transient computer-readable storage medium, and no specific limitation is made herein.

[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating real-time digital human videos, characterized in that, include: Obtain the target audio and digital human template; For the current audio frame in the target audio, feature extraction and encoding are performed on the current audio frame to obtain the current audio feature, and the current audio feature is matched with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio feature; the facial motion template includes multiple facial motion patterns; Based on the current facial motion pattern identifier, the current facial motion material is retrieved from the material database of the digital human template; the material database includes multiple pre-produced facial motion materials, and each facial motion material corresponds one-to-one with multiple facial motion patterns; The current facial motion data is fused with the digital human template to generate a digital human video in real time; The step of matching the current audio features with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio features includes: The current audio features are input into the feature mapping model to obtain the current facial motion pattern identifier output by the feature mapping model; The feature mapping model is used to calculate the association between the current audio feature and the feature information corresponding to each facial motion pattern in the facial motion template, and to determine the current facial motion pattern identifier based on the association.

2. The real-time digital human video generation method according to claim 1, characterized in that, The steps for determining the material database include: Face detection is performed on the digital human template to obtain a face image and parameter information of the face image in the digital human image of the digital human template; Based on the facial motion template, facial driving is performed on the face image to generate the facial motion material corresponding to each facial motion mode; Each facial motion data and its corresponding parameter information are stored in the data database.

3. The real-time digital human video generation method according to claim 2, characterized in that, The step of performing facial driving on the face image based on the facial motion template to generate facial motion material corresponding to each facial motion mode includes: The face image and the facial motion template are input into the facial motion driving model to obtain the facial motion materials output by the facial motion driving model; wherein, the facial motion driving model is based on a pre-trained diffusion model and is constructed using a control network framework and a local control module.

4. The real-time digital human video generation method according to claim 2, characterized in that, The digital human template is a template video; The step of performing face detection on the digital human template to obtain a face image includes: Face detection is performed on each frame of the template video to obtain the face image in each frame.

5. The real-time digital human video generation method according to claim 2, characterized in that, The step of fusing the current facial motion data with the digital human template to generate a digital human video in real time includes: Based on the parameter information corresponding to the current facial motion material, the current facial motion material is fused with the digital human template to generate the digital human video in real time.

6. The real-time digital human video generation method according to any one of claims 1-5, characterized in that, The step of retrieving current facial motion data from the digital human template's data database based on the current facial motion pattern identifier includes: Based on the current facial motion pattern identifier, determine the access information of the current facial motion material; Based on the access information, the current facial motion footage is retrieved from the media database.

7. A real-time digital human video generation device, characterized in that, include: The acquisition module is used to acquire the target audio and the digital human template; The motion pattern determination module is used to perform feature extraction and encoding on the current audio frame in the target audio to obtain the current audio feature, and to match the current audio feature with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio feature; the facial motion template includes multiple facial motion patterns; The motion material determination module is used to retrieve the current facial motion material from the material database of the digital human template based on the current facial motion pattern identifier; each facial motion material in the material database corresponds one-to-one with multiple facial motion patterns; The video generation module is used to fuse the current facial motion footage with the digital human template to generate a digital human video in real time. The step of matching the current audio features with a predefined facial motion template to obtain the current facial motion pattern identifier corresponding to the current audio features includes: The current audio features are input into the feature mapping model to obtain the current facial motion pattern identifier output by the feature mapping model; The feature mapping model is used to calculate the association between the current audio feature and the feature information corresponding to each facial motion pattern in the facial motion template, and to determine the current facial motion pattern identifier based on the association.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the real-time digital human video generation method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the real-time digital human video generation method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Virtual human video synthesis method based on voice driving and face self-driving

    CN116528019A

  • Voice-driven expression generation method and device, equipment and storage medium

    CN118071901A