Method and apparatus for generating a digital human

CN122530397APending Publication Date: 2026-08-07HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-02-06
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]由于目前方案中生成的数字人是基于大量的真人视频素材生成的数字人,用户需要进行繁琐的视频拍摄和处理,同时还要进行复杂的数字人制作,导致数字人的生成以及部署效率低,定制成本高

Benefits of technology

[0038]可以理解,上述提供的任意一种基于数字人的生成装置、计算设备、计算设备集群、计算机可读介质或计算机程序产品等所能达到的有益效果可参考对应的方法中的有益效果,此处不再赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530397A_ABST
    Figure CN122530397A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a digital human generation method and device, which are used to improve the generation efficiency of digital humans. The method of the embodiments of the present application comprises: a computing device acquires a digital human video, the digital human video is used as training data for generating a digital human model, and the data source of the digital human video includes one or more of the following: a film and television work, a game animation video, a digital human application video, and a social media video. The computing device trains the digital human model based on the training data, the digital human model includes a plurality of feature extraction networks, the plurality of feature extraction networks are used to extract different types of digital human features in the digital human video, and the digital human features include one or more of the following: facial expression features, body movement features, and scene features. The computing device performs a multi-modal interaction task based on the trained digital human model, and the multi-modal interaction task includes generating a corresponding digital human video based on multi-modal content input by a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more particularly to a method and apparatus for generating a digital human. Background Technology

[0002] With the development of digital modeling technology, the application of digital human virtual avatars is becoming increasingly widespread. For example, digital humans can serve as virtual idols or virtual anchors, providing immersive audiovisual experiences in entertainment content such as concerts, movies, and TV series, and can also be used in scenarios such as news broadcasting and weather forecasting. Furthermore, digital humans can act as intelligent customer service representatives, providing 24 / 7 online service and simulating the communication methods of real humans.

[0003] Current digital human generation solutions often require users to film numerous videos of real people, perform post-processing on these videos, and rely on technologies such as motion capture and facial expression capture to map the real people's movements and expressions onto the digital human. This is necessary to generate a digital human image with relatively natural and realistic movements and expressions. For example, users film numerous videos of real people speaking, and then use facial expression capture technology to map the changes in lip movements onto the digital human, thus obtaining a virtual digital human image with lip movements.

[0004] Because current solutions generate digital humans based on large amounts of real-life video footage, users need to perform tedious video shooting and processing, as well as complex digital human creation, resulting in low efficiency in digital human generation and deployment, and high customization costs. Furthermore, digital humans generated using motion capture and facial expression capture technologies often can only perform fixed tasks driven by voice, leading to poor interactivity between the generated digital humans and users. Summary of the Invention

[0005] This application provides a method for generating digital humans, which improves the efficiency of digital human generation. This application also provides a digital human generation apparatus, computing device, computing device cluster, computer-readable storage medium, and computer program product corresponding to this method.

[0006] In a first aspect, embodiments of this application provide a method for generating a digital human. This method can be executed by a computing device, or by a component of the computing device, such as a processor, chip, or chip system, or by a logic module or software capable of implementing all or part of the functions of the computing device. The method provided in the first aspect includes: the computing device acquiring a digital human video, which is used as training data to generate a digital human model. The data sources for the digital human video include one or more of the following: film and television works, game animation videos, digital human application videos, and social media videos. The computing device trains a digital human model based on the training data. The digital human model includes multiple feature extraction networks used to extract different types of digital human features from the digital human video. The digital human features include one or more of the following: facial expression features, body movement features, and scene features. The computing device performs a multimodal interaction task based on the trained digital human model. The multimodal interaction task includes generating a corresponding digital human video based on multimodal content input by the user.

[0007] In this embodiment, the computing device can train a digital human model based on various existing types of digital human videos and generate digital human videos based on the trained digital human model. Compared to the existing technology that uses user-captured live videos and relies on motion capture and facial expression capture technologies to create digital human videos, this embodiment allows the computing device to directly generate digital human videos based on the trained digital human model, avoiding complex video processing operations and thus improving the efficiency of digital human generation. Furthermore, in this embodiment, the digital human model can generate corresponding digital human videos according to multimodal content input by the user, enhancing the interactive capabilities of the digital human.

[0008] In one possible implementation, during the process of acquiring digital human videos, the computing device can use movie and TV series clips obtained in cooperation with film and television production companies as digital human videos, virtual character animations obtained from game companies as digital human videos, interactive videos of virtual anchors and customer service on online digital human platforms as digital human videos, or relevant video content uploaded by social media users as digital human videos.

[0009] In this embodiment, the computing device can obtain training data of the digital human model from multiple sources, thereby reducing the dependence on user-shot real-life videos for generating digital human videos, reducing the difficulty of obtaining training data, improving the efficiency of obtaining training data for the digital human model, and further improving the quality of digital human videos generated by the digital human model.

[0010] In one possible implementation, before training the digital human model based on the training data, the computing device annotates the digital human features in the digital human video based on the annotation model to obtain training data. The training data includes the digital human video and its annotated data. The annotation model includes models corresponding to different types of annotated data, including text annotation data, speech annotation data, and...

[0011] In this embodiment, the computing device annotates the digital human features in the digital human video based on the annotation model, and can generate different types of annotated data based on different annotation models, thereby improving the feasibility of the digital human model.

[0012] In one possible implementation, during the process of the computing device annotating the digital human features in the digital human video based on the annotation model, the computing device preprocesses the digital human video to obtain digital human video frames, extracts features from the digital human video frames based on the annotation model, and obtains digital human features, thereby identifying information such as people, actions and scenes in the digital human video frames. The annotation model includes multiple feature extraction networks, which are used to extract different types of digital human features. The feature extraction networks include pre-trained convolutional neural networks (CNNs).

[0013] In this embodiment, the computing device extracts features from digital human video frames based on a labeling model and labels different digital human features in the digital human video, thereby improving the feasibility of labeling different types of digital human features in the digital human video.

[0014] In one possible implementation, during the process of annotating digital human features in a digital human video based on an annotation model, the computing device analyzes the text descriptions corresponding to different digital human features in the video based on a large language model. These text descriptions are used to indicate the annotation data of the digital human video. The computing device then converts the text descriptions based on a speech model to obtain speech annotation data, and finally converts the text descriptions based on an image model to obtain image annotation data.

[0015] In this embodiment, the computing device can annotate different digital human features in digital human videos based on a large language model to obtain text descriptions corresponding to the digital human videos, thereby improving the annotation efficiency of digital human videos. At the same time, the computing device can also convert the annotation data through speech models and image models to obtain speech annotation data and image annotation data, thereby improving the richness of the annotation data.

[0016] In one possible implementation, during the annotation process of digital human videos, the computing device can directly process the digital human videos to obtain speech annotation data or image annotation data without converting the text description generated by the large language model to obtain speech annotation data and image annotation data.

[0017] In this embodiment of the application, the computing device can also directly process the digital human model through the speech model and the image model to obtain speech annotation data and image annotation data, thereby improving the efficiency of generating speech annotation data and image annotation data.

[0018] In one possible implementation, during the training of the digital human model based on training data, the computing device encodes the training data using an encoder to obtain a labeled sequence. This labeled sequence is used to indicate the digital human features in the digital human video. Subsequently, the computing device trains the digital human model based on the labeled sequence and a self-attention mechanism, which generates new labeled sequences according to the relationships between the labeled sequences.

[0019] In this embodiment of the application, the computing device can use an encoder to convert digital human videos into labeled sequences that the model can process, and use a self-attention mechanism to train the digital human model, thereby improving the quality of digital human videos generated by the digital human model.

[0020] In one possible implementation, during the process of encoding training data based on the encoder, the computing device extracts features from the digital human video based on the feature extraction network to obtain digital human features. The encoder includes multiple feature extraction networks, and different types of digital human features correspond to different feature extraction networks. The computing device converts the digital human features into a labeled sequence based on the serialization module, and there is a mapping relationship between the labels in the labeled sequence and the digital human features.

[0021] In the embodiments of this application, the computing device can utilize different feature extraction networks to process different types of digital human features during the encoding of training data based on the encoder, thereby improving the accuracy and efficiency of feature extraction during the training of the digital human model. At the same time, multiple feature extraction networks also enhance the generalization ability of the digital human model.

[0022] In one possible implementation, the computing device performs forward propagation on the labeled sequence based on a transformation model. During forward propagation, the transformation model processes the input data layer by layer and generates outputs. The computing device calculates the loss value obtained from the loss function based on the outputs and labeled data, and then performs backpropagation. During backpropagation, the computing device calculates gradients layer by layer and passes them to the previous layer until it reaches the input layer. These gradients are used to guide the direction of parameter updates in the model. The computing device adjusts the parameters of the transformation model according to the gradients to minimize the loss value. The computing device iteratively trains the digital human model by repeating the process of forward propagation, loss calculation, backpropagation, and parameter updates to obtain the trained model.

[0023] In the embodiments of this application, the computing device can generate a trained digital human model based on the process of forward propagation, loss calculation, backpropagation and parameter update of the labeled sequence using a transformation model, thereby improving the feasibility of the computing device generating a digital human model.

[0024] In one possible implementation, during the process of training the digital human model based on training data, the computing device can also evaluate and optimize the trained digital human model. For example, the computing device can evaluate the performance of the digital human model using a validation dataset and perform hyperparameter tuning on the trained digital human model. After the computing device completes the training of the digital human model, it can deploy the trained digital human model to real-world application scenarios to perform multimodal interactions of the digital human model.

[0025] After the computing device completes the training of the digital human model in this embodiment, it can also evaluate and deploy the trained digital human model, thereby improving the applicability of the digital human model to actual application scenarios.

[0026] In one possible implementation, during the process of a computing device performing a multimodal interaction task based on a trained digital human model, the computing device receives multimodal content input by the user. This multimodal content serves to instruct the digital human model on its input commands and includes one or more of the following: text, speech, and images. The computing device then generates a digital human interactive avatar corresponding to the multimodal content based on the digital human model. This digital human interactive avatar includes digital human facial expressions, digital human actions, and digital human speech corresponding to the multimodal content.

[0027] The digital human model provided in this application embodiment can interact with users and generate corresponding digital human videos based on the multimodal content input by the user, thereby improving the control and interaction capabilities of the digital human model in generating digital human videos.

[0028] Secondly, embodiments of this application provide a digital human generation apparatus, comprising an acquisition unit and a processing unit. The acquisition unit acquires digital human videos, which are used as training data for generating a digital human model. The data sources for the digital human videos include one or more of the following: film and television works, game animation videos, digital human application videos, and social media videos. The processing unit trains a digital human model based on the training data. The digital human model includes multiple feature extraction networks, which extract different types of digital human features from the digital human videos. These digital human features include one or more of the following: facial expression features, body movement features, and scene features. The processing unit also performs a multimodal interaction task based on the trained digital human model. The multimodal interaction task includes generating corresponding digital human videos based on multimodal content input by the user.

[0029] In one possible implementation, the processing unit is further configured to annotate the digital human features in the digital human video based on the annotation model to obtain training data. The annotation model includes models corresponding to different types of annotation data, including text annotation data, speech annotation data and image annotation data.

[0030] In one possible implementation, the processing unit is specifically used to analyze the text descriptions corresponding to different digital human features in the digital human video based on a large language model. The text descriptions are used to indicate the annotation data of the digital human video. The text descriptions are converted based on a speech model to obtain speech annotation data, and the text descriptions are converted based on an image model to obtain image annotation data.

[0031] In one possible implementation, the processing unit is specifically used to encode the training data based on the encoder to obtain a label sequence, the label sequence being used to indicate the digital human features in the digital human video, and to train a digital human model based on the label sequence and a self-attention mechanism, the self-attention mechanism being used to generate new label sequences according to the correlation between the label sequences.

[0032] In one possible implementation, the processing unit is specifically used to extract features from the digital human video based on the feature extraction network to obtain digital human features. Different types of digital human features correspond to different feature extraction networks. The digital human features are converted into a labeled sequence based on the serialization module. There is a mapping relationship between the labels in the labeled sequence and the digital human features.

[0033] In one possible implementation, the processing unit is specifically used to receive multimodal content input by the user. This multimodal content serves as an instruction for the digital human model, and includes one or more of the following: text, speech, and images. Based on the digital human model, a digital human interactive avatar corresponding to the multimodal content is generated. This digital human interactive avatar includes digital human facial expressions, digital human actions, and digital human speech corresponding to the multimodal content.

[0034] Thirdly, embodiments of this application provide a computing device including a processor coupled to a memory. The processor stores instructions, which, when executed by the processor, cause the computing device to perform the method described in the first aspect or any possible implementation thereof.

[0035] Fourthly, embodiments of this application provide a computing device cluster, which includes one or more computing devices. Each computing device includes a processor coupled to a memory. The processor is used to store instructions, which, when executed by the processor, cause the computing device cluster to perform the method described in the first aspect or any possible implementation thereof.

[0036] Fifthly, embodiments of this application provide a computer-readable storage medium having instructions stored thereon, which, when executed, cause a computer to perform the method described in the first aspect or any possible implementation thereof.

[0037] Sixthly, embodiments of this application provide a computer program product including instructions that, when executed, cause a computer to implement the method described in the first aspect or any possible implementation thereof.

[0038] It is understood that the beneficial effects that any of the above-mentioned digital human-based generation devices, computing equipment, computing equipment clusters, computer-readable media, or computer program products can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description

[0039] Figure 1 A schematic diagram of the system architecture of a digital human generation system provided in this application embodiment;

[0040] Figure 2 A flowchart illustrating a method for generating a digital human, as provided in an embodiment of this application;

[0041] Figure 3 A schematic diagram illustrating training data generated from digital human videos, provided as an embodiment of this application;

[0042] Figure 4 A schematic diagram illustrating the training of a digital human model as provided in an embodiment of this application;

[0043] Figure 5 A schematic diagram illustrating an application of a digital human model provided in an embodiment of this application;

[0044] Figure 6 A schematic diagram of a digital human generation device provided in an embodiment of this application;

[0045] Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0046] Figure 8 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0047] Figure 9 This is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation

[0048] This application provides a method and apparatus for generating digital humans, which improves the efficiency of digital human generation.

[0049] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0050] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0051] First, some of the terms used in the embodiments of this application are introduced to facilitate understanding of the technical solutions by those skilled in the art.

[0052] Large language models (LLMs) are artificial intelligence models with a large number of parameters and training data, capable of understanding and generating natural language. LLMs typically employ a transformer architecture, a self-attention mechanism that efficiently processes sequential data.

[0053] The transformer model is a deep learning model based on the self-attention mechanism. The transformer model consists of two parts: an encoder and a decoder. The encoder is responsible for mapping the input sequence into a series of vectors, and the decoder predicts the next output based on the encoder's output and some known outputs.

[0054] Self-attention is the core of the Transformer model. It calculates the correlations between elements within the input sequence, captures complex dependencies, and generates a weighted representation. Self-attention works by calculating three matrices: query, key, and value. These matrices are then used to calculate attention scores and weights, resulting in a weighted output.

[0055] A token is the smallest unit into which data is divided before or during model processing. Tokens are usually represented by vectors and can also be called word segments or words.

[0056] To make the technical solution of this application clearer and easier to understand, the system architecture of this application will be described below with reference to the accompanying drawings.

[0057] Please see Figure 1 , Figure 1 This application provides a schematic diagram of the system architecture for a data asset system. Figure 1 In the example shown, the digital human generation system 10 includes a data preprocessing module 101, a model building and training module 102, a model management module 103, a model deployment module 104, and a model application module 105. The specific functions of each part of the digital human generation system 10 are described below.

[0058] The data preprocessing module 101 is used to preprocess the dataset required for training the digital human model. The dataset can be digital human video data pre-collected by the user according to the actual application scenario; for example, the dataset can be real-life videos shot by the user, or it can be existing datasets obtained by the user, such as videos of movies, games, or animations. The dataset to be processed by the data preprocessing module 101 can be stored in the object storage service (OBS) of the cloud platform, and the data preprocessing module 101 can obtain the dataset from the object storage service.

[0059] The data preprocessing module 101 is also used to annotate the digital human videos in the dataset, and the annotated data carries labels. When the annotated data is used as input data to train the digital human model, the model building and training module 102 can adjust the parameters in the digital human model according to the data labels. It should be noted that traditional data annotation is usually done manually. Since the amount of data to be annotated is usually huge, it requires a lot of human resources. In this embodiment, the data preprocessing module 101 can also automatically annotate the dataset based on the trained annotation model.

[0060] The model building and training module 102 is used to build and train the digital human model. When building the digital human model, the module 102 can start with an initial transformer model selected on the AI ​​infrastructure development platform and then train the initial transformer model to obtain a digital human model that meets the user's goals. When training the digital human model, the module 102 can convert the dataset into a token sequence based on an encoder and use the transformer model to learn the dependencies between tokens. After training on a large-scale dataset, the transformer model will gradually develop the ability to generate digital human videos, thus obtaining the trained digital human video.

[0061] It should be noted that, to improve the training efficiency of the digital human model, the model building and training module 120 can employ distributed parallel training. Distributed parallel training of the digital human model includes data parallelism and model parallelism. Data parallelism refers to deploying the same digital human model to be trained on multiple nodes, and dividing the training dataset into multiple subsets, distributing them across the nodes. Model parallelism refers to splitting a digital human model into multiple sub-model parts, with different sub-model parts deployed on different nodes.

[0062] The model management module 103 is used to manage the trained digital human model in a unified manner, including evaluating, diagnosing, optimizing, and converting the digital human model. Among them, evaluating the digital human model means using at least one evaluation metric to measure the performance of the trained digital human model. For example, the accuracy of the inference results of the trained digital human model on the evaluation dataset can be calculated.

[0063] The model deployment module 104 is used to deploy the trained digital human model. The trained digital human model can be deployed on nodes in a cloud environment or on nodes in an edge environment. Nodes in the cloud environment can be virtual machine instances, container instances, physical servers, etc. When the digital human model is large in scale, the model deployment module 104 can distribute the digital human model across multiple nodes, or it can deploy the digital human model independently on multiple nodes to support a large volume of online service access. Nodes in the edge environment can be various edge devices.

[0064] The model application module 105 is used to provide digital human applications based on the deployed digital human model. The digital human applications can perform multimodal interaction tasks based on the digital human model. For example, the digital human applications generate corresponding digital human videos based on the multimodal content input by the user. The multimodal content input by the user includes text, voice, images and videos.

[0065] It should be noted that the deployed digital human model can become a digital human application or part of other applications; there are no specific limitations. For example, users can access the digital human application online through a webpage or a client app. When the digital human application is used, it can invoke the digital human model deployed in the edge or cloud environment to provide a response through online calls; there are no specific limitations.

[0066] Each module in the digital human generation system 10 in this application embodiment can be deployed on a computing device or a cluster of computing devices. Therefore, in this application embodiment, the digital human generation system 10 or each module in the digital human generation system 10 can also be referred to as a computing device or a cluster of computing devices.

[0067] based on Figure 1 The digital human generation system 10 shown in this application also provides a method for generating digital humans. The method for generating digital humans provided in this application will be described below with reference to embodiments.

[0068] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for generating a digital human, as provided in an embodiment of this application. Figure 2 In the example shown, the method includes the following steps:

[0069] 201. Obtain digital human videos. Digital human videos are used as training data to generate digital human models. The data sources for digital human videos include one or more of the following: film and television works, game animation videos, digital human application videos, and social media videos.

[0070] To train the digital human model in this embodiment, the digital human generation system 10 needs to acquire digital human videos as training data. These videos are used to generate the training data for the digital human model. The data sources for the digital human videos include one or more of the following: film and television works, game animation videos, digital human application videos, and social media videos. It is understood that the digital human generation system 10 needs to acquire a large number of digital human videos to ensure that the trained digital human model can generate digital humans in various scenarios.

[0071] In this embodiment of the application, the digital human generation system 10 can directly acquire digital human videos or indirectly acquire them from third parties during the acquisition process, without any specific limitation. Direct acquisition of digital human videos includes, for example, shooting a large amount of real-life video footage. Indirect acquisition from third parties includes, for example, collaborating with film and television production companies to acquire video clips of digital human characters from movies and TV series. Another example is acquiring animated videos of virtual characters from game companies. Yet another example is collecting daily interactive videos from existing digital human platforms, such as videos of virtual anchors and virtual customer service representatives. Finally, another example is collecting user-uploaded digital human-related video content from social media platforms.

[0072] In this embodiment, digital human videos are obtained from multiple sources, reducing reliance on user-shot videos and avoiding the processing of user-shot videos. Once the digital human model is trained, the digital human image can be directly generated using the large digital human model, avoiding the complex digital human synthesis process and directly simplifying the difficulty of digital human creation and customization.

[0073] In one possible implementation, after the digital human generation system 10 acquires a digital human video, it annotates the digital human features in the video based on an annotation model to obtain training data. The training data includes the digital human video and the annotated data of the video. The annotated data can also be called labels. The annotation model includes models corresponding to different types of annotated data. That is, different types of annotated data can be annotated by different annotation models. Different types of annotated data include text annotated data, speech annotated data and image annotated data. The annotation models include, for example, large language models, speech models and graphics models.

[0074] Since the digital human video annotation process in this embodiment involves the annotation of different types of data, the annotation in this embodiment can also be called multimodal annotation. Before multimodal annotation, the digital human generation system 10 first preprocesses the digital human video, so that it can extract digital human features that can be annotated based on the preprocessed digital human video. Specifically, the digital human generation system 10 segments the digital human video into video frames, and the annotation model extracts digital human features based on the segmented video frames, thereby identifying information such as people, actions, and scenes in the digital human video frames. The annotation model used to extract digital human features from video frames can be a pre-trained convolutional neural network (CNN) model. After the digital human generation system 10 extracts digital human features based on the pre-trained CNN model, it can input the extracted digital human features into a large language model (LLM), which provides the text description of the digital human video. This text description is the annotation data of the large language model.

[0075] In one possible implementation, during the process of annotating the digital human features in the digital human video based on the annotation model, the digital human generation system 10 analyzes the text descriptions corresponding to different digital human features in the digital human video based on the large language model. The text descriptions are used to indicate the annotation data of the digital human video, and the text descriptions include the character dialogue descriptions, character expression descriptions, and character action descriptions of the digital human video frames.

[0076] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating a training data generation method based on digital human video, as provided in an embodiment of this application. Figure 3 In the example shown, the computing device acquires digital human videos. These digital human videos can be sourced from multiple sources, such as movie and TV series clips acquired in cooperation with film and television production companies, virtual character animations acquired from game companies, interactive videos of virtual anchors based on online digital human platforms, and related video content uploaded by social media users.

[0077] exist Figure 3 In the example shown, after acquiring the digital human video, the computing device processes it to obtain annotated digital human videos. For example, the computing device uses a Large Language Model (LLM) to analyze the visuals and dialogues in the digital human video, obtaining text descriptions corresponding to the digital human features. These text descriptions are also called text-annotated data. Examples of text-annotated data include dialogues, facial expressions, and actions in the digital human video. The digital human video and its corresponding text descriptions acquired by the computing device can serve as the basic dataset for training the digital human model. It is understandable that the text-annotated data can be further converted into speech-annotated data and image-annotated data, which will be discussed in detail below.

[0078] In one possible implementation, after the digital human generation system 10 analyzes the text descriptions corresponding to different digital human features in the digital human video based on a large language model, the digital human generation system 10 converts the text descriptions based on a speech model to obtain speech annotation data, and converts the text descriptions based on an image model to obtain image annotation data. The digital human video, the speech annotation data corresponding to the digital human video, and the image annotation data corresponding to the digital human video can also form the training data for the digital human model.

[0079] For example, in the annotation process of digital human videos, after the digital human generation system 10 extracts the digital human features, it inputs the digital human features and the audio track of the digital human video into a large language model. The large language model can analyze the digital human features and generate text descriptions corresponding to the digital human features. These text descriptions are the text annotation data of the digital human video. After generating text annotation data based on the large language model, the digital human generation system 10 can further convert the text annotation data based on a speech model to obtain speech annotation data, and then convert the text annotation data based on an image model to obtain image annotation data.

[0080] It is understandable that, in another possible implementation, during the process of annotating 10 pairs of digital human videos in the generation of digital humans, the computing device may not need to convert the text description based on the speech model or image model to obtain speech annotation data and image annotation data. Instead, it may directly process the digital human video based on the speech model or image model to obtain speech annotation data or image annotation data, without any specific limitation.

[0081] 202. Train a digital human model based on training data. The digital human model includes multiple feature extraction networks. These networks are used to extract different types of digital human features from the digital human video. The digital human features include one or more of the following: facial expression features, body movement features, and scene features.

[0082] The digital human generation system 10 trains a digital human model based on training data, which includes digital human videos and labeled data of the digital human videos. Specifically, the digital human generation system 10 encodes the training data using an encoder to obtain a token sequence, which is used to indicate digital human features in the digital human videos. For example, the token sequence represents facial expressions, body movements, and human scenes in the digital human videos.

[0083] Please continue reading. Figure 3 ,exist Figure 3In the example shown, after the computing device annotates the digital human video to obtain training data for the digital human model, it trains the model based on this data. Specifically, the computing device uses an encoder to convert the digital human video data into a format that the model can understand. For example, the encoder processes the digital human video frame by frame, dividing it into smaller units and converting these units into preliminary tokenized results, i.e., tokens or token sequences. These tokens or token sequences can be represented by vectors; therefore, the tokenized results can also be called vector representations. It is understandable that since the digital human video encoded by the encoder has already been annotated, the tokens or token sequences encoded by the encoder also carry labels.

[0084] The encoder in this embodiment also includes multiple feature extraction networks. These feature extraction networks can be based on convolutional neural networks (CNNs) using deep learning technology. The feature extraction networks are used to extract digital human features from the digital human video. These digital human features include one or more of the following: facial expression features, body movement features, and scene features. Multiple feature extraction networks are used to extract different types of digital human features from the digital human video. For example, facial expressions in the digital human video are extracted using an expression feature extraction network, body movement features are extracted using a body movement feature extraction network, and scene features are extracted using a scene feature network.

[0085] In one possible implementation, during the encoding of training data by the encoder, the digital human generation system 10 uses a feature extraction network within the encoder to extract features from the digital human video, resulting in digital human features. Different types of digital human features correspond to different feature extraction networks. The digital human generation system 10 then uses a serialization module to convert the digital human features into a token sequence, where a mapping relationship exists between the tokens and the digital human features. For example, the digital human generation system 10 uses the serialization module to map these digital human features to tokens in a high-dimensional space. Each token represents one or more specific features in a video frame. The digital human generation system 10 uses the encoder to convert the video data into a form suitable for processing by the transformer model, i.e., a token sequence.

[0086] The digital human generation system 10, after encoding digital human videos with an encoder to obtain labeled sequences, further trains a digital human model based on the labeled sequences and a self-attention mechanism, thereby obtaining a trained digital human model. The self-attention mechanism can generate new labeled sequences based on the correlation between labeled sequences, which is described in detail below:

[0087] In the process of training the digital human model based on the labeled sequence and the self-attention mechanism, the digital human generation system 10 can use the transformer model for training. The transformer model is a deep learning model for processing sequential data, which is the encoded labeled sequence. The transformer model can capture the dependencies between various positions in the sequence through the self-attention mechanism. For example, the transformer model learns the complex relationship between digital human tokens in this embodiment through the self-attention mechanism, thereby generating new token sequences. These new token sequences are new and coherent digital human video sequences.

[0088] During the training process of the Transformer model, the Transformer model can learn how to predict the next token based on a given token sequence, thereby generating new video frames. By training on large-scale training data, the Transformer model will gradually develop the ability to generate digital human videos, capable of generating new digital human behaviors and expressions that have never appeared before, thus obtaining the trained digital human model.

[0089] Please see Figure 4 , Figure 4 This is a schematic diagram illustrating the training of a digital human model, as provided in an embodiment of this application. Figure 4 In the example shown, the computing device converts the digital human video data into an initial token sequence using an encoder, and then trains a transformer model based on the initial token sequence and labeled data. Specifically, the computing device needs to select the transformer model architecture and determine the transformer model parameters according to the task objective. The model architecture is, for example, the BERT architecture, and the model parameters are, for example, the number of model layers and the dimension of the hidden layers.

[0090] exist Figure 4 In the example shown, after the computing device determines the model architecture and parameters of the transformer model, it uses an initial token sequence as input to the transformer model. This token sequence contains key features from the digital human video, which form the basis for the transformer model's learning and prediction. The token sequence input to the transformer model undergoes forward propagation. During forward propagation, the transformer model processes the input token sequence layer by layer and generates outputs. These outputs can be used, along with labeled data, to calculate the loss function and guide the training direction of the transformer model.

[0091] exist Figure 4In the example shown, the computing device performs backpropagation based on the loss value obtained from the calculated loss function. During backpropagation, the computing device calculates gradients layer by layer and passes them to the previous layer until it reaches the input layer. These gradients guide the parameter update direction of the transformer model. The computing device adjusts the parameters of the transformer model according to the gradients to minimize the loss value. The computing device iteratively trains the transformer model by repeating the process of forward propagation, loss calculation, backpropagation, and parameter update to ensure that the transformer model can converge and reach the expected performance level.

[0092] Understandably, during the process of training a digital human model based on training data, computing devices can also evaluate and optimize the trained digital human model. For example, computing devices can evaluate the performance of the digital human model using validation datasets and perform hyperparameter tuning on the trained digital human model. After completing the training of the digital human model, the computing device can deploy the trained digital human model to real-world application scenarios to perform multimodal interactions of the digital human model.

[0093] This application provides a digital human encoder that can tokenize a large number of digital human videos and use the Transformer model to train a digital human model. This results in digital humans generated by the trained digital human model having a higher similarity to real people in terms of senses. Whether in appearance or speaking style, it can more realistically simulate real people, thereby improving the generation quality and interactivity of digital humans.

[0094] 203. Perform multimodal interaction tasks based on the trained digital human model. The multimodal interaction tasks include generating corresponding digital human videos based on multimodal content input by the user.

[0095] The digital human generation system 10 performs multimodal interaction tasks based on a trained digital human model. The multimodal interaction tasks include the trained digital human model reasoning based on multimodal content input by the user to generate corresponding digital human videos. The multimodal content input by the user is used to guide the digital human videos generated by the digital human model. The multimodal content input by the user includes control commands in the form of text, voice, images, and video.

[0096] In this embodiment, the digital human model can generate digital human videos based on different types of input. These videos feature actions and expressions corresponding to user control commands. Users can further interact with the digital human in the videos through various means such as text, voice, and video, and the digital human model continues to generate corresponding videos.

[0097] In one possible implementation, during the multimodal interaction task performed by the digital human generation system 10 based on a trained digital human model, the computing device receives multimodal content input by the user. This multimodal content serves as an instruction for the digital human model and includes one or more of the following: text, speech, images, and video. The computing device generates a digital human interactive avatar corresponding to the multimodal content based on the digital human model, such as a digital human video. This digital human interactive avatar includes digital human facial expressions, digital human actions, and digital human speech corresponding to the multimodal content.

[0098] Please see Figure 5 , Figure 5 This is a schematic diagram illustrating an application of a digital human model provided in an embodiment of this application. Figure 5 In the example shown, the computing device can use a trained digital human model to perform multimodal reasoning to generate digital human videos. For example, the user inputs a text message: "Generate a digital human avatar and say the following: Hello everyone, my name is..." After receiving the user's text, the digital human model converts the text into a token sequence through an encoder. Then, the digital human model generates a corresponding token sequence of digital human behavior and expressions. The digital human model converts these token sequences into video frames through a decoder, thereby generating a digital human video. In the digital human video, the digital image says "Hello everyone, my name is..." according to the user's input.

[0099] exist Figure 5 In the example shown, the multimodal content input by the user can also be voice, images, and videos. The digital human model can convert these inputs into initial token sequences through corresponding preprocessing steps. Then, based on these initial token sequences, the digital human model can perform inference to generate corresponding token sequences for digital human behavior and expressions. Finally, a decoder converts these token sequences into video frames, thereby generating a digital human video. For example, if a user inputs a cartoon rabbit image with the attached voice message "Generate a digital human that closely resembles the input image, and the digital human is dancing," the digital human model can process the aforementioned multimodal content input by the user and generate a digital human video in which the cartoon rabbit is dancing.

[0100] As can be seen from the above embodiments, the computing device in this application can train a digital human model based on various existing types of digital human videos, and generate digital human videos based on the trained digital human model. Since the computing device can directly generate digital human videos based on the trained digital human model, complex video processing and other operations are avoided, thereby improving the generation efficiency of digital humans. In addition, the digital human model in this application can generate corresponding digital human videos according to the multimodal content input by the user, improving the interactive capabilities of the digital human.

[0101] Based on the above method embodiments, this application also provides a digital human generation device, which is described in detail below.

[0102] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a digital human generation device provided in an embodiment of this application. Figure 6 In the example shown, the digital human generation device 600 is used to implement the various steps performed by the digital human generation system in the above embodiments. The digital human generation device 600 includes an acquisition unit 601 and a processing unit 602.

[0103] The acquisition unit 601 is used to acquire digital human videos, which are used to generate training data for the digital human model. The data sources for the digital human videos include one or more of the following: film and television works, game animation videos, digital human application videos, and social media videos. The processing unit 602 is used to train the digital human model based on the training data. The digital human model includes multiple feature extraction networks, which are used to extract different types of digital human features from the digital human videos. These digital human features include one or more of the following: facial expression features, body movement features, and scene features. The processing unit 602 is also used to perform multimodal interaction tasks based on the trained digital human model. These multimodal interaction tasks include generating corresponding digital human videos based on multimodal content input by the user.

[0104] In one possible implementation, the processing unit 602 is further configured to annotate the digital human features in the digital human video based on the annotation model to obtain training data. The annotation model includes models corresponding to different types of annotation data, including text annotation data, speech annotation data and image annotation data.

[0105] In one possible implementation, the processing unit 602 is specifically used to analyze the text descriptions corresponding to different digital human features in the digital human video based on a large language model. The text descriptions are used to indicate the annotation data of the digital human video. The text descriptions are converted based on the speech model to obtain speech annotation data, and the text descriptions are converted based on the image model to obtain image annotation data.

[0106] In one possible implementation, the processing unit 602 is specifically used to encode the training data based on the encoder to obtain a label sequence, the label sequence being used to indicate the digital human features in the digital human video, and to train a digital human model based on the label sequence and a self-attention mechanism, the self-attention mechanism being used to generate new label sequences according to the correlation between the label sequences.

[0107] In one possible implementation, the processing unit 602 is specifically used to extract features from the digital human video based on the feature extraction network to obtain digital human features. Different types of digital human features correspond to different feature extraction networks. The digital human features are converted into a labeled sequence based on the serialization module. There is a mapping relationship between the labels in the labeled sequence and the digital human features.

[0108] In one possible implementation, the processing unit 602 is specifically used to receive multimodal content input by the user. The multimodal content is used to instruct the digital human model on input commands. The multimodal content includes one or more of the following: text, speech, and images. Based on the digital human model, a digital human interactive avatar corresponding to the multimodal content is generated. The digital human interactive avatar includes digital human facial expressions, digital human actions, and digital human speech corresponding to the multimodal content.

[0109] It is understandable that the acquisition unit 601 and processing unit 602 in the digital human generation device 600 can function as functional modules. Figure 1 The various modules in the digital human generation system 10 are mapped to each other, thereby realizing the functions of each module in the digital human generation system 10.

[0110] It should be understood that the division of units in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, all units in the device can be implemented entirely through software calls from processing elements; all units can be implemented entirely in hardware; or some units can be implemented through software calls from processing elements, and others in hardware. For example, each unit can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as a program in memory, called and executed by a processing element of the device. Moreover, these units can be fully or partially integrated together, or implemented independently. The processing element mentioned here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above units can be implemented through integrated logic circuits in the processor element or through software calls from processing elements.

[0111] It is worth noting that, for the sake of simplicity, the above method embodiments are described as a series of actions. However, those skilled in the art should know that this application is not limited to the order of the described actions. Furthermore, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0112] Other reasonable combinations of steps that can be conceived by those skilled in the art based on the above description also fall within the scope of protection of this application. Furthermore, those skilled in the art should also be aware that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to this application.

[0113] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Figure 7 As shown, the computing device 700 includes a processor 701, a memory 702, a communication interface 703, and a bus 704. The processor 701, memory 702, and communication interface 703 are coupled via the bus (not shown in the figure). The memory 702 stores instructions. When the instructions in the memory 702 are executed, the computing device 700 executes the method performed by the digital human generation system in the above method embodiment.

[0114] The computing device 700 may be one or more integrated circuits configured to implement the methods described above, such as: one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these forms of integrated circuits. Furthermore, when the units in the device can be implemented in the form of a processing element scheduler, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these units may be integrated together to implement a system-on-a-chip (SOC).

[0115] The processor 701 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0116] The memory 702 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0117] The memory 702 stores executable program code, and the processor 701 executes the executable program code to implement the functions of the aforementioned units or modules, thereby realizing the above-mentioned method for generating a digital human. That is, the memory 702 stores instructions for executing the above-mentioned method for generating a digital human.

[0118] The communication interface 703 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 700 and other devices or communication networks.

[0119] In addition to the data bus, the 704 bus can also include a power bus, a control bus, and a status signal bus. The bus can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc. The bus can be divided into address bus, data bus, and control bus.

[0120] Please see Figure 8 , Figure 8 This is a schematic diagram of a computing device cluster provided in an embodiment of this application. Figure 8 As shown, the computing device cluster 800 includes at least one computing device 700.

[0121] like Figure 8 As shown, the computing device cluster 800 includes at least one computing device 700. The memory 702 of one or more computing devices 700 in the computing device cluster 800 may store the same instructions for executing the aforementioned digital human generation method.

[0122] In some possible implementations, the memory 702 of one or more computing devices 700 in the computing device cluster 800 may also store partial instructions for executing the aforementioned digital human generation method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions for executing the aforementioned digital human generation method.

[0123] It should be noted that the memories 702 in the different computing devices 700 within the computing device cluster 800 can store different instructions, each used to execute a portion of the functions of the aforementioned node load control device. That is, the instructions stored in the memories 702 of the different computing devices 700 can implement the functions of one or more modules in the acquisition unit and processing unit.

[0124] In some possible implementations, one or more computing devices 700 in the computing device cluster 800 can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc.

[0125] Please see Figure 9 , Figure 9 This is a schematic diagram illustrating the network connection of computer devices in a computer cluster, as provided in an embodiment of this application. Figure 9 As shown, the two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.

[0126] In one possible implementation, the memory in computing device 700A stores instructions for performing the fetch unit function. Meanwhile, the memory in computing device 700B stores instructions for performing the processing unit function.

[0127] It should be understood that Figure 9 The functions of computing device 700A shown can also be performed by multiple computing devices. Similarly, the functions of computing device 700B can also be performed by multiple computing devices.

[0128] In another embodiment of this application, a computer-readable storage medium is also provided, which stores computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device executes the method performed by the digital human generation system in the above method embodiment.

[0129] In another embodiment of this application, a computer program product is also provided, which includes computer-executable instructions stored in a computer-readable storage medium. When the processor of the device executes the computer-executable instructions, the device performs the method executed by the digital human generation system in the above method embodiments.

[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0131] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0132] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0133] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0134] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for generating a digital human, characterized in that, include: Acquire digital human videos, which are used to generate training data for digital human models. The data sources for the digital human videos include one or more of the following: film and television works, game animation videos, digital human application videos, and social media videos. The digital human model is trained based on the training data. The digital human model includes multiple feature extraction networks. The multiple feature extraction networks are used to extract different types of digital human features from the digital human video. The digital human features include one or more of the following: facial expression features, body movement features, and scene features. Multimodal interaction tasks are performed based on the trained digital human model, including generating corresponding digital human videos based on multimodal content input by the user.

2. The method according to claim 1, characterized in that, Before training the digital human model based on the training data, the method further includes: The digital human features in the digital human video are labeled based on the labeling model to obtain the training data. The labeling model includes models corresponding to different types of labeled data, including text labeled data, speech labeled data and image labeled data.

3. The method according to claim 2, characterized in that, The annotation of the digital human features in the digital human video based on the annotation model includes: Based on the analysis of the large language model, the text descriptions corresponding to different digital human features in the digital human video are used to indicate the labeled data of the digital human video; The text description is converted based on a speech model to obtain speech annotation data; The text description is transformed based on the image model to obtain image annotation data.

4. The method according to any one of claims 1 to 3, characterized in that, Training the digital human model based on the training data includes: The training data is encoded using an encoder to obtain a label sequence, which is used to indicate the digital human features in the digital human video; The digital human model is trained based on the labeled sequences and a self-attention mechanism, wherein the self-attention mechanism is used to generate new labeled sequences based on the correlation between the labeled sequences.

5. The method according to claim 4, characterized in that, The encoding of the training data based on the encoder includes: Based on the feature extraction network, feature extraction is performed on the digital human video to obtain the digital human features. Different types of digital human features correspond to different feature extraction networks. The digital human features are converted into a tag sequence based on the serialization module, and there is a mapping relationship between the tags in the tag sequence and the digital human features.

6. The method according to any one of claims 1 to 5, characterized in that, The multimodal interaction task performed based on the trained digital human model includes: The system receives multimodal content input by the user, the multimodal content being used to indicate input instructions for the digital human model, and the multimodal content including one or more of the following: text, voice, and image; Based on the digital human model, a digital human interactive image corresponding to the multimodal content is generated. The digital human interactive image includes digital human expressions, digital human actions, and digital human voices corresponding to the multimodal content.

7. A device for generating a digital human, characterized in that, include: The acquisition unit is used to acquire digital human videos, which are used to generate training data for digital human models. The data sources of the digital human videos include one or more of the following: film and television works, game animation videos, digital human application videos, and social media videos. A processing unit is configured to train the digital human model based on the training data. The digital human model includes multiple feature extraction networks, which are used to extract different types of digital human features from the digital human video. The digital human features include one or more of the following: facial expression features, body movement features, and scene features. The processing unit is also used to perform multimodal interaction tasks based on the trained digital human model, the multimodal interaction tasks including generating corresponding digital human videos based on multimodal content input by the user.

8. The apparatus according to claim 7, characterized in that, The processing unit is also used for: The digital human features in the digital human video are labeled based on the labeling model to obtain the training data. The labeling model includes models corresponding to different types of labeled data, including text labeled data, speech labeled data and image labeled data.

9. The apparatus according to claim 8, characterized in that, The processing unit is specifically used for: Based on the analysis of the large language model, the text descriptions corresponding to different digital human features in the digital human video are used to indicate the labeled data of the digital human video; The text description is converted based on a speech model to obtain speech annotation data; The text description is transformed based on the image model to obtain image annotation data.

10. The apparatus according to any one of claims 7 to 9, characterized in that, The processing unit is specifically used for: The training data is encoded using an encoder to obtain a label sequence, which is used to indicate the digital human features in the digital human video; The digital human model is trained based on the labeled sequences and a self-attention mechanism, wherein the self-attention mechanism is used to generate new labeled sequences based on the correlation between the labeled sequences.

11. The apparatus according to claim 10, characterized in that, The processing unit is specifically used for: Based on the feature extraction network, feature extraction is performed on the digital human video to obtain the digital human features. Different types of digital human features correspond to different feature extraction networks. The digital human features are converted into a tag sequence based on the serialization module, and there is a mapping relationship between the tags in the tag sequence and the digital human features.

12. The apparatus according to any one of claims 7 to 11, characterized in that, The processing unit is specifically used for: The system receives multimodal content input by the user, the multimodal content being used to indicate input instructions for the digital human model, and the multimodal content including one or more of the following: text, voice, and image; Based on the digital human model, a digital human interactive image corresponding to the multimodal content is generated. The digital human interactive image includes digital human expressions, digital human actions, and digital human voices corresponding to the multimodal content.

13. A computing device cluster, characterized in that, The device includes at least one computing device, the computing device including a processor coupled to a memory, the processor being used to store instructions that, when executed by the processor, cause the cluster of computing devices to perform the method of any one of claims 1 to 6.

14. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions are executed, they cause the computer to perform the method of any one of claims 1 to 6.

15. A computer program product, the computer program product comprising instructions, characterized in that, When the instructions are executed, they cause the computer to implement the method of any one of claims 1 to 6.