A multi-modal digital human generation method
By using a multimodal data fusion method, the problem of inconsistency between appearance and timbre in single-modal digital human cloning technology has been solved, achieving an improvement in the consistency and naturalness of digital human appearance and timbre, enhancing the expressiveness and interactivity of virtual humans, and promoting the development of virtual human technology.
Patent Information
- Application Number
- CN202510617627.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Existing digital human cloning technology mainly relies on single-modal data, resulting in a lack of consistency between the virtual human's appearance and vocal characteristics, affecting realism and naturalness.
A multimodal data fusion method is adopted, which establishes the mapping relationship between audio and facial features, and facial features and silent video through the Transformer architecture. Combined with VAE with frozen weights and VITS text encoder, a digital human image and voice cloning model is constructed. The model is trained using a warmup learning rate strategy and a weighted loss function to achieve accurate feature extraction and fusion of multimodal data.
It achieves a high degree of consistency and naturalness between the appearance and voice characteristics of digital humans, enhances the realism and interactivity of virtual humans, provides an immersive and personalized user experience, and promotes the widespread application of virtual human technology in fields such as virtual anchors and intelligent customer service.
Smart Images

Figure CN120526008B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of images, the field of speech and the field of digital people, and particularly relates to a digital person generation method based on multi-modal. BACKGROUND
[0002] In today's society, the fields of digital people technology, timbre synthesis and virtual image cloning have made significant progress, which is mainly due to the rapid development of artificial intelligence technology, computer vision and deep learning algorithms. Especially under the promotion of multi-modal technology, through the fusion of visual, audio and semantic multi-source data, the virtual person generation technology of highly restoring the real human image and sound characteristics has been realized. These technologies have been widely used in virtual anchors, intelligent customer service, virtual concerts and other scenarios, providing users with more immersive and personalized interactive experiences, and promoting the landing application of virtual people technology in media, entertainment, education and other fields.
[0003] Current digital person cloning technology mainly relies on single modal data, such as restoring appearance only through facial images or synthesizing timbre through audio samples. However, this single modal generation method often has limitations, making it difficult to capture the dynamic correlation between human image and sound, resulting in a lack of consistency between virtual person appearance and timbre characteristics, affecting the overall realism and naturalness. To solve the above problems, the present patent proposes a digital person generation method based on multi-modal. By fusing visual features, audio features and semantic information, the overall restoration of target person appearance, timbre and emotional expression is realized, ensuring that the virtual person achieves high consistency and naturalness in dynamic performance and sound synthesis. SUMMARY
[0004] The present application proposes a digital person generation method based on multi-modal to solve the technical problems existing in the above background technology.
[0005] In order to achieve the above purpose, the technical solution adopted by the present application is as follows:
[0006] S1, data acquisition: acquiring a talking video of different images under the same text, separating it into audio and non-talking video, extracting facial features and constructing a data set;
[0007] S2, building a digital person image cloning model: containing an audio codec module and a non-talking video generation module, establishing the mapping relationship between audio and facial features, and facial features and non-talking video through the Transformer architecture;
[0008] S3, training and testing of digital person image cloning model: training the model using the warmup learning rate strategy and weighted KL divergence and MSE as the loss function, and verifying the generation effect;
[0009] S4, build a digital human timbre cloning model: including text encoding, timbre encoding and decoding modules, through the VITS text encoder and the Transformer architecture with frozen pre-training weights to realize timbre cloning;
[0010] S5, digital human timbre cloning model training and testing: using warmup learning rate strategy and weighted LSD and MSE as loss function to train the model, and verifying the audio generation effect;
[0011] S6, integration framework: combine the image cloning model and the timbre cloning model, and realize the digital human question and answer exchange through the large language model.
[0012] As preferred, the step S1 specifically comprises: grouping the sound video according to the text content, and separating the sound video into audio and silent video; extracting the face features of the silent video, and constructing a data set D1=(txt,audio,face_features,video), wherein txt represents text information, audio is audio, vidio is silent video, and face_features represents face features; randomly cutting the audio to obtain an audio segment, and constructing a data set D2=(txt,audio,audio_stage), wherein audio_stage represents the audio segment; dividing the data set D1 and the data set D2 into a training set and a test set in a manner of 8:2.
[0013] As preferred, the audio codec module in the step S2 comprises: Patch framing, encoding and position encoding of the audio, generating an implicit space feature vector F s through a 3-layer Transformer encoder; and decoding F s into face features through a 3-layer Transformer decoder.
[0014] As preferred, the silent video generation module in the step S2 comprises: Transforme encoding of the face features and adding Gaussian noise, predicting noise distribution through Unet; encoding a single portrait through a VAE encoder with frozen weights, and performing feature transmission with Unet through a pooling operation; and obtaining the silent video through VAE decoding of the features output by the Unet with frozen weights.
[0015] Preferably, the step S3 of training and testing the digital human clone model comprises: inputting the training set of (audio, face_features) and (face_features, video) in the data set D1 into the audio codec module and the silent video generation module, respectively; using warmup as the learning rate change strategy, training for 50 rounds, and using weighted KL divergence and MSE as the loss function, wherein the loss function is calculated as follows: wherein β is the weight value, N is the sample quantity, y i is the real data, is the predicted data; inputting the test set of (audio, face_features) and (face_features, video) in the data set D1 into the audio codec module and the silent video generation module in the digital human clone model, respectively, to verify the effectiveness of the two modules, and then connecting the two modules to obtain the digital human clone model.
[0016] Preferably, the step S4 of the digital human voice clone model comprises: Token segmentation of the text description, Token encoding and position encoding of the text description, using the VITS text encoder as the main body, freezing the pre-training weight in the VITS text encoder, and finally obtaining the hidden space feature vector T a of the text, wherein a represents the vector dimension; Patch audio slicing of the audio segment, Patch encoding and position encoding of the audio segment, using two Transformer encodings to construct the voice encoder core, and finally obtaining the hidden space feature vector A a of the voice; constructing an audio decoder, using three Transformer decodings to construct the audio decoder core, and finally obtaining the cloned audio with the input voice features and the input text as the audio content.
[0017] Preferably, the step S5 of training and testing the digital human voice clone model comprises: inputting the training set of (txt, audio, audio_stage) in the data set D2 into the digital human voice clone model; using warmup as the learning rate change strategy, training for 100 rounds, and using weighted LSD and MSE as the loss function, wherein the loss function is calculated as follows: wherein β is the weight value, P(ω) is the real audio energy spectrum, is the generated audio energy spectrum; inputting the test set of (txt, audio, audio_stage) in the data set D2 into the digital human voice clone model to verify the effectiveness of the model, and finally obtaining the trained digital human voice clone model.
[0018] Compared with the prior art, the advantages and positive effects of the present application are that a multi-modal data fusion and accurate feature extraction system is constructed. Through this method, high restoration of digital human image and tone is realized, which can effectively solve the inconsistency between appearance and tone characteristics and the problem of inaccurate emotional expression, thereby improving the realism and naturalness of virtual humans. This method not only enhances the expressiveness of digital human characters in various application scenarios, but also provides more immersive and personalized user experience, promoting the wide application of digital human technology in virtual anchors, intelligent customer service, virtual concerts and other fields. In addition, through multi-modal deep fusion and optimization processing, the interactivity and emotional expression ability of virtual characters in different situations are greatly improved, promoting the further development of virtual human technology. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0020] Figure 1 It is an audio encoding and decoding module structure diagram; Figure 2 It is a silent video generation module structure diagram; Figure 3 It is a digital human image cloning model flowchart; Figure 4 It is a digital human tone model flowchart; Figure 5 It is a digital human cloning framework diagram. DETAILED DESCRIPTION
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0022] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, however, the present application can also be implemented in other ways different from those described herein, therefore, the present application is not limited to the specific embodiments disclosed in the following description.
[0023] Embodiment, in the current situation of wide application of digital human technology, the existing single modal digital human cloning technology has many defects, such as lack of consistency between appearance and tone, inaccurate emotional expression, which seriously affects user experience. In order to achieve the effect of generating highly realistic and natural digital human with accurate emotional expression and consistent appearance and tone, and solve the limitations of traditional digital human cloning technology, the present application adopts an innovative multi-modal digital human generation scheme, which is as follows:
[0024] Select a popular online education platform as the data collection source. This platform has numerous different-looking instructors explaining the same course text in audio-visual videos that cover a wide range of expressions, movements, and tone variations.
[0025] Data acquisition: Collect audio-visual videos of different-looking people under the same text, separate them into audio and silent videos, extract facial features, and build a dataset. Specifically, using professional audio-video processing software, group the audio-visual videos according to the text content, and separate the audio-visual videos into audio and silent videos; extract the facial features of the silent videos and build a dataset D1 = (txt, audio, face_features, video), where txt represents the text information, audio is the audio, vidio is the silent video, and face_features represents the facial features; randomly cut the audio to get audio clips and build a dataset D2 = (txt, audio, audio_stage), where audio_stage represents the audio clip; divide the dataset D1 and the dataset D2 into training set and test set in the ratio of 8:2.
[0026] Then build a digital human image cloning model: including an audio codec module and a silent video generation module, establish the mapping relationship between audio and facial features, and facial features and silent video through the Transformer architecture. Specifically, build the audio codec module in the digital human image cloning model, build the mapping relationship between audio and facial features, as shown in Figure 1 , perform Patch audio framing on the audio, then perform Patch encoding and position encoding, use 3 Transformer Encode to build the core of the audio encoder, and finally get the hidden space feature vector F s of the audio, s represents the vector dimension. Then use 3 Transformer Decode to build the core of the audio decoder, which receives the encoder layer-by-layer features at the same time, and finally decodes the hidden space feature vector F s of the audio to map to the facial features. Build the silent video generation module in the digital human image cloning model, build the mapping relationship between facial features and silent video, and realize the conversion of a single portrait into a silent video according to the continuously changing facial features. As shown in Figure 2As shown, three Transformer Encode are used to build the face feature encoder core to obtain the hidden space feature vector of the face feature. Noise is added to the obtained hidden space feature vector of the face feature, which needs to conform to Gaussian distribution, to construct the data required for the diffusion model. The vector is then sent to the Unet for training to predict the noise distribution. The single portrait is encoded by the VAE Encode with frozen weights, and the features are transmitted to the Unet through the pooling operation. The features output by the Unet are decoded by the VAE Decode with frozen weights to obtain the noiseless video. The audio codec module and the noiseless video generation module jointly constitute the digital human image cloning model, as shown in Figure 3 .
[0027] Training and testing of the digital human image cloning model: The model is trained using the warmup learning rate strategy and weighted KL divergence and MSE as the loss function to verify the generation effect. Specifically, the digital human image cloning model is trained, and the training set of (audio, face_features) and (face_features, video) in the data set D1 is input into the audio codec module and the noiseless video generation module, respectively. The single portrait in the noiseless video generation module is a static state of a middle figure, and the same training method is used for training. The learning rate change strategy uses warmup, and the model is trained for 50 rounds. The loss uses weighted KL divergence and MSE as the loss function to measure the difference between the generated sample distribution and the real data distribution, and to measure the difference between the generated sample and the real data. The loss function is defined as: where β is the weight value, N is the number of samples, y i is the real data, is the predicted data; the test set of (audio, face_features) and (face_features, video) in the data set D1 is input into the audio codec module and the noiseless video generation module in the digital human image cloning model, respectively, to verify the effectiveness of the two modules, and the digital human image cloning model is obtained by concatenation.
[0028] Building a digital human timbre cloning model: including text encoding, timbre encoding and decoding modules, and realizing timbre cloning through a frozen pre-training weight VITS text encoder and a Transformer architecture. Specifically, according to the given audio, the timbre is cloned to repeat more text content. The text encoding part of the digital human timbre cloning model is shown in Figure 4 , Token segmentation is performed on the text description, and Token encoding and position encoding are performed. The VITS text encoder is used as the main body, and the pre-training weights in the VITS text encoder are frozen. Finally, the hidden space feature vector T awherein a represents the vector dimension. The timbre encoder in the digital human timbre cloning model is built, the audio segment is Patch audio sliced, then Patch encoding and position encoding are performed, two Transformer Encode are used to build the timbre encoder core, and finally the hidden space feature vector A of the timbre is obtained a The audio decoder in the digital human timbre cloning model is built, three Transformer Dcode are used to build the audio decoder core, finally the cloned audio with input timbre features and audio content as input text is obtained.
[0029] Digital human timbre cloning model training and testing: the model is trained by using the warmup learning rate strategy and weighted LSD and MSE as the loss function, and the audio generation effect is verified. Specifically, the digital human timbre cloning model is trained, and the training set of (txt, audio, audio_stage) in the data set D2 is input into the digital human timbre cloning model. The learning rate change strategy of the training method is warmup, and the training is performed for 100 rounds. The loss adopts weighted LSD (Log-spectral distance) and MSE as the loss function, which is used to measure the difference between the generated audio distribution and the real audio distribution, and the difference between the generated sample and the real data. The loss function is defined as wherein β is the weight value, P(ω) is the real audio energy spectrum, is the generated audio energy spectrum; the test set of (txt, audio, audio_stage) in the data set D2 is input into the digital human timbre cloning model, and the effectiveness of the model is verified, and finally the trained digital human timbre cloning model is obtained.
[0030] Finally, the digital human cloning framework is built, as shown in Figure 5 The obtained digital human image cloning model and the obtained digital human timbre cloning model are integrated in function, the image and timbre of the person in the given video are cloned, the large language model (LLM) is driven, and the question and answer exchange of the cloned digital human is realized.
[0031] The above is only a preferred embodiment of the present application, and is not intended to limit the present application in other forms. Any skilled person in the art can modify or change the above disclosed technical content to equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made on the basis of the technical essence of the present application to the above embodiments still belongs to the protection scope of the present application.
Claims
1. A multi-modal digital human generation method, characterized by, The method comprises the following steps: S1, data acquisition: obtaining voice videos of different digital human images under the same text, separating them into audio and silent videos, extracting facial features and building a data set; S2, building a digital human image cloning model: including an audio codec module and a silent video generation module, establishing the mapping relationship between audio and facial features, and facial features and silent videos through the Transformer architecture; S3, training and testing of the digital human image cloning model: using the warmup learning rate strategy and weighted KL divergence and MSE as the loss function to train the model and verify the generation effect; S4, building a digital human voice tone cloning model: including text encoding, voice tone encoding and decoding modules, realizing voice tone cloning through a VITS text encoder with frozen pre-training weights and a Transformer architecture; S5, training and testing of the digital human voice tone cloning model: using the warmup learning rate strategy and weighted LSD and MSE as the loss function to train the model and verify the audio generation effect; S6, integration framework: combining the image cloning model and the voice tone cloning model to realize digital human question and answer exchange through a large language model; The step S1 specifically comprises: Grouping the voice videos according to the text content, and separating the voice videos into audio and silent videos; extracting facial features of the silent video, constructing a dataset wherein, denotes text information, is audio, is silent video, denotes facial features; randomly intercepting audio to obtain an audio clip, constructing a dataset wherein represents an audio clip; Dataset and dataset The dataset is divided into training and testing sets in an 8:2 ratio. The step S3 of training and testing the digital human image cloning model comprises: The training set of the data set , and is input into the audio codec module and the silent video generation module, respectively; The learning rate change strategy adopts warmup, trains for 50 rounds, and uses weighted KL divergence and MSE as the loss function, which is calculated as follows: wherein is a weight value, N is the number of samples, is the real data, is the predicted data; The test set of and and in the data set is respectively sent into the audio codec module and the silent video generation module in the digital human clone model to verify the effectiveness of the two modules respectively, and the digital human clone model is obtained in series.
2. The multi-modal digital human generation method of claim 1, wherein, The audio codec module in step S2 comprises: Patch frame, encode and position coding of audio, generate hidden space feature vector through 3-layer transformer encoder ; are decoded into facial features by a 3-layer Transformer decoder. are decoded into facial features.
3. The multi-modal digital human generation method of claim 1, wherein, The silent video generation module in step S2 comprises: Transforme encoding of facial features and adding Gaussian noise, predicting noise distribution through Unet; Using a VAE encoder with frozen weights to encode a single portrait, and performing feature transmission through a pooling operation and Unet; Unet output features are decoded by a VAE decoder with frozen weights to obtain silent videos.
4. The multi-modal digital human generation method of claim 1, wherein, The step S4 of the digital human voice tone cloning model comprises: Tokenize the text description, then Tokenize and position encode it, use VITS text encoder as the main body, freeze the pre-training weights in the VITS text encoder, and finally get the hidden space feature vector of the text where a represents the dimension of the vector Patch audio slices the audio segment, then encodes and positions it, uses two Transformer encodings to build the timbre encoder core, and finally gets the timbre hidden space feature vector ; Building an audio decoder, using three Transformer decoders to build the core of the audio decoder, and finally obtaining cloned audio with input voice tone features and audio content as input text.
5. The multi-modal digital human generation method of claim 1, wherein, The step S5 of training and testing the digital human voice tone cloning model comprises: inputting a training set of the data set into a digital human vocal color cloning model; The training mode learning rate change strategy adopts warmup, and the training is performed for 100 rounds. The weighted LSD and MSE are used as the loss function, and the calculation method of the loss function is as follows: wherein is a weight value, is a real audio energy spectrum, is a generated audio energy spectrum; The dataset in the test set of the digital human timbre cloning model is put into the digital human timbre cloning model, the effectiveness of the model is verified, and finally a trained digital human timbre cloning model is obtained.
Citation Information
Patent Citations
Method for generating digital human voice and facial animation through text
CN116863038A
Emotion synchronization 2D digital human model training method and device
CN119598346A