Digital human image generation method combining cold start driving and active learning mechanism

By combining cold-start driven and active learning mechanisms, and utilizing a few-sample emotional speech generation model and a digital human image generation model, the personalization and consistency issues of digital human image generation under data scarcity conditions are solved, achieving high-quality automatic generation of digital human images and enhancing the natural interaction capabilities of digital humans.

CN120833401BActive Publication Date: 2025-12-05SHAANXI JIEQI NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511324476.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-05
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing methods for generating digital human images struggle to achieve high quality, strong consistency, and strong individual expression capabilities under conditions of data scarcity and user cold start, failing to meet the needs of user customization and emotion-driven behavior.

Method used

This method combines cold-start driven and active learning mechanisms. By pre-setting a few-sample emotional speech generation model and a digital human image generation model, and using a cold-start quality evaluator to screen training samples, combined with a conditional generative adversarial network and an expression controller module, target audio files and images are generated to achieve consistency between emotional expression and personality expression.

Benefits of technology

The system achieves high-quality, personalized digital human avatars automatically generated under extremely limited data conditions, breaking through the modeling barriers of digital human systems, enhancing the natural interaction capabilities and style consistency of digital humans, and realizing the generation of target audio files with high fidelity and consistent emotional and personality expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833401B_ABST
    Figure CN120833401B_ABST
Patent Text Reader

Abstract

The application discloses a digital human image generation method combining cold start driving and active learning mechanism. The preset few-sample emotional speech generation model derived from the cold start process and having high personalization and dynamic adaptability is used to process input data, efficient personalized and multi-emotional speech batch generation can be realized on large-scale unlabeled input data, and a target audio file with high fidelity and consistent emotional and personal expression can be output. The preset few-sample emotional speech generation model is trained based on the first qualified sample obtained by screening the candidate training sample by using the first cold start quality evaluator, so that the model can be started under the condition of very few data, and high-quality personalized digital human image automatic generation is realized under the condition of small samples. In addition, by inputting a target emotion embedding vector and a target personality feature vector, expression synchronous generation under speech driving is realized, and the natural interaction ability and style consistency of the digital human are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method for generating digital human images that combines cold start-driven and active learning mechanisms. Background Technology

[0002] Currently, with the widespread emergence of applications such as virtual reality, augmented reality, digital humans, intelligent customer service, and virtual anchors, the demand from all sectors of society for digital humans with highly personalized, emotional expression, and synchronized voice and facial expressions is rapidly increasing.

[0003] Currently, most existing digital human avatar generation methods rely on large amounts of manually labeled data, and users can only provide a very limited amount of information upon first use, making it difficult to adapt to user customization and emotion-driven needs. Therefore, how to construct a digital human avatar generation method that can achieve high quality, strong consistency, and strong personalized expression under conditions of data scarcity and user cold start has become an urgent problem to be solved. Summary of the Invention

[0004] This application aims to at least solve the technical problems existing in the prior art. To this end, the first aspect of this application proposes a digital human image generation method that combines cold start driving and active learning mechanisms, the method comprising:

[0005] The target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector are input into a preset few-shot emotion speech generation model for processing, generating target audio files corresponding to each target digital person in the target digital population; wherein, the target audio files are consistent with the emotion expression and personality expression of the target digital person; the preset few-shot emotion speech generation model is trained based on the first qualified sample after being screened by the first cold start quality evaluator on the text-speech pair data training samples of candidate joint personality emotion;

[0006] The target audio file, target emotion embedding vector, and target personality feature vector are input into the preset digital human image generation model for processing to generate the target image corresponding to each target digital human. The target image is used to reflect the facial state consistent with the emotion and personality expression in the target audio file. The preset digital human image generation model is built on the basis of conditional generative adversarial network and includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

[0007] In one possible implementation, the generation process of the preset few-sample emotional speech generation model includes:

[0008] The text-speech pair data training samples of candidate joint personality and emotion are input into the first cold start quality evaluator for processing, and the first comprehensive quality score is calculated by the first preset evaluation formula; wherein, the first preset evaluation formula is determined based on the first clarity index, the emotion consistency index, the first diversity index, and the first anomaly confidence.

[0009] Based on the first comprehensive quality score and the first preset screening threshold, the first qualified samples are selected, and the first qualified samples are sampled in a balanced manner to obtain the training set and the validation set.

[0010] On the training set, the parameters of the initial emotional speech generation model are updated by maximizing the preset training objective to obtain the updated model parameters. The preset training objective is determined based on the Mel spectrogram, text data, emotion embedding vector, and personality feature vector corresponding to the speech data. The initial emotional speech generation model is constructed based on the Tacotron2 conditional speech synthesis model and the conditional variational autoencoder mechanism. The conditional variational autoencoder mechanism fuses the emotion embedding vector and personality feature vector and uses them as conditional variables as input.

[0011] On the validation set, the updated model parameters are validated by calculating the validation loss, and the target model parameters are obtained by backpropagating the gradient.

[0012] Based on the target model parameters, a preset few-sample emotional speech generation model is generated.

[0013] In one possible implementation, text-to-speech pair data training samples of candidate joint personality emotions are input into a first cold-start quality evaluator for processing, and a first comprehensive quality score is calculated using a first preset evaluation formula, including:

[0014] Input the text-speech pair training samples of candidate joint personality emotion into the first cold start quality evaluator to obtain the signal-to-noise ratio, speech perception quality score and speech intelligibility index corresponding to the text-speech pair training samples of candidate joint personality emotion. Calculate the first clarity index based on the signal-to-noise ratio, speech perception quality score and speech intelligibility index.

[0015] Based on the emotional Mel spectrograms of the text-speech pairs training samples of candidate joint personality emotions, the predicted emotion distribution is obtained, and the emotion consistency index is calculated based on the similarity between the predicted emotion distribution and the preset emotion labels.

[0016] The first diversity index is calculated based on the distance between the training samples of text-speech pairs of candidate joint personality emotions.

[0017] Based on the emotional Mel spectrogram and preset emotion labels, the first probability that the text-speech pair data training sample of the candidate joint personality emotion is generated as an anomaly is calculated, and the first anomaly confidence is obtained based on the first probability.

[0018] The first comprehensive quality score is calculated based on the first clarity index, the sentiment consistency index, the first diversity index, and the first anomaly confidence level.

[0019] In one possible implementation, the updated model parameters are validated by calculating the validation loss, and the target model parameters are obtained by backpropagating the gradient, including:

[0020] The updated model parameters are validated, and the validation loss is calculated.

[0021] The gradient is backpropagated based on the validation loss, meta-learning rate, loss term obtained from the score of the first cold start quality evaluator, and preset adjustment weights, and the global model initialization parameters are updated in combination with the cold start iteration mechanism to obtain the target model parameters.

[0022] In one possible implementation, the process of generating a pre-defined digital human image model includes:

[0023] Obtain a candidate training sample set; wherein, the candidate training sample set includes multiple candidate training samples;

[0024] The candidate training sample set is input into the second cold start quality evaluator for processing, and the second comprehensive quality score is calculated by the second preset evaluation formula; wherein, the second preset evaluation formula is determined based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.

[0025] Based on the second comprehensive quality score and the second preset screening threshold, the second qualified sample is obtained.

[0026] The second qualified sample is input into the initial image generation model for training, and the total training loss is calculated. The total training loss includes adversarial loss, identity preservation loss, emotion consistency loss, temporal smoothing loss, and reconstruction loss.

[0027] The model parameters are updated based on the total training loss to generate a preset digital human image generation model.

[0028] In one possible implementation, the candidate training sample set is input into a second cold-start quality evaluator for processing, and a second comprehensive quality score is calculated using a second preset evaluation formula, including:

[0029] The candidate training sample set is input into the second cold start quality evaluator to obtain the no-reference image quality evaluation index and Laplacian variance corresponding to the candidate training sample. The second sharpness index is calculated based on the no-reference image quality evaluation index and Laplacian variance.

[0030] The semantic consistency index is calculated based on the similarity between the first predicted identity embedding vector corresponding to the candidate training sample and the second predicted identity embedding vector corresponding to the original face image.

[0031] The second diversity index is calculated based on the distance between each candidate training sample;

[0032] The second probability that the candidate training sample is generated as an anomaly is calculated based on the candidate training sample and the second predicted identity embedding vector, and the second anomaly confidence is obtained based on the second probability.

[0033] The second comprehensive quality score is calculated based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.

[0034] In one possible implementation, the target audio file, the target emotion embedding vector, and the target personality feature vector are input into a preset digital human image generation model for processing to generate the target image corresponding to each target digital human, including:

[0035] The target audio file, target emotion embedding vector, and target personality feature vector are input into the preset digital human image generation model. The cold start quality assessment and playback module evaluates and filters the current input data to obtain qualified input data.

[0036] The expression controller module performs frame segmentation and encoding on the target emotion embedding vector in the qualified input data to obtain the emotion control vector that controls the expression state of the image.

[0037] The personality adjustment vector is obtained by mapping the target personality feature vector in the qualified input data through the personality adjustment module; the personality adjustment vector is used to control the facial expression style and detailed features of the target digital human.

[0038] The conditional self-adaptor module jointly maps the target emotion embedding vector and target personality feature vector in the qualified input data into a layer-by-layer style code, and outputs the self-adapted conditional style code; the self-adapted conditional style code is used to modulate the feature mapping of each layer of the image generator module.

[0039] The image generator module processes the conditional style code, the emotion control vector, and the personality adjustment vector in the qualified input data to generate the target image corresponding to each target digital human.

[0040] A second aspect of this application proposes a digital human image generation device that combines cold start-driven and active learning mechanisms, the device comprising:

[0041] The first generation module is used to input the target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector into the preset few-sample emotion speech generation model for processing, and generate the target audio file corresponding to each target digital person in the target digital population; wherein, the target audio file is consistent with the emotion expression and personality expression of the target digital person; the preset few-sample emotion speech generation model is obtained by training on the first qualified sample after being screened by the first cold start quality evaluator on the text-speech pair data training sample of the preset joint personality emotion;

[0042] The second generation module is used to input the target audio file, target emotion embedding vector, and target personality feature vector into the preset digital human image generation model for processing, and generate the target image corresponding to each target digital human; wherein, the target image is used to reflect the facial state consistent with the emotion expression and personality expression in the target audio file; the preset digital human image generation model is built on the basis of conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

[0043] In one possible implementation, the above-described device is further used for:

[0044] The text-speech pair data training samples of candidate joint personality and emotion are input into the first cold start quality evaluator for processing, and the first comprehensive quality score is calculated by the first preset evaluation formula; wherein, the first preset evaluation formula is determined based on the first clarity index, the emotion consistency index, the first diversity index, and the first anomaly confidence.

[0045] Based on the first comprehensive quality score and the first preset screening threshold, the first qualified samples are selected, and the first qualified samples are sampled in a balanced manner to obtain the training set and the validation set.

[0046] On the training set, the parameters of the initial emotional speech generation model are updated by maximizing the preset training objective to obtain the updated model parameters. The preset training objective is determined based on the Mel spectrogram, text data, emotion embedding vector, and personality feature vector corresponding to the speech data. The initial emotional speech generation model is constructed based on the Tacotron2 conditional speech synthesis model and the conditional variational autoencoder mechanism. The conditional variational autoencoder mechanism fuses the emotion embedding vector and personality feature vector and uses them as conditional variables as input.

[0047] On the validation set, the updated model parameters are validated by calculating the validation loss, and the target model parameters are obtained by backpropagating the gradient.

[0048] Based on the target model parameters, a preset few-sample emotional speech generation model is generated.

[0049] In one possible implementation, the above-described device is further used for:

[0050] Input the text-speech pair training samples of candidate joint personality emotion into the first cold start quality evaluator to obtain the signal-to-noise ratio, speech perception quality score and speech intelligibility index corresponding to the text-speech pair training samples of candidate joint personality emotion. Calculate the first clarity index based on the signal-to-noise ratio, speech perception quality score and speech intelligibility index.

[0051] Based on the emotional Mel spectrograms of the text-speech pairs training samples of candidate joint personality emotions, the predicted emotion distribution is obtained, and the emotion consistency index is calculated based on the similarity between the predicted emotion distribution and the preset emotion labels.

[0052] The first diversity index is calculated based on the distance between the training samples of text-speech pairs of candidate joint personality emotions.

[0053] Based on the emotional Mel spectrogram and preset emotion labels, the first probability that the text-speech pair data training sample of the candidate joint personality emotion is generated as an anomaly is calculated, and the first anomaly confidence is obtained based on the first probability.

[0054] The first comprehensive quality score is calculated based on the first clarity index, the sentiment consistency index, the first diversity index, and the first anomaly confidence level.

[0055] In one possible implementation, the above-described device is further used for:

[0056] The updated model parameters are validated, and the validation loss is calculated.

[0057] The gradient is backpropagated based on the validation loss, meta-learning rate, loss term obtained from the score of the first cold start quality evaluator, and preset adjustment weights, and the global model initialization parameters are updated in combination with the cold start iteration mechanism to obtain the target model parameters.

[0058] In one possible implementation, the above-described device is further used for:

[0059] Obtain a candidate training sample set; wherein, the candidate training sample set includes multiple candidate training samples;

[0060] The candidate training sample set is input into the second cold start quality evaluator for processing, and the second comprehensive quality score is calculated by the second preset evaluation formula; wherein, the second preset evaluation formula is determined based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.

[0061] Based on the second comprehensive quality score and the second preset screening threshold, the second qualified sample is obtained.

[0062] The second qualified sample is input into the initial image generation model for training, and the total training loss is calculated. The total training loss includes adversarial loss, identity preservation loss, emotion consistency loss, temporal smoothing loss, and reconstruction loss.

[0063] The model parameters are updated based on the total training loss to generate a preset digital human image generation model.

[0064] In one possible implementation, the above-described device is further used for:

[0065] The candidate training sample set is input into the second cold start quality evaluator to obtain the no-reference image quality evaluation index and Laplacian variance corresponding to the candidate training sample. The second sharpness index is calculated based on the no-reference image quality evaluation index and Laplacian variance.

[0066] The semantic consistency index is calculated based on the similarity between the first predicted identity embedding vector corresponding to the candidate training sample and the second predicted identity embedding vector corresponding to the original face image.

[0067] The second diversity index is calculated based on the distance between each candidate training sample;

[0068] The second probability that the candidate training sample is generated as an anomaly is calculated based on the candidate training sample and the second predicted identity embedding vector, and the second anomaly confidence is obtained based on the second probability.

[0069] The second comprehensive quality score is calculated based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.

[0070] In one possible implementation, the second generation module is specifically used for:

[0071] The target audio file, target emotion embedding vector, and target personality feature vector are input into the preset digital human image generation model. The cold start quality assessment and playback module evaluates and filters the current input data to obtain qualified input data.

[0072] The expression controller module performs frame segmentation and encoding on the target emotion embedding vector in the qualified input data to obtain the emotion control vector that controls the expression state of the image.

[0073] The personality adjustment vector is obtained by mapping the target personality feature vector in the qualified input data through the personality adjustment module; the personality adjustment vector is used to control the facial expression style and detailed features of the target digital human.

[0074] The conditional self-adaptor module jointly maps the target emotion embedding vector and target personality feature vector in the qualified input data into a layer-by-layer style code, and outputs the self-adapted conditional style code; the self-adapted conditional style code is used to modulate the feature mapping of each layer of the image generator module.

[0075] The image generator module processes the conditional style code, the emotion control vector, and the personality adjustment vector in the qualified input data to generate the target image corresponding to each target digital human.

[0076] A third aspect of this application provides an electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a digital human image generation method combining cold start driving and active learning mechanisms as described in the first aspect.

[0077] The fourth aspect of this application proposes a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the digital human image generation method combining cold start driving and active learning mechanisms as described in the first aspect.

[0078] The embodiments of this application have the following beneficial effects:

[0079] The digital human image generation method combining cold start driving and active learning mechanisms provided in this application includes: inputting a target digital population, target text content corresponding to the target digital population, target emotion embedding vector, and target personality feature vector into a preset few-sample emotion speech generation model for processing, generating target audio files corresponding to each target digital human in the target digital population, wherein the target audio files are consistent with the emotion and personality expressions of the target digital human; inputting the target audio files, target emotion embedding vector, and target personality feature vector into the preset digital human image generation model for processing, generating target images corresponding to each target digital human, wherein the target images are used to reflect facial states consistent with the emotion and personality expressions in the target audio files; the preset digital human image generation model is constructed based on a conditional generative adversarial network, and includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module. This solution processes input data using a pre-defined few-sample emotional speech generation model derived from the cold-start process and possessing high personalization and dynamic adaptability. It enables efficient batch generation of personalized, multi-emotional speech on large-scale unlabeled input data, outputting high-fidelity target audio files with consistent emotional and personality expressions. Furthermore, the pre-defined few-sample emotional speech generation model is trained on the first qualified samples selected by the first cold-start quality evaluator from text-speech pairs of candidate joint personality emotions, allowing for startup under minimal data conditions. Simultaneously, an automated quality evaluation mechanism continuously optimizes the generated samples, expanding the available data space and breaking through the modeling barriers of existing digital human systems. This achieves high-quality, personalized digital human image automatic generation under small-sample conditions. Additionally, by inputting the target emotion embedding vector and target personality feature vector, it achieves voice-driven synchronous expression generation, enhancing the digital human's natural interaction capabilities and style consistency. Attached Figure Description

[0080] Figure 1 A block diagram of a computer device provided in an embodiment of this application;

[0081] Figure 2 A flowchart illustrating the steps of a digital human image generation method combining cold start driving and active learning mechanisms provided in this application embodiment;

[0082] Figure 3 A flowchart illustrating the steps for generating a preset few-sample emotional speech generation model provided in this application embodiment;

[0083] Figure 4 A flowchart illustrating the steps for calculating a first comprehensive quality score is provided in this application embodiment;

[0084] Figure 5 A flowchart illustrating the steps for obtaining target model parameters is provided in this application embodiment.

[0085] Figure 6 A flowchart illustrating the steps for constructing a preset digital human image generation model is provided in this application embodiment;

[0086] Figure 7 A flowchart illustrating the steps for calculating a second comprehensive quality score is provided in this application embodiment.

[0087] Figure 8 This application provides a flowchart of steps for generating target images corresponding to various target digital humans. Detailed Implementation

[0088] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0089] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more. Furthermore, the use of "based on" or "according to" implies openness and inclusiveness, because processes, steps, calculations, or other actions "based on" or "according to" one or more of the stated conditions or values ​​may in practice be based on additional conditions or beyond the stated values.

[0090] The digital human image generation method combining cold start driving and active learning mechanism provided in this application can be applied to computer equipment (electronic devices). The computer equipment can be a server or a terminal. The server can be a single server or a server cluster composed of multiple servers. This application does not specifically limit this. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets and portable wearable devices.

[0091] Taking a computer device as an example, Figure 1 A block diagram of a server is shown, such as Figure 1As shown, the server may include a processor and memory connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. When the computer program is executed by the processor, it implements a digital human image generation method that combines cold-start driven and active learning mechanisms.

[0092] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the server to which the present application is applied. Optionally, the server may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0093] Figure 2 A flowchart illustrating the steps of a digital human image generation method combining cold start driving and active learning mechanisms, provided in this application embodiment, includes the following steps:

[0094] Step 202: Input the target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector into the preset few-sample emotion speech generation model for processing, and generate the target audio file corresponding to each target digital person in the target digital population.

[0095] Among them, the target audio file is consistent with the emotional expression and personality expression of the target digital human. The preset few-sample emotional speech generation model is trained on the first qualified sample after being screened by the first cold start quality evaluator based on the text-speech pair data training samples of candidate joint personality and emotion.

[0096] like Figure 3 As shown, Figure 3 A flowchart illustrating the steps for generating a preset few-shot emotional speech generation model, provided in this application embodiment, includes:

[0097] Step 302: Input the text-speech pair data training samples of candidate joint personality emotion into the first cold start quality evaluator for processing, and calculate the first comprehensive quality score through the first preset evaluation formula.

[0098] The first preset evaluation formula is determined based on the first clarity index, the emotional consistency index, the first diversity index, and the first anomaly confidence level.

[0099] In some alternative embodiments, such as Figure 4 As shown, Figure 4A flowchart illustrating the steps for calculating a first comprehensive quality score, provided in this application embodiment, includes:

[0100] Step 402: Input the candidate joint personality emotion text-speech pair data training samples into the first cold start quality evaluator, obtain the signal-to-noise ratio, speech perception quality score and speech intelligibility index corresponding to the candidate joint personality emotion text-speech pair data training samples, and calculate the first clarity index based on the signal-to-noise ratio, speech perception quality score and speech intelligibility index.

[0101] Step 404: Based on the emotional Mel spectrum of the text-speech pair data training samples of the candidate joint personality emotion, the predicted emotion distribution is obtained, and the emotion consistency index is calculated based on the similarity between the predicted emotion distribution and the preset emotion label.

[0102] Step 406: Calculate the first diversity index based on the distance between the text-speech pair data training samples of each candidate joint personality emotion.

[0103] Step 408: Calculate the first probability that the text-speech pair training sample of the candidate joint personality emotion is generated abnormally based on the sentiment Mel spectrogram and preset sentiment labels, and obtain the first anomaly confidence based on the first probability.

[0104] Step 410: Calculate the first comprehensive quality score based on the first clarity index, the sentiment consistency index, the first diversity index, and the first anomaly confidence level.

[0105] In acquiring text-speech pair data training samples for candidate joint personality and emotion, the following steps are taken: First, the original text and emotion tags are encoded to obtain the encoding results. Then, the encoding results are processed using a preset diffusion model to generate a Mel spectrogram. Next, the Mel spectrogram is input into a preset neural vocoder to recover the time-domain speech waveform, generating multi-emotion text-speech pair enhancement data. Finally, the personality type and interest tags are encoded using a preset one-hot coding algorithm to obtain the encoding results. These encoding results are then sequentially concatenated and normalized to generate personality enhancement data. Based on the multi-emotion text-speech pair enhancement data and the personality enhancement data, candidate joint personality and emotion text-speech pair data training samples are generated.

[0106] Specifically, the emotional speech sample set includes original speech samples, the corresponding original text, and emotion tags. The original speech samples can be speech waveforms in Pulse Code Modulation (PCM) format, the original text is a string, and the emotion tags are strings, with four emotion tags—"joy," "anger," "sorrow," and "happiness"—supported by default.

[0107] Next, data augmentation can be performed on the emotional speech sample set. This process is based on the lightweight diffusion model Grad-TTS, a speech generation model based on a diffusion process that supports high-fidelity speech synthesis and enhancement under conditional control. It is used to increase the quantity of speech and enrich emotional features from a small number of original emotional speech samples, providing diverse training samples for subsequent speech synthesis.

[0108] Optionally, the original text and emotion tag can be encoded first to obtain the encoding result. For example, if the original text is "Hello!", the encoding result obtained after converting the original text into a text embedding sequence can be [13, 27, 45]. If the emotion tag is "joy", the encoding result obtained after converting the emotion tag into a unique encoding can be [1, 0, 0, 0].

[0109] Therefore, a conditional control vector can be constructed. Since only emotion control is considered, the conditional control vector is taken as the encoding result corresponding to the emotion label. Furthermore, a conditional diffusion generation process can be constructed based on the Grad-TTS model. For each original speech sample, under the corresponding fixed original text and emotion label, multiple samplings are performed through a preset diffusion model to generate speech spectrum representations with different styles, that is, to generate preset text-speech pair data training samples with preset joint personality and emotion.

[0110] Specifically, a random noise vector can be initialized first, and then, using the text embedding sequence and conditional control vector as conditions, a Grad-TTS model is used to perform conditional diffusion generation to obtain a Mel spectrogram. This Mel spectrogram has speech feature representations with consistent semantics, specified emotions, and style differences. Next, the Mel spectrogram can be input into a preset neural vocoder to recover the time-domain speech waveform, thereby outputting multi-emotional text-speech pair augmented data.

[0111] In addition, for personality types and interest tags, a pre-defined one-hot encoding algorithm can be used to encode the personality types and interest tags to obtain the encoding results. These encoded results are then concatenated and normalized sequentially to generate personality-enhanced data. Finally, based on the multi-emotional text-speech pair enhancement data and the personality enhancement data, candidate joint personality-emotion text-speech pair training samples can be generated.

[0112] After obtaining the training samples of text-speech pairs with preset joint personality and emotion, the training samples of text-speech pairs with candidate joint personality and emotion can be input into the first cold start quality evaluator to obtain the signal-to-noise ratio, speech perception quality score and speech intelligibility index corresponding to the training samples of text-speech pairs with candidate joint personality and emotion. The first clarity index is calculated based on the signal-to-noise ratio, speech perception quality score and speech intelligibility index. The first clarity index is used to measure whether the generated speech is clear and free of artifacts. The specific calculation process is shown in formula (1).

[0113] (1)

[0114] SNR represents the signal-to-noise ratio, and a higher SNR indicates clearer speech; PESQ represents the speech perception quality score; and STOI represents the speech intelligibility index. This indicates the primary resolution indicator; , , These represent the corresponding weights.

[0115] The predicted emotion distribution can be obtained by training the emotion Mel spectrogram of the candidate joint personality emotion text-speech pair data. The emotion consistency index is calculated based on the similarity between the predicted emotion distribution and the preset emotion label. The emotion consistency index is used to measure whether the emotion features of the generated speech are consistent with the input preset emotion label. The specific calculation process is shown in formula (2).

[0116] (2)

[0117] in, Indicators of emotional consistency; Indicates cosine similarity; This represents a small sentiment classifier that is trained through self-supervision and a small number of labeled samples, and outputs a predicted sentiment distribution. This indicates a preset emotion label.

[0118] It should be noted that the above-mentioned small sentiment classifier The model comprises an input layer, a first convolutional layer, a second convolutional layer, a pooling layer, a bidirectional Long Short-Term Memory (LSTM) layer, a fully connected layer, a random dropout layer, and an output layer. Specifically, two convolutional layers extract local features from the input data, followed by downsampling through pooling layers. Next, a bidirectional LSTM layer captures temporal information, enhancing the model's ability to perceive sentiment changes. The fully connected layer further fuses the extracted features and enhances the non-linear expression using a Rectified Linear Unit (ReLU) activation function. To prevent overfitting, a random dropout layer is introduced in the small sentiment classifier, and finally, the output layer, activated by a Softmax function, classifies the sentiment category.

[0119] The first diversity index is calculated based on the distance between the training samples of text-speech pairs of candidate joint personality emotions. The first diversity index aims to calculate the distance between different samples. The greater the difference, the better the diversity. This ensures that the enhanced speech samples of different samples have stylistic differences and avoids pattern collapse. The specific calculation process is shown in formula (3).

[0120] (3)

[0121] in, Indicates the first diversity indicator; , The training samples are text-to-speech pairs representing different candidate joint personality emotions; K represents the total number of samples.

[0122] Based on the Mel-ray spectrogram of sentiment and preset sentiment labels, the first probability that the text-speech pair training samples of candidate joint personality sentiment are generated abnormally is calculated. Based on the first probability, the first anomaly confidence is obtained, and a lightweight discriminator trained using contrastive learning can be used. To identify whether the generated spectrum is abnormal, the specific calculation process is shown in formula (4).

[0123] (4)

[0124] in, Indicates the first level of anomaly confidence; Mel spectrum representing emotion; This indicates a preset emotion label.

[0125] It should be noted that the aforementioned lightweight discriminator The system comprises an input layer, a first convolutional layer, a second convolutional layer, a pooling layer, a flattening layer, a fully connected layer, a discriminant layer, and an output layer. Specifically, two convolutional layers extract local features from the input data, and a pooling layer reduces the feature dimensionality. Next, the flattening layer passes the pooled feature vectors to the fully connected layer for feature fusion. Finally, the discriminant layer uses the sigmoid activation function to output the probability of whether a sample is abnormal, and the result is output through the output layer.

[0126] Finally, the first comprehensive quality score can be calculated based on the first clarity index, the sentiment consistency index, the first diversity index and the first anomaly confidence level. The specific calculation process is shown in formula (5).

[0127] (5)

[0128] in, Indicates the first overall quality score; , , , These represent the corresponding weights.

[0129] It should be noted that the aforementioned first cold start quality evaluator is equivalent to the first comprehensive quality score. That is, it is the weighted sum of the four indicators mentioned above: first clarity index, sentiment consistency index, first diversity index, and first anomaly confidence level.

[0130] Step 304: Based on the first comprehensive quality score and the first preset screening threshold, the first qualified samples are selected, and the first qualified samples are sampled in a balanced manner to obtain the training set and the validation set.

[0131] Specifically, the first comprehensive quality score can be compared with the first preset screening threshold. If the first comprehensive quality score is less than the first preset screening threshold, the corresponding sample is discarded; otherwise, it is regarded as the first qualified sample.

[0132] Next, the first qualified samples can be sampled evenly to obtain training and validation sets for subsequent model training. The trained model can quickly and proactively learn the emotional speech style of individual users with very few samples, complete the personalized transfer of new emotions, and generate personalized speech with multiple emotions.

[0133] Step 306: On the training set, update the parameters of the initial emotional speech generation model by maximizing the preset training objective to obtain the updated model parameters.

[0134] The preset training objective is determined based on the Mel spectrogram, text data, emotion embedding vector, and personality feature vector corresponding to the speech data. The initial emotion speech generation model is constructed based on the Tacotron2 conditional speech synthesis model and the conditional variational autoencoder mechanism. The conditional variational autoencoder mechanism fuses the emotion embedding vector and personality feature vector and uses them as conditional variables as input.

[0135] Specifically, optionally, during model training, a Tacotron2-based conditional speech synthesis model can be constructed first, and emotion and personality vectors can be fused and used as conditional variables input to the encoder and decoder of the model. By introducing a conditional variational autoencoder mechanism, the model learns how to generate accurate and stylistically consistent speech given semantic and stylistic features.

[0136] Preset training objectives It is determined based on voice data, text data, emotion embedding vectors, and personality feature vectors, and can be specifically represented as: ,in, Represents voice data; This represents the corresponding text data; , Represents the emotion embedding vector, Represents a personality trait vector; Represents a vector of latent variables, sampled from the standard normal distribution. ; It is an approximate posterior distribution, output by the encoder; The prior distribution of the latent variables is usually denoted as the standard normal distribution. . It is the reconstruction error term, representing the error in the latent variables. Under the posterior distribution, can the current model reproduce real speech data? . It is the relative entropy regularization term, representing the approximate posterior distribution of the model output. With prior distribution (usually a standard normal distribution) The distance between ( ) can prevent potential space degradation.

[0137] During training, a Few-shot mechanism is employed, enabling the model to quickly adapt to new tasks with little or no training data. In this learning mode, the model doesn't require a large amount of training data; instead, it leverages its powerful learning mechanism to complete tasks with minimal or no sample guidance. Therefore, a balanced sampling of the first qualified samples is performed to obtain the training and validation sets. On the training set, the parameters of the initial emotional speech generation model are updated by maximizing the preset training objective, resulting in the updated model parameters. , ,in, This indicates the initialization of model parameters. This represents the learning rate during the loop training process within the preset task. Indicates in the training set The preset training objectives.

[0138] Step 308: On the validation set, validate the updated model parameters by calculating the validation loss, and obtain the target model parameters by backpropagating the gradient.

[0139] In some alternative embodiments, such as Figure 5 As shown, Figure 5 A flowchart of steps for obtaining target model parameters provided in this application embodiment includes:

[0140] Step 502: Validate the updated model parameters and calculate the validation loss.

[0141] Step 504: Based on the validation loss, meta-learning rate, loss term obtained from the score of the first cold start quality evaluator, and preset adjustment weights, backpropagate the gradient, and update the global model initialization parameters in combination with the cold start iteration mechanism to obtain the target model parameters.

[0142] After training on the training set, outer loop optimization is required. This can be done on the validation set by calculating the validation loss to validate the updated model parameters, and then obtaining the target model parameters by backpropagating the gradient. Specifically, this can be done on the validation set. The performance of the updated model parameters is evaluated, and the validation loss is calculated. , ,in, The expression is the same as the preset training objective; It is the first qualified sample in the validation set; This represents the expectation of the validation loss over all validation sets. Then, the gradient can be backpropagated to update the global model initialization parameters, i.e. ,in, This indicates the preset meta-learning rate. This represents the loss term in the score obtained from the first cold start quality assessment. This indicates that the preset adjustment weights can be used to obtain the target model parameters.

[0143] Step 310: Generate a preset few-sample emotional speech generation model based on the target model parameters.

[0144] Once the target model parameters are obtained, they can be deployed in the model to obtain a preset few-sample emotional speech generation model.

[0145] In some optional embodiments, after the model training is complete, based on the active learning mechanism, when facing a new user or a new emotional task, it can also be denoted as... At this time, only a small number of supporting samples (such as 5) are needed to quickly complete fine-tuning and obtain the fine-tuned model parameters. , ,in, Indicates in The preset training objectives can be used to perform subsequent processing using the model with finely tuned parameters.

[0146] In this embodiment, a minimal sample-driven modeling mechanism is used. Users only need to provide 3-5 voice recordings, 1 facial image, and 1 set of personality tags to automatically construct a preliminary digital human voice and image style, achieving personalized modeling from 0 to 1. That is, the model construction has personalized cold start capability. In addition, the model parameters are updated quickly with a small number of samples, and the quality of the generated results is improved iteratively in rounds, realizing an end-to-end automatic modeling closed loop of "input-optimization-feedback", which significantly improves the cold start experience and generation stability.

[0147] Step 204: Input the target audio file, target emotion embedding vector, and target personality feature vector into the preset digital human image generation model for processing, and generate the target image corresponding to each target digital human.

[0148] The target image is used to reflect facial states consistent with the emotional and personality expressions in the target audio file. The preset digital human image generation model is built on the basis of conditional generative adversarial networks and includes an expression controller module, a personality modifier module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

[0149] The aforementioned preset digital human image generation model also needs to be pre-built. In some optional embodiments, such as Figure 6 As shown, Figure 6 A flowchart illustrating the steps for constructing a preset digital human image generation model, as provided in this application embodiment, includes:

[0150] Step 602: Obtain the candidate training sample set.

[0151] Step 604: Input the candidate training sample set into the second cold start quality evaluator for processing, and calculate the second comprehensive quality score using the second preset evaluation formula.

[0152] Step 606: Based on the second comprehensive quality score and the second preset screening threshold, the second qualified sample is obtained.

[0153] Step 608: Input the second qualified sample into the initial image generation model for training, and calculate the total training loss.

[0154] Step 610: Update the model parameters based on the total training loss and generate the preset digital human image generation model.

[0155] The candidate training sample set includes multiple candidate training samples, which may include face image samples, emotion embedding vector samples, personality feature vector samples, etc.

[0156] Next, the candidate training sample set can be input into the second cold start quality evaluator for processing, and the second comprehensive quality score can be calculated by the second preset evaluation formula. The second preset evaluation formula is determined based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.

[0157] In some alternative embodiments, such as Figure 7 As shown, Figure 7 A flowchart illustrating the steps for calculating a second comprehensive quality score, provided in this application embodiment, includes:

[0158] Step 702: Input the candidate training sample set into the second cold start quality evaluator, obtain the no-reference image quality evaluation index and Laplacian variance corresponding to the candidate training samples, and calculate the second sharpness index based on the no-reference image quality evaluation index and Laplacian variance.

[0159] Step 704: Calculate the semantic consistency index based on the similarity between the first predicted identity embedding vector corresponding to the candidate training sample and the second predicted identity embedding vector corresponding to the original face image.

[0160] Step 706: Calculate the second diversity index based on the distance between each candidate training sample.

[0161] Step 708: Calculate the second probability that the candidate training sample is generated as an anomaly based on the candidate training sample and the second predicted identity embedding vector, and obtain the second anomaly confidence based on the second probability.

[0162] Step 710: Calculate the second comprehensive quality score based on the second clarity index, semantic consistency index, second diversity index, and second anomaly confidence.

[0163] The second sharpness index is used to measure whether the generated image is clear and free of artifacts. The specific calculation process is shown in formula (6).

[0164] (6)

[0165] in, The second sharpness index is represented by NIQE and BRISQUE, which represent different no-reference image quality assessment indices. The specific calculation process can be found in existing technologies. The lower the value, the higher the quality. LaplVar is the Laplacian variance. The larger the value, the higher the edge sharpness. , , These represent the corresponding weights.

[0166] The semantic consistency index is used to measure whether the generated image retains the original identity and the rationality of the expression. The higher the cosine similarity, the more consistent the generated image is with the original identity. The specific calculation process is shown in formula (7).

[0167] (7)

[0168] in, Indicates semantic consistency metrics; This represents a face recognition model that outputs a predicted identity embedding vector. Indicates candidate training samples The corresponding first predicted identity embedding vector; Represents the original human face image The corresponding second predicted identity embedding vector.

[0169] It should be noted that the above-mentioned face recognition model Built upon a convolutional neural network, the model comprises an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a max-pooling layer, a fully connected layer, a random dropout layer, a feature embedding layer, and an output layer. Specifically, the three convolutional layers extract low-level and high-level features from the input face image, followed by a max-pooling layer to reduce the feature dimensionality. Next, a fully connected layer fuses these features, and the ReLU activation function enhances the non-linear representation. To prevent overfitting, the face recognition model introduces a random dropout layer, and the feature embedding layer maps the extracted features to a low-dimensional space, generating a unique face feature vector, i.e., the identity embedding vector, which is then output through the output layer.

[0170] The second diversity index is used to ensure that different candidate training samples have different styles. The larger the value, the stronger the diversity. The specific calculation process is shown in formula (8).

[0171] (8)

[0172] in, This represents the second diversity indicator; , represents different candidate training samples; L represents the number of candidate training samples.

[0173] The second anomaly confidence can be obtained from a discriminator trained by self-supervised contrastive learning. The output, and the specific calculation process are shown in formula (9).

[0174] (9)

[0175] in, This indicates the second level of anomaly confidence.

[0176] It should be noted that the above discriminator The network consists of an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fully connected layer, and an output layer. Specifically, four convolutional layers extract multi-level features from the input face image, gradually reducing the spatial dimension of the image. A Leaky Rectified Linear Unit (LeakyReLU) activation function is used to increase the network's non-linearity. Next, the features are fused through a fully connected layer, and a probability value is output through a sigmoid activation function to measure whether the input image was generated anomalously. This result is output through the output layer.

[0177] Finally, the second comprehensive quality score can be calculated based on the second clarity index, the semantic consistency index, the second diversity index and the second anomaly confidence, and the specific calculation process is shown in formula (10).

[0178] (10)

[0179] in, Indicates the second overall quality score; , , , These represent the corresponding weights.

[0180] It should be noted that the aforementioned second cold start quality evaluator is equivalent to the second comprehensive quality score. That is, it is the weighted sum of the four indicators mentioned above: the second clarity indicator, the semantic consistency indicator, the second diversity indicator, and the second anomaly confidence.

[0181] Therefore, a second qualified sample can be selected based on the second comprehensive quality score and the second preset screening threshold. Specifically, the second comprehensive quality score can be compared with the second preset screening threshold. If the second comprehensive quality score is less than the second preset screening threshold, the corresponding sample is discarded; otherwise, it is taken as the second qualified sample.

[0182] Next, the second qualified sample can be input into the initial image generation model for training, and the total training loss can be calculated. The total training loss includes adversarial loss, identity preservation loss, emotion consistency loss, temporal smoothing loss, and reconstruction loss. Finally, the model parameters are updated based on the total training loss to generate the preset digital human image generation model. The initial image generation model is also built on the basis of a conditional generative adversarial network, and the model parameters of the conditional generative adversarial network are obtained after initialization.

[0183] Among them, the improved generative adversarial network discriminator PatchGAN discriminator can be used to discriminate the generated face images. Compared with the preset real face images in the training sample set The losses in the confrontation To ensure the realism of the images, .in, It is a discriminator that represents the probability that the generated face is "real". T Indicates the length of the generated face image sequence; in the first term Indicates to The expected value of the logarithm of the discriminator results from the training set is calculated. This step aims to encourage the discriminator to analyze real face images from the training set. Providing high-confidence judgments enhances the ability to identify real images; in the second item... Indicates to The expected value of the logarithm of the discriminator results after complementation is calculated. This step aims to encourage the discriminator to correctly analyze the generated face images. The system provides a judgment that "this is a forgery," thereby enhancing the ability to identify forged images.

[0184] A pre-trained face recognition network can be used to extract image features, maintain identity consistency with the input face image augmentation samples, and calculate the identity preservation loss. , .in, It is a facial feature extractor, which is also the facial recognition model mentioned above. .

[0185] A pre-defined emotion classifier can be used to perform facial expression recognition on the generated facial images, matching them with the emotions in the input speech, thereby calculating the emotion consistency loss. , .in, It is the cross-entropy loss function; It is a preset emotion classifier. Output the encoding result corresponding to the face image generated in frame t; It is a function that maps the target emotion embedding vector back to the one-hot emotion encoding. Output the encoding result corresponding to the target sentiment embedding vector in frame t.

[0186] It should be noted that the above-mentioned preset emotion classifier The system comprises an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a max-pooling layer, a fully connected layer, a random dropout layer, and an output layer. Specifically, the generated face image is input into a pre-defined emotion classifier through the input layer, and then sequentially passed through four convolutional layers to extract the emotion features of the input data. A max-pooling layer is used for dimensionality reduction to retain key information. The extracted features are fused through a fully connected layer, and the non-linear expression is enhanced using the ReLU activation function. Finally, the output layer, activated by the Softmax activation function, classifies the generated face image into different emotion categories, and the result is output through the output layer.

[0187] The temporal smoothing loss can be calculated by encouraging continuous changes between adjacent frames to prevent facial jitter. , .

[0188] The reconstruction loss can be calculated by encouraging generated images to be as close as possible to preset real face images for fine alignment. , .

[0189] Therefore, it is possible to base this on adversarial loss, identity preservation loss, emotional consistency loss, temporal smoothing loss, and reconstruction loss. The total training loss is calculated. .in, , , , All of these are corresponding preset weights. By introducing multiple losses, the subsequently generated target image retains individual characteristics while also possessing realistic and natural dynamic expressions.

[0190] Therefore, when the target audio file, target emotion embedding vector, and target personality feature vector are input into the preset digital human image generation model for processing, such as... Figure 8 As shown, Figure 8 A flowchart illustrating the steps for generating target images corresponding to various target digital humans, as provided in this application embodiment, includes:

[0191] Step 802: Input the target audio file, target emotion embedding vector, and target personality feature vector into the preset digital human image generation model. Evaluate and filter the current input data through the cold start quality assessment and playback module to obtain qualified input data.

[0192] Step 804: The expression controller module performs frame segmentation and encoding on the target emotion embedding vector in the qualified input data to obtain the emotion control vector that controls the expression state of the image.

[0193] Step 806: The personality adjustment vector is obtained by mapping the target personality feature vector in the qualified input data through the personality adjustment module.

[0194] Step 808: Using the conditional self-adaptor module, the target emotion embedding vector and target personality feature vector in the qualified input data are jointly mapped into a layer-by-layer style code, and the self-adapted conditional style code is output.

[0195] Step 810: The image generator module processes the conditional style code, the emotion control vector and personality adjustment vector in the qualified input data to generate the target image corresponding to each target digital human.

[0196] To address the issues of limited data volume and inconsistent data quality during the cold start phase, the cold start quality assessment and replay module automatically evaluates and filters the current input data during the generation process to obtain qualified input data. Specifically, it first calculates the current comprehensive quality score corresponding to the current input data, then compares this score with pre-set qualified and unqualified thresholds. Current input data exceeding the qualified threshold is retained, while data below the unqualified threshold is discarded. Current input data falling between the qualified and unqualified thresholds is resampled. Optionally, a replay buffer can be pre-set to store high-quality samples for subsequent incremental training, avoiding overfitting caused by insufficient samples.

[0197] The expression controller module performs frame segmentation and encoding on the target emotion embedding vector in the qualified input data, thereby obtaining the emotion control vector that controls the expression state of the image. The emotion control vector can determine the emotional expression of the digital human image at frame t, ensuring that the generated expression changes dynamically with the speech context.

[0198] Next, the personality adjuster module can be used to map the target personality feature vector from the qualified input data to obtain the personality adjustment vector. This personality adjustment vector controls the facial expression style and detailed features of the target digital avatar, ensuring that the generated personalized representation matches the user's personality label. Facial expression style can include facial amplitude, tension, micro-expressions, etc., while detailed features can include clothing style, etc.

[0199] During the cold start phase, due to the extremely limited sample size, directly inputting the target sentiment embedding vector and target personality feature vector into the generator can easily lead to training instability. Therefore, a conditional self-adaptor module is introduced. This module uses a pre-defined conditional fusion network to jointly map the target sentiment embedding vector and target personality feature vector from the qualified input data into layer-by-layer style codes, outputting self-adapted conditional style codes. These self-adapted conditional style codes are used to modulate the feature mappings at each layer of the image generator module, enhancing robustness and generalization ability under limited sample sizes.

[0200] Finally, the image generator module processes the conditional style code, the emotion control vector, and the personality adjustment vector in the qualified input data to generate the target image corresponding to each target digital human. This module can generate a sequence of face frames that are dynamically synchronized with the voice, ensuring that the target image corresponding to the output target digital human remains consistent in both emotion and personality.

[0201] This application provides a digital human image generation method that combines cold-start driven and active learning mechanisms. The method includes: inputting a target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector into a preset few-sample emotion speech generation model for processing, generating target audio files corresponding to each target digital human in the target digital population, wherein the target audio files are consistent with the emotional expression and personality expression of the target digital human; inputting the target audio files, the target emotion embedding vector, and the target personality feature vector into the preset digital human image generation model for processing, generating target images corresponding to each target digital human, wherein the target images are used to reflect facial states consistent with the emotional expression and personality expression in the target audio files; the preset digital human image generation model is constructed based on a conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold-start quality assessment and playback module. This solution processes input data using a pre-defined few-sample emotional speech generation model derived from the cold-start process and possessing high personalization and dynamic adaptability. It enables efficient batch generation of personalized, multi-emotional speech on large-scale unlabeled input data, outputting high-fidelity target audio files with consistent emotional and personality expressions. Furthermore, the pre-defined few-sample emotional speech generation model is trained on the first qualified samples selected by the first cold-start quality evaluator from text-speech pairs of candidate joint personality emotions, allowing for startup under minimal data conditions. Simultaneously, an automated quality evaluation mechanism continuously optimizes the generated samples, expanding the available data space and breaking through the modeling barriers of existing digital human systems. This achieves high-quality, personalized digital human image automatic generation under small-sample conditions. Additionally, by inputting the target emotion embedding vector and target personality feature vector, it achieves voice-driven synchronous expression generation, enhancing the digital human's natural interaction capabilities and style consistency.

[0202] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0203] This application embodiment also provides a digital human image generation device that combines cold start driving and active learning mechanisms, including:

[0204] The first generation module is used to input the target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector into the preset few-sample emotion speech generation model for processing, and generate the target audio file corresponding to each target digital person in the target digital population; wherein, the target audio file is consistent with the emotion expression and personality expression of the target digital person; the preset few-sample emotion speech generation model is obtained by training on the first qualified sample after being screened by the first cold start quality evaluator on the text-speech pair data training sample of the preset joint personality emotion.

[0205] The second generation module is used to input the target audio file, target emotion embedding vector, and target personality feature vector into the preset digital human image generation model for processing, and generate the target image corresponding to each target digital human; wherein, the target image is used to reflect the facial state consistent with the emotion expression and personality expression in the target audio file; the preset digital human image generation model is built on the basis of conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

[0206] Regarding the apparatus in the above embodiments, the specific methods by which each module performs its operations have been described in detail in the embodiments related to the method, and will not be elaborated upon here. Each module in the digital human image generation apparatus combining cold start driving and active learning mechanisms can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations of each module.

[0207] In one embodiment of this application, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0208] The target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector are input into a preset few-shot emotion speech generation model for processing, generating target audio files corresponding to each target digital person in the target digital population; wherein, the target audio files are consistent with the emotion expression and personality expression of the target digital person; the preset few-shot emotion speech generation model is trained based on the first qualified sample after being screened by the first cold start quality evaluator on the text-speech pair data training samples of candidate joint personality emotion;

[0209] The target audio file, target emotion embedding vector, and target personality feature vector are input into the preset digital human image generation model for processing to generate the target image corresponding to each target digital human. The target image is used to reflect the facial state consistent with the emotion and personality expression in the target audio file. The preset digital human image generation model is built on the basis of conditional generative adversarial network and includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

[0210] The computer device provided in this application embodiment has a similar implementation principle and technical effect to the above method embodiment, and will not be described again here.

[0211] In one embodiment of this application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it performs the following steps:

[0212] The target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector are input into a preset few-shot emotion speech generation model for processing, generating target audio files corresponding to each target digital person in the target digital population; wherein, the target audio files are consistent with the emotion expression and personality expression of the target digital person; the preset few-shot emotion speech generation model is trained based on the first qualified sample after being screened by the first cold start quality evaluator on the text-speech pair data training samples of candidate joint personality emotion;

[0213] The target audio file, target emotion embedding vector, and target personality feature vector are input into the preset digital human image generation model for processing to generate the target image corresponding to each target digital human. The target image is used to reflect the facial state consistent with the emotion and personality expression in the target audio file. The preset digital human image generation model is built on the basis of conditional generative adversarial network and includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

[0214] The computer-readable storage medium provided in this embodiment is similar in principle and technical effect to the method embodiment described above, and will not be repeated here.

[0215] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0216] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0217] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A digital human image generation method combining cold start driving and active learning mechanism, characterized in that, The method includes: The target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector are input into a preset few-shot emotion speech generation model for processing, generating target audio files corresponding to each target digital person in the target digital population; wherein, the target audio files are consistent with the emotion expression and personality expression of the target digital person; the preset few-shot emotion speech generation model is trained based on the first qualified samples after being screened by the first cold-start quality evaluator on the text-speech pair data training samples of candidate joint personality emotion; the generation process of the preset few-shot emotion speech generation model includes: The candidate text-speech pair data training samples of joint personality and emotion are input into the first cold-start quality evaluator for processing, and a first comprehensive quality score is calculated by a first preset evaluation formula. The first preset evaluation formula is determined based on a first clarity index, an emotion consistency index, a first diversity index, and a first anomaly confidence level. Based on the first comprehensive quality score and a first preset screening threshold, the first qualified samples are selected, and balanced sampling is performed on the first qualified samples to obtain a training set and a validation set. On the training set, the parameters of the initial emotional speech generation model are updated by maximizing a preset training objective to obtain updated model parameters. The preset training objective is determined based on the Mel spectrogram, text data, emotion embedding vector, and personality feature vector corresponding to the speech data. The initial emotional speech generation model is constructed based on the Tacotron2 conditional speech synthesis model and a conditional variational autoencoder mechanism. The conditional variational autoencoder mechanism fuses the emotion embedding vector and personality feature vector and uses them as conditional variables as input. On the validation set, the updated model parameters are validated by calculating the validation loss, and the target model parameters are obtained by backpropagating the gradient. Based on the target model parameters, the preset few-sample emotional speech generation model is generated. The step of inputting the text-speech pair data training samples of the candidate joint personality emotion into the first cold start quality evaluator for processing, and calculating the first comprehensive quality score through a first preset evaluation formula, includes: The text-to-speech pair training samples of the candidate joint personality and emotion data are input into the first cold-start quality evaluator to obtain the signal-to-noise ratio, speech perception quality score, and speech intelligibility index corresponding to the text-to-speech pair training samples of the candidate joint personality and emotion data. The first clarity index is calculated based on the signal-to-noise ratio, speech perception quality score, and speech intelligibility index. A predicted emotion distribution is obtained based on the emotion Mel spectrogram corresponding to the text-to-speech pair training samples of the candidate joint personality and emotion data. The emotion consistency index is calculated based on the similarity between the predicted emotion distribution and preset emotion labels. The first diversity index is calculated based on the distance between each text-to-speech pair training sample of the candidate joint personality and emotion data. A first probability that the text-to-speech pair training sample of the candidate joint personality and emotion data is abnormally generated is calculated based on the emotion Mel spectrogram and the preset emotion labels. The first anomaly confidence level is obtained based on the first probability. The first comprehensive quality score is calculated based on the first clarity index, the emotion consistency index, the first diversity index, and the first anomaly confidence level. The target audio file, the target emotion embedding vector, and the target personality feature vector are input into a preset digital human image generation model for processing to generate target images corresponding to each target digital human; wherein, the target image is used to reflect facial states consistent with the emotion and personality expressions in the target audio file; the preset digital human image generation model is constructed based on a conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

2. The method of claim 1, wherein, The step of validating the updated model parameters by calculating the validation loss and obtaining the target model parameters by backpropagating the gradient includes: The updated model parameters are validated, and the validation loss is calculated. The target model parameters are obtained by backpropagating the gradient based on the validation loss, meta-learning rate, loss term of the score obtained by the first cold start quality evaluator, and preset adjustment weights, and updating the global model initialization parameters in conjunction with the cold start iteration mechanism.

3. The method according to claim 1 or 2, characterized in that, The generation process of the preset digital human image generation model includes: Obtain a candidate training sample set; wherein, the candidate training sample set includes multiple candidate training samples; The candidate training sample set is input into the second cold start quality evaluator for processing, and the second comprehensive quality score is calculated by the second preset evaluation formula; wherein, the second preset evaluation formula is determined based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence. Based on the second comprehensive quality score and the second preset screening threshold, a second qualified sample is obtained. The second qualified sample is input into the initial image generation model for training, and the total training loss is calculated; the total training loss includes adversarial loss, identity preservation loss, emotion consistency loss, temporal smoothing loss, and reconstruction loss; The model parameters are updated based on the total training loss, and the preset digital human image generation model is generated.

4. The method of claim 3, wherein, The step of inputting the candidate training sample set into the second cold start quality evaluator for processing, and calculating the second comprehensive quality score using a second preset evaluation formula, includes: The candidate training sample set is input into the second cold start quality evaluator to obtain the no-reference image quality evaluation index and Laplacian variance corresponding to the candidate training sample, and the second sharpness index is calculated based on the no-reference image quality evaluation index and Laplacian variance. The semantic consistency index is calculated based on the similarity between the first predicted identity embedding vector corresponding to the candidate training sample and the second predicted identity embedding vector corresponding to the original face image. The second diversity index is calculated based on the distance between each of the candidate training samples; Based on the candidate training sample and the second predicted identity embedding vector, a second probability that the candidate training sample is generated abnormally is calculated, and a second abnormality confidence is obtained based on the second probability. The second comprehensive quality score is calculated based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.

5. The method according to claim 1 or 2, characterized in that, The step of inputting the target audio file, the target emotion embedding vector, and the target personality feature vector into a preset digital human image generation model for processing, and generating the target image corresponding to each target digital human, includes: The target audio file, the target emotion embedding vector, and the target personality feature vector are input into a preset digital human image generation model. The cold start quality assessment and playback module evaluates and filters the current input data to obtain qualified input data. The expression controller module performs frame segmentation and encoding on the target emotion embedding vector in the qualified input data to obtain the emotion control vector that controls the expression state of the image. The personality adjustment module maps the target personality feature vector in the qualified input data to obtain a personality adjustment vector; wherein, the personality adjustment vector is used to control the facial expression style and detailed features of the target digital human. The conditional self-adaptor module jointly maps the target emotion embedding vector and target personality feature vector in the qualified input data into a layer-by-layer style code, and outputs the self-adapted conditional style code; wherein, the self-adapted conditional style code is used to modulate the feature mapping of each layer of the image generator module. The image generator module processes the conditional style code, the emotion control vector, and the personality adjustment vector in the qualified input data to generate the target image corresponding to each target digital human.

6. A digital human image generation device combining cold start driving and active learning mechanisms, characterized in that, The device includes: The first generation module is configured to input a target digital population, target text content corresponding to the target digital population, a target emotion embedding vector, and a target personality feature vector into a preset few-sample emotion speech generation model for processing to generate a target audio file corresponding to each target digital person in the target digital population; wherein the target audio file is consistent with the emotional expression and personality expression of the target digital person; the preset few-sample emotion speech generation model is trained based on first qualified samples selected by a first cold start quality evaluator from text-speech pair data training samples of candidate joint personality emotions; the generation process of the preset few-sample emotion speech generation model includes: inputting the text-speech pair data training samples of the candidate joint personality emotions into the first cold start quality evaluator for processing, and calculating a first comprehensive quality score by a first preset evaluation formula; wherein the first preset evaluation formula is determined based on a first clarity index, an emotional consistency index, a first diversity index, and a first abnormal confidence; based on the first comprehensive quality score and a first preset screening threshold, the first qualified samples are screened and balanced sampling is performed on the first qualified samples to obtain a training set and a validation set; on the training set, the parameters of an initial emotion speech generation model are updated by maximizing a preset training target to obtain updated model parameters; wherein the preset training target is determined based on a mel-spectrogram corresponding to speech data, text data, an emotion embedding vector, and a personality feature vector, the initial emotion speech generation model is constructed based on a Tacotron2 conditional speech synthesis model and a conditional variational autoencoder mechanism, and the conditional variational autoencoder mechanism inputs the emotion embedding vector and the personality feature vector after fusion as a conditional variable; on the validation set, the updated model parameters are verified by calculating a validation loss, and target model parameters are obtained by backpropagation gradient; based on the target model parameters, the preset few-sample emotion speech generation model is generated; the inputting of the text-speech pair data training samples of the candidate joint personality emotions into the first cold start quality evaluator for processing and the calculation of the first comprehensive quality score by the first preset evaluation formula include: inputting the text-speech pair data training samples of the candidate joint personality emotions into the first cold start quality evaluator, obtaining a signal-to-noise ratio, a speech perceptual quality score, and a speech intelligibility index corresponding to the text-speech pair data training samples of the candidate joint personality emotions, calculating the first clarity index based on the signal-to-noise ratio, the speech perceptual quality score, and the speech intelligibility index; obtaining a predicted emotion distribution based on an emotion mel-spectrogram corresponding to the text-speech pair data training samples of the candidate joint personality emotions, and calculating the emotional consistency index based on the similarity of the predicted emotion distribution and a preset emotion label; calculating the first diversity index based on the distance between each of the text-speech pair data training samples of the candidate joint personality emotions;The text-speech pair data training sample of the candidate joint personality emotion is calculated based on the emotional Mel spectrum diagram and the preset emotion label to obtain a first probability of abnormal generation, and a first abnormal confidence is obtained based on the first probability; and the first comprehensive quality score is calculated based on the first clarity index, the emotional consistency index, the first diversity index, and the first abnormal confidence. The second generation module is used to input the target audio file, the target emotion embedding vector, and the target personality feature vector into a preset digital human image generation model for processing, and generate target images corresponding to each target digital human; wherein, the target image is used to reflect facial states consistent with the emotion expression and personality expression in the target audio file; the preset digital human image generation model is constructed based on a conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality assessment and playback module.

7. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the digital human image generation method combining cold start driving and active learning mechanisms as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the digital human image generation method combining cold start driving and active learning mechanisms as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Digital human animation generation method and device and digital human animation generation model

    CN118799461A

  • Systems and methods for personalized generalized content recommendations

    US20130290110A1