Digital human image generation method combining cold start driving and active learning mechanism
By combining cold start-driven and active learning mechanisms, the problem of high-quality personalization in digital human image generation under data scarcity conditions was solved. It achieved the generation of high-fidelity digital human audio and images with consistent emotional and personality expression with very little data, and enhanced the natural interaction capabilities of digital humans.
Patent Information
- Application Number
- CN202511324476.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing methods for generating digital human images struggle to achieve high quality, strong consistency, and strong individual expression capabilities under conditions of data scarcity and user cold start, failing to meet the needs of user customization and emotion-driven behavior.
This approach combines cold-start driven and active learning mechanisms. By pre-setting a few-sample emotional speech generation model and a digital human image generation model, target audio files and images are generated. A cold-start quality evaluator is used to screen samples, and a model is built through a conditional generative adversarial network to achieve consistency in emotional and personality expression.
Generating high-fidelity target audio files and images with consistent emotional and personality expression under extremely limited data conditions breaks through the modeling barriers of digital human systems and enhances the natural interaction capabilities and style consistency of digital humans.
Smart Images

Figure CN120833401A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a digital figure generation method combining cold start driving and active learning mechanism. BACKGROUND
[0002] At present, with the wide rise of virtual reality, augmented reality, digital people, intelligent customer service, virtual anchors and other applications, the demand for digital people with high personalization, emotional expression and voice expression synchronization ability is rapidly growing in the community.
[0003] At present, the existing digital figure generation method mostly relies on a large amount of manual annotation data, and only a small amount of information can be provided when the user uses it for the first time, which is difficult to adapt to the needs of user personalization and emotional driving. Therefore, how to build a digital figure generation method that can realize high-quality, strong consistency and strong personal expression ability under the conditions of data scarcity and user cold start has become a problem to be solved. SUMMARY
[0004] The present application aims to at least solve the technical problems existing in the prior art. To this end, the first aspect of the present application proposes a digital figure generation method combining cold start driving and active learning mechanism, which comprises: inputting the target digital people group, the target text content corresponding to the target digital people group, the target emotion embedding vector and the target personality feature vector into a preset few-sample emotional speech generation model for processing to generate target audio files corresponding to each target digital person in the target digital people group; wherein the target audio file is consistent with the emotional expression and personality expression of the target digital person; the preset few-sample emotional speech generation model is trained based on the first qualified sample selected by the first cold start quality evaluator from the text-speech pair data training sample of the candidate joint personality emotion; inputting the target audio file, the target emotion embedding vector and the target personality feature vector into a preset digital figure generation model for processing to generate target figures corresponding to each target digital person; wherein the target figure is used to reflect the facial state consistent with the emotional expression and personality expression in the target audio file; the preset digital figure generation model is constructed based on a conditional generative adversarial network, and the preset digital figure generation model comprises an expression controller module, a personality adjuster module, a conditional self-adaptor module, an image generator module and a cold start quality evaluation and playback module.
[0005] In one possible implementation, the generation process of the preset few-sample emotional speech generation model comprises: The text-speech pair data training sample of the candidate joint personality emotion is input into the first cold start quality evaluator for processing, and a first comprehensive quality score is calculated by a first preset evaluation formula; wherein the first preset evaluation formula is determined based on a first intelligibility index, an emotion consistency index, a first diversity index, and a first abnormal confidence; Based on the first comprehensive quality score and the first preset screening threshold, a first qualified sample is screened, and the first qualified sample is balanced sampled to obtain a training set and a verification set; On the training set, the parameters of the initial emotion speech generation model are updated by maximizing the preset training target to obtain updated model parameters; wherein the preset training target is determined based on the mel-spectrogram corresponding to the speech data, the text data, the emotion embedding vector, and the personality feature vector, the initial emotion speech generation model is constructed based on the Tacotron2 conditional speech synthesis model and the conditional variational autoencoder mechanism, and the emotion embedding vector and the personality feature vector are fused as a conditional variable for input; On the verification set, the updated model parameters are verified by calculating the verification loss, and the target model parameters are obtained by backpropagation gradient; Based on the target model parameters, a preset few-sample emotion speech generation model is generated.
[0006] In one possible implementation, the text-speech pair data training sample of the candidate joint personality emotion is input into the first cold start quality evaluator for processing, and a first comprehensive quality score is calculated by a first preset evaluation formula, including: The text-speech pair data training sample of the candidate joint personality emotion is input into the first cold start quality evaluator, the signal-to-noise ratio, the speech perceptual quality score, and the speech intelligibility index corresponding to the text-speech pair data training sample of the candidate joint personality emotion are obtained, and the first intelligibility index is calculated based on the signal-to-noise ratio, the speech perceptual quality score, and the speech intelligibility index; The predicted emotion distribution is obtained based on the emotion mel-spectrogram corresponding to the text-speech pair data training sample of the candidate joint personality emotion, and the emotion consistency index is calculated based on the similarity of the predicted emotion distribution and the preset emotion label; The first diversity index is calculated based on the distance between each text-speech pair data training sample of the candidate joint personality emotion; The first probability of the text-speech pair data training sample of the candidate joint personality emotion being abnormally generated is calculated based on the emotion mel-spectrogram and the preset emotion label, and the first abnormal confidence is obtained based on the first probability; The first comprehensive quality score is calculated based on the first intelligibility index, the emotion consistency index, the first diversity index, and the first abnormal confidence.
[0007] In a possible implementation, the updated model parameters are verified by calculating a verification loss, and the target model parameters are obtained by backpropagating the gradient, comprising: verifying the updated model parameters to obtain a verification loss; backpropagating the gradient based on the verification loss, the meta learning rate, the loss term of the score obtained by the first cold start quality evaluator, and a preset adjustment weight, and updating the global model initialization parameters in combination with a cold start iteration mechanism to obtain the target model parameters.
[0008] In a possible implementation, the generation process of the preset digital human image generation model comprises: obtaining a candidate training sample set; wherein the candidate training sample set comprises a plurality of candidate training samples; inputting the candidate training sample set into a second cold start quality evaluator for processing to obtain a second comprehensive quality score by a second preset evaluation formula; wherein the second preset evaluation formula is determined based on a second clarity index, a semantic consistency index, a second diversity index, and a second anomaly confidence; obtaining a second qualified sample based on the second comprehensive quality score and a second preset screening threshold; inputting the second qualified sample into an initial image generation model for training to obtain a total training loss; the total training loss comprises an adversarial loss, an identity preservation loss, an emotion consistency loss, a temporal smoothing loss, and a reconstruction loss; updating the model parameters based on the total training loss to generate the preset digital human image generation model.
[0009] In a possible implementation, the candidate training sample set is inputted into the second cold start quality evaluator for processing to obtain the second comprehensive quality score by the second preset evaluation formula, comprising: inputting the candidate training sample set into the second cold start quality evaluator to obtain a no-reference image quality evaluation index and a Laplacian variance corresponding to the candidate training sample, and calculating the second clarity index based on the no-reference image quality evaluation index and the Laplacian variance; calculating the semantic consistency index based on the similarity between the first predicted identity embedding vector corresponding to the candidate training sample and the second predicted identity embedding vector corresponding to the original human face image; calculating the second diversity index based on the distance between each candidate training sample; calculating a second probability that the candidate training sample is generated abnormally based on the candidate training sample and the second predicted identity embedding vector, and obtaining the second anomaly confidence based on the second probability; The second comprehensive quality score is calculated based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.
[0010] In a possible implementation, the target audio file, the target emotion embedding vector, and the target personality feature vector are input into a preset digital human image generation model for processing to generate target images corresponding to each target digital human, including: The target audio file, the target emotion embedding vector, and the target personality feature vector are input into a preset digital human image generation model, and the current input data is evaluated and filtered through a cold start quality evaluation and playback module to obtain qualified input data. The target emotion embedding vector in the qualified input data is frame-divided and encoded by an expression controller module to obtain an emotion control vector for controlling the expression state of the image; The target personality feature vector in the qualified input data is mapped by a personality adjuster module to obtain a personality adjustment vector; the personality adjustment vector is used to control the expression style and detailed features of the target digital human; The target emotion embedding vector and the target personality feature vector in the qualified input data are jointly mapped into a layer-by-layer style code by a conditional adapter module, and an adapted conditional style code is output; the adapted conditional style code is used to modulate the feature mapping of each layer of an image generator module; The conditional style code, the emotion control vector, and the personality adjustment vector in the qualified input data are processed by an image generator module to generate target images corresponding to each target digital human.
[0011] The second aspect of the application proposes a digital human image generation device combining a cold start driving mechanism and an active learning mechanism, which includes: A first generation module is configured to input a target digital human group, target text content corresponding to the target digital human group, a target emotion embedding vector, and a target personality feature vector into a preset few-sample emotion speech generation model for processing to generate target audio files corresponding to each target digital human in the target digital human group; the target audio file is consistent with the emotional expression and personality expression of the target digital human; the preset few-sample emotion speech generation model is trained based on first qualified samples selected by a first cold start quality evaluator from text-speech pair data training samples of a preset joint personality emotion. The second generation module is configured to input the target audio file, the target emotion embedding vector, and the target personality feature vector into a preset digital human image generation model for processing to generate a target image corresponding to each target digital human. The target image is used to reflect a facial state consistent with the expression of emotion and personality in the target audio file. The preset digital human image generation model is constructed based on a conditional generative adversarial network. The preset digital human image generation model includes an expression controller module, a personality adjuster module, a conditional self-adaptor module, an image generator module, and a cold start quality evaluation and playback module.
[0012] In a possible implementation, the apparatus is further configured to: input the text-speech pair data training sample of the candidate joint personality emotion into the first cold start quality evaluator for processing, and calculate a first comprehensive quality score by using a first preset evaluation formula. The first preset evaluation formula is determined based on a first intelligibility index, an emotion consistency index, a first diversity index, and a first abnormal confidence. based on the first comprehensive quality score and a first preset screening threshold, screen a first qualified sample, and perform balanced sampling on the first qualified sample to obtain a training set and a verification set; on the training set, update the parameters of an initial emotion speech generation model by maximizing a preset training target to obtain updated model parameters. The preset training target is determined based on a mel spectrogram corresponding to the speech data, the text data, the emotion embedding vector, and the personality feature vector. The initial emotion speech generation model is constructed based on a Tacotron2 conditional speech synthesis model and a conditional variational autoencoder mechanism. The conditional variational autoencoder mechanism inputs the emotion embedding vector and the personality feature vector after fusion as a conditional variable. on the verification set, verify the updated model parameters by calculating a verification loss, and obtain target model parameters by backpropagation of gradients; based on the target model parameters, generate a preset few-sample emotion speech generation model.
[0013] In a possible implementation, the apparatus is further configured to: input the text-speech pair data training sample of the candidate joint personality emotion into the first cold start quality evaluator, obtain a signal-to-noise ratio, a speech perceptual quality score, and a speech intelligibility index corresponding to the text-speech pair data training sample of the candidate joint personality emotion, and calculate a first intelligibility index based on the signal-to-noise ratio, the speech perceptual quality score, and the speech intelligibility index. based on the emotion mel spectrogram corresponding to the text-speech pair data training sample of the candidate joint personality emotion, obtain a predicted emotion distribution, and calculate an emotion consistency index based on the similarity between the predicted emotion distribution and a preset emotion label. a first diversity index is calculated based on the distance between the text-speech pair data training samples of each candidate joint personality emotion; a first probability that the text-speech pair data training samples of the candidate joint personality emotion are abnormally generated is calculated based on the emotion mel-spectrum graph and the preset emotion label, and a first abnormal confidence is obtained based on the first probability; a first comprehensive quality score is calculated based on the first intelligibility index, the emotional consistency index, the first diversity index, and the first abnormal confidence.
[0014] In a possible implementation, the apparatus is further configured to: verify the updated model parameters to obtain a verification loss; a gradient is back-propagated based on the verification loss, the meta-learning rate, a loss term of the score obtained by the first cold start quality evaluator, and a preset adjustment weight, and the global model initialization parameters are updated in combination with a cold start iteration mechanism to obtain target model parameters.
[0015] In a possible implementation, the apparatus is further configured to: obtain a candidate training sample set; wherein the candidate training sample set includes a plurality of candidate training samples; input the candidate training sample set into a second cold start quality evaluator for processing, and obtain a second comprehensive quality score by a second preset evaluation formula; wherein the second preset evaluation formula is determined based on a second intelligibility index, a semantic consistency index, a second diversity index, and a second abnormal confidence; obtain a second qualified sample based on the second comprehensive quality score and a second preset screening threshold; input the second qualified sample into an initial avatar generation model for training to obtain a total training loss; the total training loss includes an adversarial loss, an identity preservation loss, an emotional consistency loss, a temporal smoothing loss, and a reconstruction loss; update the model parameters based on the total training loss to generate a preset digital human avatar generation model.
[0016] In a possible implementation, the apparatus is further configured to: input the candidate training sample set into a second cold start quality evaluator to obtain a no-reference image quality evaluation index and a Laplacian variance corresponding to the candidate training sample, and calculate a second intelligibility index based on the no-reference image quality evaluation index and the Laplacian variance; calculate a semantic consistency index based on the similarity between the first predicted identity embedding vector corresponding to the candidate training sample and the second predicted identity embedding vector corresponding to the original face image; calculate a second diversity index based on the distance between each candidate training sample; a second probability that the candidate training sample is generated as an anomaly is calculated based on the candidate training sample and the second predicted identity embedding vector, and a second anomaly confidence is obtained based on the second probability; a second comprehensive quality score is calculated based on the second definition index, the semantic consistency index, the second diversity index, and the second anomaly confidence.
[0017] In a possible implementation, the second generation module is specifically configured to: input the target audio file, the target emotion embedding vector, and the target personality feature vector into a preset digital human image generation model, evaluate and filter the current input data through the cold start quality evaluation and playback module to obtain qualified input data; frame and encode the target emotion embedding vector in the qualified input data through the expression controller module to obtain an emotion control vector for controlling the image expression state; map the target personality feature vector in the qualified input data through the personality adjuster module to obtain a personality adjustment vector; wherein the personality adjustment vector is used to control the expression style and detail features of the target digital human; jointly map the target emotion embedding vector and the target personality feature vector in the qualified input data into a layer-by-layer style code through the conditional self-adaptor module, and output the self-adapted conditional style code; wherein the self-adapted conditional style code is used to modulate the feature mapping of each layer of the image generator module; process the conditional style code, the emotion control vector, and the personality adjustment vector in the qualified input data through the image generator module to generate the target image corresponding to each target digital human.
[0018] The third aspect of the present application proposes an electronic device, which includes a processor and a memory, and the memory stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to realize the digital human image generation method combining the cold start driving and the active learning mechanism as described in the first aspect.
[0019] The fourth aspect of the present application proposes a computer-readable storage medium, which stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to realize the digital human image generation method combining the cold start driving and the active learning mechanism as described in the first aspect.
[0020] The embodiments of the present application have the following beneficial effects: The digital human image generation method provided by the embodiment of the present application combines cold start driving and active learning mechanism, and the method comprises the following steps: inputting a target digital human group, target text content corresponding to the target digital human group, a target emotion embedding vector and a target personality feature vector into a preset few-sample emotional speech generation model for processing to generate a target audio file corresponding to each target digital human in the target digital human group, wherein the target audio file is consistent with the emotional expression and personality expression of the target digital human; inputting the target audio file, the target emotion embedding vector and the target personality feature vector into a preset digital human image generation model for processing to generate a target image corresponding to each target digital human, wherein the target image is used to reflect a facial state consistent with the emotional expression and personality expression in the target audio file; the preset digital human image generation model is constructed based on a conditional generative adversarial network; and the preset digital human image generation model comprises an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module and a cold start quality evaluation and playback module. The preset few-sample emotional speech generation model derived from the cold start process and having high personalization and dynamic adaptability is used to process the input data, so that efficient personalized and multi-emotional speech batch generation can be realized on large-scale unlabeled input data, thereby outputting a target audio file with high fidelity and consistent emotional and personal expression; and the preset few-sample emotional speech generation model is trained based on the first qualified sample selected by the first cold start quality evaluator from the text-speech pair data training sample of the candidate joint personality emotion, so that the model can be started under the condition of very few data, and the automatic quality evaluation mechanism is used to continuously optimize the generated sample, expand the available data space, break through the modeling barrier of the existing digital human system, and realize high-quality personalized digital human image automatic generation under the condition of small sample; in addition, by inputting the target emotion embedding vector and the target personality feature vector, synchronous expression generation under speech driving is realized, and the natural interaction ability and style consistency of the digital human are enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 A block diagram of a computer device provided by the embodiment of the present application; Figure 2 A step flowchart of a digital human image generation method combining cold start driving and active learning mechanism provided by the embodiment of the present application; Figure 3 A step flowchart of generating a preset few-sample emotional speech generation model provided by the embodiment of the present application; Figure 4 A step flowchart of calculating a first comprehensive quality score provided by the embodiment of the present application; Figure 5 A step flowchart of obtaining a target model parameter provided by the embodiment of the present application; Figure 6 A flowchart of the steps for constructing a preset digital human image generation model provided in an embodiment of the present application; Figure 7 A flowchart of the steps for calculating a second comprehensive quality score provided in an embodiment of the present application; Figure 8 A flowchart of the steps for generating a target image corresponding to each target digital human provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0023] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "multiple" means two or more. In addition, the use of "based on" or "according to" means openness and inclusiveness, because the process, steps, calculations or other actions "based on" or "according to" one or more of the conditions or values may be based on additional conditions or values beyond the stated in practice.
[0024] The digital human image generation method combining cold start drive and active learning mechanism provided in this application can be applied to computer devices (electronic devices). The computer device can be a server or a terminal. The server can be a single server or a server cluster composed of multiple servers. The embodiments of this application do not make specific restrictions on this. The terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices.
[0025] Take the computer device as an example, Figure 1 A block diagram of a server is shown, such as Figure 1As shown, the server can include a processor and a memory connected through a system bus. Among them, the processor of the server is used to provide computing and control capabilities. The memory of the server includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The computer program is executed by the processor to implement a digital human image generation method combining cold start driving and active learning mechanism.
[0026] Those skilled in the art can understand that, Figure 1 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the server to which the scheme of the present application is applied. Alternatively, the server can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0027] Figure 2 A step flow chart of a digital human image generation method combining cold start driving and active learning mechanism provided by an embodiment of the present application, the method comprising the following steps: Step 202, input the target digital human group, the target text content corresponding to the target digital human group, the target emotion embedding vector and the target personality feature vector into a preset few-sample emotion speech generation model for processing, to generate target audio files corresponding to each target digital human in the target digital human group.
[0028] Among them, the target audio file is consistent with the emotional expression and personality expression of the target digital human, and the preset few-sample emotion speech generation model is trained based on the first qualified sample selected by the first cold start quality evaluator from the text-speech pair data training sample of the candidate joint personality emotion.
[0029] As Figure 3 shown, Figure 3 A step flow chart of generating a preset few-sample emotion speech generation model provided by an embodiment of the present application, comprising: Step 302, input the text-speech pair data training sample of the candidate joint personality emotion into the first cold start quality evaluator for processing, and calculate the first comprehensive quality score by the first preset evaluation formula.
[0030] Among them, the first preset evaluation formula is determined based on the first intelligibility index, the emotional consistency index, the first diversity index, and the first abnormal confidence.
[0031] In some optional embodiments, as Figure 4 shown, Figure 4 A step flow chart of calculating the first comprehensive quality score provided by an embodiment of the present application, comprising: Step 402, input the text-speech pair data training sample of the candidate joint personality emotion into the first cold start quality evaluator, obtain the signal-to-noise ratio, speech perceptual quality score and speech intelligibility index corresponding to the text-speech pair data training sample of the candidate joint personality emotion, and calculate the first clarity index based on the signal-to-noise ratio, speech perceptual quality score and speech intelligibility index.
[0032] Step 404, obtain the predicted emotion distribution based on the emotion mel-spectrogram corresponding to the text-speech pair data training sample of the candidate joint personality emotion, and calculate the emotion consistency index based on the predicted emotion distribution and the similarity of the preset emotion label.
[0033] Step 406, calculate the first diversity index based on the distance between each text-speech pair data training sample of the candidate joint personality emotion.
[0034] Step 408, calculate the first probability that the text-speech pair data training sample of the candidate joint personality emotion is abnormally generated based on the emotion mel-spectrogram and the preset emotion label, and obtain the first abnormal confidence based on the first probability.
[0035] Step 410, calculate the first comprehensive quality score based on the first clarity index, the emotion consistency index, the first diversity index and the first abnormal confidence.
[0036] Wherein, when obtaining the text-speech pair data training sample of the candidate joint personality emotion, the emotion speech sample set corresponding to the target digital human group, the personality type and interest label corresponding to the target digital human group can be obtained; the original text and emotion label are encoded to obtain the encoding result, and the encoding result is processed through the preset diffusion model to generate a mel-spectrum graph; the mel-spectrum graph is input into the preset neural vocoder to restore the time-domain speech waveform, and the multi-emotion text-speech pair enhancement data is generated; after the personality type and interest label are encoded by the preset one-hot encoding algorithm, the encoding result is obtained, and after the encoding result is sequentially spliced and normalized, the personality enhancement data is generated; based on the multi-emotion text-speech pair enhancement data and the personality enhancement data, the text-speech pair data training sample of the candidate joint personality emotion is generated.
[0037] Specifically, the emotion speech sample set includes an original speech sample, an original text and an emotion label corresponding to the original speech sample. The original speech sample can be a speech waveform in pulse code modulation (PCM) format, the data type of the original text is a string, the data type of the emotion label is a string, and the four emotion labels of "joy, anger, sadness and happiness" are supported by default.
[0038] Then, the emotional speech sample set can be data enhanced, and the process is implemented based on a lightweight diffusion model Grad-TTS. The Grad-TTS is a speech generation model based on a diffusion process, which supports high-fidelity speech synthesis and enhancement under conditional control. It is used to enhance the number of speeches and rich emotional features from a small amount of original emotional speech samples, and to provide diversified training samples for subsequent speech synthesis.
[0039] Optionally, the original text and the emotion label can be encoded to obtain an encoding result. For example, if the original text is "Hello!", the original text is converted into a text embedding sequence to obtain an encoding result of [13, 27, 45]. If the emotion label is "happy", the encoding result obtained after converting the emotion label into a one-hot encoding can be [1, 0, 0, 0].
[0040] Thus, a conditional control vector can also be constructed, i.e., only emotional control is considered at present, so the conditional control vector is the encoding result corresponding to the emotion label. Further, a conditional diffusion generation process can be constructed based on the Grad-TTS model. For each original speech sample, under the corresponding fixed original text and emotion label, a preset diffusion model is used to generate a speech spectrum representation with different styles, i.e., a preset joint personality emotional text-speech pair data training sample.
[0041] Specifically, a random noise vector can be initialized first, and then the text embedding sequence and the conditional control vector are used as conditions to perform conditional diffusion generation using the Grad-TTS model to obtain a mel spectrogram. The mel spectrogram has consistent semantics, specified emotions, and speech feature representations with style differences. Then, the mel spectrogram can be input into a preset neural vocoder to restore the time-domain speech waveform, thereby outputting multi-emotional text-speech pair enhancement data.
[0042] In addition, for personality types and interest labels, a preset one-hot encoding algorithm can be used to encode the personality types and interest labels to obtain an encoding result, and the encoding result can be sequentially spliced and normalized to generate personality enhancement data. Finally, based on the multi-emotional text-speech pair enhancement data and the personality enhancement data, a candidate joint personality emotional text-speech pair data training sample can be generated.
[0043] After obtaining the text-speech pair data training sample of the preset joint personality emotion, the text-speech pair data training sample of the candidate joint personality emotion can be input into the first cold start quality evaluator to obtain the signal-to-noise ratio, the speech perceptual quality score and the speech intelligibility index corresponding to the text-speech pair data training sample of the candidate joint personality emotion. The first intelligibility index is calculated based on the signal-to-noise ratio, the speech perceptual quality score and the speech intelligibility index. The first intelligibility index is used to measure whether the generated speech is clear and free of artifacts. The specific calculation process is shown in formula (1).
[0044] (1) wherein, SNR represents the signal-to-noise ratio, the higher the signal-to-noise ratio, the clearer; PESQ represents the speech perceptual quality score; STOI represents the speech intelligibility index; represents the first intelligibility index. , , respectively represent the corresponding weight.
[0045] The predicted emotion distribution can be obtained based on the emotion mel-spectrogram corresponding to the text-speech pair data training sample of the candidate joint personality emotion. The emotion consistency index is calculated based on the predicted emotion distribution and the similarity of the preset emotion label. The emotion consistency index is used to measure whether the emotion features of the generated speech are consistent with the input preset emotion label. The specific calculation process is shown in formula (2).
[0046] (2) wherein, represents the emotion consistency index; represents the cosine similarity; represents a small emotion classifier, which is trained by self-supervision combined with a small amount of labeled samples, and outputs the predicted emotion distribution; represents the preset emotion label.
[0047] It should be noted that the above small emotion classifier The input layer, the first convolutional layer, the second convolutional layer, the pooling layer, the bidirectional long short-term memory (LSTM) layer, the fully connected layer, the random dropout layer, and the output layer are included. Specifically, the local features of the input data are extracted through two convolutional layers, and then down-sampling is performed through the pooling layer. Then, the time sequence information is captured through the bidirectional LSTM layer to enhance the perception ability of the model to the emotional changes. The extracted features are further fused through the fully connected layer, and the non-linear expression is enhanced through the rectified linear unit (ReLU) activation function. In order to prevent overfitting, the small-sized sentiment classifier also introduces the random dropout layer, and finally the output layer of the Softmax activation function is used to classify the emotional categories.
[0048] Based on the distance between the text-speech pair data training samples of each candidate joint personality emotion, a first diversity index is calculated, which aims to calculate the distance between different samples, and the greater the difference, the better the diversity, ensuring that the enhanced speech of different samples has differences in style and avoiding mode collapse. The specific calculation process is shown in formula (3).
[0049] (3) Wherein, represents the first diversity index; , represents the text-speech pair data training sample of different candidate joint personality emotion; K represents the total number of samples.
[0050] Based on the emotional mel-spectrum graph and the preset emotional label, a first probability that the text-speech pair data training sample of the candidate joint personality emotion is abnormally generated is calculated, and a first abnormal confidence is obtained based on the first probability. The lightweight discriminator trained by contrast learning can be used to identify whether the generated spectrum is abnormal. The specific calculation process is shown in formula (4).
[0051] (4) Wherein, represents the first abnormal confidence; represents the emotional mel-spectrum graph; represents the preset emotional label.
[0052] It should be noted that the lightweight discriminator The input layer, the first convolutional layer, the second convolutional layer, the pooling layer, the flattening layer, the fully connected layer, the discriminant layer, and the output layer are included. Specifically, the local features of the input data are extracted through two convolutional layers, and the feature dimension is reduced through the pooling layer. Then, the flattened layer transmits the feature vector after the pooling to the fully connected layer for feature fusion. Finally, the discriminant layer uses the Sigmoid activation function to output the probability of whether the sample is abnormal, and the result is output through the output layer.
[0053] Finally, the first comprehensive quality score can be calculated based on the first clarity index, the sentiment consistency index, the first diversity index, and the first abnormal confidence. The specific calculation process is shown in formula (5).
[0054] (5) Wherein, represents the first comprehensive quality score; , , , respectively represent the corresponding weight.
[0055] It should be noted that the first cold start quality evaluator is equivalent to the first comprehensive quality score , that is, the weighted sum of the first clarity index, the sentiment consistency index, the first diversity index, and the first abnormal confidence.
[0056] Step 304, based on the first comprehensive quality score and the first preset screening threshold, the first qualified sample is screened, and the first qualified sample is balanced sampled to obtain the training set and the verification set.
[0057] Wherein, the first comprehensive quality score and the first preset screening threshold can be compared, if the first comprehensive quality score is less than the first preset screening threshold, the corresponding sample is discarded, otherwise, it is taken as the first qualified sample.
[0058] Then, the first qualified sample can be balanced sampled to obtain the training set and the verification set to perform the subsequent model training process. The trained model can quickly and actively learn the emotional speech style of individual users under very few samples, complete the personal transfer of new emotions, and generate personalized speech with multiple emotions.
[0059] Step 306, on the training set, the parameters of the initial emotional speech generation model are updated by maximizing the preset training target to obtain the updated model parameters.
[0060] The preset training target is determined based on a mel spectrogram corresponding to voice data, text data, an emotion embedding vector and a personality feature vector, and the initial emotion voice generation model is constructed based on a Tacotron2 conditional voice synthesis model and a conditional variational autoencoder mechanism. The conditional variational autoencoder mechanism inputs the emotion embedding vector and the personality feature vector after fusion as a conditional variable.
[0061] Specifically, optionally, in the model training process, a Tacotron2-based conditional voice synthesis model can be constructed first, and the emotion and personality vectors are fused to be input as a conditional variable into an encoder and a decoder of the model. By introducing the conditional variational autoencoder mechanism, the model learns how to generate accurate and style-consistent voice under the condition of given semantic and style features.
[0062] The preset training target is determined based on voice data, text data, an emotion embedding vector and a personality feature vector, and can be specifically represented as: , wherein represents voice data; represents corresponding text data; , represents an emotion embedding vector, represents a personality feature vector; represents a latent variable vector, sampled from a standard normal distribution; is an approximate posterior distribution, output by the encoder; represents a prior distribution of the latent variable, usually set as a standard normal distribution. is a reconstruction error term, representing whether the current model can restore the real voice data under the posterior distribution of the latent variable . is a relative entropy regularization term, representing the distance between the approximate posterior distribution output by the model and the prior distribution (usually a standard normal distribution ), which can prevent the latent space from degenerating.
[0063] In the training process, the Few-shot mechanism is implemented, which enables the model to quickly adapt to new tasks without or with only a small amount of training data. In this learning mode, the model does not need a large amount of training data, but uses its powerful learning mechanism to complete the task with little or even no sample guidance. Thus, the first qualified sample can be balanced sampled to obtain a training set and a validation set. On the training set, the parameters of the initial emotion voice generation model are updated by maximizing the preset training target to obtain updated model parameters , wherein, represents initializing model parameters, represents a preset learning rate of an inner loop training process, represents a preset training objective on a training set .
[0064] Step 308, on a validation set, the updated model parameters are verified by calculating a validation loss, and the target model parameters are obtained by backpropagation of the gradient.
[0065] In some optional embodiments, as Figure 5 shown, Figure 5 is a step flowchart for obtaining target model parameters provided by the embodiments of the present application, comprising: Step 502, the updated model parameters are verified, and a validation loss is calculated.
[0066] Step 504, the gradient is backpropagated based on the validation loss, the meta learning rate, the loss term of the score obtained by the first cold start quality evaluator, and a preset adjustment weight, and the global model initialization parameters are updated in combination with a cold start iteration mechanism to obtain the target model parameters.
[0067] After training on the training set is completed, outer loop optimization is needed, which can be performed on a validation set by verifying the updated model parameters by calculating a validation loss, and the target model parameters are obtained by backpropagation of the gradient. Specifically, the performance of the updated model parameters can be evaluated on the validation set , and a validation loss is calculated, wherein, the expression is the same as the preset training objective; is a first qualified sample in the validation set; represents an expectation of all validation losses in the validation set. Then the global model initialization parameters can be updated by backpropagation of the gradient, that is wherein, represents a preset meta learning rate, represents a loss term of the score obtained by the first cold start quality evaluator, represents a preset adjustment weight, and finally the target model parameters can be obtained.
[0068] Step 310, based on the target model parameters, a preset few-shot emotion speech generation model is generated.
[0069] After the target model parameters are obtained, the target model parameters can be deployed in the model to obtain the preset few-shot emotion speech generation model.
[0070] In some optional embodiments, after the model training is completed, a new user or a new emotion task is faced, and both are recorded as , only a small number of support samples (such as 5) need to be provided to quickly complete fine-tuning and obtain fine-tuned model parameters , , wherein represents a preset training target under , and finally the model corresponding to the fine-tuned model parameters can be used for subsequent processing.
[0071] In this embodiment, through the extremely small sample driven modeling mechanism, the user only needs to provide 3-5 voice, 1 face image and 1 set of personality labels to automatically build the initial digital human voice and image style, realize the personal modeling from 0 to 1, that is, the model construction has the personalized cold start ability; in addition, a small number of samples are used to quickly complete the model parameter update, and the result quality is improved round by round, realizing the end-to-end automatic modeling closed loop of "input-optimization-feedback", and significantly improving the cold start experience and generation stability.
[0072] Step 204, input the target audio file, target emotion embedding vector and target personality feature vector into the preset digital human image generation model for processing to generate target images corresponding to each target digital human.
[0073] Among them, the target image is used to reflect the facial state consistent with the emotional expression and personality expression in the target audio file. The preset digital human image generation model is constructed based on the conditional generative adversarial network. The preset digital human image generation model includes an expression controller module, a personality adjuster module, a condition self-adaptor module, an image generator module, a cold start quality evaluation and playback module.
[0074] The above-mentioned preset digital human image generation model also needs to be pre-constructed. In some optional embodiments, as shown in Figure 6 , it is a step flow chart for constructing a preset digital human image generation model provided by the embodiment of the application, which includes: Figure 6 Step 602, obtain a candidate training sample set.
[0075] Step 604, input the candidate training sample set into the second cold start quality evaluator for processing, and calculate the second comprehensive quality score through the second preset evaluation formula.
[0076] Step 606, based on the second comprehensive quality score and the second preset screening threshold, the second qualified sample is screened.
[0077] Step 608, input the second qualified sample into the initial image generation model for training to calculate the total training loss.
[0078] Step 610, updating the model parameters based on the total training loss to generate a preset digital human image generation model.
[0079] Wherein, the candidate training sample set includes a plurality of candidate training samples, and the candidate training samples can include face image samples, emotion embedding vector samples, and personality feature vector samples.
[0080] Then, the candidate training sample set can be input into the second cold start quality evaluator for processing, and a second comprehensive quality score is calculated by a second preset evaluation formula, wherein the second preset evaluation formula is determined based on a second clarity index, a semantic consistency index, a second diversity index, and a second anomaly confidence.
[0081] In some optional embodiments, as shown in Figure 7 , Figure 7 A step flowchart for calculating a second comprehensive quality score provided by the embodiments of the present application includes: Step 702, inputting the candidate training sample set into the second cold start quality evaluator to obtain the no-reference image quality evaluation index and Laplacian variance corresponding to the candidate training sample, and calculating the second clarity index based on the no-reference image quality evaluation index and the Laplacian variance.
[0082] Step 704, calculating the semantic consistency index based on the similarity between the first predicted identity embedding vector corresponding to the candidate training sample and the second predicted identity embedding vector corresponding to the original face image.
[0083] Step 706, calculating the second diversity index based on the distance between each candidate training sample.
[0084] Step 708, calculating the second probability that the candidate training sample is generated abnormally based on the candidate training sample and the second predicted identity embedding vector, and obtaining the second anomaly confidence based on the second probability.
[0085] Step 710, calculating the second comprehensive quality score based on the second clarity index, the semantic consistency index, the second diversity index, and the second anomaly confidence.
[0086] Wherein, the second clarity index is used to measure whether the generated image is clear and free of artifacts, and the specific calculation process is shown in formula (6).
[0087] (6) Wherein, NIQE and BRISQUE represent different no-reference image quality evaluation indexes, and the lower the value, the higher the quality; LaplVar is the Laplacian variance, and the larger the value, the higher the edge definition; 、 、 respectively represent corresponding weights.
[0088] The semantic consistency index is used to measure whether the generated image retains the original identity and reasonable expression. The higher the cosine similarity, the more consistent the generated image is with the original identity. The specific calculation process is shown in formula (7).
[0089] (7) wherein, denotes the semantic consistency index; denotes a face recognition model, and outputs a predicted identity embedding vector; denotes a candidate training sample corresponding to a first predicted identity embedding vector; denotes an original face image corresponding to a second predicted identity embedding vector.
[0090] It should be noted that the face recognition model is constructed based on a convolutional neural network, including an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a maximum pooling layer, a fully connected layer, a random dropout layer, a feature embedding layer, and an output layer. Specifically, low-level and high-level features of the input face image are extracted through three convolutional layers, and then the feature dimension is reduced using the maximum pooling layer. Then, the features are fused through the fully connected layer, and the non-linear representation is enhanced through the ReLU activation function. In order to prevent overfitting, the face recognition model introduces a random dropout layer, and maps the extracted features to a low-dimensional space through the feature embedding layer to generate a unique face feature vector, i.e. a predicted identity embedding vector, and outputs the result through the output layer.
[0091] The second diversity index is used to ensure that the styles of different candidate training samples are different. The larger the value, the stronger the diversity. The specific calculation process is shown in formula (8).
[0092] (8) wherein, denotes the second diversity index; 、 denotes different candidate training samples; L denotes the number of candidate training samples.
[0093] The second abnormal confidence can be obtained by a discriminator trained by self-supervised contrast learning The output is specifically calculated as shown in formula (9).
[0094] (9) wherein, represents the second abnormal confidence.
[0095] It should be noted that the above discriminator includes an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fully connected layer, and an output layer. Specifically, multi-level features of the input face image are extracted through four convolutional layers, and the spatial dimension of the image is gradually reduced, and a Leaky ReLU (Leaky Rectified Linear Unit) activation function is used to increase the non-linear ability of the network. Then, the features are fused through the fully connected layer, and a probability value is output through the Sigmoid activation function, which is used to measure whether the input image is abnormal. The result is output through the output layer.
[0096] Finally, the second comprehensive quality score can be calculated based on the second intelligibility index, the semantic consistency index, the second diversity index, and the second abnormal confidence. The specific calculation process is shown in formula (10).
[0097] (10) wherein, represents the second comprehensive quality score; , , , respectively represent the corresponding weights.
[0098] It should be noted that the above second cold start quality evaluator is equivalent to the second comprehensive quality score , that is, the weighted sum of the above four indexes, namely the second intelligibility index, the semantic consistency index, the second diversity index, and the second abnormal confidence.
[0099] Thus, the second qualified sample can be obtained based on the second comprehensive quality score and the second preset screening threshold. Specifically, the second comprehensive quality score can be compared with the second preset screening threshold. If the second comprehensive quality score is less than the second preset screening threshold, the corresponding sample is discarded, otherwise it is regarded as the second qualified sample.
[0100] Then the second qualified sample can be input into the initial image generation model for training to calculate a total training loss. The total training loss includes an adversarial loss, an identity preservation loss, an emotion consistency loss, a temporal smoothing loss and a reconstruction loss. Finally, the model parameters are updated based on the total training loss to generate a preset digital human image generation model. The initial image generation model is also constructed based on a conditional generative adversarial network, and the model parameters of the conditional generative adversarial network are obtained after initialization.
[0101] The improved generative adversarial network discriminator PatchGAN discriminator can be used to distinguish the generated face image from the preset real face image in the training sample set . to ensure the image authenticity . Wherein, is a discriminator representing the probability that the generated face is "real"; T represents the length of the generated face image sequence; the first term represents the expectation of the logarithmic value of the discriminator result of the set . This term aims to encourage the discriminator to give a high confidence judgment to the preset real face image from the training set, and to strengthen the recognition ability of the real image; the second term represents the expectation of the logarithmic value of the discriminator result of the set after the complement processing. This term aims to encourage the discriminator to give a "fake" judgment to the generated face image , and to strengthen the recognition ability of the fake image.
[0102] The pre-trained face recognition network can be used to extract image features to maintain the identity consistency with the input face image enhanced sample, and to calculate the identity preservation loss . . Wherein, is a face feature extractor, which is also the face recognition model .
[0103] The preset emotion classifier can be used to perform expression recognition on the generated face image to match the input voice emotion, so as to calculate the emotion consistency loss . . Wherein, is a cross-entropy loss function; is a preset emotion classifier, outputs the encoding result corresponding to the t-th frame of the generated face image; is a function of mapping the target emotion embedding vector back to the emotion one-hot encoding, An encoding result corresponding to the target emotion embedding vector of the t-th frame is output.
[0104] It should be noted that the above-mentioned preset emotion classifier The preset emotion classifier includes an input layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a max-pooling layer, a fully connected layer, a random dropout layer, and an output layer. Specifically, the generated face image is input into the preset emotion classifier through the input layer, and then the emotion features of the input data are extracted through the four convolutional layers in sequence, and the dimensionality is reduced through the max-pooling layer to retain key information. The extracted features are fused through the fully connected layer, and the non-linear expression is enhanced through the ReLU activation function. Finally, the generated face image is classified into different emotion categories through the output layer of the Softmax activation function, and the result is output through the output layer.
[0105] The time smoothing loss can be calculated by encouraging the continuous change of adjacent frame images to prevent expression jitter , .
[0106] The reconstruction loss can be calculated by encouraging the generated image to be as close as possible to the preset real face image for fine alignment , .
[0107] Thus, the total training loss can be calculated based on the adversarial loss, the identity preservation loss, the emotion consistency loss, the time smoothing loss, and the reconstruction loss , . Among them, , , , are the corresponding preset weights. By introducing multiple losses, the subsequently generated target image not only retains individual characteristics but also has a real and natural dynamic expression.
[0108] Thus, when the target audio file, the target emotion embedding vector, and the target personality feature vector are input into the preset digital human image generation model for processing, as shown in Figure 8 , Figure 8 is a step flowchart provided by an embodiment of the present application for generating a target image corresponding to each target digital human, which includes: Step 802, input the target audio file, the target emotion embedding vector, and the target personality feature vector into the preset digital human image generation model, and evaluate and filter the current input data through the cold start quality evaluation and playback module to obtain qualified input data.
[0109] Step 804, the target emotion embedding vector in the qualified input data is frame-divided and encoded by the expression controller module to obtain an emotion control vector for controlling the expression state of the image.
[0110] Step 806, the target personality feature vector in the qualified input data is mapped by the personality adjuster module to obtain a personality adjustment vector.
[0111] Step 808, the target emotion embedding vector and the target personality feature vector in the qualified input data are jointly mapped into a layer-by-layer style code by the condition adapter module, and an adapted condition style code is output.
[0112] Step 810, the condition style code, the emotion control vector and the personality adjustment vector in the qualified input data are processed by the image generator module to generate a target image corresponding to each target digital person.
[0113] To solve the problem of small data quantity and poor quality in the cold start stage, the cold start quality evaluation and playback module automatically evaluates and filters the current input data during the generation process to obtain qualified input data. Specifically, the current comprehensive quality score corresponding to the current input data can be calculated first, and then the current comprehensive quality score is compared with the pre-set qualified threshold and unqualified threshold. The current input data greater than the qualified threshold is retained, the current input data less than the unqualified threshold is removed, and the current input data between the qualified threshold and the unqualified threshold is resampled. Optionally, a playback buffer pool can be pre-set to store high-quality samples for subsequent incremental training, avoiding overfitting caused by insufficient small samples.
[0114] The expression controller module frame-divides and encodes the target emotion embedding vector in the qualified input data to obtain an emotion control vector for controlling the expression state of the image. The emotion control vector can determine the emotional performance of the digital person image at the t-th frame, ensuring that the generated expression dynamically changes with the speech context.
[0115] Then, the target personality feature vector in the qualified input data can be mapped by the personality adjuster module to obtain a personality adjustment vector. The personality adjustment vector is used to control the expression style and detailed features of the target digital person, so as to ensure that the generated personification performance meets the user's personality label. The expression style can include expression amplitude, tension, micro-expression, etc., and the detailed features can include clothing style, etc.
[0116] In the cold start stage, directly inputting the target emotion embedding vector and the target personality feature vector into the generator may cause unstable training due to the few samples. Therefore, a conditional adapter module is introduced. The conditional adapter module maps the target emotion embedding vector and the target personality feature vector in the qualified input data into layer-by-layer style codes through a preset condition fusion network, and outputs the adapted conditional style code. The adapted conditional style code is used to modulate the feature mapping of each layer of the image generator module, which can enhance the robustness and generalization ability under small samples.
[0117] Finally, the image generator module can process the conditional style code, the emotion control vector and the personality adjustment vector in the qualified input data to generate the target image corresponding to each target digital person. This module can generate a sequence of face frames synchronized with the dynamic voice, ensuring that the output target image corresponding to the target digital person is consistent in emotion and personality.
[0118] The application provides a digital figure generation method combining cold start driving and active learning mechanism, which comprises the following steps: inputting a target digital crowd, target text content corresponding to the target digital crowd, a target emotion embedding vector and a target personality feature vector into a preset few-sample emotional speech generation model for processing to generate a target audio file corresponding to each target digital person in the target digital crowd, wherein the target audio file is consistent with the emotional expression and personality expression of the target digital person; inputting the target audio file, the target emotion embedding vector and the target personality feature vector into a preset digital figure generation model for processing to generate a target image corresponding to each target digital person, wherein the target image is used to reflect the facial state consistent with the emotional expression and personality expression in the target audio file; the preset digital figure generation model is constructed based on a conditional generative adversarial network; and the preset digital figure generation model comprises an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module and a cold start quality evaluation and playback module. According to the scheme, the input data is processed by the preset few-sample emotional speech generation model derived from the cold start process and having high personalization and dynamic adaptability, efficient personalized and multi-emotional speech batch generation can be realized on large-scale unlabeled input data, and therefore the target audio file with high fidelity and consistent emotional and personal expression can be output; the preset few-sample emotional speech generation model is trained based on the first qualified sample selected by the first cold start quality evaluator from the text-speech pair data training sample of the candidate joint personality emotion, so that the model can be started under the condition of very few data, and the automatic quality evaluation mechanism is used to continuously optimize the generated sample, expand the available data space, break through the modeling barrier of the existing digital human system, and realize the automatic generation of high-quality personalized digital human image under the condition of small sample; in addition, the target emotion embedding vector and the target personality feature vector are inputted to realize the synchronous generation of expressions under the driving of the speech, and the natural interaction ability and style consistency of the digital human are enhanced.
[0119] It should be understood that, although each step in the flowchart involved in each embodiment as described above is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.
[0120] The present application also provides a digital human image generation device that combines a cold start drive and an active learning mechanism, including: The first generation module is used to input the target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector and the target personality feature vector into the preset few-sample emotion speech generation model for processing, and generate a target audio file corresponding to each target digital person in the target digital population; wherein the target audio file is consistent with the emotion expression and personality expression of the target digital person; the preset few-sample emotion speech generation model is obtained by training based on the first qualified sample after screening the preset joint personality emotion text-speech pair data training sample using the first cold start quality evaluator.
[0121] The second generation module is used to input the target audio file, target emotion embedding vector and target personality feature vector into the preset digital human image generation model for processing to generate the target image corresponding to each target digital human; wherein, the target image is used to reflect the facial state consistent with the emotional expression and personality expression in the target audio file; the preset digital human image generation model is constructed based on the conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adapter module, an image generator module, and a cold start quality assessment and playback module.
[0122] Regarding the apparatus in the above-mentioned embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method and will not be elaborated on here. The various modules in the digital human image generation device that combines the cold start drive and active learning mechanism can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations of the above modules.
[0123] In one embodiment of the present application, a computer device is provided. The computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented: The target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector are input into a preset few-shot emotion speech generation model for processing to generate a target audio file corresponding to each target digital person in the target digital population; wherein the target audio file is consistent with the emotion expression and personality expression of the target digital person; the preset few-shot emotion speech generation model is obtained by training the first qualified sample after screening the candidate joint personality emotion text-speech pair data training sample using the first cold start quality evaluator; The target audio file, target emotion embedding vector and target personality feature vector are input into the preset digital human image generation model for processing to generate the target image corresponding to each target digital human; wherein, the target image is used to reflect the facial state consistent with the emotion expression and personality expression in the target audio file; the preset digital human image generation model is constructed based on the conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adapter module, an image generator module, and a cold start quality assessment and playback module.
[0124] The computer device provided in the embodiment of the present application has similar implementation principles and technical effects to those of the above-mentioned method embodiment, and will not be described in detail here.
[0125] In one embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: The target digital population, the target text content corresponding to the target digital population, the target emotion embedding vector, and the target personality feature vector are input into a preset few-shot emotion speech generation model for processing to generate a target audio file corresponding to each target digital person in the target digital population; wherein the target audio file is consistent with the emotion expression and personality expression of the target digital person; the preset few-shot emotion speech generation model is obtained by training the first qualified sample after screening the candidate joint personality emotion text-speech pair data training sample using the first cold start quality evaluator; The target audio file, target emotion embedding vector and target personality feature vector are input into the preset digital human image generation model for processing to generate the target image corresponding to each target digital human; wherein, the target image is used to reflect the facial state consistent with the emotion expression and personality expression in the target audio file; the preset digital human image generation model is constructed based on the conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adapter module, an image generator module, and a cold start quality assessment and playback module.
[0126] The computer-readable storage medium provided in this embodiment has similar implementation principles and technical effects to those of the above-mentioned method embodiment, and will not be described in detail here.
[0127] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, databases, or other media in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0128] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the disclosure that are deemed to fall within the general principles of the disclosure and include commonly known or customary practice in the art. The specification and examples are to be considered exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.
[0129] It should be understood that the present disclosure is not limited to the precise structures as herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the claims that follow.
Claims
1. A digital human image generation method combining cold start driving and active learning mechanism, characterized in that, The method comprises: inputting a target digital population, target text content corresponding to the target digital population, a target emotion embedding vector and a target personality feature vector into a preset few-sample emotion speech generation model for processing to generate a target audio file corresponding to each target digital person in the target digital population; wherein the target audio file is consistent with the emotional expression and personality expression of the target digital person; the preset few-sample emotion speech generation model is trained based on first qualified samples selected by a first cold start quality evaluator from text-speech pair data training samples of candidate joint personality emotions; inputting the target audio file, the target emotion embedding vector and the target personality feature vector into a preset digital human image generation model for processing to generate a target image corresponding to each target digital person; wherein the target image is used to reflect a facial state consistent with the emotional expression and personality expression in the target audio file; the preset digital human image generation model is constructed based on a conditional generative adversarial network, and comprises an expression controller module, a personality adjuster module, a conditional self-adaptor module, an image generator module and a cold start quality evaluation and playback module.
2. The method of claim 1, wherein, The generation process of the preset few-sample emotion speech generation model comprises: inputting the text-speech pair data training samples of the candidate joint personality emotions into the first cold start quality evaluator for processing to calculate a first comprehensive quality score by a first preset evaluation formula; wherein the first preset evaluation formula is determined based on a first intelligibility index, an emotional consistency index, a first diversity index and a first abnormal confidence; based on the first comprehensive quality score and a first preset screening threshold, the first qualified samples are screened and balanced sampling is performed on the first qualified samples to obtain a training set and a validation set; on the training set, the parameters of an initial emotion speech generation model are updated by maximizing a preset training target to obtain updated model parameters; wherein the preset training target is determined based on a mel spectrogram corresponding to speech data, text data, an emotion embedding vector and a personality feature vector, the initial emotion speech generation model is constructed based on a Tacotron2 conditional speech synthesis model and a conditional variational autoencoder mechanism, and the conditional variational autoencoder mechanism inputs the emotion embedding vector and the personality feature vector after fusion as a conditional variable; on the validation set, the updated model parameters are verified by calculating a verification loss, and target model parameters are obtained by backpropagation of gradients; based on the target model parameters, the preset few-sample emotion speech generation model is generated.
3. The method of claim 2, wherein, the inputting of the text-speech pair data training samples of the candidate joint personality emotions into the first cold start quality evaluator for processing to calculate a first comprehensive quality score by a first preset evaluation formula comprises: inputting the text-speech pair data training sample of the candidate joint personality emotion into the first cold start quality evaluator, obtaining a signal-to-noise ratio, a speech perceptual quality score and a speech intelligibility index corresponding to the text-speech pair data training sample of the candidate joint personality emotion, and calculating the first intelligibility index based on the signal-to-noise ratio, the speech perceptual quality score and the speech intelligibility index; obtaining a predicted emotion distribution based on the emotion mel-spectrogram corresponding to the text-speech pair data training sample of the candidate joint personality emotion, and calculating the emotion consistency index based on the predicted emotion distribution and a similarity of a preset emotion label; calculating the first diversity index based on distances between the text-speech pair data training samples of the candidate joint personality emotion; calculating a first probability that the text-speech pair data training sample of the candidate joint personality emotion is abnormally generated based on the emotion mel-spectrogram and the preset emotion label, and obtaining the first abnormal confidence based on the first probability; calculating the first comprehensive quality score based on the first intelligibility index, the emotion consistency index, the first diversity index and the first abnormal confidence.
4. The method of claim 2, wherein, The verification of the updated model parameters by calculating the verification loss, and the target model parameters obtained by backpropagation of the gradient, include: verifying the updated model parameters to calculate the verification loss; updating the global model initialization parameters based on the verification loss, a meta learning rate, a loss term of the score obtained by the first cold start quality evaluator and a preset adjustment weight, and obtaining the target model parameters in combination with a cold start iteration mechanism.
5. The method according to any one of claims 1 to 4, characterized in that, The generation process of the preset digital human image generation model includes: obtaining a candidate training sample set; wherein the candidate training sample set includes a plurality of candidate training samples; inputting the candidate training sample set into a second cold start quality evaluator for processing, and calculating a second comprehensive quality score by a second preset evaluation formula; wherein the second preset evaluation formula is determined based on a second intelligibility index, a semantic consistency index, a second diversity index and a second abnormal confidence; filtering a second qualified sample based on the second comprehensive quality score and a second preset filtering threshold; inputting the second qualified sample into an initial image generation model for training to calculate a total training loss; the total training loss includes an adversarial loss, an identity preservation loss, an emotion consistency loss, a temporal smoothing loss and a reconstruction loss; updating the model parameters based on the total training loss to generate the preset digital human image generation model.
6. The method of claim 5, wherein, The inputting of the candidate training sample set into the second cold start quality evaluator for processing to calculate the second comprehensive quality score by the second preset evaluation formula includes: inputting the candidate training sample set into the second cold start quality evaluator, obtaining a no-reference image quality evaluation index and a Laplacian variance corresponding to the candidate training sample, and calculating the second intelligibility index based on the no-reference image quality evaluation index and the Laplacian variance; The semantic consistency index is calculated based on a similarity between a first predicted identity embedding vector corresponding to the candidate training sample and a second predicted identity embedding vector corresponding to the original face image; The second diversity index is calculated based on distances between the candidate training samples; The second probability that the candidate training sample is generated abnormally is calculated based on the candidate training sample and the second predicted identity embedding vector, and the second abnormal confidence is obtained based on the second probability; The second comprehensive quality score is calculated based on the second clarity index, the semantic consistency index, the second diversity index, and the second abnormal confidence.
7. The method according to any one of claims 1 to 4, characterized in that, The target audio file, the target emotion embedding vector, and the target personality feature vector are input into a preset digital human image generation model for processing to generate target images corresponding to each of the target digital humans, including: The target audio file, the target emotion embedding vector, and the target personality feature vector are input into a preset digital human image generation model, and the current input data is evaluated and filtered by a cold start quality evaluation and playback module to obtain qualified input data; The target emotion embedding vector in the qualified input data is frame-divided and encoded by the expression controller module to obtain an emotion control vector for controlling the image expression state; The target personality feature vector in the qualified input data is mapped by the personality adjuster module to obtain a personality adjustment vector; wherein the personality adjustment vector is used to control the expression style and detail features of the target digital human; The target emotion embedding vector and the target personality feature vector in the qualified input data are jointly mapped into layer-by-layer style codes by the conditional adapter module, and the adapted conditional style codes are output; wherein the adapted conditional style codes are used to modulate the feature mapping of each layer of the image generator module; The conditional style codes, the emotion control vector, and the personality adjustment vector in the qualified input data are processed by the image generator module to generate target images corresponding to each of the target digital humans.
8. A digital human image generation device combining a cold start driving and active learning mechanism, characterized in that, The device includes: A first generation module is configured to input a target digital human group, target text content corresponding to the target digital human group, a target emotion embedding vector, and a target personality feature vector into a preset few-sample emotional speech generation model for processing to generate target audio files corresponding to each target digital human in the target digital human group; wherein the target audio file is consistent with the emotional expression and personality expression of the target digital human; and the preset few-sample emotional speech generation model is trained based on first qualified sample training after a first cold start quality evaluator is used to filter a preset joint personality emotion text-speech pair data training sample. A second generation module is configured to input the target audio file, the target emotion embedding vector, and the target personality feature vector into a preset digital human image generation model for processing to generate a target image corresponding to each target digital human; wherein the target image is used to reflect a facial state consistent with the emotional expression and personality expression in the target audio file; the preset digital human image generation model is constructed based on a conditional generative adversarial network, and the preset digital human image generation model includes an expression controller module, a personality regulator module, a conditional self-adaptor module, an image generator module, and a cold start quality evaluation and playback module.
9. An electronic device, comprising: The electronic device includes a processor and a memory, and the memory stores at least one instruction, at least one program, a code set, or an instruction set, which is loaded and executed by the processor to implement the digital human image generation method combining the cold start driving and active learning mechanism according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, a code set, or an instruction set, which is loaded and executed by the processor to implement the digital human image generation method combining the cold start driving and active learning mechanism according to any one of claims 1-7.
Citation Information
Patent Citations
Digital human animation generation method and device and digital human animation generation model
CN118799461A
Intelligent customer service training method and device of multi-modal large model, and storage medium
CN119577459A
3D digital human real-time dialogue interaction system and method
CN119783689A
Immersive virtual character interaction method based on virtual reality
CN120085753A
Method and device for automatically producing whole-body digital human data and recording medium
CN120472260A
Cited By
Intelligent algorithm-based corpus automatic collection and quality control method and system
CN122347962A