An interactive digital human generation method and system based on artificial intelligence
By employing an AI-based interactive digital human generation method, combined with style transfer and super-resolution techniques, a multi-layer neural network is constructed to fuse multimodal data, optimizing digital human performance and response strategies. This addresses the issues of emotional adaptability and interaction efficiency in digital humans, achieving high-quality digital human generation and interactive experiences.
Patent Information
- Application Number
- CN202411647512.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing digital human technologies struggle to adapt to the emotions or states of specific situations when generating facial expressions or voice responses, lacking realism and natural fluency, resulting in low efficiency in user interaction.
An AI-based interactive digital human generation method is adopted, which generates latent space representation through conditional vectors, combines style transfer and super-resolution techniques, constructs a multi-layer neural network to fuse multimodal data, uses generative adversarial networks to optimize digital human performance, and optimizes response strategies through reinforcement learning.
The generated digital human images and behaviors are visually and interactively realistic and natural, with high precision and consistency. They can adapt to diverse scenarios, achieve a natural and smooth interactive experience, and be continuously optimized.
Smart Images

Figure CN119600159B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital human interaction, more specifically, to an interactive digital human generation method and system based on artificial intelligence. BACKGROUND
[0002] In the current field of digital human technology, although there have been numerous developments and applications, there are still some key technical challenges and limitations, including the complexity of data integration, the lack of realism in generated data, the accuracy of emotion recognition, and the optimization needs of real-time response. Existing digital humans often cannot well adapt to the emotions or states of specific situations when generating facial expressions or voice responses, lacking realism and natural fluency, which affects the user's immersion experience and satisfaction. Existing systems often lack accuracy and efficiency in recognizing user emotions and generating corresponding responses, especially when faced with complex emotional expressions, the system's response is not accurate or timely enough, leading to low interaction efficiency. SUMMARY
[0003] The present application provides an interactive digital human generation method based on artificial intelligence, which can improve the financial ability and acceptance rate of customers, reduce the risk coefficient, at the same time, reduce the risk coefficient, inversely improve the financial ability and acceptance rate of customers, so as to find the yield balance point and improve the business profitability.
[0004] The present application provides an interactive digital human generation method based on artificial intelligence, comprising the following steps:
[0005] Digital human image generation step: define the input conditions of the digital human, the input conditions are represented by a condition vector, describing the basic characteristics and generation requirements of the digital human; input the condition vector into the encoder to generate a latent space representation, sample the latent vector through the decoder to generate a preliminary digital human image, use the reconstruction loss to ensure that the generated image meets the characteristics of the input conditions, and use the KL divergence to ensure the continuity of the latent space distribution. On the basis of the preliminary generated digital human image, apply style transfer technology to adapt it to the target environment or style, use a super-resolution network to improve the resolution and detail quality of the image, and output a high-resolution image final image;
[0006] Digital human behavior generation steps: Collect multimodal data and convert it into a corresponding feature vector set, where the multimodal data includes facial expressions, voice, and body movements; construct a fusion function, which is used to integrate multiple feature vectors in the feature vector set into a unified feature representation. The fusion function is a multi-layer neural network that maps data from different modalities to a unified feature space through the weight matrix in the neural network; define a loss function to evaluate the difference between the generated digital human behavior and the preset target; adjust the loss function parameters through an optimization algorithm to minimize the loss function value, where the optimization algorithm includes using stochastic gradient descent or the Adam algorithm;
[0007] Digital human performance generation steps: Set a conditional vector to represent a specific emotion or state; receive random noise, the conditional vector, and a unified feature representation of multimodal data through the generator, and generate facial expression or voice data corresponding to the set emotion or state; use the discriminator to evaluate the match between the generated data and the conditional vector and determine its authenticity; optimize the generator and discriminator by minimizing the adversarial loss function to improve the authenticity and match of the generated data, and construct a performance data synthesizer;
[0008] Digital human emotion recognition steps: Obtain a user's facial image and extract features from the user's facial image using a pre-trained convolutional neural network model; Input the extracted feature vector into a classifier to identify the user's specific expression. The classifier is a support vector machine or a multi-layer perceptron; Based on the recognized expression category, the digital human generates a corresponding facial expression or voice response;
[0009] Digital human response optimization steps: receiving user input information, which includes voice commands or text information; processing the user input and generating a corresponding response through a preset reaction function; adjusting the parameters of the reaction function through feedback to ensure the accuracy and timeliness of the response.
[0010] Preferably, in the digital human behavior generation step, feature vectors of facial expressions, voice and body movements are extracted from real-time video and audio data. The fusion function adopts the form of a combination of convolutional layers and fully connected layers to process and integrate input data of different modalities. The loss function is cross entropy or mean square error, which is used to evaluate the consistency between the output predicted by the model and the actual label or expected output.
[0011] Preferably, in the digital human response optimization step, the user's input data occurs at time t, where For voice commands or text messages, represented as ; Use the reaction function R to generate the digital human's response , the reaction function R is expressed as: ,in, is a parameter of the reaction function The parameters of the reaction function are optimized through the training data.
[0012] Preferably, the method further comprises continuously optimizing the digital human's behavior and reaction strategy using reinforcement learning, comprising the following steps:
[0013] Defining a set of possible actions that can be taken based on the current system state and the context of user interaction; Let denote the system state at time , including the current state of the digital human and the context of user interaction, denote the set of actions that can be taken in state , learn a policy using reinforcement learning algorithm selects the optimal action in state , denoted as: ; where is the parameter of the policy;
[0014] Action results in the state becoming and produces an immediate reward , denoted as:
[0015]
[0016] Using reinforcement learning algorithm, learn a policy to select actions that maximize the expected value of cumulative rewards; The goal of reinforcement learning is to optimize the policy
[0017]
[0018] where is the discount factor, representing the current value of future rewards;
[0019] Update the policy parameters based on the results of the action and the immediate reward generated, to improve future behavior decisions; The policy parameters are updated by gradient ascent method to maximize long-term returns, denoted as:
[0020]
[0021] where is the learning rate, is the gradient of with respect to , calculated by the policy gradient method;
[0022] Generate a comprehensive reaction function , denoted as: .
[0023] The application also provides an interactive digital human generation system based on artificial intelligence, based on the aforementioned interactive digital human generation method based on artificial intelligence, characterized in that it comprises the following modules:
[0024] Digital human image generation module: used to define the input conditions of the digital human, the input conditions are represented by a condition vector, which describes the basic characteristics and generation requirements of the digital human; the condition vector is input into the encoder to generate a latent space representation, the sampled latent vector is input into the decoder to generate a preliminary digital human image, the reconstruction loss is used to ensure that the generated image meets the characteristics of the input conditions, and the KL divergence is used to ensure the continuity of the latent space distribution; based on the preliminary generated digital human image, the style transfer technology is applied to adapt it to the target environment or style, the super-resolution network is used to improve the resolution and detail quality of the image, and the high-resolution image is output as the final image;
[0025] Digital human behavior generation module: used to collect multi-modal data and convert them into corresponding feature vector sets, the multi-modal data including facial expressions, voices and body movements; a fusion function is constructed, which is used to integrate multiple feature vectors in the feature vector set into a unified feature representation, the fusion function being a multi-layer neural network, and different modal data being mapped to a unified feature space through a weight matrix in the neural network; a loss function is defined to evaluate the difference between the generated digital human behavior and the preset target; the loss function parameters are adjusted by an optimization algorithm to minimize the loss function value, the optimization algorithm including using a stochastic gradient descent or Adam algorithm;
[0026] Digital human performance generation module: used to set a condition vector representing a specific emotion or state; through a generator, random noise, the condition vector and the unified feature representation of the multi-modal data are received to generate facial expressions or voice data corresponding to the set emotion or state; a discriminator is used to evaluate the matching degree of the generated data and the condition vector and judge the authenticity thereof; the generator and the discriminator are optimized by minimizing the adversarial loss function to improve the authenticity and matching degree of the generated data, and a performance data synthesizer is constructed;
[0027] Digital human emotion recognition module: used to obtain the facial image of a user, and a pre-trained convolutional neural network model is used to extract features from the facial image of the user; the extracted feature vector is input into a classifier to identify the specific expression of the user, the classifier being a support vector machine or a multi-layer perceptron; according to the identified expression category, the digital human generates a corresponding facial expression or voice response;
[0028] Digital human response optimization module: for receiving input information of the user, the input information including voice instructions or text information; processing the user input and generating a corresponding response through a preset reaction function; adjusting the parameters of the reaction function through feedback to ensure the accuracy and timeliness of the response.
[0029] The beneficial effects of the present application are: the present scheme constructs an intelligent, efficient and multifunctional digital human system through a complete digital human image, behavior, performance, emotion and response generation and optimization process. First, based on the generation of condition vectors and the structured representation of latent space, combined with style transfer and super-resolution technology. The continuity of the latent space ensures that the generated digital human image balances between consistency and diversity. The style transfer technology enables the digital human image to adapt to different virtual environments or artistic styles, with higher visual adaptability. The application of super-resolution network significantly improves the resolution and detail quality of the image, meeting the high demand for visual presentation. The generated digital human image not only has personalized features, but also can adapt to various application scenarios, with clear image details and extremely high visual appeal. Fusion and unified feature representation generation based on multi-modal data. Through the fusion function constructed by the multi-layer neural network, the multi-modal data is mapped to the common feature space, capturing the associated information between different modalities. The generated digital human behavior has high precision and consistency, can accurately express target emotions and actions, and realizes natural and smooth interaction experience, suitable for virtual hosts, education and training scenarios. The generation of random noise, emotion condition vectors and multi-modal feature representation based on the generative adversarial network framework. The generated digital human performance is natural and realistic, and the emotional performance is highly matched with the set target. BRIEF DESCRIPTION OF DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the present application will be further described below with reference to the drawings and embodiments. The drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the premise of not paying creative labor:
[0031] Figure 1 is a flowchart of an interactive digital human generation method based on artificial intelligence according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be described below clearly and completely. Obviously, the described embodiments are some embodiments of the present application, not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0033] As Figure 1 shown, the present application provides an interactive digital human generation method based on artificial intelligence, comprising the following steps:
[0034] Digital human image generation step: define the input conditions of the digital human, the input conditions are represented by a condition vector, which describes the basic characteristics and generation requirements of the digital human; input the condition vector into the encoder to generate the latent space representation, pass the sampled latent vector through the decoder to generate the preliminary digital human image, use the reconstruction loss to ensure that the generated image meets the characteristics of the input conditions, use the KL divergence to ensure the continuity of the latent space distribution, on the basis of the preliminary generated digital human image, apply style transfer technology to adapt it to the target environment or style, use super-resolution network to improve the resolution and detail quality of the image, and output the high-resolution image final image;
[0035] Digital human behavior generation step: collect multi-modal data and convert them into corresponding feature vector sets, the multi-modal data including facial expressions, voices and body movements; construct a fusion function, the fusion function is used to integrate multiple feature vectors in the feature vector set into a unified feature representation, the fusion function is a multi-layer neural network, and different modal data is mapped to a unified feature space through the weight matrix in the neural network; define a loss function for evaluating the difference between the generated digital human behavior and the preset target; adjust the loss function parameters through an optimization algorithm to minimize the loss function value, the optimization algorithm includes using stochastic gradient descent or Adam algorithm;
[0036] Digital human performance generation step: set a condition vector representing a specific emotion or state; through the generator, receive random noise, condition vector and unified feature representation of multi-modal data, generate facial expressions or voice data corresponding to the set emotion or state; use the discriminator to evaluate the matching degree of the generated data and the condition vector, and judge its authenticity; optimize the generator and discriminator by minimizing the adversarial loss function to improve the authenticity and matching degree of the generated data, and build a performance data synthesizer;
[0037] Digital human emotion recognition step: obtain the user's facial image, and use a pre-trained convolutional neural network model to extract features from the user's facial image; input the extracted feature vector into a classifier to identify the user's specific expression, the classifier is a support vector machine or a multi-layer perceptron; according to the identified expression category, the digital human generates the corresponding facial expression or voice response;
[0038] Digital human response optimization step: receive user input information, the input information includes voice instructions or text information; process user input and generate corresponding responses through a pre-set reaction function; adjust the parameters of the reaction function through feedback to ensure the accuracy and timeliness of the response.
[0039] This solution builds an intelligent, efficient and multi-functional digital human system through a complete digital human image generation and optimization process, including appearance, behavior, performance, emotion and response. First, based on the generation of conditional vectors and the structured representation of latent space, combined with style transfer and super-resolution technology. The continuity of the latent space ensures that the generated digital human image balances between consistency and diversity. Style transfer technology enables digital human images to adapt to different virtual environments or artistic styles, with higher visual adaptability. The application of super-resolution network significantly improves the resolution and detail quality of the image, meeting the high demand for visual presentation. The generated digital human image not only has personalized features, but also can adapt to various application scenarios (such as the metaverse, film production, virtual assistants), with clear image details and high visual appeal.
[0040] Based on the fusion and unified feature representation generation of multi-modal data (facial expressions, speech, body movements). The fusion function constructed by multi-layer neural network maps multi-modal data to a common feature space, capturing the correlation information between different modalities. The generated digital human behavior has high precision and consistency, can accurately express target emotions and actions, and realizes natural and smooth interaction experience, suitable for virtual hosts, education and training scenarios.
[0041] Combined with the generation of random noise, emotion condition vectors and multi-modal feature representation, the framework of the generative adversarial network. The generated digital human performance is natural and realistic, and the emotional performance is highly matched with the set target, suitable for emotional interaction scenarios.
[0042] Through convolutional neural network to extract user facial features, and combined with support vector machine or multilayer perceptron classifier to identify specific expression categories. Accurate emotion recognition algorithm provides the basis for digital human to generate targeted facial expressions or speech responses. Digital human can quickly perceive user emotions and make adaptive performance, improving the individualization of interaction and the satisfaction of user experience.
[0043] Based on the reaction function input by the user and the reinforcement learning strategy optimization, to ensure that the response of the digital human continuously adapts to the user's needs. Adjust the parameters of the reaction function through real-time feedback to ensure the real-time and accuracy of the response. Use reinforcement learning to optimize the behavior and response strategy in the long term, maximize the long-term benefits of interaction. Digital human not only can quickly respond to user input, but also can continuously improve the interaction effect through learning, forming a self-adaptive and intelligent response mechanism, suitable for intelligent assistants, virtual assistants and other application scenarios that require long-term learning. This solution seamlessly integrates image generation, behavior generation, performance generation, emotion recognition and response optimization modules to build a complete digital human generation and optimization system.
[0044] This solution not only applies to static scenes (such as advertising, artistic creation), but also supports dynamic interactive scenes (such as virtual meetings, education and training, and metaverse interaction), with strong scalability.
[0045] This solution builds an intelligent, personalized, and adaptable digital human system through clear creation points and optimization techniques at each stage of digital human generation, with the following outstanding advantages:
[0046] High-quality image generation: clear and realistic visual effects, diverse styles.
[0047] Precise behavior performance: natural coordination, matching user emotions.
[0048] Intelligent response optimization: real-time response, continuous learning, and enhanced interactive experience.
[0049] In the digital human behavior generation step of the embodiment, the feature vectors of facial expressions, speech, and body movements are extracted from real-time video and audio data, the fusion function adopts a combination of convolutional layers and fully connected layers, which is used to process and integrate input data of different modalities, and the loss function is cross-entropy or mean square error, which is used to evaluate the consistency of the model's predicted output with the actual label or expected output.
[0050] In the digital human behavior generation step of the embodiment, the multi-modal data includes facial expressions E (feature vectors extracted from video frames by facial recognition algorithms, each feature vector represents the facial expression features at a time), speech S (speech features extracted by audio processing algorithms, such as mel-frequency cepstral coefficients), and body movements A (skeletal movements or joint angle features extracted by motion capture technology, each represents a feature vector of a movement frame), represented as:
[0051]
[0052] where represent the feature vectors of the collected facial expressions, speech, and body movements, respectively; multi-modal data comes from different collection frequencies, and needs to be synchronized with the time axis of different modal data through a time alignment algorithm to ensure that the features at corresponding time points can be fused; the multi-modal data is fused through a fusion module;
[0053] A fusion function is constructed, which is used to fuse feature vectors into a unified representation: , the fusion function is a multi-layer neural network that maps data of different modalities to a common feature space through a weight matrix , specifically represented as:
[0054]
[0055] wherein are weight matrices corresponding to facial expression, speech, body action respectively, is a bias term, is an activation function; the activation function Activation can select a nonlinear function (such as ReLU, Tanh or Sigmoid) for introducing nonlinear capabilities;
[0056] The network structure of the fusion model includes an input layer (receiving three kinds of modal features), a hidden layer (several neural network layers for learning the correlation between different modalities), and an output layer (generating a unified feature representation ).
[0057] The loss function of the fusion function is minimized defined as: wherein, is the feature representation generated by the fusion function F, is the target feature vector derived from the expected output. The above loss function is optimized using stochastic gradient descent , denoted as: ;
[0058] The training process of the fusion model is: initializing the network weights and the bias ; inputting multi-modal features ; forward propagation calculation ; calculating the loss ; backpropagation updating parameters ; repeat the training until the loss converges.
[0059] The fusion function uses a neural network to integrate multi-modal features (facial expressions, speech, and actions) to generate a unified feature representation. The generated unified feature representation more comprehensively reflects the user's emotions, intentions, and behavioral characteristics, improving the accuracy of subsequent digital human behavior generation. The loss function directly optimizes the matching degree between the fused features and the target features, continuously narrowing the feature gap during the training process, making the fused features more consistent with the real situation. Using stochastic gradient descent or algorithm to optimize network parameters, quickly converging to the optimal parameters, improving training efficiency, and adapting to large-scale data training. The deep fusion of multi-modal features provides high-quality input for the digital human behavior generation step. Through the deep fusion of multi-modal features in the behavior generation step, this scheme realizes the unified expression of multi-modal data, the precise matching of target features, and the efficient optimization of the training process. These improvements significantly improve the accuracy and coordination of digital human behavior generation, making it have obvious technical advantages and application value in terms of realism, naturalness, and adaptability.
[0060] In the digital human performance generation step of the present embodiment,
[0061] A condition vector is set to represent a specific emotion or state;
[0062] The generator is used to accept random noise , condition vector and unified feature representation of multi-modal data , output generated data , represented as: , generated data is the facial expression or voice of a digital human;
[0063] The discriminator is used to evaluate the authenticity of the generated data, specifically represented as: ;
[0064] The following adversarial loss function of the generator and discriminator is minimized, represented as:
[0065]
[0066] The objective function G is a function for generating the behavior and performance of a digital human, which accepts input features and outputs generated digital human performance, represented as: .
[0067] The condition vector represents the specific emotion or state of the digital human, such as happy, serious, angry, etc. By one-hot encoding, the emotion category is represented, or by using continuous values (such as emotion intensity) to represent the amplitude and subtle changes of emotion. Random noise is used to introduce diversity, so that the generated performance is not limited to fixed features, while avoiding mode collapse. It is usually sampled from a standard normal distribution. is a unified feature representation obtained by fusing facial expression, voice and body action features. provides the basis for the behavior of a digital human, ensuring that the generated performance conforms to the input data.
[0068] The generator structure includes: input layer (receiving z, and ), hidden layer (using multi-layer neural network to capture the relationship between multi-modal features and emotional conditions), output layer (generating target performance data ).
[0069] The discriminator structure includes: input layer (receiving and ), a hidden layer (extracting features using a multi-layer neural network), and an output layer (binary classification probability value, evaluating the authenticity of the input data).
[0070] logarithmic probability of real data , encouraging the discriminator to correctly identify real data.
[0071] logarithmic probability of generated data , the goal of the generator is to make the discriminator think that the generated data is real.
[0072] Alternately update the generator and discriminator until the generator can generate realistic performance data.
[0073] This embodiment realizes high-quality generation of digital human performance by fusing the condition vector, random noise and multi-modal features, combined with the generative adversarial network. The optimization of the adversarial loss function makes the generated data achieve a high balance in authenticity, diversity and emotional matching degree, providing an integrated solution for digital human performance generation, suitable for various application scenarios such as virtual reality, metaverse, intelligent assistants, etc.
[0074] In the digital human emotion recognition step of this embodiment, first, the user's face image is obtained, and then a pre-trained CNN model is used to extract features from the user's face image, represented as: where is the high-dimensional feature vector of the user's face image, is the parameter of the CNN model;
[0075] The extracted feature vector is input into the classifier to identify the specific emotion, represented as: where is the identified emotion category, is the parameter of the classifier;
[0076] The identified expression feature is supplemented as an additional input feature to the multi-modal data to adjust the fusion function, and the adjusted fusion function is represented as ;
[0077] The emotion feature is taken as a condition, and combined with random noise, the generator outputs facial expressions or voice data consistent with the emotion, represented as .
[0078] The user's face image face_image is captured in real time by the camera, and preprocessed to ensure the accuracy of feature extraction.
[0079] Preprocessing step:
[0080] Grayscale or normalize the image.
[0081] Adjust the image size to match the input size of the CNN (e.g., 224 x 224).
[0082] Data augmentation (e.g., rotation, flipping) to improve the generalization ability of the model.
[0083] Extract high-dimensional feature vectors F from facial images using a pre-trained convolutional neural network (CNN); the CNN model can be selected from architectures such as ResNet, VGG, or MobileNet, and the pre-trained parameters come from large-scale facial datasets (e.g., FER-2013 or AffectNet). The extracted features F include information about local facial regions (e.g., eyes, corners of the mouth, etc.), which helps accurately identify emotions.
[0084] Input the extracted feature vectors F into a classifier to identify specific emotion categories; the classifier type can be selected from support vector machines (SVM) or multilayer perceptrons (MLP). The output emotion category E represents the identified emotion, such as "happy," "angry," "neutral," etc.
[0085] The emotion feature E becomes part of the multi-modal feature through a fusion function, and together with facial expressions, speech, and action data, forms a more comprehensive and unified representation. The emotion feature E provides emotional information and works in conjunction with other modal data to improve the accuracy and naturalness of behavior generation and performance generation.
[0086] Using a pre-trained CNN to extract facial features can capture subtle changes in local regions and significantly improve the accuracy of emotion recognition. The classifier further enhances the generalization ability of emotion recognition, making it suitable for various emotional scenarios. The emotion feature is an important input that is supplemented to the multi-modal feature, enhancing the multi-modal data's ability to perceive emotional changes. By adjusting the structure of the fusion function, the synergy between the emotion feature and facial expressions, speech, and actions is significantly enhanced, improving the quality of multi-modal representation. Using the emotion feature and random noise as conditions, performance data consistent with a specific emotion is generated. The features provided by the emotion recognition module seamlessly connect with the multi-modal fusion and performance generation module, ensuring the overall coordination of the system. The generated facial expressions and speech data are consistent with the emotional conditions and have higher authenticity and naturalness.
[0087] In the digital human image generation step of the embodiment, the input conditions of the digital human are defined , the condition vector and random noise are input into the encoder to generate a latent space representation and , and a latent vector is generated in the latent space, represented as:
[0088]
[0089] The sampled latent vector Through the decoder, a preliminary digital human image is generated , denoted as:
[0090]
[0091] Using the reconstruction loss to ensure that the generated image meets the characteristics of the input conditions, denoted as:
[0092] The preliminary generated image is adjusted to the target style S, obtaining the adjusted image , and applying the content loss function to ensure that the generated image retains the content features of the preliminary image, denoted as: ;
[0093] Applying the style loss function to minimize the difference between the generated image and the reference style image in the style representation, denoted as: ;
[0094] Applying the content loss function to ensure that the generated image retains the content features of the preliminary image, denoted as: ;
[0095] Applying the style loss function to minimize the difference between the generated image and the reference style image in the style representation, denoted as: , where represents the style features of the style image , represents the Gram matrix at the layer, is the layer weight;
[0096] Combining the content loss and the style loss, the visual effect of the generated image is adjusted by optimizing the total loss function to match the specific virtual environment, denoted as: , where and are the weights of the content loss and the style loss, respectively;
[0097] Output the final static image .
[0098] The basic features of the digital human (such as gender, age, style requirements, etc.) are provided through the condition vector, and random noise is combined to achieve diversity in each generated image while meeting the conditions. The encoder generates parameters for the latent space to ensure the continuity of the latent space distribution, ensuring consistency in the input conditions and supporting flexible adjustments. The generated digital human image is highly personalized and can adapt to various application scenarios (such as virtual assistants, game characters, educational platforms, etc.). The preliminary generated image is converted into a reference style image, allowing the digital human image to adapt to different virtual environments (such as science fiction, cartoon, and realistic styles). The content loss ensures that the generated image still retains the core features of the preliminary generated image after style adjustment. The style loss minimizes the difference between the generated image and the target style image in terms of style representation, enhancing visual expression. The digital human image can naturally integrate into different scenarios, providing an immersive experience and being suitable for various scenarios such as the metaverse, film production, and social platforms. The content loss preserves the features of the preliminary image, such as facial proportions and posture, while the style loss optimizes details such as color and texture. The generated image balances content and style, has delicate visual expression and high-quality details, and meets high-resolution requirements.
[0099] In the digital human image generation step of the present embodiment, a dynamic animation is also generated based on the static image and multi-modal performance , the dynamic animation includes motion sequences and expression changes, and the specific steps are as follows:
[0100] Inputting the motion sequence condition , generating motion data , represented as:
[0101]
[0102] Combining the generated motion data with the static image and performance data to generate animation frames, represented as:
[0103] .
[0104] The embodiment realizes the generation of dynamic animation by organically combining static images with multi-modal expressions and action sequences. This design has the following beneficial effects: dynamicity and realism: the action sequence condition controls the generated action data, making the digital person's expression changes and action smooth and logical, and the generated animation frames are more natural and realistic in vision, enhancing the immersion and expressiveness of the interaction. Modularity and flexibility: the dynamic animation generation process is based on the combination of independent modules (static images, expression data, and action data), supporting flexible expansion and personalized customization, and can adapt to diverse scene requirements (such as virtual reality, education and training, metaverse, etc.). Efficiency and controllability: the action sequence condition provides explicit control input, and users can accurately design action sequences and expression changes, and the generation process is efficient and the output animation is highly predictable, suitable for real-time applications. Wide applicability: the generated dynamic animation has realism and flexibility, and can be applied to virtual anchors, film and television production, intelligent assistants, etc. Through the above design, the embodiment significantly improves the dynamic expression capability of digital person images.
[0105] In the digital person response optimization step of the embodiment, the input data of the user occurs at time t, where is a voice instruction or text information, represented as ; the response of the digital person is generated using the reaction function R , and the reaction function R is represented as: , where is a parameter of the reaction function , and the parameters of the reaction function are optimized through training data.
[0106] In the embodiment, a method for continuously optimizing the behavior and reaction strategy of the digital person using reinforcement learning is also included, which includes the following steps:
[0107] According to the current system state and the context of user interaction, a set of behaviors that can be taken is defined; let represent the system state at time , including the current state of the digital person and the context of user interaction, represent the set of actions that can be taken in state , learn a strategy through a reinforcement learning algorithm , which selects the optimal action in state , represented as: ; where is a parameter of the strategy;
[0108] The action causes the state to become and produces an immediate reward , denoted as:
[0109]
[0110] Using reinforcement learning algorithms, learn a policy for choosing actions to maximize the expected value of cumulative rewards; the goal of reinforcement learning is to optimize the policy π to maximize the expected value of cumulative rewards, denoted as:
[0111]
[0112] where, is a discount factor representing the current value of future rewards;
[0113] According to the results of the action and the immediate reward generated, update the policy parameters to improve future behavior decisions; policy parameters updated by gradient ascent method to maximize long-term returns, denoted as:
[0114]
[0115] where, is the learning rate, is the gradient of calculated by the policy gradient method;
[0116] generate a comprehensive response function , denoted as: .
[0117] This embodiment combines instant response generation based on the response function with reinforcement learning to optimize the policy, generating a comprehensive response function, achieving dynamic adaptability and long-term optimization capability of digital people to user input. The instant response function ensures that the digital person can quickly process user input and generate accurate real-time feedback to meet the needs of real-time interaction. By continuously optimizing behavior decision strategies through reinforcement learning, the optimal action is selected in combination with context information to generate a comprehensive response, making the response of the digital person more intelligent and more context-adaptive. The optimization goal focuses on cumulative rewards, balancing immediate rewards and long-term value, enabling the digital person to continuously improve its interactive performance and improve user satisfaction and system efficiency. The system adopts modular design, combining instant response and long-term strategy optimization, supporting diverse user scenarios and being applicable to intelligent assistants, virtual customer service, education and training and other application fields. Through this method, the digital person system realizes the whole process improvement from real-time interaction to intelligent behavior strategy optimization, greatly improving the interaction effect and user experience.
[0118] The present application also provides an embodiment of an interactive digital person generation system based on artificial intelligence, based on the previous embodiment, comprising the following modules:
[0119] Digital human image generation module: for defining the input conditions of the digital human, the input conditions are represented by a condition vector, which describes the basic characteristics and generation requirements of the digital human; input the condition vector into the encoder to generate the latent space representation, pass the sampled latent vector through the decoder to generate the preliminary digital human image, use the reconstruction loss to ensure that the generated image meets the characteristics of the input conditions, use the KL divergence to ensure the continuity of the latent space distribution, on the basis of the preliminary generated digital human image, apply style transfer technology to adapt it to the target environment or style, use super-resolution network to improve the resolution and detail quality of the image, and output the high-resolution image final image;
[0120] Digital human behavior generation module: for collecting multi-modal data and converting them into corresponding feature vector sets, the multi-modal data including facial expressions, voices and body movements; construct a fusion function, which is used to integrate multiple feature vectors in the feature vector set into a unified feature representation, the fusion function is a multi-layer neural network, which maps data of different modalities to a unified feature space through the weight matrix in the neural network; define a loss function for evaluating the difference between the generated digital human behavior and the preset target; adjust the loss function parameters through optimization algorithm to minimize the loss function value, the optimization algorithm includes using stochastic gradient descent or Adam algorithm;
[0121] Digital human performance generation module: for setting a condition vector representing a specific emotion or state; through the generator, receiving random noise, condition vector and unified feature representation of multi-modal data, generating facial expressions or voice data corresponding to the set emotion or state; use the discriminator to evaluate the matching degree of the generated data and the condition vector, and judge its authenticity; optimize the generator and discriminator by minimizing the adversarial loss function to improve the authenticity and matching degree of the generated data, and build the performance data synthesizer;
[0122] Digital human emotion recognition module: for obtaining the user's facial image, using a pre-trained convolutional neural network model to extract features from the user's facial image; input the extracted feature vector into the classifier to identify the user's specific expression, the classifier is a support vector machine or a multi-layer perceptron; according to the identified expression category, the digital human generates the corresponding facial expression or voice response;
[0123] Digital human response optimization module: for receiving user input information, the input information including voice instructions or text information; process user input and generate corresponding response through the preset reaction function; adjust the parameters of the reaction function through feedback to ensure the accuracy and timeliness of the response.
[0124] It is to be understood that all such modifications and variations that can occur to those skilled in the art in the light of the foregoing description are to be considered within the scope of the application as defined in the claims appended hereto.
Claims
1. An interactive digital human generation method based on artificial intelligence, characterized in that: The following steps are involved: Digital human image generation steps: define the input conditions of the digital human. The input conditions are represented by a condition vector, which describes the basic characteristics and generation requirements of the digital human. The conditional vector is input into the encoder to generate a latent space representation. The sampled latent vector is passed through the decoder to generate a preliminary digital human image. Reconstruction loss is used to ensure that the generated image meets the characteristics of the input condition. KL divergence is used to ensure the continuity of the latent space distribution. Based on the preliminary generated digital human image, style transfer technology is applied to adapt it to the target environment or style. A super-resolution network is used to improve the resolution and detail quality of the image, and the final high-resolution image is output; Digital human behavior generation steps: Collect multimodal data and convert it into a corresponding feature vector set, where the multimodal data includes facial expressions, voice, and body movements; construct a fusion function, which is used to integrate multiple feature vectors in the feature vector set into a unified feature representation. The fusion function is a multi-layer neural network that maps data from different modalities to a unified feature space through the weight matrix in the neural network; define a loss function to evaluate the difference between the generated digital human behavior and the preset target; adjust the loss function parameters through an optimization algorithm to minimize the loss function value, where the optimization algorithm includes using stochastic gradient descent or the Adam algorithm; Digital human performance generation steps: Set a conditional vector to represent a specific emotion or state; receive random noise, the conditional vector, and a unified feature representation of multimodal data through the generator, and generate facial expression or voice data corresponding to the set emotion or state; use the discriminator to evaluate the match between the generated data and the conditional vector and determine its authenticity; optimize the generator and discriminator by minimizing the adversarial loss function to improve the authenticity and match of the generated data, and construct a performance data synthesizer; Digital human emotion recognition steps: Obtain a user's facial image and extract features from the user's facial image using a pre-trained convolutional neural network model; Input the extracted feature vector into a classifier to identify the user's specific expression. The classifier is a support vector machine or a multi-layer perceptron; Based on the recognized expression category, the digital human generates a corresponding facial expression or voice response; Digital Human response optimization step: receiving user input information, which includes voice commands or text information; processing the user input and generating a corresponding response through a preset response function; adjusting the parameters of the response function through feedback to ensure the accuracy and timeliness of the response; In the digital human response optimization step, the user's input data occurs at time t, where For voice commands or text messages, represented as ; Use the reaction function R to generate the digital human's response , the reaction function R is expressed as: ,in, is the reaction function Parameters, reaction function The parameters of are optimized through training data.
2. The method for generating an interactive digital human based on artificial intelligence according to claim 1, characterized in that: In the digital human behavior generation step, feature vectors of facial expressions, voice, and body movements are extracted from real-time video and audio data. The fusion function adopts a combination of convolutional layers and fully connected layers to process and integrate input data of different modalities. The loss function is cross entropy or mean square error, which is used to evaluate the consistency between the output predicted by the model and the actual label or expected output.
3. An artificial intelligence-based interactive digital human generation system, based on the artificial intelligence-based interactive digital human generation method according to claim 1 or 2, characterized in that: Includes the following modules: Digital Human Image Generation Module: This module is used to define the input conditions for the digital human. The input conditions are represented by a conditional vector that describes the basic characteristics and generation requirements of the digital human. The conditional vector is input into the encoder to generate a latent space representation. The sampled latent vector is passed through the decoder to generate a preliminary digital human image. Reconstruction loss is used to ensure that the generated image meets the characteristics of the input condition. KL divergence is used to ensure the continuity of the latent space distribution. Based on the preliminary generated digital human image, style transfer technology is applied to adapt it to the target environment or style. A super-resolution network is used to improve the resolution and detail quality of the image, and the final high-resolution image is output; Digital Human Behavior Generation Module: This module is used to collect multimodal data and convert it into a corresponding feature vector set. The multimodal data includes facial expressions, speech, and body movements. It constructs a fusion function to integrate multiple feature vectors in the feature vector set into a unified feature representation. The fusion function is a multi-layer neural network that maps data from different modalities to a unified feature space through the weight matrix in the neural network. It also defines a loss function to evaluate the difference between the generated digital human behavior and the preset target. The loss function parameters are adjusted through an optimization algorithm to minimize the loss function value. The optimization algorithm may include stochastic gradient descent or the Adam algorithm. Digital Human Performance Generation Module: This module sets a conditional vector representing a specific emotion or state. The generator receives random noise, the conditional vector, and a unified feature representation of multimodal data to generate facial expression or voice data corresponding to the set emotion or state. A discriminator is used to evaluate the match between the generated data and the conditional vector and determine its authenticity. The generator and discriminator are optimized by minimizing the adversarial loss function to improve the authenticity and match of the generated data, thereby constructing a performance data synthesizer. Digital Human Emotion Recognition Module: This module is used to obtain a user's facial image and extract features from the user's facial image using a pre-trained convolutional neural network model. The extracted feature vector is input into a classifier, such as a support vector machine or a multi-layer perceptron, to identify the user's specific expression. Based on the recognized expression category, the digital human generates a corresponding facial expression or voice response. Digital Human Response Optimization Module: This module is used to receive user input, including voice commands or text messages; process the user input and generate a corresponding response through a preset response function; and adjust the parameters of the response function through feedback to ensure the accuracy and timeliness of the response.
Citation Information
Patent Citations
Method and system for automatically generating digital human
CN117173294A
Digital human generation method and device and storage medium
CN118916461A