Digital human generation method and system based on AI interaction information
By obtaining multimodal interaction information and real-time facial data, and combining AI models to generate dynamic images of digital human virtual images, the authenticity and personalization of interaction between digital human virtual images and users and audiences is solved, and synchronous expression of emotions and interaction is achieved.
Patent Information
- Application Number
- CN202510884750.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In the prior art, it is difficult for digital human virtual images to achieve dynamic facial expression drive that echoes the user's actual emotional state and the audience's emotional interaction, resulting in a lack of authenticity and personalization of emotional expression, which affects the user's sense of substitution and the audience's interaction experience.
By obtaining the user's multimodal interaction information, generating multi-dimensional identification features, combining pre-trained AI language model and image generation model, real-time facial acquisition data is obtained for emotional feature extraction, and dynamic images of digital human virtual images are generated based on historical emotional characteristics and barrage interaction information, and fusion decisions are made to adjust facial emotions and achieve dynamic interaction with users and audiences.
It improves the authenticity and interactivity of digital human expressions, enhances the synchronous expression of user emotions and the feedback of audience emotions, and improves the social interaction experience between digital humans and audiences.
Smart Images

Figure CN120388115A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of digital human image generation, and more specifically, to a digital human generation method and system based on AI interaction information. Background Art
[0002] In recent years, it has developed rapidly under the promotion of the technology, culture, and entertainment industries, showing a diverse and innovative development trend. Its biggest feature is that real human hosts conduct live broadcast activities in the form of virtual digital human images, and the virtual digital human image is the exclusive visual image created by the host relying on digital technology. With the rapid development of virtual human technology, digital humans generated based on artificial intelligence have also been widely used in scenarios such as virtual live broadcast, digital marketing, and remote interaction.
[0003] In the prior art, the virtual digital human image usually expresses emotions through a pre-set language model, and it is difficult to achieve dynamic facial expression driving that corresponds to the actual emotional state of the user and the emotional interaction with the audience. This results in the lack of authenticity and personalization in the emotional expression shown by the digital human, affecting the user's sense of immersion and the audience's interaction experience. How to improve the naturalness and interactivity of the virtual digital human image has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a digital human generation method and system based on AI interaction information, which can dynamically adjust the facial emotions of the digital human in response to the audience's feedback, improving the authenticity and interactivity of the digital human's expression.
[0005] In a first aspect, this application provides a digital human generation method based on AI interaction information. This method can be executed by a network device, or it can also be executed by a chip configured in the network device. This application does not make any limitations in this regard.
[0006] Specifically, this method includes:
[0007] Obtain the multimodal interaction information of the user, and generate the multi-dimensional identification features of the user based on the multimodal interaction information;
[0008] Based on the pre-trained AI language model and image generation model, generate the virtual digital human image through the multi-dimensional identification features;
[0009] Obtain the facial acquisition data of the user in real time, extract the real-time emotion features according to the facial acquisition data, and generate the first facial dynamic image of the virtual digital human image according to the real-time emotion features;
[0010] Predict the emotional trajectory based on the user's historical emotional characteristics to obtain the initial emotional prediction trajectory. Extract the emotional influence factors according to the historical barrage interaction information and historical emotional characteristics, collect the barrage distribution time window, and generate the second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotional influence factors, and the initial emotional prediction trajectory;
[0011] Based on the first facial dynamic image and the second facial dynamic image, perform a fusion decision to generate a digital human dynamic image for live broadcast push.
[0012] Combined with the first aspect, in some implementation manners of the first aspect, generating the multi-dimensional identification features of the user based on the multi-modal interaction information specifically includes:
[0013] For any group of interaction information of the multi-modal interaction information, respectively extract the original feature digital representation through the feature extraction network matching with this modality;
[0014] For the original feature digital representation corresponding to any group of modalities, respectively perform a unified mapping to the specified dimension in the feature space through feature transformation to obtain multiple modality features, and align and fuse all the modality features in the feature space according to the preset dimension to obtain the multi-dimensional feature vector of the user as the multi-dimensional identification feature.
[0015] Combined with the first aspect, in some implementation manners of the first aspect, predicting the emotional trajectory according to the user's historical emotional characteristics to obtain the initial emotional prediction trajectory specifically includes:
[0016] Real-time record the facial emotional characteristics of the user during the live broadcast to form a time series sample set of historical emotional characteristics;
[0017] Construct a moving average autoregressive model of emotional characteristics according to the time series sample set of historical emotional characteristics, and determine the initial emotional prediction trajectory based on the moving average autoregressive model of emotional characteristics.
[0018] Combined with the first aspect, in some implementation manners of the first aspect, before extracting the emotional influence factors according to the historical barrage interaction information and historical emotional characteristics, it further includes: storing the barrage information received by the user during the live broadcast to obtain the historical barrage interaction information.
[0019] Combined with the first aspect, in some implementation manners of the first aspect, extracting the emotional influence factors according to the historical barrage interaction information and historical emotional characteristics specifically includes:
[0020] Use a natural language processing model to perform emotional classification on the historical barrage interaction information, count the distribution of barrage emotional types within the set time window, and generate a barrage emotional distribution sequence;
[0021] Obtain the time series of historical emotion features, and determine the emotion influence factor based on the sequence correlation between the barrage emotion distribution sequence and the time series of the historical emotion features.
[0022] Combined with the first aspect, in some implementation manners of the first aspect, generating the second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotion influence factor, and the initial emotion prediction trajectory specifically includes:
[0023] Obtain the initial emotion prediction distribution through the initial emotion prediction trajectory;
[0024] Extract the barrage emotion features based on the barrage distribution time window, and construct an emotion influence prediction distribution based on the barrage emotion features and the emotion influence factor;
[0025] Obtain the fusion weight of the initial emotion prediction distribution and the emotion influence prediction distribution, and perform weighted fusion on the initial emotion prediction distribution and the emotion influence prediction distribution to determine the final emotion prediction result;
[0026] Input the emotion prediction result and the static digital human virtual image into an expression driving model to generate the second facial dynamic image of the digital human virtual image.
[0027] Combined with the first aspect, in some implementation manners of the first aspect, performing a fusion decision based on the first facial dynamic image and the second facial dynamic image, and generating a digital human dynamic image for live broadcast push specifically includes:
[0028] Obtain the emotion consistency of the first facial dynamic image and the second facial dynamic image;
[0029] When the emotion consistency is higher than a preset threshold, extract the motion parameters of the first facial dynamic image and the second facial dynamic image respectively, perform weighted fusion on the motion parameters through a preset weight, and generate the final facial motion data;
[0030] When the emotion consistency is lower than the preset threshold, use a fuzzy decision algorithm to select the facial dynamic image, and determine the final facial motion data based on the decision result;
[0031] Generate a digital human dynamic image based on the final facial motion data and perform live broadcast push.
[0032] In a second aspect, the present application provides a digital human generation system based on AI interaction information, which includes a digital human image generation unit, and the digital human image generation unit includes:
[0033] A user information extraction module, configured to obtain multimodal interaction information of a user and generate multi-dimensional identification features of the user based on the multimodal interaction information;
[0034] A virtual image generation module, configured to generate a digital human virtual image based on a pre-trained AI language model and an image generation model through the multi-dimensional identification features;
[0035] A dynamic image generation module, configured to obtain facial acquisition data of a user in real time, extract real-time emotion features according to the facial acquisition data, and generate a first facial dynamic image of the digital human virtual image according to the real-time emotion features;
[0036] The dynamic image generation module is further configured to predict an emotion trajectory according to historical emotion features of the user to obtain an initial emotion prediction trajectory, extract emotion influence factors according to historical barrage interaction information and historical emotion features, collect a barrage distribution time window, and generate a second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotion influence factors, and the initial emotion prediction trajectory;
[0037] The dynamic image generation module is further configured to perform a fusion decision based on the first facial dynamic image and the second facial dynamic image, generate a digital human dynamic image, and use it for live broadcast push.
[0038] In a third aspect, the present application provides a computer terminal device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned digital human generation method based on AI interaction information.
[0039] In a fourth aspect, the present application provides a computer-readable storage medium, which stores at least one computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned digital human generation method based on AI interaction information.
[0040] The technical solutions provided by the disclosed embodiments of the present application have the following beneficial effects:
[0041] In a method and system for generating a digital human based on AI interaction information provided by this application, first, multimodal interaction information of a user is obtained, and multidimensional identification features of the user are generated based on the multimodal interaction information; based on a pre-trained AI language model and an image generation model, a digital human virtual image is generated through the multidimensional identification features; facial capture data of the user is obtained in real time, real-time emotion features are extracted according to the facial capture data, and a first facial dynamic image of the digital human virtual image is generated according to the real-time emotion features; an emotion trajectory prediction is performed according to the historical emotion features of the user to obtain an initial emotion prediction trajectory, an emotion influence factor is extracted according to the historical barrage interaction information and the historical emotion features, a barrage distribution time window is collected, and a second facial dynamic image of the digital human virtual image is generated based on the barrage distribution time window, the emotion influence factor, and the initial emotion prediction trajectory; a fusion decision is made based on the first facial dynamic image and the second facial dynamic image to generate a digital human dynamic image for live broadcast push.
[0042] Therefore, it can be seen that this application extracts the current emotion features of the user from the real-time facial capture data, and uses an expression-driven model to map this emotion to generate a first facial dynamic image, realizing the synchronous expression of the user's current real emotion on the digital human, ensuring that the emotion output of the digital human stems from the real user state, being close to the immediate expression of the user himself, enhancing the authenticity and emotional immersion. Furthermore, an audience barrage interaction emotion analysis mechanism is introduced. By classifying the barrage emotions and modeling the emotion distribution within the time window, an emotion influence factor model of the audience emotions is constructed. Then, combined with the emotion trajectory prediction result, a second facial dynamic image is generated to reflect the regulatory effect of the overall audience emotion on the emotion performance of the digital human, so as to be able to perceive and respond to the audience's emotion feedback, thereby realizing a dynamic expression with social interaction awareness. Finally, through a fusion decision on the first facial dynamic image and the second facial dynamic image, the facial emotion of the digital human is dynamically balanced between the user's expression and the audience's feedback, thus avoiding conflicts and improving the naturalness and interaction consistency.
[0043] In summary, this application can dynamically adjust the facial emotion of the digital human in response to the audience feedback, improving the authenticity and interactivity of the digital human's expression. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is an exemplary flowchart of a method for generating a digital human based on AI interaction information according to some embodiments of this application;
[0045] Figure 2 is a schematic structural diagram of a digital human image generation unit according to some embodiments of this application;
[0046] Figure 3Schematic diagram of a computer terminal device for implementing a digital human generation method based on AI interaction information according to some embodiments of the present application. Detailed implementation manners
[0047] In the present application, multi-modal interaction information of a user is obtained, and multi-dimensional identification features of the user are generated based on the multi-modal interaction information; based on a pre-trained AI language model and an image generation model, a digital human virtual image is generated through the multi-dimensional identification features; facial acquisition data of the user is obtained in real time, real-time emotion features are extracted according to the facial acquisition data, and a first facial dynamic image of the digital human virtual image is generated according to the real-time emotion features; an emotion trajectory prediction is performed according to the historical emotion features of the user to obtain an initial emotion prediction trajectory, an emotion influence factor is extracted according to the historical barrage interaction information and the historical emotion features, a barrage distribution time window is collected, and a second facial dynamic image of the digital human virtual image is generated based on the barrage distribution time window, the emotion influence factor, and the initial emotion prediction trajectory; a fusion decision is made based on the first facial dynamic image and the second facial dynamic image to generate a digital human dynamic image for live broadcast push, which can dynamically adjust the facial emotion of the digital human for feedback to the audience, improving the authenticity and interactivity of the digital human expression.
[0048] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the specification drawings and specific implementation manners. Refer to Figure 1 , this figure is an exemplary flowchart of a digital human generation method based on AI interaction information according to some embodiments of the present application. The digital human generation method 100 based on AI interaction information mainly includes the following steps:
[0049] In step S101, multi-modal interaction information of the user is obtained, and multi-dimensional identification features of the user are generated based on the multi-modal interaction information.
[0050] Optionally, in some embodiments, the multi-modal interaction information includes: voice modality interaction information, text modality interaction information, image modality interaction information, and behavior modality interaction information. Among them, the voice modality interaction information can identify the user's voice through a voice recognition model, and extract the Mel-frequency cepstral coefficients of the user's voice as the voice modality interaction information; the text modality interaction information can extract semantic vectors related to the context of the user's voice text, and statistically analyze the emotion features of the language style as the text modality interaction information. The image modality interaction information includes the user's facial embedding vector, facial expression capture information, and expression state distribution probability; the behavior modality includes the operation frequency and interaction duration of the user on the live broadcast interface, etc.
[0051] Preferably, in some embodiments, generating the multi-dimensional identification features of the user based on the multi-modal interaction information specifically includes:
[0052] For the interaction information of any group of modalities in the multi-modal interaction information, the original feature digital representation is respectively extracted through a feature extraction network matching the modality.
[0053] For the original feature digital representation corresponding to any group of modalities, they are respectively mapped to the specified dimension in the feature space through feature transformation to obtain multiple modality features. All the modality features are aligned and fused in the feature space according to the preset dimension to obtain the multi-dimensional feature vector of the user as the multi-dimensional identification feature.
[0054] In specific implementation, the original feature digital representation is a digital vector or tensor that can be processed by a computer. It is the first output result of feature extraction and is the underlying feature expression form extracted from the original modality data. For example, the original feature digital representation of the speech modality can adopt a matrix with a dimension of 13×the number of time frames. The original feature digital representations corresponding to other modalities are extracted using corresponding feature extraction algorithms, which will not be elaborated in this application. Furthermore, a normalized linear transformation strategy can be adopted as the feature transformation method for the original feature digital representation for feature space mapping, and all the modality features are aligned and fused in the feature space according to the preset dimension to obtain the multi-dimensional feature vector of the user as the multi-dimensional identification feature.
[0055] Optionally, in some embodiments, a linear multi-dimensional space allocation is performed based on the information capacity of the interaction information of each modality in the multi-modal interaction information. For example, dimensions 0 to 128 of the multi-dimensional identification feature can be used to store speech modality features, dimensions 129 to 384 of the multi-dimensional identification feature can be used to store speech modality features, dimensions 385 to 896 of the multi-dimensional identification feature can be used to store facial modality features, and dimensions 897 to 960 of the multi-dimensional identification feature can be used to store behavior modality features.
[0056] In step S102, based on the pre-trained AI language model and image generation model, a digital human virtual image is generated through the multi-dimensional identification feature.
[0057] Preferably, in some embodiments, generating a digital human virtual image through the multi-dimensional identification feature based on the pre-trained AI language model and image generation model specifically includes: using the multi-dimensional identification feature of the user as a conditional control vector, inputting it into the pre-trained language expression model to generate text information containing the user's language habits, and synthesizing speech content with the user's voice characteristics by a speech synthesis model; inputting the facial modality feature in the multi-dimensional identification feature into the image generation model to generate a digital human virtual image, driving lip movement based on the speech information, and changing expressions based on real-time emotion features, and finally constructing a dynamic digital human image for live interaction applications.
[0058] In specific implementation, the pre-trained AI language model and image generation model in this application can adopt existing technologies. For example, in one embodiment of this application, the pre-trained AI language model can adopt a combination of a text generation model and a text-to-speech (TTS) model. The image generation model can combine the text-driven image algorithm of Stable Diffusion with the structure control algorithm of ControlNet to generate a digital human virtual image by fusing text descriptions and facial features. Algorithms or scripts capable of implementing speech audio synthesis and virtual image generation in other existing technologies can also be applied, and this application does not make any limitations in this regard.
[0059] In step S103, the facial acquisition data of the user is acquired in real time, the real-time emotion features are extracted according to the facial acquisition data, and the first facial dynamic image of the digital human virtual image is generated according to the real-time emotion features.
[0060] Preferably, in some embodiments, during the process of acquiring the facial acquisition data of the user in real time, the facial image data of the user is acquired in real time through the user's live camera, and the facial key point coordinate information is extracted through the facial key point detection model as the facial acquisition data of the user.
[0061] Specifically, the facial key point detection model can adopt the currently publicly available high-robustness facial expression analysis tool OpenFace, which has the ability to perform key point positioning, head pose estimation, eye movement tracking, and facial action unit (ActionUnit, AU) intensity analysis on the face region in RGB images; the OpenFace algorithm can accurately extract 68 three-dimensional facial key points in the user's facial image, including key regions such as the mouth, eyebrows, eyes, nose bridge, and face contour.
[0062] Furthermore, based on the extracted facial key point coordinates, the system can construct the current facial feature vector of the user and input it into the emotion recognition module to extract the real-time emotion state; for example, by extracting the facial action units in OpenFace, such as the intensity scores of AU06 (raising eyebrows), AU12 (smile), etc., and combining the rule model or the deep learning emotion classification model, the facial expression can be mapped into a multi-dimensional emotion vector, such as happy, surprised, bored, etc.; the emotion vector can be used as the input information to drive the change of the digital human virtual image, realizing the digital human facial dynamic expression corresponding to the user's current facial state.
[0063] In addition, to ensure the stability and real-time performance of the acquisition process, the system can set a face detection threshold and a tracking cache mechanism to adapt to the face recognition effect under different lighting conditions, angles, and backgrounds, ensuring the continuity and robustness of the extracted facial acquisition data. The acquisition and processing process can support a processing rate of no less than 25 frames per second, meeting the requirements of real-time performance and naturalness for digital human live broadcast applications.
[0064] Preferably, in some embodiments, a long short-term memory artificial neural network can be used as the emotion recognition model to extract real-time emotion features based on the facial acquisition data.
[0065] It should be noted that the emotion recognition model can adopt a long short-term memory artificial neural network (Long Short-Term Memory, LSTM). This model is suitable for processing the facial dynamic data presented by users in a continuous time series, thereby achieving more accurate real-time emotion feature extraction. Specifically, the long short-term memory artificial neural network is used to model and classify the time series features in the facial acquisition data of users. The facial acquisition data can include information such as multi-frame continuous facial key point coordinates, action unit intensities, and head postures extracted through OpenFace or similar facial key point detection algorithms. The multi-dimensional time series data can be used as the input of the long short-term memory artificial neural network.
[0066] In the training stage, the system can construct a set of facial expression datasets as training samples. Each sample contains a facial feature sequence of multiple continuous frames and its corresponding emotion label, such as: happy, surprised, angry, sad, etc. The long short-term memory artificial neural network retains long-term dependence information through its unique gating mechanism (input gate, forget gate, output gate), effectively capturing the subtle changes in the user's facial expressions over time, and improving the recognition ability for complex emotion states, such as shyness, boredom, and confusion.
[0067] In the actual application stage, when the system receives the real-time facial data stream from the camera, it will first convert it into a standardized time series tensor form, such as [time length × feature dimension], and input it into the trained long short-term memory artificial neural network. The output of the model is an emotion feature vector or a probability distribution vector, representing the emotion state and corresponding confidence at the current time point. For example, the model outputs the following vector:
[0068] [Happy: 0.72, Surprised: 0.10, Sad: 0.08, Angry: 0.05, Calm: 0.05].
[0069] The system can select the current dominant emotion according to this vector, which is further used to drive the facial dynamic expression of the digital human virtual image. Further, to enhance the model's ability to capture the timing of complex expressions, a long short-term memory artificial neural network structure can be combined with a convolutional neural network (CNN) to form a CNN-LSTM emotion recognition architecture, where the convolutional neural network is used to extract spatial features from facial images or AU matrices, and the long short-term memory artificial neural network is responsible for processing its time evolution process; this combined structure has been proven to have high recognition accuracy for micro-expressions and atypical emotional states in existing public research and is suitable for real-time digital human emotion recognition tasks in live interaction scenarios.
[0070] Preferably, in some embodiments, generating the first facial dynamic image of the digital human virtual image according to the real-time emotion characteristics specifically includes: inputting the real-time emotion characteristics and the static digital human virtual image into an expression driving model, and generating the first facial dynamic image through the expression driving model.
[0071] In specific implementation, the user emotion feature vector extracted in real time and the preset static digital human virtual image are input into the expression driving model, and the first facial dynamic image matching the user's current emotion state is generated through this model, so as to realize the dynamic deduction of the synchronization of the digital human's facial expression and the user's emotion. For example, the expression driving model can adopt technical structures such as a generative adversarial network, an expression transfer network, or a 3D expression reconstruction model, and preferably use mature deep generative models such as Audio2Face, which has the ability to map expression parameters or emotion vectors into dynamic face images.
[0072] The input of the model includes: Static digital human virtual image: namely, the basic head image or 3D head model of the target digital human, used as the appearance reference for the generation process; Real-time emotion feature vector: the current user emotion state extracted by models such as the aforementioned long short-term memory artificial neural network, such as the confidence distribution vector or low-dimensional embedding feature of seven types of emotion labels, representing the category and intensity of facial expression performance. Among them, to achieve continuous and natural transition of expressions, the system can first post-process the input emotion feature vector, such as: standardization and smoothing filtering to eliminate mutations; Feature mapping, mapping the high-dimensional emotion vector to the facial expression action space, such as the AU coding space; Temporal modeling, smoothing the transition changes through frame interpolation. In the expression driving stage, the system generates dynamic images in the following way: input the emotion feature as a control signal into the action control branch of the expression driving model; The model internally combines the facial structure, material, and texture features of the static digital human image to simulate the expression changes and reconstruct the image; Finally, output a sequence of facial images of consecutive frames to form a dynamic digital human expression reflecting the current user's emotion, such as smiling, frowning, being surprised, etc.; Preferably, if the static digital human virtual image is a 3D model, the emotion feature can also be mapped to facial blendshape weights or 3D point position transformation parameters to drive the deformation of the virtual facial mesh and achieve structural-level dynamic expression synthesis.
[0073] In some specific embodiments of the present application, to meet the real-time requirement of live broadcast push, the generation rate of the first facial dynamic image may not be lower than 20 frames per second, and the output image size can be configured as 128×128, 256×256 or higher resolution according to application requirements, and is efficiently transmitted by encapsulating it into the live broadcast frame stream through a compression algorithm, such as H.264 / H.265.
[0074] In the solution of the present application, through real-time facial acquisition data such as facial key points, expression action units, etc., the emotion recognition model is used to extract the current emotion features of the user, and the expression driving model is used to map this emotion to generate the first facial dynamic image, realizing the synchronous expression of the user's real emotion at this moment on the digital human.
[0075] In step S104, based on the historical emotion features of the user, emotion trajectory prediction is performed to obtain an initial emotion prediction trajectory. Emotion influence factors are extracted according to the historical barrage interaction information and historical emotion features, and a barrage distribution time window is collected. Based on the barrage distribution time window, the emotion influence factors, and the initial emotion prediction trajectory, a second facial dynamic image of the digital human virtual image is generated.
[0076] Preferably, in some embodiments, performing emotion trajectory prediction based on the historical emotion features of the user to obtain an initial emotion prediction trajectory specifically includes:
[0077] Record the facial emotion features of users during the live broadcast in real time to form a time series sample set of historical emotion features;
[0078] Construct a moving average autoregressive model of emotion features based on the time series sample set of the historical emotion features, and determine an initial emotion prediction trajectory based on the moving average autoregressive model of emotion features.
[0079] Preferably, in some embodiments, the time series of the historical emotion features is obtained by recording the historical real-time emotion features and the corresponding time tags, and then obtaining the dominant emotion type of the real-time emotion features for digital mapping to obtain emotion feature values, and forming a time series of the historical emotion features according to the emotion feature values and the corresponding time tags, wherein the initial emotion prediction trajectory is a sequence of emotion feature values for a future period of time determined based on the moving average autoregressive model.
[0080] Optionally, in some embodiments, before extracting the emotion influence factor according to the historical barrage interaction information and the historical emotion features, it further includes: storing the barrage information received by the user during the live broadcast to obtain the historical barrage interaction information.
[0081] It should be noted that the emotion influence factor in this application represents the intensity of the effect of the external barrage emotion on the internal emotion change of the user. Preferably, in some embodiments, extracting the emotion influence factor according to the historical barrage interaction information and the historical emotion features specifically includes:
[0082] Use a natural language processing model to classify the emotion of the historical barrage interaction information, count the distribution of the barrage emotion types within a set time window, and generate a barrage emotion distribution sequence;
[0083] Obtain the time series of the historical emotion features, and determine the emotion influence factor based on the sequence correlation between the barrage emotion distribution sequence and the time series of the historical emotion features.
[0084] In specific implementation, the BERT-wwm-ext algorithm can be adopted as a natural language processing model to classify the emotions of the historical barrage interaction information, and generate a barrage emotion distribution value based on the distribution of barrage emotion types. Among them, the barrage information is divided into time periods according to a set time granularity, such as every second, and mapped based on the threshold range where the barrage emotion distribution is located in different time periods to obtain the barrage emotion distribution value. Optionally, in some embodiments, the barrage emotion classification types can be calibrated as four emotion distribution types: positive, negative, encouraging, and ironic. Then, the distribution ratios of different types of emotions are weighted and fused according to the preset emotion weights of different types of emotions to obtain the final barrage emotion distribution value and the corresponding time tag, and a barrage emotion distribution sequence is formed according to the time sequence. For example, in some specific embodiments of the present application, the weights of the four emotion distribution types of positive, negative, encouraging, and ironic are +1.0, -0.8, +0.8, and -0.6 respectively.
[0085] Preferably, in some embodiments, in the process of determining the emotion influence factor based on the sequence correlation between the barrage emotion distribution sequence and the time sequence of the historical emotion features, a difference sequence of the barrage emotion distribution sequence is obtained based on the time tag. After sequence alignment and normalization of the difference sequence of the barrage emotion distribution sequence and the time sequence of the historical emotion features, the Pearson correlation coefficient between the difference sequence of the barrage emotion distribution sequence and the time sequence of the historical emotion features is used as the emotion influence factor.
[0086] In some specific embodiments of the present application, the initial emotion prediction distribution and the emotion influence prediction distribution are calculated based on information from different sources, and the final emotion prediction result is obtained through fusion. Specifically, the initial emotion prediction distribution is generated according to the user's facial emotion features at the target moment and the prediction result of the emotion prediction model, while the emotion influence prediction distribution is calculated by combining the barrage emotion features and the emotion influence factor within the barrage distribution time window considering the potential influence of the audience's emotions on the digital human virtual image.
[0087] It should be noted that the barrage distribution time window in the present application refers to a fixed-length time interval containing the statistical distribution of barrage types set for time series analysis and feature extraction when processing the barrage information sent by the audience in a live broadcast or video. Optionally, in some embodiments, generating the second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotion influence factor, and the initial emotion prediction trajectory specifically includes:
[0088] Obtaining the initial emotion prediction distribution through the initial emotion prediction trajectory;
[0089] Extract the barrage emotion features based on the barrage distribution time window, and construct an emotion influence prediction distribution based on the barrage emotion features and the emotion influence factor;
[0090] Obtain the fusion weight of the initial emotion prediction distribution and the emotion influence prediction distribution, and perform weighted fusion on the initial emotion prediction distribution and the emotion influence prediction distribution to determine the final emotion prediction result;
[0091] Input the emotion prediction result and the static digital human virtual image into an expression driving model to generate a second facial dynamic image of the digital human virtual image.
[0092] It should be noted that in this application, the initial emotion prediction distribution is the prediction probability distribution of the user's facial emotion at the target moment, which reflects the user's emotion tendency at the target moment. Among them, the initial emotion prediction trajectory contains emotion feature values within a period of time in the future. By reflecting the emotion feature value corresponding to the target moment, the corresponding initial emotion prediction distribution can be obtained. Among them, the proportion of the emotion type distribution in the barrage distribution time window closest to the target moment is statistically composed into an emotion distribution vector, and the emotion influence factor is used as a proportional coefficient to correct the emotion distribution vector to obtain the final mapped emotion prediction distribution.
[0093] In step S105, based on the first facial dynamic image and the second facial dynamic image, perform a fusion decision to generate a digital human dynamic image and use it for live broadcast push.
[0094] It should be noted that the first facial dynamic image is a dynamic facial expression image generated according to real-time emotion features, which reflects the facial expression of the user in the current emotional state. The second facial dynamic image is an image generated by weighting historical emotion features, barrage emotion features and emotion influence factors, taking into account the emotional feedback of the audience and the emotional trajectory of the user.
[0095] Optionally, in some embodiments, performing a fusion decision based on the first facial dynamic image and the second facial dynamic image to generate a digital human dynamic image and use it for live broadcast push specifically includes:
[0096] Obtain the emotion consistency of the first facial dynamic image and the second facial dynamic image;
[0097] When the emotion consistency is higher than a preset threshold, respectively extract the action parameters of the first facial dynamic image and the second facial dynamic image, perform weighted fusion on the action parameters through a preset weight, and generate the final facial action data.
[0098] When the emotional consistency is lower than a preset threshold, a fuzzy decision-making algorithm is used to select facial dynamic images, and the final facial action data is determined based on the decision result;
[0099] Generate a digital human dynamic image based on the final facial action data and perform live broadcast push.
[0100] Preferably, in some embodiments, a CNN+LSTM neural network model can be used to classify the emotions of the first facial dynamic image and the second facial dynamic image, obtain corresponding emotion labels (such as: happy, surprised, angry, etc.) and confidence distributions, and then use a cosine similarity model to judge the emotional consistency of the emotion labels corresponding to the first facial dynamic image and the second facial dynamic image. For example, the reciprocal of the cosine similarity between the confidence distributions of the emotion labels is used as the emotional consistency.
[0101] Optionally, in some embodiments, when the emotional consistency is higher than a preset threshold, the OpenFace facial motion capture model can be used to extract the motion parameters of the first facial dynamic image and the second facial dynamic image, set the real-time emotion weight of the first facial dynamic image to 0.6, set the emotion influence weight of the second facial dynamic image to 0.4, perform weighted fusion on the motion parameters of the first facial dynamic image and the second facial dynamic image, and input the fused motion parameters into the FLAME expression driving model or other existing expression driving models to generate the final dynamic image frame.
[0102] The following gives a specific embodiment of the application using a fuzzy decision-making algorithm to select facial dynamic images: Build a fuzzy logic inference system with three inputs and one output, specifically including:
[0103] Input variables:
[0104] Confidence 1: The emotion recognition confidence of the first facial dynamic image, which represents the credibility of the emotion obtained based on real-time user facial data analysis. In this application, the preset emotion recognition confidence of the first facial dynamic image is 0.7. Confidence 2: The emotion recognition confidence of the second facial dynamic image, which represents the credibility of the emotion generated based on bullet screen deduction. It is determined based on the cosine similarity between the motion parameters of the second facial dynamic image and the real-time collected user facial motion parameters. Audience emotion weight: Used to measure the emotional influence intensity of the audience interaction content, which can be quantified by the value of the emotion influence factor.
[0105] Output variable:
[0106] Image priority selection value A: The value range is [0, 1], indicating the weight preference for the final image selection. A approaching 0 means preferring to adopt the first facial dynamic image, A approaching 1 means preferring the second facial dynamic image; the intermediate value means image fusion.
[0107] During the process of setting the fuzzy rules, the system makes inference and judgment based on the following typical fuzzy rules:
[0108] If the confidence level 1 is high and the audience emotion weight is low, the image priority selection value A tends to 0 (select the first image); if the confidence level 2 is high and the audience emotion weight is high, the image priority selection value A tends to 1 (select the second image); if the difference between the confidence level 1 and the confidence level 2 is not large, or the audience emotion weight is medium, the image priority selection value A takes the intermediate value (perform image fusion); in some other embodiments of the present application, the membership functions of high, medium, and low can be further defined, and the inference is completed using the fuzzy inference method, and the present application does not limit this.
[0109] In addition, on the other hand of the present application, in some embodiments, the present application provides a digital human generation system based on AI interaction information. The system includes a digital human image generation unit. Refer to Figure 2 , which is a schematic diagram of the exemplary hardware and / or software structure of the digital human image generation unit shown according to some embodiments of the present application. The digital human image generation unit 200 includes: a user information extraction module 201, a virtual image generation module 202, and a dynamic image generation module 203, which are described as follows:
[0110] The user information extraction module 201 is used to obtain the multimodal interaction information of the user and generate the multi-dimensional identification features of the user based on the multimodal interaction information;
[0111] The virtual image generation module 202 is used to generate a digital human virtual image based on the pre-trained AI language model and image generation model through the multi-dimensional identification features;
[0112] The dynamic image generation module 203 is used to obtain the facial acquisition data of the user in real time, extract the real-time emotion features according to the facial acquisition data, and generate the first facial dynamic image of the digital human virtual image according to the real-time emotion features;
[0113] The dynamic image generation module 203 is further used to predict the emotion trajectory according to the historical emotion features of the user to obtain an initial emotion prediction trajectory, extract the emotion influence factor according to the historical barrage interaction information and historical emotion features, collect the barrage distribution time window, and generate the second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotion influence factor, and the initial emotion prediction trajectory;
[0114] The dynamic image generation module 203 is further configured to make a fusion decision based on the first facial dynamic image and the second facial dynamic image, generate a digital human dynamic image and use it for live broadcast push.
[0115] The above has introduced in detail an example of a digital human generation method and system based on AI interaction information provided by the embodiments of the present application. It can be understood that, in order to implement the above functions, the corresponding device includes a hardware structure and / or software module that executes each function.
[0116] Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function in the application is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Therefore, professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0117] In addition, the present application also provides a computer terminal device, the computer terminal device includes a memory and a processor, the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned digital human generation method based on AI interaction information.
[0118] In some embodiments, refer to Figure 3 , this figure is a schematic structural diagram of a computer terminal device for implementing a digital human generation method based on AI interaction information according to some embodiments of the present application. The above-mentioned digital human generation method based on AI interaction information in the above embodiments can be implemented by Figure 3 The shown computer terminal device, the computer terminal device 300 includes at least one communication bus 301, a communication interface 302, a processor 303, and a memory 304.
[0119] The processor 303 can be a general central processing unit (CPU), or a specific application integrated circuit (ASIC), or one or more used to control the execution of a digital human generation method based on AI interaction information in the present application.
[0120] The communication bus 301 may include a path for transmitting information between the above components.
[0121] The memory 304 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 304 can exist independently and be connected to the processor 303 through the communication bus 301. The memory 304 can also be integrated with the processor 303.
[0122] Among them, the memory 304 is used to store the program code for executing the solution of this application and is controlled by the processor 303 for execution. The processor 303 is used to execute the program code stored in the memory 304. The program code can include one or more software modules. The generation of the first facial dynamic image in the above embodiments can be implemented by one or more software modules in the program code in the processor 303 and the memory 304.
[0123] The communication interface 302 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0124] Optionally, the above computer terminal device 300 can further include a power supply 305 for supplying power to various components or circuits in the real-time computer terminal device.
[0125] In a specific implementation, as an embodiment, the computer terminal device can include multiple processors, and each of the processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0126] The above computer terminal device can be a general computer terminal device or a dedicated computer terminal device. In specific implementations, the computer terminal device can be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of the present application do not limit the type of the computer terminal device.
[0127] In addition, in other aspects of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores at least one computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the above-mentioned method for generating a digital human based on AI interaction information.
[0128] In summary, in the method and system for generating a digital human based on AI interaction information disclosed in the embodiments of the present application, first, multi-modal interaction information of a user is obtained, and multi-dimensional identification features of the user are generated based on the multi-modal interaction information; based on a pre-trained AI language model and an image generation model, a digital human virtual image is generated through the multi-dimensional identification features; facial acquisition data of the user is obtained in real time, real-time emotion features are extracted according to the facial acquisition data, and a first facial dynamic image of the digital human virtual image is generated according to the real-time emotion features; an initial emotion prediction trajectory is obtained by predicting an emotion trajectory according to the historical emotion features of the user, an emotion influence factor is extracted according to the historical barrage interaction information and the historical emotion features, a barrage distribution time window is collected, and a second facial dynamic image of the digital human virtual image is generated based on the barrage distribution time window, the emotion influence factor, and the initial emotion prediction trajectory; a fusion decision is made based on the first facial dynamic image and the second facial dynamic image to generate a digital human dynamic image for live broadcast push, which can dynamically adjust the facial emotion of the digital human for feedback to the audience, improving the authenticity and interactivity of the digital human expression.
[0129] The above are only the embodiments of the present application. Specific technical solutions or common knowledge such as well-known features are not described in detail herein. It should be noted that for those skilled in the art, without departing from the technical solutions of the present application, several modifications and improvements can be made, which should also be regarded as the protection scope of the present application and will not affect the implementation effect of the present application and the practicality of the patent.
[0130] The protection scope required by the present application should be based on the content of its claims. The specific implementation manners described in the specification can be used to explain the content of the claims. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present invention. Thus, if the modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include the modifications and variations.
Claims
1. A digital life generation method based on AI interaction information, characterized in that, Including: Obtain the multi-modal interaction information of the user, and generate the multi-dimensional identification features of the user based on the multi-modal interaction information; Based on the pre-trained AI language model and image generation model, generate the digital human virtual image through the multi-dimensional identification features; Obtain the facial acquisition data of the user in real time, extract the real-time emotion features according to the facial acquisition data, and generate the first facial dynamic image of the digital human virtual image according to the real-time emotion features; Predict the emotion trajectory according to the historical emotion features of the user to obtain the initial emotion prediction trajectory, extract the emotion influence factors according to the historical barrage interaction information and historical emotion features, collect the barrage distribution time window, and generate the second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotion influence factor and the initial emotion prediction trajectory; Based on the first facial dynamic image and the second facial dynamic image, perform fusion decision-making to generate the digital human dynamic image and use it for live broadcast push.
2. The method according to claim 1, characterized in that, Generating the multi-dimensional identification features of the user based on the multi-modal interaction information specifically includes: For any group of modal interaction information in the multi-modal interaction information, respectively extract the original feature digital representation through the feature extraction network matching the modality; For the original feature digital representation corresponding to any group of modalities, respectively map them to the specified dimension in the feature space through feature transformation to obtain multiple modal features, and align and fuse all modal features in the feature space according to the preset dimension to obtain the multi-dimensional feature vector of the user as the multi-dimensional identification feature.
3. The method according to claim 1, wherein Predicting the emotion trajectory according to the historical emotion features of the user to obtain the initial emotion prediction trajectory specifically includes: Record the facial emotion features of the user during the live broadcast in real time to form a time series sample set of historical emotion features; Construct a moving average autoregressive model of emotion features according to the time series sample set of historical emotion features, and determine the initial emotion prediction trajectory based on the moving average autoregressive model of emotion features.
4. The method according to claim 1, characterized in that Before extracting the emotion influence factor according to the historical barrage interaction information and historical emotion features, it also includes: storing the barrage information received by the user during the live broadcast to obtain the historical barrage interaction information.
5. The method according to claim 1, characterized in that, Extracting the emotion influence factor according to the historical barrage interaction information and historical emotion features specifically includes: Use the natural language processing model to classify the emotions of the historical barrage interaction information, count the distribution of barrage emotion types within the set time window, and generate the barrage emotion distribution sequence; Obtain the time series of historical emotion features, and determine the emotion influence factor based on the sequence correlation between the barrage emotion distribution sequence and the time series of historical emotion features.
6. The method according to claim 1, wherein Generating the second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotion influence factor and the initial emotion prediction trajectory specifically includes: Obtain the initial emotion prediction distribution through the initial emotion prediction trajectory; Extract the barrage emotion features based on the barrage distribution time window, and construct an emotion influence prediction distribution based on the barrage emotion features and the emotion influence factor; Obtain the fusion weights of the initial emotion prediction distribution and the emotion influence prediction distribution, and perform weighted fusion on the initial emotion prediction distribution and the emotion influence prediction distribution to determine the final emotion prediction result; Input the emotion prediction result and the static digital human virtual image into the expression driving model to generate the second facial dynamic image of the digital human virtual image.
7. The method according to claim 1, characterized in that, Based on the first facial dynamic image and the second facial dynamic image for fusion decision-making, generating a digital human dynamic image for live broadcast push specifically includes: Obtain the emotion consistency of the first facial dynamic image and the second facial dynamic image; When the emotion consistency is higher than the preset threshold, respectively extract the motion parameters of the first facial dynamic image and the second facial dynamic image, perform weighted fusion on the motion parameters through preset weights, and generate the final facial motion data; When the emotion consistency is lower than the preset threshold, use the fuzzy decision-making algorithm to select the facial dynamic image, and determine the final facial motion data based on the decision result; Generate a digital human dynamic image based on the final facial motion data and perform live broadcast push.
8. A digital human generation system based on AI interaction information, comprising a digital human image generation unit, wherein the digital human image generation unit is used to execute the method for generating a digital human based on AI interaction information according to any one of claims 1 to 7, characterized in that, The digital human image generation unit includes: A user information extraction module, configured to obtain the multimodal interaction information of the user, and generate the multi-dimensional identification features of the user based on the multimodal interaction information; A virtual image generation module, configured to generate a digital human virtual image based on the pre-trained AI language model and image generation model through the multi-dimensional identification features; A dynamic image generation module, configured to obtain the facial acquisition data of the user in real time, extract the real-time emotion features according to the facial acquisition data, and generate the first facial dynamic image of the digital human virtual image according to the real-time emotion features; The dynamic image generation module is further configured to predict the emotion trajectory according to the historical emotion features of the user to obtain the initial emotion prediction trajectory, extract the emotion influence factors according to the historical barrage interaction information and historical emotion features, collect the barrage distribution time window, and generate the second facial dynamic image of the digital human virtual image based on the barrage distribution time window, the emotion influence factors and the initial emotion prediction trajectory; The dynamic image generation module is further configured to perform fusion decision-making based on the first facial dynamic image and the second facial dynamic image, generate a digital human dynamic image and use it for live broadcast push.
9. A computer terminal device, characterized in that, The computer terminal device includes a memory and a processor, the memory stores code, and the processor is configured to obtain the code and execute a method for generating a digital human based on AI interaction information according to any one of claims 1 to 7.
10. A computer-readable storage medium storing at least one computer program, characterized in that, The computer program is loaded and executed by the processor to implement the operations performed by a method for generating a digital human based on AI interaction information according to any one of claims 1 to 7.
Citation Information
Patent Citations
Live broadcast interaction method and system based on AI digital human
CN119071521A
Method and system for generating digital human video
CN119180895A
AI digital human interaction system based on emotion recognition
CN119473003A
Virtual digital human interaction system based on AI
CN119902625A
Live broadcast system based on AI interaction
CN120091164A