Digital human dynamic interaction method and system based on deep learning
Through multimodal feature fusion and deep reinforcement learning, the digital human system can accurately understand user emotions and provide personalized, natural and smooth dynamic interactions, solving the problems of stiff interactions and poor adaptability in existing technologies. It is suitable for fields such as intelligent customer service, virtual assistants and online education.
Patent Information
- Application Number
- CN202511100843.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing digital human interaction systems have difficulty accurately capturing complex emotions, their behavioral patterns lack flexibility and adaptability, their real-time responsiveness is poor, and they lack the ability to continuously learn, resulting in stiff and mechanical interactions and the inability to achieve natural and smooth dynamic interactions.
It adopts multimodal feature fusion, deep reinforcement learning strategy and real-time rendering architecture, obtains real-time user interaction data, pre-processes it, and uses a multimodal sentiment analysis model to extract emotional feature vectors. It combines the deep reinforcement learning algorithm to generate a dynamic response strategy for the digital human, and updates the behavioral decision-making model through reinforcement learning.
It achieves accurate emotional insights, intelligent dynamic responses, provides personalized interactive experience, and improves the natural fluency and adaptability of digital human interaction. It is suitable for multiple fields such as intelligent customer service, virtual assistants, and online education, improving user experience and service quality.
Smart Images

Figure CN120688535A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of AI digital humans, and in particular to a method and system for dynamic interaction of digital humans based on deep learning. Background Art
[0002] With the development of artificial intelligence and virtual reality technology, the application of digital human interaction is becoming increasingly widespread, but the existing interaction system has shortcomings.
[0003] Traditional methods mostly rely on simple keyword matching or preset rule libraries to understand user intentions and emotions, which makes it difficult to accurately capture complex emotions. The dimension of emotional understanding is single, and the interaction appears stiff and mechanical. When the user's words imply deep emotions, the digital human often cannot recognize and respond appropriately; when faced with diverse and situational user needs, its behavior patterns lack flexibility and adaptability, and it is impossible to achieve truly natural and smooth dynamic interaction, which limits user experience and application expansion; dynamic response is poor in real time. The decision-making system based on the rule library of traditional methods requires a large number of pre-defined interaction templates. When faced with complex dialogue scenarios, the response delay increases significantly, and the measured average delay exceeds 800ms; behavioral performance is mechanical and stiff: existing action generation schemes mostly use linear interpolation animation, lack modeling of the nonlinear relationship between emotion intensity and action amplitude, resulting in digital human actions that do not conform to human social habits; lack of continuous learning ability: static models find it difficult to optimize interaction strategies based on user feedback, cannot achieve personalized adaptation, and lack online update mechanisms. Summary of the Invention
[0004] To solve the above problems, the present invention provides a digital human dynamic interaction method and system based on deep learning. Through multimodal feature fusion, deep reinforcement learning strategy optimization and real-time rendering architecture, it effectively overcomes the technical defects of traditional methods such as the single dimension of digital human emotional understanding, the stiff and mechanical interaction and the lack of conformity with human social habits.
[0005] To achieve the above object, the technical solution adopted by the present invention is: In a first aspect of the present invention, a method for dynamic interaction of a digital human based on deep learning is provided, comprising the following steps: S1. Acquire and pre-process real-time interaction data input by a user, wherein the interaction data includes at least text information, voice information, and video image information; S2. Analyzing the preprocessed interaction data using a multimodal sentiment analysis model to extract a user emotion feature vector, wherein the emotion feature vector includes an emotion polarity value, an emotion intensity value, and an intent category label; S3. Inputting the emotional feature vector into the digital human behavior decision model, the digital human behavior decision model generates a dynamic response strategy of the three-dimensional digital human through a deep reinforcement learning algorithm. The response strategy includes a set of expression parameters, a sequence of body movements, and voice feedback content; S4. Drive the digital human to interact in real time according to the dynamic response strategy, synchronously monitor user feedback data and establish a reinforcement learning reward function, and dynamically update the network parameters of the behavior decision model.
[0006] Preferably, in S1, the text information includes character-level text data input by the user, and the text data includes but is not limited to chat records, instruction sentences and comment content; the voice information includes the user's voice data, and the voice data includes acoustic physical features, rhythmic emotional features and speech semantic features; the video image information includes the user's video data, and the video data includes facial dynamic features, body behavior features and scene context features.
[0007] Preferably, in S2, the multimodal sentiment analysis model construction includes: The text processing module uses the BERT-TextCNN hybrid architecture to extract semantic features; The image processing module extracts dynamic features of facial micro-expressions through the 3D-CNN network; The speech processing module uses Mel spectrogram features combined with Attention-LSTM to extract acoustic features; The feature fusion layer uses a gated attention mechanism to dynamically weight multimodal features.
[0008] Preferably, in S3, the emotional feature vector is input into the digital human behavior decision model to understand the user's real-time emotions and intentions, and generate corresponding digital human dynamic response behaviors according to preset rules and strategies.
[0009] Preferably, the dynamic response behaviors of the digital human include but are not limited to changes in facial expressions, body movements, voice intonation adjustments, and dialogue content generation. The dynamic response strategy score of the digital human is calculated based on user feedback. The specific calculation formula for the dynamic response strategy score is as follows: , Among them, S represents the dynamic response strategy score, Re w represents the text deviation coefficient of user feedback, e is a constant, α1 represents the text interaction weight, d w represents the ability score of the multimodal sentiment analysis model on the calculated preliminary text data, α2 represents the voice interaction weight, and d s Represents the ability score of the multimodal sentiment analysis model on the calculated preliminary speech data, Re srepresents the voice deviation coefficient of user feedback, α3 represents the video interaction weight, d v Represents the ability score of the multimodal sentiment analysis model on the calculated preliminary video data, v Denotes the video deviation coefficient representing user feedback.
[0010] Preferably, in S4, the user emotion feature vector and dynamic response strategy score are input into a recommendation model based on reinforcement learning, and the optimal response strategy recommendation generated for each user based on the current dynamic response strategy and historical behavior is output. The specific process is as follows: , Among them, r i Indicates the The optimal personalized recommendation for each user, f represents the recommendation strategy generated by the reinforcement learning model, Indicates the User emotion feature vector, Indicates the Dynamic response strategy scoring of digital humans.
[0011] Better yet, the optimal response strategy recommendation is combined with the system execution strategy. During the strategy execution process, each recommended operation will be converted into actual digital human actions or language, and a reinforcement learning reward function will be introduced. When the emotion matching accuracy and user satisfaction are low, the system automatically adjusts the recommendation strategy to select a digital human interaction mode that is more suitable for the current interaction state.
[0012] In a second aspect of the present invention, a digital human dynamic interaction system based on deep learning is provided, comprising: Data acquisition and processing module, used to obtain real-time interactive data input by users and perform preprocessing; Multimodal sentiment analysis module, used to parse pre-processed interaction data based on the multimodal sentiment analysis model and extract user emotion feature vectors; The digital human behavior decision generation module is used to input the emotional feature vector into the digital human behavior decision model. The digital human behavior decision model generates a dynamic response strategy for the three-dimensional digital human through a deep reinforcement learning algorithm, combines the optimal response strategy recommendation with the system execution strategy, and drives the digital human to interact in real time. The feedback update module is used to synchronously monitor user feedback data and establish a reinforcement learning reward function, dynamically updating the network parameters of the behavior decision model.
[0013] The beneficial effects of the present invention are: 1. Accurate emotional insight: Deeply analyzing the emotions used in user words can accurately grasp the user's emotional state. Compared with traditional methods, the recognition rate of complex and implicit emotions is greatly improved, enabling digital humans to better understand user emotions.
[0014] 2. Intelligent Dynamic Response: Understand user emotions and intentions in real time and quickly generate dynamic responses tailored to the situation. This includes intelligent adjustments to expressions, movements, voice, and conversation content, making interactions as natural and coherent as real-life conversations.
[0015] 3. Personalized interactive experience: Customize exclusive interaction modes based on different user emotions and needs, provide personalized services for each user, and enhance user stickiness and satisfaction.
[0016] 4. Wide applicability: It can be applied to multiple fields such as intelligent customer service, virtual assistants, and online education, providing the industry with efficient and intelligent interactive solutions, helping enterprises improve service quality and efficiency and expand business boundaries. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a digital human dynamic interaction method based on deep learning in the present invention.
[0018] Figure 2 This is a block diagram of a digital human dynamic interaction system based on deep learning in the present invention. DETAILED DESCRIPTION
[0019] See also Figure 1 As shown, in a first aspect of the present invention, a digital human dynamic interaction method based on deep learning is provided, comprising the following steps: S1. Acquire and pre-process real-time interaction data input by a user, wherein the interaction data includes at least text information, voice information, and video image information; The core goal of step S1 is to preprocess the real-time interaction data (text, voice, and video images) input by users and convert it into a unified input feature set to facilitate subsequent modeling. Interaction data comes from multiple sources with varying timestamps, data granularity, and formats. Therefore, time alignment, normalization, and missing value interpolation are required to ensure that the data can be fused into a standardized feature vector.
[0020] The text information includes character-level text data input by the user, and the text data includes but is not limited to chat records, command sentences and comment content; the voice information includes the user's voice data, and the voice data includes acoustic physical features, rhythmic emotional features and speech semantic features; the video image information includes the user's video data, and the video data includes facial dynamic features, body behavior features and scene context features.
[0021] Interaction data includes: : Text data, including but not limited to chat logs, command statements, and comments; : Speech data, including acoustic physical features, prosodic emotional features, and speech semantic features; : Video data, including facial dynamic features, body behavior features, scene context features and other data.
[0022] The data preprocessing process is divided into three core steps: time alignment, normalization and missing value interpolation to ensure data quality and consistency.
[0023] Time Alignment: Since interaction data comes from multiple different sources, their sampling frequencies and timestamps may not be completely aligned. To ensure that data from different data sources can be effectively combined during model training, time alignment is crucial. The specific steps are as follows: Video data ( ) records the start and end time of user behavior. Each interaction has a clear time period, which will be used to align with the device state data.
[0024] Text data ( ) and voice data ( ) often have different sampling frequencies. For example, voice data might be sampled every second, while text data might be sampled every 5 minutes. Align all data using timestamps. Specifically, if the sampling interval for device data is 1 second, sample the data to the start timestamp of each user action and then interpolate based on the time window of the action duration. For text data, align to the most recent interaction moment. If missing data exists, use linear interpolation to fill in the missing data.
[0025] Normalization processing: The purpose of normalization is to eliminate the dimensional differences between different data sources and different features so that they can be compared and calculated on the same scale.
[0026] The benefit of normalization is that it ensures that there is no dimensional difference between different features, so that subsequent model training will not be affected by the scale of a feature being too large or too small.
[0027] Missing value imputation: During the actual data collection process, some data may be missing due to equipment failure or external factors. In order to ensure the continuity and stability of subsequent processing, missing value interpolation is an essential step. The specific interpolation method is as follows: The specific method is to fill the missing values by calculating the average of several adjacent data points within the time window. This can effectively avoid deviations caused by sudden changes or erroneous data.
[0028] The output of this step S1 is a structured behavior sample set , obtained by merging after preprocessing, the sample set is a , where N is the number of samples (i.e., the number of car washes), and each row represents the characteristic data of a user interaction behavior. This output will be used as the input for user behavior modeling in subsequent steps. Specifically, Each row in contains three types of information: user text information, voice information, and video image information. This output will be used as input for subsequent steps to model user behavior.
[0029] Through this step S1, the multi-source data from user operations, device status and external environment are effectively aligned, normalized and missing value interpolated to generate a unified format input feature vector .
[0030] During this processing, we ensured data consistency and physical plausibility, avoiding instabilities caused by time misalignment, feature scale differences, or missing data. This provided high-quality, reliable input data for subsequent user behavior modeling, ensuring the stability and operability of the entire system.
[0031] S2. Analyzing the preprocessed interaction data using a multimodal sentiment analysis model to extract a user emotion feature vector, wherein the emotion feature vector includes an emotion polarity value, an emotion intensity value, and an intent category label; The multimodal sentiment analysis model includes a text processing module, an image processing module, and a speech processing module.
[0032] The output dimension of the text processing module is a 256-dimensional semantic vector. The training data uses the CMU-MOSEI dataset, which includes 25,000 annotated samples. The BERT-TextCNN hybrid architecture is used to extract semantic features. The image processing module uses SlowFast 3D-CNN (β=0.125 version) to extract facial action units (AUs). Its key parameters are as follows: Input resolution: 224×224×16 (16 consecutive frames of video clip); Output features: AU activation intensity vector (dimension 52); The speech processing module uses Mel spectrogram features combined with Attention-LSTM to extract acoustic features; The feature fusion layer uses a gated attention mechanism to dynamically weight multimodal features.
[0033] The core objective of step S2 is to extract user emotion feature vectors from the structured user interaction data obtained in step S1 using a multimodal sentiment analysis model. These feature vectors include emotion polarity values, emotion intensity values, and intent category labels, while ensuring privacy protection. These feature vectors help the system understand and predict user behavior, thereby supporting tasks such as personalized interaction recommendations and service optimization for digital humans. To protect user privacy, differential privacy techniques and incremental learning strategies are combined to ensure that the user emotion feature vectors are continuously optimized as data changes, and that user privacy is protected throughout the entire process.
[0034] First, the input data comes from the output of step S1 - the user behavior feature matrix Each line represents an interactive behavior sample, which contains the following three features: text data ( ), voice data( ), and video data ( ). The input matrix The size is , where N is the number of user samples and 3 is the feature dimension of each sample, which is input into the multimodal sentiment analysis model.
[0035] During the extraction of user emotion feature vectors, we perform sentiment classification and annotation based on the collected interactive text data and its corresponding feedback data. Based primarily on user satisfaction ratings, we annotate text samples with positive sentiment (corresponding to "very satisfied" and "satisfied"), negative sentiment (corresponding to "dissatisfied" and "very dissatisfied"), or neutral sentiment (corresponding to "neutral"), forming a training sample dataset with emotion labels.
[0036] Through training, the model can extract deep personalized features from user behavior data and generate personalized emotional feature vectors for each user. , whose dimension is d. This vector will become the input of the subsequent recommendation system.
[0037] To ensure model performance and privacy protection, this step S2 introduces an innovative regularization mechanism. A privacy regularization term is designed to protect user personal information and prevent the leakage of sensitive data during training.
[0038] This regularization term introduces noise into the loss function, making it impossible for an attacker to infer detailed information about a single user from the model update. Specifically, during training, Gaussian noise is injected into each gradient update. The intensity of the noise is controlled by a hyperparameter This ensures a balance between the strength of privacy protection and model fitting accuracy.
[0039] On this basis, the model adopts an incremental learning strategy. This means that when new user behavior data arrives, the model does not retrain from scratch, but instead learns by fine-tuning existing model parameters. This incremental update approach enables the model to promptly adapt to changes in user behavior, improving its efficiency and flexibility in practical applications.
[0040] In terms of privacy protection, this step ensures that each user's behavioral data is not leaked through differential privacy and noise injection mechanisms. In addition, the introduction of privacy regularization terms further enhances the strength of privacy protection and ensures the security of user data.
[0041] Finally, the output of this step S2 is the user's personalized emotion feature vector , whose dimensions are Each row represents a user's personalized features. This user's personalized emotional feature vector will serve as input to the subsequent digital human behavioral decision-making model, helping the system accurately understand the user's needs and preferences and provide personalized decision-making models and recommendations.
[0042] Through this step S2, the entire system can accurately generate the user's personalized emotion feature vector while protecting the user's privacy, and ensure that as new data is added, the model is continuously optimized and adapted to changes in user behavior.
[0043] S3. Inputting the emotional feature vector into the digital human behavior decision model, the digital human behavior decision model generates a dynamic response strategy of the three-dimensional digital human through a deep reinforcement learning algorithm. The response strategy includes a set of expression parameters, a sequence of body movements, and voice feedback content; The core goal of this step S3 is to establish a behavioral decision-making model for the data person and input the emotional feature vector into the behavioral decision-making model. The task of the behavioral decision-making model is to understand the user's real-time emotions and intentions, and generate corresponding dynamic response behaviors of the digital person according to preset rules and strategies.
[0044] An innovative decision model design is used to dynamically evaluate user interaction behaviors to ensure that the system can operate efficiently and accurately.
[0045] The emotional feature vector is input into the digital human behavior decision-making model to understand the user's real-time emotions and intentions, and generate corresponding digital human dynamic response behaviors based on preset rules and strategies.
[0046] First, the emotion feature vector generated in step S2 and the video data in step S1 Merge into a new feature matrix , as the input of the equipment capability prediction model. The dimension of the combined feature matrix is ,in It is the dimension of the video data (for example, facial dynamic features, body behavior features, and scene context features, etc.).
[0047] After the data is prepared, the feature matrix will be input next Input into the behavioral decision-making model for training to generate the corresponding dynamic response behavior of the digital human.
[0048] The dynamic response behaviors of the digital human include but are not limited to changes in facial expressions, body movements, voice intonation adjustments, and dialogue content generation. The dynamic response strategy score of the digital human is calculated based on user feedback. The specific calculation formula for the dynamic response strategy score is as follows: , Among them, S represents the dynamic response strategy score, Re w represents the text deviation coefficient of user feedback, e is a constant, α1 represents the text interaction weight, d w represents the ability score of the multimodal sentiment analysis model on the calculated preliminary text data, α2 represents the voice interaction weight, and d s Represents the ability score of the multimodal sentiment analysis model on the calculated preliminary speech data, Re s represents the voice deviation coefficient of user feedback, α3 represents the video interaction weight, d v Represents the ability score of the multimodal sentiment analysis model on the calculated preliminary video data, v Denotes the video deviation coefficient representing user feedback.
[0049] The dynamic response strategy score is used to score the user's feedback after the digital human interacts with the user. If the user repeats the last interaction action and denies the digital human's interaction action, it means that the user does not approve of the interaction action. When the dynamic response strategy score is less than the preset threshold, the user's intention is re-analyzed through the multimodal sentiment analysis model, and a better dynamic response strategy is regenerated through new intention recognition and required to be executed by the digital human.
[0050] Suppose the digital human performs well in the initial stages of interaction. However, as usage increases, the number of text interactions increases, the number of voice interactions increases, and the excessive video data interaction leads to a decrease in system efficiency. In this case, the computational effort increases over time, dynamically adjusting the dynamic response strategy scoring threshold to ensure that the dynamic response strategy score more accurately reflects the device's actual performance.
[0051] To ensure that the behavior decision model can adapt to changes in interaction status, an incremental learning strategy is introduced. Specifically, when new interaction data arrives, the model is not trained from scratch, but is updated by fine-tuning the existing model. This reduces computing resource consumption while ensuring the real-time performance of the model.
[0052] The key to incremental learning is to fine-tune the model using only new data as it arrives. With each update, the model updates its weights through mini-batch learning, rather than retraining the entire model. This allows the device capability model to quickly adapt to changes in device status, improving prediction accuracy.
[0053] The output of this step is the dynamic response strategy score S, which is a The vector represents the score of each interaction of the digital human. The higher the score, the better the dynamic response strategy score.
[0054] S4. Drive the digital human to interact in real time according to the dynamic response strategy, simultaneously monitor user feedback data and establish a reinforcement learning reward function, and dynamically update the network parameters of the behavior decision model; The user's emotional feature vector and dynamic response strategy score are input into the recommendation model based on reinforcement learning, and the optimal response strategy recommendation generated for each user based on the current dynamic response strategy and historical behavior is output. The specific process is as follows: , Among them, r i represents the optimal personalized recommendation for the i-th user, f represents the recommendation strategy generated by the reinforcement learning model, represents the user's emotion feature vector, Represents the dynamic response strategy score of the digit person.
[0055] After generating a recommendation, the next step is to combine it with the device's execution strategy. The device's execution strategy must not only be based on the recommended action, but also be adjusted in real time based on the digital human's current interaction state.
[0056] To achieve this goal, the policy gradient method from reinforcement learning is applied to the execution strategy. During the policy execution process, each recommended action is translated into an actual digital human interaction. During the execution of the interaction, user input (such as text, voice, and video images) influences the assignment of tasks. The dynamic response strategy score S is used as one of the inputs to the execution strategy, ensuring that the system does not overload when device performance is insufficient.
[0057] The optimal response strategy recommendation is combined with the system execution strategy. During the strategy execution process, each recommended operation will be converted into actual digital human actions or language, and a reinforcement learning reward function will be introduced. When the emotion matching accuracy and user satisfaction are low, the system automatically adjusts the recommendation strategy to select a digital human interaction mode that is more suitable for the current interaction state.
[0058] The core objective of step S4 is to combine the user emotion feature vector generated in step S2 with the response strategy and dynamic response strategy score obtained in step S3, generating personalized recommendations using reinforcement learning (RL) methods and closely integrating them with the strategy execution process. This process aims to dynamically adjust recommendations and execute tasks based on the user's behavioral characteristics and the digital human's current interactive capabilities, thereby maximizing system efficiency, improving user experience, and optimizing system usage. The ultimate goal is to achieve intelligent and personalized services in the self-service digital human system through intelligent decision-making.
[0059] See also Figure 2 As shown, in a second aspect of the present invention, a digital human dynamic interaction system based on deep learning is provided, comprising: Data acquisition and processing module, used to obtain real-time interactive data input by users and perform preprocessing; Multimodal sentiment analysis module, used to parse pre-processed interaction data based on the multimodal sentiment analysis model and extract user emotion feature vectors; The digital human behavior decision generation module is used to input the emotional feature vector into the digital human behavior decision model. The digital human behavior decision model generates a dynamic response strategy for the three-dimensional digital human through a deep reinforcement learning algorithm, combines the optimal response strategy recommendation with the system execution strategy, and drives the digital human to interact in real time. The feedback update module is used to synchronously monitor user feedback data and establish a reinforcement learning reward function, dynamically updating the network parameters of the behavior decision model.
[0060] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering technicians in this field should fall within the scope of protection determined by the claims of the present invention.
Claims
1. A digital human dynamic interaction method based on deep learning, characterized in that: The following steps are involved: S1. Acquire and pre-process real-time interaction data input by a user, wherein the interaction data includes at least text information, voice information, and video image information; S2. Analyzing the preprocessed interaction data using a multimodal sentiment analysis model to extract a user emotion feature vector, wherein the emotion feature vector includes an emotion polarity value, an emotion intensity value, and an intent category label; S3. Inputting the emotional feature vector into the digital human behavior decision model, the digital human behavior decision model generates a dynamic response strategy of the three-dimensional digital human through a deep reinforcement learning algorithm. The response strategy includes a set of expression parameters, a sequence of body movements, and voice feedback content; S4. Drive the digital human to interact in real time according to the dynamic response strategy, synchronously monitor user feedback data and establish a reinforcement learning reward function, and dynamically update the network parameters of the behavior decision model.
2. The method for dynamic interaction of digital humans based on deep learning according to claim 1, characterized in that: In S1, the text information includes character-level text data input by the user, and the text data includes but is not limited to chat records, instruction sentences and comment content; the voice information includes the user's voice data, and the voice data includes acoustic physical features, rhythmic emotional features and speech semantic features; the video image information includes the user's video data, and the video data includes facial dynamic features, body behavior features and scene context features.
3. The method for dynamic interaction of digital humans based on deep learning according to claim 1, characterized in that: In S2, the multimodal sentiment analysis model construction includes: The text processing module uses the BERT-TextCNN hybrid architecture to extract semantic features; The image processing module extracts dynamic features of facial micro-expressions through the 3D-CNN network; The speech processing module uses Mel spectrogram features combined with Attention-LSTM to extract acoustic features; The feature fusion layer uses a gated attention mechanism to dynamically weight multimodal features.
4. The method for dynamic interaction of digital humans based on deep learning according to claim 1, characterized in that: In S3, the emotional feature vector is input into the digital human behavior decision model to understand the user's real-time emotions and intentions, and generate corresponding digital human dynamic response behaviors according to preset rules and strategies.
5. The method for dynamic interaction of digital humans based on deep learning according to claim 4, characterized in that: The dynamic response behaviors of the digital human include but are not limited to changes in facial expressions, body movements, voice intonation adjustments, and dialogue content generation. The dynamic response strategy score of the digital human is calculated based on user feedback. The specific calculation formula of the dynamic response strategy score is as follows: , Among them, S represents the dynamic response strategy score, Re w represents the text deviation coefficient of user feedback, e is a constant, α1 represents the text interaction weight, d w represents the ability score of the multimodal sentiment analysis model on the calculated preliminary text data, α2 represents the voice interaction weight, and d s Represents the ability score of the multimodal sentiment analysis model on the calculated preliminary speech data, Re s represents the voice deviation coefficient of user feedback, α3 represents the video interaction weight, d v Represents the ability score of the multimodal sentiment analysis model on the calculated preliminary video data, v Denotes the video deviation coefficient representing user feedback.
6. The method for dynamic interaction of digital humans based on deep learning according to claim 1, characterized in that: In S4, the user's emotional feature vector and dynamic response strategy score are input into the recommendation model based on reinforcement learning, and the optimal response strategy recommendation generated for each user based on the current dynamic response strategy and historical behavior is output. The specific process is as follows: , Among them, r i represents the optimal personalized recommendation for the i-th user, f represents the recommendation strategy generated by the reinforcement learning model, represents the i-th user emotion feature vector, Represents the dynamic response strategy score of the i-th digital person.
7. The method for dynamic interaction of digital humans based on deep learning according to claim 6, characterized in that: The optimal response strategy recommendation is combined with the system execution strategy. During the strategy execution process, each recommended operation will be converted into actual digital human actions or language, and a reinforcement learning reward function will be introduced. When the emotion matching accuracy and user satisfaction are low, the system automatically adjusts the recommendation strategy to select a digital human interaction mode that is more suitable for the current interaction state.
8. A digital human dynamic interaction system based on deep learning, characterized in that: Includes sequentially connected: Data acquisition and processing module, used to obtain real-time interactive data input by users and perform preprocessing; Multimodal sentiment analysis module, used to parse pre-processed interaction data based on the multimodal sentiment analysis model and extract user emotion feature vectors; The digital human behavior decision generation module is used to input the emotional feature vector into the digital human behavior decision model. The digital human behavior decision model generates a dynamic response strategy for the three-dimensional digital human through a deep reinforcement learning algorithm, combines the optimal response strategy recommendation with the system execution strategy, and drives the digital human to interact in real time. The feedback update module is used to synchronously monitor user feedback data and establish a reinforcement learning reward function, dynamically updating the network parameters of the behavior decision model.
Citation Information
Patent Citations
Intelligent mobile AI digital human interaction method and system based on transparent display device
CN117908683A
Intelligent emotional interaction method for service robot
CN119150099A
Digital human control method and device, equipment and storage medium
CN119884644A
AI digital human interaction system
CN120045069A
Adaptively Modifying Dialog Output by an Artificial Intelligence Engine During a Conversation with a Customer
US20220270594A1
Cited By
Intelligent customer interaction method and system based on cloud computing and Al
CN121258524A
Multi-scene digital human interaction management method based on artificial intelligence
CN121501151A