Customized virtual pet interactive simulation method and system
By perceiving user emotions and behavioral characteristics, combining historical data, Markov decision-making process and deep reinforcement learning are used to optimize virtual pet behavior, solving the problem of virtual pet lacking personalized adaptability in long-term interactions, and achieving stable and personalized interactive experience.
Patent Information
- Application Number
- CN202510325870.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing virtual pet system is difficult to continuously optimize according to users' long-term emotional changes and behavioral habits, resulting in a lack of personalized adaptability and long-term stability of the interaction model.
By perceiving the user's emotional state and behavioral characteristics, combining historical interactive data, Markov's decision-making process and deep reinforcement learning are used to optimize the behavioral strategy of virtual pets, multi-modal emotion analysis and long-term memory network predict user emotional trends, and personalized behavioral decisions and feedback adjustments are made.
It has achieved that virtual pets can continuously monitor user emotional changes, predict long-term trends, optimize behavioral strategies, improve interaction coherence and naturalness, and form a personalized experience in line with users' long-term preferences.
Smart Images

Figure CN120258034A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent human-computer interaction, and specifically to a customized virtual pet interaction simulation method and system. Background Art
[0002] In the development of modern human-computer interaction technology, virtual pets, as an emotional companionship system, have gradually been widely applied in multiple fields such as social entertainment, mental health management, and intelligent assistants. Virtual pets can not only interact with users but also gradually establish personalized interaction patterns during long-term use, thereby providing a companionship experience that is more in line with the emotional needs of users. However, since the emotional expression and behavior habits of users change over time, how to enable virtual pets to have long-term adaptability, so that they can continuously optimize their own behaviors and form an interaction pattern that conforms to the long-term preferences of users has become a key technical issue.
[0003] In the prior art, some virtual pet systems already have a certain degree of emotional adaptation ability and can adjust their own interaction methods based on the user's immediate emotional input, such as expressions, voices, texts, etc. Some technical solutions use short-term reinforcement learning to enable virtual pets to quickly adjust their behaviors according to the user's feedback and improve the flexibility of interaction. In addition, there are also technical solutions that use a rule-setting method. Through a pre-set interaction logic, virtual pets can provide corresponding interaction patterns for different types of user needs. These methods have improved the adaptability of virtual pets to user emotions to a certain extent and enhanced the user's interaction experience.
[0004] However, there are still some deficiencies in the prior art. First of all, the prior art mainly focuses on short-term emotional matching and rarely considers the changes in the user's long-term interaction habits, resulting in unstable or rigid behavior patterns in virtual pets during long-term use. Secondly, although the interaction method relying on short-term emotional feedback can provide good adaptability in a single interaction, it is difficult to establish an optimization strategy for the user's long-term needs. Virtual pets will not be able to make behavior choices that truly conform to the user's personalized preferences due to the lack of accumulation of long-term emotional trends. In addition, although some reinforcement learning-based systems can adjust behaviors through immediate rewards, they usually do not fully consider the influence of long-term emotional trends, resulting in virtual pets being easy to learn high-frequency interaction patterns in a short time, but over time, these patterns no longer meet the real needs of users. Finally, although the rule-setting method provides a predefined interaction logic, its inherent static nature makes it difficult to adapt to the gradual changes in user behavior habits, resulting in a lack of flexibility and pertinence in the interaction experience during long-term use. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a customized virtual pet interaction simulation method and system, which solves the problem that virtual pets in the prior art are difficult to be continuously optimized according to the long-term emotional changes and behavior habits of users, resulting in a lack of personalized adaptability and long-term stability in the interaction mode.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A customized virtual pet interaction simulation method, comprising the following steps:
[0007] S1. Sense the user's emotional state, receive the user's voice, facial expression and touch input information through sensors, and obtain the user's real-time emotional state and behavior characteristics;
[0008] S2. Analyze the user's behavior and emotions, based on the sensed information, analyze the user's emotional state through an emotion computing module, and combine historical interaction data to extract the user's long-term behavior pattern;
[0009] S3. Establish a virtual pet behavior model, model the behavior of the virtual pet as a Markov decision process, and generate a behavior decision model for the virtual pet according to the user's current emotional state and historical behavior;
[0010] S4. Optimize the behavior selection, through the linear quadratic regulation algorithm, optimize the immediate behavior of the virtual pet based on the behavior model, and generate an optimal behavior decision;
[0011] S5. Deep reinforcement learning optimization, through proximal policy optimization and hierarchical reinforcement learning algorithms, optimize the long-term behavior strategy of the virtual pet based on the user's long-term interaction data;
[0012] S6. Execute the virtual pet behavior, generate and execute the specific behavior actions of the virtual pet according to the optimal behavior decision, and adjust its voice and expression;
[0013] S7. Feedback and adjust the behavior strategy, receive the user's behavior feedback and emotional reaction through the feedback module, and use the feedback information to adjust the decision parameters in the virtual pet behavior model to achieve long-term optimization.
[0014] Preferably, the sensing of the user's emotional state includes:
[0015] Through speech recognition, extract emotional information from the user's voice through a microphone, identify the emotional type and calculate the emotional intensity;
[0016] Through facial expression analysis, identify the user's facial expression through a camera and convert it into an emotional label;
[0017] Through touch sensing, receive the user's touch input information through a touch sensor, and infer the user's current emotional state and interaction intention;
[0018] Combined with historical interaction records, a long short-term memory network is used to model the user's emotional trend, predict the user's current emotional state based on time series, and optimize the accuracy of multi-modal sentiment analysis.
[0019] Preferably, the analysis of user behavior and emotion includes:
[0020] Perform sentiment classification on user input information, perform sentiment classification on speech, facial expressions, and touch information through a deep learning model, and generate multi-modal sentiment state data;
[0021] Establish a user behavior feature model, extract the long-term rules of user behavior, including emotional fluctuations and interaction preferences, according to the user's interaction history data through a behavior analysis algorithm;
[0022] Dynamically update emotion and behavior data, combine real-time interaction information and historical data, and dynamically adjust the model parameters of user emotion and behavior through weighted average or other time series analysis methods.
[0023] Preferably, the establishment of the virtual pet behavior model includes:
[0024] Use the Markov decision process for modeling, and establish a decision model for virtual pet behavior according to the state space, action space, and reward function of the virtual pet;
[0025] Definition of the state space, define the state space of the virtual pet as a multi-dimensional space including the current emotional state of the pet, the user's emotional state, and historical interaction data;
[0026] Definition of the action space, define the behavior of the virtual pet as multiple discrete actions, including actions such as playing, soothing, or resting;
[0027] Definition of the reward function, define the reward function of the virtual pet behavior according to the user's emotional reaction and behavior preference, and promote the pet's behavior to better meet the user's expectations.
[0028] Preferably, the optimization of behavior selection includes:
[0029] Apply the linear quadratic regulation algorithm, calculate the optimal control law through the Riccati equation, and ensure that the immediate behavior of the virtual pet can maximize the user's emotional satisfaction and minimize the behavior error;
[0030] Optimization of behavior error and emotion matching, adjust the behavior selection of the virtual pet through the linear quadratic regulation algorithm to ensure that the degree of matching between its behavior and the user's current emotional state is maximized.
[0031] Preferably, the deep reinforcement learning optimization includes:
[0032] Use the Proximal Policy Optimization (PPO) algorithm to optimize the long-term behavior policy of the virtual pet, ensuring that its behavior can improve user satisfaction over multiple interaction cycles;
[0033] The hierarchical reinforcement learning method divides the behavior learning of the virtual pet into a high-level policy and a low-level policy. The high-level policy is based on the semi-Markov decision process, and the low-level policy uses Proximal Policy Optimization.
[0034] Preferably, the execution of the virtual pet behavior includes:
[0035] Generate an action sequence according to the optimal behavior decision, and use an action generation neural network to convert the optimized behavior decision into a specific action sequence, including the animation, voice, and expressions of the virtual pet;
[0036] Control the performance of the virtual pet according to the action sequence, and achieve multi-modal behavior performance through the animation engine, speech synthesis system, and expression rendering module of the virtual pet.
[0037] Preferably, the feedback and adjustment of the behavior policy include:
[0038] Receive the emotional feedback data of the user, and obtain the emotional reaction of the user to the virtual pet behavior in real time through the interaction feedback module, including the user's mood changes, behavior preferences, and subjective evaluation of the virtual pet behavior;
[0039] Update the user behavior model, and update the long-term behavior pattern of the user according to the emotional feedback data, including the emotional fluctuation model and the behavior preference model, to ensure that the decision-making layer can make more accurate decisions based on the latest user data;
[0040] Adjust the virtual pet behavior policy, use inverse reinforcement learning to compare the feedback data with the existing decision model, and adjust the parameters of the state transition probability and reward function in the MDP modeling module, so as to optimize the behavior decision of the virtual pet.
[0041] The present invention also provides a customized virtual pet interaction simulation system, and the system includes:
[0042] A perception module for receiving and processing the input information of the user, and analyzing the emotional state and behavior characteristics of the user;
[0043] A decision module for generating the behavior decision of the virtual pet based on the emotional data and behavior pattern of the user, including an MDP modeling module, an LQR optimization module, and a PPO+HRL reinforcement learning module;
[0044] An execution module for generating and executing the actions of the virtual pet according to the behavior decision, and adjusting the decision-making process through the feedback module;
[0045] A feedback module for receiving user feedback information and feeding it back to the decision-making module to optimize the behavior strategy of the virtual pet.
[0046] Preferably, the perception module includes:
[0047] Voice information collection, which obtains the user's voice signal through a microphone and uses a voice feature extraction method to extract parameters such as intonation, pitch, and speech rate to identify the emotional information in the voice;
[0048] Facial expression information collection, which obtains the user's facial image through a camera and uses computer vision methods to extract the key points and expression features of the user's face;
[0049] Touch information collection, which records the user's touch methods, including click, slide, and long-press operations, through a touch sensor or device interaction data, and infers the user's emotional state based on the analysis of the touch pattern;
[0050] Multi-modal emotion analysis, which classifies the user's real-time emotional state based on a multi-modal deep learning model that fuses voice, facial expression, and touch information, and generates the user's emotion label and emotion intensity value.
[0051] The present invention provides a customized virtual pet interaction simulation method and system. It has the following beneficial effects:
[0052] 1. The present invention adopts a long-term emotional state tracking mechanism, and the virtual pet can continuously monitor the emotional changes of the user and make adjustments according to the long-term trend. Compared with the prior art that only relies on short-term emotional feedback, the present invention can predict the long-term emotional state changes of the user, avoid the random fluctuations of the virtual pet's behavior, and improve the coherence and naturalness of the interaction.
[0053] 2. The present invention introduces a personalized behavior pattern recognition method, which extracts stable personalized behavior preferences by analyzing the user's long-term interaction data. Different from the existing behavior adaptation methods based on fixed rules, this method enables the virtual pet to recognize the user's interaction habits and actively optimize its behavior in future interactions, forming a truly personalized experience that meets the user's needs.
[0054] 3. The present invention adopts a long-term reward correction mechanism of reinforcement learning, so that when the virtual pet makes a behavior decision, it not only pays attention to the immediate feedback but also considers the long-term emotional trend of the user. Compared with the prior art that only relies on short-term reinforcement learning methods, the present invention can make the behavior adjustment more targeted by dynamically adjusting the reward weight, avoid the behavior imbalance caused by short-term optimization, and ensure that the learning process of the virtual pet is more stable and intelligent.
[0055] 4. The present invention combines a long-term feedback adjustment mechanism to optimize the behavior selection of virtual pets based on the user's long-term interaction data, enabling them to continuously evolve as the user's needs change. Compared with the traditional fixed behavior logic, the present invention allows virtual pets to continuously learn and adapt during long-term interactions and form an interaction pattern that better conforms to the user's long-term preferences, avoiding the rigidity and ineffectiveness of the interaction pattern. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flowchart of the method of the present invention;
[0057] Figure 2 is a module architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the specification of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] Please refer to the attached Figure 1 , the embodiments of the present invention provide a customized virtual pet interaction simulation method, including the following steps:
[0060] S1. Sense the user's emotional state, receive the user's voice, facial expression, and touch input information through sensors, and obtain the user's real-time emotional state and behavior characteristics;
[0061] The main objective of S1 is to extract emotional information from data such as voice, facial expression, and touch, and combine historical interaction data to form a stable emotional state estimate. Generally, the user's emotions are expressed by multiple signals together, and single-modal emotion recognition is often easily affected by noise or special scenarios, resulting in recognition errors. Therefore, the present invention adopts multi-modal emotion analysis technology, combines methods such as deep learning, statistical modeling, and sequence analysis, and fuses information from different data sources to improve the accuracy and robustness of emotion perception.
[0062] As an option, the output data of the perception layer can be used for the behavior decision-making modeling of virtual pets to ensure the rationality of interaction behaviors. Specifically, the perception layer not only provides the user's immediate emotional state but also can combine historical interaction data to predict the trend of emotional changes. The present invention adopts an emotion time series modeling method, enabling virtual pets to adjust their own interaction strategies based on the user's long-term emotional data, thereby enhancing the intelligent level of the interaction experience.
[0063] In this embodiment, for user emotion perception, multimodal fusion calculation is performed using voice, facial expressions, and touch data.
[0064] In a possible implementation, the voice emotion recognition module uses a microphone array to collect the user's voice signal and preprocesses it, including operations such as band-pass filtering, mean normalization, and short-time Fourier transform (STFT). Generally, the emotional state of speech can be characterized by various audio features, including:
[0065] Speech pitch, which measures the fundamental frequency change of speech and can reflect the ups and downs of emotions. Its calculation method is usually based on the autocorrelation function to estimate the periodicity of the signal:
[0066]
[0067] where: \(R_x(\tau)\) represents the autocorrelation value at delay \(\tau\); \(x(t)\) represents the speech signal at sampling time \(t\); \(\tau\) represents the time delay; \(T\) is the total duration of the speech signal; and \(x(t + \tau)\) is the sampling value of the speech signal after delaying \(\tau\) time steps.
[0068] Short-Term Energy (STE), which is used to evaluate the energy change of the speech signal and can capture the change in emotional intensity:
[0069]
[0070] where: \(E(n)\) represents the short-term energy of the \(n\)th frame; \(x(n + m)\) is the sampling value of the speech signal at time point \(n + m\) within the window range; \(M\) is the window size; and \(n\) is the time point.
[0071] As an option, the feature extraction of the speech signal can be combined with Mel Frequency Cepstral Coefficients (MFCC), and a deep learning architecture with bidirectional LSTM + attention mechanism can be used for emotion classification. Generally, LSTM is used to process the temporal characteristics of the speech signal, and the attention mechanism is used to highlight key emotional segments in long-term dependencies.
[0072] In some other embodiments, the facial expression analysis module uses a camera to obtain the user's facial image and analyzes the user's facial emotional state by combining methods such as key point detection, deep learning feature extraction, and emotion classification. Generally, facial expressions can be analyzed by combining local features and global features. Specifically:
[0073] Local feature analysis:
[0074] Face detection is performed using OpenCV, and 68 key point features are extracted using Dlib, including:
[0075] Eye region (for detecting blink frequency)
[0076] Brow region (for judging frowning degree)
[0077] Mouth region (for analyzing emotions such as smiling and anger)
[0078] Calculate the change in Euclidean distance of facial feature points:
[0079]
[0080] Where: D represents the Euclidean distance between two points; (x1, y1) and (x2, y2) are the coordinates of the feature points.
[0081] Global feature analysis:
[0082] Use ResNet-50 or EfficientNet for deep feature extraction to map the face image to a high-dimensional feature space;
[0083] Identify human face emotions through a Softmax classifier, including categories such as pleasure, sadness, and surprise.
[0084] In a possible implementation, to improve the accuracy of facial expression recognition, a method of fusing multi-scale convolutional features can be adopted, that is, combining low-level features (edges, textures) with high-level features (shapes, structures) to enhance the generalization ability of the model.
[0085] In another possible implementation, the touch behavior perception module is used to analyze the user's touch pattern to assist in emotion recognition. The acquisition methods of touch behavior include:
[0086] Capacitive sensor (suitable for smart bracelets and smartphones)
[0087] Force sensor (for analyzing pressing force)
[0088] Inertial sensor (for detecting sliding direction)
[0089] Generally, touch behavior can be classified as:
[0090] Tap: Short-term contact, indicating that the user is in a relatively relaxed mood;
[0091] LongPress: Continuous contact, indicating that the user is thinking or anxious;
[0092] Swipe: Large-range contact, indicating that the user is in an active state.
[0093] As an option, a Hidden Markov Model (HMM) can be used to model touch behavior to analyze the temporal correlation in the touch sequence. Generally, the HMM infers the user's emotional state by maximizing the probability of the observation sequence:
[0094] P(O|λ)=∑ Q P(O|Q,λ)P(Q|λ);
[0095] Where: P(O|λ) represents the probability of the observation sequence O under the model parameters λ; O is the observation sequence, representing the user's touch behavior sequence; Q represents the hidden state sequence (the user's emotional state); P(O|Q,λ represents the probability of observing the sequence O given the hidden state Q and the model parameters λ; P(Q|λ) represents the state transition probability of the HMM model.
[0096] By adopting the multimodal fusion method, weighted calculations are performed on voice, facial expressions, and touch data. Generally, the attention mechanism can be used to dynamically adjust the weights of each modal data to ensure that the most important information has the greatest impact on the final decision. Finally, the output of the multimodal emotion perception system will be used as input data and passed to the decision-making layer for the intelligent behavior generation of the virtual pet.
[0097] S2. Analyze the user's behavior and emotion. Based on the perception information, analyze the user's emotional state through the emotion calculation module, and combine the historical interaction data to extract the user's long-term behavior pattern;
[0098] S1 has completed the perception of the user's emotional state, including the real-time acquisition of voice, facial expressions, and touch data, and used deep learning and feature extraction techniques to perform emotion classification on these data. However, since the user's emotional state is not a static variable but will evolve over time, relying solely on single-time emotion perception is difficult to accurately reflect the user's long-term emotional trend.
[0099] Generally, the user's emotions are affected by factors such as the environment, interaction habits, and individual characteristics, and have obvious temporal correlation and individual characteristic differences. Therefore, in step S2, the system needs to further analyze the user's historical interaction data, extract the user's long-term behavior pattern, and combine time series analysis methods to predict the user's future emotional state so that the subsequent behavior decision-making layer can make adaptive adjustments based on the long-term trend.
[0100] As an option, the system can adopt various algorithms such as time series modeling, hidden state learning, and emotion change prediction to ensure the accuracy and stability of the modeling. Specifically, the main objectives of this step include:
[0101] Construct the user's emotional time series model and analyze the long-term change trend of the emotion data;
[0102] Identify the interaction patterns of users and summarize the behavioral characteristics of different users under emotional fluctuations;
[0103] Dynamically update personalized emotional parameters to optimize the intelligent adaptability of subsequent behavior decisions.
[0104] In this embodiment, user emotion analysis is based on time series modeling to extract long-term behavioral characteristics and predict future emotional states.
[0105] In a possible implementation, the system first serializes the emotional state of the user and models it based on the Long Short-Term Memory network (LSTM). LSTM can effectively avoid the problem of gradient disappearance when processing time series data, thus more accurately learning the long-term emotional dependence relationship of users.
[0106] LSTM adopts a gating mechanism to control the information flow to determine which historical information needs to be retained or discarded. Its core calculation process is as follows:
[0107]
[0108] where: f t is the forget gate, which controls whether to retain the information of the previous time step; i t is the input gate, which controls whether to update the new information of the current time step; is the candidate memory unit, which stores the candidate value of the new emotional state; C t is the cell state, which stores the long-term emotional state information; o t is the output gate, which controls the output of the emotional state of the current time step; h t is the final output of the emotional state, which is used for subsequent prediction; W f ,W i ,W C ,W o are the weight parameters obtained by model training; b f ,b i ,b C ,b o are the corresponding bias terms; x t is the input emotional feature of the current time step, including voice, facial expression, touch data; σ(·) is the sigmoid activation function, which is used to control the gating signal; tanh(·) is the hyperbolic tangent function, which is used to normalize the data range.
[0109] Generally, LSTM can capture the past emotional change patterns of users and predict future emotional states. This enables the virtual pet to actively adjust the interaction mode when the user is in a low mood, such as providing soothing behaviors, and enhance the interaction experience when the user is in a high mood.
[0110] In some embodiments, to improve the stability of sentiment analysis, weighted moving average (WMA) can be combined to smooth the data.
[0111] The sentiment data of users is often affected by short-term fluctuations. Therefore, before performing long-term trend prediction, data smoothing is usually required to reduce the impact of random fluctuations on sentiment modeling.
[0112]
[0113] Where: S t is the smoothed sentiment state value at time t; X t-i is the original sentiment data at the (t - 1)-th time step; w i is the weighting coefficient; N is the size of the sliding window, which determines the reference range of the system for past sentiment data.
[0114] In one possible implementation, the moving average model can be combined with LSTM. That is, first use weighted moving average to smooth the sentiment data, and then input the smoothed data into LSTM for training and prediction to improve the stability of sentiment trend prediction.
[0115] In another possible implementation, the user's interaction behavior pattern can be modeled by a Hidden Markov Model (HMM) to identify the user's typical interaction patterns.
[0116] HMM can be used to analyze the behavior sequence of users in different sentiment states and predict future possible interaction behaviors. Its core calculation process is as follows:
[0117] P(Q t |Q t-1 ) = A ij P(O t |Q t ) = B jk ;
[0118] Where: P(Q t |Q t-1 ) is the probability that the user enters state Q t at time step t, representing the transition relationship of sentiment states; A ij is an element in the state transition matrix, representing the probability of transitioning from state i to state j; P(O t |Q t ) represents the probability of observing user behavior O t in sentiment state Q t ; B jk is an element in the observation matrix, representing the distribution of user behaviors in different sentiment states.
[0119] In some embodiments, the Viterbi algorithm can be used to optimize the HMM to find the optimal emotional state transition path and predict the long-term behavior trend of the user.
[0120] In another possible implementation, the system can combine personalized emotion modeling to adapt to the long-term interaction patterns of different users.
[0121] There may be significant individual differences in the emotional expression ways of users. Therefore, the system can use a personalized weight adjustment strategy to adjust the model parameters based on the user's historical interaction data. For example:
[0122] For users with more intense emotional expressions, the time window of the LSTM memory unit can be increased to capture emotional changes in a longer time range;
[0123] For users with relatively stable emotional changes, the state transition frequency of the HMM can be reduced to reduce unnecessary prediction fluctuations.
[0124] In some embodiments, the personalized parameter adjustment can be optimized through Bayesian optimization to ensure that the parameter adjustment conforms to the long-term interaction pattern of the user.
[0125] Through a variety of time series modeling methods, a dynamic prediction mechanism for the user's emotional state is established, and combined with behavior pattern analysis, the recognition of long-term interaction patterns is achieved.
[0126] S3. Establish a virtual pet behavior model, model the behavior of the virtual pet as a Markov decision process, and generate a behavior decision model for the virtual pet according to the user's current emotional state and historical behavior;
[0127] In S2, through methods such as time series modeling, hidden Markov model (HMM), and weighted moving average (WMA), a long-term emotional state prediction model of the user is established, and the user's interaction behavior pattern is analyzed. Generally, the user's emotional state will directly affect the behavior decision of the virtual pet. Therefore, in step S3, it is necessary to formulate a behavior strategy for the virtual pet based on the foregoing analysis results, so that it can perform personalized, dynamic, and adaptive interaction adjustments in different emotional states.
[0128] As an option, the behavior decision of the virtual pet can be based on reinforcement learning, combined with the user emotional state prediction model, so that it can continuously optimize its own behavior strategy through long-term interaction. Specifically, it is necessary to establish a state-action-reward decision framework so that the virtual pet can autonomously learn the optimal behavior to adapt to the user's emotional changes. In a possible implementation, the system adopts deep reinforcement learning and combines neural networks for Q-value optimization to ensure that the virtual pet can adjust its behavior response in real time according to the changes in the user's emotional state.
[0129] In this embodiment, the behavior decision-making of the virtual pet is optimized based on reinforcement learning to achieve adaptive behavior generation for the user's emotional state.
[0130] Generally, reinforcement learning can be modeled through a Markov decision process. The MDP consists of the following basic elements:
[0131] The state space S represents the user's emotional state, such as happy, anxious, angry, calm, etc.;
[0132] The action space A represents the behaviors that the virtual pet can execute, such as interaction, comfort, companionship, observation, etc.;
[0133] The state transition function P(s′|s,a) represents the probability of transitioning from the current state s to the new state s′ after executing the action a;
[0134] The reward function R(s,a), which represents the immediate feedback obtained after executing the action a and is used to guide behavior optimization;
[0135] The discount factor γ represents controlling the balance between short-term and long-term rewards.
[0136] In a possible implementation, reinforcement learning is optimized through the Q-learning algorithm to learn the optimal policy of the state-action value function Q(s,a). Q-learning uses the following update formula:
[0137] Q(s,a)←Q(s,a)+α[R(s,a)+γmax a′ Q(s′,a′)-Q(s,a)];
[0138] Where: Q(s,a) represents the long-term reward that can be obtained by executing the action a in the state s; α is the learning rate, which represents controlling the step size of Q-value update; R(s,a) is the immediate reward, which represents the immediate feedback obtained after executing the action a and is used to guide behavior optimization; γ is the discount factor, which is used to measure the influence of future rewards on the current decision; max a′ Q(s′,a′) is the maximum future reward, which represents the best future reward that can be obtained in the new state s′.
[0139] In some embodiments, deep Q-learning (DQN) can be used to further optimize the computational efficiency of Q-learning. DQN approximates the Q-value function through a neural network and uses techniques such as experience replay and target network to improve the stability of training.
[0140] As an option, the system can also adopt the Policy Gradient method to directly optimize the policy function π(a|s) to obtain the optimal behavior:
[0141]
[0142] where: J(θ) is the policy objective function, used to measure the quality of the current policy; θ is the policy parameter, used to control the weight of the policy function; π θ (a|s) is the policy function, representing the probability of selecting action a in state s; Q(s,a) is the expected reward for executing action a in state s; is the gradient used to optimize the policy parameter to enhance the behavior decision-making ability.
[0143] In another possible implementation, the system can combine the Rule-Based System, and through predefined emotion-behavior mapping rules, enable the virtual pet to quickly execute appropriate behaviors in specific emotional states.
[0144] Generally, the following state-behavior mapping rules can be established:
[0145] When the user is in a "happy" state, the virtual pet can initiate interactions actively, such as playing, jumping, etc.;
[0146] When the user is in an "anxious" state, the virtual pet can take soothing behaviors, such as approaching the user and communicating softly;
[0147] When the user is in an "angry" state, the virtual pet can reduce active intervention and observe the user's emotional changes.
[0148] In some embodiments, fuzzy logic can be combined to handle the ambiguity of emotional states. For example, define the membership function of the happy emotional state:
[0149]
[0150] where: μ happiness (x) represents the probability that the user's current state belongs to the "happy" category; x represents the emotional state value predicted by LSTM; k is the steepness of the control function, determining the sensitivity of emotional classification; c is the critical value defining the emotional state classification; e is the base of the natural logarithm.
[0151] In a possible implementation, reinforcement learning and rule methods can be combined:
[0152] Short-term behavior optimization: Use the rule-based method to ensure that the virtual pet can respond quickly;
[0153] Long-term Behavior Optimization: Using reinforcement learning, the virtual pet optimizes its behavior strategy through experience learning.
[0154] In this embodiment, the behavior decision-making of the virtual pet comprehensively adopts reinforcement learning, rule-based methods, and fuzzy logic to ensure the adaptability and personalization of behaviors.
[0155] Generally, reinforcement learning can be used to optimize long-term strategies, rule-based methods can provide immediate behavior feedback, and fuzzy logic can be used to handle the ambiguity of emotional states.
[0156] In some embodiments, adversarial learning can be combined to enhance the virtual pet's adaptability to different users. For example:
[0157] Using user feedback as a reward signal to continuously adjust the behavior of the virtual pet;
[0158] Combining generative adversarial networks (GANs) to improve the robustness of emotion prediction.
[0159] As an option, multi-modal perception data (such as speech, facial expressions, text sentiment analysis) can be further combined to improve the accuracy of behavior decision-making.
[0160] S3 constructs an adaptive behavior decision-making system for the virtual pet through reinforcement learning, rule-based methods, and fuzzy logic. These methods will be used to guide the behavior generation of the virtual pet, enabling it to make intelligent adjustments according to the user's emotional state and improving the naturalness and personalization of long-term interactions.
[0161] S4. Optimal Behavior Selection: Optimize the immediate behavior of the virtual pet based on the behavior model through the linear quadratic regulator algorithm to generate the optimal behavior decision;
[0162] Since the user's emotional state is dynamically changing, relying solely on fixed behavior strategies may lead to instability or reduced adaptability of the interaction. Therefore, in step S4, the system needs to monitor the process of the virtual pet's behavior execution in real time and dynamically adjust the behavior strategy in combination with the user's emotional feedback to ensure the stability and accuracy of long-term interactions.
[0163] As an option, the system can adopt a reinforcement learning feedback mechanism. After the virtual pet executes an action, it can monitor the change in the user's emotional state in real time and optimize the action strategy accordingly. Specifically, the system can use the Adaptive Q-learning Algorithm to adaptively adjust the Q value based on the change value of the user's emotional state, so as to ensure that the behavior of the virtual pet can be continuously optimized during long-term interactions. In a possible implementation, the update formula of the Q-learning algorithm will introduce user emotional feedback to adjust the amplitude and direction of Q value update.
[0164] Generally, the reinforcement learning model needs to dynamically adjust the strategy according to the environmental feedback to improve the adaptability of behavior decision-making. In the interaction system of virtual pets, the change in the user's emotional state can be regarded as a kind of environmental feedback for optimizing the behavior selection of virtual pets. For this reason, the present invention adopts the Adaptive Q-learning algorithm to adjust the Q value according to the actual emotional feedback of the user after the action is executed, so as to optimize the behavior decision-making. Its core update formula is as follows:
[0165] Q(s,a)←Q(s,a)+α[R(s,a)+βΔE+γmax a′ Q(s′,a′)-Q(s,a)];
[0166] Where: Q(s,a) represents the long-term reward that can be obtained by executing action a in state s; α is the learning rate, which represents the step size for controlling the update of the Q value; R(s,a) is the immediate reward, which represents the immediate feedback obtained after executing action a and is used to guide behavior optimization; γ is the discount factor, which is used to measure the impact of future rewards on the current decision; max a′ Q(s′,a′) is the maximum future reward, which represents the best future reward that can be obtained in the new state s′; β is the emotional feedback weight, which is used to adjust the influence degree of the change in the user's emotional state on the Q value update; ΔE is the emotional state change value, which represents the change amount of the user's emotional state before and after the action is executed, and its calculation formula is as follows:
[0167] ΔE = E t+1 -E t ;
[0168] Where: E t is the current emotional state, which represents the emotional state of the user at time step t and is provided by the LSTM emotional prediction model; E t+1 is the subsequent emotional state, which represents the new emotional state of the user after the action is executed and is obtained by the real-time emotional detection module.
[0169] In some embodiments, if the user's emotional state changes significantly, the system will use an emotional deviation correction mechanism to make additional adjustments to the Q-value update to prevent instability in behavior decision-making caused by sudden emotional changes. The correction formula is as follows:
[0170] Q new (s,a) = Q(s,a) - λ1|ΔE|;
[0171] Where: Q new (s,a) is the corrected Q-value, representing the state-action value used for decision-making after adjustment; λ1 is the penalty factor, used to adjust the Q-value penalty when the emotional state changes significantly; |ΔE| is the emotional change amplitude, used to measure the degree of sudden change in the user's emotional state.
[0172] In another possible implementation, the system can combine fuzzy logic control to perform smoother control over the behavior adjustment of the virtual pet, so as to avoid instability in the interaction experience caused by sudden adjustments.
[0173] Generally, the change of the user's emotional state often has non-linear characteristics. Directly using the hard update mechanism of reinforcement learning is likely to cause violent fluctuations in behavior adjustment. Therefore, fuzzy control rules can be adopted to dynamically adjust the amplitude of behavior adjustment according to the change trend of the user's emotional state. Use the "membership function" disclosed in S3.
[0174] In one possible implementation, fuzzy logic can be used to adjust the intensity of behavior execution. For example:
[0175] When the user's emotional change is small (μ happiness (x) is low), the behavior adjustment amplitude is small to avoid unnecessary intervention;
[0176] When the user's emotional change is large (μ happiness (x) is high), the behavior adjustment amplitude is large to meet the user's needs.
[0177] In this embodiment, the behavior execution monitoring of the virtual pet adopts a reinforcement learning feedback mechanism, fuzzy logic control, and emotional deviation correction to ensure that its behavior can be dynamically optimized according to the user's emotional state.
[0178] Generally, reinforcement learning can be used for long-term behavior optimization, fuzzy logic can improve the smoothness of adjustment, and the emotional deviation correction mechanism can prevent unreasonable behavior responses when the emotional state changes suddenly.
[0179] In some embodiments, user behavior analysis can be combined to further optimize the behavior adjustment strategy of the virtual pet by observing the user's long-term interaction habits. For example:
[0180] If a user is more inclined to passive companionship rather than active interaction in an anxious state, the system can make personalized adjustments to the corresponding behavior weights;
[0181] If a user shows a higher willingness to interact in a happy state, the system can increase the initiative of the virtual pet.
[0182] As an option, multi-modal perception data (such as speech, facial expressions, text sentiment analysis) can be further combined to improve the accuracy of behavior decision-making.
[0183] S4 constructs a behavior execution monitoring system for virtual pets through a reinforcement learning feedback mechanism, fuzzy logic control, and emotional deviation correction. These methods are used to guide the behavior adjustment of virtual pets, enabling them to adaptively optimize according to the user's emotional state during long-term interaction and improve the stability and personalization of the interaction experience.
[0184] S5, Deep Reinforcement Learning Optimization, optimizes the long-term behavior strategy of virtual pets based on the user's long-term interaction data through proximal policy optimization and hierarchical reinforcement learning algorithms;
[0185] The main goal of S5 is to enable virtual pets to predict and adapt to the long-term change trend of the user's emotions by further refining and optimizing the behavior strategy, so as to continuously provide more accurate emotional responses during long-term interaction. To this end, step S5 introduces an emotional trend prediction module, which uses historical data based on the change of the user's emotional state and a deep learning model to predict future emotional fluctuations.
[0186] Specifically, the core technical solution of step S5 is to use a Long Short-Term Memory (LSTM) network, which can process time series data and capture the long-term dependence relationship of the user's emotional changes. By learning the historical emotional state of the user, LSTM can predict the change trend of the future emotional state. This prediction result will play a key role in the behavior decision-making of virtual pets. Especially when adjusting the long-term behavior strategy, virtual pets can understand the possible changes in the user's emotional fluctuations in advance and thus make more accurate responses.
[0187] Generally, the output of the LSTM network will be used as the result of emotional prediction to affect the strategy adjustment of virtual pet behavior. Specifically, the output E of LSTM t+1 will be used to correct the state-action value in the Q-learning algorithm, thereby optimizing the choice of virtual pet behavior. The core prediction formula of the LSTM network is as follows:
[0188]
[0189] where: E t+1E is the emotional state at the next moment, which indicates the emotional state of the user at the next moment output based on the current emotional state and the prediction model. This state is a prediction of future emotional changes and is used to guide the behavior adjustment of the virtual pet; t is the current emotional state, which represents the user's emotional state at time step t, provided by the real-time emotion detection module; is the user behavior pattern, which represents the user's behavior pattern at time step t, usually obtained by analyzing the user's historical interaction data and used to assist in emotional state prediction; φ is the LSTM network parameter, which represents the weight and bias terms in the LSTM model, including the weight matrix of the hidden layer and the connection weight from the input layer to the hidden layer, and is responsible for learning and predicting the input data.
[0190] In one possible implementation, the emotional state E predicted by LSTM t+1 It will be used as an emotional feedback correction factor to further affect the update of the Q value. The behavior decision of the virtual pet will combine the user's emotional state prediction and actual feedback to ensure the long-term stability and accuracy of the behavior decision. Specifically, the emotional state prediction results output by LSTM will be combined with the Q value in the Q-learning algorithm to correct the Q value update strategy, using the "Q-learning update formula" disclosed by S2.
[0191] In some embodiments, the system will dynamically adjust the Q value update according to the magnitude of the change in the user's emotional state. For example, when the user's emotions change greatly, the magnitude of the Q value update may increase so that the virtual pet can quickly adapt to the user's emotional needs. The specific dynamic adjustment mechanism is as follows:
[0192] Q adjusted (s,a)=Q(s,a)+λ|E t+1 -E t |;
[0193] Where: Q adjusted (s,a) is the adjusted Q value, which represents the state-action value after the adjustment of the emotional change amplitude, and is used for behavioral decision-making; λ is the adjustment factor, which represents the influence of the control of emotional state change on the Q value adjustment; |E t+1 -E t | is the amplitude of emotion change, which indicates the amplitude of change of the user's emotional state. The larger the value, the more drastic the emotion change and the greater the impact adjustment.
[0194] By optimizing the behavior strategy of the virtual pet based on emotional trend prediction, this embodiment further enhances the adaptability of the virtual pet's behavior, enabling it to make flexible responses and adjustments according to the long-term emotional changes of the user. This behavior optimization method can not only improve the stability of the interaction experience but also enable the virtual pet to continuously enhance the emotional resonance with the user in multiple rounds of long-term interactions.
[0195] S6. Execute the virtual pet behavior. According to the optimal behavior decision, generate and execute the specific behavior actions of the virtual pet, and adjust its voice and expression.
[0196] The core objective of S6 is to construct a real-time behavior optimization model, which combines user instant feedback analysis, reinforcement learning adaptive adjustment, and multimodal emotion fusion methods, enabling the virtual pet to instantaneously correct its behavior during the interaction process, making the behavior decision more accurate and personalized. Generally, the behavior optimization of the virtual pet is based on the cumulative adjustment of long-term emotional trend prediction and reinforcement learning. However, in the case of sudden changes in the user's emotional state, relying solely on long-term trend prediction cannot quickly adapt to the user's needs. Therefore, step S6 enables the virtual pet to adjust its behavior within a short time to match the user's current emotional state by introducing an error compensation mechanism and an instant reward adjustment strategy.
[0197] Specifically, this embodiment adopts a behavior dynamic optimization method based on error compensation in step S6. This method mainly relies on emotional state error calculation, Q-value adjustment, and adaptive behavior correction to optimize the real-time behavior decision of the virtual pet. As an option, this method uses the error between the user's current real-time emotional state and the predicted emotional state to dynamically adjust the reward function in the behavior decision-making process and optimize the learning rate of reinforcement learning according to the error magnitude.
[0198] In a possible implementation manner, the emotional state error calculation formula is as follows:
[0199] δE = E actual - E predicted ;
[0200] Where: δE is the emotional state error, representing the error between the user's actual emotional state E actual and the predicted emotional state E predicted This error is used to measure the accuracy of emotional prediction and guide behavior correction.
[0201] As an option, this error value can be used to dynamically adjust the instant reward function in reinforcement learning to optimize the behavior decision of the virtual pet, enabling it to more accurately match the user's emotional state. The instant reward adjustment formula is as follows:
[0202] R adjusted(s,a) = R(s,a) - η|δE|;
[0203] where: R adjusted (s,a) is the adjusted immediate reward, representing the immediate reward value corrected by combining the emotional state error; R(s,a) is the original immediate reward; η is the error compensation factor, used to control the influence of the error on the immediate reward value; |δE| is the absolute value of the emotional state error, used to measure the deviation between the predicted emotion and the actual emotion.
[0204] In some embodiments, in order to improve the adaptability of the virtual pet to the sudden change of the user's emotion, this embodiment further introduces a reinforcement learning adaptive adjustment mechanism. Specifically, this mechanism dynamically adjusts the learning rate of the reinforcement learning algorithm according to the change trend of the emotional state error, so that the virtual pet can accelerate learning in the case of large errors in order to better adapt to the user's current emotional state. The adjustment formula is as follows:
[0205] α adjusted = α0 + κ|δE|;
[0206] where: α adjusted is the adjusted learning rate, representing the reinforcement learning learning rate adjusted by combining the error; α0 is the base learning rate, representing the default learning rate of reinforcement learning; κ is the adaptive adjustment factor, used to control the influence of the emotional error on the learning rate adjustment; |δE| is the emotional state error, representing the error between the user's current true emotional state and the predicted emotional state, used to dynamically adjust the change range of the learning rate.
[0207] In a possible implementation, if the error |δE| exceeds the set threshold δE threshold , the system will trigger an immediate behavior correction strategy to directly adjust the behavior of the virtual pet to ensure its matching degree with the user's emotional state. For example:
[0208] If |δE| < δE threshold , the virtual pet still optimizes its behavior based on the standard adjustment process of reinforcement learning without performing additional adjustments;
[0209] If |δE| ≥ δE threshold , the virtual pet will perform immediate behavior adjustments, such as changing the intonation, increasing soothing actions or reducing activity, to quickly adapt to the user's current emotional state.
[0210] In some embodiments, this error compensation mechanism can also be combined with a multi-modal emotion fusion method to improve the detection accuracy of the system for the user's emotional state. Specifically, this method calculates the weighted emotional state value by integrating various emotional signals such as the user's voice, facial expression, and text emotion, and adjusts the error compensation strategy based on this value. The calculation formula is as follows:
[0211] E actual = ω1E speech + ω2E facial + ω3E text ;
[0212] Where: E actual is the actual emotional state obtained by weighted calculation, representing the user's emotional state calculated by integrating different modal emotional signals; E speech is the voice emotional state, the emotional state calculated based on the user's voice features; E facial is the facial emotional state, the emotional state obtained by analyzing the user's facial expressions; E text is the text emotional state, the emotional state obtained by analyzing the text input by the user; ω1, ω2, ω3 are modal weighting coefficients, representing the weights of different modalities in the calculation of emotional states.
[0213] By introducing a real-time error compensation mechanism, an immediate reward adjustment strategy, a reinforcement learning adaptive adjustment mechanism, and a multi-modal emotion fusion method, S6 enhances the virtual pet's real-time adaptation ability to the user's emotional state. This optimization scheme ensures that the virtual pet can not only adjust its behavior based on long-term trends but also quickly correct its decisions when the user's emotions change suddenly, so as to improve the naturalness and stability of the interaction.
[0214] S7, Feedback and adjustment of behavior strategies, receive the user's behavior feedback and emotional responses through the feedback module, and use the feedback information to adjust the decision-making parameters in the virtual pet behavior model to achieve long-term optimization;
[0215] The core goal of S7 is to build a long-term interaction adaptation model, which is based on the long-term trend modeling of reinforcement learning, enabling the virtual pet to continuously optimize its own behavior strategy, thereby improving the stability and personalized adaptability of user interaction. Generally, the behavior adjustment of virtual pets mainly relies on short-term reward feedback and emotional state error compensation. However, in long-term interaction scenarios, the user's behavior patterns change gradually, such as the transfer of interest points, the adjustment of emotional expression habits, or the change of interaction frequency. These long-term changes cannot be effectively adapted by the short-term emotional error correction mechanism. Therefore, it is necessary to establish an optimization strategy based on the analysis of historical interaction patterns and the prediction of long-term emotional trends.
[0216] Specifically, in this embodiment, a long-term emotional state tracking mechanism is introduced in step S7. This mechanism dynamically optimizes the behavior pattern of the virtual pet by analyzing the emotional trends of the user within different time periods. As an option, this mechanism can segment the user's interaction data based on time window division, calculate the emotional bias values in different periods, and thereby determine whether there are trend changes in the user's emotional pattern during long-term interaction. For example, if the system detects that the user's interaction frequency has gradually decreased in the recent period, it can be inferred that the current behavior pattern does not fully match the user's long-term needs. At this time, the system will appropriately adjust the behavior of the virtual pet so that it is more inclined to adopt the interaction methods that the user has preferred in future interactions.
[0217] In a possible implementation, this long-term optimization mechanism combines a personalized behavior pattern recognition method. By performing cluster analysis on the user's interaction history, typical behavior patterns are extracted and used as the basis for the virtual pet's behavior decision-making. For example, the user is more inclined to initiate active interactions during certain specific time periods, while being more inclined to passively receive emotional feedback during other time periods. The virtual pet can adjust its own interaction strategy based on these long-term behavior patterns. For example, it can enhance active interactions when the user is active and reduce unnecessary behaviors when the user has less interaction to ensure the continuity and personalization of the interaction experience.
[0218] In addition, to further improve the adaptability of the virtual pet during long-term interaction, this embodiment introduces a long-term feedback adjustment mechanism that long-term corrects the behavior of the virtual pet through the user's cumulative feedback data. Specifically, the system will regularly evaluate the user's response to the virtual pet's behavior.
[0219] In some embodiments, this long-term optimization mechanism can also combine an environmental perception adjustment strategy to optimize the virtual pet's interaction method by monitoring external environmental factors. Therefore, the system can adjust the virtual pet's interaction strategy in different environments based on factors such as the user's device usage data and daily routine habits to make it more in line with the user's actual needs.
[0220] S7 further enhances the adaptability of the virtual pet during long-term interaction by introducing a long-term behavior learning mechanism, personalized behavior pattern recognition, long-term feedback adjustment mechanism, and environmental perception adjustment strategy, enabling it to continuously optimize its own behavior strategy according to the user's long-term emotional trends and behavior patterns.
[0221] A customized virtual pet interaction simulation system described below can be correspondingly referred to the customized virtual pet interaction simulation method described above.
[0222] Please refer to the appendix Figure 2 , the present invention also provides a customized virtual pet interaction simulation system, which includes:
[0223] A perception module, configured to receive and process the input information of the user, and analyze the emotional state and behavioral characteristics of the user;
[0224] A decision-making module, configured to generate behavioral decisions for the virtual pet based on the emotional data and behavioral patterns of the user, including an MDP modeling module, an LQR optimization module, and a PPO+HRL reinforcement learning module;
[0225] An execution module, configured to generate and execute the actions of the virtual pet according to the behavioral decisions, and adjust the decision-making process through a feedback module;
[0226] A feedback module, configured to receive the feedback information of the user and feedback it to the decision-making module to optimize the behavioral strategy of the virtual pet.
[0227] The system of this embodiment can be used to execute the above method embodiment, and its principle and technical effect are similar, which will not be elaborated here.
[0228] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A customized virtual pet interaction simulation method, characterized in that, It includes the following steps: S1. Sense the user's emotional state, receive the user's voice, facial expression and touch input information through sensors, and obtain the user's real-time emotional state and behavioral characteristics; S2. Analyze the user's behavior and emotion. Based on the sensed information, analyze the user's emotional state through an emotion computing module, and combine historical interaction data to extract the user's long-term behavior pattern; S3. Establish a virtual pet behavior model, model the behavior of the virtual pet as a Markov decision process, and generate a behavior decision model for the virtual pet according to the user's current emotional state and historical behavior; S4. Optimize the behavior selection. Through the linear quadratic regulation algorithm, optimize the immediate behavior of the virtual pet based on the behavior model to generate an optimal behavior decision; S5. Deep reinforcement learning optimization. Through proximal policy optimization and hierarchical reinforcement learning algorithms, optimize the long-term behavior strategy of the virtual pet based on the user's long-term interaction data; S6. Execute the virtual pet behavior. According to the optimal behavior decision, generate and execute the specific behavior actions of the virtual pet, and adjust its voice and expression; S7. Feedback and adjust the behavior strategy. Receive the user's behavior feedback and emotional reaction through a feedback module, and use the feedback information to adjust the decision parameters in the virtual pet behavior model to achieve long-term optimization.
2. The customized virtual pet interaction simulation method according to claim 1, characterized in that, The sensing of the user's emotional state includes: Through speech recognition, extract emotional information from the user's voice through a microphone, identify the emotional type and calculate the emotional intensity; Through facial expression analysis, identify the user's facial expression through a camera and convert it into an emotion label; Through touch sensing, receive the user's touch input information through a touch sensor, and infer the user's current emotional state and interaction intention; Combined with historical interaction records, use a long short-term memory network to model the user's emotional trend, predict the user's current emotional state based on time series, and optimize the accuracy of multi-modal emotion analysis.
3. A customized virtual pet interaction simulation method according to claim 1, characterized in that The analysis of the user's behavior and emotion includes: Conduct emotion classification on the user input information. Through a deep learning model, conduct emotion classification on voice, facial expression and touch information, and generate multi-modal emotion state data; Establish a user behavior feature model. According to the user's interaction historical data, extract the long-term rules of the user's behavior through a behavior analysis algorithm, including emotional fluctuations and interaction preferences; Dynamically update the emotion and behavior data. Combine real-time interaction information and historical data, and dynamically adjust the model parameters of the user's emotion and behavior through weighted average or other time series analysis methods.
4. The customized virtual pet interaction simulation method according to claim 1, characterized in that, The establishment of the virtual pet behavior model includes: Use a Markov decision process for modeling. According to the state space, action space and reward function of the virtual pet, establish a decision model for the virtual pet behavior; Definition of the state space. Define the state space of the virtual pet as a multi-dimensional space including the current emotional state of the pet, the user's emotional state and historical interaction data; Definition of the action space. Define the behavior of the virtual pet as multiple discrete actions, including actions such as playing, soothing or resting; Definition of the reward function. According to the user's emotional reaction and behavior preference, define the reward function of the virtual pet behavior to promote the behavior of the pet to better meet the user's expectations.
5. The customized virtual pet interaction simulation method according to claim 1, wherein The optimal behavior selection includes: Applying the linear quadratic regulation algorithm, calculating the optimal control law through the Riccati equation to ensure that the immediate behavior of the virtual pet can maximize the user's emotional satisfaction and minimize the behavior error; Behavior error and emotion matching optimization, adjusting the behavior selection of the virtual pet through the linear quadratic regulation algorithm to ensure that the degree of matching between its behavior and the user's current emotional state is maximized.
6. The customized virtual pet interaction simulation method according to claim 1, characterized in that, The deep reinforcement learning optimization includes: Using the proximal policy optimization algorithm to optimize the long-term behavior policy of the virtual pet to ensure that its behavior can improve user satisfaction over multiple interaction cycles; The hierarchical reinforcement learning method divides the behavior learning of the virtual pet into a high-level policy and a low-level policy, where the high-level policy is based on the semi-Markov decision process and the low-level policy uses proximal policy optimization.
7. A customized virtual pet interaction simulation method according to claim 1, characterized in that, The execution of the virtual pet behavior includes: Generating an action sequence according to the optimal behavior decision, and using an action generation neural network to convert the optimized behavior decision into a specific action sequence, including the animation, voice, and expressions of the virtual pet; Controlling the performance of the virtual pet according to the action sequence, and realizing multi-modal behavior performance through the animation engine, speech synthesis system, and expression rendering module of the virtual pet.
8. A customized virtual pet interaction simulation method according to claim 1, characterized in that The feedback and adjustment of the behavior strategy include: Receiving the emotional feedback data of the user, and obtaining the emotional reaction of the user to the virtual pet's behavior in real time through the interactive feedback module, including the user's mood changes, behavior preferences, and the user's subjective evaluation of the virtual pet's behavior; Updating the user behavior model, updating the long-term behavior pattern of the user according to the emotional feedback data, including the emotional fluctuation model and the behavior preference model, to ensure that the decision-making layer can make more accurate decisions based on the latest user data; Adjusting the virtual pet behavior strategy, using inverse reinforcement learning to compare the feedback data with the existing decision model, and adjusting the parameters of the state transition probability and the reward function in the MDP modeling module, so as to optimize the behavior decision of the virtual pet.
9. A customized virtual pet interaction simulation system, characterized in that, Using a customized virtual pet interaction simulation method according to any one of claims 1-8, the system includes: A sensing module for receiving and processing the user's input information and analyzing the user's emotional state and behavior characteristics; A decision-making module for generating the behavior decision of the virtual pet based on the user's emotional data and behavior pattern, including an MDP modeling module, an LQR optimization module, and a PPO+HRL reinforcement learning module; An execution module for generating and executing the actions of the virtual pet according to the behavior decision and adjusting the decision-making process through the feedback module; A feedback module for receiving the user's feedback information and feeding it back to the decision-making module to optimize the behavior strategy of the virtual pet.
10. The customized virtual pet interaction simulation system according to claim 9, characterized in that, The sensing module includes: Voice information collection, obtaining the user's voice signal through a microphone, and using a voice feature extraction method to extract parameters such as intonation, pitch, and speech rate to identify the emotional information in the voice; Facial expression information collection, obtaining the user's facial image through a camera, and using computer vision methods to extract the key points and expression characteristics of the user's face; Touch information collection, which records the user's touch patterns, including click, swipe, and long-press operations, through touch sensors or device interaction data, and infers the user's emotional state based on the analysis of touch patterns; Multimodal emotion analysis, which classifies the user's real-time emotional state based on a multimodal deep learning model that fuses voice, facial expressions, and touch information, and generates the user's emotion label and emotion intensity value.
Citation Information
Cited By
Method, device and equipment for promoting retention through interactive scene prediction and storage medium
CN120744855A
A method, device, equipment and storage medium for improving retention through interactive scenario prediction
CN120744855B
Pet accompanying robot state control system and method capable of dynamically updating character
CN120921385A
Virtual pet interaction method, device and system, storage medium and program product
CN121257588A
Multi-modal perception and bionic action coordinated pet interaction device
CN121561319A