Action generation method and device based on multi-modal emotion perception, equipment and medium
By integrating multimodal data and adjusting personalized scene features, action parameters are generated, which solves the problem of insufficient emotional perception in robot emotional interaction, and achieves accurate, personalized and real-time emotional response, thereby improving the naturalness and adaptability of human-computer interaction.
Patent Information
- Application Number
- CN202511374154.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing robot emotion interaction technologies suffer from limitations such as a single dimension of emotion perception, lack of dynamic adjustment, and inability to personalize and adapt to specific scenarios. This results in delayed emotional responses and an inability to provide a natural and accurate emotional interaction experience, particularly in the fintech and healthcare sectors where they struggle to meet the demands for agile emotion recognition and personalized interaction.
By acquiring multimodal data (such as facial expressions, voice tone, body movements, and touch perception), weights are assigned to fuse emotional features, emotional states are determined, and action parameters are generated by combining personalized features and scene features. Emotional actions are then executed and the parameter mapping relationship is updated, thereby achieving closed-loop optimization of multimodal emotion perception.
It improves the accuracy, personalization, and real-time nature of emotional responses, enhances the robot's ability to interact naturally with users, and adapts to the differentiated needs of different users and scenarios.
Smart Images

Figure CN120872157A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for generating actions based on multimodal emotion perception. Background Technology
[0002] Existing robot emotional interaction technologies have several limitations, mainly manifested in the single dimension of emotional perception, the lack of dynamic adjustment of action responses to different emotional intensities and types, the lack of personalization and scene adaptation of emotional actions, and the high response delay between emotional perception and action generation, making it difficult to provide a more natural, accurate and user-friendly emotional interaction experience.
[0003] In the fintech sector, existing technologies largely rely on voice or facial expressions to determine user emotions, failing to incorporate crucial details such as body language, micro-expressions, and tactile behavior. This makes it difficult to accurately identify a customer's true emotional state during customer service. Furthermore, current emotional action parameters lack linkage with perceived emotional intensity and type, failing to achieve differentiated and adaptive emotional action adjustments based on individual user characteristics (such as age and personality) and service scenarios (such as branch offices or mobile environments), hindering the improvement of user experience and service friendliness. The response delay in emotional interaction further weakens the real-time interaction effect between the chatbot and the customer, failing to meet the demands for agile emotion recognition and timely reassurance in financial scenarios.
[0004] In the healthcare field, existing technologies for perceiving patients' emotional states also suffer from a narrow scope, particularly in their insufficient utilization of emotional information from facial micro-expressions, body movements, and physical contact, leading to biased judgments of patients' psychological states. Current robots cannot dynamically adjust their emotional responses to patients based on the perceived types and intensities of emotions. They also lack adaptation to individual patient characteristics (such as age and personality) and practical application scenarios (such as wards and outpatient clinics), impacting the personalization and professionalism of doctor-patient emotional interactions. Significant delays in emotion perception and action planning response make it difficult for robots to achieve real-time, natural interaction with patients, affecting patients' trust and comfort in nursing robots.
[0005] In the field of general human-computer interaction, the shortcomings of existing technologies lie in their weak ability to integrate multimodal data, neglecting non-traditional emotional expression elements such as the amplitude of body movements and the intensity of contact, resulting in incomplete recognition of complex user emotions. Robotic emotional actions are fixed, lacking dynamic matching with the user's current emotional state, and also lacking differentiated adaptation across different user groups and scenarios. Low efficiency in data processing and motion planning prevents low-latency responses during emotional interaction, reducing the naturalness and immersion of human-computer emotional interaction. Summary of the Invention
[0006] The main objective of this invention is to provide a method, apparatus, device, and storage medium for action generation based on multimodal emotion perception. This aims to solve the technical problem that existing technologies cannot form a unified closed loop by adapting multimodal emotion perception, personalized features to scene features, dynamically adjusting action parameters, and real-time feedback updates, resulting in a lack of comprehensiveness, adaptability, and real-time performance in robot emotional interaction.
[0007] To achieve the above objectives, the present invention provides an action generation method based on multimodal emotion perception, comprising: The system acquires multimodal data of the target object, assigns corresponding weights to the multimodal data based on the sentiment recognition of each modal data, and fuses the multimodal data based on the corresponding weights to generate fused sentiment features. Based on the fused emotional features, an emotional state including emotional type, emotional intensity value, and emotional intensity change rate is determined; Based on the mapping relationship between the emotional state and the preset action parameters, initial action parameters are generated; The personalized features of the target object and the scene features of the current scene are obtained. A personalized adjustment coefficient is generated based on the personalized features, and a scene adjustment coefficient is generated based on the scene features. The initial action parameters are processed using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters; The system executes an emotional action based on the final action parameters, obtains feedback data from the target object on the emotional action, and updates the action parameter mapping relationship based on the feedback data.
[0008] Furthermore, to achieve the above objectives, the present invention provides an action generation device based on multimodal emotion perception, comprising: The multimodal perception and fusion module is used to acquire multimodal data of the target object, assign corresponding weights to the multimodal data based on the sentiment recognition degree of each modal data, and fuse the multimodal data based on the corresponding weights to generate fused sentiment features; The emotional state analysis module is used to determine the emotional state, including emotional type, emotional intensity value, and emotional intensity change rate, based on the fused emotional features. The action parameter mapping module is used to generate initial action parameters based on the emotional state and the preset action parameter mapping relationship; The personalization and scene adaptation module is used to obtain the personalized features of the target object and the scene features of the current scene, generate a personalization adjustment coefficient based on the personalized features, and generate a scene adjustment coefficient based on the scene features. The motion parameter optimization module is used to process the initial motion parameters using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final motion parameters. The emotional action execution and feedback learning module is used to execute emotional actions according to the final action parameters, obtain feedback data of the target object on the emotional actions, and update the action parameter mapping relationship based on the feedback data.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multimodal emotion perception-based action generation program stored in the memory and executable on the processor, wherein when the multimodal emotion perception-based action generation program is executed by the processor, it implements the steps of the multimodal emotion perception-based action generation method as described above.
[0010] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an action generation program based on multimodal emotion perception, wherein the action generation program based on multimodal emotion perception, when executed by a processor, implements the steps of the action generation method based on multimodal emotion perception as described above.
[0011] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for action generation based on multimodal emotion perception, including: acquiring multiple modal data of a target object; assigning weights based on the emotion recognition degree of each modality and fusing them to generate fused emotion features; determining an emotion state including emotion type, emotion intensity value, and emotion intensity change rate; generating personalized adjustment coefficients and scene adjustment coefficients by combining the target object's personalized features and current scene features; processing initial action parameters to obtain final action parameters; executing the emotional action and acquiring feedback data from the target object; and updating the action parameter mapping relationship based on the feedback data. This invention improves the comprehensiveness of emotion perception by fusing multimodal data, enhances adaptability by dynamically adjusting action parameters by combining personalized features and scene features, and further achieves feedback learning and closed-loop optimization by updating the action parameter mapping relationship based on feedback data, thereby enhancing the accuracy, personalization, and real-time nature of emotional responses. Attached Figure Description
[0012] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an action generation method based on multimodal emotion perception in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the action generation method based on multimodal emotion perception according to the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the action generation device based on multimodal emotion perception of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0013] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0014] The action generation method based on multimodal emotion perception provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain multi-modal data of the target object from the user terminal, assign weights based on the emotional recognition of each modality, and fuse them to generate fused emotional features. It determines the emotional state, including emotional type, emotional intensity value, and emotional intensity change rate. Combining the target object's personalized features and current scene features, it generates personalized adjustment coefficients and scene adjustment coefficients, processes initial action parameters to obtain final action parameters, executes the emotional action, and obtains feedback data from the target object. Based on the feedback data, it updates the action parameter mapping relationship. This invention improves the comprehensiveness of emotional perception by fusing multi-modal data, enhances adaptability by dynamically adjusting action parameters by combining personalized features and scene features, and further achieves feedback learning and closed-loop optimization by updating the action parameter mapping relationship based on feedback data. This enhances the accuracy, personalization, and real-time nature of emotional responses. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0015] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the action generation method based on multimodal emotion perception provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0016] like Figure 2 As shown, the action generation method based on multimodal emotion perception proposed in this invention includes the following steps: S10, acquire multiple modal data of the target object, assign corresponding weights to the multiple modal data based on the sentiment recognition degree of each modal data, and fuse the multiple modal data based on the corresponding weights to generate fused sentiment features; In this embodiment, to accurately process the complex and diverse emotional expressions of users, it is first necessary to collect multiple modal data of the target object. These modal data include, but are not limited to, facial expression data, voice tone data, body movement data, and touch perception data. Facial expression data can be collected using high-definition camera equipment, with particular attention paid to the activity frequency of the eye area and the rate of change of the corners of the mouth. Eye activity frequency represents the number of blinks and gazes per unit time, while the rate of change of the corners of the mouth represents the dynamic amplitude of smiling or frowning movements. Voice tone data is collected using a high-sensitivity microphone, extracting features including the standard deviation of the fundamental frequency of the voice, reflecting the stability and variation of pitch, and the rate of change of speech rate, reflecting fluctuations in speaking rhythm. Body movement data is acquired using a 3D motion capture device, focusing on the amplitude of body swings and the frequency of movements per unit time; this information reflects the motor characteristics in an individual's emotional expression. Touch perception data is acquired using a pressure sensor array or capacitive touch sensor, focusing on contact force and contact duration. Contact force represents the degree of pressure applied by the target object during interaction, and contact duration represents the duration of physical contact.
[0017] After acquiring the aforementioned data, it is necessary to evaluate the sentiment discrimination of each modality. Sentiment discrimination refers to the ability of each modality feature to distinguish the true emotion of the target object in the current context. This can be achieved through a pre-trained confidence classification module. The confidence classification module outputs a confidence score for each modality feature based on historical interaction data and a pre-trained emotion recognition model. Sentiment discrimination is used to guide subsequent weight allocation, assigning greater weight to modalities with higher discrimination to strengthen their contribution to emotion judgment. The weight allocation process can be achieved by adjusting the weights of each modality using a normalization function so that their sum is 1. Based on this allocation result, the multimodal data is weighted and fused according to their corresponding weights. During the fusion calculation, the standardized feature value of each modality is multiplied by its weight and summed to obtain the fused sentiment feature used to describe the current emotional state. The fused sentiment feature retains the advantages of multimodal input and improves the adaptability and robustness to complex emotional expressions under the weighting strategy.
[0018] For facial expression data acquisition, a multi-angle camera array can be used to reduce the impact of lighting changes and occlusion on data accuracy. Eye movement frequency can be identified using image processing algorithms to determine pupil movement trajectories and calculate the number of movements per unit time. Mouth corner change rate can be calculated by tracking the mouth corner position change curve to determine its offset per unit time. Speech intonation data acquisition can suppress background noise by using a directional microphone array. The speech fundamental frequency standard deviation can be obtained by using an adaptive window-length Fourier transform to acquire the dominant frequency distribution of speech frames and calculate the standard deviation. Speech rate change rate can be statistically analyzed by detecting phoneme boundary intervals. Limb movement data acquisition can combine an inertial measurement unit (IMU) with an optical motion capture system to calculate the spatial displacement vector of limb nodes and obtain the swing amplitude and movement frequency. Contact perception data acquisition can accurately detect the pressure value of palm or fingertip contact using a flexible capacitive pressure array and statistically analyze the duration.
[0019] The calculation of sentiment recognition involves inputting collected feature data into a trained classifier model and outputting a confidence score. Different classifier models can include image classifiers based on convolutional neural networks, speech classifiers based on recurrent neural networks, and sequence data classifiers based on long short-term memory networks. The weight allocation corresponding to the sentiment recognition score is implemented through a softmax normalization function to ensure that the contributions of different modalities are dynamically adjusted within the interval. The calculated weights are multiplied one by one with the data from each modality and then summed in a weighted manner to form a fused sentiment feature.
[0020] Example: In the healthcare business field, the application can equip critically ill patients with cameras, microphones, mattress pressure sensors, and handheld pressure sensing devices to collect subtle facial expressions, voice, limb movements, and grip strength changes in real time, and integrate them to generate the patient's current emotional state, providing nursing staff with accurate non-verbal emotional state indicators.
[0021] In the fintech business, applications can deploy front-facing cameras, voice acquisition modules, and desktop pressure sensors on smart customer service terminals to collect data on customers' facial expressions of tension, increased speech speed, and the force and frequency of their fingers tapping the desktop. By fusing these data, customer emotional characteristics can be generated to help smart customer service identify whether customers are anxious or angry and optimize service response strategies.
[0022] This embodiment overcomes the limitations of single modality in emotion perception by acquiring multi-dimensional multimodal data and using emotion recognition-guided weight allocation. This is especially true when body movements are weak, speech is incomplete, or contact data is abnormal; the comprehensive features improve the accuracy of emotion perception in complex environments. By training a confidence module to evaluate the emotion discrimination ability of each modality and dynamically adjusting its weights, the system achieves reasonable fusion of multimodal data, enhancing its overall sensitivity to subtle changes in emotional expression.
[0023] S20, Based on the fused emotional features, determine the emotional state including emotional type, emotional intensity value and emotional intensity change rate; In this embodiment, the first step is to perform probabilistic classification calculations on the fused emotional features. The goal of probabilistic classification is to output the most probable emotional type from the fused feature vector. Emotional types can include categories such as joy, anger, sadness, surprise, and disgust. Specifically, a multi-class classification model, such as a softmax classifier, can be used. The input is the fused emotional feature vector, and the output is the probability distribution of each emotional category. The category with the highest probability is selected as the final emotional type. Probabilistic classification can fully utilize the joint information of multiple modalities in the fused emotional features, improving the accuracy of emotional classification, especially when there are conflicts among the modalities. It balances the contributions of each modality through probabilistic inference after fusion.
[0024] The fused sentiment features are then subjected to a nonlinear transformation to generate sentiment intensity values. Sentiment intensity values represent the strength of emotional expression and are typically mapped to a continuous range of 0 to 1, where 0 indicates almost no emotional expression and 1 indicates extremely strong emotional expression. The nonlinear transformation can employ a sigmoid or tanh function to map the multidimensional fused sentiment features to this range. This step addresses the scaling problem caused by differences in the numerical ranges of different modal features while maintaining the continuous adjustability of the intensity.
[0025] Further, a dynamic analysis strategy is selected based on the emotion type. This strategy calculates the changes in emotion intensity values across different time points. Calculating the rate of change in emotion intensity requires obtaining emotion intensity values at at least two time points, such as the first and second time points. The emotion intensity value at the first time point serves as a baseline, and the value at the second time point is compared to this baseline. Different calculation rules are employed based on the emotion type. For example, a smooth linear difference strategy can be used for sadness, while an exponential change model can be used for anger, reflecting the dynamic characteristics of different emotions over time. The difference in emotion intensity between two time points is calculated using the dynamic analysis strategy, and combined with the time interval between the two points, the rate of change in emotion intensity is generated. This rate of change reflects whether the emotion is strengthening, weakening, or remaining stable, helping to provide a more comprehensive description of the current emotional state.
[0026] Finally, based on the emotion type, emotion intensity value, and emotion intensity change rate, an emotion state is formed. The emotion state is a structured data set that fully describes the target object's emotion category, emotion intensity, and dynamic emotion change characteristics at the current time point, laying the foundation for the generation of subsequent action parameters.
[0027] Emotion type classification can employ a multilayer perceptron neural network as the classifier, fusing emotional features as input and processing them through one or more fully connected hidden layers to output the probability distribution of emotion types. Network training can be based on existing multimodal emotion databases, such as labeled facial expressions, speech, and action data, for supervised learning.
[0028] The generation of emotional intensity values can employ a regularized sigmoid function, projecting the linear combination of emotional features onto a 0-1 interval. The weight parameters of this linear combination can be obtained through training, ensuring that the distribution of intensity values aligns with the subjective perceptions of different user groups. For dynamic analysis strategies, different rate-of-change calculation models can be predefined for different emotional types. For example, a weighted average strategy can be used to calculate smooth changes for joy, while a time-weighted fast decay function model can be used for surprise, ensuring that the rate of change objectively reflects the dynamic characteristics of the specific emotion.
[0029] The structure of an emotional state can be defined as a vector containing three parts: the emotional type encoding, the emotional intensity value, and the emotional intensity change rate. This vector can be stored in an emotion database for subsequent rapid association and retrieval of its mapping relationship with action parameters.
[0030] Example: In the healthcare field, this can be applied to nursing robots. The robot can collect multimodal information about patients through cameras, microphones, and touch sensors, comprehensively analyze whether the patient is currently in a state of anxiety, depression, etc., and determine whether the trend of emotional intensity changes is intensifying, providing timely psychological state assessments for medical staff.
[0031] In the fintech business, it can be used for intelligent financial advisors. By analyzing the multimodal characteristics of customers' interactions with the terminal, it can identify the customer's current level of stress and its changing trends, and assist advisors in adjusting recommendation strategies or terminating transaction suggestions when customers are anxious or emotionally agitated, thereby reducing service risks and improving user experience.
[0032] This embodiment overcomes the shortcomings of traditional systems that only identify emotion categories by simultaneously extracting emotion type, emotion intensity value, and emotion intensity change rate from fused emotion features. This enables emotion recognition to include not only the type of emotion but also the intensity of the emotion and its trend over time, thereby providing rich and accurate emotion input data for downstream action planning and improving the naturalness and adaptability of human-computer interaction.
[0033] S30, Generate initial action parameters based on the mapping relationship between the emotional state and the preset action parameters; In this embodiment, it is necessary to first parse the emotion type and emotion intensity value in the emotional state. The emotion type represents the currently identified main emotion category, such as joy, sadness, anger, surprise, fear, etc., and the emotion intensity value represents the intensity of the emotion, usually a standardized numerical range used to quantify the salience of the emotion. In practical applications, the emotion type can be converted into the corresponding category label through a lookup table or encoding dictionary, and the emotion intensity value can be directly used as a continuous variable in subsequent calculations.
[0034] The action parameter mapping relationship is a pre-defined data structure that includes parameter generation strategies for different emotion types. The design of these strategies can encompass action templates, adjustment rules, and boundary conditions corresponding to different emotion categories, transforming emotional data into parameter values usable for action control. The system uses emotion type as the search key to select the parameter generation strategy matching the current emotion type from the action parameter mapping relationship. The parameter generation strategy can include linear mapping rules, nonlinear function models, and conditional constraints to ensure that the generated action parameters under different emotional states are significantly distinguishable and behaviorally reasonable.
[0035] Based on the emotion intensity value, a numerical transformation strategy is applied using the selected parameters. The transformation results in initial body movement amplitude, initial movement speed, initial voice pitch, and initial speech rate. These parameters correspond to the robot's movement amplitude, movement rhythm, pitch variation range, and speech rate adjustment amount when performing emotional expression, respectively. During the transformation process, each parameter can use an independent mapping formula. For example, for anger, the body movement amplitude may increase linearly with the emotion intensity value, while the voice pitch value uses an exponential enhancement model to ensure the vividness and consistency of the action expression.
[0036] After the initial parameters are generated, boundary constraint processing is required to prevent parameter values from exceeding the control range allowed by the physical device or generating extreme values that are unsuitable for the scenario. Specific boundary constraints can be implemented using maximum and minimum value limit functions. For example, the initial value of the movement speed can be limited to the range of 0.1 m / s to 1.0 m / s to ensure robot movement safety and accurate expression. After boundary processing, initial values for limb movement amplitude, movement speed, voice pitch, and speech rate are generated.
[0037] Finally, the four initial values are combined into a structured set of initial action parameters, which are used in subsequent action control modules or further personalization and scene adjustment processes. This set can be stored or transmitted as a standard data structure, ensuring that the action execution system can directly read and apply it.
[0038] The implementation of parameter generation strategies can include mapping rules designed for different emotions. For example, for joy, a linear positive relationship can be defined between the amplitude of body movements and the intensity of emotion; a proportionally increasing relationship can be defined between the speed of movement and the intensity of emotion; a logarithmic relationship can be defined between the pitch of voice and the intensity of emotion; and a weighted average relationship can be defined between the speed of voice and the intensity of emotion. For sadness, the mapping rules can be defined in reverse, for example, the amplitude of body movements decreases as the intensity value decreases, and the tone of voice tends to be lower. Boundary constraints can be implemented using lookup tables, dynamically adjusting the allowable range to adapt to the needs of different devices and scenarios. For example, the speed limit constraint for medical service robots may be stricter to ensure patient safety.
[0039] Alternatively, the parameter generation strategy can be dynamically adjusted based on machine learning methods. During the training phase, the input emotion type, emotion intensity value, and expert-annotated action parameters are used as training samples. The mapping relationship is learned through a regression model. During the deployment phase, the trained mapping model is called in real time to replace the fixed rules.
[0040] Example Explanation: In the healthcare field, when faced with patients' different emotional states, especially sadness, anxiety, or depression, robots can analyze the patient's current emotional type and intensity. Based on a pre-defined mapping relationship of action parameters, they can dynamically generate initial action parameters that match these emotions. For example, they can moderately reduce the amplitude of limb movements and lower the pitch and speed of speech to make the action output gentler and avoid causing further tension or discomfort to the patient. In implementation, by adjusting the mapping relationship, it is ensured that the robot can generate differentiated initial action parameters for different patient groups under subtle emotional differences, thus better meeting the needs of medical companionship scenarios.
[0041] In the fintech sector, financial service robots, when communicating with users, analyze the user's current emotional type and intensity. For example, if they identify a user's anxiety or doubt, they can select a more soothing parameter generation strategy based on a mapping relationship. This could involve appropriately reducing the amplitude of gestures and lowering the initial speech rate, resulting in a more composed response and increased user trust. In practice, this process ensures that the robot can promptly output optimized initial action parameters based on the current emotional state when facing diverse customer emotions in different financial service scenarios, making the service process smoother and more professional.
[0042] This embodiment associates emotional states with action parameters, enabling different parameter generation logics for different emotion categories. This allows action parameters to be dynamically adjusted according to emotional changes, overcoming the shortcomings of traditional fixed action parameters or those unrelated to emotional intensity. It ensures that action performance is highly consistent with emotional state, improving the naturalness and emotional adaptability of robot emotional expression.
[0043] S40, obtain the personalized features of the target object and the scene features of the current scene, generate personalized adjustment coefficients based on the personalized features, and generate scene adjustment coefficients based on the scene features; In this embodiment, firstly, the personalized characteristics of the target object are acquired, including but not limited to age information and personality traits. Age information can be extracted from user history files or registration data, and personality traits can be assessed based on historical interaction data analysis or pre-survey results. Extraversion, as a typical personality trait, is based on data derived from indicators such as social behavior frequency and historical interaction performance, possessing stability and quantifiable characteristics. The acquisition of personalized characteristics provides the foundation for generating personalized adjustment coefficients, which characterize the proportion of adjustment required for action parameters across different users. For example, older users may need to reduce the intensity of actions or the pitch of their voices; therefore, this coefficient reflects the degree of action adjustment closely related to individual user characteristics. Next, the scene characteristics of the current environment are acquired. Scene type can be identified through environmental sensors and contextual data, such as home scenes, office scenes, and public place scenes. Environmental volume characteristics can be directly collected by environmental noise sensors, representing the background noise level of the current scene. The acquisition of scene characteristics ensures that action parameters are adapted to the environmental conditions. Personalized adjustment coefficients are obtained by inputting personalized features into a preset personalized mapping relationship. This mapping relationship can be trained based on historical data or configured according to rules; for example, users with strong extroversion are mapped to larger gestures, while older users are mapped to a gentler speech rate. Scene adjustment coefficients are obtained by inputting scene features into a scene mapping relationship; for example, quiet environments are mapped to lower voice pitch values and lower speech rates, while noisy environments are mapped to higher voice pitch values and faster speech rates. Throughout the process, the personalized adjustment coefficients and scene adjustment coefficients serve as weighting factors when processing action parameters, supporting the flexibility of action adjustments for different users and in different scenes.
[0044] By binding user identity during the registration process, the system automatically collects age information and accesses historical interaction data between users and the system. Analysis of interaction frequency, proactive dialogue rate, and voice tone characteristics helps determine extroversion levels. The system can also sample the current background volume level in real time using an ambient microphone as an environmental volume feature, and determine the current environment category using contextual location data (such as GPS or network location) or visual analysis as a scene type feature. Personalized mapping relationships can be pre-defined in tabular or parametric function form. For example, linear interpolation can map age groups to action intensity adjustment factors, and personality scores to voice adjustment factors. Scene mapping relationships can be maintained through a rule base; for example, mapping ambient volume above a preset threshold to increased speech rate. Through these methods, action parameters can adapt to specific combinations of user and environment, achieving differentiated adjustments.
[0045] Example: In the healthcare field, nursing robots can automatically reduce the range of motion and tone of voice based on the age characteristics and lower extroversion of elderly patients to reduce patient discomfort and psychological stress. At the same time, based on the relatively quiet environment of the ward, the speaking speed and voice intensity can be further reduced to optimize the human-computer interaction experience.
[0046] In the fintech business, customer service robots can identify young users with high extroversion and adjust their voice to be more enthusiastic. At the same time, they can adapt to higher intonation and faster speech speed in the noisy environment of the business hall, making the service clearer and more efficient in the financial consultation process.
[0047] This embodiment combines the personalized characteristics of the target object with the scene characteristics of the current scene to generate personalized adjustment coefficients and scene adjustment coefficients respectively, providing a fine-grained action parameter adaptation mechanism to ensure that the action output not only conforms to the user's acceptance habits but also adapts to the current environmental state, thereby improving the naturalness, comfort, and situational adaptability of emotional actions.
[0048] S50, the initial action parameters are processed by the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters; In this embodiment, during the adjustment of action parameters, the previously calculated personalized adjustment coefficient and scene adjustment coefficient are first obtained. These two coefficients reflect the adjustment needs of individual user differences and current scene characteristics, respectively. The sources of the personalized adjustment coefficient and scene adjustment coefficient can be traced back to the mapping calculation of user profiles, interaction behavior analysis, and environmental perception data, ensuring that they are consistent with the actual needs of users and scenes. Subsequently, the personalized adjustment coefficient and scene adjustment coefficient are mathematically operated on, for example, by multiplication to form a comprehensive adjustment coefficient. The comprehensive adjustment coefficient is used to uniformly adjust various action parameters, so that the adjustment result takes into account both user characteristics and scene adaptability. After the comprehensive adjustment coefficient is calculated, the initial action parameters are adjusted one by one. Specifically, the initial action parameters include the initial values of limb movement amplitude, movement speed, voice pitch, and voice speed. These are multiplied by the comprehensive adjustment coefficient to obtain the final values of limb movement amplitude, final movement speed, final voice pitch, and final voice speed. In this process, it is ensured that the adjusted result of each parameter reflects both user preferences and adapts to the requirements of the current environment. Finally, the four final action parameters are recombined into a structured parameter set to form the final action parameters, providing directly usable standardized input for subsequent action execution.
[0049] The personalized adjustment coefficient and the scene adjustment coefficient can be calculated using a floating-point multiplication unit to generate a single comprehensive adjustment coefficient. Then, each initial motion parameter is input into a hardware multiplier, and the final motion parameter value is calculated in real time based on the comprehensive adjustment coefficient. A parallel computing unit can also be used to process the four parameters in parallel, improving computational efficiency and reducing overall latency. The range of the comprehensive adjustment coefficient can be dynamically adjusted through parameter configuration, such as defining upper and lower thresholds to prevent abnormally small or large values after adjustment. The adjusted parameter value can be directly output to the motion controller and voice controller to achieve consistent real-time adaptation and adjustment effects.
[0050] Example: In the healthcare business, nursing robots combine the personalized adjustment coefficient (smaller) for elderly patients with the scene adjustment coefficient (smaller) for quiet ward environments to generate a comprehensive adjustment coefficient. Ultimately, this adjusts the range of limb movements to be smaller, the speed of movements to be slower, the tone of voice to be softer, and the speech rate to be slower, making the interaction more soothing and gentle, and meeting the needs of patients.
[0051] In the fintech business, customer service robots identify young, outgoing users in noisy business hall scenarios. The personalized adjustment coefficient (relatively large) is combined with the scenario adjustment coefficient (relatively large) to form a comprehensive adjustment coefficient. The final action parameters after adjustment are characterized by large body movements, faster action speed, higher tone of voice, and faster speech rate, making the service performance more dynamic and clear, and improving the customer experience.
[0052] This embodiment uses personalized adjustment coefficients and scene adjustment coefficients to form a comprehensive adjustment coefficient, and then adjusts and recombines the initial action parameters one by one to ensure that the action output achieves a dynamic balance between user personalized preferences and scene characteristics, thereby improving the adaptability and acceptability of action generation, reducing user discomfort, and improving the accuracy and naturalness of emotional response.
[0053] S60, perform an emotional action according to the final action parameters, obtain feedback data from the target object on the emotional action, and update the action parameter mapping relationship based on the feedback data.
[0054] In this embodiment, when performing emotional actions, the amplitude, speed, pitch, and speed of the limb movements in the final action parameters are first transmitted to the action control module and the speech synthesis module, respectively. The controller drives the robot's joint actuators, facial expression module, and speech synthesizer to perform specific emotional actions based on these parameter values. After execution, a feedback data acquisition process is initiated to assess the effect of the emotional action. The feedback data includes facial expression response signals and skin conductance response signals collected in real time by biosensors. The facial expression response signals reflect the degree of change in the user's expression, and the skin conductance response signals reflect the intensity of the user's physiological response. These data can objectively reflect the user's immediate emotional acceptance. The collected feedback data is converted into a standardized numerical form by the signal processing module and used as a comprehensive feedback index in the calculation of the action parameter mapping relationship update. The operation of updating the action parameter mapping relationship is based on the correlation between the feedback data and the current action parameters. A parameter update algorithm (such as gradient adjustment or incremental update mechanism) is used to appropriately adjust the parameter coefficients of each parameter generation strategy in the mapping relationship, so that the initial action parameters generated under the same emotional state can better meet the user's needs and adapt to the scene conditions.
[0055] The final motion parameters can be directly mapped to the servo control commands for each robot joint and the input parameters for the speech synthesis module via the driver program, achieving synchronized output of motion and speech. Feedback data acquisition utilizes a high frame rate camera and conductivity sensors to monitor facial expression changes and skin conductivity responses in real time, respectively. The embedded signal processing unit then filters, normalizes, and extracts feature values from this data. The module that updates the motion parameter mapping relationship can use a weighted average method or an adaptive adjustment formula to ensure that the updated parameters more accurately reflect the user's immediate emotional state. Furthermore, storing historical feedback records allows for individualized long-term adjustments, further optimizing system adaptability.
[0056] Example: In the healthcare field, after a companion robot performs comforting physical actions and verbal reassurances for a recovering patient, it collects the patient's facial expression changes and skin conductance in real time. If the feedback data shows low emotional receptivity, the system adjusts the amplitude of the physical actions and tone parameters corresponding to sadness in the mapping relationship, making subsequent actions gentler and slower, and more in line with the needs of the recovery scenario.
[0057] In the fintech business, when customer service robots provide consultation services to users in the business hall, they express enthusiasm through actions and voice. If the facial expression response and skin conduction signal collected indicate that the customer's acceptance is low, the system will adjust the speech rate and tone strategy in the parameter mapping relationship associated with happy emotions in a timely manner, so that the language expression in subsequent services is softer and closer to the acceptance preferences of different customer groups.
[0058] This embodiment combines the collection and analysis of action execution and multi-dimensional feedback data to dynamically adjust and update the mapping relationship of action parameters. This enables the system to continuously learn user preferences and reaction characteristics in multiple rounds of interaction, enhancing the personalized adaptability and situational adaptability of action generation, thereby improving the accuracy, naturalness and user satisfaction of emotional responses.
[0059] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a method, apparatus, device, and medium for action generation based on multimodal emotion perception, comprising: acquiring multiple modal data of a target object; assigning weights based on the emotion recognition degree of each modality and fusing them to generate fused emotion features; determining an emotion state including emotion type, emotion intensity value, and emotion intensity change rate; generating personalized adjustment coefficients and scene adjustment coefficients by combining the target object's personalized features and current scene features; processing initial action parameters to obtain final action parameters; executing the emotional action and acquiring feedback data from the target object; and updating the action parameter mapping relationship based on the feedback data. This invention improves the comprehensiveness of emotion perception by fusing multimodal data, enhances adaptability by dynamically adjusting action parameters by combining personalized features and scene features, and further achieves feedback learning and closed-loop optimization by updating the action parameter mapping relationship based on feedback data, thereby enhancing the accuracy, personalization, and real-time nature of emotion responses.
[0060] In one embodiment, step S10 includes: S101, Collect facial expression data of the target object, and extract the frequency of eye movement and the rate of change of the corner of the mouth from the facial expression data; S102, Collect the speech and intonation data of the target object, and extract the speech fundamental frequency standard deviation and speech rate change rate from the speech and intonation data; S103, Collect the limb movement data of the target object, and extract the limb swing amplitude and movement frequency from the limb movement data; S104, Collect contact sensing data of the target object, and extract the contact force and contact duration from the contact sensing data; S105, through the confidence classification module connected to the biosignal sensor, determine the emotional recognition degree of the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, limb swing amplitude, movement frequency, contact force and contact duration; S106, assign corresponding weights to the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, body swing amplitude, movement frequency, contact force and contact duration based on the emotion recognition score; S107, Based on the corresponding weights, the frequency of eye movement, the rate of change of mouth corners, the standard deviation of speech fundamental frequency, the rate of change of speech rate, the amplitude of body swaying, the frequency of movement, the contact force and the contact duration are fused to generate fused emotional features.
[0061] In this embodiment, collecting multiple modal data of the target object requires the synchronous establishment of a data path in a multidimensional information flow. Facial expression data is acquired through a high-resolution image acquisition device. Extracting the eye activity frequency requires using image processing algorithms to detect and perform time-series analysis on the eyelid opening and closing changes in each frame. The eye activity frequency is obtained by counting the number of eyelid opening and closing times per unit time. The extraction of the mouth corner change rate is achieved by marking the areas on both sides of the mouth corner using a key point localization algorithm. Combined with displacement changes and time difference calculations, the average rate of mouth corner movement is obtained. Speech pitch data is acquired through a microphone array. The raw audio signal is input to a spectrum analysis module. When extracting the speech fundamental frequency standard deviation, a short-time Fourier transform is used to calculate the fundamental frequency curve of the audio frame sequence, and the standard deviation of the curve is calculated to measure pitch fluctuation. The speech rate change rate is obtained by using a syllable segmentation recognition algorithm to count the change in the number of syllables per unit time and calculate the time gradient. Limb movement data is acquired using an inertial measurement unit (IMU) or a 3D motion capture system. Limb swing amplitude is obtained by tracking changes in the spatial coordinates of key limb points and calculating the spatial displacement amplitude. Movement frequency is obtained by counting complete swing cycles within a time window. Contact perception data is acquired using a distributed pressure sensor array. Contact force is calculated by weighted averaging of pressure values at each sensing point. Contact duration is calculated by recording the trigger and relaxation times.
[0062] In the confidence classification module, the trained sentiment classification model is introduced into each of the eight feature categories mentioned above. The confidence level of each feature matching the sentiment category is calculated, and this confidence level is the sentiment recognition score. Each sentiment recognition score will serve as the basis for the importance weight allocation of that feature in multimodal fusion. The weight allocation operation is standardized to ensure that the sum of all weights is 1, so as to ensure that the relative contribution of each feature to the final result in the subsequent weighted fusion is consistent with its recognition effect.
[0063] Based on the assigned weights, a weighted fusion operation is performed. During the fusion process, each feature value is multiplied by its corresponding weight, and all weighted results are summed to generate the final fused sentiment feature. The fused sentiment feature is output as a multi-dimensional vector, serving as input for subsequent sentiment state determination. The above processing requires strict synchronization in the time dimension to ensure that the timestamps of the data from each modality are consistent, so that the fusion result reflects the comprehensive sentiment signal of users within the same time period.
[0064] This embodiment achieves the selection of the most identifiable features from multi-source information and their integration with dynamic weights through synchronous acquisition of multi-modal data, precise feature extraction, classification confidence calculation, and weighted fusion. This not only avoids the one-sidedness caused by the limitation of a single modality, but also enhances the ability to distinguish complex emotional states. This enables the system to form a comprehensive representation of the target object's real and multi-dimensional emotional state, thereby providing a high-confidence emotional input basis for subsequent emotional responses.
[0065] In one embodiment, step S20 above includes: S201, Perform probability classification processing on the fused emotional features to determine the emotional type; S202, Perform nonlinear transformation processing on the fused emotional features to generate the emotional intensity value; S203, Select the corresponding dynamic analysis strategy according to the emotion type; S204, obtain the first emotional intensity value at the first moment and the second emotional intensity value at the second moment; S205, Based on the dynamic analysis strategy, determine the difference in emotional intensity between the second emotional intensity value and the first emotional intensity value; S206, Based on the time interval between the second time point and the first time point and the difference in emotional intensity, generate the rate of change of emotional intensity; S207, Based on the emotion type, emotion intensity value and emotion intensity change rate, generate an emotion state.
[0066] In this embodiment, firstly, probabilistic classification processing is performed on the fused emotional features. This requires a pre-trained multi-class emotion recognition model. This model accepts a multi-dimensional vector input of the fused emotional features and outputs a probability distribution of the emotion categories. The classification result uses the emotion label corresponding to the highest probability as the emotion type. In this process, the emotion type is limited to a set of discrete, high-level abstract labels, such as joy, sadness, anger, calmness, anxiety, etc. The classification model can use convolutional neural networks, recurrent neural networks, or multilayer perceptrons enhanced with attention mechanisms. During training, it is based on a cross-modal emotion annotation dataset to fully adapt to the multimodal fused features.
[0067] Next, a nonlinear transformation process is used to generate emotion intensity values. Specifically, the fused emotion features are input into a nonlinear activation function module, such as using a sigmoid or tanh function, to map the original multidimensional continuous values to a predefined range, such as between 0 and 1 or between -1 and 1, to represent the intensity of the current emotion. This mapping ensures that the influence of inconsistencies in different modal scales on the intensity values is normalized and adjusted. Furthermore, the curvature and threshold of the transformation curve can be adjusted according to different emotion types. For example, a steeper sigmoid curve is used for anger, which is easily affected by external interference, to highlight the sensitivity to changes in emotion intensity.
[0068] Based on the identified emotion type, a corresponding dynamic analysis strategy is selected. This strategy is a set of analytical functions associated with different emotion types, used to explain the pattern of emotion intensity changes over time. For example, a linear fitting strategy is used for sadness to capture the slow increasing or decreasing trend, while an exponential fitting strategy is used for anger to reflect its suddenness and rapid dissipation. The dynamic analysis strategies are stored in a configurable strategy library, and the system directly calls them based on the emotion type as the search key.
[0069] When acquiring the sentiment intensity values at the first and second moments, time-series data points are extracted through a historical data buffer to ensure strict timestamp alignment. The first and second sentiment intensity values are read from a sentiment intensity value storage unit that is synchronized with time and input into subsequent processing.
[0070] Based on a dynamic analysis strategy, the difference between the second and first emotion intensity values is calculated to obtain the emotion intensity difference, which reflects the magnitude of change in emotion intensity between the two moments. During the difference calculation, adjustments can be made according to the operational rules defined by different strategies; for example, a smoothing filter can be introduced for high-frequency fluctuating emotions to eliminate abnormal noise.
[0071] The rate of change of emotional intensity is obtained by dividing the difference in emotional intensity by the time interval between two moments. This rate of change is a standard quantitative indicator of the speed of dynamic change in emotion. It has a time-standardized attribute, which ensures comparability across different time scales.
[0072] Finally, by combining the emotion type, emotion intensity value, and emotion intensity change rate, an emotion state is generated. This state is represented by structured data, such as a JSON object or a collection of key-value pairs, which contains the emotion tag, the current intensity value, and the rate of change. This information is used as the parameter input for subsequent emotion actions, ensuring that subsequent modules can perform targeted action planning and output based on this state.
[0073] This embodiment, by fusing probabilistic classification, nonlinear intensity mapping, type-based dynamic analysis, time series difference calculation, and rate of change normalization, can accurately extract the subjective emotion category and intensity level expressed by the target object based on multimodal information, and further capture the trend of emotion change over time. Thus, it achieves a full-dimensional and highly dynamic characterization of the emotional state as a whole, laying the foundation for subsequent adaptive adjustment and personalized response of emotional actions, and solving the problems of traditional single-point emotion recognition being unable to perceive the change of emotion over time and lacking dynamic adaptation.
[0074] In one embodiment, step S30 above includes: S301, Obtain the emotion type and emotion intensity value in the emotional state; S302, Select the corresponding parameter generation strategy from the preset action parameter mapping relationship according to the emotion type; S303, The emotional intensity value is processed by the parameter generation strategy to generate initial body movement amplitude value, initial movement speed value, initial voice pitch value and initial voice speed value; S304, perform boundary constraint processing on the initial limb movement amplitude value, initial movement speed value, initial voice pitch value and initial voice speed value to generate initial values for limb movement amplitude, initial movement speed, initial voice pitch and initial voice speed. S305, combine the initial values of the limb movement amplitude, movement speed, voice pitch, and voice speed to generate initial movement parameters.
[0075] In this embodiment, when obtaining the emotion type and emotion intensity value from the emotional state, it is first necessary to extract fields from the structured data of the emotional state. The emotion type is qualitative descriptive data, such as categories like joy, sadness, and anger. The emotion intensity value is a quantitative value corresponding to the type, usually a real number between 0 and 1, used to describe the intensity of the emotion. The emotion type is encoded in the form of a string or enumeration value, and the emotion intensity value is expressed in floating-point format. Data extraction can be achieved by parsing the structured data format, such as JSON parsing or database query interface.
[0076] When selecting a parameter generation strategy from a preset action parameter mapping relationship, the mapping relationship is stored in the form of a mapping table or a relational database. Emotion type serves as the key field for the query, and each emotion type corresponds to a set of strategy rules. These strategy rules define how to calculate the initial values of different action parameters based on the emotion intensity value. The parameter generation strategies include function models for different emotion types. For example, a linear scaling strategy is used for the joy type, a decreasing adjustment strategy for the sadness type, and an exponential enhancement strategy for the anger type. These function models can be invoked through lookup tables or by calling function interfaces.
[0077] The emotional intensity value is processed through a parameter generation strategy, calculating initial body movement amplitude, initial movement speed, initial voice pitch, and initial speech rate. The calculation formulas for each initial value can be independent. For example, body movement amplitude can be calculated by linearly scaling the intensity value; movement speed can be adjusted non-linearly based on the intensity value; voice pitch can be calculated using a function positively correlated with the intensity value; and speech rate can be determined by weighted calculations of the intensity values. This process requires calling a mathematical function module to calculate each of the four initial values for the emotional intensity value.
[0078] When performing boundary constraint processing, upper and lower limits are applied to each initial value to ensure that the output value is within a safe and acceptable range. For example, the amplitude of body movements is limited to a minimum of 0.1 and a maximum of 1.0, the speed of movements is limited to 0.2 and 2.0, the pitch of speech is limited to a pitch-related range, and the speech rate is limited to a range suitable for human hearing. Boundary constraints are implemented through minimum and maximum value comparison operations, using the mathematical min and max functions to enclose the initial values.
[0079] Finally, the initial values of the limb movement amplitude, movement speed, voice pitch, and speech rate are combined to generate initial action parameters. These initial action parameters are expressed in the form of a data structure, such as an object containing four fields. Each of the four fields records the values of the four action parameters, which serve as input for the next action adjustment and execution module.
[0080] This embodiment combines the mapping relationship between emotional state and action parameters, enabling different emotional types and intensity values to be mapped to the initial values of four key action parameters. By limiting the controllable range of each parameter through individual boundary constraints, it ensures that the initial action parameters have a detailed discriminativeness and adaptive adjustment capability for the input emotion. This solves the problem in existing methods that cannot dynamically adjust the amplitude, speed, tone and rate of speech according to different emotional types and intensities, thereby improving the diversity and personalization of the output emotional actions.
[0081] In one embodiment, step S40 above includes: S401, Extract the age characteristics and personality extroversion characteristics of the target object as personalized characteristics; S402, extract the current scene's type features and ambient volume features as scene features; S403, Based on the age characteristics and personality extroversion characteristics, a personalized adjustment coefficient is generated through a personalized mapping relationship; S404, Based on the type features and ambient volume features, generate scene adjustment coefficients through scene mapping relationships.
[0082] In this embodiment, when extracting the target object's age and extroversion characteristics as personalized features, the age feature is obtained by querying the user's historical records or user-provided personal data, and is usually represented in the form of an integer or timestamp. The extroversion characteristic is calculated through methods such as questionnaire ratings, historical interaction behavior analysis, and social behavior pattern mining, and can be represented by interval values, label values, or percentage ranks. Data extraction requires calling interfaces from structured data storage or user behavior logs, and then performing format conversion to adapt to subsequent calculation modules.
[0083] When extracting the type features and ambient volume features of the current scene as scene features, the scene type features are determined through multi-source information such as context recognition, location service data, and environmental description information. For example, it distinguishes between situations such as home, office, and public places, and is usually coded with classification labels. The ambient volume features are collected in real time by connecting to environmental microphones and sound sensor devices, and are expressed in decibels (dB) or linear units. The collection process combines real-time sampling, noise filtering, and short-time statistical processing techniques to ensure the accuracy and representativeness of the scene volume data.
[0084] Based on age and extroversion characteristics, personalized adjustment coefficients are generated through personalized mapping relationships. These personalized mapping relationships are implemented using a rule engine or function model. For example, a milder adjustment coefficient is applied to younger users, while a more significant adjustment coefficient is applied to users with higher extroversion. The mapping relationships are stored using mathematical expressions or lookup tables. The calculation module applies the corresponding function rules to the input personalized feature values and outputs a personalized adjustment coefficient within a preset range, which serves as an adjustment factor for subsequent parameter corrections.
[0085] Based on scene type features and ambient volume features as scene features, when generating scene adjustment coefficients through scene mapping relationships, the scene mapping relationships rely on standard adaptation rules for different types of scenes. For example, a low-volume scene adjustment coefficient should be adapted for a library scene, and a high-activity scene adjustment coefficient should be adapted for a party scene. The scene type label and ambient volume value are used as inputs, and the scene mapping relationship outputs a scene adjustment coefficient through multi-condition branch logic or weighted calculation method to ensure that the parameter adjustment can be adaptively associated with the current environmental context.
[0086] This embodiment maps the target's age and extroversion level as personalized adjustment coefficients, and combines scene type and ambient volume as scene adjustment coefficients. This allows the action parameters to adapt to both individual user characteristics and current environmental characteristics, solving the problem of lack of user personalization and scene adaptation in existing systems. This significantly improves the acceptability and scene adaptability of emotional actions, especially when facing diverse users and complex scenes, maintaining the flexibility and naturalness of emotional expression.
[0087] In one embodiment, step S50 above includes: S501, Obtain the initial values of limb movement amplitude, movement speed, voice pitch, and voice speed from the initial action parameters; S502, Multiply the personalized adjustment coefficient by the scene adjustment coefficient to generate a comprehensive adjustment coefficient; S503, Multiply the comprehensive adjustment coefficient by the initial value of the limb movement amplitude to generate the final limb movement amplitude value; S504, Multiply the comprehensive adjustment coefficient by the initial value of the motion speed to generate the final motion speed value; S505, Multiply the comprehensive adjustment coefficient by the initial value of the voice pitch to generate the final voice pitch value; S506, Multiply the comprehensive adjustment coefficient by the initial value of the speech rate to generate the final speech rate value; S507, combine the final limb movement amplitude value, final movement speed value, final voice pitch value, and final voice speed value to generate the final movement parameters.
[0088] In this embodiment, the initial action parameters include initial values for limb movement amplitude, movement speed, voice pitch, and speech rate. These values are generated by the preceding module based on the emotional state and represent the basic emotional expression intention. The initial value for limb movement amplitude represents the quantified amplitude of spatial displacement in the limb movement trajectory, usually expressed as a normalized amplitude value; the initial value for movement speed represents the degree of completion of limb movement per unit time, which can be expressed in standard speed units or normalized proportions; the initial value for voice pitch corresponds to the subjective pitch baseline of the audio signal, expressed in Hertz (Hz) or relative proportions; the initial value for speech rate is the rhythmic speed of speech output per unit time, commonly described in words per minute (WPM) or relative units. These initial values are stored in a parameter list format and passed to the subsequent adjustment module as input to the parameter matrix.
[0089] The personalized adjustment coefficient and the scene adjustment coefficient are output as separate scalar parameters for dynamically correcting the initial action parameters. They are calculated using a multiplier to generate a comprehensive adjustment coefficient, which reflects the combined influence weight of individual characteristics and the current scene. The calculation of the comprehensive adjustment coefficient does not involve data order dependencies, can be parallelized, improves overall computational efficiency, and is suitable for low-latency scenarios.
[0090] The multiplication operation between the overall adjustment coefficient and the initial value of the body movement amplitude is a single numerical multiplication used to adjust the body movement amplitude to match the user's preferences and scenario requirements. The same calculation logic applies to the multiplication of the initial values of movement speed, voice pitch, and speech rate with the overall adjustment coefficient. The result of each multiplication is a real-time adjustment value of the target output parameter, ensuring that personalization and scenario adaptation are reflected in the final output.
[0091] After calculation, the final limb movement amplitude value, final movement speed value, final speech pitch value, and final speech rate value are packaged into a final action parameter set by a data encapsulator. The set format is consistent with the subsequent actuator interface standard and supports standard serialization formats such as JSON or binary protocols, which facilitates decoding and calling by the robot motion control system and speech synthesis module.
[0092] This embodiment integrates personalized adjustment coefficients and scene adjustment coefficients into a comprehensive adjustment coefficient, which is then applied to the initial values of body movement amplitude, movement speed, voice pitch, and speech rate, respectively, achieving fine-grained dynamic adjustment of emotional movement parameters. This approach overcomes the problem of fixed, template-based emotional movements in traditional methods, enabling movement parameters to adapt to changes in time, individual differences, and environment. It allows for the generation of more personalized and context-specific movement expressions in real-time interactions, enhancing the robot's adaptability and naturalness in emotional interactions with diverse users and in complex environments.
[0093] In one embodiment, step S60 above includes: S601, control the robot's joint motors and speech synthesizer to perform emotional actions according to the final motion parameters; S602 collects facial expression response signals and skin conductance response signals of the target object through a biosensor; S603, Generate an emotion receptivity score based on the facial expression response signal; S604, Generate an emotional resonance score based on the skin conductance response signal; S605, the emotional acceptance score and the emotional resonance score are weighted and fused to generate a comprehensive feedback score; S606, when the comprehensive feedback score is lower than the preset score threshold, adjust the coefficients in the action parameter mapping relationship.
[0094] In this embodiment, the final motion parameters are the output of the previously adjusted parameter module, comprising four sets of parameters: limb motion amplitude, motion speed, voice pitch, and voice speed. These parameters correspond to the specific execution commands for robot motion execution and voice output. When controlling the robot's joint motors to perform emotional actions, the limb motion amplitude and motion speed values are used to set the target rotation angle and angular velocity commands for the motors, respectively. The range of motion and speed of each joint are precisely controlled by the corresponding fields in the parameters, ensuring that the actions are coherent and conform to the intended emotional expression. The voice pitch and voice speed values are used as inputs and mapped to the baseband control and speech rate modulation modules of the speech synthesizer, respectively, to ensure that the tone and rhythm of the voice output are consistent with the action expression, achieving multimodal synchronous emotional expression.
[0095] While the robot performs emotional actions, the biosensor module is activated. Facial expression response signals are collected by a visual sensor array, including facial muscle movement features and key point motion trajectories. The data format is standardized into a feature vector sequence for subsequent calculations. Skin conductivity response signals are measured by an electrode array attached to the user's skin surface. The signals are recorded as conductivity change curves, reflecting the user's level of emotional arousal.
[0096] The generation of the emotional receptivity score employs a pattern matching algorithm for facial expression response signals. It calculates the similarity between real-time collected facial expression key point motion features and various emotional expression templates in an emotional standard template library. After standardization of the similarity score, the emotional receptivity score is output, quantifying the degree of fit between the current facial expression and the expected emotional action. The generation of the emotional resonance score is based on the dynamic feature extraction of skin conductance response signals. By weighting and combining indicators such as the average rate of ascent and fluctuation amplitude of the signal curve, the correlation between the user's current physiological response and the target emotional arousal state is quantified, and the resulting score is the resonance score.
[0097] The comprehensive feedback score is generated using a weighted fusion algorithm. Emotional receptivity and emotional resonance scores serve as weighted inputs, with weights that can be preset or dynamically adjusted to reflect the current user feedback's emphasis. When the comprehensive feedback score falls below a preset threshold, adaptive parameter adjustment logic is executed. This includes adjusting the coefficient entries corresponding to various emotional types in the action parameter mapping relationship. The adjustment process uses the deviation between the feedback score and the threshold as the weighting basis for the adjustment magnitude. Coefficient updates employ a recursive update formula, ensuring that the parameter output for future similar emotional states in the mapping relationship more closely matches user preferences.
[0098] The entire process operates on a real-time closed-loop mechanism. Each round of action execution, feedback collection, score calculation, and coefficient adjustment is completed on a millisecond timescale, ensuring that the robot dynamically corrects the deviation between the action expression and user feedback during continuous interaction. This allows the action parameter mapping relationship to gradually adapt to the user's personality and the scene environment. The comprehensive design ensures that the final emotional actions not only accurately respond to the user's emotional state but also have the ability to gradually optimize based on historical interaction experience, enhancing the naturalness and personalization of the human-computer interaction process.
[0099] Example Explanation: In the healthcare field, such as in the application of emotion-assisting robots for long-term care or rehabilitation patients, the robot first collects multiple modal data from the patient by integrating visual cameras, voice sensors, tactile sensors, and motion sensors. Specifically, this includes facial expression data, voice tone data, body movement data, and contact perception data during patient interaction with the device. From the collected facial expression data, eye movement frequency and the rate of change of the corners of the mouth are extracted as important facial features. From the voice tone data, the standard deviation of the fundamental frequency and the rate of change of speech rate are extracted to reflect the emotional state of the voice. From the body movement data, the amplitude and frequency of body movements are extracted to determine the patient's physical activity level. From the contact perception data, the contact force and duration are extracted to supplement the patient's intention to actively interact. Through a confidence classification module connected to a biosignal sensor, the contribution of these features to emotion recognition is calculated. Based on the emotion recognition score, corresponding weights are assigned, and the weighted fusion of each modal feature generates a fused emotion feature.
[0100] Based on the generated fused emotional features, the robotic system performs probabilistic classification using a built-in emotional classification model to determine the patient's current emotional type, such as anxiety, depression, or pleasure. Furthermore, the system quantifies the emotional intensity value from the fused emotional features using a nonlinear mapping algorithm. Combining this with the patient's intensity changes over time, and selecting a matching dynamic analysis strategy based on different emotional types, the system calculates the current emotional intensity change rate, ultimately determining the emotional state, including emotional type, emotional intensity value, and emotional intensity change rate.
[0101] For different emotional states, the robot selects the appropriate parameter generation strategy based on a preset mapping relationship of motion parameters. For example, anxiety corresponds to slow and steady body movements and a soft tone of voice, while joy corresponds to dynamic movements and a high-pitched tone of voice. The system uses the emotional intensity value as input to calculate the initial body movement amplitude value, initial movement speed value, initial tone of voice pitch value, and initial speech rate value, and applies boundary constraints to each value, such as ensuring that the movement speed is not lower than a certain minimum safety threshold and the speech rate is not faster than the upper limit of comprehensibility. After boundary constraints, the above parameters are combined to form the initial motion parameters used to drive the robot's movements and speech output.
[0102] Simultaneously, the system extracts the patient's personalized characteristics, such as age and extroversion level, and combines these with the type of the current medical setting (e.g., ward or rehabilitation training room) and ambient volume characteristics. It then generates personalized adjustment coefficients and scene adjustment coefficients through personalized and scene mapping relationships, respectively. For example, the system automatically reduces the amplitude of movements and the volume of speech for elderly, introverted patients in a nighttime ward environment. The personalized adjustment coefficient is multiplied by the scene adjustment coefficient to form a comprehensive adjustment coefficient. This comprehensive adjustment coefficient is used to adjust the four values of the initial action parameters, yielding the final limb movement amplitude, final movement speed, final speech pitch, and final speech rate. These four values are then combined to generate the final action parameters.
[0103] Based on the final action parameters, the robot controls joint motors and a speech synthesizer to output corresponding emotional actions and speech expressions to the patient, such as a gentle tone and soothing gestures. Subsequently, the robot uses biosensors to monitor the patient's actual feedback to these emotional actions, collecting facial expression response signals (such as raised eyebrows or drooping eyes) and skin conductance response signals (reflecting the level of autonomic nervous activity). Algorithms calculate emotional receptivity scores and emotional resonance scores, which are then weighted and fused to obtain a comprehensive feedback score. If the comprehensive feedback score is lower than a preset threshold, the system adjusts the parameter coefficients in the action parameter mapping relationship to make the next round of action parameters more suitable for the patient's personality and current emotional response characteristics.
[0104] Through this closed-loop process, the emotion-assisting robot can not only perceive the patient's emotional state in real time using multimodal methods, but also dynamically adjust its output of emotional actions and speech expressions according to individual differences and scenario requirements, collecting and analyzing patient feedback for continuous optimization. This adaptive adjustment capability enhances the emotional naturalness, acceptability, and comfort of the robot's interaction with patients in the healthcare environment, providing a precise and intelligent solution for patient rehabilitation support and emotional care.
[0105] In the fintech business, such as scenarios involving intelligent assistants at financial service counters or online intelligent financial advisors, the system first uses multimodal acquisition devices to perceive customers in real time, collecting various modal data, including facial expression data, voice tone data, body movement data, and contact perception data during customer interaction with the terminal device. From facial expression data, it extracts eye movement frequency and mouth corner change rate; from voice tone data, it extracts the fundamental frequency standard deviation and speech rate change rate; from body movement data, it extracts body swing amplitude and movement frequency; and from contact perception data, it extracts contact force and contact duration. All data is input into a confidence classification module connected to a biosignal sensor, calculating the contribution of these features to emotion judgment as the emotion recognition score. Based on the emotion recognition score, a weight is assigned to each feature, and then the modalities are weighted and fused to form a fused emotion feature, thus comprehensively reflecting the customer's current emotional state.
[0106] Building upon this foundation, the financial intelligent assistant uses a probabilistic classification algorithm to determine the customer's emotional type, such as tension, anxiety, or confidence. A nonlinear transformation method calculates the current emotional intensity value, and a dynamic analysis strategy combines the intensity changes over continuous time points to calculate the rate of change in emotional intensity, comprehensively determining the emotional state to reflect the customer's emotional dynamics when faced with financial product recommendations or risk warnings.
[0107] Based on a pre-defined mapping relationship for action parameters, the system selects different parameter generation strategies according to the determined emotion type and intensity value in the emotional state. For example, when a customer exhibits high anxiety, the system chooses gentler gestures and lowers the tone and speed of speech. Using the emotion intensity value as input, the system calculates initial body movement amplitude, initial movement speed, initial tone of voice, and initial speech speed. Boundary constraints are applied to each initial value, such as ensuring that the speech speed does not fall below a comprehensible threshold or that the gestures do not appear excessively slow. The four constrained parameters are combined to form the initial action parameters for subsequent interactive output.
[0108] Simultaneously, the system collects personalized customer characteristics, such as age and extroversion level, to address the varying acceptance of communication methods among customers of different age groups. It also collects the type characteristics of the current scenario (e.g., high-end financial hall vs. regular branch) and ambient volume characteristics. Personalized adjustment coefficients and scenario adjustment coefficients are generated through personalized and scenario mapping relationships, respectively. For example, for older, introverted customers in quiet environments, the amplitude of movements and voice volume are adjusted to a gentler level. The personalized adjustment coefficient and scenario adjustment coefficient are multiplied to form a comprehensive adjustment coefficient, used to adjust the initial action parameters, generating final body movement amplitude values, final action speed values, final voice pitch values, and final voice speed values, which are then combined to form the final action parameters.
[0109] Ultimately, the action parameters drive the physical behavior and voice output of the financial intelligent assistant, such as adjusting tone of voice, smile amplitude, and gesture frequency to make the service behavior more in line with the customer's current emotional state. Subsequently, the system continuously collects the customer's facial expression response signals (such as smiling or frowning) and skin conductance response signals (reflecting tension levels) through biosensors, calculates emotional receptiveness scores and emotional resonance scores, and weights and fuses them into a comprehensive feedback score. If the comprehensive feedback score is lower than a preset threshold, the system adjusts the coefficients of the action parameter mapping relationship to make the actions and tone of voice in subsequent services more consistent with the current customer's personality and emotional characteristics.
[0110] In fintech business scenarios, this intelligent service assistant can achieve personalized, emotion-sensitive financial advice interactions through real-time multimodal perception and adaptive adjustment. It can reduce customer anxiety, enhance customer comfort and trust in financial service processes, and improve the professionalism and humanization of financial institutions' intelligent services, especially demonstrating higher service precision when dealing with high-net-worth clients or risk-sensitive clients.
[0111] This embodiment directly links the final motion parameters to the execution of the robot's joint motors and speech synthesizer, and combines the synchronous acquisition and processing of facial expression response signals and skin conductance response signals to achieve refined quantification of multimodal user feedback. This accurately characterizes the actual emotional acceptance and physiological resonance level of the target user. Based on this, a weighted fusion of emotional acceptance and emotional resonance scores is used to dynamically calculate the comprehensive feedback score. Whether the comprehensive feedback score falls below a preset threshold is used as a condition to adjust the coefficients in the motion parameter mapping relationship in a timely manner, making subsequent motion parameter generation more closely aligned with the user's individual needs and contextual requirements. This process is real-time, and the closed-loop adaptive update ensures that emotional and motor expressions continuously align with the user's actual reactions. Through this approach, the personalization, adaptability, and accuracy of the robot's emotional expression are effectively improved, reducing the deviation between the robot's motor and speech outputs and the user's emotional state, and enhancing the comfort and naturalness of the user's interactive experience.
[0112] In one embodiment, a multimodal emotion-based action generation device is provided, which corresponds one-to-one with the multimodal emotion-based action generation method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the action generation device based on multimodal emotion perception of the present invention. The modules include: multimodal perception and fusion module 10, emotion state analysis module 20, action parameter mapping module 30, personalization and scene adaptation module 40, action parameter optimization module 50, and emotion action execution and feedback learning module 60. Detailed descriptions of each functional module are as follows: The multimodal perception and fusion module 10 is used to acquire multiple modal data of the target object, assign corresponding weights to the multiple modal data based on the sentiment recognition degree of each modal data in the multiple modal data, and fuse the multiple modal data based on the corresponding weights to generate fused sentiment features; The emotional state analysis module 20 is used to determine the emotional state, including emotional type, emotional intensity value and emotional intensity change rate, based on the fused emotional features. The action parameter mapping module 30 is used to generate initial action parameters based on the emotional state and the preset action parameter mapping relationship; The personalization and scene adaptation module 40 is used to obtain the personalized features of the target object and the scene features of the current scene, generate a personalized adjustment coefficient based on the personalized features, and generate a scene adjustment coefficient based on the scene features. The motion parameter optimization module 50 is used to process the initial motion parameters through the personalized adjustment coefficient and the scene adjustment coefficient to generate the final motion parameters. The emotional action execution and feedback learning module 60 is used to execute emotional actions according to the final action parameters, obtain feedback data of the target object on the emotional actions, and update the action parameter mapping relationship based on the feedback data.
[0113] In one embodiment, the multimodal sensing and fusion module 10 is specifically used for: Collect facial expression data of the target object, and extract the frequency of eye movement and the rate of change of the corner of the mouth from the facial expression data; Collect speech and intonation data of the target object, and extract the speech fundamental frequency standard deviation and speech rate change rate from the speech and intonation data; Collect limb movement data of the target object, and extract the limb swing amplitude and movement frequency from the limb movement data; Collect contact sensing data of the target object, and extract the contact force and contact duration from the contact sensing data; The confidence classification module connected to the biosignal sensor determines the emotional recognition of the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, limb swing amplitude, movement frequency, contact force and contact duration. Based on the emotional recognition score, assign corresponding weights to the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, body sway amplitude, movement frequency, contact force, and contact duration. Based on the corresponding weights, the frequency of eye movements, the rate of change of the corner of the mouth, the standard deviation of the fundamental frequency of speech, the rate of change of speech rate, the amplitude of body swaying, the frequency of movement, the contact strength and the contact duration are fused to generate fused emotional features.
[0114] In one embodiment, the emotion state analysis module 20 is specifically used for: The fused emotional features are subjected to probability classification to determine the emotional type; The fused emotional features are subjected to a nonlinear transformation to generate the emotional intensity value. Select the corresponding dynamic analysis strategy based on the emotion type; Obtain the first emotional intensity value at the first moment and the second emotional intensity value at the second moment; The difference in emotional intensity between the second emotional intensity value and the first emotional intensity value is determined based on the dynamic analysis strategy. Based on the time interval between the second moment and the first moment and the difference in emotional intensity, an emotional intensity change rate is generated; An emotional state is generated based on the emotional type, emotional intensity value, and emotional intensity change rate.
[0115] In one embodiment, the action parameter mapping module 30 is specifically used for: Obtain the emotion type and emotion intensity value in the emotional state; From the preset action parameter mapping relationship, select the corresponding parameter generation strategy according to the emotion type; The emotional intensity value is processed by the parameter generation strategy to generate initial body movement amplitude value, initial movement speed value, initial voice pitch value and initial voice speed value. Boundary constraint processing is performed on the initial limb movement amplitude value, initial movement speed value, initial voice pitch value, and initial voice speed value to generate initial values for limb movement amplitude, initial movement speed, initial voice pitch, and initial voice speed. The initial motion parameters are generated by combining the initial values of the limb movement amplitude, the initial value of the movement speed, the initial value of the voice pitch and the initial value of the voice speed.
[0116] In one embodiment, the personalization and scene adaptation module 40 is specifically used for: The age and extroversion characteristics of the target object are extracted as personalized features; Extract the current scene's type features and ambient volume features as scene features; Based on the age characteristics and personality extroversion characteristics, a personalized adjustment coefficient is generated through a personalized mapping relationship; Based on the aforementioned type characteristics and environmental volume characteristics, scene adjustment coefficients are generated through scene mapping relationships.
[0117] In one embodiment, the motion parameter optimization module 50 is specifically used for: Obtain the initial values of limb movement amplitude, movement speed, voice pitch, and voice speed from the initial action parameters; Multiply the personalized adjustment coefficient by the scene adjustment coefficient to generate a comprehensive adjustment coefficient; Multiply the comprehensive adjustment coefficient by the initial value of the limb movement amplitude to generate the final limb movement amplitude value; Multiply the overall adjustment coefficient by the initial value of the motion speed to generate the final motion speed value; Multiply the comprehensive adjustment coefficient by the initial value of the voice pitch to generate the final voice pitch value; The final speech rate value is generated by multiplying the comprehensive adjustment coefficient by the initial speech rate value. The final limb movement amplitude value, final movement speed value, final voice pitch value, and final voice speed value are combined to generate the final movement parameters.
[0118] In one embodiment, the emotional action execution and feedback learning module 60 is specifically used for: The robot's joint motors and speech synthesizer are controlled to perform emotional actions based on the final motion parameters. Facial expression response signals and skin conductivity response signals of the target object are collected using biosensors; An emotion receptivity score is generated based on the facial expression response signal; An emotional resonance score is generated based on the skin conductance response signal. The emotional receptivity score and emotional resonance score are weighted and fused together to generate a comprehensive feedback score; When the overall feedback score is lower than a preset score threshold, the coefficients in the action parameter mapping relationship are adjusted.
[0119] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side action generation method based on multimodal emotion perception.
[0120] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements user-side functions or steps of a multimodal emotion-based action generation method.
[0121] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: The system acquires multimodal data of the target object, assigns corresponding weights to the multimodal data based on the sentiment recognition of each modal data, and fuses the multimodal data based on the corresponding weights to generate fused sentiment features. Based on the fused emotional features, an emotional state including emotional type, emotional intensity value, and emotional intensity change rate is determined; Based on the mapping relationship between the emotional state and the preset action parameters, initial action parameters are generated; The personalized features of the target object and the scene features of the current scene are obtained. A personalized adjustment coefficient is generated based on the personalized features, and a scene adjustment coefficient is generated based on the scene features. The initial action parameters are processed using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters; The system executes an emotional action based on the final action parameters, obtains feedback data from the target object on the emotional action, and updates the action parameter mapping relationship based on the feedback data.
[0122] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: The system acquires multimodal data of the target object, assigns corresponding weights to the multimodal data based on the sentiment recognition of each modal data, and fuses the multimodal data based on the corresponding weights to generate fused sentiment features. Based on the fused emotional features, an emotional state including emotional type, emotional intensity value, and emotional intensity change rate is determined; Based on the mapping relationship between the emotional state and the preset action parameters, initial action parameters are generated; The personalized features of the target object and the scene features of the current scene are obtained. A personalized adjustment coefficient is generated based on the personalized features, and a scene adjustment coefficient is generated based on the scene features. The initial action parameters are processed using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters; The system executes an emotional action based on the final action parameters, obtains feedback data from the target object on the emotional action, and updates the action parameter mapping relationship based on the feedback data.
[0123] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0124] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0125] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0126] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for action generation based on multimodal emotion perception, characterized in that, Includes the following steps: The system acquires multimodal data of the target object, assigns corresponding weights to the multimodal data based on the sentiment recognition of each modal data, and fuses the multimodal data based on the corresponding weights to generate fused sentiment features. Based on the fused emotional features, an emotional state including emotional type, emotional intensity value, and emotional intensity change rate is determined; Based on the mapping relationship between the emotional state and the preset action parameters, initial action parameters are generated; The personalized features of the target object and the scene features of the current scene are obtained. A personalized adjustment coefficient is generated based on the personalized features, and a scene adjustment coefficient is generated based on the scene features. The initial action parameters are processed using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters; The system executes an emotional action based on the final action parameters, obtains feedback data from the target object on the emotional action, and updates the action parameter mapping relationship based on the feedback data.
2. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, Acquire multi-modal data of the target object, assign corresponding weights to the multi-modal data based on the sentiment recognition degree of each modality, and fuse the multi-modal data based on the corresponding weights to generate fused sentiment features, including: Collect facial expression data of the target object, and extract the frequency of eye movement and the rate of change of the corner of the mouth from the facial expression data; Collect speech and intonation data of the target object, and extract the speech fundamental frequency standard deviation and speech rate change rate from the speech and intonation data; Collect limb movement data of the target object, and extract the limb swing amplitude and movement frequency from the limb movement data; Collect contact sensing data of the target object, and extract the contact force and contact duration from the contact sensing data; The confidence classification module connected to the biosignal sensor determines the emotional recognition of the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, limb swing amplitude, movement frequency, contact force and contact duration. Based on the emotional recognition score, assign corresponding weights to the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, body sway amplitude, movement frequency, contact force, and contact duration. Based on the corresponding weights, the frequency of eye movements, the rate of change of the corner of the mouth, the standard deviation of the fundamental frequency of speech, the rate of change of speech rate, the amplitude of body swaying, the frequency of movement, the contact strength and the contact duration are fused to generate fused emotional features.
3. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, Based on the fused emotional features, an emotional state is determined, including emotional type, emotional intensity value, and emotional intensity change rate, including: The fused emotional features are subjected to probability classification to determine the emotional type; The fused emotional features are subjected to a nonlinear transformation to generate the emotional intensity value. Select the corresponding dynamic analysis strategy based on the emotion type; Obtain the first emotional intensity value at the first moment and the second emotional intensity value at the second moment; The difference in emotional intensity between the second emotional intensity value and the first emotional intensity value is determined based on the dynamic analysis strategy. Based on the time interval between the second moment and the first moment and the difference in emotional intensity, an emotional intensity change rate is generated; An emotional state is generated based on the emotional type, emotional intensity value, and emotional intensity change rate.
4. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, Based on the mapping relationship between the emotional state and preset action parameters, initial action parameters are generated, including: Obtain the emotion type and emotion intensity value in the emotional state; From the preset action parameter mapping relationship, select the corresponding parameter generation strategy according to the emotion type; The emotional intensity value is processed by the parameter generation strategy to generate initial body movement amplitude value, initial movement speed value, initial voice pitch value and initial voice speed value. Boundary constraint processing is performed on the initial limb movement amplitude value, initial movement speed value, initial voice pitch value, and initial voice speed value to generate initial values for limb movement amplitude, initial movement speed, initial voice pitch, and initial voice speed. The initial motion parameters are generated by combining the initial values of the limb movement amplitude, the initial value of the movement speed, the initial value of the voice pitch and the initial value of the voice speed.
5. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, The process involves acquiring the personalized features of the target object and the scene features of the current scene, generating personalized adjustment coefficients based on the personalized features, and generating scene adjustment coefficients based on the scene features, including: The age and extroversion characteristics of the target object are extracted as personalized features; Extract the current scene's type features and ambient volume features as scene features; Based on the age characteristics and personality extroversion characteristics, a personalized adjustment coefficient is generated through a personalized mapping relationship; Based on the aforementioned type characteristics and environmental volume characteristics, scene adjustment coefficients are generated through scene mapping relationships.
6. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, The initial action parameters are processed using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters, including: Obtain the initial values of limb movement amplitude, movement speed, voice pitch, and voice speed from the initial action parameters; Multiply the personalized adjustment coefficient by the scene adjustment coefficient to generate a comprehensive adjustment coefficient; Multiply the comprehensive adjustment coefficient by the initial value of the limb movement amplitude to generate the final limb movement amplitude value; Multiply the overall adjustment coefficient by the initial value of the motion speed to generate the final motion speed value; Multiply the comprehensive adjustment coefficient by the initial value of the voice pitch to generate the final voice pitch value; The final speech rate value is generated by multiplying the comprehensive adjustment coefficient by the initial speech rate value. The final limb movement amplitude value, final movement speed value, final voice pitch value, and final voice speed value are combined to generate the final movement parameters.
7. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, Execute an emotional action based on the final action parameters, obtain feedback data from the target object on the emotional action, and update the action parameter mapping relationship based on the feedback data, including: The robot's joint motors and speech synthesizer are controlled to perform emotional actions based on the final motion parameters. Facial expression response signals and skin conductivity response signals of the target object are collected using biosensors; An emotion receptivity score is generated based on the facial expression response signal; An emotional resonance score is generated based on the skin electrical conductivity response signal; The emotional receptivity score and the emotional resonance score are weighted and fused together to generate a comprehensive feedback score; When the overall feedback score is lower than a preset score threshold, the coefficients in the action parameter mapping relationship are adjusted.
8. A motion generation device based on multimodal emotion perception, characterized in that, The action generation device based on multimodal emotion perception includes: The multimodal perception and fusion module is used to acquire multimodal data of the target object, assign corresponding weights to the multimodal data based on the sentiment recognition degree of each modal data, and fuse the multimodal data based on the corresponding weights to generate fused sentiment features; The emotional state analysis module is used to determine the emotional state, including emotional type, emotional intensity value, and emotional intensity change rate, based on the fused emotional features. The action parameter mapping module is used to generate initial action parameters based on the emotional state and the preset action parameter mapping relationship; The personalization and scene adaptation module is used to obtain the personalized features of the target object and the scene features of the current scene, generate a personalization adjustment coefficient based on the personalized features, and generate a scene adjustment coefficient based on the scene features. The motion parameter optimization module is used to process the initial motion parameters using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final motion parameters. The emotional action execution and feedback learning module is used to execute emotional actions according to the final action parameters, obtain feedback data of the target object on the emotional actions, and update the action parameter mapping relationship based on the feedback data.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal emotion-aware action generation program stored in the memory and executable on the processor. When the multimodal emotion-aware action generation program is executed by the processor, it implements the steps of the multimodal emotion-aware action generation method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores an action generation program based on multimodal emotion perception, which, when executed by a processor, implements the steps of the action generation method based on multimodal emotion perception as described in any one of claims 1-7.
Citation Information
Patent Citations
Intelligent accompanying human-type robot based on multi-modal emotion interaction and sensing method
CN119839861A
Robot behavior mode dynamic adjustment method based on multi-mode perception
CN120257050A
Intelligent teaching interaction feedback method and system driven by multi-modal sentiment analysis
CN120408515A
Dynamic self-adaptive multi-modal sentiment analysis fusion method and system
CN120429691A
Network media video data analysis and supervision system
CN120614475A
Cited By
Virtual image model construction method and system based on image cloning
CN121349311A