Motion generation method and device based on multi-modal emotion perception, equipment and medium

By using multimodal emotion perception technology, which integrates data from multiple modalities to generate emotion features, and dynamically adjusts action parameters based on personalization and scene characteristics, the problem of single perception and delayed response in robot emotional interaction is solved, achieving accurate, personalized and real-time emotional response.

CN120872157BActive Publication Date: 2025-12-16PING AN TECH (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511374154.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-12-16
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing robot emotional interaction technologies suffer from limitations such as a single dimension of emotional perception, lack of dynamic adjustment, and inability to personalize and adapt to specific scenarios, resulting in delayed emotional responses and an inability to provide a natural and accurate emotional interaction experience.

Method used

By acquiring multimodal data, assigning weights to fuse emotional features, determining emotional state, generating adjustment coefficients by combining personalized and scene features, processing initial action parameters, executing emotional actions, and updating parameter mapping relationships.

Benefits of technology

It improves the comprehensiveness and adaptability of emotional perception, achieves the accuracy, personalization and real-time nature of emotional response, and enhances the naturalness of human-computer interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872157B_ABST
    Figure CN120872157B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, can be applied to business scenes such as financial technology and medical health, and discloses a motion generation method, device and equipment based on multi-modal emotion perception and a medium, which comprises the following steps: acquiring multi-modal data of a target object, assigning weights based on emotion recognition degrees and fusing to generate fused emotion features, determining emotion types, emotion intensity values and change rates, combining personalized features and scene features of the target object to generate personalized adjustment coefficients and scene adjustment coefficients, processing initial motion parameters to obtain final motion parameters, executing an emotional action and acquiring feedback data, and updating a motion parameter mapping relationship based on the feedback data. The application improves the comprehensiveness of emotion perception by fusing multi-modal data, dynamically adjusts motion parameters in combination with personalized features and scene features, updates the motion parameter mapping relationship based on feedback data, and improves the accuracy, personalization and real-time performance of emotional response.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a motion generation method and device based on multi-modal emotion perception, equipment and a storage medium. BACKGROUND

[0002] The existing robot emotion interaction technology has many limitations, mainly manifested in single emotion perception dimension, lack of dynamic adjustment of action response to different emotion intensity and types, lack of personalization and scene adaptation of emotional action, and high response delay between emotion perception and action generation, which makes it difficult to provide more natural, accurate and user demand-oriented emotional interaction experience.

[0003] In the field of financial technology business, the existing technology relies on voice or facial expression to judge user emotions, and cannot combine key details such as user body movements, micro-expressions and contact behaviors, making it difficult to accurately identify the real emotional state of customers in customer service. In addition, the current emotional action parameters lack linkage with the perceived emotion intensity and type, and cannot achieve differentiated and adapted emotional action adjustment for different user personal characteristics (such as age, personality) and service scenarios (such as business outlets, mobile environments), making it difficult to improve user experience and service affinity. The response delay of emotional interaction further weakens the real-time interaction effect of robots and customers, and cannot meet the demand for agile emotion recognition and timely pacification in financial scenarios.

[0004] In the field of medical and health business, the existing technology also has the problem of narrow dimension in perceiving the emotional state of patients, especially lacking full use of emotional information such as facial micro-expression, body movement and contact behavior, resulting in one-sided judgment of the psychological state of patients. The emotional action expression of the current robot cannot be dynamically adjusted according to the perceived different emotion types and intensity, and also lacks adaptive adjustment combined with the personal characteristics of patients (such as age, personality) and the actual application scenarios (such as ward, outpatient department), affecting the personalization and specialization of doctor-patient emotional interaction. The response delay between emotion perception and action planning is large, making it difficult for the robot to achieve real-time and natural interaction with the patient, affecting the patient's trust and comfort in the nursing robot.

[0005] In the field of general human-computer interaction, the existing technology is deficient in weak multi-modal data integration capability, ignoring non-traditional emotional expression elements such as body movement amplitude and contact force, resulting in incomplete recognition of complex user emotions. The robot emotional action output is fixed, lacking dynamic matching with the current emotional state of the user, and lacking differentiated adaptation in different user groups and scenarios. The data processing and action planning efficiency is low, and it is difficult to achieve low-delay response in the process of emotional interaction, reducing the naturalness and immersion of human-computer emotional interaction. SUMMARY

[0006] The main purpose of the present application is to provide a multi-modal emotion perception based action generation method, device, equipment and storage medium, aiming at solving the technical problems that the prior art cannot adapt multi-modal emotion perception, personalized features and scene features, dynamically adjust action parameters and form a unified closed loop through real-time feedback update, resulting in lack of comprehensiveness, adaptability and real-time of robot emotional interaction.

[0007] To achieve the above purpose, the present application provides a multi-modal emotion perception based action generation method, comprising:

[0008] Obtain multi-modal data of a target object, and assign corresponding weights to the multi-modal data based on the emotion recognition degree of each modal data in the multi-modal data, and generate fused emotion features based on the corresponding weights of the multi-modal data fusion;

[0009] Based on the fused emotion features, determine an emotion state containing emotion type, emotion intensity value and emotion intensity change rate;

[0010] Based on the emotion state and the preset action parameter mapping relationship, generate initial action parameters;

[0011] Obtain the personalized features of the target object and the scene features of the current scene, generate a personalized adjustment coefficient based on the personalized features, and generate a scene adjustment coefficient based on the scene features;

[0012] Process the initial action parameters through the personalized adjustment coefficient and the scene adjustment coefficient to generate final action parameters;

[0013] According to the final action parameters, perform emotional action, and obtain the feedback data of the target object to the emotional action, and update the action parameter mapping relationship based on the feedback data.

[0014] Further, to achieve the above purpose, the present application provides a multi-modal emotion perception based action generation device, comprising:

[0015] Multi-modal perception and fusion module, for obtaining multi-modal data of a target object, and assigning corresponding weights to the multi-modal data based on the emotion recognition degree of each modal data in the multi-modal data, and generating fused emotion features based on the corresponding weights of the multi-modal data fusion;

[0016] Emotion state analysis module, for determining an emotion state containing emotion type, emotion intensity value and emotion intensity change rate based on the fused emotion features;

[0017] Action parameter mapping module, for generating initial action parameters based on the emotion state and the preset action parameter mapping relationship;

[0018] a personalization and scene adaptation module, configured to obtain a personalization feature of the target object and a scene feature of a current scene, generate a personalization adjustment coefficient based on the personalization feature, and generate a scene adjustment coefficient based on the scene feature;

[0019] an action parameter optimization module, configured to process the initial action parameter by the personalization adjustment coefficient and the scene adjustment coefficient, and generate a final action parameter;

[0020] an emotional action execution and feedback learning module, configured to execute an emotional action according to the final action parameter, obtain feedback data of the target object on the emotional action, and update the action parameter mapping relationship based on the feedback data.

[0021] Further, to achieve the above object, the present application also provides a computer device, which comprises a memory, a processor, and a multi-modal emotion perception based action generation program stored in the memory and executable on the processor, and the multi-modal emotion perception based action generation program, when executed by the processor, implements the steps of the multi-modal emotion perception based action generation method.

[0022] Further, to achieve the above object, the present application also provides a computer readable storage medium, which stores a multi-modal emotion perception based action generation program, and the multi-modal emotion perception based action generation program, when executed by a processor, implements the steps of the multi-modal emotion perception based action generation method.

[0023] Beneficial effects: The present application relates to the field of artificial intelligence, and can be applied to business scenarios such as financial technology and medical health, and discloses a multi-modal emotion perception based action generation method, device, equipment and medium, which comprises: obtaining multi-modal data of a target object, assigning weights based on the emotion recognition degrees of the modal data and generating fused emotion features, determining an emotion state containing an emotion type, an emotion intensity value and an emotion intensity change rate, generating a personalization adjustment coefficient and a scene adjustment coefficient in combination with personalization features and current scene features of the target object, processing initial action parameters to obtain final action parameters, executing an emotional action and obtaining feedback data of the target object, and updating an action parameter mapping relationship based on the feedback data. The present application improves the comprehensiveness of emotion perception by fusing multi-modal data, improves the adaptability by dynamically adjusting action parameters in combination with personalization features and scene features, further updates the action parameter mapping relationship based on feedback data to realize feedback learning and closed-loop optimization, and can enhance the accuracy, personalization and real-time performance of emotional response. BRIEF DESCRIPTION OF DRAWINGS

[0024] The present application will be further described below in conjunction with the accompanying drawings and embodiments, wherein:

[0025] Figure 1 An application environment schematic diagram of the action generation method based on multi-modal emotion perception in an embodiment of the present application;

[0026] Figure 2 A flowchart of the action generation method based on multi-modal emotion perception in an embodiment of the present application;

[0027] Figure 3 A functional module schematic diagram of the action generation device based on multi-modal emotion perception in a preferred embodiment of the present application;

[0028] Figure 4 A structure schematic diagram of a computer device in an embodiment of the present application;

[0029] Figure 5 Another structure schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0030] It should be understood that the specific embodiments described herein are merely illustrative of the present application and do not limit the present application.

[0031] The action generation method based on multi-modal emotion perception provided by the embodiments of the present application can be applied in an application environment as shown in Figure 1 , wherein the user end communicates with the service end through the network. The service end can obtain multi-modal data of a target object through the user end, distribute weights based on the emotion recognition degrees of the modal data and generate fusion emotion features, determine an emotion state containing an emotion type, an emotion intensity value and an emotion intensity change rate, generate a personalized adjustment coefficient and a scene adjustment coefficient in combination with personalized features and current scene features of the target object, process initial action parameters to obtain final action parameters, execute an emotional action and obtain feedback data of the target object, and update the action parameter mapping relationship based on the feedback data. The present application improves the comprehensiveness of emotion perception by fusing multi-modal data, improves the adaptability by dynamically adjusting the action parameters in combination with the personalized features and the scene features, further realizes feedback learning and closed-loop optimization by updating the action parameter mapping relationship based on the feedback data, and can enhance the precision, personalization and real-time performance of the emotional response. The user end can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The service end can be realized by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0032] Please refer to Figure 2 , Figure 2A flowchart of an embodiment of the action generation method based on multi-modal emotion perception provided by the present application is shown. It should be noted that although a logical sequence is shown in the flowchart, in some cases, the steps shown or described can be performed in an order different from that shown.

[0033] As shown in Figure 2 The action generation method based on multi-modal emotion perception provided by the present application includes the following steps:

[0034] S10, acquiring multi-modal data of a target object, and assigning corresponding weights to the multi-modal data based on the emotion recognition degree of each modal data in the multi-modal data, and generating fused emotion features based on the corresponding weights and the fusion of the multi-modal data;

[0035] In this embodiment, in order to accurately process the complex and diverse emotional expressions of users, first, multi-modal data of the target object needs to be collected, which includes but is not limited to facial expression data, voice tone data, body movement data, and contact perception data. Facial expression data can be collected by a high-definition camera device, with particular attention to the activity frequency of the eye area and the change rate of the mouth corner area. The eye activity frequency represents the number of blinks and gazes per unit time, and the mouth corner change rate represents the dynamic amplitude of smiling or frowning actions. Voice tone data is collected by a high-sensitivity microphone, and the extracted features include the standard deviation of the voice fundamental frequency, which is used to reflect the stability and change of the tone, and the speech rate change rate, which is used to reflect the fluctuation of the speaking rhythm. Body movement data is obtained by a three-dimensional motion capture device, focusing on the amplitude of body swing and the action frequency per unit time, which can reflect the motion characteristics in individual emotional expression. Contact perception data is obtained by a pressure sensor array or a capacitive touch sensor, focusing on the contact force and the contact duration. The contact force represents the degree of pressure applied by the target object in the interaction, and the contact duration represents the duration of maintaining physical contact.

[0036] After obtaining the above data, it is necessary to evaluate the emotion recognition degree of each modality data. The emotion recognition degree refers to the distinguishing ability of each modality feature to recognize the true emotion of the target object in the current situation, which can be realized by a trained confidence classification module. The confidence classification module outputs the credibility score of each modality feature according to historical interaction data and a pre-trained emotion recognition model. The emotion recognition degree is used to guide the subsequent weight allocation, and the modality with higher recognition degree is given greater weight to strengthen its contribution in emotion judgment. The weight allocation process can adjust the weights of each modality by a normalization function so that their sum is 1. Through this allocation result, multiple modal data are weighted and fused according to their corresponding weights. In the fusion calculation process, the standardized feature value of each modality is multiplied by its weight and summed to obtain the fused emotion feature that describes the current emotional state. The fused emotion feature retains the advantages of multi-modal input and improves the adaptability and robustness to complex emotional expressions under the weighting strategy.

[0037] When collecting facial expression data, a multi-angle camera array can be used to reduce the influence of light changes and occlusions on data accuracy. Eye movement frequency can be identified by tracking the pupil movement trajectory and calculating the number of movements per unit time. The rate of change of the mouth corner is calculated by tracking the change curve of the mouth corner position and calculating the displacement per unit time. Speech tone data collection can suppress background noise by setting a directional microphone array. The standard deviation of the speech fundamental frequency is obtained by using the Fourier transform with an adaptive window length to obtain the main frequency distribution of the speech frame and calculating the standard deviation. The speech rate change rate is calculated by detecting the phoneme boundary interval to calculate the change range. The body movement data collection can combine an inertial measurement unit (IMU) and an optical motion capture system to calculate the spatial displacement vector of the body node and obtain the swing amplitude and movement frequency. Contact sensing data collection can accurately detect the pressure value of the palm or fingertip contact through a flexible capacitive pressure array and calculate the duration.

[0038] The calculation of emotion recognition degree inputs the collected feature data into the trained classifier model and outputs the confidence score. Different classifier models can include convolutional neural network-based image classifiers, recurrent neural network-based speech classifiers, and long short-term memory network-based sequence data classifiers. The weight allocation corresponding to the emotion recognition degree is realized by a softmax normalization function, which dynamically adjusts the contribution of different modalities within a certain range. The calculated weights are multiplied by each modality data and weighted to form a fused emotion feature.

[0039] Example: In the medical health business field, the application can configure a camera, a microphone, a mattress pressure sensor, and a handheld pressure sensing device for a critically ill patient to collect the patient's weak facial expressions, speech, body swings, and grip changes in real time, and fuse to generate the patient's current emotional state, providing precise non-verbal emotional state indicators for nursing staff.

[0040] In the field of financial technology business, the application can deploy a front camera, a voice collection module and a desktop pressure sensing device in an intelligent customer service terminal to collect the customer's facial tension expression, accelerated speech, the intensity and frequency of finger tapping on the desktop, generate customer emotion features by fusion, and assist the intelligent customer service to identify whether the customer is anxious or angry, and optimize the service response strategy.

[0041] The embodiment can overcome the limitations of a single modality in emotion perception through multi-dimensional collection of multi-modal data and weight distribution guided by emotion recognition degree. Especially when the body movement is weak, the speech expression is incomplete, and the contact data is abnormal, the comprehensive features improve the accuracy of emotion perception in complex environments. Through the confidence module trained to evaluate the emotion distinguishing ability of each modality data and dynamically adjust the weight, the reasonable fusion of multi-modal data is realized, and the sensitivity of the overall system to the details of emotion expression is enhanced.

[0042] S20, determining an emotion state containing an emotion type, an emotion intensity value and an emotion intensity change rate based on the fused emotion features;

[0043] In this embodiment, first, probability classification calculation is needed for the fused emotion features. The goal of probability classification is to output the most possible emotion type from the fused feature vector. The emotion type can include categories such as joy, anger, sadness, surprise, and disgust. In specific implementation, a multi-class classification model such as a softmax classifier can be used to input the fused emotion feature vector and output the probability distribution of each emotion category, and the category corresponding to the maximum probability is selected as the final emotion type. Probability classification can fully utilize the joint information of multi-modal in the fused emotion features, improve the accuracy of emotion classification, and especially when there is conflict between multi-modal, balance the contribution of each modality through the way of fusion probability inference.

[0044] Then, the fused emotion features are subjected to nonlinear transformation to generate emotion intensity values. Emotion intensity values represent the intensity of emotion expression, usually mapped as continuous values from 0 to 1, 0 representing almost no emotion expression, and 1 representing extremely strong emotion expression. Nonlinear transformation can use sigmoid function or tanh function to map multi-dimensional fused emotion features to the interval. This step can solve the scaling problem caused by the difference in the numerical range of different modal features, while maintaining the continuous adjustability of intensity.

[0045] Further according to the emotion type selection dynamic analysis strategy, the dynamic analysis strategy is used to calculate the change of the emotion intensity value between different time points. The calculation of the emotion intensity change rate requires obtaining the emotion intensity values of at least two time points, such as the first time and the second time. The emotion intensity value at the first time is used as the baseline, and the emotion intensity value at the second time is used for comparison with the baseline. The dynamic analysis strategy adopts different calculation rules according to different emotion types, such as a smooth linear difference strategy for sadness type and an exponential change model for anger type, reflecting the dynamic characteristics of different emotions over time. Through the dynamic analysis strategy, the difference value of the emotion intensity at the two times is calculated, and then combined with the time interval between the two times, the emotion intensity change rate is generated. The change rate can reflect whether the emotion is increasing, decreasing or remaining stable, which helps to more comprehensively describe the current emotional state.

[0046] Finally, based on the emotion type, the emotion intensity value and the emotion intensity change rate, the emotion state is formed comprehensively. The emotion state is a structured data set that completely describes the emotional category, emotional intensity and emotional dynamic change characteristics of the target object at the current time point, laying a foundation for subsequent action parameter generation.

[0047] The classification of emotion type can use a multi-layer perceptron neural network as a classifier, fuse emotion features as input, pass through one or more fully connected hidden layers for calculation, and output the probability distribution of emotion type. Network training can be based on existing multi-modal emotion database, such as containing labeled expression, speech, action data, for supervised learning.

[0048] The generation of emotion intensity value can use a regularized sigmoid function to project the linear combination of fused emotion features to the interval of 0 to 1, wherein the weight parameters of linear combination can be obtained through training, so that the distribution of intensity value conforms to the subjective perception of different user groups. For dynamic analysis strategy, different change rate calculation models can be predefined for different emotion types, such as using a weighted average strategy to calculate the smooth change for joy, and using a time-weighted fast decay function model for surprise, to ensure that the change rate can objectively reflect the dynamic characteristics of specific emotions.

[0049] The structure of emotion state can be defined as a vector containing three parts, recording emotion type code, emotion intensity value and emotion intensity change rate. This vector can be stored in the emotion database for subsequent fast association and retrieval of the mapping relationship between action parameters.

[0050] Example: In the medical and health business field, it can be applied to nursing robots. The robot can collect multi-modal information of the patient through the camera, microphone and touch sensor, comprehensively analyze whether the patient is currently in a state of anxiety, depression, etc., judge whether the emotion intensity change trend is intensified, and provide timely psychological state evaluation for medical staff.

[0051] In the field of financial technology business, it can be used for intelligent financial consultants. By analyzing the multi-modal features of the customer's interaction with the terminal, the current tension of the customer and its trend are identified, which helps the consultant to adjust the recommendation strategy or terminate the transaction suggestion when the customer is anxious or emotional, reduces the service risk and improves the user experience.

[0052] By simultaneously extracting the emotion type, emotion intensity value and emotion intensity change rate from the fused emotion features, the embodiment can overcome the deficiency of traditional systems that only identify emotion categories, so that emotion recognition not only includes the type of emotion, but also the intensity of emotion and its trend over time, thereby providing rich and accurate emotional input data for downstream action planning, and improving the naturalness and adaptability of human-computer interaction.

[0053] S30, generating initial action parameters based on the emotion state and the preset action parameter mapping relationship;

[0054] In this embodiment, the emotion type and emotion intensity value in the emotion state need to be analyzed first. The emotion type represents the main emotion category currently identified, such as joy, sadness, anger, surprise, fear, etc., and the emotion intensity value represents the intensity of the emotion, which is usually a standardized numerical range, used to quantitatively express the saliency of the emotion. In actual application, the emotion type can be converted into a corresponding category label through a lookup table or an encoding dictionary, and the emotion intensity value is directly used as a continuous variable in subsequent operations.

[0055] The action parameter mapping relationship is a preset data structure, and its content includes parameter generation strategies for different emotion types. The design of the parameter generation strategy can cover action templates, adjustment rules, boundary conditions, etc. corresponding to different emotion categories, which are used to convert emotion data into parameter values that can be used for action control. The system selects the parameter generation strategy that matches the current emotion type from the action parameter mapping relationship by using the emotion type as the retrieval key. The content of the parameter generation strategy can include linear mapping rules, nonlinear function models, conditional constraints, etc. to ensure that the action parameters generated under different emotion states have significant distinguishability and behavior rationality.

[0056] Based on the emotion intensity value, the selected parameter generation strategy is called to perform numerical conversion. The conversion result is the initial body movement amplitude value, the initial action speed value, the initial voice tone value and the initial voice speed value, which correspond to the movement amplitude, action rhythm, tone change range and speed adjustment amount of the robot when performing emotion expression respectively. In the conversion process, each parameter can use an independent mapping formula, for example, for the anger emotion, the body movement amplitude value may increase linearly with the emotion intensity value, while the voice tone value uses an exponential enhancement model, so as to ensure the vividness and consistency of the action expression.

[0057] After the initial parameters are generated, boundary constraint processing is still needed to avoid parameter values exceeding the control range allowed by the physical device or producing extreme values that do not adapt to the scene. Specific boundary constraints can be implemented by maximum and minimum value limiting functions, for example, limiting the initial value of the action speed to the range of 0.1 m / s to 1.0 m / s to ensure the safety and accuracy of the robot motion. After boundary processing, the initial values of the limb action amplitude, action speed, voice tone, and voice speed are formed.

[0058] Finally, the above four initial values are combined into a structured initial action parameter set, which is used in subsequent action control modules or further personalized and scene adjustment processing steps. This set can be stored or transmitted as a standard data structure, and ensures that the action execution system can directly read and apply it.

[0059] The implementation of the parameter generation strategy can include mapping rules designed for different emotions, for example, defining a linear positive relationship between the limb action amplitude value and the emotion intensity value for the happy emotion, a proportional enhancement relationship between the action speed value and the emotion intensity value, a logarithmic relationship between the voice tone value and the emotion intensity value, and a weighted average relationship between the voice speed value and the emotion intensity value. For the sad emotion, the mapping rules can be defined in reverse, for example, the limb action amplitude decreases with the intensity value, and the voice tone tends to be low. Boundary constraints can be implemented through lookup tables, dynamically adjusting the allowed range to adapt to different devices and scene requirements, for example, the speed upper limit constraint for medical service robots may be more stringent to ensure patient safety.

[0060] The parameter generation strategy can also be dynamically adjusted based on machine learning methods. In the training phase, the emotion type, emotion intensity value, and expert-labeled action parameters are input as training samples, and a regression model is used to learn the mapping relationship. In the deployment phase, the trained mapping model is called in real time to replace the fixed rules.

[0061] Example: In the medical health business field, when the robot faces different emotional states of the patient, especially sadness, anxiety, or depression, it can analyze the current emotion type and emotion intensity value of the patient, dynamically generate initial action parameters that match these emotions according to the pre-set action parameter mapping relationship, for example, moderately reducing the limb action amplitude, lowering the voice tone value and speed value, making the action output more gentle, and avoiding triggering further tension or discomfort in the patient. In implementation, by adjusting the mapping relationship, the robot can generate differentiated initial action parameters for different patient groups under subtle emotional differences, thus better meeting the needs of the medical care scene.

[0062] In the field of financial technology business, when communicating with users, the financial service robot can select a parameter generation strategy with a more calming effect based on a mapping relationship by analyzing the current user emotion type and intensity, such as identifying that the user has a state of anxiety or doubt, for example, appropriately reducing the action amplitude value and reducing the initial value of the speech speed, so that the answer is more calm, and the user's trust is improved. In implementation, this processing step can ensure that the robot can output the initial action parameters optimized based on the current emotional state in the face of the diversified emotions of customers in different financial service scenarios, helping the service process to be smoother and more professional.

[0063] By associating the emotion state with the action parameter mapping relationship, the embodiment can realize different parameter generation logics corresponding to different emotion categories, dynamically adjust the action parameters according to the emotion changes, overcome the deficiencies of the traditional fixed action parameters or the irrelevant with the emotion intensity, ensure that the action performance is highly consistent with the emotion state, and improve the naturalness and emotion adaptability of the robot emotion expression.

[0064] S40, obtaining the individualized features of the target object and the scene features of the current scene, generating an individualized adjustment coefficient based on the individualized features, and generating a scene adjustment coefficient based on the scene features;

[0065] In this embodiment, first, the personalized features of the target object are acquired, including but not limited to age information and personality characteristics. The age information can be extracted from user historical archives or registration data. The personality characteristics can be evaluated according to historical interaction data analysis or pre-questionnaire survey results. The personality extroversion degree is a typical personality characteristic, and its data is derived from social behavior frequency, historical interaction performance and other indicators, which has stability and quantifiable characteristics. The acquisition of personalized features provides basic support for the subsequent generation of personalized adjustment coefficients. The personalized adjustment coefficient is used to represent the amplitude ratio of the action parameter that needs to be adjusted among different users. For example, older users may need to reduce the action intensity and lower the voice tone. Therefore, the coefficient reflects the action adjustment degree closely related to the individual characteristics of the user. Next, the scene features of the current scene are acquired. The scene type can be identified through environmental sensors and context data, such as home scene, office scene, and public place scene. The environmental volume feature can be directly collected by the environmental noise sensor, representing the current environmental background noise level. The collection of scene features ensures that the action parameters are adapted to the environmental state. The personalized adjustment coefficient is obtained by inputting the personalized features into the preset personalized mapping relationship. The mapping relationship can be trained according to historical data or configured by rules. For example, users with high extroversion are mapped to larger action amplitude, and older users are mapped to gentler speech speed. The scene adjustment coefficient is obtained by inputting the scene features into the scene mapping relationship. For example, in a quiet environment, it is mapped to a smaller voice tone value and a lower speech speed. In a noisy environment, it is mapped to a higher voice tone value and a faster speech speed. In the whole process, the personalized adjustment coefficient and the scene adjustment coefficient are used as weight factors in the subsequent processing of the action parameter, providing support for the flexibility of action adjustment for different users and different scenes.

[0066] The age information can be automatically collected through the user registration link and identity binding. At the same time, the historical interaction data of past users and the system is called to analyze the interaction frequency, active dialogue ratio, voice tone characteristics and other dimensions to determine the personality extroversion degree. The current background volume level can be sampled in real time through the environmental microphone as the environmental volume feature. At the same time, the current environment category is determined through the context location data (such as GPS or network location) or visual analysis as the scene type feature. The personalized mapping relationship can be preset in the form of a table or a parameter function, such as mapping the age range to the action intensity adjustment factor and mapping the personality score to the voice adjustment factor through linear interpolation. The scene mapping relationship can be maintained through a rule base, such as mapping to an increased speech speed when the environmental volume is higher than a preset threshold. In this way, the action parameters can be adapted to the specific combination of the user and the environment, and differential adjustment is achieved.

[0067] Example: In the medical health business field, the nursing robot can automatically reduce the action amplitude and voice tone according to the age characteristics and lower extroversion of the elderly patients, so as to reduce the discomfort and psychological pressure of the patients, and further lower the speech speed and voice intensity according to the relatively quiet environment in the ward, so as to optimize the human-computer interaction experience.

[0068] In the financial technology business field, the customer service robot can identify young users with high extroversion characteristics, adjust the voice expression to a more enthusiastic tone, and according to the noisy environment in the business hall, adapt a higher voice tone and a faster speech speed, so that the service is clearer and more efficient in the financial consulting process.

[0069] The embodiment provides a fine-grained action parameter adaptation mechanism by combining the personalized characteristics of the target object and the scene characteristics of the current scene, ensures that the action output conforms to the acceptance habit of the user and adapts to the current environment state, and thus improves the naturalness, comfort and situational adaptability of the emotional action.

[0070] S50, processing the initial action parameters by the personalized adjustment coefficient and the scene adjustment coefficient to generate final action parameters;

[0071] In the embodiment, in the action parameter adjustment process, the personalized adjustment coefficient and the scene adjustment coefficient calculated previously need to be obtained first, and the two coefficients respectively reflect the adjustment requirements of the user individual difference and the current scene characteristics on the action parameters. The source of the personalized adjustment coefficient and the scene adjustment coefficient can be traced back to the mapping calculation of the user profile, the interaction behavior analysis and the environment perception data, so as to ensure that it is consistent with the actual requirements of the user and the scene. Then, the personalized adjustment coefficient and the scene adjustment coefficient are mathematically operated, for example, a comprehensive adjustment coefficient is formed by multiplication operation, and the comprehensive adjustment coefficient is used to uniformly adjust various action parameters, so that the adjustment result takes into account both the user characteristics and the scene adaptability. After the comprehensive adjustment coefficient is calculated, the initial action parameters are adjusted one by one. Specifically, the initial action parameters include the initial value of the body action amplitude, the initial value of the action speed, the initial value of the voice tone and the initial value of the voice speed, which are multiplied by the comprehensive adjustment coefficient to obtain the final body action amplitude value, the final action speed value, the final voice tone value and the final voice speed value. In this process, it is ensured that each parameter after adjustment reflects the user preference and adapts to the current environment requirements. Finally, the four final action parameters are recombined into a structured parameter set to form the final action parameters, which provides direct and available standardized input for subsequent action execution.

[0072] The personalized adjustment coefficient and the scene adjustment coefficient can be calculated by a floating-point multiplication unit to generate a single comprehensive adjustment coefficient. Then, each initial action parameter is input into a hardware multiplier, and the final action parameter value is calculated in real time by combining the comprehensive adjustment coefficient. A parallel computing unit can also be used to process the four parameters in parallel, improving the operation efficiency and reducing the overall delay. The value range of the comprehensive adjustment coefficient can be dynamically adjusted by parameter configuration, for example, defining upper and lower threshold values to prevent abnormal values of the adjusted parameters from being too small or too large. The adjusted parameter value can be directly output to the motion controller and the voice controller to achieve consistent real-time adaptive adjustment effect.

[0073] Example: In the medical health business field, the nursing robot combines the personalized adjustment coefficient (small) of the elderly patient with the scene adjustment coefficient (small) in the quiet ward scene to generate a comprehensive adjustment coefficient, which finally adjusts the small amplitude of the limb movement, the slow movement speed, the low and soft voice tone, and the slow speech speed, making the interaction more soothing and gentle, and meeting the needs of the patient.

[0074] In the financial technology business field, the customer service robot identifies a young and outgoing user in a noisy business hall scene, and combines the personalized adjustment coefficient (large) with the scene adjustment coefficient (large) to form a comprehensive adjustment coefficient. The final action parameters after adjustment are characterized by large amplitude of limb movement, fast movement speed, high voice tone, and fast speech speed, making the service performance more energetic and clear, and improving the customer experience.

[0075] The embodiment uses the personalized adjustment coefficient and the scene adjustment coefficient to form a comprehensive adjustment coefficient, and adjusts the initial action parameters one by one and then recombines them, to ensure that the action output realizes dynamic balance between user personalized preferences and scene characteristics, improve the adaptability and acceptability of action generation, reduce user discomfort, and improve the accuracy and naturalness of emotional response.

[0076] S60, performing an emotional action according to the final action parameter, and obtaining feedback data of the target object to the emotional action, updating the action parameter mapping relationship based on the feedback data.

[0077] In this embodiment, when performing emotional actions, first, the limb action amplitude value, action speed value, voice tone high and low value, and voice speed value in the final action parameter are transmitted to the action control module and the speech synthesis module, respectively, and the controller drives the robot joint actuator, facial expression module, and speech synthesizer according to these parameter values to perform specific emotional actions. After the execution is completed, the feedback data collection process is started for the effect of the emotional action. The feedback data includes the facial expression response signal and the skin conductance response signal collected in real time through the biological sensor, the facial expression response signal reflects the degree of user expression change, and the skin conductance response signal reflects the intensity of user physiological response. These data can objectively reflect the user's immediate emotional acceptance condition. The collected feedback data is converted into a standardized numerical form by the signal processing module and used as a comprehensive feedback index to participate in the update calculation of the action parameter mapping relationship. The operation of updating the action parameter mapping relationship is based on the correlation between the feedback data and the current action parameter, and uses a parameter update algorithm (such as gradient adjustment or incremental update mechanism) to adjust the parameter coefficients of the parameter generation strategy in the mapping relationship, so that the initial action parameters generated in the same emotional state in the subsequent process can better meet the user's needs and adapt to the scene conditions.

[0078] The final action parameters can be directly mapped to the robot joint servo control instructions and the speech synthesis module input parameters through the driver to realize the synchronous output of actions and speech. The feedback data collection can use a high-frame-rate camera and a conductance sensor to monitor the facial expression change and skin conductance response in real time, and filter, normalize, and extract feature values of these data through an embedded signal processing unit. The calculation module for updating the action parameter mapping relationship can use a weighted average method or an adaptive adjustment formula to make the updated parameters more accurately reflect the user's immediate emotional state, and can also realize individual long-term adjustment through the storage of feedback history records to further optimize the system adaptability.

[0079] Example: In the medical and health business field, after the accompanying robot performs comforting limb actions and voice soothing for the rehabilitation patient, the patient's facial expression change and skin conductance response are collected in real time. If the feedback data shows that the emotional acceptance is low, the system adjusts the limb action amplitude and tone parameters corresponding to the sad emotion in the mapping relationship, so that the subsequent action is more gentle and slow, and more in line with the needs of the rehabilitation scene.

[0080] In the financial technology business field, when the customer service robot provides consultation services for users in the business hall, it expresses enthusiasm through actions and speech. If the facial expression response and skin conductance signal collected show that the customer acceptance is low, the system adjusts the speed and tone strategy in the parameter mapping relationship associated with the happy emotion in a timely manner, so that the language expression in the subsequent service is softer and closer to the acceptance preferences of different customer groups.

[0081] The embodiment combines action execution and multi-dimensional feedback data collection and analysis, dynamically adjusts and updates the action parameter mapping relationship, enables the system to continuously learn user preferences and reaction characteristics in multiple rounds of interaction, enhances the personalized adaptability and situational adaptability of action generation, and thus improves the accuracy, naturalness and user satisfaction of emotional response.

[0082] The application relates to the technical field of artificial intelligence, and can be applied to business scenarios such as financial technology and medical health, and discloses a multi-modal emotion perception-based action generation method, device, equipment and medium, which comprises the following steps: acquiring multi-modal data of a target object, assigning weights based on emotion recognition degrees of the multi-modal data and fusing to generate fused emotion features, determining an emotion state containing an emotion type, an emotion intensity value and an emotion intensity change rate, combining personalized features and current scene features of the target object to generate a personalized adjustment coefficient and a scene adjustment coefficient, processing initial action parameters to obtain final action parameters, executing an emotional action and acquiring feedback data of the target object, and updating an action parameter mapping relationship based on the feedback data. The application improves the comprehensiveness of emotion perception by fusing multi-modal data, improves the adaptability by dynamically adjusting action parameters in combination with personalized features and scene features, further updates the action parameter mapping relationship based on feedback data to realize feedback learning and closed-loop optimization, and can enhance the accuracy, personalization and real-time performance of emotional response.

[0083] In one embodiment, the above step S10 comprises:

[0084] S101, facial expression data of a target object are collected, and eye movement frequency and mouth corner change rate in the facial expression data are extracted;

[0085] S102, speech intonation data of the target object are collected, and speech fundamental frequency standard deviation and speech speed change rate in the speech intonation data are extracted;

[0086] S103, body action data of the target object are collected, and body swing amplitude and action frequency in the body action data are extracted;

[0087] S104, contact perception data of the target object are collected, and contact force and contact duration in the contact perception data are extracted;

[0088] S105, a confidence classification module connected with a biological signal sensor is used to determine emotion recognition degrees of the eye movement frequency, the mouth corner change rate, the speech fundamental frequency standard deviation, the speech speed change rate, the body swing amplitude, the action frequency, the contact force and the contact duration;

[0089] S106, assigning the corresponding weights of the eye movement frequency, the mouth corner change rate, the voice fundamental frequency standard deviation, the speech speed change rate, the body swing amplitude, the action frequency, the contact force, and the contact duration based on the emotion recognition degree;

[0090] S107, generating a fused emotion feature based on the corresponding weights of the eye movement frequency, the mouth corner change rate, the voice fundamental frequency standard deviation, the speech speed change rate, the body swing amplitude, the action frequency, the contact force, and the contact duration.

[0091] In the embodiment, collecting the multi-modal data of the target object requires establishing a data path in a multi-dimensional information flow in synchronization, wherein the facial expression data is acquired by a high-resolution image acquisition device, the eye movement frequency is extracted by detecting and time series analyzing the opening and closing changes of the eyelids in each frame of picture by using an image processing algorithm, and the eye movement frequency is obtained by counting the number of opening and closing of the eyelids per unit time. The mouth corner change rate is extracted by marking the regions on both sides of the mouth corner by using a key point positioning algorithm, and the average rate of the movement of the mouth corner region is obtained by combining the displacement change and the time difference. The speech tone data is collected by using a microphone array, the original audio signal is input into a frequency spectrum analysis module, the voice fundamental frequency standard deviation is extracted by using a short-time Fourier transform to calculate the fundamental frequency curve of the audio frame sequence, and the standard deviation of the curve is calculated to measure the pitch fluctuation. The speech speed change rate is calculated by counting the number of syllables per unit time and calculating the time gradient by using a syllable segmentation recognition algorithm. The body movement data is collected by using an inertial measurement unit (IMU) or a three-dimensional motion capture system, the body swing amplitude is obtained by tracking the spatial coordinate changes of the key points of the body and calculating the spatial displacement amplitude, and the action frequency is obtained by counting the complete swing cycles in a time window. The contact perception data is collected by using a distributed pressure sensor array, the contact force is obtained by calculating the weighted average of the pressure values of each sensing point, and the contact duration is calculated by recording the trigger and relaxation time points.

[0092] In the confidence level classification module, a trained emotion classification model is introduced for each of the above eight types of features, and the confidence level of each feature matching the emotion category is calculated. The confidence level is the emotion recognition degree. Each emotion recognition degree value will be used as the basis for the importance weight distribution of the feature in the multi-modal fusion. The weight distribution operation is processed by standardization, so that the sum of all weights is 1, to ensure that the relative contribution of each type of feature to the final result in the subsequent weighted fusion conforms to its recognition effect.

[0093] Based on the assigned corresponding weights, a weighted fusion operation is performed. During the fusion process, each feature value is multiplied by its corresponding weight, and then all the weighted results are summed up to generate a fused sentiment feature. The fused sentiment feature is output in the form of a multi-dimensional vector as the input for subsequent sentiment state determination. The above processing process requires strict synchronization in the time dimension to ensure that the timestamps of the modal data are consistent, so that the fusion result reflects the user's comprehensive emotional signal in the same time period.

[0094] The embodiment realizes the selection of the most recognizable features from multi-source information and the integration with dynamic weights based on the synchronous acquisition of multi-modal data, accurate feature extraction, classification confidence calculation and weighted fusion. This not only avoids one-sidedness caused by limitation of a single mode, but also enhances the ability to distinguish complex emotional states, so that the system can form a comprehensive representation of the target object's real, multi-dimensional emotional state, thereby providing a high-confidence emotional input basis for subsequent emotional response.

[0095] In one embodiment, the above step S20 includes:

[0096] S201, performing probability classification processing on the fused sentiment feature to determine the sentiment type;

[0097] S202, performing nonlinear transformation processing on the fused sentiment feature to generate the sentiment intensity value;

[0098] S203, selecting a corresponding dynamic analysis strategy according to the sentiment type;

[0099] S204, obtaining a first sentiment intensity value at a first time and a second sentiment intensity value at a second time;

[0100] S205, determining a sentiment intensity difference between the second sentiment intensity value and the first sentiment intensity value based on the dynamic analysis strategy;

[0101] S206, generating a sentiment intensity change rate based on the time interval between the second time and the first time and the sentiment intensity difference;

[0102] S207, generating a sentiment state based on the sentiment type, the sentiment intensity value and the sentiment intensity change rate.

[0103] In this embodiment, first, a probability classification process is performed on the fused sentiment features, which requires a pre-trained multi-class sentiment recognition model. The model accepts a multi-dimensional vector of fused sentiment features as input and outputs a probability distribution of sentiment categories. The classification result is the sentiment label corresponding to the maximum probability. In this process, the sentiment type is limited to a set of discrete high-level abstract labels, such as joy, sadness, anger, calm, and anxiety. The classification model can use a convolutional neural network, a recurrent neural network, or a multi-layer perceptron enhanced by an attention mechanism. During training, the model is based on a cross-modal sentiment annotation dataset to ensure that it can fully adapt to multi-modal fused features.

[0104] Next, a nonlinear transformation process is used to generate a sentiment intensity value. Specifically, the fused sentiment features are input into a nonlinear activation function module, such as a sigmoid function or a tanh function, to map the original multi-dimensional continuous values to a predefined interval range, such as between 0 and 1 or between -1 and 1, to represent the intensity of the current sentiment. This mapping ensures that the influence of different modal scale inconsistencies on the intensity value is normalized and adjusted. Additionally, the curvature and threshold of the transformation curve can be adjusted according to different sentiment types, such as using a steeper sigmoid curve for the anger sentiment which is easily affected by external interference, to highlight the sensitivity of sentiment intensity changes.

[0105] According to the determined sentiment type, a corresponding dynamic analysis strategy is selected. This strategy is a set of analysis functions associated with different sentiment types, used to explain the law of sentiment intensity change over time. For example, a linear fitting strategy is used for sadness to capture the slow increasing or decreasing trend, and an exponential fitting strategy is used for anger to reflect its suddenness and rapid decline characteristics. The dynamic analysis strategy is stored in a configurable strategy library, and the system directly calls it according to the sentiment type as the retrieval key.

[0106] When obtaining the first and second sentiment intensity values, time series data points are extracted from the historical data buffer to ensure that the timestamps are strictly aligned. The first and second sentiment intensity values are read out from the sentiment intensity value storage unit synchronized with time and input into subsequent processing.

[0107] Based on the dynamic analysis strategy, the difference between the second sentiment intensity value and the first sentiment intensity value is calculated to obtain a sentiment intensity difference value, which reflects the change amplitude of the sentiment intensity between the two time points. During the difference calculation process, the operation rules can be adjusted according to different strategies, such as introducing a smoothing filter for high-frequency fluctuation type emotions to eliminate abnormal noise.

[0108] The sentiment intensity change rate is obtained by dividing the sentiment intensity difference value by the time interval between the two time points. This change rate is a standard quantitative indicator of the dynamic change speed of the sentiment, with time standardization properties, ensuring comparability on different time scales.

[0109] Finally, the emotion type, the emotion intensity value and the emotion intensity change rate are integrated to generate an emotion state, which is in a structured data representation form, such as a JSON object or a key-value pair set, and contains an emotion label, a current intensity value and a change rate, which are used as parameter inputs for subsequent emotion actions to ensure that the subsequent modules can perform targeted action planning and output according to the state.

[0110] The embodiment can accurately extract the subjective emotion category and intensity level expressed by the target object on the basis of multi-modal information and further capture the change trend of the emotion over time, thereby realizing a full-dimensional and high-dynamic-precision description of the emotion state as a whole, laying a foundation for subsequent adaptive adjustment and personalized response of emotion actions, and solving the problems that the traditional single-point emotion recognition cannot perceive the change of emotion over time and lacks dynamic adaptation.

[0111] In one embodiment, the step S30 includes:

[0112] S301, obtaining an emotion type and an emotion intensity value in the emotion state;

[0113] S302, selecting a corresponding parameter generation strategy from the preset action parameter mapping relationship according to the emotion type;

[0114] S303, processing the emotion intensity value by using the parameter generation strategy to generate an initial body movement amplitude value, an initial action speed value, an initial voice tone value and an initial voice speed value;

[0115] S304, performing boundary constraint processing on the initial body movement amplitude value, the initial action speed value, the initial voice tone value and the initial voice speed value to generate a body movement amplitude initial value, an action speed initial value, a voice tone initial value and a voice speed initial value;

[0116] S305, combining the body movement amplitude initial value, the action speed initial value, the voice tone initial value and the voice speed initial value to generate an initial action parameter.

[0117] In this embodiment, when obtaining the emotion type and emotion intensity value in the emotional state, it is necessary to first extract the fields from the emotional state structured data. The emotion type is qualitative description data, such as categories of joy, sadness, anger, etc. The emotion intensity value is a quantitative value corresponding to the type, usually a real number between 0 and 1, used to describe the intensity of emotion. The emotion type is encoded in the form of a string or an enumeration value, and the emotion intensity value is expressed in the form of a floating-point number. Data extraction can be achieved by parsing the structured data format, such as JSON parsing or database query interface implementation.

[0118] When selecting a parameter generation strategy from the preset action parameter mapping relationship, the mapping relationship is stored in the form of a mapping table or an associated database. The emotion type is used as a query key field, and each emotion type corresponds to a set of strategy rules. The strategy rules define how to calculate different action parameter initial values based on the emotion intensity value. The parameter generation strategy includes function models for different emotion types, such as linear scaling strategy for joy type, decreasing adjustment strategy for sadness type, and exponential enhancement strategy for anger type. These function models can be called through lookup table or function interface.

[0119] The emotion intensity value is processed by the parameter generation strategy to calculate the initial limb movement amplitude value, the initial action speed value, the initial voice tone value, and the initial voice speed value. The calculation formulas of each initial value can be independent of each other. For example, the limb movement amplitude value can be calculated by linear amplification of the intensity value, the action speed value can be adjusted by non-linear scaling based on the intensity value, the voice tone value can be calculated by a function positively related to the intensity value, and the voice speed value can be determined by weighting operation on the intensity value. This processing requires calling mathematical function modules to calculate the four initial values one by one.

[0120] When performing boundary constraint processing, an upper limit and a lower limit are applied to each initial value to ensure that the output value is within a safe and acceptable range. For example, the limb movement amplitude value is limited between a minimum value of 0.1 and a maximum value of 1.0, the action speed value is limited between 0.2 and 2.0, the voice tone value is limited within a range related to the pitch, and the voice speed value is limited to a speed range suitable for human hearing. Boundary constraints are implemented through minimum and maximum value comparison operations using mathematical min and max functions to envelope the initial values.

[0121] Finally, the initial action parameters are generated by combining the initial limb movement amplitude value, the initial action speed value, the initial voice tone value, and the initial voice speed value. The initial action parameters are expressed in the form of a data structure, such as an object containing four fields, which record the four action parameter values as input for the next action adjustment and execution module.

[0122] The embodiment combines the emotional state and action parameter mapping relationship, so that different emotional types and intensity values can be mapped to the initial values of four types of key action parameters, and the controllable range of each parameter is limited through separate boundary constraints, ensuring that the initial action parameters have detailed distinguishability and adaptive adjustment capability for input emotions, solving the problem that the existing method cannot dynamically adjust the action amplitude, speed, tone and speed according to different emotional types and intensity, thereby improving the diversity and individualization level of output emotional actions.

[0123] In one embodiment, the above step S40 comprises:

[0124] S401, extracting the age feature and personality extroversion degree feature of the target object as individualization features;

[0125] S402, extracting the type feature and environmental volume feature of the current scene as scene features;

[0126] S403, generating an individualization adjustment coefficient based on the age feature and personality extroversion degree feature through an individualization mapping relationship;

[0127] S404, generating a scene adjustment coefficient based on the type feature and environmental volume feature through a scene mapping relationship.

[0128] In the embodiment, when the age feature and personality extroversion degree feature of the target object are extracted as individualization features, the age feature is obtained by querying the user's historical archives or the user's provided personal data, and is usually represented in the form of an integer or a timestamp, and the personality extroversion degree feature is calculated through questionnaire scoring, historical interaction behavior analysis, social behavior pattern mining, etc., and can be represented by interval value, label value or percentage grade. Data extraction needs to call interfaces from structured data storage or user behavior logs to obtain and format conversion to adapt to subsequent calculation modules.

[0129] When the type feature and environmental volume feature of the current scene are extracted as scene features, the scene type feature is determined through multi-source information such as context recognition, location service data, and environment description information, for example, distinguishing between home, office, and public places, and is usually encoded with a classification label, and the environmental volume feature is collected in real time through a connected environmental microphone and sound sensor device, and is represented by a sound pressure level unit decibel (dB) or a linear unit, and the collection process combines real-time sampling, noise filtering, and short-time statistical processing technology to ensure the accuracy and representativeness of the scene volume data.

[0130] When the personalized mapping relationship is used to generate the personalized adjustment coefficient based on the age feature and the personality extroversion feature, the personalized mapping relationship is implemented by a rule engine or a function model. For example, a more moderate adjustment coefficient is adapted for a user with a smaller age, and a more significant adjustment coefficient is adapted for a user with a high personality extroversion. The mapping relationship is stored in a mathematical expression or a lookup table. The calculation module applies a corresponding function rule to the input personalized feature value, and outputs a personalized adjustment coefficient in a preset range as an adjustment factor for subsequent parameter correction.

[0131] When the scene mapping relationship is used to generate the scene adjustment coefficient based on the scene type feature and the environmental volume feature as the scene features, the scene mapping relationship relies on standard adaptation rules for different types of scenes. For example, a low-volume scene adjustment coefficient should be adapted in a library scene, and a high-activity scene adjustment coefficient should be adapted in a party scene. The scene type label and the environmental volume value are used as inputs. The scene mapping relationship outputs a scene adjustment coefficient through multi-condition branching logic or weighted calculation, so as to ensure that the parameter adjustment is adaptively associated with the current environmental context.

[0132] In this embodiment, the age and personality extroversion of the target object are mapped into the personalized adjustment coefficient, and the scene type and environmental volume are mapped into the scene adjustment coefficient, so that the motion parameters can be adapted to the individual features of the user and the current environmental features. The problem of lack of user personalization and scene adaptation in the existing system is solved, thereby significantly improving the acceptance and scene adaptability of the emotional motion, especially when facing various users and complex scenes, the flexibility and naturalness of emotional expression can be maintained.

[0133] In one embodiment, the above step S50 includes:

[0134] S501, obtaining the initial values of the initial motion parameters, including the initial value of the body motion amplitude, the initial value of the motion speed, the initial value of the voice tone, and the initial value of the voice speed;

[0135] S502, multiplying the personalized adjustment coefficient and the scene adjustment coefficient to generate a comprehensive adjustment coefficient;

[0136] S503, multiplying the comprehensive adjustment coefficient and the initial value of the body motion amplitude to generate a final body motion amplitude value;

[0137] S504, multiplying the comprehensive adjustment coefficient and the initial value of the motion speed to generate a final motion speed value;

[0138] S505, multiplying the comprehensive adjustment coefficient and the initial value of the voice tone to generate a final voice tone value;

[0139] S506, multiplying the comprehensive adjustment coefficient and the initial value of the voice speed to generate a final voice speed value.

[0140] S507, combine the final limb movement amplitude value, the final movement speed value, the final voice tone value and the final voice speed value to generate a final action parameter.

[0141] In this embodiment, the initial action parameters include limb movement amplitude initial value, movement speed initial value, voice tone initial value and voice speed initial value, which are generated by the pre-module based on the emotional state, representing the basic emotional expression intention. The limb movement amplitude initial value represents the amplitude quantization value of the spatial displacement in the limb movement trajectory, usually expressed in normalized amplitude value; the movement speed initial value represents the degree of completion of the limb movement per unit time, which can be expressed in standard speed unit or normalized proportion; the voice tone initial value corresponds to the subjective pitch baseline of the audio signal, expressed in hertz (Hz) or relative proportion; the voice speed initial value is the rhythm speed of voice output per unit time, usually described in words per minute (WPM) or relative unit. These initial values are stored in the form of parameter list and passed to the subsequent adjustment module as input of the parameter matrix.

[0142] The personalized adjustment coefficient and the scene adjustment coefficient are output as separate scalar parameters, respectively, for dynamic correction of the initial action parameters. The two are calculated by a multiplication operator to generate a comprehensive adjustment coefficient, which reflects the comprehensive influence weight of individual characteristics and the current scene. The calculation of the comprehensive adjustment coefficient does not involve data order dependence and can be processed in parallel, improving the overall calculation efficiency and being suitable for low-latency scenarios.

[0143] The multiplication operation of the comprehensive adjustment coefficient and the limb movement amplitude initial value is a single numerical multiplication, which is used to adjust the limb movement amplitude to match the user's preference and scene requirements. The same calculation logic is applicable to the multiplication processing of the movement speed initial value, the voice tone initial value, the voice speed initial value and the comprehensive adjustment coefficient. The result of each multiplication processing is the real-time adjustment value of the target output parameter, ensuring that the personalization and scene adaptation are reflected in the final output.

[0144] The final limb movement amplitude value, the final movement speed value, the final voice tone value and the final voice speed value are packaged into a final action parameter set by a data packager after calculation, and the set format is consistent with the interface standard of the subsequent executor, supporting standard serialization formats such as JSON or binary protocol, facilitating decoding and calling of the robot motion control system and the voice synthesis module.

[0145] The embodiment realizes fine-grained dynamic adjustment of emotional action parameters by fusing the personalized adjustment coefficient and the scene adjustment coefficient into a comprehensive adjustment coefficient and applying it to the initial values of the limb action amplitude, action speed, voice tone, and voice speed. This method overcomes the problem of traditional emotional action fixed templates, enables the action parameters to have adaptive ability in time, individual differences, and environmental changes, and can generate action expressions that are more in line with the user's personality and the current scene requirements in real-time interaction, thereby improving the adaptability and naturalness of the robot's emotional interaction with diverse users and complex environments.

[0146] In one embodiment, the above step S60 comprises:

[0147] S601, controlling the robot joint motor and the speech synthesizer to perform emotional action according to the final action parameters;

[0148] S602, collecting facial expression response signals and skin conductance response signals of the target object through a biological sensor;

[0149] S603, generating an emotional receptivity score based on the facial expression response signals;

[0150] S604, generating an emotional resonance score based on the skin conductance response signals;

[0151] S605, weighting and fusing the emotional receptivity score and the emotional resonance score to generate a comprehensive feedback score;

[0152] S606, when the comprehensive feedback score is lower than a preset score threshold, adjusting the coefficients in the action parameter mapping relationship.

[0153] In the embodiment, the final action parameters are the output of the previous parameter adjustment module and include four sets of parameters: limb action amplitude value, action speed value, voice tone value, and voice speed value. These parameters correspond to specific execution instructions for robot motion execution and voice output. When controlling the robot joint motor to perform emotional action, the limb action amplitude value and the action speed value are used to set the target rotation angle and angular velocity instructions of the motor, respectively. The motion range and speed of each joint are accurately controlled by the corresponding fields in the parameters, ensuring that the action is coherent and consistent with the intended emotional expression intention. The voice tone value and the voice speed value are input into the fundamental frequency control and speech speed modulation modules of the speech synthesizer, respectively, to ensure that the tone and rhythm of the voice output are consistent with the action expression, realizing multi-modal synchronous emotional expression.

[0154] When the robot performs the emotional action, the biosensor module is started, the facial expression response signal is collected through the visual sensor array, including facial muscle movement characteristics and key point movement trajectories, the data format is standardized as a feature vector sequence for subsequent calculation. The skin conductance response signal is measured by the electrode array attached to the user's skin surface, and the signal is recorded in the form of a conductivity change curve, reflecting the user's emotional arousal level.

[0155] The generation process of the emotional acceptance score adopts a pattern matching algorithm for facial expression response signals, and similarity calculations are performed between the facial expression key point movement characteristics collected in real time and various emotional expression templates in the emotional standard template library. The similarity score is standardized to output the emotional acceptance score, quantifying the degree of fit between the current facial expression and the expected emotional action. The generation of the emotional resonance score is based on dynamic feature extraction of the skin conductance response signal, which quantifies the correlation between the user's current physiological response and the target emotional arousal state by weighting and combining indicators such as the average rising rate and fluctuation amplitude of the signal curve, and outputs as the resonance score.

[0156] The generation of the comprehensive feedback score adopts a weighted fusion algorithm, with the emotional acceptance score and the emotional resonance score as weighted inputs. The weight value can be pre-set or dynamically adjusted to reflect the focus of the current user feedback evaluation. When the comprehensive feedback score is lower than the pre-set score threshold, the parameter self-adaptive adjustment logic is executed, including the adjustment of the coefficient table entries corresponding to various emotional types in the action parameter mapping relationship. The adjustment process uses the deviation of the feedback score from the threshold as the weight basis for the adjustment amplitude, and the coefficient update uses a recursive update formula to make the mapping relationship output parameters closer to the user's preferences for similar emotional states in the future.

[0157] The entire process operates in a real-time closed-loop mechanism, with each round of action execution, feedback collection, score calculation, and coefficient adjustment completed within milliseconds, ensuring that the robot dynamically corrects the deviation between action expression and user feedback in continuous interaction, and gradually adapts the action parameter mapping relationship to user individuality and scene environment. The comprehensive design ensures that the final emotional action not only accurately responds to the user's emotional state, but also has the ability to gradually optimize historical interaction experience, enhancing the naturalness and individualization of the human-computer interaction process.

[0158] Example: In the field of medical health, for example, in the application of emotional auxiliary robots for long-term care patients or rehabilitation patients, the robot first collects various modal data of the patient by integrating visual cameras, voice sensors, touch sensors, and motion sensors, including facial expression data, voice tone data, body movement data, and contact perception data when the patient interacts with the device. Extract the eye movement frequency and mouth corner change rate as important facial features from the collected facial expression data, extract the voice fundamental frequency standard deviation and speech rate change rate from the voice tone data to reflect the emotional voice, extract the body swing amplitude and movement frequency from the body movement data to determine the patient's physical activity state, and extract the contact force and contact duration from the contact perception data as a supplement to the patient's active interaction intention. Through the confidence classification module connected with the biological signal sensor, calculate the contribution of these features to emotion recognition, and use the emotion recognition degree as a basis to assign corresponding weights. After weighting and fusing the modal features, the fusion emotion features are generated.

[0159] Based on the generated fusion emotion features, the robot system performs probabilistic classification processing through the built-in emotion classification model to determine the current emotional type of the patient, such as anxiety, depression, or joy, etc. Further, the system quantifies the emotional intensity value from the fusion emotion features through a nonlinear mapping algorithm, combines the intensity changes before and after the patient, selects a matching dynamic analysis strategy based on different emotional types, calculates the current emotional intensity change rate, and finally determines the emotional state including emotional type, emotional intensity value, and emotional intensity change rate.

[0160] For different emotional states, the robot selects the corresponding parameter generation strategy based on the preset action parameter mapping relationship. For example, the anxiety state corresponds to slow and smooth body movements and soft voice tone, and the joy state corresponds to high-tension movements and high-pitched tone. The system uses the emotional intensity value as input to calculate the initial body movement amplitude value, initial movement speed value, initial voice tone value, and initial voice speed value, and performs boundary constraints on each value, such as ensuring that the movement speed is not lower than a certain minimum safety threshold and the speech speed is not faster than the upper limit that can be understood. After boundary constraint, combine the above parameters to form the initial action parameters for driving the robot's movements and voice output.

[0161] Meanwhile, the system extracts personalized features of the patient, such as age and personality extroversion, and combines the type of current medical scene (such as a ward or a rehabilitation training room) and environmental volume features to generate personalized adjustment coefficients and scene adjustment coefficients through personalized mapping relationships and scene mapping relationships, respectively. For example, the system automatically reduces the motion amplitude and voice volume for an elderly introverted patient in a night environment in a ward. The personalized adjustment coefficients and the scene adjustment coefficients are multiplied to form a comprehensive adjustment coefficient, which is used to adjust the four values of the initial motion parameters to obtain the final limb motion amplitude value, the final motion speed value, the final voice tone value, and the final voice speed value, and then the four values are combined to generate the final motion parameters.

[0162] Based on the final motion parameters, the robot outputs corresponding emotional actions and voice expressions, such as soft tone and gentle gestures of condolence, to the patient by controlling joint motors and voice synthesizers. Subsequently, the robot monitors the actual feedback of the patient to the emotional action using biosensors to collect facial expression response signals (such as eyebrow lifting or eye corner drooping) and skin conductance response signals (reflecting the level of autonomic nervous activity). The algorithm calculates the emotional receptivity score and the emotional resonance score, respectively, and weights them to obtain a comprehensive feedback score. If the comprehensive feedback score is lower than a preset score threshold, the system adjusts the parameter coefficients in the motion parameter mapping relationship so that the next round of motion parameters is more suitable for the patient's personality and current emotional response characteristics.

[0163] Through this closed-loop process, the emotional assistance robot not only can real-time multi-modal perceive the emotional state of the patient, but also can dynamically adjust the output emotional actions and voice expressions according to the personalized differences and scene needs, collect and analyze the feedback information of the patient for continuous optimization. This adaptive adjustment capability improves the emotional naturalness, receptivity and comfort of the robot in the interaction with the patient in the medical and health environment, and provides a precise and intelligent solution for the rehabilitation support and emotional care of the patient.

[0164] In the field of financial technology business, such as the scene of financial service counter intelligent assistant or online intelligent financial consultant, the system first performs real-time perception on the customer through multi-modal acquisition devices to collect multi-modal data, including customer facial expression data, voice tone data, body movement data, and contact perception data when the customer interacts with the terminal device. The eye activity frequency and the mouth corner change rate are extracted from the facial expression data, the voice fundamental frequency standard deviation and the speech rate change rate are extracted from the voice tone data, the body swing amplitude and the motion frequency are extracted from the body movement data, and the contact force and the contact duration are extracted from the contact perception data. All data are input into the confidence classification module connected with the biosignal sensor to calculate the contribution of these features to the emotional judgment as the emotional recognition degree. Based on the emotional recognition degree, the weight of each feature is assigned, and then the weighted fusion of each modality is performed to form the fused emotional features, which comprehensively reflect the current emotional state of the customer.

[0165] The financial intelligent assistant determines the customer sentiment type, such as nervousness, anxiety or confidence, through a probabilistic classification algorithm on this basis. A non-linear transformation method calculates the current sentiment intensity value, and a dynamic analysis strategy combines the intensity changes at consecutive time points to calculate the sentiment intensity change rate, and comprehensively determines the sentiment state, which is used to reflect the emotional dynamics of the customer when facing the recommendation of financial products or risk prompts.

[0166] In combination with the preset action parameter mapping relationship, the system selects different parameter generation strategies according to the sentiment type and sentiment intensity value determined in the sentiment state. For example, when the customer shows high anxiety, the system selects a soothing gesture, lowers the voice tone and speech speed. Using the sentiment intensity value as input, the initial limb movement amplitude value, initial movement speed value, initial voice tone value and initial voice speed value are calculated. Boundary constraint processing is performed on each initial value, such as ensuring that the speech speed is not lower than a threshold value that can be understood or the gesture movement does not appear too slow. The four constrained parameters are combined to form the initial action parameters, which are used for subsequent interaction output.

[0167] At the same time, the system collects the customer's personalized features, such as age, personality extroversion degree, etc., to cope with the acceptance differences of different age groups of customers to communication methods. It also collects the type characteristics of the current scene (such as high-end financial hall, ordinary business hall) and the environmental volume characteristics. Through personalized mapping relationship and scene mapping relationship, personalized adjustment coefficients and scene adjustment coefficients are generated respectively, for example, for older and introverted customers in a quiet environment, the movement amplitude and voice volume are adjusted to a more moderate level. The personalized adjustment coefficient and the scene adjustment coefficient are multiplied to form a comprehensive adjustment coefficient, which is used to adjust the initial action parameters to generate the final limb movement amplitude value, the final movement speed value, the final voice tone value and the final voice speed value, and then combined into the final action parameters.

[0168] The final action parameters drive the physical behavior and voice output of the financial intelligent assistant, such as adjusting the tone of speech, smile amplitude, movement frequency, so that the service behavior is more in line with the current emotional state of the customer. Thereafter, the system continuously collects the customer's facial expression response signals (such as smiling or frowning) and skin conductance response signals (reflecting tension level) through biosensors, calculates the sentiment acceptance score and the sentiment resonance score, and weights and fuses them into a comprehensive feedback score. If the comprehensive feedback score is lower than the preset score threshold, the system will adjust the coefficients of the action parameter mapping relationship, so that the movement and speech tone in subsequent services are more in line with the current customer's personality and emotional characteristics.

[0169] The intelligent service assistant can realize personalized and emotion-sensing sensitive financial consulting interaction through real-time multi-modal perception and adaptive adjustment in a financial technology business scene. It can reduce customer anxiety, improve customer comfort and trust in the financial service process, and enhance the professionalism and humanization level of intelligent services of financial institutions, especially when facing high net worth customers or risk-sensitive customers, showing higher service precision.

[0170] The embodiment realizes fine quantization of multi-modal user feedback by directly associating the final action parameter with the execution of the robot joint motor and the voice synthesizer, and synchronously collecting and processing facial expression response signals and skin conductance response signals, which can accurately depict the actual emotional acceptance degree and physiological resonance level of the target object. On this basis, the weighted fusion of emotional acceptance score and emotional resonance score is adopted to dynamically calculate the comprehensive feedback score, and whether the comprehensive feedback score is lower than the preset score threshold is used as a condition to timely adjust the coefficient in the action parameter mapping relationship, so that the subsequent action parameter generation is closer to the user's individuality and situational needs. The process has real-time performance, and the closed-loop adaptive update makes the emotional action expression and the user's real reaction consistent. Through the above-mentioned manner, the individuality, adaptability and accuracy of the robot's emotional expression are effectively improved, the deviation between the action and voice output and the user's emotional state is reduced, and the comfort and naturalness of the user's interactive experience are improved.

[0171] In an embodiment, a multi-modal emotion perception-based action generation device is provided, which corresponds to the multi-modal emotion perception-based action generation method in the above-mentioned embodiments. Referring to Figure 3 , Figure 3 A functional module schematic diagram of a preferred embodiment of the multi-modal emotion perception-based action generation device of the present application. The multi-modal perception and fusion module 10, the emotion state analysis module 20, the action parameter mapping module 30, the individualization and scene adaptation module 40, the action parameter optimization module 50, and the emotional action execution and feedback learning module 60. The detailed description of each functional module is as follows:

[0172] The multi-modal perception and fusion module 10 is used to obtain multi-modal data of a target object, and assign corresponding weights to the multi-modal data based on the emotion recognition degree of each modal data in the multi-modal data, and generate fused emotion features based on the corresponding weights of the multi-modal data;

[0173] The emotion state analysis module 20 is used to determine an emotion state containing an emotion type, an emotion intensity value and an emotion intensity change rate based on the fused emotion features;

[0174] The action parameter mapping module 30 is used to generate initial action parameters based on the emotion state and a preset action parameter mapping relationship.

[0175] a personalization and scenario adaptation module 40 configured to obtain a personalization feature of the target object and a scenario feature of a current scenario in which the target object is located, generate a personalization adjustment coefficient based on the personalization feature, and generate a scenario adjustment coefficient based on the scenario feature;

[0176] an action parameter optimization module 50 configured to process the initial action parameter by using the personalization adjustment coefficient and the scenario adjustment coefficient, and generate a final action parameter;

[0177] an emotional action execution and feedback learning module 60 configured to execute an emotional action according to the final action parameter, obtain feedback data of the target object on the emotional action, and update the action parameter mapping relationship based on the feedback data.

[0178] In an embodiment, the multi-modal perception and fusion module 10 is specifically configured to:

[0179] collect facial expression data of the target object, and extract an eye movement frequency and a mouth corner change rate in the facial expression data;

[0180] collect speech intonation data of the target object, and extract a speech fundamental frequency standard deviation and a speech speed change rate in the speech intonation data;

[0181] collect limb action data of the target object, and extract a limb swing amplitude and an action frequency in the limb action data;

[0182] collect contact perception data of the target object, and extract a contact force and a contact duration in the contact perception data;

[0183] determine emotional recognition degrees of the eye movement frequency, the mouth corner change rate, the speech fundamental frequency standard deviation, the speech speed change rate, the limb swing amplitude, the action frequency, the contact force, and the contact duration by using a confidence classification module connected to a biological signal sensor;

[0184] assign corresponding weights of the eye movement frequency, the mouth corner change rate, the speech fundamental frequency standard deviation, the speech speed change rate, the limb swing amplitude, the action frequency, the contact force, and the contact duration based on the emotional recognition degrees;

[0185] fuse the eye movement frequency, the mouth corner change rate, the speech fundamental frequency standard deviation, the speech speed change rate, the limb swing amplitude, the action frequency, the contact force, and the contact duration based on the corresponding weights, and generate a fused emotional feature.

[0186] In an embodiment, the emotional state analysis module 20 is specifically configured to:

[0187] perform a probability classification process on the fused emotional feature, and determine the emotional type.

[0188] performing nonlinear transformation on the fusion emotional feature to generate the emotional intensity value;

[0189] selecting a corresponding dynamic analysis strategy according to the emotional type;

[0190] obtaining a first emotional intensity value at a first time and a second emotional intensity value at a second time;

[0191] determining an emotional intensity difference value between the second emotional intensity value and the first emotional intensity value based on the dynamic analysis strategy;

[0192] generating an emotional intensity change rate based on a time interval between the second time and the first time and the emotional intensity difference value;

[0193] generating an emotional state based on the emotional type, the emotional intensity value and the emotional intensity change rate.

[0194] In an embodiment, the action parameter mapping module 30 is specifically configured to:

[0195] obtain an emotional type and an emotional intensity value in the emotional state;

[0196] select a corresponding parameter generation strategy according to the emotional type from the preset action parameter mapping relationship;

[0197] generate an initial body movement amplitude value, an initial action speed value, an initial voice tone value and an initial voice speed value by processing the emotional intensity value through the parameter generation strategy;

[0198] perform boundary constraint processing on the initial body movement amplitude value, the initial action speed value, the initial voice tone value and the initial voice speed value to generate a body movement amplitude initial value, an action speed initial value, a voice tone initial value and a voice speed initial value;

[0199] combine the body movement amplitude initial value, the action speed initial value, the voice tone initial value and the voice speed initial value to generate an initial action parameter.

[0200] In an embodiment, the personalization and scene adaptation module 40 is specifically configured to:

[0201] extract an age feature and an extroversion degree feature of the target object as personalization features;

[0202] extract a type feature and an environment volume feature of the current scene as scene features;

[0203] generate a personalization adjustment coefficient through a personalization mapping relationship based on the age feature and the extroversion degree feature;

[0204] generate a scene adjustment coefficient through a scene mapping relationship based on the type feature and the environmental volume feature.

[0205] In an embodiment, the action parameter optimization module 50 is specifically configured to:

[0206] obtain an initial value of a body movement amplitude, an initial value of a movement speed, an initial value of a voice tone, and an initial value of a voice speed in the initial action parameter;

[0207] multiply the individualized adjustment coefficient and the scene adjustment coefficient to generate a comprehensive adjustment coefficient;

[0208] multiply the comprehensive adjustment coefficient and the initial value of the body movement amplitude to generate a final value of the body movement amplitude;

[0209] multiply the comprehensive adjustment coefficient and the initial value of the movement speed to generate a final value of the movement speed;

[0210] multiply the comprehensive adjustment coefficient and the initial value of the voice tone to generate a final value of the voice tone;

[0211] multiply the comprehensive adjustment coefficient and the initial value of the voice speed to generate a final value of the voice speed;

[0212] combine the final value of the body movement amplitude, the final value of the movement speed, the final value of the voice tone, and the final value of the voice speed to generate a final action parameter.

[0213] In an embodiment, the emotional action execution and feedback learning module 60 is specifically configured to:

[0214] control a robot joint motor and a voice synthesizer to execute an emotional action according to the final action parameter;

[0215] collect facial expression response signals and skin conductance response signals of a target object through a biosensor;

[0216] generate an emotional receptivity score based on the facial expression response signals;

[0217] generate an emotional resonance score based on the skin conductance response signals;

[0218] weight and fuse the emotional receptivity score and the emotional resonance score to generate a comprehensive feedback score;

[0219] when the comprehensive feedback score is lower than a preset score threshold, adjust a coefficient in the action parameter mapping relationship.

[0220] In one embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as follows:Figure 4 As shown in the figure. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide determination and control capability. The memory of the computer device includes non-volatile and / or volatile storage medium, internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with the external user terminal through the network connection. The computer program is executed by the processor to realize the functions or steps of the server side of the action generation method based on multi-modal emotion perception.

[0221] In one embodiment, a computer device is provided, which can be a user terminal, and its internal structure diagram can be as shown in the figure. Figure 5 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide determination and control capability. The memory of the computer device includes non-volatile storage medium, internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with the external server through the network connection. The computer program is executed by the processor to realize the functions or steps of the user terminal side of the action generation method based on multi-modal emotion perception.

[0222] In one embodiment, a computer device is provided, including a memory, a processor and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to realize the following steps:

[0223] Obtain multi-modal data of a target object, and assign corresponding weights to the multi-modal data based on the emotion recognition degree of each modal data in the multi-modal data, and generate a fused emotion feature by fusing the multi-modal data based on the corresponding weights;

[0224] Based on the fused emotion feature, determine an emotion state containing an emotion type, an emotion intensity value and an emotion intensity change rate;

[0225] Based on the emotion state and a preset action parameter mapping relationship, generate an initial action parameter;

[0226] Obtain individualized features of the target object and scene features of the current scene, generate an individualized adjustment coefficient based on the individualized features, and generate a scene adjustment coefficient based on the scene features;

[0227] processing the initial action parameter through the individualized adjustment coefficient and the scene adjustment coefficient, to generate a final action parameter;

[0228] performing an emotional action according to the final action parameter, and obtaining feedback data of the target object to the emotional action, and updating the action parameter mapping relationship based on the feedback data.

[0229] In one embodiment, a computer readable storage medium is provided, and the computer program is stored on the computer readable storage medium, and the computer program is executed by a processor to implement the following steps:

[0230] obtaining a plurality of modal data of a target object, and assigning a corresponding weight to each modal data in the plurality of modal data based on the emotion recognition degree of the modal data, and generating a fused emotional feature by fusing the plurality of modal data based on the corresponding weight;

[0231] determining an emotional state including an emotional type, an emotional intensity value and an emotional intensity change rate based on the fused emotional feature;

[0232] generating an initial action parameter based on the emotional state and a preset action parameter mapping relationship;

[0233] obtaining individualized features of the target object and scene features of a current scene, generating an individualized adjustment coefficient based on the individualized features, and generating a scene adjustment coefficient based on the scene features;

[0234] processing the initial action parameter through the individualized adjustment coefficient and the scene adjustment coefficient, to generate a final action parameter;

[0235] performing an emotional action according to the final action parameter, and obtaining feedback data of the target object to the emotional action, and updating the action parameter mapping relationship based on the feedback data.

[0236] It should be noted that the functions or steps of the above computer readable storage medium or computer device can be referred to the related description of the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0237] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0238] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0239] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for action generation based on multimodal emotion perception, characterized in that, Includes the following steps: The system acquires multimodal data of the target object, assigns corresponding weights to the multimodal data based on the sentiment recognition of each modal data, and fuses the multimodal data based on the corresponding weights to generate fused sentiment features. Based on the fused emotional features, an emotional state including emotional type, emotional intensity value, and emotional intensity change rate is determined; Based on the mapping relationship between the emotional state and the preset action parameters, initial action parameters are generated; The personalized features of the target object and the scene features of the current scene are obtained. A personalized adjustment coefficient is generated based on the personalized features, and a scene adjustment coefficient is generated based on the scene features. The initial action parameters are processed using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters; The emotional action is executed according to the final action parameters, and the feedback data of the target object to the emotional action is obtained. Based on the feedback data, the action parameter mapping relationship is updated. Based on the fused emotional features, an emotional state is determined, including emotional type, emotional intensity value, and emotional intensity change rate, including: The fused emotional features are subjected to probability classification to determine the emotional type; The fused emotional features are subjected to a nonlinear transformation to generate the emotional intensity value. Select the corresponding dynamic analysis strategy based on the emotion type; Obtain the first emotional intensity value at the first moment and the second emotional intensity value at the second moment; The difference in emotional intensity between the second emotional intensity value and the first emotional intensity value is determined based on the dynamic analysis strategy. Based on the time interval between the second moment and the first moment and the difference in emotional intensity, an emotional intensity change rate is generated; An emotional state is generated based on the emotional type, emotional intensity value, and emotional intensity change rate.

2. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, Acquire multi-modal data of the target object, assign corresponding weights to the multi-modal data based on the sentiment recognition degree of each modality, and fuse the multi-modal data based on the corresponding weights to generate fused sentiment features, including: Collect facial expression data of the target object, and extract the frequency of eye movement and the rate of change of the corner of the mouth from the facial expression data; Collect speech and intonation data of the target object, and extract the speech fundamental frequency standard deviation and speech rate change rate from the speech and intonation data; Collect limb movement data of the target object, and extract the limb swing amplitude and movement frequency from the limb movement data; Collect contact sensing data of the target object, and extract the contact force and contact duration from the contact sensing data; The confidence classification module connected to the biosignal sensor determines the emotional recognition of the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, limb swing amplitude, movement frequency, contact force and contact duration. Based on the emotional recognition score, assign corresponding weights to the eye movement frequency, mouth corner change rate, speech fundamental frequency standard deviation, speech rate change rate, body sway amplitude, movement frequency, contact force, and contact duration. Based on the corresponding weights, the frequency of eye movements, the rate of change of the corner of the mouth, the standard deviation of the fundamental frequency of speech, the rate of change of speech rate, the amplitude of body swaying, the frequency of movement, the contact force and the contact duration are fused to generate fused emotional features.

3. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, Based on the mapping relationship between the emotional state and preset action parameters, initial action parameters are generated, including: Obtain the emotion type and emotion intensity value in the emotional state; From the preset action parameter mapping relationship, select the corresponding parameter generation strategy according to the emotion type; The emotional intensity value is processed by the parameter generation strategy to generate initial body movement amplitude value, initial movement speed value, initial voice pitch value and initial voice speed value. Boundary constraint processing is performed on the initial limb movement amplitude value, initial movement speed value, initial voice pitch value, and initial voice speed value to generate initial values ​​for limb movement amplitude, initial movement speed, initial voice pitch, and initial voice speed. The initial motion parameters are generated by combining the initial values ​​of the limb movement amplitude, the initial value of the movement speed, the initial value of the voice pitch and the initial value of the voice speed.

4. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, The process involves acquiring the personalized features of the target object and the scene features of the current scene, generating personalized adjustment coefficients based on the personalized features, and generating scene adjustment coefficients based on the scene features, including: The age and extroversion characteristics of the target object are extracted as personalized features; Extract the current scene's type features and ambient volume features as scene features; Based on the age characteristics and personality extroversion characteristics, a personalized adjustment coefficient is generated through a personalized mapping relationship; Based on the aforementioned type characteristics and environmental volume characteristics, scene adjustment coefficients are generated through scene mapping relationships.

5. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, The initial action parameters are processed using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final action parameters, including: Obtain the initial values ​​of limb movement amplitude, movement speed, voice pitch, and voice speed from the initial action parameters; Multiply the personalized adjustment coefficient by the scene adjustment coefficient to generate a comprehensive adjustment coefficient; Multiply the comprehensive adjustment coefficient by the initial value of the limb movement amplitude to generate the final limb movement amplitude value; Multiply the overall adjustment coefficient by the initial value of the motion speed to generate the final motion speed value; Multiply the comprehensive adjustment coefficient by the initial value of the voice pitch to generate the final voice pitch value; The final speech rate value is generated by multiplying the comprehensive adjustment coefficient by the initial speech rate value. The final limb movement amplitude value, final movement speed value, final voice pitch value, and final voice speed value are combined to generate the final movement parameters.

6. The action generation method based on multimodal emotion perception as described in claim 1, characterized in that, Execute an emotional action based on the final action parameters, obtain feedback data from the target object on the emotional action, and update the action parameter mapping relationship based on the feedback data, including: The robot's joint motors and speech synthesizer are controlled to perform emotional actions based on the final motion parameters. Facial expression response signals and skin conductance response signals of the target object are collected using biosensors; An emotion receptivity score is generated based on the facial expression response signal; An emotional resonance score is generated based on the skin conductance response signal. The emotional receptivity score and emotional resonance score are weighted and fused together to generate a comprehensive feedback score; When the overall feedback score is lower than a preset score threshold, the coefficients in the action parameter mapping relationship are adjusted.

7. A motion generation device based on multimodal emotion perception, characterized in that, The action generation device based on multimodal emotion perception includes: The multimodal perception and fusion module is used to acquire multimodal data of the target object, assign corresponding weights to the multimodal data based on the sentiment recognition degree of each modal data, and fuse the multimodal data based on the corresponding weights to generate fused sentiment features; The emotional state analysis module is used to determine the emotional state, including emotional type, emotional intensity value, and emotional intensity change rate, based on the fused emotional features. The action parameter mapping module is used to generate initial action parameters based on the emotional state and the preset action parameter mapping relationship; The personalization and scene adaptation module is used to obtain the personalized features of the target object and the scene features of the current scene, generate a personalization adjustment coefficient based on the personalized features, and generate a scene adjustment coefficient based on the scene features. The motion parameter optimization module is used to process the initial motion parameters using the personalized adjustment coefficient and the scene adjustment coefficient to generate the final motion parameters. The emotional action execution and feedback learning module is used to execute emotional actions according to the final action parameters, obtain feedback data of the target object on the emotional actions, and update the action parameter mapping relationship based on the feedback data; Based on the fused emotional features, an emotional state is determined, including emotional type, emotional intensity value, and emotional intensity change rate, including: The fused emotional features are subjected to probability classification to determine the emotional type; The fused emotional features are subjected to a nonlinear transformation to generate the emotional intensity value. Select the corresponding dynamic analysis strategy based on the emotion type; Obtain the first emotional intensity value at the first moment and the second emotional intensity value at the second moment; The difference in emotional intensity between the second emotional intensity value and the first emotional intensity value is determined based on the dynamic analysis strategy. Based on the time interval between the second moment and the first moment and the difference in emotional intensity, an emotional intensity change rate is generated; An emotional state is generated based on the emotional type, emotional intensity value, and emotional intensity change rate.

8. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal emotion-aware action generation program stored in the memory and executable on the processor. When the multimodal emotion-aware action generation program is executed by the processor, it implements the steps of the multimodal emotion-aware action generation method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The storage medium stores an action generation program based on multimodal emotion perception, which, when executed by a processor, implements the steps of the action generation method based on multimodal emotion perception as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Intelligent accompanying human-type robot based on multi-modal emotion interaction and sensing method

    CN119839861A

  • Robot behavior mode dynamic adjustment method based on multi-mode perception

    CN120257050A

  • Dynamic self-adaptive multi-modal sentiment analysis fusion method and system

    CN120429691A