A multimodal perception-based large model emotion analysis method, device and medium
Patent Information
- Application Number
- CN202610959527.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-18
AI Technical Summary
然而,单一模态信息往往存在局限性:面部表情可被有意掩饰,语音易受环境噪声干扰,难以准确反映用户的真实情绪状态
[0013] This invention provides a large-scale emotion analysis method based on multimodal perception. First, by embedding capacitive and resistive sensor arrays within a flexible electronic skin substrate, raw capacitance and resistance signals generated during interaction are collected, respectively, to analyze the user's tactile behavior data. Through dual-modal collaborative acquisition and analysis of capacitance and resistance, transient and steady-state touch information can be captured, achieving refined perception of multiple types of tactile behaviors and providing accurate and rich tactile feature data for subsequent emotion analysis. Then, the tactile behavior data is aligned with the collected visual data to construct a multimodal emotion feature vector, thereby enhancing the tactile behavior data. This provides high-quality aligned input for large-scale emotion analysis. Multimodal emotion feature vectors, along with pre-acquired interaction history features and scene context encoding, are input into a pre-defined large-scale model for emotion analysis, achieving deep semantic association between tactile behavior and visual expressions. By fusing interaction history and scene context, the large-scale model accurately infers the user's true emotional state. Finally, based on the identified emotional state expressions, corresponding emotion description text is generated to control the robot to execute corresponding emotional interaction feedback strategies. This invention, by fusing tactile and visual multimodal perception information and utilizing a large-scale model for deep emotion reasoning, enables accurate identification and intelligent interactive feedback of the user's true emotional state.
Smart Images

Figure CN122776983A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent robot technology, and in particular to a method, device and medium for large-scale emotion analysis based on multimodal perception. Background Technology
[0002] With the rapid development of human-computer interaction technology, emotional interaction has become an important research direction in the field of intelligent robotics. Traditional emotion recognition methods mainly rely on single-modal information, such as computer vision recognition based on facial expressions or emotion analysis based on speech. However, single-modal information often has limitations: facial expressions can be intentionally masked, speech is easily interfered with by environmental noise, and it is difficult to accurately reflect the user's true emotional state.
[0003] Touch, as an important channel for human emotional expression, contains rich emotional information during physical interactions with robots. For example, stroking usually expresses friendliness and comfort, while tapping may convey dissatisfaction or anger. Although existing robotic electronic skin technology can sense simple physical parameters such as pressure and deformation, it lacks the ability to deeply understand the semantics of tactile behavior and perform cross-modal correlation analysis, making it difficult to accurately identify the user's emotional state in complex emotional scenarios. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a method, device, and medium for large-scale emotion analysis based on multimodal perception. By fusing tactile and visual multimodal perception information and utilizing a large model for deep emotion reasoning, it is possible to accurately identify the user's true emotional state.
[0005] According to a first aspect of the present invention, a large-scale sentiment analysis method based on multimodal perception is provided, comprising the following steps:
[0006] S1 collects the original capacitance and resistance signals generated by the user during the interaction process through the capacitance sensor array and resistance sensor array embedded in the flexible electronic skin matrix of the robot, and analyzes the tactile behavior data of the user during the interaction process with the robot based on the original capacitance and resistance signals.
[0007] S2, synchronously collects the user's visual data through a camera pre-installed on the robot body, and performs feature alignment and temporal synchronization processing on the parsed tactile behavior data and visual data to construct a multimodal emotion feature vector; the visual data includes facial expressions, body postures and dynamic gesture data.
[0008] S3, input the multimodal emotion feature vector, the pre-acquired interaction history features and the scene context encoding into the preset large model for emotion analysis, and infer the user's current emotion state label and the corresponding label confidence.
[0009] S4. Based on the inferred emotional state labels and corresponding label confidence levels, a structured emotional description text is generated to control the robot to execute the corresponding emotional interaction feedback strategy.
[0010] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the above-described large-model sentiment analysis method based on multimodal perception.
[0011] According to a third aspect of the present invention, an electronic device is provided, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0012] The present invention has at least the following beneficial effects:
[0013] This invention provides a large-scale emotion analysis method based on multimodal perception. First, by embedding capacitive and resistive sensor arrays within a flexible electronic skin substrate, raw capacitance and resistance signals generated during interaction are collected, respectively, to analyze the user's tactile behavior data. Through dual-modal collaborative acquisition and analysis of capacitance and resistance, transient and steady-state touch information can be captured, achieving refined perception of multiple types of tactile behaviors and providing accurate and rich tactile feature data for subsequent emotion analysis. Then, the tactile behavior data is aligned with the collected visual data to construct a multimodal emotion feature vector, thereby enhancing the tactile behavior data. This provides high-quality aligned input for large-scale emotion analysis. Multimodal emotion feature vectors, along with pre-acquired interaction history features and scene context encoding, are input into a pre-defined large-scale model for emotion analysis, achieving deep semantic association between tactile behavior and visual expressions. By fusing interaction history and scene context, the large-scale model accurately infers the user's true emotional state. Finally, based on the identified emotional state expressions, corresponding emotion description text is generated to control the robot to execute corresponding emotional interaction feedback strategies. This invention, by fusing tactile and visual multimodal perception information and utilizing a large-scale model for deep emotion reasoning, enables accurate identification and intelligent interactive feedback of the user's true emotional state. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1The flowchart illustrates a large-scale emotion analysis method based on multimodal perception provided in this embodiment of the invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] This invention provides a large-scale sentiment analysis method based on multimodal perception, such as... Figure 1 As shown, the method includes the following steps:
[0018] S1 collects the original capacitance and resistance signals generated by the user during the interaction process through the capacitance sensor array and resistance sensor array embedded in the flexible electronic skin matrix of the robot, and analyzes the tactile behavior data of the user during the interaction process with the robot based on the original capacitance and resistance signals.
[0019] Specifically, the capacitive sensor array adopts a cross-electrode structure, consisting of upper and lower flexible metal electrodes and a middle dielectric layer. The upper and lower flexible metal electrodes can be silver nanowires or liquid metal, and the middle dielectric layer can be made of silicone composite material. That is, each capacitive sensor uses a variable electrode spacing principle, sensitive to instantaneous, short-duration, high-frequency touch behaviors such as tapping, clicking, and slapping. When such touch behaviors occur, the electrode spacing changes rapidly, causing a sudden change in capacitance. The capacitance value of each capacitive sensor can be collected using a high-speed capacitance detection chip, such as the FDC2214 chip, with a sampling rate set to no less than 100Hz.
[0020] Furthermore, the resistive sensor array employs a piezoresistive composite material to form a mesh structure; the piezoresistive composite material can be PDMS material doped with graphene. Each resistive sensor is a variable resistor; when subjected to continuous action such as touching, pressing, or rubbing, the resistance value changes. The voltage signal can be acquired through a low-power analog-to-digital converter, and the steady-state resistance value can be extracted, with a corresponding sampling rate range of 10-50Hz.
[0021] Specifically, the process of parsing tactile behavior data during user-robot interaction based on the original capacitance and resistance signals includes the following steps:
[0022] S101: Based on the pre-acquired reference capacitance values of each capacitance sensor in the capacitance sensor array and the reference resistance values of each resistance sensor in the resistance sensor array, combined with the acquired original capacitance and resistance signals, the effective change in capacitance and the effective change in resistance are calculated. That is, the difference between the original capacitance signal and the reference capacitance value is the effective change in capacitance; the difference between the original resistance signal and the reference resistance value is the effective change in resistance.
[0023] Specifically, the reference capacitance and reference resistance values are data collected periodically when the robot is unloaded.
[0024] S102, based on the pre-calibrated temperature coefficients of each capacitance sensor and each resistance sensor, and the current temperature of the flexible electronic skin substrate, performs temperature drift compensation calculations on the effective change in capacitance and the effective change in resistance, respectively, to obtain the compensated capacitance signal and the compensated resistance signal.
[0025] Specifically, the formula for calculating the compensated capacitance signal C is as follows:
[0026] C=C0-ρ C ×(ρ-ρ0), where C0 is the effective change in capacitance, ρ C ρ is the temperature coefficient of the pre-calibrated capacitive sensor, ρ is the current temperature of the flexible electronic skin substrate, and ρ0 is the preset reference temperature, such as 25℃.
[0027] Furthermore, the formula for calculating the compensated resistance signal R is as follows:
[0028] R=R0-ρ R ×(ρ-ρ0), where R0 is the effective change in resistance, ρ R The temperature coefficient of the pre-calibrated resistance sensor.
[0029] S103 performs transient event detection based on the compensated capacitance signal and steady-state event detection based on the compensated resistance signal, extracts corresponding preset index data for the detected events, and records the data extraction results.
[0030] Specifically, step S103 includes the following steps:
[0031] S1031, calculate the rate of change of the capacitance signal based on the compensated capacitance signal. When the rate of change of the capacitance signal exceeds a preset transient rate of change threshold, determine that the current tactile behavior is a transient event, and extract the trigger time, trigger position, and impact intensity level corresponding to the current transient event. Those skilled in the art can set the preset transient rate of change threshold according to actual needs, or adaptively update it based on the noise standard deviation σ and baseline slow drift μ under the robot's no-load state. Specifically, the preset transient rate of change threshold θ = 3 × σ + μ.
[0032] Specifically, the trigger time refers to the moment when the capacitance change is detected to be greater than the preset rate of change threshold; the trigger position is the centroid position calculated from the coordinates of the triggered capacitance sensor.
[0033] Furthermore, the impact strength level is positively correlated with the maximum rate of change of the capacitance signal.
[0034] S1032, the resistance dispersion of the resistance signal is calculated based on the compensated resistance signal. When the resistance dispersion is lower than a preset steady-state threshold and its duration exceeds a preset minimum steady-state time, the current tactile behavior is determined to be a steady-state event, and the mean resistance and resistance relaxation rate corresponding to the current steady-state event are extracted. The resistance dispersion is obtained by calculating the resistance variance. Those skilled in the art can set the preset steady-state threshold and preset minimum steady-state time according to actual needs.
[0035] Specifically, the rate of change of resistance relaxation is η, where η = 1 / γ × (R t -R0), where γ is a preset time constant, R t Let Rt be the resistance value at time t, and R0 be the steady-state resistance value.
[0036] Furthermore, if the conditions for determining both transient and steady-state events are met simultaneously within the same time period, the tactile behavior is split into transient and steady-state events and recorded separately in chronological order; if neither condition is met, the current tactile behavior is determined to be an invalid event, and the corresponding signal data is discarded.
[0037] By detecting the differences in the rate of change of capacitance signal and the degree of resistance dispersion, transient impact events and steady-state continuous events are accurately identified respectively. This enables the subdivision of tactile behavior types, such as knocking and tapping, and long-term behaviors such as stroking and pressing, providing rich tactile feature inputs for subsequent multimodal emotion analysis.
[0038] S104, parse the preset behavior data corresponding to each event from the data extraction results to form tactile behavior data; the preset behavior data includes touch position, pressure amplitude, duration of action, touch frequency, spatial distribution characteristics, and touch type. Among them, the pressure amplitude is obtained by mapping the normalized resistance change to a pressure level from 0 to 100%; the duration of action is the start and end time difference of the original capacitance signal exceeding the corresponding preset threshold; the spatial distribution characteristics refer to the touch heat map matrix constructed based on the changes in the original capacitance signal; and the touch type is a tactile behavior type such as tapping, tapping, or pressing.
[0039] As described above, through the dual-modal collaborative acquisition and analysis of capacitive and resistive sensor arrays, transient impact characteristics and steady-state continuous characteristics in user touch behavior are captured respectively. In this process, environmental interference is eliminated through baseline calibration and temperature drift compensation, ensuring the authenticity and stability of tactile signals. This enables refined perception of various types of tactile behaviors such as tapping, touching, stroking, and pressing, providing accurate and rich tactile feature data for subsequent emotion analysis.
[0040] S2 synchronously collects the user's visual data through a camera pre-installed on the robot body, and performs feature alignment and temporal synchronization processing on the parsed tactile behavior data and visual data to construct a multimodal emotion feature vector.
[0041] Specifically, the visual data includes, but is not limited to, facial expressions, body postures, dynamic gesture data, relative motion distances, motion speeds, and environmental context. The camera's visual acquisition module uses a miniature RGB camera with a visual acquisition frame rate set to 15-30fps. A lightweight pose detection network extracts the coordinates of several facial key points, the relative positions of limb key points, and the dynamic trajectories corresponding to hand key points. The lightweight pose detection network can adopt an OpenPose architecture.
[0042] Specifically, the multimodal emotion feature vector is constructed through the following steps:
[0043] S201 assigns a unified global timestamp to tactile behavior data and visual data to align the time references of tactile behavior data and visual data.
[0044] S202, using the visual acquisition frame rate when collecting user visual data as a reference, performs linear interpolation on the tactile behavior data to obtain tactile behavior data aligned with the visual data; this can be understood as: using cubic spline interpolation to resample the tactile behavior data in the temporal domain, keeping the temporal sampling rate consistent with the visual acquisition frame rate; those skilled in the art know the specific implementation steps of the cubic spline interpolation method, and will not elaborate further here.
[0045] S203 uses a pre-calibrated coordinate system transformation matrix to project and transform the touch position coordinates in the tactile behavior data, mapping the touch position coordinates to the camera image coordinate system, thus achieving spatial alignment between the touch position and the camera image coordinate system.
[0046] Specifically, the formula for the projection transformation is:
[0047] U = K·(H·J+G), where U is the projected coordinate when the touch position coordinates are mapped to the camera image coordinate system, K is the camera intrinsic parameter matrix, H is the rotation matrix from the skin coordinate system to the camera image coordinate system, J is the touch position coordinates, and G is the three-dimensional offset of the origin of the skin coordinate system relative to the origin of the camera image coordinate system.
[0048] S204. Tactile behavior feature vectors and visual feature vectors are extracted from the aligned tactile behavior data and visual data respectively, and the vectors are concatenated to obtain multimodal emotion feature vectors.
[0049] Specifically, tactile behavior feature vectors refer to converting tactile behavior data into numerical vector form, such as encoding touch position into coordinate form, normalizing pressure amplitude into floating-point numbers between 0 and 1, and mapping touch type into one-hot encoding, etc.
[0050] Furthermore, the visual feature vector includes facial action unit intensity calculated based on facial keypoint coordinates, limb posture features obtained based on limb keypoint positions, and dynamic gesture embedding features generated by encoding a two-dimensional coordinate sequence of hand keypoints. Here, facial action unit intensity refers to the value obtained by normalizing the Euclidean distance between the current keypoint coordinates and the reference coordinates of the same keypoint during a neutral expression for each keypoint.
[0051] The above-mentioned triple calibration of timestamp alignment, interpolation processing and projection transformation achieves accurate alignment of tactile behavior data and visual data in the temporal and spatial dimensions, eliminates the temporal deviation and feature misalignment problems between heterogeneous data, and constructs a multimodal emotion feature vector with complementary and enhanced information and unified dimensions, providing high-quality aligned input for large-scale model emotion analysis, thereby reducing the emotion misjudgment rate caused by the lack of single modality information.
[0052] S3, input the multimodal emotion feature vector, the pre-acquired interaction history features and the scene context encoding into the preset large model for emotion analysis, and infer the user's current emotion state label and the corresponding label confidence.
[0053] Specifically, step S3 includes the following steps:
[0054] S301: The multimodal emotion feature vectors are organized into an emotion feature sequence according to the order of events corresponding to the tactile behavior data. The emotion feature sequence, interaction history features, and scene context encoding are then input into a pre-defined large model. This can be understood as: splitting the multimodal emotion feature vectors into emotion features corresponding to each event in chronological order to obtain the emotion feature sequence. For example, the events are converted into structured text, such as: tapping the hand with 80% force for 0.5 seconds, stroking the back with 30% force for 2.3 seconds.
[0055] Specifically, the data corresponding to the interaction history features are the data from the most recent rounds of interaction. Each round of interaction includes, but is not limited to, user behavior, multimodal emotion features, robot response, and probability distribution of emotion output, and is encoded as temporal interaction history features.
[0056] Furthermore, the data corresponding to the scene context encoding includes the robot's operating status, environmental information, and user identification, all uniformly encoded into a fixed-length context vector. The operating status includes task mode, battery level, temperature, etc.; the environmental information includes lighting, noise, scene type, etc.
[0057] S302, based on a pre-defined large model, extracts several tactile behavioral feature sub-vectors and several visual feature sub-vectors from the emotion feature sequence, forming tactile vector sequences and visual vector sequences respectively. Then, using a cross-attention mechanism, the tactile vector sequence is used as the query, and the visual vector sequence as the key and value for cross-modal fusion, generating a tactile enhancement vector sequence. In other words, after inputting the multimodal emotion feature vectors into the pre-defined large model, it is parsed into tactile and visual features, and the cross-attention mechanism performs intermodal interaction calculations to form a vector sequence that enhances tactile sensation through vision.
[0058] Specifically, the pre-defined large model can adopt a multimodal large model based on GPT-4V structure or Video-LLaMA. The core computation includes modality encoding, cross-attention fusion, and self-attention decoding operations. In the encoding stage, tactile behavioral features in the emotion feature sequence are mapped to a tactile vector sequence through a Transformer lightweight temporal encoder, and visual features are mapped to a visual vector sequence through a visual encoder, such as a ViT or ResNet encoder. Then, the fusion of tactile and visual modal information is completed through the internal cross-attention mechanism of the pre-defined large model. In the decoding stage, the fused features are combined with interaction history features, scene context encoding, and pre-defined prompt words, and input into the pre-defined multi-layer Transformer decoder.
[0059] S303 concatenates the fused haptic enhancement vector sequence with interaction history features, scene context encoding, and preset prompt word templates, and inputs it into several layers of decoders within a preset large model. These layers are then mapped to preset emotion state category spaces through fully connected layers, outputting the final emotion state label and corresponding label confidence. This can be understood as: inputting the hidden state vector containing semantically condensed representations output by the last decoder layer into a fully connected layer, and outputting the emotion state label through a classifier.
[0060] Specifically, the preset emotional state category space includes emotional state labels such as happy, joyful, angry, anxious, tense, relaxed, coquettish, dependent, lonely, seeking comfort, false politeness, and tentative interaction.
[0061] As described above, by constructing multimodal emotion feature vectors into a temporal sequence and completing cross-modal fusion through a cross-attention mechanism, a deep semantic association between tactile behavior and visual expression is achieved. Furthermore, by integrating interaction history and scene context, and utilizing the cross-modal understanding and reasoning capabilities of a large model, the user's true emotional state can be accurately identified. This effectively solves the problem of emotional ambiguity when single-modal information is insufficient or contradictory, and achieves reliable multi-dimensional emotion perception.
[0062] S4 generates structured emotion description text based on the inferred emotion state labels and their corresponding label confidence scores to control the robot to execute corresponding emotion interaction feedback strategies. This can be understood as follows: the emotion interaction feedback strategy includes preset interaction feedback actions for each emotion state label, such as LED light strip color synchronization or automatic adjustment of the robot's voice tone. In specific implementations, the default output is the emotion state label with the highest label confidence score and its corresponding label confidence score.
[0063] Specifically, structured emotion description text refers to the explanatory text for the output emotion state labels. For example, if a user taps the robot's head while frowning and closing their lips, and this is repeated multiple times within 30 seconds, it can be inferred that the user is angry with 87% confidence. In practice, a greedy search algorithm can be used to generate natural language emotion description text.
[0064] As described above, structured emotional description text is automatically generated based on the results of emotion analysis, and the robot is driven to execute differentiated interactive feedback strategies to achieve a closed loop from emotion perception to behavioral feedback, thereby improving the naturalness of human-computer interaction and adapting to users' personalized emotional needs.
[0065] In one specific embodiment, step S4 further includes the following steps:
[0066] S401, if the confidence level of the label corresponding to the emotional state label is lower than a preset confidence threshold, then an uncertain emotional state determination result is output, and the robot is controlled to execute an interactive feedback strategy of questioning or re-collecting data. Those skilled in the art can set the preset confidence threshold according to actual needs, for example, 0.6.
[0067] S402, if the confidence level of the label corresponding to the emotional state label is not lower than the preset confidence threshold, and the emotional state label is a preset complex emotion category, the preset complex emotion category is verified.
[0068] Furthermore, when the emotional state label is a preset complex emotion category, the mutual information between facial expression valence and tactile arousal is calculated. When the mutual information is lower than a first preset threshold, the preset complex emotion category is verified. Those skilled in the art can set the first preset threshold according to actual needs, for example, to 0.2.
[0069] Specifically, the formula for calculating the valence V of facial expressions is: V = sigmoid(w v T ×f1+b v ), where w v To map the hidden states input to a fully connected layer to a weight matrix along the valence dimension, T denotes the transpose, and b v is the learnable bias term for the branch corresponding to the valence dimension. f1 is the facial feature vector extracted from the visual feature vector from the input to the fully connected layer, including facial landmarks and facial action unit intensity values. The value range of facial expression valence is mapped from the 0 to 1 interval corresponding to the sigmoid activation function to the -1 to 1 interval, with negative values representing negative emotions and positive values representing positive emotions.
[0070] Specifically, the formula for calculating tactile arousal A is: A = sigmoid(w a T ×f2+b a ), where w a To map the hidden states input to the fully connected layer to a weight matrix representing the arousal dimension, f2 is the tactile behavior feature vector input to the fully connected layer, and b... a This is the learnable bias term for the branch corresponding to the arousal dimension. The value range of haptic arousal is mapped from the 0 to 1 interval corresponding to the sigmoid activation function to the -1 to 1 interval, with negative values representing calm touch and positive values representing excited touch.
[0071] Furthermore, the formula for calculating the mutual information (MI) between facial expression valence and tactile arousal is as follows:
[0072] Where P(x, y) is the joint probability of facial expression valence x and tactile arousal y, P(x) is the probability of facial expression valence x, and P(y) is the probability of tactile arousal y. x and y are values in the corresponding sets obtained by discretizing facial expression valence and tactile arousal respectively. P(x, y), P(x), and P(y) are obtained from pre-collected sample data.
[0073] If the preset complex emotion category is false politeness, then the label confidence level corresponding to false politeness is recalculated based on the preset confidence level discrimination formula. When the label confidence level corresponding to false politeness is greater than the preset confidence level threshold, the current emotion state label is maintained; otherwise, the current emotion state label is corrected to "emotion uncertain" and then output.
[0074] Specifically, the preset confidence level discrimination formula is as follows:
[0075] F = sigmoid(α×MI+β×D), where F is a preset confidence discriminant formula, D is the degree of tactile-visual contradiction, and α and β are preset weighting coefficients. The degree of tactile-visual contradiction is obtained by dividing the absolute value of the difference between facial expression valence and tactile emotional polarity by 2, used to measure whether the emotional direction of tactile behavior is consistent with that of facial expression. Tactile emotional polarity is mapped according to touch type; for example, stroking corresponds to a tactile emotional polarity of 0.7, patting to 0.3, grasping to -0.5, and tapping to -0.8.
[0076] If the preset complex emotion category is tentative interaction, then when the duration of the tactile action is less than the second preset threshold, the rate of change of pressure amplitude exceeds the third preset threshold, and there is no subsequent touch action, and eye movement speed is detected to exceed the fourth preset threshold, the current emotion state label is maintained; otherwise, the current emotion state label is corrected to "emotion uncertain" before being output. The eye movement speed is obtained by removing pupil positions between adjacent frames at time intervals. Those skilled in the art can set the second, third, and fourth preset thresholds according to actual needs; for example, the second preset threshold is 0.5 seconds.
[0077] The above-mentioned dual judgment mechanism of confidence threshold and complex emotion category directly reflects the uncertainty of emotion for low confidence output to avoid misjudgment, and performs secondary verification for easily confused emotions such as false politeness and tentative interaction under high confidence, which effectively improves the accuracy of emotion recognition in complex social scenarios and realizes multimodal emotion contradiction detection.
[0078] Furthermore, the method further includes the following steps:
[0079] S10 connects to a voice acquisition module, a heart rate sensor, and a temperature sensor via a pre-installed universal expansion interface on the robot body to collect the user's voice audio, heart rate signal, and skin temperature data, respectively. For example, a microphone is used to collect voice audio via an I2S interface, a heart rate sensor is connected via an I2S interface to obtain heart rate signals, and a temperature sensor is connected via an I2C interface to collect skin temperature.
[0080] S20, based on the user's voice audio, heart rate signal, and skin temperature data, obtains voice emotion embedding features, heart rate value, heart rate variability features, and temperature change rate features, and then performs standardization processing on each feature before concatenating them into a joint feature vector. This can be understood as the standardization process unifying each feature dimension to a preset length. Specifically, when any modality data is missing, the corresponding feature position is filled with a zero vector.
[0081] S30, the joint feature vector is fused with the multimodal emotion feature vector to form an enhanced multimodal emotion feature vector, which is then used to replace the original multimodal emotion feature vector and input into a preset large model for emotion analysis.
[0082] As mentioned above, by reserving a general-purpose extension interface, the fusion of tactile, visual, and physiological multimodal signals can be achieved. This allows for adaptation to the perception needs of different scenarios without modifying the core architecture, further improving the reliability and robustness of recognizing complex emotional states.
[0083] Embodiments of the present invention also provide a non-transitory computer-readable storage medium that can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a method in the method embodiments, wherein the at least one instruction or the at least one program is loaded and executed by the processor to implement the method provided in the above embodiments.
[0084] Embodiments of the present invention also provide an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.
[0085] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. It should also be understood that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.
Claims
1. A large-scale sentiment analysis method based on multimodal perception, characterized in that, The method includes the following steps: S1, through the capacitive sensor array and resistive sensor array embedded in the flexible electronic skin matrix of the robot, collects the original capacitance signal and original resistance signal generated by the user during the interaction process, and analyzes the tactile behavior data of the user during the interaction process with the robot based on the original capacitance signal and original resistance signal. S2, synchronously collects the user's visual data through a camera pre-installed on the robot body, and performs feature alignment and temporal synchronization processing on the parsed tactile behavior data and visual data to construct a multimodal emotion feature vector; the visual data includes facial expressions, body postures and dynamic gesture data; S3, input the multimodal emotion feature vector, the pre-acquired interaction history features and the scene context encoding into the preset large model for emotion analysis, and infer the user's current emotion state label and the corresponding label confidence; S4. Based on the inferred emotional state labels and corresponding label confidence levels, a structured emotional description text is generated to control the robot to execute the corresponding emotional interaction feedback strategy.
2. The large-scale sentiment analysis method based on multimodal perception according to claim 1, characterized in that, In step S1, the step of parsing the tactile behavior data during the interaction between the user and the robot based on the original capacitance signal and the original resistance signal includes the following steps: S101, based on the reference capacitance values of each capacitance sensor in the pre-acquired capacitance sensor array and the reference resistance values of each resistance sensor in the resistance sensor array, combined with the acquired original capacitance signal and original resistance signal, calculate the effective change in capacitance and the effective change in resistance. S102, based on the pre-calibrated temperature coefficients of each capacitance sensor and each resistance sensor, and the current temperature of the flexible electronic skin substrate, performs temperature drift compensation calculations on the effective change in capacitance and the effective change in resistance, respectively, to obtain the compensated capacitance signal and the compensated resistance signal. S103, performs transient event detection based on the compensated capacitance signal and steady-state event detection based on the compensated resistance signal, extracts corresponding preset index data for the detected events and records the data extraction results; S104, parse the preset behavior data corresponding to each event from the data extraction results to form tactile behavior data; the preset behavior data includes touch position, pressure amplitude, duration of action, touch frequency, spatial distribution characteristics and touch type.
3. The large-scale sentiment analysis method based on multimodal perception according to claim 2, characterized in that, Step S103 includes the following steps: S1031, calculate the rate of change of the capacitance signal based on the compensated capacitance signal. When the rate of change of the capacitance signal exceeds the preset transient rate of change threshold, determine that the current tactile behavior is a transient event, and extract the trigger time, trigger position and impact intensity level corresponding to the current transient event. S1032, calculate the resistance dispersion of the resistance signal based on the compensated resistance signal. When the resistance dispersion is lower than the preset steady-state threshold and the duration exceeds the preset minimum steady-state time, determine that the current tactile behavior belongs to a steady-state event, and extract the mean resistance and resistance relaxation rate corresponding to the current steady-state event.
4. The large-scale sentiment analysis method based on multimodal perception according to claim 1, characterized in that, In step S2, a multimodal emotion feature vector is constructed through the following steps: S2 01. Assign a unified global timestamp to tactile behavior data and visual data to align the time reference of tactile behavior data and visual data; S202, using the visual acquisition frame rate when acquiring user visual data as a reference, perform linear interpolation on the tactile behavior data to obtain tactile behavior data aligned with the visual data; S203 uses a pre-calibrated coordinate system transformation matrix to project and transform the touch position coordinates in the tactile behavior data, mapping the touch position coordinates to the camera image coordinate system, thereby achieving spatial alignment between the touch position and the camera image coordinate system. S204. Tactile behavior feature vectors and visual feature vectors are extracted from the aligned tactile behavior data and visual data respectively, and the vectors are concatenated to obtain multimodal emotion feature vectors.
5. The large-scale sentiment analysis method based on multimodal perception according to claim 1, characterized in that, Step S3 includes the following steps: S301, organize the multimodal emotion feature vectors into an emotion feature sequence according to the order of events corresponding to the tactile behavior data, and input the emotion feature sequence, interaction history features and scene context encoding into the preset large model; S302, based on a pre-set large model, extracts several tactile behavior feature sub-vectors and several visual feature sub-vectors from the emotion feature sequence, forming tactile vector sequences and visual vector sequences respectively, and uses a cross-attention mechanism to perform cross-modal fusion of the tactile vector sequence as a query and the visual vector sequence as a key and value to generate a tactile enhancement vector sequence; S303 concatenates the fused haptic enhancement vector sequence with interaction history features, scene context encoding, and preset prompt word templates, and inputs it into several layers of decoders within a preset large model. Through fully connected layers, these layers are mapped to preset emotion state category spaces, and the final emotion state label and corresponding label confidence are output.
6. The large-scale sentiment analysis method based on multimodal perception according to claim 1, characterized in that, Step S4 also includes the following steps: S401, if the confidence level of the label corresponding to the emotional state label is lower than the preset confidence threshold, then output the determination result of emotional uncertainty, and control the robot to execute the interactive feedback strategy of questioning or re-collection. S402, if the confidence level of the label corresponding to the emotion state label is not lower than a preset confidence threshold, and the emotion state label is a preset complex emotion category, then the preset complex emotion category is verified; wherein... If the preset complex emotion category is false politeness, then the label confidence level corresponding to false politeness is recalculated based on the preset confidence level discrimination formula. When the label confidence level corresponding to false politeness is greater than the preset confidence level threshold, the current emotion state label is maintained; otherwise, the current emotion state label is corrected to "emotion uncertain" and then output. If the preset complex emotion category is tentative interaction, then when the duration of the tactile behavior is less than the second preset threshold, the rate of change of pressure amplitude exceeds the third preset threshold and there is no subsequent touch action, and the eye movement speed is detected to exceed the fourth preset threshold, the current emotion state label is maintained; otherwise, the current emotion state label is corrected to "emotion uncertain" before being output.
7. The large-scale sentiment analysis method based on multimodal perception according to claim 1, characterized in that, The method further includes the following steps: S10 connects to a voice acquisition module, heart rate sensor, and temperature sensor via a universal expansion interface reserved on the robot body, and collects the user's voice audio, heart rate signal, and skin temperature data respectively. S20, based on the user's voice audio, heart rate signal and skin temperature data, obtains voice emotion embedding features, heart rate value, heart rate variability features and temperature change rate features, and then performs standardization processing on each feature before concatenating them into a joint feature vector; S30, the joint feature vector is fused with the multimodal emotion feature vector to form an enhanced multimodal emotion feature vector, which is then used to replace the original multimodal emotion feature vector and input into a preset large model for emotion analysis.
8. A non-transitory computer-readable storage medium, wherein the storage medium stores at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program segment is loaded and executed by the processor to implement the large model sentiment analysis method based on multimodal perception as described in any one of claims 1-7.
9. An electronic device, characterized in that, Includes a processor and the non-transitory computer-readable storage medium as described in claim 8.