Emotion analysis method, device and equipment based on body intelligence and medium
Through a multimodal analysis method that collects and fuses physiological signal data and behavioral performance data, the problem of low accuracy of emotion analysis under traditional single mode data is solved, and higher accuracy of emotion recognition is achieved.
Patent Information
- Application Number
- CN202510669658.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-26
AI Technical Summary
Traditional emotions analysis methods rely mostly on single modal data, resulting in low accuracy and difficulty in accurately identifying the complex emotional states of human beings.
Multimodal data of the target user, including physiological signal data and behavioral performance data, multimodal feature vectors are obtained through quantitative processing, and the preset model is input to identify the emotion category after fusion.
Through the fusion recognition of multimodal data, the accuracy of emotion recognition is improved, and users' emotions categories can be more accurately identified from multiple dimensions.
Smart Images

Figure CN120541780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biometric technology, and is applied to scenarios such as emotion recognition and psychological medicine, and in particular to an emotion analysis method, device, equipment and medium based on embodied intelligence. Background Art
[0002] Sentiment analysis is the process of identifying, understanding, and quantifying human emotions through technical means. It is a core research area of affective computing. Traditional sentiment analysis methods rely on single-modal data, such as facial expressions or text. Specifically, they identify a person's emotional state by analyzing facial expressions or text content.
[0003] The inventors realized that traditional emotion methods mostly rely on single-modal data (such as text or facial expressions), which makes it difficult to accurately identify complex human emotional states, resulting in low accuracy. Summary of the Invention
[0004] The present invention provides an artificial intelligence-based emotion analysis method, apparatus, computer equipment, and medium to solve the problem of low accuracy of traditional emotion analysis methods.
[0005] First, a sentiment analysis method based on embodied intelligence is provided, including:
[0006] Collecting multimodal data of the target user, wherein the multimodal data includes physiological signal data and behavioral performance data;
[0007] quantizing the multimodal data to obtain a multimodal feature vector;
[0008] fusing the multimodal feature vectors to obtain a fused vector;
[0009] The fusion vector is input into a preset model so that the preset model outputs the emotion category of the target user.
[0010] In a second aspect, a sentiment analysis device based on embodied intelligence is provided, comprising:
[0011] An acquisition module, configured to acquire multimodal data of a target user, wherein the multimodal data includes physiological signal data and behavioral performance data;
[0012] A quantization module, configured to perform quantization processing on the multimodal data to obtain a multimodal feature vector;
[0013] A fusion module, configured to fuse the multimodal feature vectors to obtain a fusion vector;
[0014] An output module is used to input the fusion vector into a preset model so that the preset model outputs the emotion category of the target user.
[0015] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned emotion analysis method based on embodied intelligence when executing the computer program.
[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned emotion analysis method based on embodied intelligence are implemented.
[0017] In the scheme implemented by the above-mentioned emotion analysis method, device, computer equipment and storage medium based on embodied intelligence, multimodal data of the target user is collected, wherein the multimodal data includes physiological signal data and behavioral performance data; the multimodal data is quantized to obtain a multimodal feature vector; the multimodal feature vector is fused to obtain a fusion vector; the fusion vector is input into a preset model so that the preset model outputs the emotion category of the target user. In the present invention, physiological signal data and behavioral performance data can be collected to form multimodal data, and then the multimodal data can be quantized to obtain a multimodal feature vector, and multiple multimodal feature vectors are fused into a fusion vector. Finally, the fusion vector is input into the preset model, and the preset model then outputs the emotion category, so that the user's emotion category can be identified through multiple dimensions, thereby improving the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0019] Figure 1 is a flow chart of a method for emotion analysis based on embodied intelligence in one embodiment of the present invention;
[0020] Figure 2 yes Figure 1 A schematic flow chart of a specific implementation of step S10;
[0021] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step S20;
[0022] Figure 4 yes Figure 1Another specific implementation flow diagram of step S20;
[0023] Figure 5 yes Figure 1 A flow chart of another specific implementation of step S20;
[0024] Figure 6 yes Figure 1 A schematic flow chart of a specific implementation of step S30;
[0025] Figure 7 yes Figure 1 A schematic flow chart of a specific implementation of step S40;
[0026] Figure 8 is a structural diagram of an emotion analysis device based on embodied intelligence in one embodiment of the present invention;
[0027] Figure 9 is a structural diagram of a computer device in one embodiment of the present invention;
[0028] Figure 10 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] The embodied intelligence-based emotion analysis method provided in embodiments of the present invention can be applied to computer devices, including but not limited to various personal computers and laptop computers. This method can collect physiological signal data and behavioral performance data to obtain multimodal data, then quantize the multimodal data to obtain a multimodal feature vector. Multiple multimodal feature vectors are then fused to obtain a fused vector. Finally, this fused vector is input into a preset model, which is used to determine the target user's emotion category. The present invention is described in detail below using specific embodiments.
[0031] See also Figure 1 As shown, Figure 1 A flowchart of a sentiment analysis method based on embodied intelligence provided by an embodiment of the present invention includes the following steps:
[0032] S10: Collecting multimodal data of the target user, wherein the multimodal data includes physiological signal data and behavioral performance data.
[0033] The embodied intelligence-based emotion analysis method provided by this invention can be applied to emotion analysis in various application scenarios, such as nursing and psychological healthcare. In nursing scenarios, such as nursing homes, emotion analysis can be used to identify the user's emotional state, facilitating targeted care. In psychological healthcare, psychologists can use emotion analysis to identify patients' emotional states in response to different descriptions, facilitating targeted medical treatment.
[0034] Multimodal data can include physiological signal data and behavioral performance data, both of which can be collected through corresponding sensors or calculated from the collected data. For example, physiological signal data can be obtained through a heart rate monitor and skin conductivity sensor to obtain heart rate data and skin conductivity data, respectively. Behavioral performance data can be captured by a camera to capture changes in the target user's facial expressions and joint range of motion. In addition to physiological signal data and behavioral performance data, it can also include text data, such as user-entered comments.
[0035] It should be understood that physiological signal data may include heart rate data and skin conductivity data, and behavioral performance data may include facial expression image data and human body movement posture data. Figure 2 As shown, step S10, i.e. collecting multimodal data of the target user, includes the following steps:
[0036] S11: Collecting the heart rate data and skin conductivity data of the target user to obtain the physiological signal data.
[0037] S12: Collecting facial expression image data and body movement posture data of the target user to obtain the behavior performance data.
[0038] For steps S11-S13, physiological signal data may include heart rate data and skin conductivity data. Heart rate data can be obtained through heart rate monitoring devices, such as smart bracelets, heart rate chest straps, etc. Taking smart bracelets as an example, their built-in photoelectric sensors measure heart rate by emitting green light and detecting changes in reflected light. When the heart beats, the blood flow and color in the blood vessels change, causing the intensity of the reflected light to change. The sensor converts this light intensity change into an electrical signal, and after algorithm processing, outputs the heart rate value. Skin conductivity data is used to reflect changes in the secretion activity of sweat glands on the skin surface, and the secretion activity of sweat glands is closely related to mood swings. Skin conductivity data can be obtained through a skin conductivity sensor.
[0039] Behavioral performance data can include facial expression image data and human motion and posture data. Facial expression image data can be obtained using facial expression recognition cameras, which use computer vision technology to identify and analyze facial expressions. After capturing facial images, the camera uses image preprocessing, feature extraction, and classification to identify different facial expressions, such as happiness, sadness, anger, and surprise. Common facial expression recognition algorithms are based on deep learning. By learning from a large number of annotated facial expression images, the trained model can accurately predict the expression category in the input image. Human motion and posture data can be collected through motion capture systems, such as optical motion capture and inertial motion capture. Optical motion capture uses reflective markers attached to key parts of the body and captures images from different angles using multiple cameras. Based on the changes in the position of the markers in the different images, the three-dimensional coordinates of the human joints are calculated to reconstruct human movements. Inertial motion capture uses inertial sensors worn on various parts of the body to measure physical quantities such as acceleration and angular velocity, thereby inferring human motion and posture.
[0040] S20: quantizing the multimodal data to obtain a multimodal feature vector.
[0041] After obtaining the multimodal data, the multimodal data can be quantized to obtain multiple multimodal feature vectors. For example, if the multimodal data includes heart rate data, skin conductivity data, facial expression image data, and human motion posture data, the heart rate data, skin conductivity data, facial expression image data, and human motion posture data can be quantized to obtain a heart rate feature vector, a skin conductivity feature vector, a facial expression feature vector, and a human motion feature vector, respectively.
[0042] Among them, such as Figure 3 As shown, step S20, i.e. collecting multimodal data of the target user, includes the following steps:
[0043] S21: Calculate the heart rate variability using a first preset formula to obtain a heart rate feature vector, wherein the first preset formula is:
[0044]
[0045] HRV is heart rate variability, RR i is the interval between adjacent heartbeats, is the average interval;
[0046] S22: Calculate the change amplitude using a second preset formula or calculate the change rate using a third preset formula to obtain a skin conductivity feature vector, wherein the second preset formula is:
[0047] Δσ=σ2-σ1(2)
[0048] Δσ is the amplitude of change, σ2 and σ1 are the conductivity values in the time interval [t1, t2];
[0049] The third preset formula is:
[0050]
[0051] is the rate of change, σ2 and σ1 are the conductivity values in the time interval [t1, t2].
[0052] Formula (1) is used to calculate the heart rate variability to complete the quantification of the heart rate data, and then the heart rate feature vector can be obtained. Formula (1) includes RR i and Two parameters, among which RR i is the interval between adjacent heartbeats, is the average interval, n is the number of heartbeats, if RR1=800, RR2=750, RR3=700, RR1=720, RR1=850, the unit is ms, then Calculate the difference between each RR and the mean and square it to get The variance is 3730, and the standard deviation is 61.07, that is, HRV = 61.07.
[0053] Formula (2) is used to calculate the change in skin conductivity. For example, if a user is watching a movie, during a quiet segment (t1 = 0), σ1 = 2. When a scary segment suddenly appears (t2 = 2), σ2 = 8. Substituting this into formula (2), we get Δσ = 8 - 2 = 6.
[0054] Formula (3) is used to calculate the rate of change of skin conductivity. Taking the above example as an example, substituting the data in the above example into formula (3) yields
[0055] It can be understood that both formula (2) and formula (3) can be used to obtain the skin conductivity feature vector. In actual use, any method can be selected to calculate the skin conductivity feature vector.
[0056] Furthermore, the collected physiological signal data may contain noise and outliers. For example, heart rate data may contain abnormally high or low heart rate values due to improper device wear or external interference. For these outliers, statistical methods can be used to detect and process them. For example, a reasonable heart rate range can be set (e.g., an adult's resting heart rate is between 40-180 BPM), and values outside this range are considered outliers. Linear interpolation is then used to replace adjacent normal data. For skin conductivity data, a sliding average method can be used to remove short-term fluctuation noise.
[0057] Among them, Figure 3 As shown, step S20, i.e. collecting multimodal data of the target user, further includes the following steps:
[0058] S23: Acquire the facial expression image data, and extract facial key points of the facial expression image data;
[0059] S24: Calculate the displacement change of each facial key point to obtain displacement data, and perform weighted summation on all displacement data to obtain a facial expression feature vector.
[0060] The facial expression of the target user can be captured in real time through a high-precision camera (such as an RGB camera or a depth camera), and then the face area can be located using MTCNN (Multi-Task Convolutional Neural Network) or Dlib library. The key facial feature points (such as eyes, eyebrows, mouth corners, nose bridge, etc.) are extracted and the coordinate sequence is output: P = [(x1, y1), (x2, y2), ..., (x n ,y n )]. Let the key point coordinates of the reference emotional state be P neutral , the key point coordinates of the current facial expression are P current , we can calculate the displacement distance of each point di=(x i -x i 0 ) 2 +(y i -y i 0 ) 2 , (i=1,2,…,n). After obtaining the displacement distance of each point, we can perform a weighted summation of the distances of each point. The weight of each part can be set by the contribution of muscle movement to the expression. For example, the weight of the eyes is 0.4, the weight of the mouth is 0.3, and the weight of the eyebrows is 0.3, then:
[0061]
[0062] Among them, the “maximum possible displacement” is the maximum displacement value of the same expression in the training data, which is used to normalize to [0,1].
[0063] Among them, Figure 4 As shown, step S20, i.e. collecting multimodal data of the target user, further includes the following steps:
[0064] S25: Acquire the human body motion posture data, and extract the three-dimensional coordinates of the joints of the human body motion posture data;
[0065] S26: Calculate the joint change speed according to the three-dimensional coordinates of the joint to obtain a human motion feature vector.
[0066] Motion capture can be done in two ways: optical motion capture and inertial motion capture. Optical motion capture uses reflective markers attached to human joints and uses multiple cameras (e.g., 8-16) to capture 3D coordinates in real time. Inertial motion capture requires wearing inertial sensors (accelerometer + gyroscope) and using Kalman filtering to calculate joint angles, which is suitable for dynamic scenes. Motion capture can obtain multiple 3D coordinates of joints, such as P t =(x t ,y t ,z t ), it can be understood that each joint has its own corresponding three-dimensional joint coordinates. Then, the speed of the joint moving from position P1 to P2 in the time interval Δt is calculated to obtain the human body motion feature vector.
[0067] S30: Fusing the multimodal feature vectors to obtain a fused vector.
[0068] After obtaining the multimodal feature vector, multiple multimodal feature vectors can be fused to obtain a fused vector.
[0069] Among them, Figure 5 As shown, in step S30, that is, fusing the multimodal feature vectors to obtain a fused vector, the following steps are included:
[0070] S31: performing normalization processing on the multimodal feature vector to obtain a plurality of normalized data;
[0071] S32: performing splicing processing on the plurality of normalized data to obtain the fusion vector.
[0072] The multimodal feature vectors may include a heart rate feature vector, a skin conductivity feature vector, a facial expression feature vector, and a human motion feature vector. During normalization, the heart rate feature vector, the skin conductivity feature vector, the facial expression feature vector, and the human motion feature vector need to be scaled to a uniform range, for example, between 0 and 1. It will be appreciated that the normalization method is a well-established technical solution in this field, and its details will not be elaborated upon here.
[0073] After obtaining the normalized data, multiple normalized data are sequentially spliced to obtain a new feature vector, also known as the fusion vector. For example, V fusion =[heart rate feature vector, skin conductivity feature vector, facial expression feature vector, human body movement feature vector].
[0074] S40: Inputting the fusion vector into a preset model so that the preset model outputs the emotion category of the target user.
[0075] After obtaining the fusion vector, the fusion vector is input into the preset model, and the preset model outputs the emotion category with the highest probability based on the fusion vector.
[0076] Alternatively, the output of a pre-set model can be evaluated. For example, the contribution of different modal data to the final prediction result can be calculated. The importance of a modal data can be measured by gradually removing it and observing the decline in model performance. Assuming the accuracy of the original model on the test set is Acc1, and the accuracy of the model after removing a modal data (such as physiological data) is Acc2, the contribution C of that modal data to the accuracy can be expressed as: C = 1 - Acc2 / Acc1. The closer C is to 1, the greater the impact of that modality on model performance, while the closer C is to 0, the smaller the impact. The evaluation method also uses a combination of K-fold cross-validation and independent test set evaluation. In K-fold cross-validation, the multimodal data in each fold is ensured to be evenly distributed, and the correlations between the different modal data are properly reflected. In the independent test set evaluation, the model's generalization ability on unseen multimodal data is comprehensively assessed, including the recognition accuracy of different emotion categories and its ability to analyze complex emotional scenarios.
[0077] Based on the evaluation results of the preset model, the preset model can also be optimized. For example, multimodal data can be augmented. For physiological data, data diversity can be increased by performing random translation and scaling operations on the data. For example, the baseline value of heart rate data can be randomly adjusted within a certain range. Facial expression and motion data can be augmented using methods such as image transformation (such as rotation and cropping) and motion interpolation. For example, a small rotation of the collected facial expression images can be performed to generate new expression samples. Data augmentation can be used to expand the training set size and improve the model's generalization ability. Alternatively, the model structure can be adjusted. If a particular modality is found to contribute little to model performance, the number of processing layers or neurons in the model can be reduced to simplify the model structure and reduce computational complexity. Conversely, if a modality is closely associated with emotional value but the model fails to fully learn its characteristics, the complexity of the data processing module for that modality can be appropriately increased, such as by adding convolutional or fully connected layers to improve the model's feature extraction ability for that modality. Alternatively, methods such as grid search or random search can be used to comprehensively optimize the model's hyperparameters.
[0078] Among them, Figure 6 As shown, in step S40, that is, inputting the fusion vector into a preset model so that the preset model outputs the emotion category of the target user, the following steps are included:
[0079] S41: Obtain the fusion vector, and calculate the score of each emotion category based on a preset weight matrix;
[0080] S42: Calculate the probability of each emotion category according to the score of each emotion category, and output the emotion category with the highest probability.
[0081] Assume that the weight matrix of the fully connected layer is: W∈R C×D , C is the number of emotion categories, for example, 6 categories: happy, sad, angry, fearful, surprised, neutral, bias b∈R C , input fusion vector V∈R D , then the output Z=W*v+b, where each element z i Represents the model's "score" for the i-th category of emotion. The higher the score, the greater the possibility of belonging to that category. For example, let the fusion vector V fusion = [0.57, 0.9, 0.9, 0.6], where 0.57 is the normalized heart rate feature vector, the first 0.9 is the normalized skin conductivity feature vector, the second 0.9 is the normalized facial expression feature vector, and 0.6 is the human motion feature vector. C = 3 (calm, nervous, fear), and the weight matrix is:
[0082]
[0083] Bias b = [0.1, -0.2, 0.3], then Z = [-0.14, 0.28, 1.25]. Substituting Z into the probability calculation formula yields:
[0084]
[0085] Similarly, we can get P 紧张 =0.25, P 恐惧 =0.6, which means that the probability of fear is the highest, and the emotion category output by the preset model is fear.
[0086] It can be seen that in the above scheme, multimodal data can be obtained by collecting physiological signal data and behavioral performance data, and then multimodal feature vectors and fusion vectors are obtained in succession based on the multimodal data. The fusion vectors are then input into the preset model, and the emotion category is confirmed through the preset model, so that emotions can be identified from multiple dimensions, thereby improving the accuracy.
[0087] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0088] In one embodiment, a device for emotion analysis based on embodied intelligence is provided, which corresponds one-to-one to the method for emotion analysis based on embodied intelligence in the above embodiment. Figure 8 As shown, the emotion analysis device based on embodied intelligence includes a collection module 101, a quantification module 102, a fusion module 103 and an output module 104. The functional modules are described in detail as follows:
[0089] The acquisition module 101 is used to collect multimodal data of the target user, wherein the multimodal data includes physiological signal data and behavioral performance data;
[0090] a quantization module 102, configured to perform quantization processing on the multimodal data to obtain a multimodal feature vector;
[0091] A fusion module 103 is configured to fuse the multimodal feature vectors to obtain a fusion vector;
[0092] The output module 104 is configured to input the fusion vector into a preset model so that the preset model outputs the emotion category of the target user.
[0093] In one embodiment, the acquisition module 101 is specifically configured to:
[0094] collecting heart rate data and skin conductivity data of the target user to obtain the physiological signal data;
[0095] The facial expression image data and human body movement posture data of the target user are collected to obtain the behavior performance data.
[0096] In one embodiment, the quantization module 102 is specifically configured to:
[0097] The heart rate variability is calculated using a first preset formula to obtain a heart rate feature vector, wherein the first preset formula is:
[0098]
[0099] HRV is heart rate variability, RR i is the interval between adjacent heartbeats, is the average interval;
[0100] The skin conductivity feature vector is obtained by calculating the change amplitude using a second preset formula or calculating the change rate using a third preset formula, wherein the second preset formula is:
[0101] Δσ=σ2-σ1(2)
[0102] Δσ is the amplitude of change, σ2 and σ1 are the conductivity values in the time interval [t1, t2];
[0103] The third preset formula is:
[0104]
[0105] is the rate of change, σ2 and σ1 are the conductivity values in the time interval [t1, t2].
[0106] In one embodiment, the quantization module 102 is further configured to:
[0107] Acquire the facial expression image data, and extract facial key points of the facial expression image data;
[0108] The displacement change of each facial key point is calculated to obtain displacement data, and all the displacement data are weighted summed to obtain a facial expression feature vector.
[0109] In one embodiment, the quantization module 102 is further configured to:
[0110] Acquire the human body motion posture data, and extract the three-dimensional coordinates of the joints of the human body motion posture data;
[0111] The joint change speed is calculated according to the three-dimensional coordinates of the joint to obtain the human body motion feature vector.
[0112] In one embodiment, the fusion module 103 is configured to:
[0113] Normalizing the multimodal feature vector to obtain a plurality of normalized data;
[0114] The plurality of normalized data are spliced together to obtain the fusion vector.
[0115] In one embodiment, the output module 104 is specifically configured to:
[0116] Obtaining the fusion vector and calculating the score of each emotion category based on a preset weight matrix;
[0117] Calculate the probability of each emotion category based on the score of each emotion category, and output the emotion category with the highest probability.
[0118] The present invention provides an emotion analysis device based on embodied intelligence, which can obtain multimodal data by collecting physiological signal data and behavioral performance data, and then obtain multimodal feature vectors and fusion vectors based on the multimodal data. The fusion vectors are then input into a preset model, and the emotion category is confirmed by the preset model, so that emotions can be identified from multiple dimensions, thereby improving the accuracy.
[0119] For the specific definition of the embodied intelligence-based emotion analysis device, please refer to the definition of the embodied intelligence-based emotion analysis method above, and will not be repeated here. The various modules in the above-mentioned embodied intelligence-based emotion analysis device can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.
[0120] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a sentiment analysis method based on embodied intelligence.
[0121] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 10As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the client side of a sentiment analysis method based on embodied intelligence.
[0122] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0123] Collecting multimodal data of the target user, wherein the multimodal data includes physiological signal data and behavioral performance data;
[0124] quantizing the multimodal data to obtain a multimodal feature vector;
[0125] fusing the multimodal feature vectors to obtain a fused vector;
[0126] The fusion vector is input into a preset model so that the preset model outputs the emotion category of the target user.
[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0128] Collecting multimodal data of the target user, wherein the multimodal data includes physiological signal data and behavioral performance data;
[0129] quantizing the multimodal data to obtain a multimodal feature vector;
[0130] fusing the multimodal feature vectors to obtain a fused vector;
[0131] The fusion vector is input into a preset model so that the preset model outputs the emotion category of the target user.
[0132] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0133] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0134] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0135] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A sentiment analysis method based on embodied intelligence, characterized in that: include: Collecting multimodal data of the target user, wherein the multimodal data includes physiological signal data and behavioral performance data; quantizing the multimodal data to obtain a multimodal feature vector; fusing the multimodal feature vectors to obtain a fused vector; The fusion vector is input into a preset model so that the preset model outputs the emotion category of the target user.
2. The method according to claim 1, wherein The collecting of multimodal data of the target user includes: collecting heart rate data and skin conductivity data of the target user to obtain the physiological signal data; The facial expression image data and human body movement posture data of the target user are collected to obtain the behavior performance data.
3. The method according to claim 2, wherein The quantizing process of the multimodal data to obtain a multimodal feature vector includes: The heart rate variability is calculated using a first preset formula to obtain a heart rate feature vector, wherein the first preset formula is: HRV is heart rate variability, RR i is the interval between adjacent heartbeats, is the average interval; The skin conductivity feature vector is obtained by calculating the change amplitude using a second preset formula or calculating the change rate using a third preset formula, wherein the second preset formula is: Δσ=σ2-σ1 Δσ is the amplitude of change, σ2 and σ1 are the conductivity values in the time interval [t1, t2]; The third preset formula is: is the rate of change, σ2 and σ1 are the conductivity values in the time interval [t1, t2].
4. The method according to claim 2, wherein The quantizing process of the multimodal data to obtain a multimodal feature vector further includes: Acquire the facial expression image data, and extract facial key points of the facial expression image data; The displacement change of each facial key point is calculated to obtain displacement data, and all the displacement data are weighted summed to obtain a facial expression feature vector.
5. The method according to claim 2, wherein The quantizing process of the multimodal data to obtain a multimodal feature vector further includes: Acquire the human body motion posture data, and extract the three-dimensional coordinates of the joints of the human body motion posture data; The joint change speed is calculated according to the three-dimensional coordinates of the joint to obtain the human body motion feature vector.
6. The method according to claim 1, wherein The fusing the multimodal feature vectors to obtain a fused vector includes: Normalizing the multimodal feature vector to obtain a plurality of normalized data; The plurality of normalized data are spliced together to obtain the fusion vector.
7. The method according to claim 1, wherein The preset model outputs the emotion category of the target user, including: Obtaining the fusion vector and calculating the score of each emotion category based on a preset weight matrix; Calculate the probability of each emotion category based on the score of each emotion category, and output the emotion category with the highest probability.
8. An emotional analysis device based on embodied intelligence, characterized in that: include: An acquisition module, configured to acquire multimodal data of a target user, wherein the multimodal data includes physiological signal data and behavioral performance data; A quantization module, configured to perform quantization processing on the multimodal data to obtain a multimodal feature vector; A fusion module, configured to fuse the multimodal feature vectors to obtain a fusion vector; An output module is used to input the fusion vector into a preset model so that the preset model outputs the emotion category of the target user.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the emotion analysis method based on embodied intelligence are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the emotion analysis method based on embodied intelligence are implemented.
Citation Information
Patent Citations
Multi-mode XR emotion interaction method, system and device and storage medium
CN119251438A
Head-up display adjusting method, system and device based on emotion recognition and medium
CN119502686A