Emotion processing method and device based on multi-modal data
Through multimodal data acquisition and sentiment analysis models, combined with sentiment intervention strategies, the problems of low accuracy and lack of intervention mechanism in single-modal emotion recognition systems are solved, high-accuracy emotion recognition and active intervention are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510773087.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
Existing emotion recognition systems rely on a single modality, have limited accuracy and generalization capabilities, and lack emotion intervention mechanisms, which can lead to long-term negative emotions in users and induce psychological problems.
Multimodal data is used to obtain users' visual, language and physiological signals, emotions are identified through feature extraction and sentiment analysis models, and proactive intervention is carried out in combination with preset strategies to improve the accuracy of emotion recognition and provide emotional intervention.
Multimodal data fusion improves the accuracy of emotion recognition, enhances user experience, provides an active emotion intervention mechanism, and reduces the duration of negative emotions.
Smart Images

Figure CN120670948A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of emotion processing, and in particular to emotion processing methods, devices, equipment, and computer-readable storage media based on multimodal data. Background Art
[0002] Emotion recognition is a technology that judges people's emotional changes. It mainly infers people's psychological state by collecting their external expressions and behavioral changes.
[0003] Currently, existing emotion recognition systems rely solely on a single modality (such as speech or image), resulting in limited accuracy and generalization capabilities, and lacking mechanisms for emotional intervention. Long-term negative emotions can lead to psychological problems in users, so a comprehensive system combining efficient multimodal recognition and proactive intervention is urgently needed. Summary of the Invention
[0004] According to the embodiments of the present application, an emotion processing solution based on multimodal data is provided, which can trigger different intervention behaviors according to the emotion recognition results. On the basis of greatly improving the accuracy of emotion recognition, an active intervention mechanism is added to further enhance the user experience.
[0005] In a first aspect of the present application, a method for emotion processing based on multimodal data is provided. The method comprises: Obtain multimodal data information; Performing feature extraction on the data information to obtain multiple features; Inputting the plurality of features into a trained sentiment analysis model to obtain target emotions including a plurality of emotion classifications; Analyze the target emotion through a preset emotion analysis strategy to determine the main emotion category; According to the main emotion category, the corresponding emotion intervention strategy is executed.
[0006] Furthermore, the multimodal data information includes visual modality data, language signal data, and / or physiological signal data.
[0007] Furthermore, the multiple features include: facial expression intensity, body movements, emotional trait scores, heart rate, heart rate variability, galvanic skin response, and / or skin temperature.
[0008] Furthermore, the expression intensity includes:
[0009] in, is the weight of the i-th key point; is the offset distance of the i-th key point relative to the static state.
[0010] Furthermore, the body movements include:
[0011] in, The degree of body closure; is the position of the i-th joint; The center of the body.
[0012] Furthermore, the heart rate variability includes:
[0013] in, is the i-th heartbeat interval; is the interval mean; N is the number of sampling times.
[0014] Furthermore, executing a corresponding emotion intervention strategy according to the main emotion category includes: Implement corresponding emotional intervention strategies based on the main emotional category and time trend.
[0015] In a second aspect of the present application, a device for processing emotions based on multimodal data is provided. The device comprises: Acquisition module, used to obtain multimodal data information; An extraction module, configured to extract features from the data information to obtain a plurality of features; A classification module, configured to input the plurality of features into a trained sentiment analysis model to obtain target sentiments comprising a plurality of sentiment classifications; A determination module, configured to analyze the target emotion using a preset emotion analysis strategy to determine a primary emotion category; The execution module is used to execute the corresponding emotion intervention strategy according to the main emotion category.
[0016] In a third aspect of the present application, an electronic device is provided, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the program.
[0017] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present application is implemented.
[0018] The emotion processing method based on multimodal data provided in the embodiment of the present application obtains multimodal data information; performs feature extraction on the data information to obtain multiple features; inputs the multiple features into a trained emotion analysis model to obtain target emotions including multiple emotion classifications; analyzes the target emotions through a preset emotion analysis strategy to determine the main emotion category; and executes the corresponding emotion intervention strategy according to the main emotion category. On the basis of greatly improving the emotion recognition accuracy, an active intervention mechanism is added to further enhance the user experience.
[0019] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of the present application, nor are they intended to limit the scope of the present application. Other features of the present application will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other features, advantages and aspects of the embodiments of the present application will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein: Figure 1 is a flowchart of a method for emotion processing based on multimodal data according to an embodiment of the present application; Figure 2 is a block diagram of an emotion processing device based on multimodal data according to an embodiment of the present application; Figure 3 A schematic diagram of the structure of a terminal device or server suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION
[0021] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present disclosure.
[0022] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0023] Figure 1 A flowchart of a method for processing emotions based on multimodal data according to an embodiment of the present disclosure is shown. The method includes: S110, acquiring multimodal data information.
[0024] In some embodiments, multimodal data information can be acquired via wired or wireless means; the multimodal data information includes visual modality data, speech signal data, and / or physiological signal data. Specifically, the user's current explicit and implicit emotional cues can be simultaneously collected through multiple sensory devices, including visual modalities (facial expressions, body movements), speech signals (speech rate, pitch, etc.), and / or physiological indicators (HR, HRV, GSR, body temperature, etc.).
[0025] For example, visual modality data can be acquired through a camera; language signal data can be collected through a microphone; and physiological signal data can be collected through the user's wearable device.
[0026] In some embodiments, the timing of the buffer queue and the linear interpolation can be kept consistent by unifying the timestamp, that is, the consistency of the collected data is maintained.
[0027] S120: Extract features from the data information to obtain multiple features.
[0028] In some embodiments, feature extraction of visual modality data includes facial and body recognition.
[0029] The expression intensity (facial muscle movement analysis) can be calculated as follows:
[0030] in, is the weight of the i-th key point (e.g., the weight of the upturned corner of the mouth is large); is the offset distance of the i-th key point relative to the static state; Body movement extraction can be performed in the following ways:
[0031] in, The degree of physical closure, with less closure generally indicating more negative emotions (e.g., sadness); is the position of the i-th joint; For the center of the body (such as the center of the chest cavity); In some embodiments, multiple emotion-related features can be extracted from language signal data using the librosa tool: Specifically, MFCC (Mel Frequency Cepstral Coefficients): Represents the speech spectrum envelope, suitable for modeling language content; Extract 13-39 dimensional MFCC features, commonly used for deep network input; Calculation method: Short-time Fourier transform (STFT) can be performed first, followed by Mel filter bank and DCT; Furthermore, Pitch (fundamental frequency) and its rate of change (Pitch_std): Characterizes pitch fluctuations and reflects the degree of emotional activity; High Pitch_std usually occurs in high activation states such as anger and surprise; Further, Prosody (rhythmic features): Including speech rate (number of words per second), pause length distribution, energy envelope, etc. Standard speech fluency, rhythm and emotional rhythm; Furthermore, F0 (fundamental frequency trace) and volume (RMS): Detecting emotional trends through rising / falling pitch; High frequency + high volume corresponds to anger / excitement, and high frequency + low volume corresponds to anxiety. Further, the emotional feature score calculation formula is as follows
[0032] in, 、 is the empirical weight coefficient, and its value can be pre-set according to the actual application scenario (such as 0.5); is the standard deviation of pitch, which is used to reflect emotional fluctuations in speaking; The amplitude of the change in sound characteristics (emotional fluctuations).
[0033] In some embodiments, features may be extracted from physiological signal data in the following manner.
[0034] Heart rate: HR=60 / RR Where RR is the interval between two consecutive heartbeats (unit: seconds); HR is the heartbeats per unit event (bmp); Heart rate variability includes:
[0035] in, is the i-th heartbeat interval; is the interval mean; N is the number of sampling times; Galvanic skin response (GSR), which can be directly obtained through sensors, shows a rapid increase in skin conductance when emotions are highly activated. Characteristics include GSR average, GSR peak response rate, and derivative change rate. Skin temperature can be directly obtained through sensors. When people are emotionally excited, their blood vessels contract and their temperature drops. This can be used as a basis for judging anxiety and tension. The unit is usually ℃, and the sampling period is 1Hz.
[0036] S130: Input the multiple characteristics into a trained emotion analysis model to obtain target emotions including multiple emotion classifications. In some embodiments, the sentiment analysis model may be trained as follows: Generate a training sample set, wherein the training sample includes a script file with annotation information; the annotation information is an emotion label; The sentiment analysis model is trained using samples in the training sample set, with the script file as input and the sentiment label as output. When the unification rate of the output sentiment label and the marked sentiment label meets the preset threshold, the training of the sentiment analysis model is completed.
[0037] In some embodiments, the fused feature vector F = [Ef, Cb, Eu, HR, HRV, GSR, Temp] can be input into a deep model (e.g., Transformer + MLP) to output a softmax classification result; Preferably, the emotion classification output and reference value range include: Happy, softmax output reference value: 0.60~0.90; Anger, softmax output reference value: 0.50~0.85; Sadness, softmax output reference value: 0.50~0.80; Surprised, softmax output reference value: 0.60~0.95; Anxiety, softmax output reference value: 0.45~0.75; Relax, softmax output reference value: 0.55~0.85; Neutral, softmax output reference value: 0.50~0.70; Boring, softmax output reference value: 0.50~0.75; Fatigue: Softmax output reference value: 0.50~0.78.
[0038] S140: Analyze the target emotion using a preset emotion analysis strategy to determine a primary emotion category.
[0039] In some embodiments, the following method can be used to distinguish overlapping emotions (such as anxiety / anger, boredom / tiredness) in the softmax output. Usually, the Top-1 maximum probability attribution method can be used for initial judgment, and the maximum softmax probability is taken as the current main emotion L=arg max( ),For example: 0.62, .59 => judged as "happy" Furthermore, if the difference between the first two softmax probabilities is less than a threshold δ (e.g., 0.5), the current state is recorded as a mixed emotion state, and intervention is delayed or a comprehensive strategy is adopted, for example: .53, .52 => Classified as a mixed state of tiredness and boredom If the softmax falls into the overlapping area or the edge of the intervention threshold, it can be corrected as follows: max
[0040] in, It is an auxiliary indicator of this type of emotion (such as HRV decrease, GSR fluctuation, etc.); is the weighting coefficient (preferably in the range of 0.1 to 0.3); For example: “Anxiety vs. anger” is distinguished by GSR / HRV; “Bored vs. tired” judged by posture and speech speed; Furthermore, modal confidence weighting (system robustness) can be used to assign different modal weights based on acquisition quality: =
[0041] in, is the final weighted score of the k-th emotion label after modal fusion, with an output value of 0 to 1, which is used to determine the overall emotion judgment; is the image modal confidence weight, is the modal confidence coefficient, and its value range is 0~1, which is used to reflect the quality and credibility of image data; The softmax probability (k-th category emotion) of the image modality (such as facial expression, body posture) is the input variable and is output by the image recognition network; is the speech modal confidence weight, and is the modal confidence coefficient. For example, if the speech signal is clear, the value can be 0.4~0.5; The softmax probability of the k-th category information identified by the speech modality is the input variable and is calculated based on speech features such as MFCC and Pitch; is the physiological modal confidence weight, is the modal confidence coefficient, and the weight is determined by whether the GSR / HR acquisition is complete; The probability of the kth emotion identified by physiological modalities (such as heart rate, GSR, HRV, etc.) is the input variable and comes from the softmax output of the physiological feature model.
[0042] Furthermore, in order to avoid the influence of single modality damage (such as image blur, language interruption) on the results, the confidence level can be adjusted dynamically in real time. , ensuring stable output.
[0043] S150: Execute a corresponding emotion intervention strategy according to the main emotion category.
[0044] In some embodiments, intelligent intervention strategies may be initiated based on the current emotional state and time trends.
[0045] Specifically, determine whether the intervention trigger conditions are met. If any of the following conditions are met, the intervention is triggered: Probability threshold, softmax probability >ϴ, e.g., ϴ = 0.65; Continuity rule: the same label is output continuously ≥ N times; Trend weighting, the top-N labels fall into the negative set; Physiological support, including HRV reduction, GSR surge and other physiological characteristics support; Active expression, where users display negative emotions in conversation or behavior; Furthermore, intervention strategies can be implemented for various negative emotions, with the duration preferably set by default to 30 to 120 seconds: When the emotion is anger, the intervention goal is to reduce excitement, and the strategy is natural sound; When the emotion is anxiety, the intervention goal is to relax the sympathetic nervous system, and the strategies are meditation and breathing guidance. When the emotion is sadness, the intervention target is emotional arousal, and the strategies are encouraging videos and positive language; When the emotion is boredom, the intervention goal is to activate interaction, and the strategy is to interact with small games and recommend interesting content; When the emotion is fatigue, the intervention goal is to guide rest, and the strategies are dim light mode and soothing music; Furthermore, the intervention exit judgment can be performed in the following ways: After the intervention, observe for a period of time (e.g., 3 minutes). When the mood turns to "neutral / relaxed / happy" or the physiological indicators return to stability or the user actively provides feedback, the intervention is exited.
[0046] According to the embodiments of the present disclosure, the following technical effects are achieved: Through multimodal data collection methods, the user's facial expressions, body movements, voice intonation, physiological signals (such as heart rate, skin electricity) and other data are collected. Based on the multimodal emotion recognition fusion algorithm disclosed in this invention, the user's current emotional state is identified, and corresponding emotion intervention strategies are output according to the detection results, such as playing music, guiding breathing, comforting conversations or making social reminders, etc. On the basis of greatly improving the accuracy of emotion recognition, the user experience is enhanced, and it can be widely used in many fields such as medical health.
[0047] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0048] The above is an introduction to the method embodiment. The following is a device embodiment to further illustrate the solution described in this application.
[0049] Figure 2 FIG2 shows a block diagram 200 of an emotion processing apparatus based on multimodal data according to an embodiment of the present application, as shown in FIG200. Figure 2 Shown include: An acquisition module 210 is used to acquire multimodal data information; An extraction module 220 is used to extract features from the data information to obtain multiple features; a classification module 230 for inputting the plurality of features into a trained sentiment analysis model to obtain target sentiments including a plurality of sentiment classifications; A determination module 240 is configured to analyze the target emotion using a preset emotion analysis strategy to determine a primary emotion category; Execution module 250, for executing the corresponding emotion intervention strategy according to the main emotion category Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0050] Figure 3A schematic diagram of the structure of a terminal device or server suitable for implementing an embodiment of the present application is shown.
[0051] like Figure 3 As shown, the terminal device or server includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage part 308 into the random access memory (RAM) 303. Various programs and data required for the operation of the terminal device or server are also stored in the RAM 303. The CPU 301, ROM 302 and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0052] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, and the like; an output section 307 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 308 including a hard disk; and a communication section 309 including a network interface card such as a LAN card or a modem. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. Removable media 311, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 310 as needed, so that computer programs read therefrom can be installed into the storage section 308 as needed.
[0053] In particular, according to an embodiment of the present application, the above method flow steps can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication part 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above-mentioned functions defined in the system of the present application are executed.
[0054] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0055] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the aforementioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0056] The units or modules involved in the embodiments described in this application may be implemented in software or hardware. The units or modules described may also be provided in a processor. The names of these units or modules do not, in certain circumstances, constitute limitations on the units or modules themselves.
[0057] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device. The computer-readable storage medium stores one or more programs, which, when used by one or more processors, execute the method described in the present application.
[0058] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of application involved in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the aforementioned application concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions applied for in this application.
Claims
1. A method for emotion processing based on multimodal data, characterized in that: include: Obtain multimodal data information; Performing feature extraction on the data information to obtain multiple features; Inputting the plurality of features into a trained sentiment analysis model to obtain target emotions including a plurality of emotion classifications; Analyze the target emotion through a preset emotion analysis strategy to determine the main emotion category; According to the main emotion category, the corresponding emotion intervention strategy is executed.
2. The method according to claim 1, characterized in that The multimodal data information includes visual modality data, language signal data, and / or physiological signal data.
3. The method according to claim 2, characterized in that The features include: facial expression intensity, body movements, emotional trait scores, heart rate, heart rate variability, galvanic skin response, and / or skin temperature.
4. The method according to claim 3, characterized in that The expression intensity includes: in, is the weight of the i-th key point; is the offset distance of the i-th key point relative to the static state.
5. The method according to claim 3, characterized in that The body movements include: in, The degree of body closure; is the position of the i-th joint; The center of the body.
6. The method according to claim 3, characterized in that The heart rate variability includes: in, is the i-th heartbeat interval; is the interval mean; N is the number of sampling times.
7. The method according to claim 1, characterized in that The executing of the corresponding emotion intervention strategy according to the main emotion category includes: Implement corresponding emotional intervention strategies based on the main emotional category and time trend.
8. An emotion processing device based on multimodal data, characterized in that: include: Acquisition module, used to obtain multimodal data information; An extraction module, configured to extract features from the data information to obtain a plurality of features; A classification module, configured to input the plurality of features into a trained sentiment analysis model to obtain target sentiments comprising a plurality of sentiment classifications; A determination module, configured to analyze the target emotion using a preset emotion analysis strategy to determine a primary emotion category; The execution module is used to execute the corresponding emotion intervention strategy according to the main emotion category.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-mode XR emotion interaction method, system and device and storage medium
CN119251438A
Cited By
Interaction method based on emotional recognition and electronic equipment
CN121561761A
Emotion regulation system and device with self-adaptive concession mechanism
CN122297866A