Emotion evaluation system for old people based on multi-modal physiological and behavior signals

By developing an emotional assessment system for the elderly based on multimodal emotion analysis, using multimodal signal acquisition and data fusion model based on attention mechanism, the problem that traditional diagnostic methods are difficult to accurately identify emotional problems in the elderly is solved, and accurate assessment and personalized management of the emotional state of the elderly are achieved.

CN120154337APending Publication Date: 2025-06-17SOUTHEAST UNIV

Patent Information

Application Number
CN202510447987.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Emotional problems such as depression, anxiety and loneliness in the elderly are common, and traditional diagnostic methods are difficult to fully and objectively reflect the true emotional state of the elderly.

Method used

A system for emotional evaluation of elderly people based on multimodal emotion analysis is developed. Through the multimodal signal acquisition module, emotion-induced task paradigm module and multimodal feature processing module, multimodal feature processing module, multimodal signal processing module, etc., the multimodal data fusion model based on attention mechanism is used to extract emotions-related features and achieve accurate identification of emotional states.

Benefits of technology

It realizes accurate assessment and identification of the emotional state of the elderly, overcomes the limitations of traditional methods, and provides a personalized emotional health management solution, which is characterized by high efficiency, low cost and portability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120154337A_ABST
    Figure CN120154337A_ABST
Patent Text Reader

Abstract

The invention discloses an elderly emotion evaluation system based on multi-modal physiological and behavior signals. The elderly emotion evaluation system is composed of a multi-modal signal acquisition module, an emotion induction task design module and a multi-modal feature processing module. By collecting multi-mode signals such as electroencephalogram, electrocardio, voice and facial expressions, the system can synchronously record physiological and behavior changes of a user in a specific situation. The emotion induction task adopts scientifically designed audio and video contents, and six typical emotions such as happiness, tension, boring, anger, calm and sadness are guided. Then, the system extracts emotion related features from the multi-modal signals by using a deep learning method, models a cooperative relationship among the multi-modal signals by using an attention mechanism fusion technology, and generates unified emotion feature representation; by comprehensively analyzing and identifying the emotional state of the elderly user, the system provides important reference for early monitoring of emotional health of the elderly, psychological state evaluation and personalized intervention, and has social and clinical values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The field

[0002] The present invention relates to the technical field of emotion recognition and multimodal fusion, and in particular to an emotion assessment system for the elderly based on multimodal physiological and behavioral signals. Background Art

[0003] As the global aging trend intensifies, the emotional health problems of the elderly population have gradually attracted widespread attention from all walks of life. Studies have shown that emotional problems such as depression, anxiety and loneliness are common among the elderly and significantly affect their quality of life and social functions. According to data from the World Health Organization, more than 260 million people worldwide suffer from depression, among which the elderly are one of the important affected groups. Emotional problems in the elderly not only accelerate the decline of cognitive ability, but are also closely related to the occurrence of cardiovascular disease, decreased immune function and other chronic diseases. In addition, problems such as social isolation, loss of close relationships and decreased ability to take care of themselves make the elderly a high-risk group for emotional problems.

[0004] In China, the prevalence of depression among the elderly is about 20%-30%, and the proportion of those suffering from severe depression is even higher. Because the manifestations of emotional problems in the elderly are often hidden, many patients fail to receive timely diagnosis and intervention, which increases the burden on society and families. Emotional problems are closely related to the behavior, language and physiological reactions of the elderly, but traditional diagnostic methods mainly rely on questionnaires and doctors' subjective judgments. This method is easily affected by cultural background and individual differences, and it is difficult to fully and objectively reflect the true emotional state of the elderly.

[0005] In view of the complexity and diversity of emotional problems of the elderly, it is particularly important to propose an emotion recognition system based on multimodal sentiment analysis. The system we developed makes full use of the synchronous acquisition and analysis capabilities of multimodal data (such as EEG, ECG, voice and facial expressions) to make up for the limitations of traditional methods. The system adopts an advanced emotion induction paradigm to induce typical emotional responses of the elderly (such as happiness, tension, anger, calmness and sadness) through multiple sensory stimuli (such as video and audio), and comprehensively captures and analyzes the changes in the user's emotional state through multimodal data fusion technology. Summary of the invention

[0006] The present invention proposes an emotion recognition system based on multimodal sentiment analysis, aiming to provide an accurate and efficient solution for the emotional health management of the elderly.

[0007] The specific plan is as follows:

[0008] An elderly emotional assessment system based on multi-modal physiological and behavioral signals, characterized in that: it includes a multi-modal signal acquisition module, an emotion induction task paradigm module, and a multi-modal feature processing module; the multi-modal signal acquisition module is used to collect multi-modal signals of physiological signals and behavioral signals of elderly users in a multi-task scenario; the emotion induction task paradigm module guides the target emotion through multimedia content; the multi-modal feature processing module uses a multi-modal data fusion model based on the attention mechanism to extract emotion-related features from multi-modal signals, analyze the different contributions of each modal signal to the emotional state, and analyze the synergistic relationship between multi-modal signals, generate a unified emotional feature representation, and achieve precise recognition of the emotional state of elderly users by analyzing the influence degree of different modalities on the emotional state. This application focuses on exploring the synergistic relationship between multi-modal data, analyzing the different contributions of each modal signal to the emotional state, theoretically revealing the interaction mechanism between the emotional state and multi-modal signals, and forming a complete emotional modeling framework. The system comprehensively collects data through the multi-modal signal acquisition module, including synchronous acquisition of electroencephalogram (EEG), electrocardiogram (ECG), voice signals, and facial expressions, to ensure high-precision recording of physiological and behavioral signals. Combined with a scientific emotion induction task paradigm, the system can induce six typical emotions of users (happiness, tension, boredom, anger, calmness, and sadness), and ensure the efficiency and consistency of data through a standardized process.

[0009] Furthermore, the emotion induction task paradigm module efficiently induces the target emotion through multimedia content and in combination with an emotion induction scenario, wherein the multimedia content includes video, audio, and text. The target emotions include six types: happiness, tension, boredom, anger, calmness, and sadness. The emotion induction task paradigm module quantifies the emotional response intensity based on SAM (Self-Assessment Manikin) from the valence validity and activation degree, and provides unified test conditions for the emotional state based on the standard process of psychological experiments. Among them, the standard process includes clarifying research questions, designing experiments, collecting data, analyzing results, and writing reports. The whole process needs to strictly follow ethical norms, and verify hypotheses through controlling variables and statistical analysis. The emotion induction task paradigm of the system pays attention to the user experience and uses multimedia content such as video, audio, and text as stimulation forms to provide a rigorous standardized design for the emotion induction experiment. By combining the theoretical basis of psychological experiments, the system can accurately guide the target emotional response of users, and at the same time support adaptation and adjustment according to the cultural background and age differences of users.

[0010] Furthermore, the physiological signals include electrocardiogram (ECG) and electroencephalogram (EEG) signals, and the behavioral signals include voice and facial expression signals. The multimodal signal acquisition module constructs a precise acquisition strategy and data standardization process through the synchronous acquisition of multimodal signals of ECG, EEG, voice, and facial expressions. Among them, the acquisition strategy includes device calibration before signal acquisition and real-time monitoring of data quality. The data standardization process is a preprocessing method after signal acquisition, including filtering, downsampling, and normalization, which is used to ensure the high stability and quality of the signals. The multimodal signal acquisition module explores the dynamic correlation between signals in time series by revealing the time-domain and frequency-domain characteristics and time-varying characteristics of different modal signals in the emotional state. Here, the dynamic correlation refers to the characteristic that the correlation between two or more signals over time changes with time, emphasizing that the correlation is dynamic rather than static.

[0011] Furthermore, the emotion induction task paradigm module uses a multi-stage test for the emotional state, including a pre-test stage and a formal test stage. In the pre-test stage, short videos and lightweight tasks are played to guide users to familiarize themselves with the experimental process, reducing the sense of strangeness and operation difficulty at the initial stage of the experiment. In the formal test stage, six emotion induction scenarios are adopted, including visual and auditory stimuli, combined with the dynamic recording of the physiological and behavioral signals of elderly users, to deeply analyze the correlation of the emotional state in multimodal signals, that is, the characteristics and individual differences of the emotional state changing over time.

[0012] Furthermore, the system ensures the scientificity and repeatability of experimental data by constructing a standardized emotion induction scenario. The emotion induction scenario includes the fixed position of the video playback device, a noise-isolated laboratory environment, the screen height at the same level as the user's line of sight, and the precisely controlled playback speed of audio-visual content. By strictly controlling the experimental conditions, the system can ensure the controllability and consistency of the emotion induction task, and at the same time, based on the relationship between the response intensity of the multimodal signals of elderly users using SAM (Self-Assessment Manikin) and the effectiveness of the emotion induction task.

[0013] Furthermore, a physiological and behavioral signal feature encoder is adopted to study the specific contributions of different modal signals in emotion assessment. Among them, a unified transformer framework (ESET-net) is proposed to synchronously model multimodal signals. For physiological and voice signals, they are first transformed into spatio-temporal spectrograms through short-time Fourier transform. For behavioral signals, facial features are extracted through a pre-trained convolutional neural network, and then multimodal fusion is performed on the time-frequency features between different modal signals.

[0014] Furthermore, a multi-modal data fusion model based on the attention mechanism is adopted to conduct collaborative analysis and dynamic weighting on different modal signals. The model consists of N identical layer stacks, and each layer is composed of a multi-head self-attention mechanism and a fully connected feed-forward network. Each sub-layer adopts a residual connection and layer normalization:

[0015] LayerNorm(x + Sublayer(x))(1)

[0016] Where: Sublayer(x) is the function implemented by the sub-layer itself; LayerNorm(x) is the layer normalization function. The attention mechanism maps the query vector Q and the key-value pair vector K-V to the output value:

[0017]

[0018] Where: d k is the dimension of the query and the key. Multi-head attention enables the model to jointly attend to information in different representation subspaces at different positions:

[0019]

[0020] Where: W O is the output weight of the multi-head attention; W i Q 、 W i V are the corresponding weight matrices of Q, K, and V. Thus, the model optimizes the feature extraction and fusion strategy by dynamically adjusting the weights of different modal signals and learning the interdependence between multi-modal signals, generating a high-precision emotional feature representation.

[0021] And based on the above model, the information interaction between different modalities is promoted through the knowledge distillation method, and different models mutually form the teacher network and the student network. The architectures of the student network and the teacher network are similar. The method of pruning the model is to halve the number of layers and initialize the student network layer from the teacher network layer. The model alternates between a full copy layer and an ignore layer, and preferentially copies the top layer or the bottom layer. The student model is supervised and trained with the sample true labels and the teacher network, so as to imitate the teacher network and be as close as possible to its output. By learning the influence degree of different modal signals on the emotional state, the accurate recognition of the emotional state of elderly users is realized. This fusion strategy not only improves the overall accuracy of emotion recognition, but also can adaptively optimize according to the emotional feature differences of different individuals, providing technical support for personalized emotion analysis.

[0022] In addition, the system has a real-time monitoring function. By combining the user's historical data and multi-modal signals collected in real time, it quantifies the changing trend of the emotional state and provides personalized emotional health advice. The system hardware adopts a wearable design, featuring low cost and strong portability, which is convenient for popularization and application in community and home environments. This comprehensive solution provides an effective tool for grass-roots mental health screening and also offers important technical support for the emotional health management of the elderly population.

[0023] Compared with the prior art, the beneficial effects of this invention patent are as follows:

[0024] 1. Comprehensiveness and multi-modal fusion: By integrating multi-modal signals such as electroencephalogram, electrocardiogram, voice, and facial expressions, this system realizes the all-round monitoring and analysis of the emotional state, overcoming the limitations of traditional single-modal assessment methods.

[0025] 2. Precision and dynamic adaptation: By introducing a multi-modal fusion model based on the attention mechanism, the system can dynamically adjust the weights of modal signals according to different situations, thereby achieving the precise recognition of individualized emotional characteristics.

[0026] 3. Real-time and convenience: The system adopts real-time monitoring technology and combines with the design of portable wearable devices, enabling easy deployment in homes and communities, providing a low-cost and efficient solution for grass-roots emotional health management.

[0027] 4. Standardization and repeatability: The system designs a standardized emotional induction task paradigm and experimental environment to ensure the scientific nature of data collection and the repeatability of results, providing reliable support for the scientific research of emotional health.

[0028] 5. Wide application value: This system is not only applicable to the emotional health management of the elderly, but also can provide strong support for mental health screening, clinical intervention, and social psychological services, with significant social and economic benefits. Brief Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the system of the present invention.

[0030] Figure 2 It is a flowchart of experimental acquisition.

[0031] Figure 3 It is a flowchart of the evaluation paradigm.

[0032] Figure 4 It is a multi-modal attention fusion model diagram.

[0033] Figure 5 It is a schematic diagram of multi-modal knowledge distillation. Detailed Description of the Invention

[0034] The present invention will be described in more detail below in conjunction with the accompanying drawings and embodiments.

[0035] As Figure 1 shown, the multimodal signal acquisition system consists of multiple synchronous sensors, including an electroencephalogram (EEG) acquisition device, an electrocardiogram (ECG) recorder, a voice pick-up device, and a high-definition camera, which are used to capture facial expression features. These sensors are connected to the central processing device wirelessly or wiredly to form an efficient multimodal data acquisition network. The system integrates an automatic calibration function for real-time detection and adjustment of the working state of the sensors to ensure the accuracy of signal acquisition. In addition, the acquisition module also adopts advanced noise filtering and signal enhancement algorithms, effectively improving the data quality and providing an effective signal basis for subsequent analysis.

[0036] As Figure 2 and Figure 3 shown, the experimental acquisition process designs a phased emotion induction task paradigm, including pre-test and formal test sessions. In the pre-test stage, short videos or audio are played to help users familiarize themselves with the experimental process and reduce psychological stress; in the formal test stage, targeted multimedia content (such as specific emotion induction videos or audio) is used to stimulate the target emotional response. The whole process is carried out strictly in accordance with the standardized operation process to ensure a high degree of consistency in experimental conditions. At the same time, the system records multimodal signals such as the EEG, ECG, voice, and facial expressions of users in real time, and the data acquisition covers multidimensional physiological and behavioral characteristics. In addition, the built-in discrete emotion selection and dimensional emotion scoring modules in the system can quantify the intensity of emotional responses through subjective feedback from users, laying a foundation for the comprehensive analysis of subjective and objective data.

[0037] As Figure 4 shown, the system adopts a multimodal data fusion model based on the attention mechanism for collaborative analysis and feature optimization of the acquired multimodal signals. This model dynamically adjusts the weights of each modal signal, deeply explores the potential dependence relationships between modalities, and generates a unified emotional feature representation. Through the attention mechanism, the system can accurately identify individualized emotional features and improve the recognition performance for diverse emotional states. The fused features are used to construct a high-precision emotion analysis model, providing scientific support for the comprehensive assessment of users' emotional states, and at the same time enhancing the applicability and reliability of the system in actual application scenarios.

[0038] Figure 5 It is a multimodal knowledge distillation method that promotes information interaction between different modalities through distillation, and different models mutually form a teacher network and a student network. Furthermore, by learning the influence degree of different modal signals on the emotional state, the accurate recognition of the emotional state of elderly users is realized. Table 6 shows the 5-classification results of the multimodal emotion monitoring dataset based on wearable devices.

[0039] Table 6 Emotion Classification Results Based on the WESAD Database

[0040]

[0041] The above table shows the performance of emotion recognition and classification under different time-frequency feature extractions. When using the Hanming window (denoising), the accuracy rate (78.52%) and recall rate (98.14%) are the highest, but the precision rate (30.32%) and F1 score (46.32%) are relatively low, indicating that there are certain false alarms in the model. The performance of the Rectangular window is close to that of the Hanming window. When the window length increases from 128 to 256, the accuracy rate and recall rate decrease, but the precision rate and F1 score increase slightly. Generally speaking, the Hanming window (denoising) shows the best overall performance.

[0042] The combination of these figures will effectively illustrate the specific implementation scheme of the present invention and its practical applications, demonstrating its innovation and practicality in the assessment of cognitive impairment.

[0043] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. An elderly emotion assessment system based on multimodal physiological and behavioral signals, characterized by: It includes a multimodal signal acquisition module, an emotion inducing task paradigm module and a multimodal feature processing module; The multimodal signal acquisition module is used to collect multimodal signals of elderly users, including physiological signals and behavioral signals; the emotion induction task paradigm module guides the target emotions through multimedia content; the multimodal feature processing module adopts a multimodal data fusion model based on the attention mechanism to extract emotion-related features from multimodal signals to analyze the different contributions of each modal signal to the emotional state, and analyze the synergistic relationship between multimodal signals to generate a unified emotion feature representation, and accurately identify the emotional state of elderly users by analyzing the degree of influence of different modal signals on the emotional state.

2. The elderly emotion assessment system based on multimodal physiological and behavioral signals as claimed in claim 1, characterized in that: The emotion induction task paradigm module achieves efficient induction of target emotions through multimedia content combined with emotion induction scenarios, wherein the multimedia content includes video, audio and text, and the target emotions include six types: happiness, tension, boredom, anger, calmness and sadness. The emotion induction task paradigm module quantifies the intensity of emotional response based on valence validity and activation measurement based on SAM, and provides unified testing conditions for emotional states based on the standard process of psychological experiments; wherein the standard process includes clarifying research questions, designing experiments, collecting data, analyzing results and writing reports. The entire process must strictly follow ethical standards and verify hypotheses through controlling variables and statistical analysis.

3. The elderly emotion assessment system based on multimodal physiological and behavioral signals as claimed in claim 2, characterized in that: The physiological signals include electrocardiogram and electroencephalogram signals, and the behavioral signals include speech and facial expression signals. The multimodal signal acquisition module constructs an accurate acquisition strategy and data standardization process through the synchronous acquisition of multimodal signals of electrocardiogram, electroencephalogram, speech and facial expression; wherein the acquisition strategy includes equipment calibration before signal acquisition and real-time monitoring of data quality; the data standardization process is a preprocessing method after signal acquisition, including filtering, downsampling and normalization, which is used to ensure high stability and high quality of the signal; the multimodal signal acquisition module explores the dynamic correlation of signals between time series by revealing the time domain and frequency domain characteristics and time-varying characteristics of different modal signals in emotional states, wherein dynamic correlation refers to the characteristic that the temporal correlation of two or more signals changes with time, emphasizing that the correlation is dynamic rather than static.

4. The elderly emotion assessment system based on multimodal physiological and behavioral signals as claimed in claim 3, characterized in that: The emotion induction task paradigm module adopts a multi-stage test of emotional state, including a pre-test stage and a formal test stage. The pre-test stage guides users to familiarize themselves with the experimental process by playing short videos and lightweight tasks, thereby reducing the unfamiliarity and operational difficulty in the early stage of the experiment. The formal test stage adopts six emotion induction scenarios, including visual and auditory stimulation, combined with dynamic recording of physiological and behavioral signals of elderly users, to deeply analyze the correlation of emotional state in multimodal signals, that is, the characteristics of emotional state changes over time and individual differences.

5. The elderly emotion assessment system based on multimodal physiological and behavioral signals as claimed in claim 4, characterized in that: The emotion-inducing scenario includes a fixed video playback device position, a noise-isolated laboratory environment, a screen height that is level with the user’s line of sight, and a precisely controlled audio and video content playback speed; By controlling the experimental conditions, the controllability and consistency of the emotion induction task can be ensured, and the relationship between the response intensity of multimodal signals of elderly users and the effectiveness of the emotion induction task can be quantified based on SAM.

6. The elderly emotion assessment system based on multimodal physiological and behavioral signals as claimed in claim 5, characterized in that: Physiological and behavioral signal feature encoders are used to study the specific contributions of different modal signals in emotion assessment. Among them, a unified transformer framework ESET-net is proposed to simultaneously model multimodal signals. For physiological and speech signals, they are first converted into spatiotemporal spectrograms through short-time Fourier transform. For behavioral signals, facial features are extracted through pre-trained convolutional neural networks, and then multimodal fusion is performed on the time-frequency features between different modal signals.

7. The elderly emotion assessment system based on multimodal physiological and behavioral signals as claimed in claim 1, characterized in that: A multimodal data fusion model based on attention mechanism is used to perform collaborative analysis and dynamic weighting of different modal signals; The model consists of N identical layer stacks, each layer consists of a multi-head self-attention mechanism and a fully connected feed-forward network; each sub-layer uses residual connections and layer normalization: LayerNorm(x+Sublayer(x))(1) Where: Sublayer(x) is the function implemented by the sublayer itself; LayerNorm(x) is the layer normalization function; the attention mechanism is to map the query vector Q and the key-value pair vector KV to the output value: Where: d k is the dimension of query and key; multi-head attention enables the model to jointly pay attention to information in different representation subspaces at different positions: Where: W O Output weights for multi-head attention; are the corresponding weight matrices of Q, K, and V; thus, the model optimizes the feature extraction and fusion strategies by dynamically adjusting the weights of different modal signals and learning the interdependence between multimodal signals to generate high-precision emotional feature representation.

8. The elderly emotion assessment system based on multimodal physiological and behavioral signals as claimed in claim 1, characterized in that: The knowledge distillation method is used to promote information interaction between different modalities. Different models form a teacher network and a student network. The student network has a similar architecture to the teacher network. The method of trimming the model is to halve the number of layers and initialize the student network layer from the teacher network layer. The model alternates between a fully copied layer and an ignored layer, and preferentially copies the top or bottom layer. The student model is trained using the true label of the sample and the teacher network supervision to imitate the teacher network and get as close to its output as possible. By learning the influence of different modal signals on emotional states, accurate identification of the emotional states of elderly users can be achieved.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method and system based on confidence fusion

    CN117591967A

  • Emotion analysis method and system based on multi-modal fusion

    CN119272224A

  • Zipper slide and puller structure

    KR102756435B1

  • Determining a psychological state of a subject

    US20040210159A1

Cited By

  • Facial composite emotion recognition method and system based on scene induction and medium

    CN121096005A

  • Emotion recognition method based on online cross-modal knowledge distillation

    CN121960702A

  • Robot with emotion recognition and monitoring functions

    CN122208147A