Emotional state evaluation method and system based on multiple events and multiple tasks

By analyzing EEG signals and EMG signals, combining psychological test results, a multimodal fusion model and multitasking classifier are constructed, which solves the problem that the existing technology cannot directly reflect the internal cognitive state and emotional activities, and achieves a more accurate and comprehensive assessment of the emotional state of college students, improving the objectivity and efficiency of the assessment.

CN119943407AActive Publication Date: 2025-05-06QINXUN ZHIXIN TECHNOLOGY (ZHEJIANG) CO LTD

Patent Information

Application Number
CN202510123241.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-06
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

The existing deep learning-based mental health assessment methods rely too much on behavioral data, voice data and video data, and cannot directly reflect internal cognitive states and emotional activities. They have limitations such as inconsistent performance with internal states, lack of in-depth analysis of cognitive processes, and insufficient expression of emotional information.

Method used

By comprehensively analyzing EEG signals and EMG signals, a multimodal fusion model is constructed, and an emotional state label is generated based on psychological test results, and a multi-task classifier is constructed to achieve a more objective, comprehensive and accurate assessment of the emotional state of college students.

Benefits of technology

It improves the objectivity and comprehensiveness of mental health assessment, enhances the accuracy and interpretability of analysis, improves the efficiency and convenience of assessment, and provides a scientific basis for mental health education and intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943407A_ABST
    Figure CN119943407A_ABST
Patent Text Reader

Abstract

The invention provides an emotional state assessment method and system based on multiple events and multiple tasks, and relates to the technical field of mental health assessment and artificial intelligence. The method specifically comprises the steps of collecting physiological data of an evaluation object, performing psychological testing on the evaluation object, and generating an emotional state label of the evaluation object according to a psychological testing result; respectively constructing an electroencephalogram signal neural network model and an electromyographic signal neural network model, and performing fusion to obtain a multi-modal fusion model; constructing a multi-task classifier according to the emotional state label; and inputting the physiological data of the evaluation object into the multi-modal fusion model for feature fusion, generating a multi-modal joint feature representation, inputting the multi-modal joint feature representation into the multi-task classifier for evaluation, and obtaining an emotional state evaluation result of the evaluation object. According to the method, the defect of insufficient analysis precision of a single data source is overcome, the psychological health condition of the college students can be evaluated more accurately and comprehensively, and powerful support is provided for early discovery and intervention of psychological problems such as depression and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of mental health assessment and artificial intelligence technology, and in particular to a method and system for emotional state assessment based on multiple events and multiple tasks. Background Art

[0002] In today's fast-paced and competitive social environment, the mental health issues of college students are increasingly receiving widespread attention. Mental health problems such as depression, anxiety, and social phobia not only affect their learning efficiency and quality of life, but are also likely to lead to a series of serious consequences, such as the occurrence of some extreme behaviors. However, early detection and timely intervention of mental health problems among college students is an extremely challenging task. Traditional mental health assessment methods have many shortcomings, such as being time-consuming and labor-intensive, relying on subjective reports, and being limited by medical resources, which makes it difficult to meet the needs of efficient and accurate assessment.

[0003] With the rapid development of artificial intelligence technology, its application in the field of college students' mental health assessment has broad prospects. Deep learning models can objectively analyze multi-source data from multiple dimensions, comprehensively capture clues of individual emotions, moods and psychological states, and improve the accuracy of assessments. However, existing mental health assessment methods based on deep learning are too dependent on explicit information such as behavioral data, voice data and video data. These data reflect people's external performance and cannot directly reflect the internal cognitive state and emotional activities. There are limitations such as inconsistency between performance and internal state, lack of in-depth analysis of cognitive processes, and insufficient expression of emotional information. Summary of the invention

[0004] In view of the above-mentioned shortcomings of the prior art, the present invention proposes an emotional state assessment method and system based on multi-event and multi-task by comprehensively analyzing EEG signals and EMG signals, aiming to further improve the accuracy of the analysis by obtaining more objective, comprehensive and reliable information, so as to help college students identify potential mental health problems as early as possible, so as to take effective measures to address and improve them.

[0005] The first aspect of the present invention provides a method and system for emotional state assessment based on multi-event and multi-task, the method comprising the following steps:

[0006] Step 1: Collect the physiological data of the evaluation object, conduct psychological tests on the evaluation object, and generate the emotional state label of the evaluation object according to the psychological test results;

[0007] Step 2: Preprocess the physiological data of the evaluation object and construct a multimodal dataset, and divide the multimodal dataset into a training set and a test set according to a set ratio;

[0008] Step 3: Construct the EEG neural network model and the EMG neural network model respectively;

[0009] Step 4: Fuse the EEG neural network model of the EEG signal and the EMG neural network model of the EMG signal to obtain a multimodal fusion model;

[0010] Step 5: Build a multi-task classifier based on the emotional state labels;

[0011] Step 6: Use the training set to iteratively train the multimodal fusion model and the multi-task classifier. When the total loss function of the multi-task classifier converges or reaches the preset number of training rounds or meets the preset classification target, the iteration is stopped to obtain the trained multimodal fusion model and multi-task classifier.

[0012] Step 7: Use the test set to verify the trained multimodal fusion model and multi-task classifier to obtain the final multimodal fusion model and multi-task classifier;

[0013] Step 8: For any evaluation object, the physiological data of the evaluation object is collected and input into the final multimodal fusion model for feature fusion, and then the generated multimodal joint feature representation is input into the multi-task classifier for evaluation to obtain the emotional state evaluation result of the evaluation object;

[0014] The method for collecting physiological data in step 1 is: performing multimodal physiological and behavioral tests on the evaluation subject, and recording the physiological data of the evaluation subject in real time; wherein the process of the multimodal physiological and behavioral tests is: wearing an electroencephalogram sensor and an electromyography sensor on the evaluation subject to record the electroencephalogram signal and electromyography signal of the evaluation subject respectively, and using the electroencephalogram signal and electromyography signal as the physiological data of the evaluation subject;

[0015] Conducting several behavioral tests on the subject in turn, and recording the subject's physiological data in real time during each behavioral test; each behavioral test is any one of a self-introduction test, a text reading test, or a gait test;

[0016] The self-introduction test is as follows: the subject of evaluation introduces himself according to a specified topic; the text reading test is as follows: the subject of evaluation reads a specified text; the gait test is as follows: the subject of evaluation walks in an orderly manner in a specified area according to a predetermined route and direction;

[0017] The psychological test is: using a personality scale, a general happiness scale, a depression screening scale and a depressive symptom scale to conduct psychological tests on the evaluation object, and using the questionnaire scores of each scale as the emotional state label of the evaluation object;

[0018] The emotional state label includes: five personality dimensions of the assessed subject and the psychological state of the assessed subject;

[0019] The step 2 further comprises:

[0020] Step 2.1: Based on the timestamp information recorded in the physiological data, the start and end time points of the evaluation subject during different activity stages are marked, and the physiological data is segmented into several segments according to the marked start and end time points to obtain the segmented physiological data;

[0021] Step 2.2: downsampling the segmented physiological data to obtain downsampled physiological data;

[0022] Step 2.3: filtering the downsampled physiological data to obtain filtered physiological data; wherein a fourth-order Butterworth filter is used to band-pass filter the EEG signal in the physiological data; and a band-pass filter is used to band-pass filter the EMG signal in the physiological data;

[0023] Step 2.4: Extract features from the filtered physiological data to obtain time domain features and frequency domain features of the physiological data;

[0024] Step 2.5: for any physiological data, the time domain features and frequency domain features of the physiological data are combined into a feature vector of the physiological data, a behavioral test is regarded as an event, and the feature vectors of all physiological data in each event are regarded as a multimodal data sequence of the event, and a multimodal data sequence of all events is used to construct a multimodal data set; wherein the multimodal data set includes: an EEG data set of an electroencephalogram signal and an EMG data set of an electromyography signal;

[0025] Step 2.6: Divide the multimodal dataset into training set and test set according to the preset ratio;

[0026] The EEG neural network model for electroencephalogram signals and the EMG neural network model for electromyography signals in step 3 are both single-modality neural network models;

[0027] The single modality neural network model is: using a convolutional neural network (CNN) model to extract local features from an input multimodal data sequence, wherein the convolutional neural network (CNN) model is a layer of convolution operation;

[0028] The feature representation extracted by the convolutional neural network CNN model is input into the multi-head self-attention mechanism for processing, and the attention score of each physiological data in the multimodal data sequence is calculated. The attention score of the physiological data is used to calculate the value vector V i Perform weighted summation to obtain weighted feature representation of physiological data;

[0029] For any event, the weighted feature representation of the physiological data in the event is residually connected with the multimodal data sequence of the event, and the result of the residual connection is layer-normalized. The normalized features are then nonlinearly transformed through a feedforward network to generate a feature sequence of the event.

[0030] The feature sequence of the event is input into the bidirectional long short-term memory network Bi-LSTM, and the Bi-LSTM is used to process the forward time dependency and reverse time dependency in the feature sequence respectively, and the hidden state of the last time step in the feature sequence is taken as the output of Bi-LSTM to obtain the feature vector of the event;

[0031] The process of fusing the EEG neural network model of the electroencephalogram signal and the EMG neural network model of the electromyography signal described in step 4 is divided into two parts: a single-modal fusion process based on a channel attention module and a multi-modal fusion process based on an interactive attention mechanism;

[0032] The unimodal fusion process based on the channel attention module is as follows: constructing a joint feature representation of multiple events using the feature vectors output by the EEG neural network model of the electroencephalogram signal and the EMG neural network model of the electromyography signal, respectively, and performing global average pooling and global maximum pooling on the joint feature representation of multiple events, respectively, performing dimension reduction and activation processing on the obtained global average pooling features and global maximum pooling features through a two-layer fully connected network, obtaining the attention weight of the event and then fusing it with the joint feature representation of multiple events to generate a weighted multi-event feature representation;

[0033] For the EEG neural network model, the weighted multi-event feature representation obtained after the feature vector output by the model is fused by single modality is used as the multi-event fusion feature H of the EEG signal. EEG For the EMG neural network model, the weighted multi-event feature representation obtained after the single-modal fusion of the feature vector output by the model is used as the multi-event fusion feature H of the EMG signal. EMG ;

[0034] The multimodal fusion process based on the interactive attention mechanism is as follows: a double-branch cross attention module is used to calculate the multi-event fusion feature H of the EEG signal respectively. EEG and the multi-event fusion feature H of electromyographic signals EMG The feature correlation between them is used to obtain the attention output from the EEG modality to the EMG modality. and attention output from EMG modality to EEG modality Then by and Splice and generate EEG features respectively and electromyographic characteristics And integrate multimodal joint feature representation;

[0035] Furthermore, the specific content of the single-modal fusion process based on the channel attention module is:

[0036] The feature vectors of each event in the multimodal dataset are fused along the channel dimension to construct a joint feature representation of multiple events;

[0037] The joint feature representation of multiple events is represented by global average pooling and global maximum pooling respectively, and the global average pooling feature H is obtained. t,avg And the global maximum pooling feature H t,max , the first layer of fully connected network is used to respectively average the global pooling features H t,avg And the global maximum pooling feature H t,max Perform dimensionality reduction processing, and then use the second-layer fully connected network to activate the reduced global average pooling features and global maximum pooling features respectively, and then merge the activated global average pooling features and global maximum pooling features to obtain the attention weight of the event;

[0038] The attention weight of the event is expressed as:

[0039] α event =Sigmoid(W1(W0(H t,avg ))+W1(W0(H t,max )))

[0040] where α event Represents the attention weight of the event; Sigmoid represents the activation function; W1 is the recovery matrix; W0 is the dimension reduction matrix;

[0041] Use element-by-element multiplication to convert the event's attention weight α event The joint feature representation H of multiple events t Fusion is performed to obtain the weighted multi-event feature representation H event ;

[0042] Furthermore, the dual-branch cross attention module is: one attention branch uses the multi-event fusion feature H of the EEG signal EEG To query Query, the multi-event fusion feature H of the electromyographic signal EMG For the key Key and value Value, calculate H EEG With H EMG The feature correlation between EEG With H EMG Feature correlation between EEG modality and EMG modality attention output It is expressed as:

[0043]

[0044] Where Softmax represents the normalized exponential function; Indicates the multi-event fusion feature H of the EEG signal EEG Map to Query space; Indicates the multi-event fusion feature H of the EEG signal EEG Mapped to Key space; V t (e→m) Indicates the multi-event fusion feature H of the EEG signal EEG Mapped to the Value space; e represents the EEG modality; m represents the EMG modality; d represents the feature dimension of Query and Key, which is used to prevent the attention score from being too large in high-dimensional space, resulting in gradient disappearance or numerical instability;

[0045] Another attention branch is to use the multi-event fusion feature H of the electromyographic signal EMG To query Query, the multi-event fusion feature H of EEG signal EEG For the key Key and value Value, calculate H EMG With H EEG The feature correlation between EMG With H EEG Feature correlation between EMG and EEG modalities It is expressed as:

[0046]

[0047] in Indicates the multi-event fusion feature H of the electromyographic signal EMG Map to Query space; Indicates the multi-event fusion feature H of the electromyographic signal EMG Mapped to Key space; V t (m→e) Indicates the multi-event fusion feature H of the electromyographic signal EMG Map to Value space;

[0048] The method for constructing the multi-task classifier in step 5 is as follows: define two classification tasks according to the emotional state label, wherein one of the classification tasks is a three-level classification problem of psychological state; define the three-level classification problem of psychological state as a three-level classification task, including three categories: normal, mild problem and severe problem; the other classification task is a two-level classification problem of five personalities; define the two-level classification problem of five personalities as five two-level classification tasks, including five categories: openness, conscientiousness, extroversion, agreeableness and neuroticism;

[0049] A classifier is defined for each classification task, for evaluating the emotional state of the physiological data corresponding to the multimodal joint feature representation according to the input multimodal joint feature representation;

[0050] The total loss function L of the multi-task classifier described in step 6 total for:

[0051] L total =βL m +γ(L1+L2+L3+L4+L5)

[0052] Where L m represents the cross entropy loss of the three-level classification problem of psychological state; β represents the weight coefficient of the three-level classification problem of psychological state; L1, L2, L3, L4, L5 represent the cross entropy loss of the two-level classification problem of five personalities; γ represents the weight coefficient of the two-level classification problem of five personalities;

[0053] Where L m , the cross entropy losses of L1, L2, L3, L4 and L5 are all expressed as:

[0054]

[0055] Where L represents the cross entropy loss; N represents the total number of data in the training set; C represents the total number of categories; y lj is the true label of the lth sample in the jth category; is the evaluation result of the lth sample;

[0056] The second aspect of the present invention proposes an emotional state assessment system based on multi-event multi-task, which is used to implement the emotional state assessment method based on multi-event multi-task, and the system includes: a data acquisition module, a data processing module, a feature fusion module and a classification assessment module;

[0057] The data acquisition module is used to collect physiological data of the evaluation object and obtain the emotional state label of the evaluation object, and transmit the collected physiological data to the data processing module, and transmit the obtained emotional state label to the classification evaluation module;

[0058] The data processing module is used to perform segmentation, downsampling, filtering, feature extraction and splicing preprocessing on the received physiological data in sequence, and construct a multimodal data set using the preprocessed physiological data, and transmit the multimodal data set to the feature fusion module;

[0059] The feature fusion module is used to extract features of a single modality from data of different modalities in a multimodal data set, and sequentially perform single-modality fusion based on a channel attention module and multimodal fusion based on an interactive attention mechanism on the extracted feature vectors to generate a multimodal joint feature representation and transmit it to the classification evaluation module;

[0060] The classification and evaluation module is used to construct a classifier according to the emotional state label, and input the multimodal joint feature representation into the classifier to generate an emotional state evaluation result of the physiological data corresponding to the multimodal joint feature representation.

[0061] The beneficial effects of adopting the above technical solution are:

[0062] 1) Improved objectivity and comprehensiveness: The method of the present invention first obtains users' psychological state and emotional state label data such as personality through a scale survey, and at the same time uses electroencephalographic sensors and electromyographic sensors to record their nerve and muscle activity signals to form multimodal data. By collecting and integrating multimodal information such as physiological data, the method of the present invention can comprehensively and objectively evaluate the psychological status of college students and avoid the limitations of a single subjective test.

[0063] 2) Enhanced analysis accuracy and interpretability: The method of the present invention eliminates noise and extracts features by preprocessing multimodal data, and then uses deep learning technology to establish single-modal models respectively, and fuses the outputs of different single-modal models to construct a multimodal fusion model. The method of the present invention uses deep learning technology to perform feature representation learning and modeling on multimodal data, thereby improving the accuracy of the analysis. At the same time, through interpretability techniques such as gradient / attention visualization, it is possible to explain the contribution of different modal data to the prediction results, thereby improving the transparency of the model.

[0064] 3) Improved evaluation efficiency and convenience: The method of the present invention is based on a data-driven model evaluation method, which reduces the workload of manual evaluation and improves the automation level and efficiency of the evaluation process. At the same time, there is no need to collect privacy information, which is convenient for large-scale promotion and application.

[0065] 4) Basis for mental health education and intervention: Based on the analysis results output by the model, the method of the present invention can understand the overall psychological status of the college students and provide a basis for the school to formulate targeted mental health education content. At the same time, for high-risk groups, timely identification and intervention counseling can be carried out.

[0066] In summary, the method of the present invention can comprehensively analyze the psychological state and personality traits of college students by constructing a multimodal fusion model, and deeply explore the intrinsic connection between multimodal data and mental health status. The method of the present invention overcomes the defect of insufficient analysis accuracy of a single data source, can more accurately and comprehensively evaluate the mental health status of college students, and provide strong support for early detection of psychological problems and subsequent intervention, and has broad application prospects in the fields of mental health, medical care, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is a flow chart of an emotional state assessment method based on multi-event and multi-task in this implementation mode;

[0068] Figure 2 FIG. 1 is a diagram of the experimental scheme of multimodal physiological and behavioral testing in this embodiment;

[0069] Figure 3 A diagram for constructing a multimodal fusion model in this implementation;

[0070] Figure 4 This is a structural diagram of an emotional state assessment system based on multiple events and multiple tasks in this embodiment. DETAILED DESCRIPTION

[0071] In order to facilitate the understanding of the present application, the specific embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thoroughly understood.

[0072] In this embodiment, a method for assessing emotional state based on multiple events and multiple tasks is provided. Figure 1 As shown, the method comprises the following steps:

[0073] Step 1: Collect the physiological data of the evaluation object, conduct psychological tests on the evaluation object, and generate the emotional state label of the evaluation object according to the psychological test results.

[0074] In this embodiment, a multimodal data collection method is used to comprehensively record the user's physiological and behavioral responses in various situations, including physiological data such as electroencephalogram (EEG) and electromyogram (EMG).

[0075] The physiological data collection method is: performing multimodal physiological and behavioral tests on the evaluation object, and recording the physiological data of the evaluation object in real time.

[0076] The process of the multimodal physiological and behavioral test is as follows: the evaluation subject is equipped with an electroencephalogram sensor and an electromyography sensor to respectively record the electroencephalogram signal and electromyography signal of the evaluation subject, and the electroencephalogram signal and electromyography signal are used as the physiological data of the evaluation subject.

[0077] The evaluation object is subjected to several behavioral tests in sequence, and the physiological data of the evaluation object is recorded in real time during each behavioral test; each behavioral test is any one of the self-introduction test, text reading test or gait test.

[0078] In this embodiment, if Figure 2 As shown in the figure, the multimodal physiological and behavioral test includes: test preparation stage, self-introduction test, text reading test and gait test. In the test preparation stage, the evaluation subject needs to wear EEG sensors and EMG sensors to record the evaluation subject's nerve and muscle activity signals. The coordinated operation of these devices is aimed at comprehensively and accurately collecting the subject's physiological signals during the experiment, laying the foundation for subsequent data analysis and modeling.

[0079] The self-introduction test is as follows: the assessment subject introduces himself / herself according to a prescribed topic.

[0080] In this embodiment, the content of the self-introduction is required to include: sharing troubles, happy and sad things, and discussing the areas where one wants to improve and the areas where one is satisfied.

[0081] The text reading test is as follows: the evaluation subject reads aloud a designated text.

[0082] In this embodiment, the evaluation subject is required to read aloud the designated texts "The North Wind and the Sun" and "The Wolf and the Lamb".

[0083] The gait test is as follows: the subject of evaluation walks in an orderly manner in a designated area according to a predetermined route and direction.

[0084] The psychological test is: using a personality scale, a general well-being scale (General Well-Being Schedule Scale, GWB), a depression screening scale and a depressive symptom scale to conduct psychological tests on the evaluation subjects, and using the questionnaire scores of each scale as the emotional state label of the evaluation subject.

[0085] In this embodiment, a personality scale is used to assess personality traits, such as the Big-5 personality scale (simplified version), which can assess five main personality dimensions, including openness, conscientiousness, extroversion, agreeableness, and neuroticism. The General Well-Being Schedule Scale (GWB) is used to measure the user's level of subjective well-being. A depression screening scale is used to assess the user's depressive tendencies, such as the PHQ-9 depression scale. The depressive symptom scale is used to further refine the quantitative analysis of depressive symptoms, such as the BDI depression scale. After scale evaluation, quantitative indicators reflecting individual depressive tendencies, personality traits, and subjective well-being levels can be obtained, and the questionnaire scores of each scale are used as emotional state labels of the evaluation objects to provide supervision information for subsequent deep learning models.

[0086] Step 2: Preprocess the physiological data of the evaluation object and construct a multimodal dataset, and divide the multimodal dataset into a training set and a test set according to a set ratio.

[0087] In this embodiment, the collected EEG signals and EMG signals are standardized to improve signal quality, reduce noise and redundant information, and lay the foundation for the subsequent deep learning model construction.

[0088] The step 2 further comprises:

[0089] Step 2.1: Based on the timestamp information recorded in the physiological data, the start and end time points of the evaluation subject during different activity stages are marked, and the physiological data is segmented into several segments according to the marked start and end time points to obtain segmented physiological data.

[0090] In this embodiment, during the collection of EEG signals and EMG signals, since the evaluation subject will go through different activity stages, including task execution, rest or waiting, and there are some activity periods irrelevant to the test in these stages, it is necessary to segment the recorded physiological data, determine the start and end time points of each activity stage according to the timestamp information recorded during data collection, and then divide the continuous data into corresponding segments.

[0091] Step 2.2: Downsample the segmented physiological data to obtain downsampled physiological data.

[0092] In this embodiment, the collected EEG signals and EMG signals usually have a high sampling rate. Although the high sampling rate can retain more signal details, it will significantly increase the storage and processing costs. Therefore, it is necessary to downsample the EEG signals and EMG signals to 200Hz. This sampling rate can retain sufficient signal details while reducing the amount of calculation for subsequent analysis and processing.

[0093] Step 2.3: Filter the downsampled physiological data to obtain filtered physiological data; wherein a fourth-order Butterworth filter is used to band-pass filter the EEG signal in the physiological data; and a band-pass filter is used to band-pass filter the EMG signal in the physiological data.

[0094] In this embodiment, the original signal often contains high-frequency noise such as environmental interference and low-frequency drift such as baseline drift, which can significantly reduce the quality of the signal. Therefore, after downsampling, the physiological data needs to be filtered. For EEG signals, a fourth-order Butterworth filter is used for bandpass filtering to filter out five frequency bands: delta (1-3 Hz), theta (4-7 Hz), alpha (8-13 Hz), beta (14-30 Hz), and gamma (31-50 Hz). For electromyographic signals, a 1-45 Hz bandpass filter is used to filter out low-frequency interference such as high-frequency noise and baseline drift.

[0095] Step 2.4: Extract features from the filtered physiological data to obtain time domain features and frequency domain features of the physiological data.

[0096] In this embodiment, feature extraction includes two categories: time domain features and frequency domain features. The time domain features include the mean, variance, peak-to-peak value, etc. of the signal, which can reflect the basic statistical properties of the signal; the frequency domain features extract the power spectrum, frequency band energy, etc. through fast Fourier transform FFT, which can reveal the frequency distribution of the signal.

[0097] Step 2.5: For any physiological data, the time domain features and frequency domain features of the physiological data are combined into a feature vector of the physiological data. A behavioral test is regarded as an event, and the feature vectors of all physiological data in each event are regarded as the multimodal data sequence of the event. The multimodal data sequences of all events are used to construct a multimodal data set.

[0098] The multimodal data set includes: an EEG data set and an EMG data set;

[0099] The multimodal dataset is represented as:

[0100]

[0101] Where X represents a multimodal dataset. If the multimodal dataset only represents an EEG dataset, X is represented as X EEG ; If the multimodal dataset only represents the EMG dataset, then X is represented as X EMG ; A multimodal data sequence representing the first event E1; A multimodal data sequence representing the second event E2; Represents the i-th event E i Multimodal data sequences; Indicates the nth event E n is a multimodal data sequence; n is the total number of events, and in this embodiment, n=4.

[0102] In this embodiment, a has four events, which respectively represent self-introduction, reading the text "The North Wind and the Sun", reading the text "The Wolf and the Lamb", and gait test. The multimodal dataset contains physiological data in two physiological modalities, namely, EEG and EMG. EEG represents the physiological data under the EEG physiological mode, X EMG Represents physiological data under the EMG physiological modality.

[0103] Step 2.6: Divide the multimodal dataset into training set and test set according to the preset ratio.

[0104] In this embodiment, the samples are randomly divided into a training set and a test set according to a ratio of 4:1. The training set is used to train the model, and the test set is used to evaluate the model performance.

[0105] Step 3: Construct the EEG neural network model and the EMG neural network model respectively.

[0106] In this embodiment, deep learning technology is used to construct corresponding analysis models for EMG and EEG modal data respectively. For the construction of the neural network model, it mainly includes two core links, namely feature representation learning and modality-specific modeling and model training optimization and evaluation interpretation; wherein the feature representation learning and modality-specific modeling include: using self-supervised learning methods to allow the model to learn high-quality feature representations from preprocessed data and capture the inherent patterns and laws of the data. At the same time, according to the attributes of different modal data, a reasonable network structure is designed, such as using convolution / attention sequence models for time series modalities and convolution / graph convolution networks for spatial modalities. The network output is adjusted according to downstream tasks, and mechanisms such as multi-head attention and gated recurrent units are introduced to enhance the model's expression ability. The model training optimization and evaluation interpretation include: designing reasonable supervised training targets, such as classification cross entropy loss, regression mean square error loss, etc., and using efficient optimization algorithms such as Adam for training. Regularization strategies such as L1, L2 regularization, dropout, etc. are introduced to prevent overfitting, and techniques such as gradient clipping and cyclic regularization are used to improve training stability. The generalization ability of the model is fully evaluated on the reserved data set, using indicators such as accuracy, F1, and root mean square error, and cross-validation is performed to evaluate robustness. The internal behavior of the model is analyzed through techniques such as gradient / attention visualization, and the model is explained in combination with domain knowledge to improve model transparency.

[0107] The EEG neural network model for electroencephalogram signals and the EMG neural network model for electromyography signals are both single-modality neural network models.

[0108] The single modality neural network model is: using a convolutional neural network (CNN) model to extract local features from an input multimodal data sequence, wherein the convolutional neural network (CNN) model is a layer of convolution operation, expressed as:

[0109]

[0110] in It represents the feature representation obtained by extracting local features from the multimodal data sequence of the i-th event; Sigmoid is an activation function used to map the convolution result to the range of (0,1); Conv1D represents a one-dimensional convolution operation.

[0111] In this embodiment, the input multimodal data sequence includes features of two modalities: feature vectors of EEG signals and feature vectors of EMG signals. The features of the two modalities are respectively input into two independent convolutional neural network (CNN) models for local feature extraction. Both CNN models are one-layer convolution operations, and the size of their convolution kernels is 3 and the step size is 1, so as to extract pattern information of different scales and different abstract levels in the input features, and finally obtain a higher-level and more abstract feature representation.

[0112] The feature representation extracted by the convolutional neural network CNN model is input into the multi-head self-attention mechanism for processing, and the attention score of each physiological data in the multimodal data sequence is calculated. The attention score of the physiological data is used to calculate the value vector V i Perform weighted summation to obtain the weighted feature representation of the physiological data.

[0113] The calculation method of the attention score is:

[0114] MultiHead(Q,K,V)=[e1,...,e r ]W K (3)

[0115]

[0116] Where MultiHead(Q,K,V) represents the output of the multi-head attention mechanism; Q, K, and V represent the query feature matrix, key feature matrix, and value feature matrix, respectively; e1 represents the output of the first attention head in the multi-head attention mechanism; e r represents the output of each attention head of the rth in the multi-head attention mechanism; e irepresents the output of each attention head of the i-th in the multi-head attention mechanism; Softmax is the activation function, which is used to normalize the dot product similarity of all keys and queries to obtain the attention weight of each key; d is the dimension of the key vector, and Indicates scaling the dot product result to prevent numerical instability when the dimension is too large; Q i Represents the query vector, which is used to compare with the key vector K i Perform dot product to calculate attention weight; K i is the key vector, which is used to represent the features of each position in the input sequence and matches the query vector; V i is a value vector, which is used to represent the basis for the final output obtained through the weighted attention mechanism.

[0117] In this embodiment, the feature representations from the two CNN models are processed separately through the multi-head self-attention mechanism. The multi-head self-attention mechanism can capture the correlation between features at different positions in the input sequence, thereby better modeling long-term dependencies. The multi-head self-attention mechanism calculates an attention score for each time step, which is used to weight the importance of features at different positions, thereby improving the expressiveness of the model.

[0118] For any event, the weighted feature representation of the physiological data in the event is residually connected with the multimodal data sequence of the event, and the result of the residual connection is layer-normalized. The normalized features are then nonlinearly transformed through a feedforward network to generate a feature sequence of the event.

[0119] The feature sequence of the event is input into the bidirectional long short-term memory network Bi-LSTM, and the Bi-LSTM is used to process the forward time dependency and reverse time dependency in the feature sequence respectively. The hidden state of the last time step in the feature sequence is taken as the output of Bi-LSTM to obtain the feature vector of the event.

[0120]

[0121] in Represents the i-th event E i The feature vector of ; t represents the time step; represents bidirectional LSTM; h t-1 and h t+1 Represent the hidden states at time steps t-1 and t+1 respectively; Indicates event E i The feature sequence obtained after being processed by the multi-head self-attention mechanism.

[0122] In this embodiment, the two output sequences processed by the multi-head self-attention mechanism are sent to two long short-term memory networks (LSTM) for further processing. For each sequence input into the LSTM, the hidden state output of the last time step is taken as the final output of the LSTM, so that the model can integrate the semantic information of the entire sequence and further enhance the expressiveness of the features. LSTM selectively retains or discards information through its unique gating structure, including input gate, forget gate and output gate, so as to effectively handle dependencies in long time series. By combining the forward and backward hidden states, Bi-LSTM can capture long-term dependency information in event sequences. The event representations of different modalities are processed by the multi-head attention mechanism in the previous step, and then represented by residual connection, layer normalization and feedforward network. By combining the forward and backward hidden states, Bi-LSTM can capture the long-term dependency information in the event sequence.

[0123] Step 4: Fuse the EEG neural network model and the EMG neural network model to obtain a multimodal fusion model.

[0124] In this embodiment, the hidden representations of the two modal corresponding models are fused to construct a multimodal fusion model, such as Figure 3 As shown in the figure, it mainly includes the following two key steps: 1) Modal alignment and feature fusion, aligning data streams from different modalities to the same time scale, dealing with problems such as data loss and inconsistent sampling rates. Then the feature representations of different modalities are fused. Simple concatenation and weighted summation can be used at the feature layer, or attention mechanism can be used for adaptive fusion. In addition, external memory, gated loop mechanism, etc. can be used to simulate the interaction between modalities and capture modal feedback at different time steps. 2) Model fusion and joint training optimization, fusion of the output of single modal models to build a multimodal fusion model. Design a reasonable multi-task or multi-standard loss function, and use a suitable optimization strategy to deal with modal redundancy and noise. Finally, the performance of the fusion model is evaluated on the reserved test set, and appropriate indicators are used to compare with the single modality. The modal contribution is analyzed through attention visualization, and the model interpretability is improved by combining domain knowledge.

[0125] The process of fusing the EEG neural network model of the electroencephalogram signal and the EMG neural network model of the electromyography signal is divided into two parts: a unimodal fusion process based on a channel attention module and a multimodal fusion process based on an interactive attention mechanism.

[0126] The unimodal fusion process based on the channel attention module is as follows: the feature vectors output by the EEG neural network model of the electroencephalogram signal and the EMG neural network model are used to construct a joint feature representation of multiple events, and global average pooling and global maximum pooling are performed on the joint feature representation of multiple events respectively. The obtained global average pooling features and global maximum pooling features are subjected to dimensionality reduction and activation processing through a two-layer fully connected network, and the attention weights of the events are obtained and then fused with the joint feature representation of multiple events to generate a weighted multi-event feature representation.

[0127] For the EEG neural network model, the weighted multi-event feature representation obtained after the feature vector output by the model is fused by single modality is used as the multi-event fusion feature H of the EEG signal. EEG For the EMG neural network model, the weighted multi-event feature representation obtained after the single-modal fusion of the feature vector output by the model is used as the multi-event fusion feature H of the EMG signal. EMG .

[0128] The specific content of the single-modal fusion process based on the channel attention module is:

[0129] The feature vectors of each event in the multimodal dataset are fused along the channel dimension to construct a joint feature representation of multiple events.

[0130] The joint features of the multiple events are expressed as:

[0131]

[0132] Among them, H t represents the joint feature representation of multiple events, and represents the concatenation operation of the feature representations of four events; h dim Indicates the dimension size of the event feature vector; Concat represents the concatenation operation, which is used to fuse the feature vectors of multiple events along the channel dimension; is the feature vector of the first event E1; is the i-th event E i The eigenvector of is the nth event E n The feature vector of .

[0133] The joint feature representation of multiple events is represented by global average pooling and global maximum pooling respectively, and the global average pooling feature H is obtained. t,avg And the global maximum pooling feature H t,max , the first layer of fully connected network is used to respectively average the global pooling features H t,avg And the global maximum pooling feature H t,maxAfter dimensionality reduction, the second-layer fully connected network is used to activate the reduced global average pooling features and the global maximum pooling features respectively, and then the activated global average pooling features and the global maximum pooling features are merged to obtain the attention weight of the event.

[0134] The attention weight of the event is expressed as:

[0135] α event =Sigmoid(W1(W0(H t,avg ))+W1(W0(H t,max ))) (7)

[0136] where α event Represents the attention weight of the event, which is used to dynamically reflect the importance of each event; Sigmoid represents the activation function; W1 is the recovery matrix; W0 is the dimensionality reduction matrix.

[0137] On this basis, the Channel Attention Mechanism (CAM) is introduced to optimize the fusion process of event information and dynamically adjust the contribution weight of each event in the joint representation. t Perform global average pooling and global max pooling respectively to obtain the feature H after average pooling t,avg And the feature H after maximum pooling t,max Then, the pooled feature H t,avg and H t,max Dimensionality reduction and activation processing are performed through a two-layer fully connected network.

[0138] Use element-by-element multiplication to convert the event's attention weight α event The joint feature representation H of multiple events t Fusion is performed to obtain the weighted multi-event feature representation H event .

[0139] The weighted multi-event feature representation H event for:

[0140] H event =α event ⊙H t (8)

[0141] Where ⊙ represents a fusion operation.

[0142] After the above processing, the fused unimodal representation can more accurately reflect the importance of information between events and provide higher quality feature representation for subsequent multimodal interaction fusion.

[0143] In order to capture the complex correlation between EEG signals and EMG signals, this embodiment proposes a multimodal fusion method based on an interactive attention mechanism, which achieves deep interaction and integration of the two modal features through a dual-branch cross-attention module.

[0144] The multimodal fusion process based on the interactive attention mechanism is as follows: a double-branch cross attention module is used to calculate the multi-event fusion feature H of the EEG signal respectively. EEG and the multi-event fusion feature H of electromyographic signals EMG The feature correlation between them is used to obtain the attention output from the EEG modality to the EMG modality. and attention output from EMG modality to EEG modality Then by and Splice and generate EEG features respectively and electromyographic characteristics And integrate multimodal joint feature representation.

[0145] The dual-branch cross attention module is: one attention branch uses the multi-event fusion feature H of the EEG signal EEG To query Query, the multi-event fusion feature H of the electromyographic signal EMG For the key Key and value Value, calculate H EEG With H EMG The feature correlation between EEG With H EMG Feature correlation between EEG modality and EMG modality attention output It is expressed as:

[0146]

[0147] Where Softmax represents the normalized exponential function; Indicates the multi-event fusion feature H of the EEG signal EEG Map to Query space; Indicates the multi-event fusion feature H of the EEG signal EEG Map to Key space; Indicates the multi-event fusion feature H of the EEG signal EEG Mapped to the Value space; e represents the EEG modality; m represents the EMG modality; d represents the feature dimension of Query and Key, which is used to prevent the attention score from being too large in high-dimensional space, resulting in gradient disappearance or numerical instability.

[0148] Another attention branch is to use the multi-event fusion feature H of the electromyographic signal EMGTo query Query, the multi-event fusion feature H of EEG signal EEG For the key Key and value Value, calculate H EMG With H EEG The feature correlation between EMG With H EEG Feature correlation between EMG and EEG modalities It is expressed as:

[0149]

[0150] in Indicates the multi-event fusion feature H of the electromyographic signal EMG Map to Query space; Indicates the multi-event fusion feature H of the electromyographic signal EMG Mapped to Key space; V t (m→e) Indicates the multi-event fusion feature H of the electromyographic signal EMG Mapped to the Value space.

[0151] In this embodiment, the multi-event fusion feature of EEG is and EMG multi-event fusion features As input, one branch takes EEG features as query, EMG features as key and value; the other branch takes EMG features as query, EEG features as key and value.

[0152] For the first branch, the query, key, and value calculations are defined as follows:

[0153]

[0154] in represents the weight matrix from the EEG modality, used to transform H EEG Convert to Query; represents the weight matrix from the EEG modality, used to transform H EEG Convert to Key; represents the weight matrix from the EEG modality, used to transform H EEG Convert to Value.

[0155] For the second branch, the roles are reversed and defined as:

[0156]

[0157] in represents the weight matrix from the EMG modality, used to transform H EMG Convert to Query; represents the weight matrix from the EMG modality, used to transform H EMG Convert to Key; represents the weight matrix from the EMG modality, used to transform H EMG Convert to Value.

[0158] Then, the multi-head attention outputs of each branch are concatenated to generate interactive feature representations of EEG and EMG, that is, EEG features are generated separately. and electromyographic characteristics

[0159] The EEG characteristics Defined as:

[0160]

[0161] Concat means concatenation.

[0162] The myoelectric characteristics Defined as:

[0163]

[0164] Finally, the EEG features and electromyographic characteristics Further integration, we get the multimodal joint feature representation H c for:

[0165]

[0166] Through the above-mentioned multimodal interactive fusion mechanism, the deep integration of EEG and EMG features is achieved, which significantly improves the richness and accuracy of multimodal feature representation and provides strong support for multimodal emotional state analysis.

[0167] Step 5: Build a multi-task classifier based on the emotional state labels.

[0168] The construction method of the multi-task classifier is as follows: two classification tasks are defined according to the emotional state label, wherein one of the classification tasks is a three-level classification problem of psychological state; the three-level classification problem of psychological state is defined as a three-level classification task, including three categories: normal, mild problem and severe problem; the other classification task is a two-level classification problem of five personalities; the two-level classification problem of five personalities is defined as five two-level classification tasks, including five categories: openness, conscientiousness, extroversion, agreeableness and neuroticism.

[0169] A classifier is defined for each classification task to evaluate the emotional state of the physiological data corresponding to the multimodal joint feature representation based on the input multimodal joint feature representation.

[0170] The classifiers corresponding to the two classification tasks are defined as:

[0171] P i = Dropout(ReLU(W i H c +b i )) (16)

[0172]

[0173] Where P i is the hidden layer feature representation; W i , b i , W k , b k are weight matrices for linear transformation and bias terms respectively; Dropout is a regularization method that prevents overfitting by randomly inactivating some neurons during training; ReLU is an activation function that introduces nonlinear characteristics and improves the classifier's ability to express complex features; argmax means finding the category with the highest probability from the output of the classifier; b t represents the bias term; The evaluation result output by the classifier.

[0174] In this embodiment, a three-level classification problem of psychological state is defined, and its classification targets include "normal", "mild problem" and "serious problem". A two-level classification problem of five personalities is defined, which respectively correspond to the evaluation of the Big Five personality traits, namely openness, conscientiousness, extraversion, agreeableness and neuroticism.

[0175] In order to further improve the accuracy of emotional state assessment, this implementation method designs a multi-task framework to use the generated multimodal fusion model to comprehensively analyze the psychological state and personality characteristics of college students. The framework uses a classifier of the personality traits of a task individual output by the model to consider the impact of individual differences. The framework is based on the multimodal multi-event fusion feature representation H c, with the help of the auxiliary information of the personality trait labels of the previous psychological test results, the accuracy of the evaluation of the three-level classification problem of the psychological state and the secondary classification problem of the five personalities is improved. In this embodiment, the process of comprehensively analyzing the psychological state and personality characteristics of college students using a multimodal fusion model is mainly divided into the following two aspects: 1) The interpretation of the model evaluation results is a very important link: First, the output results of the model include indicators such as psychological state and personality characteristics. It is necessary to fully interpret the prediction results of the model for indicators such as the psychological state and personality characteristics of the college student group to understand the overall distribution. For example, analyze the proportion of poor psychological state, the distribution of different personality traits, etc. Secondly, it is necessary to combine the contribution analysis of multimodal features to deeply explore the role of different modes in the evaluation results. Through interpretable methods such as gradients, quantify the degree of influence of each feature on the final evaluation. For example, the correlation between the number of negative emotional words in the text and the psychological state, the correlation between facial micro-expressions and personality extroversion, etc. This helps to clarify the value of multimodal data in psychological analysis and provide a basis for subsequent intervention and counseling. 2) The application of model prediction results is the ultimate goal: First, schools can provide focused attention and intervention for college students in need based on the results of psychological assessment. For high-risk groups such as those with poor psychological status and paranoid personality traits, early intervention can be made to develop personalized psychological counseling plans. Second, schools can optimize and improve the content and form of college students' mental health education based on the overall analysis results.

[0176] Step 6: Use the training set to iteratively train the multimodal fusion model and the multi-task classifier. Stop the iteration when the total loss function of the multi-task classifier converges or reaches the preset number of training rounds or meets the preset classification target, and obtain the trained multimodal fusion model and multi-task classifier.

[0177] The total loss function of the multi-task classifier is:

[0178] L total =βL m +γ(L1+L2+L3+L4+L5) (18)

[0179] Where L m represents the cross entropy loss of the three-level classification problem of psychological state; β represents the weight coefficient of the three-level classification problem of psychological state; L1, L2, L3, L4, L5 represent the cross entropy loss of the two-level classification problem of five personalities; γ represents the weight coefficient of the two-level classification problem of five personalities.

[0180] Where L m , the cross entropy losses of L1, L2, L3, L4 and L5 are all expressed as:

[0181]

[0182] Where L represents the cross entropy loss; N represents the total number of data in the training set; C represents the total number of categories; y lj is the true label of the lth sample in the jth category; is the evaluation result of the lth sample.

[0183] In this embodiment, in order to optimize the training of the classifier, cross entropy is used as the loss function of the classification task, and the cross entropy loss is defined as shown in formula (19). In order to comprehensively optimize the performance of psychological state and personality trait assessment, this embodiment proposes a weighted multi-task loss function, which weights and sums the losses of each classification task, as shown in formula (18), and β and γ are used to balance the importance of each task. The loss function is used to measure the gap between the model prediction and the true label, and provide a target for model optimization. When the loss converges, reaches the preset number of training rounds, or reaches the target performance. The training process is forward propagation, calculation of loss, back propagation, and parameter update. The output of the trained classifier is the predicted category and probability distribution. The use process is to input a new sample, obtain the predicted type, and use the result for the actual task.

[0184] Step 7: Use the test set to verify the trained multimodal fusion model and multi-task classifier to obtain the final multimodal fusion model and multi-task classifier.

[0185] Step 8: For any evaluation object, the physiological data of the evaluation object is collected and input into the final multimodal fusion model for feature fusion, and then the generated multimodal joint feature representation is input into the multi-task classifier for evaluation to obtain the emotional state evaluation result of the evaluation object.

[0186] The emotional state assessment system based on multiple events and multiple tasks of this embodiment is used to implement the emotional state assessment method based on multiple events and multiple tasks, such as Figure 4 As shown, the system includes: a data acquisition module, a data processing module, a feature fusion module and a classification evaluation module.

[0187] The data acquisition module is used to collect physiological data of the evaluation object and obtain the emotional state label of the evaluation object, and transmit the collected physiological data to the data processing module, and transmit the obtained emotional state label to the classification evaluation module.

[0188] The data processing module is used to perform segmentation, downsampling, filtering, feature extraction and splicing preprocessing on the received physiological data in sequence, and to construct a multimodal data set using the preprocessed physiological data, and transmit the multimodal data set to the feature fusion module.

[0189] The feature fusion module is used to extract features of a single modality from data of different modalities in a multimodal data set, and to perform single-modality fusion based on a channel attention module and multimodal fusion based on an interactive attention mechanism on the extracted feature vectors, to generate a multimodal joint feature representation and transmit it to the classification evaluation module.

[0190] The classification and evaluation module is used to construct a classifier according to the emotional state label, and input the multimodal joint feature representation into the classifier to generate an emotional state evaluation result of the physiological data corresponding to the multimodal joint feature representation.

[0191] In this embodiment, an emotional state assessment method based on multi-event multi-task is experimentally evaluated, and the data set is divided into a training set and a test set in a ratio of 4:1 to ensure that the samples are completely independent and there are no overlapping individuals, thereby avoiding the impact of data leakage on the results.

[0192] The experimental results show that the M 3 The ADD framework, i.e., the multi-task classifier, outperforms all baseline models in the three-level classification of mental states. For the three-level classification of mental states, mental state classification is implemented from two dimensions, negative mentality and positive emotion. The accuracy and F1 score of negative mentality classification reached 86.84% and 87.09% respectively; the accuracy and F1 score of positive emotion classification reached 95.12% and 95.26% respectively. This result verifies the applicability and stability of the multimodal, multi-task, and multi-event analysis framework, and significantly improves the accuracy of emotional state assessment.

[0193] In terms of multi-task learning, experimental results show that adding personality trait classification tasks can significantly improve model performance, with accuracy and F1 score increased by at least 5%. This indicates that the personality classification task provides the model with additional contextual information, which helps to more accurately understand individual differences and optimize detection effects.

[0194] In terms of multimodal fusion, comparative experiments show that for the three-level classification of psychological states, the performance of multimodal fusion methods is significantly better than that of single-modal methods. The joint features of EEG and EMG are always better than single-modal features, further proving the importance of multimodal integration in emotional state assessment.

[0195] In terms of multi-event fusion, the multi-event mechanism improves model performance compared to single event analysis, with accuracy increased by more than 3% and F1 score increased by more than 5%. Further analysis shows that different events have significant differences in their contribution to the model.

[0196] To verify the role of channel attention CAM and cross-attention IA mechanism, ablation experiments were conducted. The results show that removing CAM or IA mechanism will significantly reduce the model performance, with accuracy and F1 score decreasing by 3%-5% respectively. When both mechanisms are removed at the same time, the performance decreases most significantly, which shows that both mechanisms are indispensable in feature extraction and modality interaction.

[0197] In general, the M of this embodiment 3 The ADD framework makes full use of multimodal, multitask, and multi-event information, and significantly improves the classification performance on the three-level classification problem of psychological states. The experimental results verify the effectiveness of the method of the present invention, provide technical support for the accurate identification and analysis of emotional states, and have important application value.

[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for assessing emotional state based on multi-event and multi-task, characterized in that: The method comprises the following steps: Step 1: Collect the physiological data of the evaluation object, conduct psychological tests on the evaluation object, and generate the emotional state label of the evaluation object according to the psychological test results; Step 2: Preprocess the physiological data of the evaluation object and construct a multimodal dataset, and divide the multimodal dataset into a training set and a test set according to a set ratio; Step 3: Construct the EEG neural network model and the EMG neural network model respectively; Step 4: Fuse the EEG neural network model of the EEG signal and the EMG neural network model of the EMG signal to obtain a multimodal fusion model; Step 5: Build a multi-task classifier based on the emotional state labels; Step 6: Use the training set to iteratively train the multimodal fusion model and the multi-task classifier. When the total loss function of the multi-task classifier converges or reaches the preset number of training rounds or meets the preset classification target, the iteration is stopped to obtain the trained multimodal fusion model and multi-task classifier. Step 7: Use the test set to verify the trained multimodal fusion model and multi-task classifier to obtain the final multimodal fusion model and multi-task classifier; Step 8: For any evaluation object, the physiological data of the evaluation object is collected and input into the final multimodal fusion model for feature fusion, and then the generated multimodal joint feature representation is input into the multi-task classifier for evaluation to obtain the emotional state evaluation result of the evaluation object.

2. The method for assessing emotional state based on multiple events and multiple tasks according to claim 1, characterized in that: The method for collecting the physiological data in step 1 is: performing multimodal physiological and behavioral tests on the evaluation subject, and recording the physiological data of the evaluation subject in real time; The process of the multimodal physiological and behavioral test is as follows: the subject is provided with an electroencephalogram sensor and an electromyography sensor to respectively record the electroencephalogram signal and the electromyography signal of the subject, and the electroencephalogram signal and the electromyography signal are used as the physiological data of the subject; Conducting several behavioral tests on the subject in turn, and recording the subject's physiological data in real time during each behavioral test; each behavioral test is any one of a self-introduction test, a text reading test, or a gait test; The self-introduction test is as follows: the subject of evaluation introduces himself according to a specified topic; the text reading test is as follows: the subject of evaluation reads a specified text; the gait test is as follows: the subject of evaluation walks in an orderly manner in a specified area according to a predetermined route and direction; The psychological test is: using a personality scale, a general happiness scale, a depression screening scale and a depressive symptom scale to conduct psychological tests on the evaluation object, and using the questionnaire scores of each scale as the emotional state label of the evaluation object; The emotional state label includes: five personality dimensions of the assessed object and the psychological state of the assessed object.

3. The method for assessing emotional state based on multiple events and multiple tasks according to claim 2, characterized in that: The step 2 further comprises: Step 2.1: Based on the timestamp information recorded in the physiological data, the start and end time points of the evaluation subject when experiencing different activity stages are marked, and the physiological data is segmented into several segments according to the marked start and end time points to obtain the segmented physiological data; Step 2.2: downsampling the segmented physiological data to obtain downsampled physiological data; Step 2.3: filtering the downsampled physiological data to obtain filtered physiological data; wherein a fourth-order Butterworth filter is used to band-pass filter the EEG signal in the physiological data; and a band-pass filter is used to band-pass filter the EMG signal in the physiological data; Step 2.4: Extract features from the filtered physiological data to obtain time domain features and frequency domain features of the physiological data; Step 2.5: for any physiological data, the time domain features and frequency domain features of the physiological data are combined into a feature vector of the physiological data, a behavioral test is regarded as an event, and the feature vectors of all physiological data in each event are regarded as a multimodal data sequence of the event, and a multimodal data sequence of all events is used to construct a multimodal data set; wherein the multimodal data set includes: an EEG data set of an electroencephalogram signal and an EMG data set of an electromyography signal; Step 2.6: Divide the multimodal dataset into training set and test set according to the preset ratio.

4. The method for assessing emotional state based on multiple events and multiple tasks according to claim 3, characterized in that: The EEG neural network model for electroencephalogram signals and the EMG neural network model for electromyography signals in step 3 are both single-modality neural network models; The single modality neural network model is: using a convolutional neural network (CNN) model to extract local features from an input multimodal data sequence, wherein the convolutional neural network (CNN) model is a layer of convolution operation; The feature representation extracted by the convolutional neural network CNN model is input into the multi-head self-attention mechanism for processing, and the attention score of each physiological data in the multimodal data sequence is calculated. The attention score of the physiological data is used to calculate the value vector V i Perform weighted summation to obtain weighted feature representation of physiological data; For any event, the weighted feature representation of the physiological data in the event is residually connected with the multimodal data sequence of the event, and the result of the residual connection is layer-normalized. The normalized features are then nonlinearly transformed through a feedforward network to generate a feature sequence of the event. The feature sequence of the event is input into the bidirectional long short-term memory network Bi-LSTM, and the Bi-LSTM is used to process the forward time dependency and reverse time dependency in the feature sequence respectively. The hidden state of the last time step in the feature sequence is taken as the output of Bi-LSTM to obtain the feature vector of the event.

5. The method for assessing emotional state based on multiple events and multiple tasks according to claim 4, characterized in that: The process of fusing the EEG neural network model of the electroencephalogram signal and the EMG neural network model of the electromyography signal described in step 4 is divided into two parts: a single-modal fusion process based on a channel attention module and a multi-modal fusion process based on an interactive attention mechanism; The unimodal fusion process based on the channel attention module is as follows: constructing a joint feature representation of multiple events using the feature vectors output by the EEG neural network model of the electroencephalogram signal and the EMG neural network model of the electromyography signal, respectively, and performing global average pooling and global maximum pooling on the joint feature representation of multiple events, respectively, performing dimension reduction and activation processing on the obtained global average pooling features and global maximum pooling features through a two-layer fully connected network, obtaining the attention weight of the event and then fusing it with the joint feature representation of multiple events to generate a weighted multi-event feature representation; For the EEG neural network model, the weighted multi-event feature representation obtained after the feature vector output by the model is fused by single modality is used as the multi-event fusion feature H of the EEG signal. EEG For the EMG neural network model, the weighted multi-event feature representation obtained after the single-modal fusion of the feature vector output by the model is used as the multi-event fusion feature H of the EMG signal. EMG ; The multimodal fusion process based on the interactive attention mechanism is as follows: a double-branch cross attention module is used to calculate the multi-event fusion feature H of the EEG signal respectively. EEG and the multi-event fusion feature H of electromyographic signals EMG The feature correlation between them is used to obtain the attention output from the EEG modality to the EMG modality. and attention output from EMG modality to EEG modality Then by and Splice and generate EEG features respectively and electromyographic characteristics And integrate multimodal joint feature representation.

6. The method for assessing emotional state based on multiple events and multiple tasks according to claim 5, characterized in that: The specific content of the single-modal fusion process based on the channel attention module is: The feature vectors of each event in the multimodal dataset are fused along the channel dimension to construct a joint feature representation of multiple events; The joint feature representation of multiple events is represented by global average pooling and global maximum pooling respectively, and the global average pooling feature H is obtained. t,avg And the global maximum pooling feature H t,max , the first layer of fully connected network is used to respectively average the global pooling features H t,avg And the global maximum pooling feature H t,max Perform dimensionality reduction processing, and then use the second-layer fully connected network to activate the reduced global average pooling features and global maximum pooling features respectively, and then merge the activated global average pooling features and global maximum pooling features to obtain the attention weight of the event; The attention weight of the event is expressed as: α event =Sigmoid(W1(W0(H t,avg ))+W1(W0(H t,max ))) where α event Represents the attention weight of the event; Sigmoid represents the activation function; W1 is the recovery matrix; W0 is the dimension reduction matrix; Use element-by-element multiplication to convert the event's attention weight α event The joint feature representation H of multiple events t Fusion is performed to obtain the weighted multi-event feature representation H event .

7. The method for assessing emotional state based on multiple events and multiple tasks according to claim 6, characterized in that: The dual-branch cross attention module is: one attention branch uses the multi-event fusion feature H of the EEG signal EEG To query Query, the multi-event fusion feature H of the electromyographic signal EMG For the key Key and value Value, calculate H EEG With H EMG The feature correlation between EEG With H EMG Feature correlation between EEG modality and EMG modality attention output It is expressed as: Where Softmax represents the normalized exponential function; Indicates the multi-event fusion feature H of the EEG signal EEG Map to Query space; Indicates the multi-event fusion feature H of the EEG signal EEG Map to Key space; Indicates the multi-event fusion feature H of the EEG signal EEG Mapped to the Value space; e represents the EEG modality; m represents the EMG modality; d represents the feature dimension of Query and Key, which is used to prevent the attention score from being too large in high-dimensional space, resulting in gradient disappearance or numerical instability; Another attention branch is to use the multi-event fusion feature H of the electromyographic signal EMG To query Query, the multi-event fusion feature H of EEG signal EEG For the key Key and value Value, calculate H EMG With H EEG The feature correlation between EMG With H EEG Feature correlation between EMG and EEG modalities It is expressed as: in Indicates the multi-event fusion feature H of the electromyographic signal EMG Map to Query space; Indicates the multi-event fusion feature H of the electromyographic signal EMG Map to Key space; Indicates the multi-event fusion feature H of the electromyographic signal EMG Mapped to the Value space.

8. The method for assessing emotional state based on multiple events and multiple tasks according to claim 7, characterized in that: The method for constructing the multi-task classifier in step 5 is as follows: defining two classification tasks according to the emotional state labels, wherein one of the classification tasks is a three-level classification problem of psychological states; The three-level classification problem of psychological state is defined as a three-level classification task, including three categories: normal, mild problem and severe problem; another classification task is a two-level classification problem of five personalities; the two-level classification problem of five personalities is defined as five two-level classification tasks, including five categories: openness, conscientiousness, extraversion, agreeableness and neuroticism; A classifier is defined for each classification task to evaluate the emotional state of the physiological data corresponding to the multimodal joint feature representation based on the input multimodal joint feature representation.

9. The method for assessing emotional state based on multiple events and multiple tasks according to claim 8, characterized in that: The total loss function L of the multi-task classifier described in step 6 total for: L total =βL m +γ(L1+L2+L3+L4+L5) Where L m represents the cross entropy loss of the three-level classification problem of psychological state; β represents the weight coefficient of the three-level classification problem of psychological state; L1, L2, L3, L4, L5 represent the cross entropy loss of the two-level classification problem of five personalities; γ represents the weight coefficient of the two-level classification problem of five personalities; Where L m , the cross entropy losses of L1, L2, L3, L4 and L5 are all expressed as: Where L represents the cross entropy loss; N represents the total number of data in the training set; C represents the total number of categories; y lj is the true label of the lth sample in the jth category; is the evaluation result of the lth sample.

10. An emotional state assessment system based on multiple events and multiple tasks, used to implement the emotional state assessment method based on multiple events and multiple tasks as described in claim 1, characterized in that: The system includes: a data acquisition module, a data processing module, a feature fusion module and a classification evaluation module; The data acquisition module is used to collect physiological data of the evaluation object and obtain the emotional state label of the evaluation object, and transmit the collected physiological data to the data processing module, and transmit the obtained emotional state label to the classification evaluation module; The data processing module is used to perform segmentation, downsampling, filtering, feature extraction and splicing preprocessing on the received physiological data in sequence, and construct a multimodal data set using the preprocessed physiological data, and transmit the multimodal data set to the feature fusion module; The feature fusion module is used to extract features of a single modality from data of different modalities in a multimodal data set, and sequentially perform single-modality fusion based on a channel attention module and multimodal fusion based on an interactive attention mechanism on the extracted feature vectors to generate a multimodal joint feature representation and transmit it to the classification evaluation module; The classification and evaluation module is used to construct a classifier according to the emotional state label, and input the multimodal joint feature representation into the classifier to generate an emotional state evaluation result of the physiological data corresponding to the multimodal joint feature representation.

Citation Information

Patent Citations

  • Multi-mode intelligent control system and method based on electroencephalogram and myoelectricity information

    CN107957783A

  • Multi-modal emotion recognition method and system based on wearable device

    CN117520826A

  • Multi-modal emotion recognition method and system based on regularization fusion

    CN118656745A

  • Depression auxiliary diagnosis system based on electroencephalogram and voice and data analysis method thereof

    CN119235314A

  • Detection Of Disease Conditions And Comorbidities

    US20170251985A1

Cited By

  • Electromyographic signal identification and classification method and system

    CN120832569A

  • Preoperative multi-complication risk prediction method and system based on structured clinical data

    CN120998507A

  • A preoperative multiple comorbidity risk prediction method and system based on structured clinical data

    CN120998507B