Emotion recognition method, device and system based on multi-modal physiological signals
Patent Information
- Application Number
- CN202611308756.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-27
- Publication Date
- 2026-09-25
AI Technical Summary
但是,皮肤电等外周生理信号在可穿戴采集场景中容易受到电极接触、出汗、运动伪迹和环境变化影响,信号质量在不同样本之间波动较大
[0016]本申请技术方案通过对外周生理信号的特征序列进行池化处理得到样本级可靠性门控系数;根据样本级可靠性门控系数对外周生理信号的特征序列进行重标定得到重标定外周特征序列;将重标定外周特征序列和其余模态生理信号的特征序列输入预设情绪分类器得到情绪状态识别结果;通过样本级可靠性门控系数重标定外周生理信号的特征序列,降低了运动伪迹、电极接触不良等干扰带来的负面影响,提高情绪状态识别精度。
Smart Images

Figure CN122805274A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing, and in particular to an emotion recognition method, device, and system based on multimodal physiological signals. Background Technology
[0002] Emotion recognition has significant application value in scenarios such as human-computer interaction. With the development of the Internet of Things and wearable body area networks, continuous emotion recognition based on physiological signals such as electroencephalography (EEG), electrooculography (EOG), and electrodermal conductance (EDC) is gradually becoming an important technical approach for emotion recognition.
[0003] Currently, most emotion recognition solutions employ a non-discriminatory feature stitching method to integrate physiological signals such as EEG, EEG, and TENS. However, peripheral physiological signals like TENS are easily affected by electrode contact, sweating, motion artifacts, and environmental changes in wearable data acquisition scenarios, resulting in significant fluctuations in signal quality across different samples. The non-discriminatory feature stitching method treats both reliable and distorted peripheral physiological signals equally, making the emotion recognition results susceptible to interference from distorted peripheral physiological signals.
[0004] Therefore, existing emotion recognition schemes suffer from low accuracy in identifying emotional states. Summary of the Invention
[0005] The main objective of this application is to propose an emotion recognition method, device, and system based on multimodal physiological signals, aiming to improve the accuracy of emotion state recognition.
[0006] To achieve the above objectives, this application proposes an emotion recognition method based on multimodal physiological signals, including: Obtain the feature sequence of multimodal physiological signals corresponding to the target object. The multimodal physiological signals include at least central nervous system physiological signals and peripheral physiological signals. The feature sequences of the peripheral physiological signals are pooled to obtain sample-level reliability gating coefficients; The peripheral physiological signal feature sequence is recalibrated based on the sample-level reliability gating coefficient to obtain the recalibrated peripheral feature sequence. The recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals are input into a preset emotion classifier to obtain the emotion state recognition result.
[0007] In some embodiments, the multimodal physiological signals further include electrooculogram (EOG) signals, the central nervous system physiological signals include electroencephalogram (EEG) signals, and the peripheral physiological signals include electrodermal (ED) signals. The feature sequence of the multimodal physiological signal corresponding to the target object includes: Collect multimodal physiological signals of the target object within a preset time period; Timing alignment is performed based on the start timestamp, sampling frequency, or sampling frame rate of each modality's physiological signal; The time-aligned physiological signals of each modality are subjected to at least one of the following signal normalization processes: filtering, baseline correction, artifact removal, standardization, and resampling. Based on the preset window length and preset window step size, the physiological signals of each modality after signal normalization are divided and processed to obtain time window samples of each modality of physiological signals. The time window samples of each modality physiological signal are input into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal.
[0008] In some embodiments, the modality-specific coding network includes a first coding network for the electroencephalogram (EEG) signals, a second coding network for the electrooculogram (EOG) signals, and a third coding network for the electrodermal (ED) signals. The step of inputting time window samples of each modality's physiological signal into the corresponding modality-specific coding network to obtain the feature sequence of each modality's physiological signal includes: The time window samples of the EEG signal are input into the first coding network to extract the features of the EEG signal, the time window samples of the EOL signal are input into the second coding network to extract the features of the EOL signal, and the time window samples of the EDS signal are input into the third coding network to extract the features of the EDS signal. The features of the electroencephalogram (EEG), electrooculogram (EOG), and electrodermal (ED) signals are mapped to time token sequences of uniform length and dimension to obtain feature sequences of the EEG, EOG, and EED signals.
[0009] In some embodiments, the pooling process of the feature sequence of the peripheral physiological signal to obtain the sample-level reliability gating coefficient includes: Pooling is performed on the feature sequences of the peripheral physiological signals along the time dimension to obtain a sample-level global statistical vector; The sample-level global statistical vector is input into the gating subnetwork to obtain the sample-level reliability gating coefficient, which has a value range of 0 to 1.
[0010] In some embodiments, the step of recalibrating the feature sequence of the peripheral physiological signal according to the sample-level reliability gating coefficient to obtain the recalibrated peripheral feature sequence includes: The sample-level reliability gating coefficient is multiplied sample by sample by sample with the feature sequence of the peripheral physiological signal to adjust the feature weights of the feature sequence of the peripheral physiological signal, thereby obtaining the recalibrated peripheral feature sequence.
[0011] In some embodiments, the emotion recognition method based on multimodal physiological signals further includes: At least one of Gaussian noise, time masking, time shifting, random dropping, and packet loss perturbation is applied to the feature sequence of the peripheral physiological signal to obtain the interfering peripheral feature sequence; Based on the interference peripheral feature sequence, sample-level reliability gating coefficients are generated, and then recalibrated and input into a preset emotion classifier to obtain the gating recognition result. The interference peripheral feature sequence is input into a preset emotion classifier to obtain a recognition result without gating. By comparing the recognition results with and without gating, the magnitude of the decrease in recognition results is determined.
[0012] This application further proposes an emotion recognition device based on multimodal physiological signals, comprising: The acquisition unit is used to acquire the feature sequence of multimodal physiological signals corresponding to the target object. The multimodal physiological signals include at least central nervous system physiological signals and peripheral physiological signals. A pooling unit is used to perform pooling processing on the feature sequences of the peripheral physiological signals to obtain sample-level reliability gating coefficients. The recalibration unit is used to recalibrate the feature sequence of the peripheral physiological signal according to the sample-level reliability gating coefficient to obtain a recalibrated peripheral feature sequence. The classification unit is used to input the recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals into a preset emotion classifier to obtain the emotion state recognition result.
[0013] In some embodiments, the multimodal physiological signals further include electrooculogram (EOG) signals, the central nervous system physiological signals include electroencephalogram (EEG) signals, and the peripheral physiological signals include electrodermal (ED) signals. The acquisition unit is specifically used for: Collect multimodal physiological signals of the target object within a preset time period; Timing alignment is performed based on the start timestamp, sampling frequency, or sampling frame rate of each modality's physiological signal; The time-aligned physiological signals of each modality are subjected to at least one of the following signal normalization processes: filtering, baseline correction, artifact removal, standardization, and resampling. Based on the preset window length and preset window step size, the physiological signals of each modality after signal normalization are divided and processed to obtain time window samples of each modality of physiological signals. The time window samples of each modality physiological signal are input into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal.
[0014] In some embodiments, the modality-specific coding network includes a first coding network for the electroencephalogram (EEG) signals, a second coding network for the electrooculogram (EOG) signals, and a third coding network for the electrodermal (ED) signals. The acquisition unit, when executing the process of inputting time window samples of each modality physiological signal into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal, is specifically used for: The time window samples of the EEG signal are input into the first coding network to extract the features of the EEG signal, the time window samples of the EOL signal are input into the second coding network to extract the features of the EOL signal, and the time window samples of the EDS signal are input into the third coding network to extract the features of the EDS signal. The features of the electroencephalogram (EEG), electrooculogram (EOG), and electrodermal (ED) signals are mapped to time token sequences of uniform length and dimension to obtain feature sequences of the EEG, EOG, and EED signals.
[0015] This application further proposes an emotion recognition system based on multimodal physiological signals, which includes a data acquisition sensor and the aforementioned emotion recognition device based on multimodal physiological signals.
[0016] The technical solution of this application obtains sample-level reliability gating coefficients by pooling the feature sequences of peripheral physiological signals; recalibrates the feature sequences of peripheral physiological signals based on the sample-level reliability gating coefficients to obtain recalibrated peripheral feature sequences; inputs the recalibrated peripheral feature sequences and the feature sequences of other modal physiological signals into a preset emotion classifier to obtain the emotion state recognition result; by recalibrating the feature sequences of peripheral physiological signals through sample-level reliability gating coefficients, the negative impact of interference such as motion artifacts and poor electrode contact is reduced, thereby improving the accuracy of emotion state recognition. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an embodiment of the emotion recognition method based on multimodal physiological signals according to this application. Figure 2 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application; Figure 3 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application; Figure 4 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application; Figure 5 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application; Figure 6 This is a schematic diagram of the structure of an embodiment of the emotion recognition device based on multimodal physiological signals according to this application; Figure 7 This is a schematic diagram of an embodiment of the emotion recognition system based on multimodal physiological signals according to this application. Detailed Implementation
[0018] The solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments in this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] It should be noted that all directional indicators (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.
[0020] It should also be noted that when a component is described as "fixed to" or "set on" another component, it can be directly on the other component or there may be an intervening component present. When a component is described as "connected to" another component, it can be directly connected to the other component or there may be an intervening component present.
[0021] Furthermore, the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. When the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.
[0022] This application proposes an emotion recognition method based on multimodal physiological signals, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the emotion recognition method based on multimodal physiological signals according to this application. In some embodiments, the emotion recognition method based on multimodal physiological signals includes: Step S110: Obtain the feature sequence of the multimodal physiological signals corresponding to the target object. The multimodal physiological signals include at least central nervous system physiological signals and peripheral physiological signals. Step S120: Pooling is performed on the feature sequences of peripheral physiological signals to obtain sample-level reliability gating coefficients; Step S130: Recalibrate the feature sequence of peripheral physiological signals according to the sample-level reliability gating coefficient to obtain the recalibrated peripheral feature sequence; Step S140: Input the recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals into the preset emotion classifier to obtain the emotion state recognition result.
[0023] In this embodiment, as Figure 1 As shown, the emotion recognition method based on multimodal physiological signals proposed in this application can be configured as a running program, functional software, or packaged as firmware, control driver, or independent functional module; then it can be deployed to an emotion recognition device based on multimodal physiological signals, so that the emotion recognition device based on multimodal physiological signals can run the emotion recognition method based on multimodal physiological signals. The emotion recognition device based on multimodal physiological signals can be simply referred to as the recognition device.
[0024] When running an emotion recognition method based on multimodal physiological signals, the recognition device can be used to identify the emotions of the target object. First, the recognition device can acquire the feature sequence of multimodal physiological signals corresponding to the target object. These multimodal physiological signals include at least central nervous system physiological signals and peripheral physiological signals. For example, central nervous system physiological signals can include electroencephalogram (EEG) signals, which can stably characterize the body's internal emotional neural activity. Peripheral physiological signals can include electrodermal signaling (EDS) signals, which are used to characterize autonomic nervous system arousal-related responses. The recognition device can first collect the multimodal physiological signals corresponding to the target object, such as collecting both central nervous system and peripheral physiological signals. Then, it performs feature extraction on the collected multimodal physiological signals in chronological order to obtain the feature sequence of the multimodal physiological signals.
[0025] After obtaining the feature sequences of multimodal physiological signals, pooling can be performed on the feature sequences of peripheral physiological signals to obtain sample-level reliability gating coefficients. For example, peripheral physiological signals are easily interfered with by sweating, loose electrode contact, and limb movement in wearable acquisition scenarios, resulting in significant differences in signal quality among different test samples. Therefore, after obtaining the feature sequences of multimodal physiological signals, the recognition device can perform pooling on the feature sequences of peripheral physiological signals to obtain sample-level reliability gating coefficients. For instance, the recognition device performs global pooling on the feature sequences of peripheral physiological signals along the time dimension, aggregating all temporal feature information within a single sample to obtain a sample-level global statistical vector that can characterize the overall signal quality of the current sample. This global statistical vector is then input into a preset gating subnetwork, and through nonlinear mapping and activation function constraints, sample-level reliability gating coefficients are output.
[0026] After obtaining the sample-level reliability gating coefficients, the feature sequences of peripheral physiological signals can be recalibrated using these coefficients to obtain recalibrated peripheral feature sequences. For example, the recognition device recalibrates the feature sequences of peripheral physiological signals based on the sample-level reliability gating coefficients. When the signal sample quality of the feature sequences of peripheral physiological signals is high and noise interference is low, a relatively complete feature sequence of the peripheral physiological signals can be retained. When the signal distortion of the feature sequences of peripheral physiological signals is severe and the reliability is low, the weight contribution of the feature sequences of peripheral physiological signals is suppressed to avoid noise features interfering with the overall recognition effect; thus, the recalibrated peripheral feature sequences are obtained.
[0027] After obtaining the recalibrated peripheral feature sequence, the recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals can be input into a preset emotion classifier to obtain the emotion state recognition result. For example, the recognition device can be configured with a preset emotion classifier. The preset emotion classifier can be customized by the developer according to the actual situation. The preset emotion classifier can be configured with an emotion classification table, which can include the emotion state recognition results corresponding to various modal physiological signals. The recognition device can input the recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals into the preset emotion classifier. The preset emotion classifier can then match the corresponding emotion state recognition result according to the recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals. Among them, the emotion state recognition result includes pleasure level, arousal level, stress level, positive / neutral / negative emotion category, etc.
[0028] The technical solution of this application obtains sample-level reliability gating coefficients by pooling the feature sequences of peripheral physiological signals; recalibrates the feature sequences of peripheral physiological signals based on the sample-level reliability gating coefficients to obtain recalibrated peripheral feature sequences; inputs the recalibrated peripheral feature sequences and the feature sequences of other modal physiological signals into a preset emotion classifier to obtain the emotion state recognition result; by recalibrating the feature sequences of peripheral physiological signals through sample-level reliability gating coefficients, the negative impact of interference such as motion artifacts and poor electrode contact is reduced, thereby improving the accuracy of emotion state recognition.
[0029] Reference Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application. In some embodiments, the multimodal physiological signals also include electrooculogram (EOG) signals, the central nervous system physiological signals include electroencephalogram (EEG) signals, and the peripheral physiological signals include electrodermal signals. The aforementioned characteristic sequences of the multimodal physiological signals corresponding to the target object include: Step S150: Collect multimodal physiological signals of the target object within a preset time period; Step S151: Perform time alignment processing based on the start timestamp, sampling frequency, or sampling frame rate of each modality's physiological signal; Step S152: Perform at least one of the following signal normalization processes on the time-aligned physiological signals of each modality: filtering, baseline correction, artifact removal, standardization, and resampling. Step S153: Based on the preset window length and preset window step size, the physiological signals of each modality after signal normalization are divided to obtain time window samples of each modality of physiological signals. Step S154: Input the time window samples of each modality physiological signal into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal.
[0030] In this embodiment, as Figure 2 As shown, during step S110, multimodal physiological signals of the target object can be collected within a preset time period. The recognition device can collect multimodal physiological signals of the target object within the preset time period. These multimodal physiological signals include electrooculogram (EOG) signals; EOG signals are used to characterize eye movements, blinking, or periocular behavioral responses. Central nervous system physiological signals include electroencephalogram (EEG) signals, and peripheral physiological signals include electrodermal (EDS) signals. For example, the recognition device can be connected to a wearable head-mounted device, EOG acquisition electrodes, and EDS sensors. The recognition device can control the wearable head-mounted device to collect EOG signals of the target object within the preset time period, control the EOG acquisition electrodes to collect EOG signals of the target object within the preset time period, and control the EDS sensors to collect EDS signals of the target object within the preset time period. The preset time can be customized by the developer according to actual conditions. For example, the preset time could be 6:00 AM to 6:10 AM, 12:00 PM to 12:30 PM, 5:00 PM to 6:00 PM, etc.
[0031] Timing alignment is performed based on the start timestamp, sampling frequency, or sampling frame rate of each modality's physiological signals. For example, EEG, EOS, and ESK signals are acquired independently by different acquisition devices, with inconsistent sampling frequencies, data transmission delays, and start times. The original EEG, EOS, and ESK signals have issues such as timing shifts, inconsistent sampling step sizes, and time axis misalignments. Direct use of these signals can lead to spatiotemporal misalignment of multimodal features and fusion failure. Therefore, after obtaining multimodal physiological signals, the identification device will use a unified built-in clock as a reference and rely on the start timestamps carried by each modality of physiological signals to calibrate the true start time of acquisition of each modality of physiological signals; combined with the sampling frequency or sampling frame rate of each modality of physiological signals, it will calculate the time interval of single frame data of different modalities of physiological signals and construct a unique original time axis for each modality of physiological signals; then, through timestamp calibration, time delay offset compensation and linear interpolation matching, it will uniformly map the three heterogeneous physiological signals of EEG signals, EOS signals and EKS signals to the same global time axis, eliminate the timing deviation caused by multi-device acquisition, and ensure the timing alignment of each modality of physiological signals.
[0032] After time-aligned physiological signals, the device performs at least one of the following signal normalization processes: filtering, baseline correction, artifact removal, standardization, and resampling. For example, EEG signals are easily contaminated with power frequency noise and EMG artifacts; EOS signals are easily affected by blinking and large eye movements; and EDS signals are easily affected by electrode sliding, sweating, and fluctuations in contact resistance. After time-aligning the EEG, EOS, and EDS signals, the recognition device further performs at least one of these signal normalization processes on each modality of physiological signal. Filtering removes power frequency interference and high-frequency random noise; baseline correction eliminates baseline drift; artifact removal removes abrupt changes and motion interference artifacts; standardization maps physiological signals with different amplitude ranges to the same numerical range, eliminating dimensional differences; and resampling unifies the multimodal heterogeneous sampling frequencies to a fixed standard sampling rate.
[0033] Based on a preset window length and a preset window step size, the pre-normalized physiological signals of each modality are divided into time window samples to obtain time window samples of each modality. For example, the preset window length and preset window step size can be customized by the developer according to actual needs. The preset window length represents the total number of time-series sampling points contained within a single sliding time window; the preset window step size represents the number of sampling points that two adjacent sliding windows cross while sliding forward. The recognition device can divide the pre-normalized physiological signals of each modality based on the preset window length and preset window step size to obtain time window samples of each modality.
[0034] In a preferred embodiment, the time window samples of each modality of physiological signal are as follows: in, This represents a time window sample of the EEG signal. This represents a time window sample of the electrooculogram (EOG) signal. This represents a time window sample of the electrodermal signal. This indicates the number of channels for acquiring EEG signals. This indicates the number of channels for acquiring electrooculogram (EOG) signals. The number of channels for collecting electrodermal signals is represented by T, the preset window length is represented by R, and all real numbers are represented by R.
[0035] There can be multiple time window samples for each modality of physiological signal. Each time window sample of the modality of physiological signal can independently represent a short-term emotional physiological response feature, which not only preserves the local temporal change pattern of physiological signal, but also realizes the structured segmentation of dataset samples to adapt to the batch input inference requirements of modality-specific encoding network.
[0036] After obtaining time window samples of each modality's physiological signals, these samples can be input into the corresponding modality-specific encoding network to obtain the feature sequences of each modality's physiological signals. For example, the recognition device can use the modality-specific encoding network corresponding to each modality's physiological signal to process the time window samples of each modality's physiological signal to obtain the feature sequences of each modality's physiological signal. Specifically, the modality-specific encoding network corresponding to EEG signals can include one or more of temporal convolution, spatial convolution, separable convolution, graph convolution, Transformer encoder, or linear projection; the modality-specific encoding networks corresponding to EOS signals and ESK signals can include one or more of one-dimensional convolution, bidirectional temporal modeling, temporal convolutional networks, pooling layers, and linear projection.
[0037] Reference Figure 3 , Figure 3 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application. In some embodiments, the modality-specific coding network includes a first coding network for electroencephalogram (EEG) signals, a second coding network for electrooculogram (EOG) signals, and a third coding network for electrodermal (ED) signals. The aforementioned inputting time window samples of each modality's physiological signal into the corresponding modality-specific coding network to obtain the feature sequences of each modality's physiological signal includes: Step S160: Input the time window samples of the EEG signal into the first coding network to extract the features of the EEG signal, input the time window samples of the EOL signal into the second coding network to extract the features of the EOL signal, and input the time window samples of the EDS signal into the third coding network to extract the features of the EDS signal. Step S161: Map the features of the EEG signal, the EEG signal, and the EKD signal into time token sequences of uniform length and uniform dimension to obtain the feature sequences of the EEG signal, the EEG signal, and the EKD signal.
[0038] In this embodiment, as Figure 3 As shown, during step S154, the physiological signals of each modality can be input into the corresponding modality-specific encoding network. The modality-specific encoding network includes a first encoding network for EEG signals, a second encoding network for EOS signals, and a third encoding network for EEG signals. The first encoding network includes one or more of temporal convolution, spatial convolution, separable convolution, graph convolution, a Transformer encoder, or linear projection. The second encoding network includes one or more of one-dimensional convolution, bidirectional temporal modeling, temporal convolutional networks, pooling layers, and linear projection. The third encoding network includes one or more of one-dimensional convolution, bidirectional temporal modeling, temporal convolutional networks, pooling layers, and linear projection.
[0039] The recognition device can input time window samples of EEG signals into a first encoding network to extract EEG signal features, input time window samples of EOL signals into a second encoding network to extract EOL signal features, and input time window samples of EDS signals into a third encoding network to extract EDS signal features. For example, the signal characteristics, channel dimensions, and temporal variation patterns of different modal physiological signals vary significantly. A dedicated network structure is adapted for each modality of physiological signal to avoid the loss of modal features and representation degradation caused by a uniform encoding method.
[0040] Among them, EEG signals are central nervous system signals with multiple channels, strong spatial correlation, and complex temporal fluctuations. Therefore, the first coding network can integrate spatial convolution and temporal convolution structures to not only mine the spatial topological correlation features between different electrode channels, but also capture the dynamic changes of EEG temporal sequence. In addition, it can be used in conjunction with the Transformer encoder to achieve long-distance temporal dependency modeling and fully extract the features of deep and high-dimensional EEG signals.
[0041] Electrooculography (EOG) signals are physiological signals with few channels, strong temporal continuity, and relatively stable variation patterns. They mainly reflect behavioral characteristics such as blinking and eye movements that accompany emotional changes. Therefore, the second encoding network can use one-dimensional convolution combined with bidirectional temporal modeling to preserve the short-term abrupt changes and long-term temporal trends of EOG while reducing network computational overhead and accurately extracting the features of EOG signals corresponding to eye movement behaviors.
[0042] Electrodermal (EDS) signals are single-channel, slowly varying temporal signals, characterized by gradual signal changes, significant noise interference, and weak effective features. Therefore, the third encoding network can employ a temporal convolutional network combined with pooling layers to filter high-frequency invalid noise, enhance the emotional arousal features corresponding to the slow fluctuations in EDS, and stably output the characteristics of the EDS signal.
[0043] The recognition device maps the features of EEG signals, Eoptometry signals, and Electrodermal Signals (EDS) signals into time token sequences of uniform length and dimension, resulting in feature sequences for EEG, EOS, and EDS signals. For example, the device maps the features of EEG, EOS, and EDS signals uniformly to a token representation space of the same dimension, ultimately outputting feature sequences for EEG, EOS, and EDS signals that are dimensionally unified, temporally aligned, and retain modal feature differences.
[0044] In a preferred embodiment, the calculation formulas for the characteristic sequences of electroencephalogram (EEG) signals, electrooculogram (EOG) signals, and electrodermal (ED) signals are as follows: Where e represents electroencephalogram (EEG) signal, o represents electrooculogram (EOG) signal, and g represents electrodermal signal. This represents the characteristic sequence of the modal physiological signal corresponding to m. This represents the time window sample of the modal physiological signal corresponding to m. Let m represent the modality-specific encoding network for the physiological signal corresponding to the modality m, L represent the length of the token representation space, D represent the dimension of the token representation space, and R represent all real numbers.
[0045] Reference Figure 4 , Figure 4 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application. In some embodiments, the aforementioned pooling process of the feature sequences of peripheral physiological signals to obtain sample-level reliability gating coefficients includes: Step S170: Pool the feature sequences of peripheral physiological signals along the time dimension to obtain a sample-level global statistical vector; Step S171: Input the sample-level global statistical vector into the gated subnetwork to obtain the sample-level reliability gate coefficient with a value range of 0 to 1.
[0046] In this embodiment, as Figure 4As shown, in step S120, the feature sequence of the peripheral physiological signal can be pooled along the time dimension first. Pooling the feature sequence of the peripheral physiological signal along the time dimension yields a sample-level global statistical vector. For example, the recognition device performs global temporal pooling on the feature sequence of the electrodermal signal along the time dimension (the token represents the length L of the space), aggregating all temporal feature information within the entire time window, eliminating single-point fluctuation interference, and extracting a sample-level global statistical vector that can characterize the overall quality of the current sample.
[0047] The sample-level global statistical vector is input into the gating subnetwork to obtain sample-level reliability gating coefficients ranging from 0 to 1. For example, the gating subnetwork may include a first linear layer, a nonlinear activation function, a second linear layer, and a sigmoid function; the sigmoid function maps any real number to the range 0 to 1. The recognition device inputs the sample-level global statistical vector into the gating subnetwork to obtain sample-level reliability gating coefficients ranging from 0 to 1. When the response of the peripheral physiological signal feature sequence is unstable, the sample-level reliability gating coefficient tends to a smaller value, thereby reducing the matching weight of the peripheral physiological signal feature sequence in the preset emotion classifier. When the response of the peripheral physiological signal feature sequence is stable and has high emotion-related information, the sample-level reliability gating coefficient tends to a larger value, thereby preserving the matching weight of the peripheral physiological signal feature sequence in the preset emotion classifier.
[0048] In a preferred embodiment, the formula for calculating the sample-level reliability gating coefficient is as follows: in, Characteristic sequences representing skin electrical signals, This indicates that pooling operations are performed along the time dimension. Let B represent the sample-level global statistical vector, D represent the token representation space dimension, and R represent the set of all real numbers. Representation layer normalization operation, This represents the sample-level global statistical vector after layer normalization. This represents the sample-level reliability gating coefficient. Indicates the first linear layer. Represents a non-linear activation function. Indicates the second linear layer. This represents the Sigmoid function.
[0049] In some embodiments, the aforementioned recalibration of the peripheral physiological signal feature sequence based on the sample-level reliability gating coefficient to obtain the recalibrated peripheral feature sequence includes: The sample-level reliability gating coefficient is multiplied sample by sample by sample with the feature sequence of the peripheral physiological signal to adjust the feature weights of the feature sequence of the peripheral physiological signal, thus obtaining the recalibrated peripheral feature sequence.
[0050] In this embodiment, when performing step S130, the sample-level reliability gating coefficient can be multiplied sample-by-sample with the feature sequence of the peripheral physiological signal. The identification device multiplies the sample-level reliability gating coefficient with the feature sequence of the peripheral physiological signal sample-by-sample to adjust the feature weights of the feature sequence of the peripheral physiological signal, thereby obtaining a recalibrated peripheral feature sequence.
[0051] In a preferred embodiment, the formula for recalibrating the peripheral feature sequence is: in, This indicates the recalibration of peripheral feature sequences. This represents the sample-by-sample multiplication operator.
[0052] Reference Figure 5 , Figure 5 This is a flowchart illustrating another embodiment of the emotion recognition method based on multimodal physiological signals according to this application. In some embodiments, the aforementioned emotion recognition method based on multimodal physiological signals further includes: Step S180: Apply at least one of Gaussian noise, time masking, time offset, random dropout, and packet loss perturbation to the feature sequence of the peripheral physiological signal to obtain the perturbed peripheral feature sequence. Step S181: Generate sample-level reliability gating coefficients based on the interference peripheral feature sequence, and perform recalibration processing and input them into a preset emotion classifier in sequence to obtain the gating recognition result. Step S182: Input the interference peripheral feature sequence into a preset emotion classifier to obtain the recognition result without gating. Step S183: Compare the recognition results with and without gating to determine the magnitude of the decrease in recognition results.
[0053] In this embodiment, as Figure 5 As shown, when implementing emotion recognition methods based on multimodal physiological signals, interference can also be applied to the feature sequences of peripheral physiological signals. The recognition device can perform comparative testing during the offline training and performance verification stages of the model to simulate various abnormal interferences in real-world wearable device usage scenarios, and quantify the effectiveness of the sample-level reliability gating coefficient recalibration mechanism in suppressing noise and distortion signals.
[0054] At least one of Gaussian noise, time masking, time offset, random discarding, and packet loss perturbation is applied to the feature sequence of peripheral physiological signals to obtain an interfering peripheral feature sequence. For example, an identification device applies at least one of Gaussian noise, time masking, time offset, random discarding, and packet loss perturbation to the feature sequence of peripheral physiological signals to simulate real acquisition defects, thereby obtaining an interfering peripheral feature sequence. Gaussian noise is used to simulate random small-amplitude noise introduced by sensor circuits and environmental electromagnetic fields, and normally distributed random values are superimposed on the entire feature matrix. Time masking is used to randomly mask a continuous time token to simulate brief signal loss during acquisition. Time offset is used to shift the time token back and forth as a whole to simulate misalignment errors caused by asynchronous acquisition timing of multiple devices. Random discarding is used to randomly zero out some feature channel values to simulate intermittent electrode contact and signal attenuation. Packet loss perturbation is used to randomly delete some time segments in batches to simulate packet loss problems in data transmission of wireless wearable devices.
[0055] Based on the perturbation peripheral feature sequence, sample-level reliability gating coefficients are generated, and then recalibrated and input into a preset emotion classifier to obtain the gating recognition result. For example, the recognition device can, based on the perturbation peripheral feature sequence, perform mean pooling along the time dimension to obtain a global statistical vector, calculate the sample-level reliability gating coefficients in the 0 to 1 interval through layer normalization, two linear layers, and Sigmoid activation; then perform adaptive recalibration by broadcasting element-wise multiplication to obtain the noise-reduced and optimized recalibrated peripheral features; input the recalibrated peripheral features and the unperturbed EEG and EEG features into the trained emotion classifier, and output the gating recognition result.
[0056] The peripheral feature sequence of interference is input into a preset emotion classifier to obtain a recognition result without gating. For example, the recognition device skips the pooling, gating coefficient calculation, and recalibration adaptive correction process and directly inputs the noisy and distorted peripheral feature sequence of interference and EEG and EEG features into the same preset emotion classifier for prediction, and outputs a recognition result without gating correction and completely retaining the interference defects.
[0057] By comparing the recognition results with and without gating, the magnitude of the decline in recognition results is determined. For example, the recognition device extracts and quantifies the emotion classification accuracy, category confidence, and number of missamples for both gating and non-gating results. The accuracy decay and confidence decrease difference of the non-gating results relative to the gating results are calculated as the magnitude of the decline in recognition results. The sample-level reliability gating coefficient's inhibitory effect on the degradation of peripheral physiological signals is evaluated based on the magnitude of the decline.
[0058] The technical solution of this application obtains sample-level reliability gating coefficients by pooling the feature sequences of peripheral physiological signals; recalibrates the feature sequences of peripheral physiological signals based on the sample-level reliability gating coefficients to obtain recalibrated peripheral feature sequences; inputs the recalibrated peripheral feature sequences and the feature sequences of other modal physiological signals into a preset emotion classifier to obtain the emotion state recognition result; by recalibrating the feature sequences of peripheral physiological signals through sample-level reliability gating coefficients, the negative impact of interference such as motion artifacts and poor electrode contact is reduced, thereby improving the accuracy of emotion state recognition.
[0059] This application further proposes an emotion recognition device based on multimodal physiological signals, referring to... Figure 6 , Figure 6 This is a schematic diagram of the structure of an embodiment of the emotion recognition device based on multimodal physiological signals according to this application. In some embodiments, the emotion recognition device 30 based on multimodal physiological signals includes: The acquisition unit 300 is used to acquire the feature sequence of multimodal physiological signals corresponding to the target object. The multimodal physiological signals include at least central nervous system physiological signals and peripheral physiological signals. Pooling unit 310 is used to perform pooling processing on the feature sequence of the peripheral physiological signal to obtain sample-level reliability gating coefficients. The recalibration unit 320 is used to recalibrate the feature sequence of the peripheral physiological signal according to the sample-level reliability gating coefficient to obtain a recalibrated peripheral feature sequence. The classification unit 330 is used to input the recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals into a preset emotion classifier to obtain the emotion state recognition result.
[0060] In some embodiments, the multimodal physiological signals further include electrooculogram (EOG) signals, the central nervous system physiological signals include electroencephalogram (EEG) signals, and the peripheral physiological signals include electrodermal (ED) signals. The acquisition unit 300 is specifically used for: Collect multimodal physiological signals of the target object within a preset time period; Timing alignment is performed based on the start timestamp, sampling frequency, or sampling frame rate of each modality's physiological signal; The time-aligned physiological signals of each modality are subjected to at least one of the following signal normalization processes: filtering, baseline correction, artifact removal, standardization, and resampling. Based on the preset window length and preset window step size, the physiological signals of each modality after signal normalization are divided and processed to obtain time window samples of each modality of physiological signals. The time window samples of each modality physiological signal are input into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal.
[0061] In some embodiments, the modality-specific coding network includes a first coding network for the electroencephalogram (EEG) signals, a second coding network for the electrooculogram (EOG) signals, and a third coding network for the electrodermal (ED) signals. The acquisition unit 300, when executing the process of inputting time window samples of each modality physiological signal into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal, is specifically used for: The time window samples of the EEG signal are input into the first coding network to extract the features of the EEG signal, the time window samples of the EOL signal are input into the second coding network to extract the features of the EOL signal, and the time window samples of the EDS signal are input into the third coding network to extract the features of the EDS signal. The features of the electroencephalogram (EEG), electrooculogram (EOG), and electrodermal (ED) signals are mapped to time token sequences of uniform length and dimension to obtain feature sequences of the EEG, EOG, and EED signals.
[0062] In some embodiments, the pooling unit 310 is specifically used for: Pooling is performed on the feature sequences of the peripheral physiological signals along the time dimension to obtain a sample-level global statistical vector; The sample-level global statistical vector is input into the gating subnetwork to obtain the sample-level reliability gating coefficient, which has a value range of 0 to 1.
[0063] In some embodiments, the recalibration unit 320 is specifically used for: The sample-level reliability gating coefficient is multiplied sample by sample by sample with the feature sequence of the peripheral physiological signal to adjust the feature weights of the feature sequence of the peripheral physiological signal, thereby obtaining the recalibrated peripheral feature sequence.
[0064] In some embodiments, the emotion recognition device 30 based on multimodal physiological signals further includes: The interference unit is used to apply at least one of Gaussian noise, time masking, time offset, random dropout, and packet loss perturbation to the feature sequence of the peripheral physiological signal to obtain an interfering peripheral feature sequence. The pooling unit, recalibration unit, and classification unit are also used to generate sample-level reliability gating coefficients based on the interference peripheral feature sequence, and then perform recalibration processing and input them into a preset emotion classifier in sequence to obtain the gating recognition result. The classification unit is also used to input the interference peripheral feature sequence into a preset emotion classifier to obtain a recognition result without gating. The comparison unit is used to compare the recognition result with the gating and the recognition result without gating to determine the magnitude of the decrease in the recognition result.
[0065] This application further proposes an emotion recognition system based on multimodal physiological signals, referring to... Figure 7 , Figure 7 This is a schematic diagram of the structure of an embodiment of the emotion recognition system based on multimodal physiological signals according to this application. In some embodiments, the emotion recognition system based on multimodal physiological signals includes a data acquisition sensor 40 and the aforementioned emotion recognition device based on multimodal physiological signals 30.
[0066] In this embodiment, as Figure 7 As shown, the emotion recognition system based on multimodal physiological signals includes a data acquisition sensor 40 and the aforementioned emotion recognition device 30 based on multimodal physiological signals. The data acquisition sensor 40 may include a wearable headband device, electrooculography (EOG) electrodes, and a skin conductance sensor.
[0067] The above description is only a part or preferred embodiment of this application. Neither the text nor the drawings should limit the scope of protection of this application. All equivalent structural transformations made using the content of this application's specification and drawings under the overall concept of this application, or direct / indirect applications in other related technical fields, are included within the scope of protection of this application.
Claims
1. An emotion recognition method based on multimodal physiological signals, characterized in that, include: Obtain the feature sequence of multimodal physiological signals corresponding to the target object. The multimodal physiological signals include at least central nervous system physiological signals and peripheral physiological signals. The feature sequences of the peripheral physiological signals are pooled to obtain sample-level reliability gating coefficients; The peripheral physiological signal feature sequence is recalibrated based on the sample-level reliability gating coefficient to obtain the recalibrated peripheral feature sequence. The recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals are input into a preset emotion classifier to obtain the emotion state recognition result.
2. The emotion recognition method based on multimodal physiological signals according to claim 1, characterized in that, The multimodal physiological signals also include electrooculogram (EOG) signals, the central nervous system physiological signals include electroencephalogram (EEG) signals, and the peripheral physiological signals include electrodermal (ED) signals. The feature sequence of the multimodal physiological signal corresponding to the target object includes: Collect multimodal physiological signals of the target object within a preset time period; Timing alignment is performed based on the start timestamp, sampling frequency, or sampling frame rate of each modality's physiological signal; The time-aligned physiological signals of each modality are subjected to at least one of the following signal normalization processes: filtering, baseline correction, artifact removal, standardization, and resampling. Based on the preset window length and preset window step size, the physiological signals of each modality after signal normalization are divided and processed to obtain time window samples of each modality of physiological signals. The time window samples of each modality physiological signal are input into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal.
3. The emotion recognition method based on multimodal physiological signals according to claim 2, characterized in that, The modality-specific coding network includes a first coding network for the electroencephalogram (EEG) signals, a second coding network for the electrooculogram (EOG) signals, and a third coding network for the electrodermal (ED) signals. The step of inputting time window samples of each modality's physiological signal into the corresponding modality-specific coding network to obtain the feature sequence of each modality's physiological signal includes: The time window samples of the EEG signal are input into the first coding network to extract the features of the EEG signal, the time window samples of the EOL signal are input into the second coding network to extract the features of the EOL signal, and the time window samples of the EDS signal are input into the third coding network to extract the features of the EDS signal. The features of the electroencephalogram (EEG), electrooculogram (EOG), and electrodermal (ED) signals are mapped to time token sequences of uniform length and dimension to obtain feature sequences of the EEG, EOG, and EED signals.
4. The emotion recognition method based on multimodal physiological signals according to claim 1, characterized in that, The pooling process of the feature sequences of the peripheral physiological signals to obtain the sample-level reliability gating coefficients includes: Pooling is performed on the feature sequences of the peripheral physiological signals along the time dimension to obtain a sample-level global statistical vector; The sample-level global statistical vector is input into the gating subnetwork to obtain the sample-level reliability gating coefficient, which has a value range of 0 to 1.
5. The emotion recognition method based on multimodal physiological signals according to claim 4, characterized in that, The step of recalibrating the feature sequence of the peripheral physiological signal according to the sample-level reliability gating coefficient to obtain the recalibrated peripheral feature sequence includes: The sample-level reliability gating coefficient is multiplied sample by sample by sample with the feature sequence of the peripheral physiological signal to adjust the feature weights of the feature sequence of the peripheral physiological signal, thereby obtaining the recalibrated peripheral feature sequence.
6. The emotion recognition method based on multimodal physiological signals according to any one of claims 1 to 5, characterized in that, The emotion recognition method based on multimodal physiological signals also includes: At least one of Gaussian noise, time masking, time shift, random dropout, and packet loss perturbation is applied to the feature sequence of the peripheral physiological signal to obtain the interfering peripheral feature sequence; Based on the interference peripheral feature sequence, sample-level reliability gating coefficients are generated, and then recalibrated and input into a preset emotion classifier to obtain the gating recognition result. The interference peripheral feature sequence is input into a preset emotion classifier to obtain a recognition result without gating. By comparing the recognition results with and without gating, the magnitude of the decrease in recognition results is determined.
7. An emotion recognition device based on multimodal physiological signals, characterized in that, include: The acquisition unit is used to acquire the feature sequence of multimodal physiological signals corresponding to the target object. The multimodal physiological signals include at least central nervous system physiological signals and peripheral physiological signals. A pooling unit is used to perform pooling processing on the feature sequence of the peripheral physiological signal to obtain sample-level reliability gating coefficients. The recalibration unit is used to recalibrate the feature sequence of the peripheral physiological signal according to the sample-level reliability gating coefficient to obtain a recalibrated peripheral feature sequence. The classification unit is used to input the recalibrated peripheral feature sequence and the feature sequences of other modal physiological signals into a preset emotion classifier to obtain the emotion state recognition result.
8. The emotion recognition device based on multimodal physiological signals according to claim 7, characterized in that, The multimodal physiological signals also include electrooculogram (EOG) signals, the central nervous system physiological signals include electroencephalogram (EEG) signals, and the peripheral physiological signals include electrodermal (ED) signals. The acquisition unit is specifically used for: Collect multimodal physiological signals of the target object within a preset time period; Timing alignment is performed based on the start timestamp, sampling frequency, or sampling frame rate of each modality's physiological signal; The time-aligned physiological signals of each modality are subjected to at least one of the following signal normalization processes: filtering, baseline correction, artifact removal, standardization, and resampling. Based on the preset window length and preset window step size, the physiological signals of each modality after signal normalization are divided and processed to obtain time window samples of each modality of physiological signals. The time window samples of each modality physiological signal are input into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal.
9. The emotion recognition device based on multimodal physiological signals according to claim 8, characterized in that, The modality-specific coding network includes a first coding network for the electroencephalogram (EEG) signals, a second coding network for the electrooculogram (EOG) signals, and a third coding network for the electrodermal (ED) signals. The acquisition unit, when executing the process of inputting time window samples of each modality physiological signal into the corresponding modality-specific coding network to obtain the feature sequence of each modality physiological signal, is specifically used for: The time window samples of the EEG signal are input into the first coding network to extract the features of the EEG signal, the time window samples of the EOL signal are input into the second coding network to extract the features of the EOL signal, and the time window samples of the EDS signal are input into the third coding network to extract the features of the EDS signal. The features of the electroencephalogram (EEG), electrooculogram (EOG), and electrodermal (ED) signals are mapped to time token sequences of uniform length and dimension to obtain feature sequences of the EEG, EOG, and EED signals.
10. An emotion recognition system based on multimodal physiological signals, characterized in that, The emotion recognition system based on multimodal physiological signals includes a data acquisition sensor and an emotion recognition device based on multimodal physiological signals as described in any one of claims 7 to 8.