Emotion recognition method for multi-modal physiological signals based on multi-scale convolution and joint attention mechanism
This method for multimodal physiological signal emotion recognition, which employs lightweight multi-scale convolution and joint attention mechanisms, solves the problems of cumbersome models and simple fusion mechanisms in existing technologies, and achieves efficient and real-time multimodal emotion monitoring on wearable devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
- Filing Date
- 2026-04-15
- Publication Date
- 2026-08-04
AI Technical Summary
Existing emotion recognition methods based on multimodal physiological signals are bulky, have simple fusion mechanisms, are difficult to deploy on resource-constrained wearable devices, and have high computational costs, making it impossible to achieve long-term real-time monitoring.
A multimodal physiological signal emotion recognition method employing lightweight multi-scale convolution and joint attention mechanisms is proposed. This method includes preprocessing, lightweight multi-scale convolution feature extraction, and lightweight joint attention feature fusion, constructing an end-to-end ultra-lightweight network architecture suitable for wearable devices.
While ensuring high recognition accuracy, it achieves a small number of model parameters and low computational cost, enabling real-time multimodal emotion monitoring on wearable devices.
Smart Images

Figure CN122498844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of physiological signal processing technology, and in particular to a multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanisms. Background Technology
[0002] Emotion recognition has wide applications in fields such as healthcare and human-computer interaction. Physiological signals, due to their objectivity and difficulty in being faked, serve as a reliable basis for emotion recognition. Existing emotion recognition methods based on multimodal physiological signals mainly suffer from three problems: bulky models, simple fusion mechanisms, and a lack of wearable design.
[0003] Specifically, based on Transformer Mainstream deep convolutional neural network models have large parameter counts and high computational costs, making them difficult to deploy on resource-constrained wearable devices. Furthermore, fusion mechanisms often employ static methods such as splicing and weighted summation, which cannot effectively handle the heterogeneity and complementarity of different modal physiological signals. They also lack optimization for sensor channel simplification and fail to consider hardware resource constraints such as microcontrollers, making long-term real-time monitoring difficult. While some lightweight methods have attempted to reduce model complexity, their designs are mostly single-modal, and the computational cost remains high for wearable devices, while also lacking efficient multimodal fusion mechanisms.
[0004] Therefore, how to achieve multimodal emotion recognition with small parameters and low computational cost while ensuring high recognition accuracy is a technical problem that urgently needs to be solved in this field.
[0005] Therefore, existing technologies still need to be improved and enhanced. Summary of the Invention
[0006] The main objective of this invention is to provide a multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism, which aims to solve the problems of bulky models and simple fusion mechanisms that are common in existing multimodal physiological signal emotion recognition methods.
[0007] To achieve the aforementioned objective, a first aspect of the present invention provides a method for multimodal physiological signal emotion recognition based on multi-scale convolution and joint attention mechanisms, wherein the method comprises: Acquire one or more sets of multimodal physiological signals, including electroencephalogram (EEG) signals and peripheral physiological signals; The multimodal physiological signals are input into a target classification model, which includes a preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module. The multimodal physiological signal is preprocessed based on the preprocessing module to obtain a multimodal input tensor; The feature extraction module performs lightweight multi-scale convolution feature extraction on the multimodal input tensor to obtain the target convolution feature; The target convolutional features are fused using a lightweight joint attention feature fusion method based on the feature fusion module to obtain the target fused features; Based on the emotion classification module, the target fusion features are classified to obtain the target emotion classification result.
[0008] In a second aspect, the present invention provides a terminal, wherein the terminal includes: a memory, a processor, and a multimodal physiological signal emotion recognition program based on a multi-scale convolution and joint attention mechanism stored in the memory and executable on the processor, wherein when the multimodal physiological signal emotion recognition program based on the multi-scale convolution and joint attention mechanism is executed by the processor, it implements the steps of the multimodal physiological signal emotion recognition method based on the multi-scale convolution and joint attention mechanism as described above.
[0009] A third aspect of the present invention provides a computer-readable storage medium, wherein the computer storage medium stores one or more programs that can be executed by one or more processors to implement the steps of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism described in any of the preceding claims.
[0010] Beneficial Effects: Compared with existing technologies, this invention provides a method for emotion recognition based on multimodal physiological signals using multi-scale convolution and joint attention mechanisms. In this method, when performing emotion recognition based on multimodal physiological signals, one or more sets of multimodal physiological signals are first acquired. These signals include electroencephalogram (EEG) signals and peripheral physiological signals. Then, the multimodal physiological signals are input into a target classification model, which includes a preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module. Next, the multimodal physiological signals are preprocessed using the preprocessing module to obtain a multimodal input tensor. Lightweight multi-scale convolution feature extraction is performed on the multimodal input tensor using the feature extraction module to obtain target convolution features. Lightweight joint attention feature fusion is performed on the target convolution features using the feature fusion module to obtain target fused features. Finally, emotion classification is performed on the target fused features using the emotion classification module to obtain the target emotion classification result. This invention provides users with a multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanisms, solving the problems of cumbersome models and simple fusion mechanisms commonly found in existing multimodal physiological signal emotion recognition methods. In this invention, through the collaborative design of lightweight multi-scale convolution feature extraction and lightweight joint attention feature fusion, high-precision multimodal physiological signal emotion recognition is achieved with extremely low parameter and computational requirements. Attached Figure Description
[0011] Figure 1 A flowchart illustrating an embodiment of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided by the present invention; Figure 2 The overall flowchart of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided by the present invention is shown below. Figure 3 The model architecture diagram of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided by this invention is shown below. Figure 4 The feature fusion module architecture diagram of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided by this invention is shown below. Figure 5 Results analysis of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this invention. Figure 1 ; Figure 6 Results analysis of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this invention. Figure 2; Figure 7 Results analysis of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this invention. Figure 3 ; Figure 8 Results analysis of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this invention. Figure 4 ; Figure 9 Results analysis of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this invention. Figure 5 ; Figure 10 A schematic diagram of the operating environment of an embodiment of the terminal provided by the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0013] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0014] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0015] This invention provides a multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism, which can be applied to terminals with computing capabilities. The terminal can execute the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided by this invention to evaluate the emotions of one or more users based on their multimodal physiological signals.
[0016] Example 1 This embodiment provides a multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism, belonging to the field of physiological signal processing technology. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this embodiment utilizes innovative lightweight multi-scale convolution (… LMSC ), lightweight joint attention ( LJA ) and dynamic feature fusion network ( DFFN The collaborative design of the model ensures or surpasses the recognition accuracy of existing advanced models while reducing the number of parameters and computational complexity to an extremely low level, enabling it to run on wearable embedded devices (such as microcontrollers) with strictly limited resources and achieve real-time and accurate emotion monitoring. Among existing technologies, emotion recognition is a key technology for understanding human cognitive states and decision-making behaviors, and it has broad application prospects in fields such as healthcare, human-computer interaction, and psychological monitoring. Physiological signals, because they objectively reflect the activity of the central and autonomic nervous systems and are difficult to fake subjectively, have advantages over non-physiological signals such as facial expressions and voice in special scenarios such as deep-sea diving, elderly care, and the diagnosis of mood disorders. Therefore, emotion recognition based on physiological signals has become a research hotspot.
[0017] Currently, emotion recognition research based on physiological signals is mainly divided into two categories: unimodal and multimodal. Unimodal research primarily uses electroencephalography (EEG), and while some progress has been made, a single signal is insufficient to comprehensively capture the multidimensional features of complex emotional states. Multimodal research improves the robustness and accuracy of recognition by integrating multiple physiological signals; however, existing technologies still have the following limitations: First, the models have high computational complexity and a large number of parameters. Most current mainstream models are based on... Transformer First, some methods rely on deep convolutional neural networks, typically with millions or even tens of millions of parameters, resulting in massive computational demands and making them difficult to deploy on resource-constrained wearable devices. Second, the multimodal fusion mechanisms are simplistic and cannot effectively handle signal heterogeneity. Most methods employ static fusion techniques such as splicing and weighted summation, failing to dynamically capture the complementary relationships between modalities and exhibiting weak processing capabilities for the heterogeneous characteristics between EEG signals and peripheral physiological signals. Third, there is a lack of design for wearable devices. Existing studies mostly use standard 32-channel EEG for experiments, without considering the impact of sensor channel simplification on model performance. Furthermore, the models are not optimized for resource-constrained hardware such as microcontrollers, making long-term, real-time emotion monitoring difficult.
[0018] To address the above problems, existing lightweight methods include... Bi - CapsNet , LResCapsule Techniques such as binarization and lightweight residual networks are used to reduce model complexity, but their designs are mostly aimed at single-modality applications, especially EEG, and the computational cost remains high for some wearable devices. Furthermore, while attention-based multimodal fusion methods can improve fusion results, their fusion modules themselves still have high computational overhead, making it difficult to achieve model lightweighting while maintaining recognition accuracy.
[0019] Therefore, how to achieve multimodal emotion recognition with small model parameters, low computational cost, and wearable device compatibility while ensuring high recognition accuracy is a technical problem that urgently needs to be solved in this field.
[0020] Specifically, such as Figure 1 As shown in this embodiment, one example of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this embodiment includes the following steps: S 100. Acquire one or more sets of multimodal physiological signals, wherein the multimodal physiological signals include electroencephalogram (EEG) signals and peripheral physiological signals.
[0021] Specifically, in emotion recognition, one or more sets of multimodal physiological signals are first acquired, including electroencephalogram (EEG) signals and peripheral physiological signals. The EEG signals are neurophysiological signals used to reflect the electrical activity of the central nervous system; the peripheral physiological signals include at least one of electrooculogram (EOG), electromyogram (EMG), electroskin response (ESR), electrocardiogram (ECG), respiratory signals, blood volume signals, and temperature signals, used to reflect the activity of the autonomic nervous system and the peripheral physiological state.
[0022] In more embodiments, the multimodal physiological signals are acquired via wearable sensors for subsequent multimodal physiological signal emotion recognition.
[0023] Before acquiring one or more sets of multimodal physiological signals, the method further includes: An initial classification model is constructed, which includes a preprocessing module, a feature extraction module, a feature fusion module, and a sentiment classification module. The feature extraction module is configured to perform lightweight multi-scale convolutional feature extraction, and the feature fusion module is configured to perform lightweight joint attention feature fusion. After obtaining the target dataset and constructing labels for the target dataset based on a preset classification task, the target dataset after label construction is divided into a first data set and a second data set based on a subject-dependent strategy and a full dataset strategy. Both the first data set and the second data set contain a training set and a test set. The initial classification model is trained and validated based on the first dataset and the second dataset, respectively, to obtain the target classification model.
[0024] Specifically, in this embodiment, before emotion recognition is required for one or more sets of multimodal physiological signals, an initial classification model is first constructed and trained to obtain a target classification model.
[0025] Specifically, the initial classification model includes a preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module. The feature extraction module is configured to perform lightweight multi-scale convolutional feature extraction to extract multi-scale spatial features from physiological signals of each modality. The feature fusion module is configured to perform lightweight joint attention feature fusion to model intra-modal dependencies and inter-modal complementarities among the features of each modality, and outputs the fused features to the emotion classification module. The emotion classification module is configured to classify the fused features and output the emotion recognition result.
[0026] Reference Figure 2 During the training phase, the input data consists of multimodal physiological signals from a publicly available dataset. Specifically, the target dataset is obtained by acquiring and preprocessing the multimodal physiological signals from the publicly available dataset. The multimodal physiological signals include electroencephalogram (EEG) signals and peripheral physiological signals.
[0027] In this embodiment, the public dataset includes DEAP Dataset DREAMER Datasets and WESAD Dataset. Among them, DEAP The dataset contains multimodal physiological signals, including electroencephalogram (EEG) signals, recorded by 32 participants while watching 40 video clips. EEG ), electrooculogram (EOG) EOG ), electromyographic signals ( EMG ), skin conductance signal ( GSR ), respiratory signals ( Resp Blood volume measured by plethysmography (PQC) Pleth ) and body temperature signal ( Temp All signals were downsampled to 128. Hz ; DREAMER The dataset contains multimodal physiological signals, including electroencephalogram (EEG) and electrocardiogram (ECG) signals, recorded by 23 participants while watching 18 video clips. The EEG signal sampling rate was 256. Hz The ECG signal sampling rate is 128. Hz ; WESADThe dataset contains chest and wrist physiological signals from 15 subjects. Chest signals include electrocardiogram (ECG), electrodermal activity (EDA), electromyography (EMG), respiratory signals, and body temperature signals, with a sampling rate of 700. Hz Wrist signals include blood volume and pulse signals ( BVP ), skin electrical activity signals ( EDA (and body temperature signals, with sampling rates ranging from 4 to 64) Hz .
[0028] Then, the obtained public dataset is preprocessed. For DEAP The dataset uses modalities including electroencephalogram (EEG) signals (32 channels), electrooculogram (EOG) signals (2 channels), electromyogram (EMG) signals (2 channels), and electrodermal conductance (EDC) signals (1 channel); for DREAMER The dataset uses modalities including electroencephalogram (EEG) signals (14 channels) and electrocardiogram (ECG) signals (2 channels); for WESAD The dataset uses modalities including ECG, electrodermal activity, and electromyography (EMG) signals from the wrist. Baseline correction was performed on each modality's data. DEAP Datasets and DREAMER Subtract the mean baseline period of 3 seconds from the dataset. WESAD The dataset does not undergo baseline processing according to the official documentation.
[0029] Next, labels are constructed for the preprocessed target dataset based on preset classification tasks, and the dataset is divided using two partitioning strategies. The preset classification tasks include binary classification, ternary classification, and quaternary classification. DEAP The dataset includes a binary classification task based on a subject-dependent strategy, using a 5-point threshold to divide valence and arousal scores into positive / high and negative / low categories. A three-class classification task uses the entire dataset as a strategy, classifying scores of 1-3 as negative / low, 4-6 as neutral, and 7-9 as positive / high. A four-class classification task, also based on a subject-dependent strategy, uses a 5-point threshold to classify scores into four categories: low valence / low arousal, low valence / high arousal, high valence / low arousal, and high valence / high arousal, collectively known as the valence-arousal four-class classification task (VA). DREAMER The dataset is divided into four categories: binary classification (valence and arousal) with a threshold of 3; triadic classification (1-2 points for negative / low, 3 points for neutral, and 4-5 points for positive / high); and quadruple classification (3 points for four categories). WESAD The dataset includes three-class classification (baseline, stress, leisure) and four-class classification (baseline, stress, leisure, meditation).
[0030] The target dataset, after label construction, is divided into a first dataset and a second dataset based on a subject-dependent strategy and a full dataset strategy. Both the first dataset and the second dataset contain training and testing sets. Specifically, when dividing based on the subject-dependent strategy, data from the same subject is divided into training and testing sets; when dividing based on the full dataset strategy, data from all subjects are mixed before being divided into training and testing sets.
[0031] Finally, the initial classification model is trained using the training sets of the first dataset and the second dataset, respectively, and the trained model is validated using the test sets of the first dataset and the second dataset, respectively. In this way, the target classification model is obtained.
[0032] S 200. The multimodal physiological signals are input into a target classification model, which includes a preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module.
[0033] Specifically, one or more sets of the multimodal physiological signals, after being collected and preprocessed, are input into the pre-constructed and trained target classification model. The target classification model adopts an end-to-end ultra-lightweight network architecture, consisting of a cascaded preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module, as detailed below. Figure 3 .exist Figure 3 Chinese abbreviation LMSC This represents a lightweight multi-scale convolution module, which is the feature extraction module. LJA This represents the lightweight joint attention module, also known as the feature fusion module. Figure 3 The figure in a This describes the overall structure and workflow of the target classification model. Figure 3 The figure in b ) is the feature extraction module LMSC Structure and examples; Figure 3 The figure in c ) is the feature fusion module LJA The structure of the module.
[0034] The preprocessing module performs data cleaning, format unification, and dimensionality transformation on the physiological signals of each modality, outputting the input tensors of each modality. The feature extraction module uses a lightweight multi-scale convolutional structure to extract multi-scale spatial features from the input tensors of each modality, outputting the convolutional features of each modality. The feature fusion module uses a lightweight joint attention mechanism to model intra-modal dependencies and inter-modal complementarities of the convolutional features of each modality, outputting the fused multimodal features. The emotion classification module performs classification mapping on the fused multimodal features, outputting the final emotion recognition result.
[0035] S 300. Based on the preprocessing module, the multimodal physiological signal is preprocessed to obtain a multimodal input tensor.
[0036] Specifically, the collected multimodal physiological signals are input into the preprocessing module, which performs preprocessing operations on each modality of physiological signal.
[0037] In this embodiment, the preprocessing operations include sampling rate unification, baseline correction, and dimensionality transformation. In other embodiments, the preprocessing may include, but is not limited to, more methods.
[0038] Specifically, refer to Figure 2 During the model training phase, the input data consists of multimodal physiological signals from a publicly available dataset. After the target classification model is trained, the input data can be a validation or test set derived from the publicly available dataset, or it can be one or more sets of multimodal physiological signals extracted in real-time or non-real-time from one or more target users after testing. These multimodal physiological signals include electroencephalogram (EEG) signals and peripheral physiological signals. In more applications, the input data can also be a set of multimodal physiological signals collected in real-time from the target user's wearable device, used to identify the target user's emotions in real-time based on the target classification model mounted on the wearable device. In other words, based on the target classification model provided in this embodiment, emotion recognition of multiple sets of multimodal physiological signals can be performed simultaneously. Alternatively, the target classification model can be lightweighted and mounted on a wearable device to identify the multimodal physiological signals acquired in real-time on the wearable device, thereby obtaining the target user's emotional state in real-time.
[0039] Specifically, after acquiring one or more sets of the multimodal physiological signals, the sampling rate of each modal physiological signal is first uniformized to unify the signals with different sampling frequencies to a preset target sampling frequency, so as to ensure the alignment of each modal signal in the time dimension and facilitate subsequent multimodal fusion processing.
[0040] Then, baseline correction processing is performed on the physiological signals of each modality after the sampling rate is unified, and the average baseline period of the preset time length is subtracted to eliminate the influence of static offset and individual differences in the signal on emotion recognition.
[0041] Finally, the physiological signals of each modality, after sampling rate unification and baseline correction, are converted into a preset data format to obtain the input tensor for each modality. The input tensor includes a batch dimension, a time dimension, and a signal channel dimension, where the signal channel dimension corresponds to the number of acquisition channels for each modality's physiological signals, and the time dimension corresponds to the number of sampling points. Through the above preprocessing operations, the preprocessing module outputs multimodal input tensors with a unified structure and standardized format for subsequent feature extraction module processing.
[0042] S 400. Based on the feature extraction module, perform lightweight multi-scale convolution feature extraction on the multimodal input tensor to obtain the target convolution feature; The feature extraction module performs lightweight multi-scale convolution feature extraction on the multimodal input tensor to obtain target convolution features, including: S 410. The multimodal input tensor is input to the feature extraction module, which includes at least two parallel convolutional branches. Each convolutional branch uses a convolutional kernel of a different size to extract spatial features of different receptive fields along the signal channel dimension. S 420. Perform convolution operations on the multimodal input tensor through the at least two parallel convolution branches to obtain the spatial features of each convolution branch; S 430. After normalizing and applying activation functions to the spatial features of each convolutional branch, the features are fused using matrix addition to obtain the target convolutional features of each modality in the multimodal physiological signal.
[0043] Specifically, the modal input tensors output from the preprocessing module are input to the feature extraction module. The feature extraction module employs a lightweight multi-scale convolutional structure to extract multi-scale spatial features from the physiological signals of each modality. The feature extraction module includes at least two parallel convolutional branches, each using a convolutional kernel of a different size, to extract spatial features of different receptive fields along the signal channel dimension. Through this module, multi-scale features of physiological signals in the spatial dimension can be captured with extremely low computational complexity, providing rich intra-modal feature representations for subsequent multimodal fusion.
[0044] The lightweight multi-scale convolutional feature extraction specifically includes the following sub-steps: S410. The multimodal input tensor is input to the feature extraction module, which includes at least two parallel convolutional branches. Each convolutional branch uses a convolutional kernel of a different size to extract spatial features of different receptive fields along the signal channel dimension.
[0045] In a preferred embodiment of the present invention, the feature extraction module includes five parallel convolutional branches, each employing convolutional kernels of sizes 1, 3, 5, 7, and 9. Each convolutional kernel performs convolutional operations along the signal channel dimension (i.e., the spatial dimension), and in the temporal dimension, each one-dimensional convolutional kernel integrates information from all time steps. This design enables the model to simultaneously capture both local fine features and global macroscopic features, and avoids parameter inflation caused by deep networks due to the use of a single-layer parallel structure.
[0046] S 420. Perform convolution operations on the multimodal input tensor through at least two parallel convolution branches to obtain the spatial features of each convolution branch.
[0047] Specifically, physiological signals exhibit dynamic changes in both time and channel dimensions. Traditional convolutional neural networks have limited feature extraction capabilities, while existing multi-scale spatiotemporal designs are typically accompanied by high computational complexity. In contrast, the feature extraction module proposed in this embodiment is a lightweight multi-scale convolution module. It adopts a simpler single-layer architecture and, through unidirectional multi-scale convolution operations, not only efficiently extracts features and enhances the robustness and generalization ability of the model, but also consciously maintains low computational complexity.
[0048] In this embodiment, the input tensors of each modality are fed into each convolutional branch of the feature extraction module, and each branch independently performs a one-dimensional convolution operation. The convolution operation slides along the signal channel dimension to extract spatial features within the receptive field at that scale. For each branch, the convolution output retains the batch size and temporal dimension length of the input tensor, while the channel dimension is transformed according to the number of convolutional kernels. After convolution, each branch outputs the corresponding spatial feature map.
[0049] Specifically, let Representing modes The input tensor, where For batch size, Number of signal channels This is the time series length (corresponding to the sampling frequency). It is input to the feature extraction module. LMSC Before the module, first of all Dimensional transformation The output after processing by a lightweight multi-scale convolution module. Given by the following formula:
[0050] in . This represents a multi-scale convolution kernel with a size set of {1, 3, 5, 7, 9}, and each parallel convolution branch uses a convolution kernel of a different size.
[0051] The feature extraction module LMSC The module is a single-layer, five-branch network structure. These convolutional kernels of different sizes define a series of receptive fields of different sizes along the physiological signal channel dimension (i.e., the spatial dimension), an operation independent of the temporal sampling rate. In the temporal dimension, each one-dimensional convolutional kernel integrates information from all time steps. Finally, the feature matrices extracted by the five parallel branches are fused through matrix addition.
[0052] Then, the spatial features output by each convolutional branch are sequentially subjected to batch normalization and... GELU Activation function processing accelerates model convergence and introduces non-linear expressive power. Batch normalization is used to stabilize the training process. GELU Activation functions, acting as nonlinear transformation layers, enhance the model's representational capabilities. Optionally, additional processing can be added after activation function processing. Dropout Layers are used to prevent overfitting.
[0053] After the above processing, the outputs of each branch are fused element-wise using matrix addition, that is, the corresponding positions of the feature maps of each branch are added together to obtain the fused features of that modality. The fused features are then processed again... Dropout The process yields the final target convolutional features for each modality. These target convolutional features remain unchanged in the spatial dimension, while the temporal dimension is compressed to a preset length (e.g., 30 time points).
[0054] Specifically, each modality data is processed by the feature extraction module. LMSC The output after module processing can be represented as:
[0055] Here, BN Indicates batch normalization ( BatchNorm ), Conv This indicates a convolution operation. LMSC After processing, spatial dimensions Remain unchanged, while the time dimension Then it is reduced to 30.
[0056] Through the above lightweight multi-scale convolutional feature extraction, the feature extraction module outputs the target convolutional features of each modality, which are then used by the subsequent feature fusion module for joint attention fusion within and between modalities.
[0057] S500. Based on the feature fusion module, perform lightweight joint attention feature fusion on the target convolutional features to obtain the target fused features.
[0058] The target convolutional features are fused using a lightweight joint attention feature fusion method based on the feature fusion module to obtain the target fused features, including: S 510. Input the target convolutional features of each modality into the feature fusion module, which includes multiple intra-modal attention units and multiple inter-modal attention units.
[0059] Specifically, refer to Figure 4 The feature fusion module employs a lightweight joint attention mechanism, including multiple intra-modal attention units and multiple inter-modal attention units. Specifically, the feature extraction module... LMSC After feature extraction by the module, the features are processed through an efficient, lightweight joint attention mechanism. LJA The feature fusion module further processes the data through a multimodal fusion mechanism to more effectively capture complementary cross-modal information and enhance the integration capability of multimodal features. LJA The structure mainly consists of multiple intramodal attention units. Intra - LJA With multiple intermodal attention units Inter - LJA It consists of its interaction mechanism. Furthermore, the main difference between these two key modules lies in the input type: Intra - LJA It is primarily used for processing single-modal data, while Inter - LJA Used for processing cross-modal data. That is, Intra - LJA and Inter - LJA These are used to capture the time-channel dependencies within a modality and the complementarity relationships between modalities, respectively, and finally, a weighted fusion is used to output a comprehensive multimodal feature.
[0060] S 520. Based on the intra-modal attention unit, perform intra-modal dependency modeling on the target convolutional features of each modality to obtain intra-modal enhancement features of each modality, wherein the intra-modal attention unit includes a multi-head attention mechanism and a dynamic feature fusion network, and the dynamic feature fusion network includes parallel linear transformation branches and one-dimensional convolution branches; Specifically, based on the intra-modal attention unit, intra-modal dependency modeling is performed on the target convolutional features of each modality to obtain the intra-modal enhancement features of each modality, including: S521. After reshaping the target convolutional features of each modality into the target shape, add positional encoding to obtain the positional encoding features of each modality; S 522. Input the location encoding feature into the first multi-head attention subunit, calculate the attention weight of the location encoding feature through multiple attention heads in the first multi-head attention subunit, and concatenate the attention weights output by each attention head to obtain the attention concatenation feature of each modality; S 523. Input the attention splicing features of each modality into the first dynamic feature fusion network, the first dynamic feature fusion network including parallel linear transformation branches and one-dimensional convolution branches; S 524. The attention splicing features of each modality are extracted and fused using the first dynamic feature fusion network to obtain intra-modal enhancement features for each modality.
[0061] In this embodiment, the intramodal attention unit is... Intra - LJA Let's take an example to explain in detail. Before entering the intramodal attention unit, the features... Reshaped into the target shape: Then, the input is given to the intramodal attention unit for intramodal attention fusion.
[0062] Features After adding positional encoding, the positional encoding features of each modality are obtained. ,in This represents a position vector. The modified position encoding feature. They are then fed into multiple first multi-head attention sub-units. Each first multi-head attention sub-unit contains... Each encoder layer contains a multi-head attention mechanism. Size. To ensure emotion recognition performance while keeping the model as lightweight as possible, in this embodiment, the following settings are made: , In many other embodiments, and It can be set to larger or smaller to make the target classification model more accurate or more lightweight.
[0063] Specifically, the position encoding features are input into the first encoder layer of the first multi-head attention subunit. In this embodiment, the input data of the first encoder layer is denoted as... In the There are encoder layers, and the data is represented as... The first multi-head attention subunit is the first... The latent features learned by the size, that is, the first The attention weight of the positional encoding feature corresponding to each head is denoted as... ,in This feature is calculated using the following formula:
[0064] in, , and These are the weight matrices for query, key, and value transformations, respectively. This is the dimension of the key in each attention head. u All in each encoder layer J The outputs of each head are concatenated to form the attention concatenation feature, which is calculated using the following formula:
[0065] Application of the attention splicing features to the spliced data Dropout Then, the "Add & LayerNorm" operation is performed, which can be expressed as:
[0066] in, LN Representation layer normalization.
[0067] After obtaining the normalized attention-concatenated features of each modality, the attention-concatenated features of each modality are input into the first dynamic feature fusion network. The first dynamic feature fusion network includes parallel linear transformation branches and one-dimensional convolution branches, used to replace traditional... Transformer Feedforward networks in [the context of] [the network].
[0068] Specifically, although the standard Transformer Feedforward networks are widely used, but they still suffer from high complexity in multimodal emotion recognition and struggle to effectively model the intermodal relationships of heterogeneous physiological signals. Therefore, in this embodiment, a dual-branch feedforward network is introduced. This structure replaces the single feedforward network in the bottleneck architecture with parallel linear transformation branches and one-dimensional convolution branches, creating a parameter-efficient and highly adaptable first dynamic feature fusion network for multimodal fusion. DFFN . DFFN It consists of two linear branches, two one-dimensional convolutional branch structures, and two stages, each stage containing a linear transformation branch. LB A one-dimensional convolution branch CB And a dynamic residual structure.
[0069] Specifically, in the first stage, the input attention splicing features The output of the first dynamic feature fusion network It is given by the following formula:
[0070] Here, will and and The results of the linear transformations are added together, and then... LN Normalization is performed. LB is the linear branch, and CB is the one-dimensional convolution branch. In the second stage, the output... It is given by the following formula:
[0071] LB The structure contains two linear layers and two Dropout Layer and one GELU Activation function. For the first stage, it is calculated as follows:
[0072] Similarly, CB The structure consists of two one-dimensional convolutional layers and two... Dropout Layer and one GELU The activation function consists of the following components. The calculation for the first stage is as follows:
[0073] CB The one-dimensional convolutional layers in the structure use a convolutional kernel of size 1 in the first stage and a convolutional kernel of size 3 in the second stage, which enables the network to capture a wider range of temporally dependent features. LB and CB The bottleneck architecture of the structure reduces the number of parameters and computational complexity while ensuring performance, thus enhancing multi-task applicability and model efficiency. (This is achieved through the...) After one encoder layer, the internal LJA The module's final output is the intra-modal enhancement features for each modality. .
[0074] S 530. Based on the intermodal attention unit, cross-modal dependency modeling is performed on the target convolutional features of each modality to obtain intermodal complementary features of each modality. The intermodal attention unit uses the target convolutional features of the target modality as the query and the target convolutional features of other modalities as the key and value, and extracts cross-modal features through a multi-head attention mechanism and a dynamic feature fusion network. Specifically, based on the inter-modal attention unit, cross-modal dependency modeling is performed on the target convolutional features of each modality to obtain inter-modal complementary features of each modality, including: S 531. Use the target convolutional features of each modality as the target query, and the target convolutional features of other modalities as the target key and target value; S 532. After adding position codes to the target query, the target key and the target value respectively, they are input into the second multi-head attention subunit. The cross-modal attention weights of each modality are calculated by multiple attention heads in the second multi-head attention subunit, and the outputs of each attention head are concatenated to obtain the cross-modal attention concatenation features of each modality. S 533. The cross-modal attention splicing features are input into the second dynamic feature fusion network, which includes parallel linear transformation branches and one-dimensional convolution branches; S 534. The cross-modal attention splicing features are extracted and fused through the second dynamic feature fusion network to obtain intermodal complementary features of each modality.
[0075] It can be seen that the intermodal attention unit Inter - LJA With the intramodal attention unit Intra - LJA The steps for modeling intra-modal dependencies are largely similar, but the key difference lies in the inter-modal attention unit. Inter - LJA Focusing on cross-modal attention fusion, it can effectively integrate complementary information from different physiological modalities. In this embodiment, multiple feature extraction modules can be used in the multi-scale feature extraction stage. LMSC The module independently extracts the temporal features of each modality. The features of any modality are denoted as... The features of all other modalities are represented as For example, regarding the aforementioned electroencephalogram (EEG) signals, i.e. EEG Modality, its own characteristics are Other modalities are characterized as follows :
[0076] Each of the intermodal attention units Inter - LJA by and As input, these inputs undergo positional encoding and embedding before being processed by the multi-head attention mechanism. In the inter-modal attention unit... Inter - LJA middle, As a query, and As keys and values, this module learns cross-modal relationships to address challenges such as differences in feature distributions and temporal misalignment. LJA The module's final output is .
[0077] exist LJA The final stage of feature fusion, Intra - LJA and Inter - LJA The outputs are integrated through weighted summation. and Redefined and The outputs of the four modes are concatenated as follows:
[0078]
[0079] The final output after weighted fusion is:
[0080] in, It is used for fusion and The weight matrix allows each modality to refine emotion-related features while integrating complementary information.
[0081] S 540. The intra-modal enhancement features and inter-modal complementary features of each modality are weighted and fused to obtain the target fused features.
[0082] Specifically, the intramodal enhancement features of all modalities are concatenated into a first feature set, and the intermodal complementary features of all modalities are concatenated into a second feature set. Then, the two feature sets are weighted and fused using a learnable weight matrix to output the final target fused feature for use by the subsequent emotion classification module.
[0083] S 600. Based on the emotion classification module, perform emotion classification on the target fusion features to obtain the target emotion classification result.
[0084] Specifically, in this step, the target fusion features are input into the emotion classification module. The emotion classification module includes a one-dimensional convolutional submodule and a fully connected submodule. The one-dimensional convolutional submodule refines and transforms the dimensionality of the target fusion features, outputting intermediate features; the fully connected submodule flattens, linearly transforms, and maps the intermediate features, outputting the predicted probability or classification label for each emotion category. The target emotion classification result includes at least one of binary, triadic, or quadruple classification results; the specific classification type used is determined based on the preset classification task.
[0085] Specifically, after the feature fusion module LJA After feature fusion, the target fused features are... They were then moved to the emotion classification stage. This stage consists of a... Conv The module refers to the one-dimensional convolutional submodule and a FC The module is composed of the fully connected sub-modules. Conv The module further refines the features and converts the channel dimensions to a fixed size for subsequent processing. FC The module performs the final classification.
[0086] Specifically, Conv The module includes one-dimensional convolution and batch normalization. BN )and Dropout Layer. Its calculation is as follows:
[0087] Conv Module output Depend on FC Module processing. FC The module consists of one flattened layer, two linear layers, and one... BN Layer and one Dropout Layer composition:
[0088] Here, This represents the prediction result, where For batch size, This represents the number of emotion categories.
[0089] During the training phase, these predictions can be compared with the true labels to evaluate classification performance, and the target classification model can be further optimized based on the evaluated classification performance.
[0090] Furthermore, in this embodiment, the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism further includes: EEG signals were acquired using a simplified channel acquisition scheme to obtain simplified EEG signals. The simplified channel includes 11 channels, which cover the frontal lobe, temporal lobe, and occipital lobe brain regions, respectively. Based on the feature fusion module, the simplified EEG signal and the peripheral physiological signal are fused in a multimodal manner to compensate for the feature loss caused by the reduction of EEG channels; The target classification model is designed to be lightweight, resulting in a lightweight classification model whose parameter count, computational cost, and inference latency meet the resource constraints of the target wearable device.
[0091] Specifically, in this embodiment, in order to address the core issues of wearable devices such as limited resources (memory, computing power) and limited sensor deployment (number of channels, size), targeted optimizations are made from two dimensions: channel simplification design and hardware adaptation optimization, so that the target classification model can meet the deployment and practical application requirements of wearable devices.
[0092] First, based on neuroscience research on emotion-related brain regions, in one embodiment, EEG signals can be acquired using a simplified channel acquisition scheme to obtain a simplified EEG signal. Specifically, brain regions strongly related to emotion perception and processing, such as the prefrontal cortex, temporal lobe, and occipital lobe, are selected, abandoning the traditional 32-channel EEG signal acquisition method. EEG Redundant channels were used to design an 11-channel EEG signal processing system. EEG The lightweight configuration solution, specifically the simplified channel, includes: Fp 1. Fp 2. F 7. F 8. T 7. T 8. P 7. P 8. O 1. O 2. Oz It covers the frontal, temporal, and occipital lobes of the brain. These 11 channels cover core emotion-related brain regions, significantly reducing sensor size, wearing cost, and signal acquisition power consumption.
[0093] Then, based on the feature fusion module, the simplified EEG signal and the peripheral physiological signal are fused in a multimodal manner, and the feature loss caused by the reduction of EEG channels is compensated by the peripheral physiological signal through the intermodal attention mechanism.
[0094] Specifically, this applies to single-modal EEG signals after channel simplification. EEG To address the issue of feature loss, a multimodal fusion strategy is employed to streamline the EEG signals from multiple channels. EEG With peripheral physiological signals PPS *(Electrooscopic signals) EOG Electromyographic signals EMGSkin conductance signals GSR Signal fusion, through the feature fusion module LJA The cross-modal complementarity mechanism compensates for the EEG signal EEG Feature loss due to reduced channels.
[0095] Finally, the target classification model is designed to be lightweight, resulting in a lightweight classification model. Specifically, the model lightweighting steps include: based on the hardware constraints of mainstream wearable microcontrollers (e.g., the hardware constraint of the target series of wearable devices is a maximum of 1.4). MRAM The target classification model is designed to be lightweight, ensuring that the total number of model parameters is controlled at 0.60. M Within, the amount of calculation ( FLOPs Controlled at 9.33 M Within this range, it is fully compatible with the memory and computing power of the target series of wearable devices, meeting the resource constraints of wearable devices.
[0096] Furthermore, this embodiment also includes inference latency optimization through a single-level feature extraction module. LMSC Lightweight feature fusion module LJA and dynamic feature fusion network DFFN The bottleneck structure reduces model inference latency to 31.10. ms It basically meets the real-time emotion recognition needs of wearable devices.
[0097] As can be seen, this embodiment includes the feature extraction module. LMSC It is a lightweight multi-scale convolution module, which includes a single-layer multi-branch design and uses a set of convolution kernels of a specific size to achieve effective multi-scale spatial feature extraction with extremely low complexity.
[0098] Furthermore, the feature fusion module LJA It includes a lightweight joint attention mechanism and a dynamic feature fusion network. DFFN . LJA pass Intra and Inter Attention-based integration effectively processes intramodal and intermodal information; DFFN As its core component, replacing the traditional feedforward network with an innovative parallel linear / convolutional branch structure is key to achieving ultra-lightweight and high-performance multi-task adaptability.
[0099] Furthermore, in this embodiment, an end-to-end ultralight frame for wearable devices is also included. ULER ) and its channel reduction strategy. The total number of parameters in the entire framework (0.60) M ) and computational cost (9.33) MFLOPs Extremely low, short inference delay (~31.10) ms ), can be deployed in STM 32 H 7. Microcontroller. Specially designed to reduce... EEG The configuration of the number of channels (such as reducing it from 32 to 11) further meets the limitations of wearable hardware while maintaining high accuracy.
[0100] The above LMSC , LJA (including) DFFN The specific network structure, connection method, and computation process; ULER The method for constructing the overall framework, as well as the lightweight design and channel optimization strategies for wearable devices, all fall within the protection scope of this invention.
[0101] The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this embodiment has been applied to three internationally recognized public datasets ( DEAP , DREAMER , WESAD Extensive experimental validation was conducted on 8 emotion recognition tasks. Eleven of the best models from the past three years were also used for comparison.
[0102] Specifically, in this embodiment, two strategies were employed: subject-dependent and full-dataset-based. Tables 1, 2, and 3 show the accuracy comparison between the stated target classification model and 11 best models from the past three years. PPS This represents all peripheral physiological signals in the corresponding dataset. PPS * indicates a portion of the peripheral physiological signals from the corresponding dataset used in this embodiment. Results show that the target classification model achieves optimal accuracy across multiple tasks.
[0103] Table 1:
[0104] Table 2:
[0105] Table 3:
[0106] Furthermore, in this embodiment, in DEAP The dataset also compared the number of parameters in comparable models during a binary classification task involving subjects. FLOPs In terms of inference time, its performance is found to be the best among comparable models, as detailed in Table 4. * indicates that the target classification model does not include... DFFN structure. M It means one million. ms Represents milliseconds. Inference time is based on... STM 32H Performance characteristics estimation of 7 sequences (maximum operating frequency 600) MHz In addition, some original literature for certain methods did not explicitly report the number of parameters or... FLOPs Missing values were estimated based on information provided in the original paper.
[0107] Table 4:
[0108] Furthermore, in this embodiment, in DEAP Data set EEG The channels are reduced, specifically including two solutions: Extended 11-channel settings (…). Fp 1. Fp 2. F 7. F 8. T 7. T 8. P 7. P 8. O 1. O 2. Oz (located in the frontal, temporal, and occipital lobes respectively); and a minimized 2-channel version. EEG set up( Fp 1. Fp 2. Located in the frontal lobe). Both approaches significantly reduce [the risk of infection] while maintaining high accuracy. EEG Number of channels. In particular, the 11-channel setup, while maintaining lightweight wearable characteristics, has achieved accuracy that meets or even surpasses the latest standards. SOTA Model performance on different tasks (and) PPS *The effect is more significant when used in combination.
[0109] Reference Figure 5 , Figure 6 , Figure 7 , Figure 8 and Figure 9 In this embodiment, the results were also analyzed, including the accuracy performance in different tasks. DEAP and DREAMER The accuracy performance of each subject under the subject-dependent strategy in the dataset. Also, the accuracy of each task. t - SNE Feature visualization. These features further enhance the reliability of the target classification model in the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanisms provided in this embodiment.
[0110] As can be seen, the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism provided in this embodiment exhibits superior emotion classification performance and extremely low resource consumption in the target classification model: DEAP , DREAMER , WESAD On three benchmark datasets, across eight tasks, ULER In achieving the highest or most competitive recognition accuracy (e.g.) DEAP While achieving a titer / awakening binary classification accuracy of 99.34% / 99.46%, the parameter quantity (0.60) was also low. M )and FLOPs (9.33) M (far lower than similar products) SOTA Model (usually several) M Up to dozens M (Parameters), achieving the optimal balance between performance and efficiency.
[0111] Furthermore, in this embodiment, the target classification model has wearable deployment capability: the model size is smaller than that of common microcontrollers (such as...). STM 32 H 7 series) RAM Capacity, measured inference latency approximately 31.10. ms This meets real-time requirements. This is the case for most existing... SOTA The model cannot achieve this.
[0112] Furthermore, the method provided in this implementation achieves efficient multimodal fusion: through LMSC + LJA + DFFN The joint design utilizes modal complementarity more efficiently than existing attention fusion methods, while significantly reducing the complexity of the fusion module.
[0113] A general multitasking architecture: unified ULER The framework is through DFFN It enhances multi-tasking adaptability, eliminating the need for significant architecture adjustments for different tasks (binary classification, quadri-class classification, etc.).
[0114] In summary, this embodiment provides a method for emotion recognition based on multimodal physiological signals using multi-scale convolution and joint attention mechanisms. When performing emotion recognition based on multimodal physiological signals, one or more sets of multimodal physiological signals are first acquired, including electroencephalogram (EEG) signals and peripheral physiological signals. Then, the multimodal physiological signals are input into a target classification model, which includes a preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module. Next, the multimodal physiological signals are preprocessed using the preprocessing module to obtain a multimodal input tensor. Lightweight multi-scale convolution feature extraction is performed on the multimodal input tensor using the feature extraction module to obtain target convolution features. Lightweight joint attention feature fusion is performed on the target convolution features using the feature fusion module to obtain target fused features. Finally, emotion classification is performed on the target fused features using the emotion classification module to obtain the target emotion classification result. This embodiment provides users with a multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanisms, solving the problems of cumbersome models and simple fusion mechanisms that are common in existing multimodal physiological signal emotion recognition methods. In this embodiment, through the collaborative design of lightweight multi-scale convolution feature extraction and lightweight joint attention feature fusion, high-precision multimodal physiological signal emotion recognition is achieved with extremely low parameter and computational costs.
[0115] It should be understood that although the steps in the flowcharts shown in the accompanying drawings are displayed sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of the steps in this invention, and these steps can be executed in other orders. Moreover, at least a portion of the steps in this invention may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0116] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program using signal-related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM). ROM Programmable ROM ( PROM ), electrically programmable ROM ( EPROM Electrically erasable programmable ROM ( EEPROM ) or flash memory. Volatile memory may include random access memory (RAM) RAM Alternatively, an external cache memory. This is for illustrative purposes only and not as a limitation. RAM It can be obtained in various forms, such as static RAM ( SRAM ),dynamic RAM ( DRAM ),synchronous DRAM ( SDRAM ), double data rate SDRAM ( DDR SDRAM ), Enhanced SDRAM ( ESDRAM ), Synchronization Link ( Synchlink ), DRAM ( SLDRAM ), memory bus ( Rambus )direct RAM ( RDRAM ), Direct Memory Bus Dynamics RAM ( DRDRAM ), and memory bus dynamics RAM ( RDRAM )wait.
[0117] Example 2 like Figure 10 As shown, based on the above-mentioned multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 10 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0118] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as the terminal's hard drive or memory. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard drive or smart memory card equipped on the terminal. SmartMediaCard , SMC ), Secure Digital ( SecureDigital , SD ) card, flash memory card ( FlashCard Furthermore, the memory 20 may include both internal storage units and external storage devices of the terminal. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a multimodal physiological signal emotion recognition program 40 based on a multi-scale convolution and joint attention mechanism. This multimodal physiological signal emotion recognition program 40 based on a multi-scale convolution and joint attention mechanism can be executed by the processor 10, thereby realizing the multimodal physiological signal emotion recognition method based on a multi-scale convolution and joint attention mechanism in this application.
[0119] In some embodiments, the processor 10 may be a central processing unit (CPU). CentralProcessingUnit , CPU The microprocessor or other data processing chip is used to run the program code stored in the memory 20 or process data, such as executing the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism.
[0120] The display 30 may be, in some embodiments, LED Monitors, LCD monitors, touch LCD monitors and OLED ( OrganicLight - EmittingDiode The display 30 includes components such as organic light-emitting diodes (OLEDs) and touchscreens. It displays information on the terminal and provides a visual user interface. The components 10-30 of the terminal communicate with each other via a system bus.
[0121] In one embodiment, when the processor 10 executes the multimodal physiological signal emotion recognition program 40 based on multi-scale convolution and joint attention mechanisms stored in the memory 20, the following steps are performed: Acquire one or more sets of multimodal physiological signals, including electroencephalogram (EEG) signals and peripheral physiological signals; The multimodal physiological signals are input into a target classification model, which includes a preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module. The multimodal physiological signal is preprocessed based on the preprocessing module to obtain a multimodal input tensor; The feature extraction module performs lightweight multi-scale convolution feature extraction on the multimodal input tensor to obtain the target convolution feature; The target convolutional features are fused using a lightweight joint attention feature fusion method based on the feature fusion module to obtain the target fused features; Based on the emotion classification module, the target fusion features are classified to obtain the target emotion classification result.
[0122] Example 3 The present invention also provides a computer-readable storage medium having stored one or more programs thereon, which can be executed by one or more processors to implement the steps of the multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism described in the above embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism, characterized in that, The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism includes: Acquire one or more sets of multimodal physiological signals, including electroencephalogram (EEG) signals and peripheral physiological signals; The multimodal physiological signals are input into a target classification model, which includes a preprocessing module, a feature extraction module, a feature fusion module, and an emotion classification module. The multimodal physiological signal is preprocessed based on the preprocessing module to obtain a multimodal input tensor; The feature extraction module performs lightweight multi-scale convolution feature extraction on the multimodal input tensor to obtain the target convolution feature; The target convolutional features are fused using a lightweight joint attention feature fusion method based on the feature fusion module to obtain the target fused features; Based on the emotion classification module, the target fusion features are classified to obtain the target emotion classification result.
2. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism according to claim 1, characterized in that, Before acquiring one or more sets of multimodal physiological signals, the method further includes: An initial classification model is constructed, which includes a preprocessing module, a feature extraction module, a feature fusion module, and a sentiment classification module. The feature extraction module is configured to perform lightweight multi-scale convolutional feature extraction, and the feature fusion module is configured to perform lightweight joint attention feature fusion. After obtaining the target dataset and constructing labels for the target dataset based on a preset classification task, the target dataset after label construction is divided into a first data set and a second data set based on a subject-dependent strategy and a full dataset strategy. Both the first data set and the second data set contain a training set and a test set. The initial classification model is trained and validated based on the first dataset and the second dataset, respectively, to obtain the target classification model.
3. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism according to claim 2, characterized in that, The peripheral physiological signals include at least one of the following: electrooculogram (EOG), electromyography (EMG), electroskin response (ESR), electrocardiogram (ECG), respiratory signals, blood volume signals, and temperature signals.
4. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism according to claim 1, characterized in that, The feature extraction module performs lightweight multi-scale convolution feature extraction on the multimodal input tensor to obtain target convolution features, including: The multimodal input tensor is input to the feature extraction module, which includes at least two parallel convolutional branches. Each convolutional branch uses a convolutional kernel of a different size to extract spatial features of different receptive fields along the signal channel dimension. The multimodal input tensor is convolved by at least two parallel convolutional branches to obtain the spatial features of each convolutional branch. After normalizing and applying activation functions to the spatial features of each convolutional branch, the features are fused using matrix addition to obtain the target convolutional features of each modality in the multimodal physiological signal.
5. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism according to claim 4, characterized in that, The target convolutional features are fused using a lightweight joint attention feature fusion method based on the feature fusion module to obtain the target fused features, including: The target convolutional features of each modality are input into the feature fusion module, which includes multiple intra-modal attention units and multiple inter-modal attention units; Based on the intramodal attention unit, intramodal dependency modeling is performed on the target convolutional features of each modality to obtain intramodal enhancement features of each modality. The intramodal attention unit includes a multi-head attention mechanism and a dynamic feature fusion network. The dynamic feature fusion network includes parallel linear transformation branches and one-dimensional convolution branches. Based on the intermodal attention unit, cross-modal dependency modeling is performed on the target convolutional features of each modality to obtain intermodal complementary features of each modality. The intermodal attention unit uses the target convolutional features of the target modality as the query and the target convolutional features of other modalities as the key and value, and extracts cross-modal features through a multi-head attention mechanism and a dynamic feature fusion network. The intra-modal enhancement features and inter-modal complementary features of each modality are weighted and fused to obtain the target fused features.
6. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism according to claim 5, characterized in that, Based on the intra-modal attention unit, intra-modal dependency modeling is performed on the target convolutional features of each modality to obtain the intra-modal enhancement features of each modality, including: After reshaping the target convolutional features of each modality into the target shape, positional encoding is added to obtain the positional encoding features of each modality; The location encoding feature is input into the first multi-head attention subunit. The attention weights of the location encoding feature are calculated by multiple attention heads in the first multi-head attention subunit. The attention weights output by each attention head are concatenated to obtain the attention concatenation features of each modality. The attention-separated features of each modality are input into the first dynamic feature fusion network, which includes parallel linear transformation branches and one-dimensional convolution branches. The first dynamic feature fusion network extracts and fuses multi-path features from the attention splicing features of each modality to obtain intra-modal enhancement features for each modality.
7. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism according to claim 5, characterized in that, Based on the inter-modal attention unit, cross-modal dependency modeling is performed on the target convolutional features of each modality to obtain inter-modal complementary features of each modality, including: The target convolutional features of each modality are used as the target query, and the target convolutional features of other modalities are used as the target key and target value; After adding position codes to the target query, the target key, and the target value, they are input into the second multi-head attention subunit. The cross-modal attention weights of each modality are calculated by multiple attention heads in the second multi-head attention subunit, and the outputs of each attention head are concatenated to obtain the cross-modal attention concatenation features of each modality. The cross-modal attention splicing features are input into a second dynamic feature fusion network, which includes parallel linear transformation branches and one-dimensional convolution branches. The second dynamic feature fusion network is used to extract and fuse multi-path features of the cross-modal attention splicing features to obtain intermodal complementary features of each modality.
8. The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism according to claim 1, characterized in that, The multimodal physiological signal emotion recognition method based on multi-scale convolution and joint attention mechanism also includes: EEG signals were acquired using a simplified channel acquisition scheme to obtain simplified EEG signals. The simplified channel includes 11 channels, which cover the frontal lobe, temporal lobe, and occipital lobe brain regions, respectively. Based on the feature fusion module, the simplified EEG signal and the peripheral physiological signal are fused in a multimodal manner to compensate for the feature loss caused by the reduction of EEG channels; The target classification model is designed to be lightweight, resulting in a lightweight classification model whose parameter count, computational cost, and inference latency meet the resource constraints of the target wearable device.
9. A smart terminal, characterized in that, The smart terminal includes a memory, a processor, and a multimodal physiological signal emotion recognition program based on a multi-scale convolution and joint attention mechanism, which is stored in the memory and can run on the processor. When the multimodal physiological signal emotion recognition program based on a multi-scale convolution and joint attention mechanism is executed by the processor, it implements the steps of the multimodal physiological signal emotion recognition method based on a multi-scale convolution and joint attention mechanism as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multimodal physiological signal emotion recognition program based on a multi-scale convolution and joint attention mechanism. When the multimodal physiological signal emotion recognition program based on a multi-scale convolution and joint attention mechanism is executed by a processor, it implements the steps of the multimodal physiological signal emotion recognition method based on any one of claims 1-8.