An electroencephalogram emotion analysis method based on a parallel training multi-task hybrid model
Patent Information
- Application Number
- CN202310273961.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-21
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-03-21
AI Technical Summary
首先,脑图并不能一比一完全复刻现实中电极在大脑上的位置信息,而是按照相对位置拼凑在矩阵中,因此电极位置本身就存在着一定的偏差
[0045]1. In the process of constructing the mind map, a multi-level data processing interpolation algorithm oriented towards network architecture was adopted to improve image resolution while reducing the impact of discrete electrodes due to individual position deviations, thus achieving normalization of the mind map.
Smart Images

Figure CN116421200B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biosignal processing technology, and in particular to a brainwave emotion analysis method based on a multi-task hybrid model trained in parallel. Background Technology
[0002] Emotions are physiological or psychological responses to external stimuli, whether conscious or unconscious. They can be broadly categorized into positive and negative emotions. Positive emotions promote physical and mental health, while negative emotions often lead to stress, irritability, depression, and even mental breakdown. Emotions originate in the central and peripheral nervous systems, primarily resulting from the synchronized execution of neurons, constituting a temporal activity. Emotions can be captured through various physiological signals, including electroencephalography (EEG), body temperature, electrocardiography (ECG), electromyography (EMG), and skin responses. They can also be expressed through non-physiological signals such as facial expressions, body movements, and tone of voice. However, non-physiological signals are easily influenced by subjective will, thus posing a significant challenge to the reliability of emotion recognition based on non-physiological signals. Physiological signals, on the other hand, are unaffected by subjective will and are more likely to express genuine emotions.
[0003] Electroencephalograms (EEGs) are physiological signals directly generated by the human nervous system, reflecting a person's true emotions in a more objective way. Currently, methods primarily use statistical features in the time and frequency domains to create artificial features or extract features directly from processed raw EEG signals, then train emotion recognition models using machine learning or deep learning methods based on these features. However, most of these methods do not consider the fusion of features from multiple domains. Mind maps, which integrate frequency, spatial, and temporal features, effectively address this issue. However, mind maps still face two problems: a limited number of detection electrodes and spatial errors caused by individual differences among subjects. Meanwhile, existing model networks often focus only on the analysis of features from a single domain and train the model on an individual basis, resulting in neglecting multifaceted information and poor generalization ability across objects.
[0004] The automatic extraction and analysis of EEG signal features directly impacts the accuracy of emotional experience recognition. However, current brain maps constructed by fusing frequency and spatial features suffer from three main problems. First, brain maps cannot perfectly replicate the actual electrode positions on the brain; instead, they are pieced together in a matrix based on relative positions, resulting in inherent inaccuracies in electrode placement. Second, the number of electrodes on a brain map is limited; relying solely on 32 or 62 electrodes to record feature information leads to significant data gaps. Third, each subject's brain exhibits variations in size and subtle differences, thus brain maps constructed using only proportional scaling ignore individual variations.
[0005] Therefore, how to reduce the differences in mind maps caused by individual differences while maximizing data filling and constructing a network that can simultaneously focus on features in multiple domains such as space, frequency, and time and is compatible with cross-object differences to optimize the accuracy of emotion recognition has become a technical problem that needs to be solved. Summary of the Invention
[0006] The purpose of this invention is to address the shortcomings of the prior art by providing a brainwave emotion analysis method based on a multi-task hybrid model trained in parallel.
[0007] The technical solution provided by this invention is: a measurement method for EEG emotion analysis based on a multi-task hybrid model trained in parallel, comprising the following steps:
[0008] 1) EEG data processing: subset segmentation and brain map generation;
[0009] For multi-input parallel training, in the data processing part, the training set is divided into n subsets. The data of a single trial is divided into multiple subsets by data partitioning and parallel training is carried out in the network using a parallel architecture to improve efficiency. Except for the data content, the data subsets are kept consistent in other control variables. The subset data are only different in time. The data comes from the same subject watching the same stimulus material. Then, the spatial frequency domain features of the EEG signal are extracted and the raw EEG signal is converted into a brain map.
[0010] 2) Spatial frequency domain feature learning: The frequency and spatial features of the signal are learned using a hierarchical adjustable concave network with bidirectional multi-scale sampling;
[0011] Each subset of data is input into a parallel hybrid network model for multi-task learning. The hybrid model stacks multiple network models, including a hierarchical adjustable concave network architecture that is good at extracting frequency and spatial features, and a network architecture that can effectively extract temporal features, namely sinusoidal position embedding and forward and backward sequence relation algorithms.
[0012] 3) Temporal feature learning: After adding positional embeddings to provide time series information, the stacked forward and backward sequence relationship algorithm is used to learn temporal features;
[0013] 4) Multi-task learning: Based on the emotion classification module of multi-task learning, the final sentiment analysis results are obtained;
[0014] Each data subset's mind map is processed by its respective spatial frequency domain feature learning module and temporal feature learning module to obtain intermediate result vectors, which are then averaged and integrated in the averaging layer of the multi-task learning module. The integrated result vector is used as input for multi-task learning to achieve multi-dimensional information sharing and obtain the final multi-dimensional task classification result.
[0015] Preferably, step 1) includes:
[0016] To better extract the frequency and spatial features of EEG, this invention employs an interpolation algorithm for multi-level data processing oriented towards network architecture to construct a brain map;
[0017] The first level of data processing is: sinusoidal position embedding and forward-backward sequence relationship algorithm, using a non-overlapping sliding window time window to slice the data of each frequency band. In this invention, an 11S window slice is used for SEED data and a 2S window slice is used for DEAP dataset.
[0018] The second level of data processing involves analyzing the network architecture for time-series features and calculating the differential entropy and density entropy features of each slice.
[0019] The third level of data processing is: oriented towards the spatial frequency domain feature analysis network architecture, mapping the differential entropy and density entropy values calculated from the same frequency band and the same time period from each channel to the corresponding positions on the brain map according to the electrode positions of the 10-20 international system.
[0020] The fourth level of data processing involves adjusting the concave network architecture according to its hierarchy, using an interpolation algorithm based on the optimal pixel. Specifically, this involves using 16 nearest neighbors to obtain the values in the scaled mind map through weighted summation within the neighborhood of the mapped point.
[0021] The first step is to determine the correspondence between the point coordinates of the original image with size a×b and the target image with size A×B. Compared with the original image coordinates The correspondence is as follows: Thus, for each point in the target image, the corresponding target location can be found in the original image;
[0022] Secondly, a basis function W(d) is constructed. The basis function W(d) determines the weight of the neighboring point relative to the target point based on the relative position of the target point and the neighboring points in the original image. W(d) is as follows:
[0023]
[0024] Among them, (1) and (2) Let C represent the relative distances between the target point and the nearest neighbor points on the x-axis and y-axis, respectively, with C = -0.5. Then, the weights of the 16 nearest neighbor points on the x and y axes are calculated using the basis functions. Finally, the weighted sum of these 16 points is calculated using the following formula:
[0025]
[0026] R and C represent the x and y coordinates of a point in the target image, respectively; r and c represent the x and y coordinates of a point in the original image; and i and j are the x and y numbers of the 16 neighboring points, respectively.
[0027] Further preferably, step 2) includes:
[0028] Due to individual differences in EEG signal data, training data from different subjects can lead to model degradation. However, in order to improve the generalization ability of the model, it is necessary to train the model using data from all individuals. Therefore, the hierarchical adjustment concave network adds batch normalization and pruning operations between each convolutional layer and max pooling layer of the bidirectional multi-scale sampling network.
[0029] Further preferably, step 3) includes:
[0030] In the temporal feature learning module, this invention uses a stacked forward-backward sequence relation algorithm to capture the relationships between long-distance EEG signal sequences. To supplement the distance-positional relationships in the temporal information, this invention uses sinusoidal position embedding to supplement the positional information before the stacked forward-backward sequence relation algorithm.
[0031] The sequence relation algorithm is implemented through three relations: input relation (3)(4)(5), preorder relation (6), and output relation (7)(8), and iterates according to the following formula:
[0032] (3)
[0033] (4)
[0034] (5)
[0035] (6)
[0036] (7)
[0037] (8)
[0038] The input relationship first requires selecting the signal information to be stored (3). It is the signal information from the previous moment. It is the input signal information at the current moment, and the output is the preceding sequence. Next, you need to select the information to be stored (4) and (5) and input the signal information from the previous moment. and the input signal information at the current moment Output the current value and cache information The calculation of the preorder relation (6) is based on the input of the current value. Preorder sequence Cache information and cached information from the previous moment Output the cache information at the current time. Output relationship (7) (8) Input data information from the previous moment Input information at the current moment Current cache information Output sequence relation value and current time signal information ;
[0039] in addition, Let represent the sigmoid activation function, tanh represent the tangent activation function, and Q represent the weight matrix. It is the input vector. , , , It's about bias, while the forward and backward sequence relation algorithms need to combine the data information from different time points obtained by the two sequence relation algorithms, forward and backward. Then, the parts are assembled.
[0040] More preferably, step 4) includes:
[0041] Dimensional data can be directly processed using multi-task learning methods, but discrete data is a single-task multi-class dataset that does not have multiple related dimensional classification tasks.
[0042] Therefore, a transformation method from single-task to multi-task learning is adopted: the single-task multi-class label is transformed, and a unique binary label is created for each class label. Thus, during multi-task learning training, each class label has a unique classification task. Therefore, for a certain category classification and recognition task, the classification results of other tasks will be weighed more comprehensively to achieve accurate classification results.
[0043] This invention provides a brainwave emotion analysis method based on a parallel-trained multi-task hybrid model, which can effectively extract information from multiple domains including the frequency, spatial, and temporal domains. Therefore, the constructed neural network itself needs to contain modules that are adept at deeply mining information in the spatial, frequency, and temporal domains. This invention proposes combining multiple different types of network architectures and finally using a parallel hybrid model multi-task learning approach to share associated information and achieve accurate recognition.
[0044] In summary, the present invention has the following main beneficial effects:
[0045] 1. In the process of constructing the mind map, a multi-level data processing interpolation algorithm oriented towards network architecture was adopted to improve image resolution while reducing the impact of discrete electrodes due to individual position deviations, thus achieving normalization of the mind map.
[0046] This invention uses an interpolation algorithm for multi-level data processing oriented towards network architecture. Therefore, the size of the target image is an important parameter. The hierarchical adjustment concave network needs to connect the channel data with the reconstructed data. The reconstructed data is obtained by pooling and deconvolution of the original data. Therefore, the data size will be magnified or reduced by a factor of two in each operation. Thus, the pixel size of the target image needs to be a power of 2. Otherwise, it will lead to a situation where the data dimensions do not match and cannot be connected.
[0047] 2. The hybrid network architecture proposed in this invention can effectively analyze both the spatial and temporal features of EEG. Therefore, the performance of the parallel training multi-task hybrid network architecture of this invention surpasses the current best recognition model.
[0048] 3. For cross-object training, a hierarchical adjustment concave network architecture is proposed, which improves the generalization ability of the entire model and reduces model degradation.
[0049] 4. Employing a multi-task learning approach to achieve information sharing improves the model's classification performance. Specifically, for discrete datasets, a method is proposed to transform discrete data from single-task to continuous data from multiple-task. Multiple association tasks are constructed on discrete sentiment datasets, expanding the applicability of multi-task learning methods and enhancing the compatibility of this approach. Attached Figure Description
[0050] Figure 1 The overall flowchart of the EEG emotion analysis method based on a multi-task hybrid model with parallel training provided by the present invention;
[0051] Figure 2 The present invention provides a hierarchical modulated concave network diagram for spatial frequency domain feature learning in an EEG emotion analysis method based on a parallel-trained multi-task hybrid model.
[0052] Figure 3 The flowchart of the single-task to multi-task learning conversion method in the EEG emotion analysis method based on a parallel training multi-task hybrid model provided by the present invention;
[0053] Figure 4 A comparison of the interpolation algorithm before and after processing in the EEG emotion analysis method based on a parallel-trained multi-task hybrid model provided by this invention, which is a multi-level data processing method for network architecture.
[0054] Figure 5The influence of time window size on DEAP results in an EEG emotion analysis method based on a parallel-trained multi-task hybrid model provided by this invention;
[0055] Figure 6 The influence of time window size on SEED results in an EEG emotion analysis method based on a parallel-trained multi-task hybrid model provided by this invention;
[0056] Figure 7 A before-and-after comparison of the beneficial modules invented on the DEAP dataset in the EEG emotion analysis method based on a parallel-trained multi-task hybrid model provided by this invention;
[0057] Figure 8 The image shows a comparison of the beneficial modules invented on the SEED dataset before and after ablation in the EEG emotion analysis method based on a parallel-trained multi-task hybrid model provided by this invention. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0059] A measurement method for EEG emotion analysis based on a parallel-trained multi-task hybrid model includes the following steps:
[0060] 1) EEG data processing: subset segmentation and brain map generation;
[0061] For multi-input parallel training, in the data processing part, the training set is divided into n subsets. The data of a single trial is divided into multiple subsets by data partitioning and parallel training is carried out in the network using a parallel architecture to improve efficiency. Except for the data content, the data subsets are kept consistent in other control variables. The subset data are only different in time. The data comes from the same subject watching the same stimulus material. Then, the spatial frequency domain features of the EEG signal are extracted and the raw EEG signal is converted into a brain map.
[0062] 2) Spatial frequency domain feature learning: The frequency and spatial features of the signal are learned using a hierarchical adjustable concave network with bidirectional multi-scale sampling;
[0063] Each subset of data is input into a parallel hybrid network model for multi-task learning. The hybrid model stacks multiple network models, including a hierarchical adjustable concave network architecture that is good at extracting frequency and spatial features, and a network architecture that can effectively extract temporal features, namely sinusoidal position embedding and forward and backward sequence relation algorithms.
[0064] 3) Temporal feature learning: After adding positional embeddings to provide time series information, the stacked forward and backward sequence relationship algorithm is used to learn temporal features;
[0065] 4) Multi-task learning: Based on the emotion classification module of multi-task learning, the final sentiment analysis results are obtained;
[0066] Each data subset's mind map is processed by its respective spatial frequency domain feature learning module and temporal feature learning module to obtain intermediate result vectors, which are then averaged and integrated in the averaging layer of the multi-task learning module. The integrated result vector is used as input for multi-task learning to achieve multi-dimensional information sharing and obtain the final multi-dimensional task classification result.
[0067] Preferably, step 1) includes:
[0068] To better extract the frequency and spatial features of EEG, this invention employs an interpolation algorithm for multi-level data processing oriented towards network architecture to construct a brain map;
[0069] The first level of data processing is: sinusoidal position embedding and forward-backward sequence relationship algorithm, using a non-overlapping sliding window time window to slice the data of each frequency band. In this invention, an 11S window slice is used for SEED data and a 2S window slice is used for DEAP dataset.
[0070] The second level of data processing involves analyzing the network architecture for time-series features and calculating the differential entropy and density entropy features of each slice.
[0071] The third level of data processing is: oriented towards the spatial frequency domain feature analysis network architecture, mapping the differential entropy and density entropy values calculated from the same frequency band and the same time period from each channel to the corresponding positions on the brain map according to the electrode positions of the 10-20 international system.
[0072] The fourth level of data processing involves adjusting the concave network architecture according to its hierarchy, using an interpolation algorithm based on the optimal pixel. Specifically, this involves using 16 nearest neighbors to obtain the values in the scaled mind map through weighted summation within the neighborhood of the mapped point.
[0073] The first step is to determine the correspondence between the point coordinates of the original image with size a×b and the target image with size A×B. Compared with the original image coordinates The correspondence is as follows: Thus, for each point in the target image, the corresponding target location can be found in the original image;
[0074] Secondly, a basis function W(d) is constructed. The basis function W(d) determines the weight of the neighboring point relative to the target point based on the relative position of the target point and the neighboring points in the original image. W(d) is as follows:
[0075]
[0076] Among them, (1) and (2) Let C represent the relative distances between the target point and the nearest neighbor points on the x-axis and y-axis, respectively, with C = -0.5. Then, the weights of the 16 nearest neighbor points on the x and y axes are calculated using the basis functions. Finally, the weighted sum of these 16 points is calculated using the following formula:
[0077]
[0078] R and C represent the x and y coordinates of a point in the target image, respectively; r and c represent the x and y coordinates of a point in the original image; and i and j are the x and y numbers of the 16 neighboring points, respectively.
[0079] Further optimization, such as Figure 2 As shown, step 2) includes:
[0080] Due to individual differences in EEG signal data, training data from different subjects can lead to model degradation. However, in order to improve the generalization ability of the model, it is necessary to train the model using data from all individuals. Therefore, the hierarchical adjustment concave network adds batch normalization and pruning operations between each convolutional layer and max pooling layer of the bidirectional multi-scale sampling network.
[0081] Further preferably, step 3) includes:
[0082] In the temporal feature learning module, this invention uses a stacked forward-backward sequence relation algorithm to capture the relationships between long-distance EEG signal sequences. To supplement the distance-positional relationships in the temporal information, this invention uses sinusoidal position embedding to supplement the positional information before the stacked forward-backward sequence relation algorithm.
[0083] The sequence relation algorithm is implemented through three relations: input relation (3)(4)(5), preorder relation (6), and output relation (7)(8), and iterates according to the following formula:
[0084] (3)
[0085] (4)
[0086] (5)
[0087] (6)
[0088] (7)
[0089] (8)
[0090] The input relationship first requires selecting the signal information to be stored (3). It is the signal information from the previous moment. It is the input signal information at the current moment, and the output is the preceding sequence. Next, you need to select the information to be stored (4) and (5) and input the signal information from the previous moment. and the input signal information at the current moment Output the current value and cache information The calculation of the preorder relation (6) is based on the input of the current value. Preorder sequence Cache information and cached information from the previous moment Output the cache information at the current time. Output relationship (7) (8) Input data information from the previous moment Input information at the current moment Current cache information Output sequence relation value and current time signal information ;
[0091] in addition, Let represent the sigmoid activation function, tanh represent the tangent activation function, and Q represent the weight matrix. It is the input vector. , , , It's about bias, while the forward and backward sequence relation algorithms need to combine the data information from different time points obtained by the two sequence relation algorithms, forward and backward. Then, the parts are assembled.
[0092] Further optimization, such as Figure 3 As shown, step 4) includes:
[0093] Dimensional data can be directly processed using multi-task learning methods, but discrete data is a single-task multi-class dataset that does not have multiple related dimensional classification tasks.
[0094] Therefore, a transformation method from single-task to multi-task learning is adopted: the single-task multi-class label is transformed, and a unique binary label is created for each class label. Thus, during multi-task learning training, each class label has a unique classification task. Therefore, for a certain category classification and recognition task, the classification results of other tasks will be weighed more comprehensively to achieve accurate classification results.
[0095] Example 1
[0096] 1) EEG data processing: subset segmentation and brain map generation;
[0097] Raw EEG signals were acquired and preprocessed (using the DEAP dataset as an example in this embodiment). The data was sampled to 128Hz and bandpass filtered from 4 to 45Hz to remove noise and artifacts. A Butterworth filter was used to decode each subset of data, which was then decoded into multiple frequency bands such as theta, alpha, beta, and gamma. A 2-second non-overlapping sliding window was used to slice each frequency band, and the differential entropy and density entropy features of each slice were calculated. The differential entropy and density entropy values calculated from the same frequency band and time period from each channel were mapped to the corresponding positions on the brain map according to the electrode positions of the 10-20 international system. A multi-level data processing interpolation algorithm oriented towards network architecture was used to process each brain map, and the data was segmented into subsets based on time.
[0098] Since the DEAP dataset is being used, both types of labels (arousal and valence) need to be divided into two categories based on the median value of the upper and lower limits of each label during the splitting process. If a single-task multi-class dataset is being used, a single-task to multi-task conversion method is required to convert the labels.
[0099] 2) Spatial frequency domain feature learning: The frequency and spatial features of the signal are learned using a hierarchical adjustable concave network with bidirectional multi-scale sampling;
[0100] The constructed subset mind maps are fed into the branches of each parallel architecture to learn spatial frequency domain features using a hierarchical adjustable concave network.
[0101] Meanwhile, the data sent to each branch all come from the same experiment in the subset segmentation stage to ensure that the subjects, stimulus materials and all other independent variables are the same.
[0102] 3) Learning of temporal characteristics;
[0103] After adding positional embeddings, time series information is provided, and then a stacked forward and backward sequence relationship algorithm is used to learn temporal features;
[0104] 4) Multi-task learning;
[0105] The mean vector is fed into the multi-task learning module to perform correlation prediction for each type of label.
[0106] This embodiment uses the public dataset DEAP. The experiment required 32 participants to watch 40 one-minute music videos to evoke emotions. Each experiment involved watching a 60-second video, with a 3-second baseline relaxation signal data period before the experiment. Each experiment was labeled with four tags (1-9): arousal, valence, dominance, and liking, as emotional references related to the stimulus. Forty physiological or EEG signal data were recorded, including 32 EEG signals and 8 other peripheral physiological signals. EEG signal acquisition used a 32-channel Biosemi ActiveTwo device and a 10-20 international system. Noise was denoised using EMG and EOG recordings and a 4-45 Hz bandpass filter, then downsampled to 128 Hz. Therefore, the final data format for each participant is shown in Table 1. The experiment only tested the arousal and valence dimensions using a multi-task learning approach. Additionally, if the dataset is a single-task, multi-class discrete dataset, the conversion method described in the technical solution is required to perform multi-task conversion of the labels.
[0107] Table 1: Data for Each Participant
[0108]
[0109] The experimental environment consisted of Python 3.8, an NVIDIA GeForce GTX 3090 GPU processor, and the entire network was implemented using the Tensorflow architecture. The dataset used 10-fold cross-validation, and the training lasted for 200 epochs. The learning rate decayed at a rate of 1 / 2 every 8 epochs. The batch size was 64 samples, and the learning rate was 0.001.
[0110] To verify the performance of the present invention, it was compared with different models, and the comparison results are shown in Table 2.
[0111] Table 2: Performance comparison with other methods on DEAP
[0112]
[0113] The results show that the present invention achieved classification accuracy of 99.62% and 99.66% in the two dimensions of the DEAP dataset, respectively. The present invention achieves high classification performance for both arousal and valence while maintaining a low standard deviation, indicating that the method has strong compatibility and can cope with the influence of differences among subjects.
[0114] Example 2
[0115] 1) EEG data processing: subset segmentation and brain map generation;
[0116] Raw EEG signals were acquired and preprocessed. Data was sampled to 200Hz and bandpass filtered from 0 to 75Hz to remove noise and artifacts. Butterworth filters were used to decode subsets of data into multiple frequency bands, including delta, theta, alpha, beta, and gamma. An 11-second non-overlapping sliding window was used to slice the data for each frequency band, and the differential entropy and density entropy features of each slice were calculated. The differential entropy and density entropy values calculated for the same frequency band and time period from each channel were mapped to the corresponding positions on the brain map according to the electrode positions of the 10-20 international system. A multi-level data processing interpolation algorithm oriented towards network architecture was used to process each brain map. Data was subsetted by time, and the labels were converted from single-task to multi-task using a single-task to multi-task conversion method.
[0117] The original three-category task (positive, negative, neutral) is transformed into three specific binary labels for use in subsequent multi-task learning: the first binary label determines whether the emotion is positive, the second binary label determines whether the emotion is negative, and so on for other tasks.
[0118] 2) Spatial frequency domain feature learning: learning the frequency and spatial features of concave networks;
[0119] The constructed subset brain maps are fed into the branches of the parallel architecture respectively, and spatial frequency domain features are learned using an improved concave network.
[0120] Meanwhile, the data sent to each branch all come from the same experiment in the subset segmentation stage to ensure that the subjects, stimulus materials and all other independent variables are the same.
[0121] 3) Learning of temporal characteristics;
[0122] After adding positional embeddings, time series information is provided, and then a stacked forward and backward sequence relationship algorithm is used to learn temporal features;
[0123] 4) Multi-task learning;
[0124] The mean vector is fed into the multi-task learning module to perform correlation prediction for each type of label.
[0125] To verify the performance of the present invention on the discrete dataset SEED, it was compared with the comparison method in Example 1. The comparison results are shown in Table 3.
[0126] Table 3: Performance comparison of the method in Example 1 on SEED
[0127]
[0128] Experimental results show that the EEG emotion analysis method based on a parallel-trained multi-task hybrid model has an accuracy rate of 99.61%, which is significantly better than other comparative methods. This effectively verifies that the EEG emotion analysis method based on a parallel-trained multi-task hybrid model can achieve better recognition accuracy.
[0129] In the process of constructing the mind map, a multi-level data processing interpolation algorithm oriented towards network architecture was adopted to improve image resolution while reducing the impact of discrete electrodes due to individual positional deviations, thus achieving normalization of the mind map.
[0130] Since this invention uses a multi-level data processing interpolation algorithm oriented towards network architecture, the size of the target image is an important parameter. The improved concave network needs to connect the channel data with the reconstructed data. The reconstructed data is obtained by pooling and deconvolution of the original data; therefore, the data size is scaled up or down by a factor of two during each operation. Thus, the pixel size of the target image needs to be a power of 2; otherwise, a data dimension mismatch will occur, preventing connection. Experiments show that the optimal pixel size of the target image is 16×16. The experimental results are shown in Table 4. A comparison before and after using the multi-level data processing interpolation algorithm oriented towards network architecture is provided. Figure 4 As shown.
[0131] Table 4: The Influence of Target Image Pixel Size on Results
[0132]
[0133] To achieve better results, this invention also experimented with different time windows during the preliminary preparation for constructing the mind map. Figure 5 and 6 The model performance on the DEAP and SEED datasets is shown for different time window sizes. Figure 5 and Figure 6 It can be seen that a time window that is too small will prevent the extracted features from acquiring more information, while a time window that is too large will lead to data sparsity, thus greatly reducing the model's generalization ability. From Figure 5 and Figure 6 It can be seen that the recognition accuracy of the DEAP and SEED datasets reaches its peak when the time window is 2s and 11s, respectively. Furthermore, this invention notes that the classification accuracy suddenly drops when the time window breaks the inflection point. Overall, this indicates that the window size plays a balancing role between temporal and frequency features.
[0134] For mind map input, the hybrid network architecture proposed in this invention can effectively analyze both the spatial and temporal features of EEG. Therefore, the performance of the parallel-trained multi-task hybrid network architecture of this invention surpasses the current best recognition model. Specific comparison results are shown in Table 5.
[0135] Table 5: Comparison of Results from Different Models
[0136]
[0137] The results show that the method of this invention achieves the best performance in both dimensions of the DEAP dataset, improving upon the second-best results by 3% and 3.38% in arousal and valence, respectively. Compared with all other methods, the method of this invention has the smallest standard deviation in arousal and valence studies, indicating good stability. On the SEED dataset, the method of this invention achieves a recognition accuracy of 99.61%, improving upon the second-best method by 0.42%. In other words, the experimental results of DEAP and SEED demonstrate the superiority of the method of this invention.
[0138] For cross-object training, the architecture of the concave network was improved, thereby enhancing the generalization ability of the overall model and reducing model degradation. Comparative results can be found in [link to comparison]. Figure 7 and Figure 8 .
[0139] The comparison results show that the concave network has a significant impact on model performance before and after the improvement. On the SEED dataset, the overall performance of the concave network dropped sharply by 5.42% after losing batch normalization and pruning in each layer. On the DEAP dataset, the impact of the missing batch normalization and pruning is even more pronounced. Because the model cannot accommodate the influence of individual differences across objects, the entire model fails to converge.
[0140] Employing a multi-task learning approach to achieve information sharing improves the model's classification performance. Specifically, for discrete datasets, a method is proposed to transform discrete data from single-task to continuous data from multiple-task approaches. Multiple association tasks are constructed on discrete sentiment datasets, expanding the applicability and compatibility of multi-task learning methods.
[0141] Figure 7 and Figure 8 The results showed that the accuracy of SEED decreased by nearly 2.3% after ablating the single-task to multi-task conversion method. When the interpolation algorithm for multi-level data processing oriented towards the network architecture was ablated, the accuracy of DEAP and SEED decreased by 0.9% and 0.5%, respectively. Other beneficial modules of this method all showed a significant decrease in recognition accuracy after ablation.
[0142] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.
[0143] It should be understood that the present invention is not limited to the precise structure shown in the above description, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A brainwave emotion analysis method based on a multi-task hybrid model trained in parallel, characterized in that, Includes the following steps: 1) EEG data processing: subset segmentation and brain map generation; For multi-input parallel training, in the data processing part, the training set is divided into n subsets. The data of a single trial is divided into multiple subsets by data partitioning and parallel training is carried out in the network using a parallel architecture to improve efficiency. Except for the data content, the data subsets are kept consistent in other control variables. The subset data are only different in time. The data comes from the same subject watching the same stimulus material. Then, the spatial frequency domain features of the EEG signal are extracted and the raw EEG signal is converted into a brain map. 2) Spatial frequency domain feature learning: The frequency and spatial features of the signal are learned using a hierarchical adjustable concave network with bidirectional multi-scale sampling; Each subset of data is input into a parallel hybrid network model for multi-task learning. The hybrid model stacks multiple network models, including a hierarchical adjustable concave network architecture that is good at extracting frequency and spatial features, and a network architecture that can effectively extract temporal features, namely sinusoidal position embedding and forward and backward sequence relation algorithms. 3) Temporal feature learning: After adding positional embeddings to provide time series information, the stacked forward and backward sequence relationship algorithm is used to learn temporal features; 4) Multi-task learning: Based on the emotion classification module of multi-task learning, the final sentiment analysis results are obtained; Each data subset's mind map is processed by its respective spatial frequency domain feature learning module and temporal feature learning module to obtain intermediate result vectors, which are then averaged and integrated in the averaging layer of the multi-task learning module. The integrated result vector is used as input for multi-task learning to achieve multi-dimensional information sharing and obtain the final multi-dimensional task classification result.
2. The measurement method of the EEG emotion analysis method based on a parallel-trained multi-task hybrid model according to claim 1, characterized in that, Step 1) includes: To better extract the frequency and spatial features of EEG, an interpolation algorithm oriented towards network architecture and multi-level data processing is used to construct a brain map. The first level of data processing is: sinusoidal position embedding and forward-backward sequence relationship algorithm, using non-overlapping sliding time windows to slice the data of each frequency band, using 11S window slicing for SEED data and 2S window slicing for DEAP dataset. The second level of data processing involves analyzing the network architecture for time-series features and calculating the differential entropy and density entropy features of each slice. The third level of data processing is: oriented towards the spatial frequency domain feature analysis network architecture, mapping the differential entropy and density entropy values calculated from the same frequency band and the same time period from each channel to the corresponding positions on the brain map according to the electrode positions of the 10-20 international system. The fourth level of data processing involves adjusting the concave network architecture according to its hierarchy, using an interpolation algorithm based on the optimal pixel. Specifically, this involves using 16 nearest neighbors to obtain the values in the scaled mind map through weighted summation within the neighborhood of the mapped point. The first step is to determine the correspondence between the point coordinates of the original image with size a×b and the target image with size A×B. Compared with the original image coordinates The correspondence is as follows: Thus, for each point in the target image, the corresponding target location can be found in the original image; Secondly, a basis function W(d) is constructed. The basis function W(d) determines the weight of the neighboring point relative to the target point based on the relative position of the target point and the neighboring points in the original image. W(d) is as follows: ; Among them, (1) and (2) Let C represent the relative distances between the target point and the nearest neighbor points on the x-axis and y-axis, respectively, with C = -0.
5. Then, the weights of the 16 nearest neighbor points on the x and y axes are calculated using the basis functions. Finally, the weighted sum of these 16 points is calculated using the following formula: ; R and C represent the x and y coordinates of a point in the target image, respectively; r and c represent the x and y coordinates of a point in the original image; and i and j are the x and y numbers of the 16 neighboring points, respectively.
3. The measurement method of the EEG emotion analysis method based on a parallel-trained multi-task hybrid model according to claim 1, characterized in that, Step 2) includes: Due to individual differences in EEG signal data, training data from different subjects can lead to model degradation. However, in order to improve the generalization ability of the model, it is necessary to train the model using data from all individuals. Therefore, the hierarchical adjustment concave network adds batch normalization and pruning operations between each convolutional layer and max pooling layer of the bidirectional multi-scale sampling network.
4. The measurement method of the EEG emotion analysis method based on a parallel-trained multi-task hybrid model according to claim 1, characterized in that, Step 3) includes: In the temporal feature learning module, a stacked forward-backward sequence relation algorithm is used to capture the relationships between long-distance EEG signal sequences; to supplement the distance-position relationship in the temporal information, sinusoidal position embedding is used to supplement the positional information before the stacked forward-backward sequence relation algorithm; The sequence relation algorithm is implemented through three relations: input relation (3)(4)(5), preorder relation (6), and output relation (7)(8), and iterates according to the following formula: (3) (4) (5) (6) (7) (8) The input relationship first requires selecting the signal information to be stored (3). It is the signal information from the previous moment. It is the input signal information at the current moment, and the output is the preceding sequence. Next, you need to select the information to be stored (4) and (5) and input the signal information from the previous moment. and the input signal information at the current moment Output the current value and cache information The calculation of the preorder relation (6) is based on the input of the current value. Preorder sequence Cache information and cached information from the previous moment Output the cache information at the current time. Output relationship (7) (8) Input data information from the previous moment Input information at the current moment Current cache information Output sequence relation values and current time signal information ; in addition, Let represent the sigmoid activation function, tanh represent the tangent activation function, and Q represent the weight matrix. It is the input vector. , , , It's about bias, while the forward and backward sequence relation algorithms need to combine the data information from different time points obtained by the two sequence relation algorithms, forward and backward. Then, the parts are assembled.
5. The measurement method of the EEG emotion analysis method based on a parallel-trained multi-task hybrid model according to claim 1, characterized in that, Step 4) includes: Dimensional data can be directly processed using multi-task learning methods, but discrete data is a single-task multi-class dataset that does not have multiple related dimensional classification tasks. Therefore, a transformation method from single-task to multi-task learning is adopted: the single-task multi-class label is transformed, and a unique binary label is created for each class label. Thus, during multi-task learning training, each class label has a unique classification task. Therefore, for a certain category classification and recognition task, the classification results of other tasks will be weighed more comprehensively to achieve accurate classification results.