Electroencephalogram emotion recognition method based on data enhancement and text cue word mechanism

By introducing multi-scale sliding window technology and text prompt word mechanism into the emotion recognition method, the problems of insufficient accuracy of emotion recognition and complex feature extraction in the prior art are solved, and efficient and accurate emotion recognition is achieved.

CN120052898APending Publication Date: 2025-05-30SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510225703.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing emotion recognition methods are affected by individual differences and external environment, and the recognition accuracy is insufficient. The technology based on physiological signals has problems such as complex feature extraction, long-distance dependence and difficulty in local context capture.

Method used

Data diversity is enhanced by introducing multi-scale sliding window technology, and text prompt word mechanism is used to integrate EEG signals with text prompt words to avoid artificial feature extraction process, and enhance the model's long-distance dependence on EEG signals and the capture ability of local context.

Benefits of technology

It significantly improves the accuracy and efficiency of emotion recognition, solves the problems of complex feature extraction and scarcity in traditional methods, and realizes end-to-end integrated design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120052898A_ABST
    Figure CN120052898A_ABST
Patent Text Reader

Abstract

The invention discloses an electroencephalogram emotion recognition method based on data enhancement and a text cue word mechanism, and belongs to the technical field of emotion calculation. Comprising the steps that an electroencephalogram data set is acquired and preprocessed; carrying out data enhancement on the preprocessed electroencephalogram data set by adopting a multi-scale sliding window technology; carrying out electroencephalogram signal embedding and text cue word embedding on the data of the electroencephalogram data set after data enhancement; carrying out feature fusion on the electroencephalogram signal embedding and the text cue word embedding; constructing an emotion classification model, and training the emotion classification model by using the fused features; and acquiring electroencephalogram data to be classified, and inputting the electroencephalogram data into the trained emotion classification model to complete emotion recognition. Through a multi-scale sliding window technology and a text cue word generation mechanism, the data diversity and the ability of the emotion classification model to capture long-distance dependence on electroencephalogram signals and local context are enhanced, and therefore the accuracy of emotion recognition is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of affective computing, and particularly relates to an electroencephalogram emotion recognition method based on data augmentation and text prompt mechanism. Background Art

[0002] Emotion recognition technology plays an increasingly important role in social activities, which is related to an individual's mental health, social adaptability, and daily behavior and work efficiency. With the development of artificial intelligence and human-computer interaction technology, emotion recognition technology has become a key research direction for improving user experience and has shown broad application potential in many fields such as mental health monitoring, emotion regulation, and education.

[0003] However, existing emotion recognition methods, especially those relying on facial expressions, speech, and physiological signals, are often affected by individual differences and external environments, resulting in insufficient recognition accuracy. Physiological signal-based technologies, such as electroencephalogram and electrocardiogram, although providing objective information about emotional states, have problems in practical applications such as complex feature extraction, difficulty in capturing long-distance dependencies and local contexts, and data scarcity. The high-dimensional, non-linear, and non-stationary characteristics of electroencephalogram signals further increase the difficulty of emotion recognition. Although the application of deep learning technology has improved the accuracy of emotion recognition, most studies have not fully utilized its end-to-end characteristics, limiting the improvement of model performance. Summary of the Invention

[0004] Aiming at the defects of the existing technology, the present invention enhances data diversity by introducing a multi-scale sliding window technique; by introducing a text prompt mechanism, it ingeniously combines electroencephalogram signals with text prompts, not only avoiding the cumbersome artificial feature extraction process in traditional methods, but also enhancing the long-distance dependence and local context capture ability of the emotion classification model for electroencephalogram signals; significantly improving the efficiency and accuracy of emotion recognition.

[0005] To achieve the above object, the present invention provides an electroencephalogram emotion recognition method based on data augmentation and text prompt mechanism, including the following steps:

[0006] (1) Obtain an electroencephalogram dataset and perform preprocessing;

[0007] (2) For the preprocessed electroencephalogram dataset, use the multi-scale sliding window technique for data augmentation;

[0008] (3) For the data of the electroencephalogram dataset after data augmentation, perform electroencephalogram signal embedding and text prompt embedding;

[0009] (4) Perform feature fusion on the electroencephalogram signal embedding and text prompt embedding;

[0010] (5) Construct an emotion classification model and train the emotion classification model using the fused features;

[0011] (6) Obtain the EEG data to be classified, input it into the trained emotion classification model, and complete emotion recognition.

[0012] Further, the preprocessing includes:

[0013] Filtering: Filter out the low-frequency and high-frequency noises in the EEG dataset through a band-pass filter;

[0014] Artifact removal: Use independent component analysis technology to remove the eye movement artifacts and muscle artifacts in the EEG dataset;

[0015] Standardization: Perform standardization processing on the data in the EEG dataset after the above processing so that the data of each channel has the same dimension.

[0016] Further, the specific content of step (2) is: Slide and segment each EEG data in the EEG dataset into multiple scales according to K time windows with a predetermined overlap rate to generate a corresponding set of signal segments; After splicing the segments of each scale along the sample dimension, perform secondary block processing through non-overlapping sliding windows, and take each block as an EEG data to obtain an enhanced dataset.

[0017] Further, the text prompt embedding is specifically:

[0018] For each EEG data in the EEG dataset after data augmentation, extract its statistical characteristics, including the maximum value, minimum value, median, and trend;

[0019] Combine the statistical characteristics with the set prompt word template to obtain the final text prompt;

[0020] Tokenize and encode the obtained text prompt through a pre-trained BERT model to obtain the text prompt embedding E propmt

[0021] Further, the EEG signal embedding is specifically:

[0022] Perform block operation on each EEG data in the EEG dataset after data augmentation according to a preset fixed length;

[0023] Use Conv1D to perform convolution on the segmented data blocks to extract temporal features;

[0024] Perform data dimensionality reduction through average pooling and linear operations;

[0025] Adopt Dropout regularization processing;

[0026] Retain the timing information of the signal through position encoding;

[0027] Adjust the data dimension through the Broadcasting operation to obtain the final EEG signal embedding E EEG .

[0028] Furthermore, the step (4) is specifically: input the EEG signal embedding E EEG and the text prompt embedding E propmt into the multi-head attention mechanism for feature fusion. The EEG signal embedding E EEG is used as the query vector, and the text prompt embedding E prompt is used as the key vector and value vector to obtain the final features for classification;

[0029]

[0030] where: Q, K, and V are the query, key, and value vectors respectively, and d k is the scaling factor.

[0031] Furthermore, the emotion classification model includes a fully connected layer and a Softmax layer;

[0032] The input data generates the probability distribution of emotion categories through the fully connected layer:

[0033] y = W * h + b

[0034] where: W is the weight matrix, h represents the emotion feature vector, which is obtained by averaging the high-dimensional features output by the embedding interaction layer, and b is the bias term;

[0035] Through the Softmax operation, the model outputs the probabilities of different emotion categories:

[0036]

[0037] where: y i represents the linear output of class i, is the predicted probability of class i;

[0038] Finally, select the class with the highest probability as the emotion recognition result of the model.

[0039] The present invention also provides an EEG emotion recognition system based on data augmentation and text prompt mechanism, including:

[0040] A data acquisition module, used to acquire the EEG data set and perform preprocessing;

[0041] A data augmentation module, used to perform data augmentation on the preprocessed EEG data set by using the multi-scale sliding window technology proposed in this article;

[0042] A feature extraction module, which is used to perform electroencephalogram (EEG) signal embedding and text prompt embedding on the EEG dataset data after data augmentation, and fuse the EEG signal embedding and text prompt embedding;

[0043] A modeling and training module, which is used to build an emotion classification model and train the emotion classification model using the fused features;

[0044] An emotion recognition module, which is used to obtain EEG data to be classified, input the trained emotion classification model, and complete emotion recognition.

[0045] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the EEG emotion recognition method based on data augmentation and text prompt mechanism as described above.

[0046] The present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the EEG emotion recognition method based on data augmentation and text prompt mechanism as described above.

[0047] Advantages of the present invention:

[0048] (1) By using the multi-scale sliding window technology, the present invention enhances data diversity, thus significantly improving the accuracy of emotion recognition.

[0049] (2) By introducing the text prompt mechanism, the present invention skillfully fuses EEG signals and text prompts, provides semantic context guidance for the model, enables it to more accurately associate physiological features with emotion categories, thus solving the inherent limitation that EEG signals lack clear semantic expression, and at the same time avoiding the cumbersome manual feature extraction process in traditional methods, realizing an end-to-end integrated design and achieving efficient emotion recognition.

[0050] (3) The method of the present invention has high processing speed and accuracy, and can be widely applied to various practical application scenarios such as mental health monitoring, human-computer interaction, and emotion regulation. Description of the Drawings

[0051] Figure 1 is a schematic flowchart of the EEG emotion recognition method based on data augmentation and text prompt mechanism according to an embodiment of the present invention.

[0052] Figure 2 is a schematic diagram of the EEG signal data augmentation technology according to an embodiment of the present invention.

[0053] Figure 3It is a schematic diagram of electroencephalogram (EEG) signal embedding, text prompt embedding and feature fusion in an embodiment of the present invention.

[0054] Figure 4 It is a schematic structural diagram of an EEG emotion recognition system based on data augmentation and text prompt mechanism in an embodiment of the present invention. Specific embodiments

[0055] The present invention will be further explained below with reference to the accompanying drawings and embodiments.

[0056] As Figure 1 shown, the present invention provides an EEG emotion recognition method based on data augmentation and text prompt mechanism, including the following steps:

[0057] S101. Obtain an EEG dataset and perform preprocessing.

[0058] Obtain the EEG data of multiple subjects to obtain an EEG dataset, and perform preprocessing to remove noise and improve data quality. The preprocessing includes filtering, artifact removal, and normalization.

[0059] Filtering process: Filter out the low-frequency and high-frequency noise of the data in the EEG dataset through a band-pass filter, and retain the main features of the EEG signal.

[0060] Artifact removal: Use independent component analysis (ICA) technology to remove the eye movement artifacts and muscle artifacts in the data of the EEG dataset to ensure the purity of the EEG signal.

[0061] Normalization: Perform normalization processing on the data in the EEG dataset after the above processing, so that the data of each channel has the same dimension to improve the training efficiency of the model.

[0062] S102. For the preprocessed EEG dataset, use the multi-scale sliding window technology for data augmentation.

[0063] Each EEG data in the EEG dataset is segmented into K time windows at a predetermined overlap rate respectively by multi-scale sliding segmentation to generate a corresponding set of signal segments; after the segments of each scale are concatenated along the sample dimension, secondary block processing is performed through non-overlapping sliding windows, and each block is used as new EEG data to obtain an augmented dataset. Through this technology, the dataset is expanded.

[0064] As Figure 2 shown, in an embodiment of the present invention, the EEG signal of each subject is segmented into multiple time windows (with lengths of 0.25 seconds, 0.5 seconds, and 0.75 seconds), and the overlap rate overlapping_rate is set to 35%. The generated signal segments are then concatenated along the sample dimension.

[0065] The number of signal segments generated for each time window is calculated by the following formula:

[0066]

[0067] where: tw is the current window length, and step is the step size. The calculation formula is:

[0068] step = [tw × overlapping_rate]

[0069] The core of data augmentation lies in horizontal augmentation. By using the sliding window technique, more time-point sequences are generated to enrich the diversity of a single trial signal in the time-point dimension.

[0070] However, the data samples after horizontal augmentation still have a long time series length. To further increase the number of samples, a non-overlapping sliding window method is used to cut the horizontally augmented data into blocks of a predetermined length, that is, vertical augmentation. This operation not only ensures that the dimensions of each sub-sample are consistent, increases the number of samples, but also improves the training effect of the model.

[0071] S103. For the EEG dataset data after bidirectional data augmentation, perform EEG signal embedding and text prompt embedding.

[0072] As Figure 3 shown, according to the dynamic characteristics of the EEG signal, text prompts related to emotions are generated and combined with the set prompt template to finally obtain the text prompt embedding E prompt , The embedding of text prompts provides additional context information for emotion classification, helps the model better understand the classification task, and improves the accuracy of emotion recognition. Specifically:

[0073] A1: Calculate the following statistical characteristics for each EEG data in the EEG dataset after data augmentation:

[0074] MaximumValue: Used to measure the peak characteristics of the signal in this time period and reflect the extreme value changes of the signal.

[0075] MinimumValue: Used to reflect the lower bound characteristics of the signal and represent the lowest amplitude state of the signal.

[0076] MedianValue: By calculating the median value of the signal amplitude in the time period, it reflects the overall level of the signal.

[0077] Trend: Used to quantify the change direction of the signal in the time period. The calculation formula is:

[0078] Trend = x emd -xstart

[0079] Among them, x start represents the value of the first time point in this time period, and x end represents the value of the last time point. A positive value indicates an upward trend of the signal, while a negative value reflects a downward trend.

[0080] A2: Combine the obtained statistical characteristics with the set prompt word template to obtain the final text prompt word.

[0081] Predefined prompt word templates are such as:

[0082] “Based on the provided EEG data, classify the underlying emotion. Statistics: minimum value=MIN, maximum value=MAX, median value=MEDIN. The overall trend of the data is TREND.”

[0083] Among them: MIN, MAX, MEDIN, and TREND correspond to the above 4 statistical characteristics.

[0084] A3: Perform word segmentation and encoding on the obtained text prompt word through a pre-trained BERT model (freeze the BERT model parameters) to obtain the text prompt word embedding E propmt .

[0085] The EEG signal embedding is specifically as follows:

[0086] B1: Perform chunking operations on each EEG data in the augmented EEG dataset according to a fixed length.

[0087] The number of signal segments N after chunking can be calculated by the following formula:

[0088]

[0089] Among them, W is the length of each segment, S is the sliding step, and L is the data length.

[0090] B2: Use Conv1D to perform convolution on the chunked data blocks to extract temporal features.

[0091] The specific operation is as follows:

[0092] E EEG = Conv1D(X p )

[0093] Among them: X p ∈R(B×N,C,W) is the input signal.

[0094] B3: Data dimensionality reduction is performed through average pooling and linear operations.

[0095] B4: To prevent overfitting, Dropout regularization is adopted.

[0096] E EEG = Dropout(Linear(Mean(E EEG )))

[0097] B5: The temporal information of the signal is retained through positional encoding.

[0098]

[0099] where pos is the position index of the segment (from 0 to N - 1), i is the index of the embedding dimension, and d bert represents the dimension of the BERT embedding.

[0100] B6: The data dimension is adjusted through Broadcasting operation to obtain the final EEG signal embedding E EEG .

[0101] Finally, the data dimension is adjusted through Broadcasting operation so that each feature meets the input requirements of the subsequent model, and where B is the batch size.

[0102] S104. Perform feature fusion on the EEG signal embedding and the text prompt embedding.

[0103] Input the EEG signal embedding E EEG and the text prompt embedding E propmt into the multi-head attention mechanism for feature fusion. The multi-head attention mechanism can automatically weight and fuse features from different sources to obtain the final features for classification.

[0104] In the embodiment of the present invention, the EEG signal embedding E EEG is used as the query vector (Query), the text prompt embedding E prompt is used as the key vector (Key) and the value vector (Value), and feature fusion is achieved through the multi-head attention mechanism:

[0105]

[0106] where: Q, K, and V are the query, key, and value vectors respectively, and d k is the scaling factor.

[0107] S105. Build an emotion classification model and train the emotion classification model using the fused features.

[0108] The emotion classification model includes a fully connected layer and a Softmax layer.

[0109] The fully connected layer is responsible for processing the fused features to generate a probability distribution of emotion categories:

[0110] y = W * h + b

[0111] Where: W is the weight matrix, h represents the emotion feature vector, which is obtained by averaging the high-dimensional features output by the embedding interaction layer, and b is the bias term.

[0112] Through the Softmax operation, the model outputs the probabilities of different emotion categories, and finally selects the category with the highest probability as the emotion recognition result of the model.

[0113]

[0114] Where: y i represents the linear output of class i, is the predicted probability of class i.

[0115] The Softmax function can convert the linear output into a probability distribution, enabling the model to use the category with the highest probability as the final prediction result.

[0116] Use the fused features of the enhanced EEG dataset to train the emotion classification model to obtain a trained emotion classification model.

[0117] S106. Obtain the EEG data to be classified, input the trained emotion classification model, and complete emotion recognition.

[0118] As Figure 4 shown, an EEG emotion recognition system based on data augmentation and text prompt mechanism is also provided in an embodiment of the present invention, including:

[0119] A data acquisition module 410, configured to obtain an EEG dataset and perform preprocessing;

[0120] A data augmentation module 420, configured to perform data augmentation on the preprocessed EEG dataset by using the multi-scale sliding window technology proposed in this article;

[0121] A feature extraction module 430, configured to perform EEG signal embedding and text prompt embedding on the EEG dataset data after data augmentation, and fuse the EEG signal embedding and text prompt embedding;

[0122] The modeling and training module 440 is used to build an emotion classification model and train the emotion classification model using the fused features;

[0123] The emotion recognition module 450 is used to obtain the EEG data to be classified, input the trained emotion classification model, and complete emotion recognition.

[0124] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the above-mentioned EEG emotion recognition method based on data augmentation and text prompt mechanism.

[0125] An embodiment of the present invention also provides an electronic device, including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above-mentioned EEG emotion recognition method based on data augmentation and text prompt mechanism.

Claims

1. A method for EEG emotion recognition based on data enhancement and text prompt word mechanism, characterized in that: The steps include: (1) Obtain EEG data sets and perform preprocessing; (2) for the preprocessed EEG dataset, a multi-scale sliding window technique is used for data enhancement; (3) performing EEG signal embedding and text prompt word embedding on the EEG data set after data enhancement; (4) performing feature fusion on the EEG signal embedding and the text prompt word embedding; (5) constructing an emotion classification model, and using the fused features to train the emotion classification model; (6) Obtain the EEG data to be classified, input the trained emotion classification model, and complete emotion recognition.

2. The method for EEG emotion recognition based on data enhancement and text prompt word mechanism according to claim 1 is characterized in that: The pre-processing comprises: Filtering: filtering out low-frequency and high-frequency noise in the EEG data set through a bandpass filter; Artifact removal: The independent component analysis technique is used to remove the eye movement artifacts and muscle artifacts in the EEG data set; Standardization: Standardize the data in the EEG data set after the above processing so that the data of each channel has the same dimension.

3. The EEG emotion recognition method based on data enhancement and text prompt word mechanism according to claim 1 is characterized in that: The step (2) is specifically as follows: each EEG data in the EEG data set is subjected to multi-scale sliding segmentation according to K time windows with a predetermined overlap rate to generate a corresponding set of signal segments; each scale segment is spliced ​​along the sample dimension, and then secondary block processing is performed through a non-overlapping sliding window, and each block is used as an EEG data to obtain an enhanced data set.

4. The method for EEG emotion recognition based on data enhancement and text prompt word mechanism according to claim 1, characterized in that: The text prompt word embedding is specifically: For each EEG data in the EEG data set after data enhancement, extract its statistical characteristics, including maximum value, minimum value, median value and trend; Combining the statistical characteristics with the set prompt word template to obtain the final text prompt word; The obtained text prompt words are segmented and encoded through the pre-trained BERT model to obtain the text prompt word embedding E propmt .

5. The method for EEG emotion recognition based on data enhancement and text prompt word mechanism according to claim 1, characterized in that: The EEG signal embedding is specifically as follows: For each EEG data in the EEG data set after data enhancement, a block operation is performed according to a preset fixed length; Use Conv1D to convolve the divided data blocks to extract time series features; Data dimensionality reduction is performed through average pooling and linear operations; Use Dropout regularization; Preserve the timing information of the signal through position encoding; The final EEG signal embedding E is obtained by adjusting the data dimension through the Broadcasting operation. EEG .

6. The method for EEG emotion recognition based on data enhancement and text prompt word mechanism according to claim 1, characterized in that: The step (4) is specifically as follows: embedding the EEG signal into E EEG and text hint word embedding E propmt Input into the multi-head attention mechanism for feature fusion, the EEG signal is embedded in E EEG is used as the query vector, the text hint word embedding E prompt Used as key vector and value vector to obtain the final features used for classification; Where: Q, K, V are query, key and value vectors respectively, d k is the scaling factor.

7. The method for EEG emotion recognition based on data enhancement and text prompt word mechanism according to claim 1, characterized in that: The emotion classification model includes a fully connected layer and a Softmax layer; The input data passes through the fully connected layer to generate the probability distribution of emotion categories: y=W*h+b Where: W is the weight matrix, h represents the emotion feature vector, which is obtained by averaging the high-dimensional features output by the embedding interaction layer, and b is the bias term; Through the Softmax operation, the model outputs the probability of different emotion categories: Where: y i represents the linear output of category i, is the predicted probability of category i; Finally, the category with the highest probability is selected as the emotion recognition result of the model.

8. An EEG emotion recognition system based on data enhancement and text prompt word mechanism, characterized in that: include: Data acquisition module, used to obtain EEG data sets and perform preprocessing; A data enhancement module, used for performing data enhancement on the preprocessed EEG data set by using the multi-scale sliding window technology proposed in this paper; A feature extraction module, used for performing EEG signal embedding and text prompt word embedding on the EEG data set after data enhancement, and performing feature fusion on the EEG signal embedding and text prompt word embedding; A modeling and training module, used to build a sentiment classification model and train the sentiment classification model using the fused features; The emotion recognition module is used to obtain the EEG data to be classified, input the trained emotion classification model, and complete emotion recognition.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the EEG emotion recognition method based on data enhancement and text prompt word mechanism as described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the EEG emotion recognition method based on data enhancement and text prompt word mechanism as described in any one of claims 1-7.