Electroencephalogram emotion estimation method and device based on time-frequency domain characteristics and language prompts

By employing an EEG emotion estimation method based on time-frequency domain features and language cues, the problems of low signal-to-noise ratio and non-stationarity of EEG signals are solved, achieving highly accurate emotion recognition and improving the reliability and objectivity of EEG emotion recognition.

CN120983051AActive Publication Date: 2025-11-21HUAZHONG NORMAL UNIV

Patent Information

Application Number
CN202511040033.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-21
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

Existing EEG-based emotion recognition technologies suffer from low signal-to-noise ratios, susceptibility to environmental noise interference, and non-stationarity, making it difficult to effectively extract and identify emotion features.

Method used

An EEG emotion estimation method based on time-frequency domain features and language cues is adopted. The method generates a time-frequency spectrum by segmenting, filtering and wavelet transforming the EEG signal, extracting time-frequency features by combining deep networks, calculating differential entropy to generate spatiotemporal features, and using a large language model to generate language cue embeddings for emotion classification.

Benefits of technology

It significantly improves the accuracy and robustness of emotion recognition, effectively extracts emotional features from EEG signals, and improves classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120983051A_ABST
    Figure CN120983051A_ABST
Patent Text Reader

Abstract

The invention provides an electroencephalogram emotion estimation method and device based on time-frequency domain features and language prompts, and belongs to the technical field of computers.The method comprises the steps that segmentation, normalization processing and frequency band filtering are conducted on an input original electroencephalogram signal, and a signal of a preset frequency band is reserved so that a time-frequency spectrogram can be generated through continuous wavelet transform; based on the time-frequency spectrogram, extracting a time-frequency feature vector through a preset deep network; calculating a difference entropy of a preset frequency band based on the original electroencephalogram signal, and generating a spatial-temporal feature vector; the text description reflecting the emotional state of the subject is used for generating language prompts to be embedded through a large language model; and fusing the time-frequency feature vector and the spatial-temporal feature vector, embedding the fused feature and a language prompt for modal alignment, and outputting an emotion classification result through a classifier. According to the method, key modes of electroencephalogram signals of subjects with different emotions can be captured, and the accuracy, objectivity and reliability of evaluation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer technology, and in particular to an electroencephalogram emotion estimation method and device based on time-frequency domain features and language prompts. BACKGROUND

[0002] Emotion recognition, as a core topic of interdisciplinary research, has shown important value in multiple fields. In the field of psychology, it provides quantitative basis for the study of emotion-related behaviors; in engineering, it promotes the optimization and upgrading of human-computer interaction, enabling machines to recognize, understand, and even simulate human emotions. Current emotion recognition technology is mainly based on two types of signal sources: physiological signals and non-physiological signals. Compared to facial expressions, gestures, and speech, which are easily influenced by subjective factors, electroencephalography (EEG) as a typical physiological signal has higher objectivity and reliability. Therefore, in recent years, EEG-based emotion recognition technology has attracted widespread attention from researchers.

[0003] However, EEG-based emotion recognition still has certain limitations. On the one hand, the signal-to-noise ratio (SNR) of EEG signals is low and is easily disturbed by environmental noise such as electromyography and electrooculography artifacts; on the other hand, EEG signals have time asymmetry and non-stationary characteristics, and their statistical properties change over time, which makes it difficult for traditional analysis methods based on steady-state assumptions to be directly applicable, thereby increasing the difficulty of feature extraction and pattern recognition. Therefore, developing robust signal processing algorithms to deal with noise and non-stationary interference is one of the key challenges to improve the reliability of EEG emotion recognition. This study innovatively proposes an EEG emotion estimation method based on time-frequency domain features and language prompts, and uses large language model prompts to assist emotion evaluation, in order to improve the accuracy and robustness of emotion recognition. SUMMARY

[0004] The present application provides an EEG emotion estimation method and device based on time-frequency domain features and language prompts to solve the defects existing in the prior art.

[0005] In a first aspect, the present application provides an electroencephalogram emotion estimation method based on time-frequency domain features and language prompts, comprising: segmenting, normalizing and filtering the input raw electroencephalogram signal in a preset frequency band to retain the signal, and generating a time-frequency spectrogram through continuous wavelet transform; extracting a time-frequency feature vector through a preset deep network based on the time-frequency spectrogram; calculating the differential entropy of the preset frequency band based on the raw electroencephalogram signal to generate a space-time feature vector; using a large language model to generate a language prompt embedding reflecting the emotional state of the subject; fusing the time-frequency feature vector and the space-time feature vector, aligning the fused features and the language prompt embedding in modalities, and outputting an emotional classification result through a classifier.

[0006] According to the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts provided by the present application, the input raw electroencephalogram signal is segmented, normalized and filtered in a preset frequency band to retain the signal, and a time-frequency spectrogram is generated through continuous wavelet transform, comprising: dividing the raw electroencephalogram signal into multiple segments of a preset length; normalizing and filtering each segment to retain the signal in a preset frequency band; and performing continuous wavelet transform on the signal in the preset frequency band using a Morlet wavelet basis function to convert the signal energy distribution into a time-frequency spectrogram.

[0007] According to the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts provided by the present application, a time-frequency feature vector is extracted through a preset deep network based on the time-frequency spectrogram, comprising: replacing the first layer convolutional network of ResNet18 with GaborNet and initializing three Gabor filters to extract local texture features of the time-frequency spectrogram; integrating the local texture features into global features through the residual structure and convolutional pooling operation of ResNet18; setting a graph convolutional network at the end of global feature extraction to learn the dynamic functional connectivity pattern between brain regions and generate a time-frequency feature vector in combination with the global features.

[0008] According to the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts provided by the present application, a space-time feature vector is generated based on the differential entropy of a preset frequency band calculated from the raw electroencephalogram signal, comprising: calculating the differential entropy of the raw electroencephalogram signal in the preset frequency band to form a differential entropy feature vector; and sequentially passing the differential entropy feature vector through a convolutional layer and an average pooling layer to generate a space-time feature vector.

[0009] According to the brain electrical emotion estimation method based on time-frequency domain features and language prompts provided by the application, a large language model is used to generate language prompt embedding from a text description reflecting the emotional state of a subject, comprising: using CLIP containing a tokenizer and a token embedder as a large language model, all parameters in the CLIP are frozen; inputting the text description reflecting the emotional state of the subject into the large language model, and processing it through the tokenizer and the token embedder in sequence, and outputting the language prompt embedding.

[0010] According to the brain electrical emotion estimation method based on time-frequency domain features and language prompts provided by the application, the time-frequency feature vector and the space-time feature vector are fused, the fused features and the language prompt embedding are modally aligned, and the emotion classification result is output through the classifier, comprising: fusing the time-frequency feature vector and the space-time feature vector to generate fused features; inputting the fused features into the classifier for processing; calculating the loss value by modally aligning the fused features and the language prompt embedding; and optimizing the parameters of the classifier based on the loss value.

[0011] According to the brain electrical emotion estimation method based on time-frequency domain features and language prompts provided by the application, the preset frequency band is the following five frequency bands: 0.5-4 Hz, 4-8 Hz, 8-12 Hz, 12-30 Hz and 30-100 Hz.

[0012] In a second aspect, the application further provides a brain electrical emotion estimation device based on time-frequency domain features and language prompts, comprising: A time-frequency spectrum generation module is configured to segment, normalize and filter the input raw electroencephalogram signal, and retain the signal of the preset frequency band to generate a time-frequency spectrum through continuous wavelet transform. A time-frequency feature extraction module is configured to extract a time-frequency feature vector based on the time-frequency spectrum through a preset deep network. A space-time feature extraction module is configured to calculate the differential entropy of the preset frequency band based on the raw electroencephalogram signal to generate a space-time feature vector. A language prompt module is configured to use a large language model to generate language prompt embedding from a text description reflecting the emotional state of a subject. An emotion recognition module is configured to fuse the time-frequency feature vector and the space-time feature vector, modally align the fused features and the language prompt embedding, and output the emotion classification result through the classifier.

[0013] In a third aspect, the application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the steps of the brain electrical emotion estimation method based on time-frequency domain features and language prompts as described above.

[0014] In a fourth aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of any of the above-mentioned emotion estimation methods based on time-frequency domain features and language cues.

[0015] The emotion estimation method and device based on time-frequency domain features and language cues provided by the present application have the following beneficial effects compared with the prior art: (1) The recognition method based on time-frequency domain features and language cues proposed by the present application can capture the key patterns of EEG signals of different emotional subjects, covering not only the complex dynamic changes in the time-frequency domain, but also further revealing the spatiotemporal characteristics of EEG signals through the introduction of differential entropy (DE) features, language cues and other auxiliary operations, thus significantly improving the accuracy, objectivity and reliability of the evaluation.

[0016] (2) Through the training of a large amount of data, the model can dynamically optimize the direction and scale parameters of the Gabor filter, and at the same time adjust the weights of the subsequent network, so as to efficiently extract the local texture information and global features in the time-frequency graph.

[0017] (3) The emotion recognition method of the present application has achieved a classification accuracy of 98.62% and 98.48% respectively on two public EEG data sets (Mumtaz data set and MODMA data set), which is superior to the prior art. This shows that the system can effectively extract the key features related to emotions in the EEG signal, providing a new idea for emotion evaluation. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0019] Figure 1 is one of the flowcharts of the emotion estimation method based on time-frequency domain features and language cues provided by the present application; Figure 2 is another flowchart of the emotion estimation method based on time-frequency domain features and language cues provided by the present application; Figure 3 is a structural schematic diagram of the emotion estimation device based on time-frequency domain features and language cues provided by the present application; Figure 4 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0020] The technical solutions and advantages of the present application will be described clearly below in conjunction with the drawings in the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0021] It should be noted that, in the description of the embodiments of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitation, the element defined by the statement "comprising a" does not exclude the presence of another identical element in the process, method, article or device comprising the element. The specific meaning of the above terms in the present application can be understood by those skilled in the art according to the specific circumstances.

[0022] The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second" and the like are generally a class, not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally represents that the front and rear associated objects are in a "or" relationship.

[0023] The embodiments of the present application will be described below in conjunction with Figures 1-4 The present application provides a method and device for estimating electroencephalogram emotion based on time-frequency domain features and language prompts.

[0024] Figure 1 is one of the flowcharts of the method for estimating electroencephalogram emotion based on time-frequency domain features and language prompts provided by the present application, as shown in Figure 1 includes but is not limited to the following steps: Step 101: segmenting, normalizing and frequency band filtering the input raw electroencephalogram signal, and retaining the signal of a preset frequency band to generate a time-frequency spectrum by continuous wavelet transform.

[0025] Specifically, step 101 includes: segmenting the raw electroencephalogram EEG signal into a plurality of segments of a preset time length; optionally, the preset time length is 3 seconds.

[0026] The normalization processing and the filtering processing are performed on each segment to retain signals of preset frequency bands; wherein, the algorithm of the normalization processing and the filtering processing can be set according to needs.

[0027] The Morlet wavelet base function is used for continuous wavelet transform on the signals of the preset frequency bands, so as to convert the signal energy distribution into a time-frequency spectrum.

[0028] Step 102: based on the time-frequency spectrum, a time-frequency feature vector is extracted through a preset deep network.

[0029] Specifically, step 102 includes: The first layer convolutional network of ResNet18 is replaced by GaborNet, and three Gabor filters are initialized to extract local texture features of the time-frequency spectrum; The local texture features are integrated into global features through the residual structure and convolutional pooling operation of ResNet18; A graph convolutional network is set at the end of the global feature extraction to learn the dynamic functional connection mode between brain regions, and the time-frequency feature vector is generated in combination with the global features.

[0030] Step 103: based on the original electroencephalogram signals, differential entropy of the preset frequency bands is calculated to generate a space-time feature vector.

[0031] Specifically, step 103 includes: The differential entropy of the original electroencephalogram signals is calculated in the preset frequency bands respectively to form a differential entropy feature vector; The differential entropy feature vector is sequentially passed through a convolutional layer and an average pooling layer to generate a space-time feature vector.

[0032] Step 104: a language prompt embedding is generated by using a large language model to reflect the text description of the emotional state of the subject.

[0033] Specifically, step 104 includes: The CLIP containing a tokenizer and a token embedder is used as a large language model, and all parameters in the CLIP are frozen; The text description reflecting the emotional state of the subject is input into the large language model, and is processed by the tokenizer and the token embedder in sequence to output the language prompt embedding.

[0034] Step 105: the time-frequency feature vector and the space-time feature vector are fused, the fused features and the language prompt embedding are modality aligned, and an emotional classification result is output through a classifier.

[0035] Specifically, step 105 includes: The time-frequency feature vector is fused with the space-time feature vector to generate a fused feature; The fused feature is input to a classifier for processing; The fused feature and the language prompt embedding are subjected to modal alignment to calculate a loss value; The parameters of the classifier are optimized based on the loss value, so that an optimized classifier is used to output an emotion classification result.

[0036] Figure 2 is a flowchart of a second embodiment of the EEG emotion estimation method based on time-frequency domain features and language prompts provided by the present application, as shown in Figure 2 The method comprises the following steps: (1) Time-frequency spectrogram generation step The original EEG signal is first segmented into 3-second long segments. Then, each segment is subjected to Z-score normalization and frequency band filtering in turn. Specifically, the Z-score processed signal is processed by a third-order Butterworth band-pass filter, and only the signal components of the five key frequency bands of delta (0.5-4 Hz), theta (4-8 Hz), alpha (8-12 Hz), beta (12-30 Hz) and gamma (30-100 Hz) are retained. After that, the processed EEG segment is further decomposed into time-frequency components using continuous wavelet transform (CWT). CWT can accurately capture the energy distribution of the signal at a specific time and frequency by adjusting the scale parameter and position parameter of the wavelet basis function. After CWT is completed, a complex matrix is obtained, and the modulus value of each element in the matrix represents the intensity of the EEG signal activity at the corresponding time and frequency. The modulus values are taken as the intensity values of the pixel points, and the time-frequency spectrogram of the EEG signal is generated. In this study, Morlet wavelet is selected as the wavelet basis function of CWT, because it achieves an ideal balance between time and frequency resolution, and is very suitable for analyzing complex signals such as EEG signals with non-stationary characteristics.

[0037] (2) Time-frequency feature extraction step The present research designs a deep network framework based on improved ResNet18. Specifically, in order to better capture the texture features of the time-frequency spectrogram, the first layer of the original ResNet18 is replaced with GaborNet with biologically inspired characteristics, which initializes three learnable Gabor filters with a size of 5*5. These filters are specially designed to process the input EEG time-frequency spectrogram, by constantly optimizing the key parameters such as the direction, wavelength and phase of the filter, simulating the frequency selectivity and direction sensitivity of the human visual system, the filter adaptively focuses on the high response area (such as the highlighted part) in the spectrogram, which usually corresponds to the significant features of brain activity (such as event-related synchronization / desynchronization, ERS / ERD), by giving these key areas higher weights and more explicit directions, GaborNet accurately captures the detailed patterns of EEG signals in the time-frequency domain, and outputs the local texture feature vectors of the time-frequency spectrogram with strong discriminability. Subsequently, these local texture feature vectors are input into the deep architecture of ResNet18, which further integrates them into global electroencephalogram representation patterns (i.e. global features) through its residual structure and hierarchical convolution-pooling operations. The residual learning mechanism of ResNet18 uses cross-layer identity mapping, effectively alleviating the gradient vanishing problem of deep networks, enabling the network to be trained stably and gradually integrating multi-scale time-frequency domain information. This hierarchical feature extraction strategy can adaptively model the nonlinear dynamic evolution process of EEG signals in the time-frequency domain, and finally output global feature vectors containing local details of the spectrogram. To model the spatial dependence of multi-channel EEG signals, a double-layer graph convolutional network (Graph Convolutional Network, GCN) is incorporated at the end of the feature extraction step. It constructs a graph structure based on the anatomical location relationship of the EEG electrodes (nodes correspond to electrode channels, edge weights reflect spatial adjacency relationships), and learns the dynamic functional connectivity patterns between brain regions through graph convolution operations. GCN receives two inputs simultaneously: 1) global feature vectors extracted by the previous network; 2) a topological adjacency matrix constructed based on the international 10-20 system standard electrode positions. Through the feature propagation and aggregation mechanism, GCN effectively integrates local texture features, time-frequency dynamic patterns, and channel neural coupling information, and finally outputs enhanced time-frequency feature vectors with high representation ability.

[0038] (3) Spatio-temporal feature extraction step The application calculates a differential entropy (DE) feature of an original EEG signal to capture its time-space characteristics. In order to be consistent with the frequency bands contained in the time-frequency spectrum, the application calculates DE values of the signal in five key frequency bands of delta (0.5-4 Hz), theta (4-8 Hz), alpha (8-12 Hz), beta (12-30 Hz) and gamma (30-100 Hz) respectively. These DE values together form a differential entropy feature vector of the EEG signal.

[0039] In order to further extract the time domain characteristics of the EEG signal, the differential entropy feature vector is processed through a convolution layer and an average pooling layer to generate a final time-space feature vector. The differential entropy feature effectively reflects the dynamic changes of the EEG signal and provides important supplementary information for emotion recognition.

[0040] (4) Language prompting step The application selects a CLIP containing a tokenizer and a token embedder as a large language model (LLM) to generate a robust prompt embedding. All parameters in the CLIP are frozen. First, the application creates four text descriptions to reflect the emotional state of the subject. Then, the text description is processed by the tokenizer to convert it into a text token that can be understood by the model, and then embedded into a high-dimensional space by the token embedder to generate a prompt embedding containing semantic details.

[0041] (5) Emotion recognition step The application first adopts a feature fusion strategy to deeply fuse the time-frequency feature vector and the time-space feature vector to generate a unified feature embedding. Then, the fused feature is sent to the designed classifier for classification. The classifier adopts a multi-layer structure, in which the feature vector passes through a linear transformation layer, a batch normalization layer, and a ReLU activation function for non-linear conversion in turn, so as to introduce a non-linear factor and enhance the expressiveness of the model. In addition, by randomly discarding some neurons through the Dropout layer, the application effectively prevents the overfitting problem of the model. At the same time, the application aligns the fused feature with the text embedding, and uses the generated loss value to assist the model to more accurately classify emotions.

[0042] Figure 3 is a structural schematic diagram of the electroencephalogram emotion estimation device based on time-frequency domain features and language prompts provided by the application, comprising: The time-frequency spectrum generation module 310 is configured to segment, normalize and filter the frequency bands of the input original electroencephalogram signal, retain the signal of the preset frequency band, and generate a time-frequency spectrum through continuous wavelet transform. The time-frequency feature extraction module 320 is configured to extract a time-frequency feature vector based on the time-frequency spectrogram through a preset deep network. The space-time feature extraction module 330 is configured to calculate differential entropy of a preset frequency band based on the original electroencephalogram signal to generate a space-time feature vector. The language prompt module 340 is configured to generate a language prompt embedding from a text description reflecting the emotional state of the subject by using a large language model. The emotion recognition module 350 is configured to fuse the time-frequency feature vector and the space-time feature vector, perform modal alignment on the fused features and the language prompt embedding, and output an emotion classification result through a classifier.

[0043] It should be noted that the electroencephalogram emotion estimation device based on time-frequency domain features and language prompts provided by the embodiments of the present application can execute the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts described in any of the above embodiments when actually running, and the present embodiment will not be repeated here.

[0044] In summary, the electroencephalogram emotion estimation method and device based on time-frequency domain features and language prompts provided by the present application have the following beneficial effects compared with the prior art: (1) The recognition method based on time-frequency domain features and language prompts provided by the present application can capture the key patterns of electroencephalogram signals of subjects with different emotions, covering not only the complex dynamic changes in the time-frequency domain, but also revealing the space-time characteristics of electroencephalogram signals through the introduction of differential entropy (DE) features and language prompts, thus significantly improving the accuracy, objectivity and reliability of the evaluation.

[0045] (2) The model can dynamically optimize the direction and scale parameters of the Gabor filter and adjust the weights of the subsequent network through training of a large amount of data, so as to efficiently extract local texture information and global features in the time-frequency graph.

[0046] (3) The emotion recognition method of the present application has achieved a classification accuracy of 95.83% and 94.34% respectively on two public EEG data sets (Mumtaz data set and MODMA data set), which is better than the prior art. This shows that the system can effectively extract key features related to emotions in EEG signals and provides a new way for emotion evaluation.

[0047] Figure 4 is a structural schematic diagram of an electronic device provided by the present application, as Figure 4As shown, the electronic device can include a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 complete communication with each other through the communications bus 440. The processor 410 can invoke a logical instruction in the memory 430 to execute the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts, which includes segmenting, normalizing, and frequency band filtering an input raw electroencephalogram signal, retaining signals of a preset frequency band to generate a time-frequency spectrogram through continuous wavelet transform; extracting a time-frequency feature vector through a preset deep network based on the time-frequency spectrogram; calculating a differential entropy of the preset frequency band based on the raw electroencephalogram signal to generate a space-time feature vector; embedding a language prompt reflecting a text description of a subject's emotional state using a large language model; fusing the time-frequency feature vector and the space-time feature vector, aligning the fused features and the language prompt embedding in modalities, and outputting an emotional classification result through a classifier.

[0048] In another aspect, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, enable the computer to execute the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts provided by each of the above embodiments, which includes segmenting, normalizing, and frequency band filtering an input raw electroencephalogram signal, retaining signals of a preset frequency band to generate a time-frequency spectrogram through continuous wavelet transform; extracting a time-frequency feature vector through a preset deep network based on the time-frequency spectrogram; calculating a differential entropy of the preset frequency band based on the raw electroencephalogram signal to generate a space-time feature vector; embedding a language prompt reflecting a text description of a subject's emotional state using a large language model; fusing the time-frequency feature vector and the space-time feature vector, aligning the fused features and the language prompt embedding in modalities, and outputting an emotional classification result through a classifier.

[0049] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the electroencephalogram emotion estimation method based on time-frequency domain features and language cues provided by the above embodiments, the method comprising: segmenting, normalizing and frequency band filtering an input raw electroencephalogram signal to retain signals of a preset frequency band to generate a time-frequency spectrogram through continuous wavelet transform; extracting a time-frequency feature vector through a preset deep network based on the time-frequency spectrogram; calculating a differential entropy of the preset frequency band based on the raw electroencephalogram signal to generate a space-time feature vector; generating a language cue embedding of a text description reflecting the emotional state of a subject using a large language model; fusing the time-frequency feature vector and the space-time feature vector, modality aligning the fused features and the language cue embedding, and outputting an emotional classification result through a classifier.

[0050] Those skilled in the art can clearly understand from the description of the above embodiments that each embodiment can be implemented by means of software and the necessary general hardware platform, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in terms of contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0051] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features thereof; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A brain electrical emotion estimation method based on time-frequency domain features and language cues, characterized in that, The method comprises the following steps: segmenting, normalizing and filtering the input raw electroencephalogram signal to retain signals in a preset frequency band, and generating a time-frequency spectrogram through continuous wavelet transform; extracting a time-frequency feature vector through a preset deep network based on the time-frequency spectrogram; calculating the differential entropy of the preset frequency band based on the raw electroencephalogram signal to generate a space-time feature vector; embedding a language prompt generated by a large language model based on a text description reflecting the emotional state of a subject; fusing the time-frequency feature vector and the space-time feature vector, aligning the fused features and the language prompt embedding in modalities, and outputting an emotional classification result through a classifier. 2.The electroencephal emotion estimation method based on time-frequency domain features and language cues according to claim 1, characterized in that, The method comprises the following steps: segmenting, normalizing and filtering the input raw electroencephalogram signal to retain signals in a preset frequency band, and generating a time-frequency spectrogram through continuous wavelet transform, including: segmenting the raw electroencephalogram signal into multiple segments of a preset length; normalizing and filtering each segment to retain signals in a preset frequency band; 3.The electroencephal emotion estimation method based on time-frequency domain features and language cues according to claim 1, characterized in that, applying Morlet wavelet basis function to the signals in the preset frequency band for continuous wavelet transform to convert the signal energy distribution into a time-frequency spectrogram. The method comprises the following steps: replacing the first layer of convolutional network of ResNet18 with GaborNet and initializing three Gabor filters to extract local texture features of the time-frequency spectrogram; integrating the local texture features into global features through the residual structure and convolutional pooling operation of ResNet18; 4.The electroencephal emotion estimation method based on time-frequency domain features and language cues according to claim 1, characterized in that, setting a graph convolutional network to learn the dynamic functional connectivity pattern between brain regions and generate a time-frequency feature vector in combination with the global features. The method comprises the following steps: calculating the differential entropy of the raw electroencephalogram signal in the preset frequency band to form a differential entropy feature vector; 5.The electroencephal emotion estimation method based on time-frequency domain features and language cues according to claim 1, characterized in that, passing the differential entropy feature vector through a convolutional layer and an average pooling layer in turn to generate a space-time feature vector. The method comprises the following steps: using a large language model to generate a language prompt embedding based on a text description reflecting the emotional state of a subject, including: 6.The electroencephal emotion estimation method based on time-frequency domain features and language cues according to claim 1, characterized in that, using CLIP containing a tokenizer and a token embedder as the large language model, and all parameters in CLIP are frozen; inputting the text description reflecting the emotional state of the subject into the large language model, and processing it through the tokenizer and the token embedder in sequence to output the language prompt embedding. The method comprises the following steps: fusing the time-frequency feature vector and the space-time feature vector, aligning the fused features and the language prompt embedding in modalities, and outputting an emotional classification result through a classifier, including: fusing the time-frequency feature vector and the space-time feature vector to generate fused features; 7.The electroencephal emotion estimation method based on time-frequency domain features and language cues according to claim 1, characterized in that, inputting the fused features into the classifier for processing; calculating the loss value by aligning the fused features and the language prompt embedding in modalities; 8. An electroencephalogram emotion estimation device based on time-frequency domain features and language cues, characterized in that, optimizing the parameters of the classifier based on the loss value to output the emotional classification result using the optimized classifier. The preset frequency band is the following five frequency bands: 0.5-4 Hz, 4-8 Hz, 8-12 Hz, 12-30 Hz and 30-100 Hz. The method comprises the following steps: The time-frequency spectrum generation module is configured to segment, normalize and frequency band filter the input raw electroencephalogram signal, retain the signal of a preset frequency band, and generate a time-frequency spectrum by continuous wavelet transform. The time-frequency feature extraction module is configured to extract a time-frequency feature vector by a preset deep network based on the time-frequency spectrum. The space-time feature extraction module is configured to calculate a differential entropy of a preset frequency band based on the raw electroencephalogram signal, and generate a space-time feature vector. The language prompt module is configured to generate a language prompt embedding by using a large language model to convert a text description reflecting the emotional state of the subject. The emotion recognition module is configured to fuse the time-frequency feature vector and the space-time feature vector, perform modal alignment on the fused features and the language prompt embedding, and output an emotion classification result by a classifier.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program is executed by the processor to implement the steps of the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts according to any one of claims 1 to 7. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the electroencephalogram emotion estimation method based on time-frequency domain features and language prompts according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Suicide emotion perception method based on multi-modal fusion of voice and micro-expressions

    CN112101096A

  • Chinese continuous language text reconstruction method based on electroencephalogram signals

    CN119847336A

  • Brain electrical emotion recognition method and device based on collaborative optimization network guidance

    CN120267303A

  • CLIP-based multi-modal dynamic facial expression recognition method

    CN120356253A

  • Multimodal dynamic attention fusion

    US20220392637A1

Cited By

  • Time-frequency aligned electroencephalogram signal analysis method, device, equipment and medium

    CN121176923A