Cognitive evaluation system based on electroencephalogram signals and multi-scale regional task perception learning

By performing regional division and multi-scale feature extraction on EEG signals, combined with feature fusion of dynamically adjusting regional weights, the problem of insufficient generalization performance of existing methods in cross-task recognition is solved, and higher accuracy of cognitive state recognition is achieved.

CN120501426APending Publication Date: 2025-08-19NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510580291.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing EEG signal processing methods lack universality and are difficult to effectively capture regionally unique features and dynamic patterns of different cognitive states, resulting in insufficient generalization performance in cross-task recognition.

Method used

The spatiotemporal data division module is used to divide the EEG signals by region, and the multi-scale region independent feature extraction module and task-aware feature fusion module are dynamically adjusted, and combined with multi-scale data generation strategy, timing characteristics and dynamic changes are comprehensively extracted.

Benefits of technology

The classification performance of EEG signal data is improved, and the adaptability and accuracy in different cognitive state recognition tasks are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120501426A_ABST
    Figure CN120501426A_ABST
Patent Text Reader

Abstract

The invention discloses a cognitive evaluation system based on electroencephalogram signals and multi-scale regional task perception learning. The cognitive evaluation system comprises a spatio-temporal data division module, a multi-scale regional independent feature extraction module, a task perception feature fusion module and a result output module. The spatio-temporal data division module divides original data according to regions and cuts the original data into uniform lengths; a multi-scale region independent feature extraction module receives the segmented data region by region, senses a data sequence in a multi-scale manner, and fully extracts time sequence features in electroencephalogram signal data; the task-aware feature fusion module judges the region with the largest contribution according to the downstream task, dynamically adjusts the weight of the region, and fuses the features of different regions; and the result output module outputs the cognitive state evaluation result of the tested individual based on the fusion features, the emotion is' happy ',' neutral 'or' negative ', and the workload is' high pressure 'or' low pressure 'and the like. According to the method, the cognitive state can be accurately and effectively evaluated, and a contribution is made for improving the accuracy of a brain decoding task based on electroencephalogram signal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neuroscience and deep learning technology, and in particular to a cognitive assessment system based on electroencephalogram (EEG) signals and multi-scale regional task perception learning. Background Art

[0002] Cognitive assessment, including tasks such as workload, emotion, and motor imagery, has attracted the attention of researchers across various disciplines. Among the most powerful tools for cognitive state assessment is the brain-computer interface (BCI), which acts as a communication bridge between the brain and external devices. Among the various modalities commonly used in BCI, electroencephalography (EEG) has attracted widespread attention due to its portability, high temporal resolution, and relatively low cost. These advantages make EEG one of the most commonly used tools for cognitive state assessment, especially in dynamic and complex environments. Research on cognitive state assessment can not only deepen our understanding of brain function but also enhance the adaptability of models and systems to different tasks and environments, ultimately improving the performance of human-computer interaction systems. Cognitive state assessment methods with excellent performance can be applied to real-life scenarios to meet the real-world needs of stress monitoring in high-risk environments and prosthetic control, thereby improving work efficiency and living standards.

[0003] With the rapid advancement of deep learning technology, methods such as convolutional neural networks (CNNs) and Transformers have been widely used to learn EEG signal representations, resulting in significant improvements in cognitive state recognition capabilities. However, the structures and modules of these methods are often designed for specific tasks, specific datasets, or even specific subjects, lacking universal representation capabilities, which limits their generalization performance across different cognitive state recognition tasks. Recently, large models have demonstrated significant generalization capabilities in natural language processing (NLP) and computer vision (CV), attracting attention in the field of neuroscience. Researchers have proposed baseline models built on large-scale EEG datasets, pre-trained them using self-supervised learning, and fine-tuned them to adapt to different cognitive state recognition tasks. Baseline models with carefully designed pre-training processes and large amounts of pre-training data have achieved excellent performance on specific tasks.

[0004] Although the baseline models are effective, they face the challenge of achieving optimal performance in all cognitive state recognition tasks due to the significant differences in the underlying neural mechanisms of different cognitive states (such as emotion recognition and motor imagery). The main limitations are as follows: (1) Different regions undertake different cognitive processing tasks, and the corresponding EEG signals exhibit different temporal characteristics and dynamic patterns. Current whole-brain representation learning models fail to emphasize region-specific features and task-related inter-regional dependencies, limiting the adaptability and generalization of the methods across tasks. (2) Different cognitive processes are associated with specific brain oscillation patterns and have different active frequency bands. Using multi-scale perception can fully obtain relevant information. However, most baseline models mainly rely on single-resolution temporal modeling, which makes it difficult to capture the complete temporal information and dynamic changes of EEG signals. Summary of the Invention

[0005] Purpose of the invention: The present invention provides a cognitive assessment system based on EEG signals and multi-scale regional task perception learning, which can achieve better classification performance, perform better on downstream tasks, and help improve the accuracy of brain decoding tasks based on EEG signal data.

[0006] Technical solution: The cognitive assessment system based on EEG signals and multi-scale regional task perception learning described in the present invention includes: a spatiotemporal data division module, a multi-scale regional independent feature extraction module, a task-perception feature fusion module and a result output module; the spatiotemporal data division module divides the original data by region and cuts it into a uniform length; the multi-scale regional independent feature extraction module receives the divided data region by region, multi-scale perceives the data sequence, and fully extracts the temporal features in the EEG signal data; the task-perception feature fusion module determines the region with the greatest contribution according to the downstream task, dynamically adjusts the regional weight, and fuses the features of different regions; the result output module outputs the cognitive state assessment result of the individual subject based on the fused features.

[0007] Furthermore, the spatiotemporal data partitioning module divides the original data into regions and reorganizes the original data along the spatial dimension to obtain the channel data corresponding to each of the five regions; for each region k, define is the number of electrodes in the area, indicating the electrode set in the area in in Corresponding to the universal electrode group defined by the international 10-20 system; by combining All channel data of form complete regional data, expressed as After spatial partitioning, the entire dataset is represented as

[0008] Furthermore, the spatiotemporal data partitioning module divides the data of each region into segments of equal length. Assuming that each sample contains τ non-overlapping windows of time points, the total number of segments is The data of a specific region k is represented as X k = in i-th segment It can be expressed as [X k,iτ+1 ,X k,iτ+2 ,…,X k,(i+1)τ ].

[0009] Furthermore, the multi-scale region independent feature extraction module receives the segmented data, generates multi-scale data, and independently models the data of different scales and regions. Specifically, the steps include:

[0010] Step 1: Generate multi-scale data based on the segmented data. For the newly generated segment, the time point is sampled with a step length s, and the perception range is expanded to sτ while maintaining the segment length τ. The new segment From X k,i No. Starting from a time point, and sampling values at intervals of s until τ elements are collected, the mathematical expression is as follows:

[0011]

[0012] Before inputting data into LaBraM, it is necessary to split the channel dimension and convert the data segments into patches, which can be recorded as:

[0013]

[0014] in Indicates that it comes from the segment The patch of length τ belonging to channel j;

[0015] Step 2: Receive input data After that, the data is fed into the temporal encoder for initial feature extraction. The module consists of a one-dimensional convolution layer, a group normalization layer, and a GELU activation function. The output feature of the temporal encoder is represented as:

[0016]

[0017] Where d represents the dimension of the feature;

[0018] Step 3: The feature sequence is input to the Transformer encoder, where the multi-layer Transformer block performs feature extraction; then a pooling operation is applied to normalize the feature matrix dimensions across scales and regions, converting the original The matrix is reshaped into a u×v matrix (where u and v are adjustable variables); features of different scales are concatenated along the feature dimension, using three different scales, so the final output is represented as:

[0019] H k ={h k,i,j ∈R 3d |i=1,2,…,u,j=1,2,v}

[0020] where h k,i,j Represents the features at position i, j after the pooling operation is performed on the output of the Transformer encoder.

[0021] Furthermore, in step 2, in order to enable the model to perceive temporal and spatial information in patch embedding, a temporal embedding list TE = {te1,te2,…,te tmax} and space embedded list in After adding these embeddings to the output of the temporal encoder, the enhanced feature representation is:

[0022]

[0023] Furthermore, the task-aware feature fusion module includes a region selector and a cross-attention module; when all the multi-scale region-independent feature extraction modules output H = [H1, H2, ..., H K ], it is fed into a region selector to determine the core region. The region selector is a multi-layer perceptron (MLP) that generates a prediction value for each region and quantifies its relevance to the downstream task. The prediction values of each region are sorted, and the region with the highest prediction value is selected as the core region. Its feature representation is H argmax(softmax(MLP(H)) In the cross-attention module, the features of the identified core area are linearly projected to Q (query) as the main data for fusing features of other areas. The remaining features are concatenated along the feature dimension and projected to K (key) and V (value) respectively through the linear layer. The final output feature is expressed as:

[0024]

[0025] Where Q = L Q (H core ), K=L K (H\H core ), V=L V (H\H core ), H core Refers to the output of the selected core area, H\H coreRepresents the set of outputs of H excluding the core area, L represents the linear layer, and generates the linear representation of Q, K, and V.

[0026] Furthermore, the result output module receives the final feature Z output by the task-aware feature fusion module and uses a multi-layer perceptron (MLP) to complete the final classification. The specific classification category is determined by the specific downstream task.

[0027] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: the present invention divides the data into regions and independently models them, extracts internal features of the regions, fully considers the differences in temporal features and dynamic patterns of EEG signals between different regions, designs a task-aware feature fusion module, dynamically adjusts regional weights according to downstream tasks, emphasizes region-specific features and task-related inter-regional dependencies; proposes a multi-scale data generation strategy, and applies multi-scale data to feature extraction, comprehensively perceives time domain information, can capture the complete temporal information and dynamic changes of EEG signals, can achieve better classification performance, perform better on downstream tasks, and help improve the accuracy of brain decoding tasks based on EEG signal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 Schematic diagram of the system structure of the present invention.

[0029] Figure 2 This is a schematic diagram of the spatiotemporal data partitioning module of the present invention.

[0030] Figure 3 Schematic diagram of the multi-scale region independent feature extraction module of the present invention.

[0031] Figure 4 Schematic diagram of the task-aware feature fusion module of the present invention.

[0032] Figure 5 Schematic diagram of the overall structure of the system of the present invention. DETAILED DESCRIPTION

[0033] like Figure 1 and Figure 5As shown, a cognitive assessment system based on EEG signals and multi-scale regional task perception learning includes: a spatiotemporal data division module, a multi-scale regional independent feature extraction module, a task-perceived feature fusion module and a result output module; the spatiotemporal data division module divides the original data by region and cuts it into a uniform length; the multi-scale regional independent feature extraction module receives the segmented data region by region, multi-scale perceives the data sequence, and fully extracts the temporal features in the EEG signal data; the task-perceived feature fusion module determines the region with the greatest contribution based on the downstream task, dynamically adjusts the regional weight, and fuses the features of different regions; the result output module outputs the cognitive state assessment result of the individual subject based on the fused features. The specific description is as follows: we represent the input multi-channel EEG signal as X∈R C×T , where C is the number of electrodes and T is the total number of time points. The electrode set of X is expressed as in Corresponding to the universal electrode set defined by the international 10-20 system.

[0034] (1) Spatiotemporal data partitioning module

[0035] like Figure 2 As shown in the figure, different regions have significant task specificity and different functional characteristics in brain cognitive activities, highlighting their special roles in cognitive processes. Capturing the neural dynamics of these specific regions is crucial for accurately decoding brain states. The number of electrodes contained in each region varies, and the regions also vary according to different division criteria. Spatial segmentation is required to obtain regional data. In order to alleviate the impact of inconsistent sample lengths on model performance, time domain splitting is required to standardize the data segments in the time dimension and provide a standard data source for subsequent multi-scale data generation. In order to overcome these limitations and complete the above operations, a spatiotemporal data segmentation module is designed to divide the EEG signals into multiple segments along the spatial and temporal dimensions.

[0036] In the first step, the spatiotemporal data partitioning module reorganizes the original data along the spatial dimension to obtain the channel data corresponding to each of the five regions; for each region k, define is the number of electrodes in the area, indicating the electrode set in the area in Then, by combining All channel data of form complete regional data, expressed as After spatial partitioning, the entire dataset can be represented as

[0037] In the second step, we split the data of each region into segments of equal length. Assuming that each sample contains τ non-overlapping windows of time points, the total number of segments is Therefore, the data of region k can be expressed as in i-th segment It can be expressed as [X k,iτ+1 ,X k,iτ+2 ,…,X k,(i+1)τ ].

[0038] (2) Multi-scale region independent feature extraction module

[0039] like Figure 3 As shown in the figure, considering the differences in the requirements of perception tasks in different regions, it is proposed to independently learn the temporal dynamics in each region and the spatial dependencies between electrodes. In addition, in order to solve the single-scale limitation of the baseline model and enable the model to capture information across different time scales, an additional multi-scale temporal feature learning framework is introduced, using a step sampling strategy. A multi-scale region-independent feature extraction module is introduced to complete the above operations, aiming to enhance the model's ability to extract comprehensive spatiotemporal representations. This module consists of two main components: using segmented data to generate multi-scale data, and applying the LaBraM encoder to the multi-scale data to complete feature extraction and feature fusion. The detailed steps are as follows:

[0040] First, multi-scale data generation is performed based on the segmented data. For the newly generated segments, the time points are sampled with a step size s, and the perception range is expanded to sτ while maintaining the segment length τ. Specifically, the new segment From X k,i No. The mathematical expression is as follows:

[0041]

[0042] The input of LaBraM needs to segment the channel data and convert the data segments into patches, which can be recorded as:

[0043]

[0044] in Indicates that it comes from the segment of the patches of length τ belonging to channel j.

[0045] Receiving input data After that, the data is first fed into the temporal encoder for initial feature extraction. This module consists of a one-dimensional convolutional layer, a group normalization layer, and a GELU activation function. The output feature of the temporal encoder is represented as:

[0046]

[0047] Where d represents the dimension of the feature. In order to enable the model to perceive the temporal and spatial information in the patch embedding, the temporal embedding list TE = {te1,te2,…,te tmax} and space embedded list in After adding these embeddings to the output of the temporal encoder, the enhanced feature representation is:

[0048]

[0049] This feature sequence is then fed into the Transformer encoder, where multiple layers of Transformer blocks perform feature extraction. Pooling operations are then applied to normalize the dimensions of the cross-scale feature matrix, converting the original The matrix is reshaped into a u×v matrix (where u and v are adjustable variables and their values remain unchanged in different regions). Finally, features of different scales are concatenated along the feature dimension. Three different scales are used in the experiment, so the final output of this module can be expressed as:

[0050] H k ={h k,i,j ∈R 3d |i=1,2,…,u,j=1,2,v}

[0051] where h k,i,j Represents the features at position i, j after the pooling operation is performed on the output of the Transformer encoder.

[0052] (3) Task-aware feature fusion module

[0053] like Figure 4 As shown in Figure 2, considering that different cognitive tasks involve different core regions, a task-aware feature fusion module is designed to fuse features across multiple regions. Unlike traditional feature fusion methods that treat all regions equally, this module adaptively emphasizes task-related core regions, ensuring that the fused features are both biologically and functionally meaningful.

[0054] The task-aware feature fusion module consists of a region selector and a cross-attention module. When all the multi-scale region-independent feature extraction modules output H = [H1, H2, ..., H K ], they are first fed into a region selector to determine the core region. The region selector is a multi-layer perceptron (MLP) that generates a prediction value for each region and quantifies its relevance to the downstream task. The prediction values of each region are sorted, and the region with the highest prediction value is selected as the core region, and its feature representation is H argmax(softmax(MLP(H))In the cross-attention module, the features of the identified core region are linearly projected to Q (query) as the main data for fusing features of other regions. The remaining features are concatenated along the feature dimension and projected to K (key) and V (value) respectively through a linear layer. This mechanism ensures that task-related information is prioritized during the fusion process, producing feature representations with higher robustness and interpretability. The final output features can be expressed as:

[0055]

[0056] Where Q = L Q (H core ), K=L K (H\H core ), V=L V (H\H core ), H core Refers to the output of the selected core area, H\H core Represents the set of outputs of H excluding the core area, L represents the linear layer, and generates the linear representation of Q, K, and V.

[0057] (4) Result output module

[0058] The result output module receives the final feature Z output by the task-aware feature fusion module and uses a multi-layer perceptron (MLP) to complete the final classification. The specific classification category is determined by the specific downstream task.

[0059] Extensive experiments were conducted on three datasets for comprehensive evaluation: SEED, PWED, and BCIC4-1, which are public datasets for emotion recognition, workload recognition, and motor imagery recognition.

[0060] Table 1 Comparative experiments between the system and existing methods on three datasets

[0061]

[0062]

[0063] As shown in Table 1, our system outperforms the comparison methods across all experimental tasks. Compared to directly fine-tuning the LaBraM model, the accuracy metrics improve by 4.16%, 2.03%, and 4.14% in emotion recognition, workload recognition, and motor imagery recognition, respectively. Furthermore, our proposed method achieves the highest performance across all tasks, while some comparison methods, such as EEGNet and ATCNet, only achieve higher performance in certain tasks. These models are more susceptible to task type, while our proposed method demonstrates stable performance across a wide range of tasks, preserving the advantages of LaBraM.

[0064] To demonstrate the effectiveness of each component in the proposed system, we started with the original LaBraM and gradually added various modules until it became a complete system. The specific experimental results are shown in Table 2. The region-independent mechanism only divides the regions for independent modeling, without adding a multi-scale mechanism. When the task-aware feature fusion module is not used as the feature fusion method, the features of each region are sequentially input as Q into the cross-attention, and the final features are obtained through splicing. The experimental results show that the addition of each module brings improved performance, indicating that each module in the system can help extract more valuable features.

[0065] Table 2 Validation of each module in the system

[0066]

[0067]

[0068] The present invention divides the data into regions and independently models them, extracts regional internal features, fully considers the differences in EEG signal temporal features and dynamic patterns between different regions, designs a task-aware feature fusion module, dynamically adjusts regional weights according to downstream tasks, and emphasizes region-specific features and task-related inter-regional dependencies. A multi-scale data generation strategy is proposed, and multi-scale data is applied to feature extraction to fully perceive time domain information, capable of capturing the complete temporal information and dynamic changes of EEG signals. LaBraM is added to the system as a feature extraction model, fully utilizing its advantages of strong generalization performance and representation learning ability. It has achieved better performance than existing methods on downstream tasks, which helps to improve the accuracy of brain decoding tasks based on EEG signal data.

Claims

1. A cognitive assessment system based on EEG signals and multi-scale regional task perception learning, characterized by: include: The spatiotemporal data partitioning module, the multi-scale region-independent feature extraction module, the task-aware feature fusion module, and the result output module are included. The spatiotemporal data partitioning module divides the original data by region and cuts it into uniform lengths. The multi-scale region-independent feature extraction module receives the segmented data region by region, perceives the data sequence at multiple scales, and fully extracts the temporal features in the EEG signal data. The task-aware feature fusion module determines the region with the greatest contribution based on the downstream task, dynamically adjusts the regional weights, and fuses the features of different regions. The result output module outputs the cognitive status assessment results of the individual subjects based on the fusion features.

2. The cognitive assessment system based on EEG signals and multi-scale regional task perception learning according to claim 1, characterized in that: The spatiotemporal data partitioning module divides the original data into regions and reorganizes the original data along the spatial dimension to obtain the channel data corresponding to each of the five regions; for each region k, define is the number of electrodes in the area, indicating the electrode set in the area in in Corresponding to the universal electrode group defined by the international 10-20 system; by combining All channel data of form complete regional data, expressed as After spatial partitioning, the entire dataset is represented as 3. The cognitive assessment system based on EEG signals and multi-scale regional task perception learning according to claim 1, characterized in that: The spatiotemporal data partitioning module divides the data of each region into segments of equal length. Assuming that each sample contains τ non-overlapping windows of time points, the total number of segments is The data of a specific region k is expressed as in The i-th segment It can be expressed as [X k,iτ+1 ,X k,iτ+2 ,…,X k,(i+1)τ ].

4. The cognitive assessment system based on EEG signals and multi-scale regional task perception learning according to claim 1, characterized in that: The multi-scale region independent feature extraction module receives the segmented data, generates multi-scale data, and independently models the data in different scales and regions. Specifically, it includes the following steps: Step 1: Generate multi-scale data based on the segmented data. For the newly generated segments, the time points are sampled with a step size of s, while maintaining the segment length At the same time, the perception range is expanded to sτ, the new segment From X k,i No. Start at a time point and sample values at intervals of s until the collection elements, the mathematical expression is as follows: Before inputting data into LaBraM, it is necessary to split the channel dimension and convert the data segments into patches, which can be recorded as: in Indicates that it comes from the segment The patch of length τ belonging to channel j; Step 2: Receive input data After that, the data is fed into the temporal encoder for initial feature extraction. The module consists of a one-dimensional convolution layer, a group normalization layer, and a GELU activation function. The output feature of the temporal encoder is represented as: Where d represents the dimension of the feature; Step 3: The feature sequence is input to the Transformer encoder, where the multi-layer Transformer block performs feature extraction; then a pooling operation is applied to normalize the feature matrix dimensions across scales and regions, converting the original The matrix is reshaped into a u×v matrix, where u and v are adjustable variables; features of different scales are concatenated along the feature dimension, using three different scales, so the final output is represented as: H k ={h k,i,j ∈R 3d ∣i=1,2,…,u,j=1,2,v} where h k,i,j Represents the features at position i, j after the pooling operation is performed on the output of the Transformer encoder.

5. The cognitive assessment system based on EEG signals and multi-scale regional task perception learning according to claim 4, characterized in that: In step 2, in order to enable the model to perceive the temporal and spatial information in the patch embedding, the temporal embedding list TE = {te1,te2,…,te tmax } and the spatial embedding list SE = {se1,se2,…,se ∣C∣ },in After adding these embeddings to the output of the temporal encoder, the enhanced feature representation is:

6. The cognitive assessment system based on EEG signals and multi-scale regional task perception learning according to claim 1, characterized in that: The task-aware feature fusion module includes a region selector and a cross-attention module; when all the multi-scale region independent feature extraction modules output H = [H1, H2, ..., H K ], it is fed into a region selector to determine the core region. The region selector is a multi-layer perceptron MLP that generates a prediction value for each region and quantifies its relevance to the downstream task; The predicted values of each region are sorted, and the region with the highest predicted value is selected as the core region, whose characteristic is represented by H argmax(softmax(MLP(H)) In the cross-attention module, the features of the identified core area are linearly projected to Q (query) as the main data for fusing features of other areas. The remaining features are concatenated along the feature dimension and projected to K (key) and V (value) respectively through the linear layer. The final output feature is expressed as: Where Q = L Q (H core ), K=L K (H\H core ), V=L V (H\H core ), H core Refers to the output of the selected core area, H\H core Represents the set of outputs of H excluding the core area, L represents the linear layer, and generates the linear representation of Q, K, and V.

7. The cognitive assessment system based on EEG signals and multi-scale regional task perception learning according to claim 1, characterized in that: The result output module receives the final feature Z output by the task-aware feature fusion module and uses the multi-layer perceptron MLP to complete the final classification. The specific classification category is determined by the specific downstream task.

Citation Information

Patent Citations

  • Method and system for training electroencephalogram model for learning electroencephalogram general characterization

    CN118296379A

  • Electroencephalogram signal analysis method and equipment based on multi-scale electroencephalogram feature fusion

    CN118551340A

  • Unmanned aerial vehicle control system based on brain-computer interface

    CN119538982A

  • KR20250047575A