A research method for visual attention mechanisms based on EEG microstates
By collecting and analyzing the electroencephalogram (EEG) signals of subjects while watching videos, extracting saliency map features and performing EEG microstate analysis, a decoding model was constructed. This addresses the shortcomings of existing technologies in the study of visual attention mechanisms and enables accurate decoding and dynamic information acquisition of visual attention mechanisms.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN UNIV
- Filing Date
- 2023-12-25
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies have failed to effectively utilize EEG microstates to study visual attention mechanisms and lack direct methods for interpreting information perceived by the human brain from the outside world.
By collecting EEG signals from subjects while they watch videos, preprocessing them, extracting sIQR and tsIQR features from the saliency map, performing EEG microstate analysis, constructing a decoding model to decode visual attention, and utilizing the high temporal resolution of EEG microstates to extract microstate features and depth features for decoding.
This study enabled a precise exploration of the mechanisms of visual attention, obtained spatiotemporal dynamic information of whole-brain activity, verified the link between EEG microstates and video saliency, and effectively decoded visual attention.
Smart Images

Figure CN117653115B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of visual attention mechanism research technology, and in particular to a method for researching visual attention mechanism based on EEG microstates. Background Technology
[0002] Visual attention is a crucial filtering mechanism in the human visual system, but our understanding of how the human brain perceives external input information and how neurons respond consciously or unconsciously remains insufficient. Electroencephalography (EEG), a non-invasive technique for monitoring brain electrophysiological activity, records neuronal electrical activity signals through scalp electrodes, providing a more objective and direct interpretation of brain responses from a neurophysiological perspective, thus offering an efficient method for studying the mechanisms of visual attention. Among these techniques, EEG microstate analysis, a computational method effectively exploring the spatiotemporal dynamics of brain electrical activity, has been applied in numerous neuroscience studies. Furthermore, EEG microstates have been found to directly and qualitatively characterize the brain's perceptual and cognitive processes; however, they have not yet been used to study the mechanisms of visual attention. Summary of the Invention
[0003] The technical problem to be solved by this invention is to provide a method for studying the mechanism of visual attention based on brainwave microstates, so as to effectively explore visual attention.
[0004] To solve the above-mentioned technical problems, the objective of this invention is achieved through the following technical solution: A method for studying visual attention mechanisms based on EEG microstates is provided, comprising the following steps: EEG data acquisition: EEG signals are collected from several subjects while watching a preset video; the collected EEG signals are preprocessed to obtain preprocessed EEG signals; Video saliency analysis: A computational visual attention model is used to extract saliency maps from the preset video; statistical features are extracted from the saliency maps from the perspectives of spatial saliency information in the temporal domain and spatiotemporal saliency information changes in the time series, respectively, to obtain sIQR features and tsIQR features; EEG microstate analysis: The preprocessed EEG signals are transformed into a sequence of EEG topology maps; the transformed EEG topology map sequences are then analyzed... Linear spatial clustering analysis was used to extract EEG microstate templates, determine the number of EEG microstate templates, and select the final microstate templates based on the number of EEG microstate templates. The selected final microstate templates were then backfitted to the preprocessed EEG signals to obtain microstate sequences. Visual attention decoding: Microstate features and depth features were extracted from the microstate sequences. Statistical analysis methods were used to test whether there were statistical differences between microstate features and sIQR and tsIQR features. Decoding models based on microstate features or depth features were constructed. The constructed decoding models were used to perform segment decoding and video decoding, respectively, to obtain corresponding segment labels and video labels. Among them, microstate features include the time proportion, duration, frequency of occurrence, and transition probability of each final microstate template.
[0005] The beneficial technical effects of this invention are as follows: The method for studying visual attention mechanisms based on EEG microstates of this invention obtains corresponding microstate sequences by processing preprocessed EEG signals, and then obtains corresponding microstate features and depth features based on the microstate sequences. A corresponding feature decoding model is then constructed for decoding to achieve the analysis of visual attention mechanisms through EEG microstates. Utilizing the high temporal resolution of EEG microstates, it can accurately acquire spatiotemporal dynamic information of whole-brain activity, enabling a better and more effective exploration of visual attention. Furthermore, by extracting potential depth features from the microstate sequences, it can more effectively extract [the necessary information / mechanisms]. This study utilizes semantic information related to visual attention in the brain to achieve better visual attention decoding. Simultaneously, a computational visual attention model is employed to extract saliency maps from a pre-defined video. Statistical features are extracted from the obtained saliency maps from both the perspective of local spatial saliency information in the temporal domain and the perspective of global spatiotemporal saliency information changes in the time series, yielding sIQR and tsIQR features. Statistical analysis methods are used to examine whether there are statistical differences between microstate features and sIQR and tsIQR features, demonstrating the connection between EEG microstates and video saliency. This verifies that decoding based on EEG microstates can be used to study the mechanisms of visual attention. Attached Figure Description
[0006] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0007] Figure 1 A flowchart illustrating the method for studying visual attention mechanisms based on EEG microstates provided in this embodiment of the invention;
[0008] Figure 2 A flowchart illustrating the video saliency analysis process of the method for studying visual attention mechanisms based on EEG microstates provided in this embodiment of the invention.
[0009] Figure 3 for Figure 2 The saliency map and schematic diagram of the feature extraction process for video saliency analysis of the research method of visual attention mechanism based on EEG microstate are shown.
[0010] Figure 4 A schematic diagram of the structure of the deep learning LSTM-AE model for the research method of visual attention mechanism based on EEG microstates provided in the embodiments of the present invention. Detailed Implementation
[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0012] Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for studying visual attention mechanisms based on electroencephalogram (EEG) microstates, as provided in an embodiment of the present invention. The method includes the following steps:
[0013] Step S11: EEG data acquisition: Collect EEG signals from several subjects while they watch a preset video. Preprocess the collected EEG signals to obtain preprocessed EEG signals, thereby improving the signal-to-noise ratio of the EEG signals and removing artifact noise.
[0014] The preset videos are designed to elicit bottom-up visual attention activity, used to acquire EEG data induced by this activity. Up to 18 preset videos are required, all featuring real-world scenarios selected from a publicly available database of human activity videos. Each preset video has a frame width of 960 pixels, a frame height of 540 pixels, an average frame rate of 28.81 fps, and a duration of 58 seconds. The videos must be in color and cover various everyday human activities, such as dancing, playing badminton, gymnastics, and playing with pets. The videos must contain a sufficient number of visually salient objects to ensure unconscious attention can be induced, and should not contain excessive dramatic changes or unexpected scenes to avoid introducing unforeseen factors during data acquisition. The BrainAmp EEG signal acquisition system is used to collect the subjects' EEG signals at a sampling rate of 1000 Hz and 64 electrodes. The number of participants can be up to 50, with a male-to-female ratio of 1:1, and the participants' ages can be 19.940±1.878. The EEG signals are multi-channel EEG signals.
[0015] Specifically, the preprocessing of the acquired EEG signals in step S11 includes:
[0016] The acquired EEG signals are subjected to operations such as eye movement (IO) channel removal, common average rereference, filtering, independent component analysis (ICA), artifact removal, and bad channel interpolation to improve the signal-to-noise ratio and remove artifact noise, thus obtaining preprocessed EEG signals.
[0017] Step S12, Video Saliency Analysis: A computational vision attention model is used to extract saliency maps from the preset video. Statistical features are extracted from the saliency maps from both the perspective of local spatial saliency information in the temporal domain and the perspective of global spatiotemporal saliency information changes over time, resulting in sIQR (spatial interquartile range) and tsIQR (tempo-spatial interquartile range) features. sIQR and tsIQR features are local and global features, respectively. sIQR describes the dispersion of pixel values in the saliency map; a more dispersed distribution indicates richer and more complex saliency information in the image. tsIQR describes the degree to which the complexity of the saliency map's saliency information changes over time.
[0018] Step S13, EEG microstate analysis: The preprocessed EEG signal is converted into an EEG topology sequence. Spatial clustering analysis is performed on the converted EEG topology sequence to extract EEG microstate templates. The number of EEG microstate templates is determined. The final microstate template is selected based on the number of EEG microstate templates. The selected final microstate template is backfitted to the preprocessed EEG signal to obtain a microstate sequence.
[0019] Step S14, Visual Attention Decoding: Extract microstate features and depth features from the microstate sequence. Use statistical analysis to test whether there are statistical differences between microstate features and sIQR and tsIQR features. Construct a decoding model based on microstate features or a decoding model based on depth features. Use the constructed decoding models to perform segment decoding and video decoding respectively to obtain corresponding segment labels and video labels. Microstate features include the time coverage, duration, occurrence frequency, and transition probability of each final microstate template to reflect the occurrence patterns of different EEG microstates and to study potential brain working patterns. Time coverage refers to the proportion of the total duration of a single EEG microstate to the total duration of the microstate sequence; duration refers to the average duration of a single EEG microstate, in milliseconds; occurrence frequency refers to the number of times a single EEG microstate occurs within one second; and transition probability refers to the probability that a single EEG microstate transitions to another EEG microstate. EEG microstates contain the potential information of all electrode channels in the EEG signal and have high temporal resolution. Analyzing EEG microstates can accurately reflect the spatiotemporal dynamics of whole-brain activity. The proposed method for studying visual attention mechanisms based on EEG microstates involves processing preprocessed EEG signals to obtain corresponding microstate sequences. Then, by acquiring corresponding microstate features and depth features from these sequences, a corresponding feature decoding model is constructed for decoding. This allows for the analysis of visual attention mechanisms through EEG microstates. Utilizing the high temporal resolution of EEG microstates, the spatiotemporal dynamics of whole-brain activity can be accurately obtained for a more effective exploration of visual attention. Furthermore, by extracting potential depth features from the microstate sequences, the method can more effectively extract information related to brain visual attention. Force-related semantic information is used to achieve better visual attention decoding. At the same time, a computational visual attention model is used to extract saliency maps from a preset video. Statistical features are extracted from the obtained saliency maps from the perspectives of local spatial saliency information in the temporal domain and global spatiotemporal saliency information changes in the time series, to obtain sIQR features and tsIQR features. Statistical analysis methods are used to test whether there are statistical differences between microstate features and sIQR and tsIQR features, to prove the connection between EEG microstates and video saliency, thereby verifying that EEG microstates can be decoded to study the mechanism of visual attention.
[0020] Combination Figure 2 and Figure 3 , Figure 2 The flowchart of the video saliency analysis method for the research method of visual attention mechanism based on EEG microstates of the present invention is demonstrated. Figure 3The present invention demonstrates the saliency map and feature extraction process of the video saliency analysis of the method for studying visual attention mechanisms based on EEG microstates. Specifically, step S12 includes:
[0021] Step S121, Saliency Map Extraction: Saliency map extraction is performed frame by frame on the preset video according to the computational vision attention model. Saliency map extraction is performed on the preset video based on frames to obtain frame saliency maps and frame saliency map sequences. Saliency map extraction is performed on the frame saliency map sequences according to a preset duration to obtain the frame saliency map corresponding to the preset duration as the saliency map extracted on the preset video based on segments to obtain segment saliency maps and segment saliency map sequences.
[0022] Specifically, step S121 is as follows:
[0023] Each frame of a preset video is input into a computational vision attention model to obtain a saliency map for each frame. The frame saliency maps are then sorted according to the temporal order of the preset videos to obtain a frame saliency map sequence. This sequence is then segmented according to a preset duration. The average value of the frame saliency maps within each preset duration is calculated to obtain segment saliency maps for each preset duration. These segment saliency maps are then sorted according to the temporal order of the preset videos to obtain a segment saliency map sequence. Since the time interval between frames in a video is short and frame changes are minimal, segmenting the frame saliency map sequence by a preset duration allows for a better exploration of the local changes in the frame saliency maps over time and reduces redundant information. The preset duration can be 0.5 seconds.
[0024] Step S122, Saliency Feature Extraction: Extract sIQR features from the obtained segment saliency map to quantitatively describe the dispersion of pixel values in the saliency map and eliminate the influence of outliers as much as possible. Quantitatively analyze the saliency information complexity of the saliency map. Sort and integrate the obtained sIQR features according to the preset video duration to obtain the sIQR sequence. Extract tsIQR features from the obtained sIQR sequence to explore the brain's dynamic response to changes in spatial saliency information of images over time. Quantitatively analyze the global temporal-spatial saliency information changes of the saliency map in temporal order.
[0025] Specifically, step S122 includes:
[0026] To obtain sIQR features, the interquartile range of all pixels in the fragment saliency map is calculated. Alternatively, the sIQR features can be obtained by calculating the difference between the upper and lower quartiles of all pixels in the fragment saliency map. The upper quartile refers to the 75th percentile of the data distribution, and the lower quartile refers to the 25th percentile of the data distribution.
[0027] The sIQR feature can then be calculated using formula (1):
[0028] sIQR=Q3-Q1(1)
[0029] In the formula, sIQR represents the sIQR feature, Q3 represents the upper quartile of the pixel value of the fragment saliency map, and Q1 represents the lower quartile of the pixel value of the fragment saliency map.
[0030] The obtained sIQR features are sorted and integrated according to the preset video duration to obtain an sIQR sequence. The interquartile range of the sIQR sequence is calculated to extract tsIQR features. By extracting tsIQR features, the dispersion of the sIQR sequence data can be quantitatively described, thus revealing the degree of change in the spatial saliency information complexity of the saliency map in the saliency map sequence in the time domain.
[0031] The tsIQR feature can then be calculated using formula (2):
[0032] tsIQR=Q3(sIQR)-Q1(sIQR) (2)
[0033] In the formula, tsIQR represents the tsIQR feature, Q3(sIQR) represents the upper quartile of the sIQR sequence data, and Q1(sIQR) represents the lower quartile of the sIQR sequence data.
[0034] Specifically, the time proportion of each final microstate template of the microstate feature in step 14 can be calculated using formula (3):
[0035]
[0036] In the formula, Cov k N represents the time percentage of the k-th final microstate template in the microstate sequence. k T represents the number of times the k-th final microstate template in the microstate sequence continues to appear. k,n T represents the time, in seconds, when the k-th final microstate template in the microstate sequence appears for the nth time. total This represents the total duration of the microstate sequence, in seconds.
[0037] The duration of each final microstate template of the microstate feature in step 14 can be calculated using formula (4):
[0038]
[0039] In the formula, Dur k N represents the duration of the k-th final microstate template in the microstate sequence, in milliseconds. k T represents the number of times the k-th final microstate template in the microstate sequence continues to appear. k,nThis represents the time, in seconds, when the k-th final microstate template appears for the nth time in the microstate sequence.
[0040] The frequency of occurrence of each final micro-state template of the micro-state feature in step 14 can be calculated using formula (5):
[0041]
[0042] In the formula, Occ k N represents the frequency of the k-th final microstate template in the microstate sequence. k T represents the number of times the k-th final microstate template in the microstate sequence continues to appear. total This represents the total duration of the microstate sequence, in seconds.
[0043] The transition probabilities of each final microstate template of the microstate features in step 14 can be calculated using formula (6):
[0044]
[0045] In the formula, TP k,l S represents the probability that the k-th final microstate template in the microstate sequence is transformed into the l-th final microstate template. k,l S represents the total number of times the k-th final microstate template in the microstate sequence is transformed into the l-th final microstate template. k,l′ S represents the total number of times the k-th final microstate template in the microstate sequence is converted to the l′-th final microstate template, where l′>l. K represents the number of EEG microstate templates. When l′=k, no conversion is required. k,l′ =0.
[0046] Specifically, in this embodiment, the number of EEG microstate templates can be five, that is, the number of final microstate templates is five. Each final microstate template can be converted into four other final microstate templates besides itself. That is, each final microstate template has four transition probabilities. Therefore, the number of microstate features corresponding to each final microstate template is seven, including a time percentage, a duration, a frequency of occurrence, and four transition probabilities. Accordingly, the number of microstate features in the entire microstate sequence is thirty-five.
[0047] Specifically, step S14, which involves using statistical analysis to examine whether there are statistical differences between all extracted microstate features and sIQR features, and between all microstate features and tsIQR features, includes:
[0048] The median was used to perform binary classification on sIQR and tsIQR features, obtaining sIQR and tsIQR feature classification results. The sIQR feature classification results correspond to segment binary classification results, i.e., true segment labels, while the tsIQR feature classification results correspond to video binary classification results, i.e., true video labels. The corresponding classification results include high and low classes, representing video saliency; that is, sIQR feature classification results include high sIQR features and low sIQR features, and tsIQR feature classification results include high tsIQR features and low tsIQR features. This binary classification of sIQR and tsIQR features aims to explore how EEG microstates reflect the human visual system's screening process for salient objects. The sIQR and tsIQR feature classification results are shown in the table below. In the table, the mean and standard value refer to the values of the sIQR or tsIQR features.
[0049]
[0050] The extracted microstate features are binary-classified based on the time points corresponding to the sIQR and tsIQR feature classification results to obtain microstate feature classification results. The microstate features of the subjects are extracted from the EEG signals collected during video viewing and contain temporal information. Based on the corresponding time points, all subjects' microstate features can be divided into two parts according to the sIQR and tsIQR feature classification results in the table above, obtaining microstate feature classification results. The microstate feature classification results are shown in the table below and can be used as data input for subsequent modeling. The table below shows the relationship between microstate features and the high and low sIQR and tsIQR features.
[0051]
[0052] Based on the obtained microstate feature classification results, sIQR feature classification results, and tsIQR feature classification results, statistical analysis methods were used to test whether there were statistical differences between the segment-based microstate features and the video-based microstate features and the sIQR and tsIQR features. Specifically, statistical analysis methods were used to test whether there were statistical differences between the microstate features in the high and low classes of sIQR features and the high and low classes of tsIQR features.
[0053] Specifically, step S14 includes:
[0054] Microstate feature extraction: Extract segment-based microstate features and video-based microstate features from the microstate sequence respectively; where segment refers to a video segment divided according to a preset duration, and video-based refers to a complete preset video;
[0055] Deep feature extraction: Deep learning LSTM-AE method is used to extract segment-based deep features and video-based deep features from the micro-state sequence; where segment refers to video segments divided according to a preset duration, and video-based refers to the complete preset video.
[0056] Difference test: Statistical analysis methods were used to test whether there were statistical differences between all extracted microstate features and sIQR features, and between all microstate features and tsIQR features;
[0057] Prediction: Construct a decoding model based on micro-state features, and use the constructed decoding model based on micro-state features to perform segment decoding and video decoding respectively on segment-based micro-state features and video-based micro-state features to obtain the corresponding segment labels and video labels; or, construct a decoding model based on depth features, and use the constructed decoding model based on depth features to perform segment decoding and video decoding respectively on segment-based depth features and video-based depth features to obtain the corresponding segment labels and video labels.
[0058] Specifically, the prediction steps include:
[0059] The micro-state features based on segments and the micro-state features based on videos are input into a classifier and subjected to a 10-fold cross-validation binary classification test to predict video saliency and obtain the corresponding segment labels and video labels; or, the deep features based on segments and the deep features based on videos are input into a classifier and subjected to a 10-fold cross-validation binary classification test to predict video saliency and obtain the corresponding segment labels and video labels.
[0060] Specifically, the prediction step further includes:
[0061] Decoding Verification: The obtained segment tags and video tags are compared with the real segment tags and real video tags, respectively. The decoding accuracy is verified based on the comparison results. Specifically, the decoding capability and effect are demonstrated by comparing the obtained segment tags and video tags with the corresponding real tags in terms of accuracy, precision, recall, and F1 score.
[0062] Specifically, the classifier can be a Gaussian Naive Bayes classifier. Bains (GNB), decision tree (DT), support vector machine (SVM), Adaboost classifier, Bagging classifier, extremely randomized trees (ET) classifier, and K-nearest neighbors (KNN) classifier.
[0063] Specifically, the deep learning LSTM-AE method is as follows:
[0064] Construct a deep learning LSTM-AE model. The deep learning LSTM-AE model includes an encoder and a decoder. The encoder and decoder are connected, and each encoder and decoder includes several sequentially connected LSTM units. The last LSTM unit of the encoder is connected to the first LSTM unit of the decoder. Each LSTM unit of the decoder is connected to a fully connected layer. The last LSTM unit refers to the last LSTM unit, and the first LSTM unit refers to the first LSTM unit.
[0065] The microstate sequence is sequentially input into the LSTM unit of the encoder of the deep learning LSTM-AE model in chronological order; wherein, each LSTM unit receives the EEG microstate at a corresponding single moment in chronological order, and the information in each LSTM unit of the encoder is passed sequentially so that the information at the corresponding moment in each LSTM unit of the encoder is passed to the next LSTM unit for calculation. The information in each LSTM unit of the encoder includes unit state variables and hidden state variables. The unit state variables and hidden state variables of each LSTM unit of the encoder are passed to the next LSTM unit. The hidden state variables output by the last LSTM unit of the encoder are used as the depth features of the microstate sequence, thus obtaining the depth features. The unit state variables and hidden state variables are both vectors of the same dimension.
[0066] The cell state variables and hidden state variables of the last LSTM unit of the encoder are used as the initial values of the information in the first LSTM unit of the decoder. The information in each LSTM unit of the decoder is passed sequentially so that the information at the corresponding time in each LSTM unit of the decoder is passed to the next LSTM unit for calculation. The information in each LSTM unit of the decoder includes the target cell state variables and the target hidden state variables. Therefore, the cell state variables and hidden state variables of the last LSTM unit of the encoder are used as the initial values of the target cell state variables and the target hidden state variables of the first LSTM unit of the decoder.
[0067] The EEG microstate of the last time step of the microstate sequence is input into the first LSTM unit of the decoder. The output of the first LSTM unit of the decoder is the target unit state variable and the target hidden state variable of the last time step. The target unit state variable and the target hidden state variable of the corresponding time step of the previous LSTM unit of the decoder are passed to the next LSTM unit of the decoder for calculation. The target hidden state variable of the last time step output by the first LSTM unit of the decoder is transformed in dimension through the fully connected layer to obtain a feature vector with the same dimension as the EEG microstate of the last time step of the microstate sequence, which is used as the target EEG microstate of the last time step. The target EEG microstate at a given moment is input into the next LSTM unit of the decoder for computation. This process is repeated. The target hidden state variables of each LSTM unit of the decoder are transformed in dimension by a fully connected layer to obtain the target EEG microstate at the corresponding moment of that LSTM unit. The target EEG microstate at the corresponding moment of the previous LSTM unit of the decoder is used as the input of the next LSTM unit. The target EEG microstates at the corresponding moments are obtained according to the order of the LSTM units of the decoder. The reconstructed microstate sequence is obtained by arranging the corresponding target EEG microstates according to the order of the LSTM units of the decoder. The time order of the reconstructed microstate sequence is the reverse of the time order of the original microstate sequence.
[0068] The structure of the deep learning LSTM-AE model is as follows: Figure 4 As shown in the figure, h1, h2, ..., h T-1 ,h T h′ represents the hidden state variables output sequentially by each LSTM unit of the encoder. T ,h′ T-1 ,h′ T-2 h′1 represents the target hidden state variables output sequentially by each LSTM unit of the decoder, c1, c2, ..., c T-1 ,c T c′ represents the unit state variables output sequentially by each LSTM unit of the encoder. T ,c′ T-1 ,c′ T-2 c′1 represents the sequential output state variables of each LSTM unit of the decoder, m1, m2, ..., m T-1 ,m T Let m′ represent the original microstate sequence of duration T. T ,m′ T-1 ,...,m′2,m′1 represent the reconstructed microstate sequence with a duration of T.
[0069] Microstate sequences are time series containing rich brain information; however, microstate features are condensed statistical features from these sequences, which are insufficient to encompass the dynamic changes of microstates over time. By employing the deep learning LSTM-AE method to extract latent deep features from microstate sequences, we can better capture the dynamic changes of microstates over time and more effectively extract semantic information related to visual attention from these sequences, thus achieving better visual attention decoding.
[0070] In summary, the method for studying visual attention mechanisms based on EEG microstates of this invention obtains corresponding microstate sequences by processing preprocessed EEG signals. Then, it acquires corresponding microstate features and depth features from these sequences and constructs a corresponding feature decoding model for decoding. This enables the analysis of visual attention mechanisms through EEG microstates. Utilizing the high temporal resolution of EEG microstates, it can accurately acquire spatiotemporal dynamic information of whole-brain activity, allowing for a more effective exploration of visual attention. Furthermore, by extracting potential depth features from the microstate sequences, it can more effectively extract information relevant to the brain's visual attention mechanisms. This study extracts semantic information related to visual attention to achieve better visual attention decoding. Simultaneously, a computational visual attention model is used to extract saliency maps from a preset video. Statistical features are extracted from the obtained saliency maps from both the perspective of local spatial saliency information in the temporal domain and the perspective of global spatiotemporal saliency information changes in the time series, yielding sIQR and tsIQR features. Statistical analysis methods are used to examine whether there are statistical differences between microstate features and sIQR and tsIQR features, demonstrating the connection between EEG microstates and video saliency. This verifies that decoding based on EEG microstates can be used to study the mechanism of visual attention.
[0071] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for studying visual attention mechanism based on electroencephalogram microstate, characterized in that, Includes the following steps: EEG data acquisition: EEG signals were collected from several subjects while they watched a preset video. The collected EEG signals were preprocessed to obtain preprocessed EEG signals. Video saliency analysis: A computational visual attention model is used to extract saliency maps from a preset video. Statistical features are extracted from the saliency maps from the perspectives of local spatial saliency information in the temporal domain and global spatiotemporal saliency information changes in the time series, to obtain sIQR features and tsIQR features. EEG microstate analysis: The preprocessed EEG signal is transformed into an EEG topology sequence. Spatial clustering analysis is performed on the transformed EEG topology sequence to extract EEG microstate templates. The number of EEG microstate templates is determined. The final microstate template is selected based on the number of EEG microstate templates. The selected final microstate template is backfitted to the preprocessed EEG signal to obtain the microstate sequence. Visual attention decoding: Microstate features and depth features are extracted from the microstate sequence. Statistical analysis methods are used to test whether there are statistical differences between microstate features and sIQR and tsIQR features. Decoding models based on microstate features or depth features are constructed. The constructed decoding models are used to perform segment decoding and video decoding, respectively, to obtain corresponding segment labels and video labels. Among them, microstate features include the time proportion, duration, frequency of occurrence and transition probability of each final microstate template. 2.The method for researching visual attention mechanism based on electroencephalography microstate according to claim 1, characterized in that, The preprocessing of the acquired EEG signals in the EEG data acquisition step specifically includes: The acquired EEG signals were subjected to operations including eye-tracking channel removal, common-mean rereference, filtering, independent component analysis, artifact removal, and bad channel interpolation. 3.The method of claim 1, wherein, The steps of the video saliency analysis include: Saliency map extraction: Saliency maps are extracted frame by frame from the preset video using a computational vision attention model to obtain frame saliency maps and frame saliency map sequences. Saliency maps are then extracted from the frame saliency map sequences according to a preset duration to obtain segment saliency maps and segment saliency map sequences. Saliency feature extraction: Extract sIQR features based on the obtained segment saliency map, sort and integrate the obtained sIQR features according to the preset video duration to obtain sIQR sequence, and extract tsIQR features from the obtained sIQR sequence. 4.The method for researching visual attention mechanism based on electroencephalography microstate according to claim 3, characterized in that, The specific steps for extracting the saliency map are as follows: Each preset video is input frame by frame into a computational vision attention model to obtain a saliency map corresponding to each frame as a frame saliency map. The corresponding frame saliency maps are sorted according to the time order of the preset videos to obtain a frame saliency map sequence. The obtained frame saliency map sequence is segmented according to a preset duration. The average value of the frame saliency maps within each preset duration is calculated to obtain a segment saliency map within each preset duration. The corresponding segment saliency maps are sorted according to the time order of the preset videos to obtain a segment saliency map sequence.
5. The method for studying visual attention mechanisms based on EEG microstates according to claim 3, characterized in that, The steps for extracting salient features include: Calculate the interquartile range of all pixels in the fragment saliency map to obtain sIQR features; The obtained sIQR features are sorted and integrated according to the preset video duration to obtain the sIQR sequence. The interquartile range of the sIQR sequence is calculated to extract the tsIQR features.
6. The method for studying visual attention mechanisms based on EEG microstates according to claim 1, characterized in that, The step of using statistical analysis methods to examine whether there are statistical differences between all extracted microstate features and sIQR features, and between all microstate features and tsIQR features, in the visual attention decoding process specifically includes: The median was used to perform binary classification on sIQR features and tsIQR features to obtain the classification results of sIQR features and tsIQR features. Based on the time points corresponding to the sIQR feature classification results and the tsIQR feature classification results, the extracted microstate features are binary classified to obtain the microstate feature classification results. Based on the obtained microstate feature classification results, sIQR feature classification results, and tsIQR feature classification results, statistical analysis methods were used to test whether there were statistical differences between the segment-based microstate features and the video-based microstate features and the sIQR and tsIQR features.
7. The method for studying visual attention mechanisms based on EEG microstates according to claim 1, characterized in that, The visual attention decoding step includes: Microstate feature extraction: Extract segment-based microstate features and video-based microstate features from the microstate sequence, respectively; Deep feature extraction: Deep learning LSTM-AE method is used to extract segment-based deep features and video-based deep features from the micro-state sequence respectively; Difference test: Statistical analysis methods were used to test whether there were statistical differences between all extracted microstate features and sIQR features, and between all microstate features and tsIQR features; Prediction: Construct a decoding model based on micro-state features, and use the constructed decoding model based on micro-state features to perform segment decoding and video decoding respectively on segment-based micro-state features and video-based micro-state features to obtain the corresponding segment labels and video labels; or, construct a decoding model based on depth features, and use the constructed decoding model based on depth features to perform segment decoding and video decoding respectively on segment-based depth features and video-based depth features to obtain the corresponding segment labels and video labels.
8. The method for studying visual attention mechanisms based on EEG microstates according to claim 7, characterized in that, The prediction steps include: The micro-state features based on segments and the micro-state features based on videos are input into a classifier and subjected to a 10-fold cross-validation binary classification test to predict video saliency and obtain the corresponding segment labels and video labels; or, the deep features based on segments and the deep features based on videos are input into a classifier and subjected to a 10-fold cross-validation binary classification test to predict video saliency and obtain the corresponding segment labels and video labels.
9. The method for studying visual attention mechanisms based on EEG microstates according to claim 7, characterized in that, The prediction step is followed by: Decoding verification: The obtained segment tags and video tags are compared with the real segment tags and real video tags respectively, and the decoding is verified based on the comparison results.