Cognitive load classification method and system based on eye movement tracking
By combining the cognitive load classification model of bidirectional spatiotemporal convolution network and self-attention mechanism structure, the problem of insufficient feature recognition ability in the existing technology is solved, and more accurate and efficient cognitive load classification results are achieved.
Patent Information
- Application Number
- CN202510233152.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The deep learning model used in the prior art for cognitive load classification has poor feature recognition capabilities, resulting in inaccurate results of cognitive load classification.
A cognitive load classification model combining bidirectional spatiotemporal convolutional network and self-attention mechanism structure is adopted, and bidirectional eye movement time series data are used to capture the spatiotemporal characteristics in eye movement data, and key characteristics are highlighted through self-attention mechanism.
The comprehensiveness of feature extraction and the model's ability to express input data features is improved, and the accuracy and efficiency of cognitive load classification results are significantly improved.
Smart Images

Figure CN120145146A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly relates to a cognitive load classification method and system based on eye tracking. Background Art
[0002] Eye tracking is a psychophysiological research technique used to measure the position of the human eye and eye movements, so as to understand an individual's visual attention, cognitive processes, and mental state. This technique can record the eye activities of an individual when observing a certain visual stimulus, including information such as the position of eye fixation, the duration of fixation, the moving speed and path of the eyeball, etc. Eye tracking technology has a wide range of applications in multiple fields, including psychological research, human-computer interaction design, virtual reality, etc. By analyzing eye movement data, researchers can obtain important information such as the attention distribution and cognitive load of an individual when performing visual tasks.
[0003] Among them, cognitive load refers to the amount of effort or resources required by the brain when processing information. When we are performing a certain task, the brain needs to mobilize different resources to understand, analyze, and store information. If the task is too complex or the information is too much, the processing ability of the brain will be overloaded, resulting in cognitive fatigue or low efficiency. And eye tracking technology can help researchers analyze the eye movement patterns, thereby indirectly inferring the cognitive load state of the brain. Cognitive load classification based on eye tracking can provide real-time and non-invasive cognitive state assessment, helping to accurately monitor the attention and load levels of an individual in complex tasks.
[0004] However, the deep learning models currently used to achieve cognitive load classification have poor feature recognition ability, and the feature recognition is not comprehensive and accurate enough, resulting in inaccurate cognitive load classification results finally output. Summary of the Invention
[0005] In view of this, the present invention provides a cognitive load classification method and system based on eye tracking to solve the problem that the feature recognition in the prior art is not comprehensive enough and there are inaccurate cognitive load classification results.
[0006] In the first aspect, the present invention provides a cognitive load classification method based on eye tracking, and the method includes:
[0007] Obtain the original eye movement data to be analyzed, and the original eye movement data to be analyzed includes fixation data, fixation movement data, pupil data, and pupil change data;
[0008] Perform data preprocessing on the original eye movement data to be analyzed to obtain eye movement feature data;
[0009] Input the eye movement feature data into a pre-constructed cognitive load classification model to obtain the cognitive load classification result. The cognitive load classification model includes a bidirectional spatio-temporal convolutional network structure, and the bidirectional spatio-temporal convolutional network structure includes a self-attention mechanism structure. The bidirectional spatio-temporal convolutional network structure is used to extract spatio-temporal features from both the forward time series and the reverse time series simultaneously, and the self-attention mechanism structure is used to highlight key features based on the extracted spatio-temporal features.
[0010] The cognitive load classification method proposed by the present invention adopts a cognitive load classification model that combines a bidirectional spatio-temporal convolutional network and a self-attention mechanism structure. Using the forward and reverse bidirectional eye movement time series data, it allows the eye movement information to flow bidirectionally, can effectively capture the spatio-temporal features in the eye movement data, and improves the comprehensiveness of feature extraction. In addition, the channel self-attention mechanism can adaptively adjust the response intensity of each channel by learning the dependence relationship between different feature channels, enabling the model to highlight important feature channels and suppress unimportant channels, thereby enhancing the model's ability to express the features of the input data, and further effectively improving the accuracy and efficiency of the cognitive load classification result.
[0011] In an alternative embodiment, the bidirectional spatio-temporal convolutional network structure includes:
[0012] At least two convolutional kernels, each convolutional kernel includes a forward convolutional branch for the forward time series and a reverse convolutional branch for the reverse time series. Among them, the forward convolutional branch is used to extract the eye movement feature data of the forward time series, and the reverse convolutional branch is used to extract the eye movement feature data of the reverse time series; each group of convolutional branches includes convolutional blocks with different dilation rates connected in sequence; among them, the first convolutional block includes a convolutional layer and a self-attention mechanism structure connected in sequence; other convolutional blocks include a normalization layer, an activation layer, a convolutional layer, and a self-attention mechanism structure connected in sequence.
[0013] In the bidirectional spatio-temporal convolutional network provided in this embodiment, the features of the shallow layer and the deep layer are fused together, promoting the comprehensive utilization of features at different levels. This multi-scale feature fusion helps to improve the network's ability to recognize complex patterns and effectively realizes the accurate recognition of different cognitive load levels.
[0014] In an alternative embodiment, the cognitive load classification model further includes:
[0015] An eye movement data input layer, which is used to obtain the eye movement feature data and input the eye movement feature data into the forward convolutional branch;
[0016] A time series inversion structure, connected to the eye movement data input layer, which is used to invert the time series of the eye movement feature data to obtain the reverse eye movement feature data; the time series inversion structure is also used to input the reverse eye movement feature data into the reverse convolutional branch;
[0017] Pooling layer, connected to the convolutional kernel, for splicing the features extracted by the convolutional kernel and performing average pooling;
[0018] Embedding layer, connected to the pooling layer, for mapping the data after average pooling to a low-dimensional space;
[0019] Classification layer, connected to the embedding layer, for making classification decisions based on the features output by the embedding layer.
[0020] In this embodiment, an end-to-end cognitive load classification model is designed based on the original eye movement data. Through the cooperation of the multi-layer structure, this model makes full use of the temporal features of the eye movement data and effectively improves the classification accuracy of the cognitive load.
[0021] In an alternative embodiment, the self-attention mechanism structure includes:
[0022] Linear layer, for creating key, query, and value vectors;
[0023] Scaled dot-product attention mechanism layer, for generating attention scores based on the query and key, and using the attention scores as the weights of the corresponding values to obtain the weighted sum value at each position.
[0024] In this embodiment, using the self-attention mechanism structure to construct the cognitive load classification model can automatically focus on the most critical features in the input data, improve the model's attention ability to important information, effectively capture the dependency relationships by dynamically calculating the attention scores, and enhance the accuracy of the model's classification of the cognitive load results.
[0025] In an alternative embodiment, the method further includes:
[0026] Sending the cognitive load classification result predicted by the cognitive load classification model to the client;
[0027] Receiving the feedback result sent by the client and performing real-time optimization of the cognitive load classification model according to the feedback result.
[0028] In this embodiment, based on the received feedback, the model can be fine-tuned or perform online learning. By adjusting the parameters of the model, the classification accuracy can be gradually improved to achieve the real-time optimization of the model.
[0029] In an alternative embodiment, the cognitive load classification model is constructed through the following steps:
[0030] Collect eye movement data, where the eye movement data includes the cognitive information of the subject;
[0031] Clean the eye movement data to obtain the initial eye movement data;
[0032] Extract features from the initial eye movement data to obtain the initial eye movement feature data;
[0033] Normalize the initial eye movement feature data to obtain a set of eye movement data;
[0034] Construct an initial classification learning model architecture;
[0035] Train the initial classification learning model architecture based on the set of eye movement data to obtain a trained cognitive load classification model.
[0036] The model constructed in this embodiment has a certain generalization ability through training and testing on multi-tasks and multi-subject groups, can adapt to different tasks and subjects, and has high accuracy.
[0037] In a second aspect, the present invention provides a cognitive load classification system based on eye movement tracking, the system comprising:
[0038] An acquisition module for acquiring original eye movement data to be analyzed, the original eye movement data to be analyzed including fixation data, fixation movement data, pupil data, and pupil change data;
[0039] A processing module for performing data preprocessing on the original eye movement data to be analyzed to obtain eye movement feature data;
[0040] A prediction module for inputting the eye movement feature data into a pre-constructed cognitive load classification model to obtain a cognitive load classification result, wherein the cognitive load classification model includes a bidirectional spatio-temporal convolutional network structure, the bidirectional spatio-temporal convolutional network structure includes a self-attention mechanism structure, the bidirectional spatio-temporal convolutional network structure is used to extract spatio-temporal features from both the forward time series and the reverse time series, and the self-attention mechanism structure is used to highlight key features based on the extracted spatio-temporal features.
[0041] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the cognitive load classification method based on eye movement tracking according to the first aspect or any corresponding embodiment thereof.
[0042] In a fourth aspect, the present invention provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the cognitive load classification method based on eye movement tracking according to the first aspect or any corresponding embodiment thereof.
[0043] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, and the computer instructions are used to cause a computer to execute the cognitive load classification method based on eye movement tracking according to the first aspect or any corresponding embodiment thereof.
[0044] It should be noted that since the cognitive load classification system based on eye movement tracking provided by the present invention, the computer device, computer-readable storage medium, and computer program product correspond to the above-mentioned cognitive load classification method based on eye movement tracking. Therefore, for the beneficial effects of the cognitive load classification system based on eye movement tracking, computer device, computer-readable storage medium, and computer program product, please refer to the description of the corresponding beneficial effects of the cognitive load classification method based on eye movement tracking above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1 is a flowchart of the cognitive load classification method based on eye movement tracking according to an embodiment of the present invention;
[0047] Figure 2 is a schematic diagram of the architecture of the cognitive load classification model according to an embodiment of the present invention;
[0048] Figure 3 is a schematic diagram of the dilated convolution with a convolution kernel size of 3 according to an embodiment of the present invention;
[0049] Figure 4 is a schematic diagram of the convolution block structure according to an embodiment of the present invention;
[0050] Figure 5 is a schematic diagram of the attention mechanism structure according to an embodiment of the present invention;
[0051] Figure 6 is a flowchart of the training process of the cognitive load classification model according to an embodiment of the present invention;
[0052] Figure 7 is a block diagram of the structure of the cognitive load classification system based on eye movement tracking according to an embodiment of the present invention;
[0053] Figure 8 is a schematic diagram of the hardware structure of the computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] According to an embodiment of the present invention, an embodiment of a cognitive load classification method based on eye tracking is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0056] In this embodiment, a cognitive load classification method based on eye tracking is provided, which can be executed by devices such as servers, terminals, and mobile terminals. Figure 1 It is a flowchart of the cognitive load classification method based on eye tracking according to an embodiment of the present invention, as Figure 1 shown, and the process includes the following steps:
[0057] Step S101, obtain the original eye movement data to be analyzed. The original eye movement data to be analyzed includes fixation data, fixation movement data, pupil data, and pupil change data.
[0058] The original eye movement data to be analyzed in this embodiment can be collected by an eye tracking device. The eye tracking device includes one or more sensors, such as an infrared camera, an optical sensor, etc., which are used to capture the reflected light of the eyes and the eye movement, etc., to obtain eye movement data. The eye movement data can reflect an individual's cognitive load. Among them, the fixation data records the information that an individual focuses on a specific position at a certain time point, including the coordinate position (x, y) of the fixation point, the fixation duration (unit: ms), etc. The fixation movement data records the process of the eyes moving from one fixation point to another, including the movement trajectory, movement speed, movement acceleration, etc. of the eyes. The pupil data records the diameter of the pupil (unit: mm) and the size change rate, etc. The pupil change data records the speed of the pupil changing over time, etc.
[0059] In this embodiment, the speed of pupil change can be calculated using the formula: t represents time, and x represents position information.
[0060] In this embodiment, the raw eye movement data to be analyzed is stored in a time series form, with a sampling frequency of 100 Hz and a dimension of C×T (C = 8 channels, T = time step). The format of the raw eye movement data to be analyzed can be as follows:
[0061] [Timestamp, fixation coordinate_x, fixation coordinate_y, fixation velocity_x, fixation velocity_y, left pupil diameter, right pupil diameter, left pupil change velocity, right pupil change velocity].
[0062] Step S102: Perform data preprocessing on the raw eye movement data to be analyzed to obtain eye movement feature data.
[0063] Continuous eye movement data can be segmented into 5-second segments (window length T = 500) using a sliding window, with a step size of 1 second, to form a data set.
[0064] Remove noise and outliers in the eye movement data, such as incorrect data points caused by equipment failures or subject blinks; identify and remove invalid data segments, such as cases of long-term eye closure or data signal loss. Then fill in the missing values, and for missing data, cubic spline interpolation can be used for completion.
[0065] Furthermore, perform normalization processing on the data, including: spatial coordinate normalization and physiological signal normalization. Spatial coordinate normalization is to map the fixation point coordinates (x, y) to the screen resolution range (interval [0, 1]); physiological signal normalization is to perform z-score standardization on the pupil diameter, fixation data, etc. according to the individual baseline value.
[0066] Step S103: Input the eye movement feature data into a pre-constructed cognitive load classification model to obtain a cognitive load classification result. Among them, the cognitive load classification model includes a bidirectional spatio-temporal convolutional network structure, and the bidirectional spatio-temporal convolutional network structure includes a self-attention mechanism structure. The bidirectional spatio-temporal convolutional network structure is used to extract spatio-temporal features from both the forward time series and the reverse time series simultaneously, and the self-attention mechanism structure is used to highlight key features based on the extracted spatio-temporal features.
[0067] The cognitive load classification model is obtained through previous model training. Specifically, a data set can be collected and preprocessed, and the preprocessed data set is divided into a training set and a test set, and 10-fold cross-validation is adopted. Then model construction is carried out. In this embodiment, a model architecture based on a bidirectional spatio-temporal convolutional network and an attention mechanism is adopted, and the model can be constructed using the deep learning framework PyTorch. The input layer of the model receives the preprocessed eye movement data features, extracts spatio-temporal features through the bidirectional spatio-temporal convolutional network, then highlights key features through the attention mechanism, and finally outputs the cognitive load classification result through the fully connected layer. The cognitive load classification results include: mild cognitive load, moderate cognitive load, and severe cognitive load, etc.
[0068] In the related art, some cognitive load classification methods perform well on specific tasks or specific subject groups, but their performance drops significantly when generalized to new tasks or different subjects. Moreover, there is insufficient integration of spatio-temporal features. Eye movement data generally has rich spatio-temporal features, but related technologies often have difficulty effectively integrating these features, resulting in low recognition accuracy of cognitive load.
[0069] The cognitive load classification method proposed by the present invention adopts a cognitive load classification model that combines a bidirectional spatio-temporal convolutional network and a self-attention mechanism structure. By using the forward and backward bidirectional eye movement time series data, it allows the eye movement information to flow bidirectionally, and can effectively capture the spatio-temporal features in the eye movement data, improving the comprehensiveness of feature extraction. In addition, the channel self-attention mechanism can adaptively adjust the response intensity of each channel by learning the dependence relationship between different feature channels, enabling the model to highlight important feature channels and suppress unimportant channels, thereby enhancing the model's ability to express the features of the input data, and further effectively improving the accuracy and efficiency of the cognitive load classification results.
[0070] In addition, the cognitive load classification method provided by the present invention is a more accurate, real-time and non-invasive cognitive load measurement method, which has significant performance advantages in processing complex eye movement data, and can effectively reduce the interference to the subject and improve the user experience by using non-invasive eye tracking technology.
[0071] In some alternative embodiments, the bidirectional spatio-temporal convolutional network structure includes:
[0072] At least two convolutional kernels, each convolutional kernel including a forward convolution branch with a forward time series and a backward convolution branch with a backward time series. Among them, the forward convolution branch is used to extract eye movement feature data of the forward time series, and the backward convolution branch is used to extract eye movement feature data of the backward time series; each convolution branch includes convolution blocks with different dilation rates connected in sequence; among them, the first convolution block includes a convolution layer and a self-attention mechanism structure connected in sequence; other convolution blocks include a normalization layer, an activation layer, a convolution layer and a self-attention mechanism structure connected in sequence.
[0073] In this embodiment, referring to Figure 2As shown, two convolutional kernels with sizes of 3 and 5 are respectively set, enabling the network to capture features at multiple time resolutions. Before global pooling, the features extracted by the two convolutional kernels are concatenated. Each convolutional kernel consists of serialized convolutional branches, and each convolutional branch contains multiple convolutional blocks with different dilation rates. Each convolutional block includes a convolutional layer and a self-attention mechanism structure, etc. Each convolutional layer performs one-dimensional convolution on the time scale, followed by normalization, activation function, and self-attention mechanism processing. The output of each convolutional layer is connected to all subsequent layers. The processed feature map is fed into the self-attention mechanism and then concatenated with the input of the previous layer and output to the next layer.
[0074] Due to the causal characteristics of dilated convolution, the present invention designs a bidirectional time series, reverses the eye movement data in time sequence, and adds non-causal factors to the network. The bidirectional spatio-temporal convolutional network structure processes both the forward and backward sequences simultaneously to understand the bidirectional flow of eye movement information. That is, the eye movement feature data of the original forward time sequence is copied into two copies: the forward sequence (t = 0 - T) and the backward sequence (t = T - 0). The eye movement feature data of the forward time sequence and the eye movement feature data of the backward time sequence are respectively input into two parallel spatio-temporal convolutional branches, and the output features are concatenated in the channel dimension, and finally the features processed forward and backward are concatenated.
[0075] In addition, the bidirectional spatio-temporal convolutional network structure in this embodiment uses gradually increasing dilation rates (0, 1, 2, 4, 8, 16, 32) for convolution in consecutive convolutional block instances. This gradually expands the receptive field, enabling the network to capture a wider range of context information without significantly increasing the computational cost. Refer to Figure 3 As shown, it is a schematic diagram of dilated convolution with a convolutional kernel size of 3.
[0076] Refer to Figure 4 As shown, in all convolutional blocks after the first convolutional block, a "pre-activation" architecture is adopted, that is, batch normalization (BN) and PReLU activation function are performed before feature extraction to introduce necessary non-linear characteristics. Through "pre-activation", the model adjusts the data distribution and activates features before processing the data.
[0077] In the bidirectional spatio-temporal convolutional network provided in this embodiment, the features of the shallow layer and the deep layer are fused together, promoting the comprehensive utilization of features at different levels. This multi-scale feature fusion helps to improve the network's ability to recognize complex patterns and effectively realizes the accurate recognition of different cognitive load levels.
[0078] In some alternative embodiments, the cognitive load classification model further includes:
[0079] An eye movement data input layer for obtaining eye movement feature data and inputting the eye movement feature data into the forward convolutional branch.
[0080] A temporal inversion structure, connected to the eye movement data input layer, is used to invert the time sequence of the eye movement feature data to obtain reverse eye movement feature data; the temporal inversion structure is also used to input the reverse eye movement feature data into the reverse convolution branch.
[0081] A pooling layer, connected to the convolution kernel, is used to splice the features extracted by the convolution kernel and perform average pooling.
[0082] An embedding layer, connected to the pooling layer, is used to map the data after average pooling to a low-dimensional space.
[0083] A classification layer, connected to the embedding layer, is used to make classification decisions based on the features output by the embedding layer.
[0084] In this embodiment, the global average pooling layer and the classification layer also both adopt the "pre-activation" architecture.
[0085] First, the eye movement data input layer obtains the eye movement feature data and inputs it into the forward convolution branch for feature extraction; then, the temporal inversion structure inverts the time sequence of the eye movement data to generate reverse eye movement feature data and inputs it into the reverse convolution branch; subsequently, the pooling layer splices the features extracted by the forward and reverse convolution kernels and performs average pooling processing; then, the embedding layer maps the pooled data to a low-dimensional space; finally, the classification layer makes classification decisions based on the output of the embedding layer. The present invention designs an end-to-end cognitive load classification model based on the original eye movement data. Through the cooperation of multiple layers of structures, this model makes full use of the temporal features of the eye movement data and effectively improves the classification accuracy of cognitive load.
[0086] In some optional embodiments, the self-attention mechanism structure includes:
[0087] A linear layer, used to create key, query, and value vectors;
[0088] A scaled dot-product attention mechanism layer, used to generate attention scores based on the query and key and use the attention scores as weights for the corresponding values to obtain the weighted sum value at each position.
[0089] In each convolution block, there is a self-attention mechanism structure. The self-attention mechanism structure uses a linear layer (Linear) to create key, query, and value vectors. The channel self-attention mechanism enables the features extracted by the network on the time scale to be enhanced after calculation. Refer to Figure 5 As shown, each self-attention mechanism structure consists of three main components: query Q, key K, and value V. The attention scores generated by the interaction between query Q and key K are used as weights for the corresponding value V. Q, K, and V are obtained by linear transformation of the input.
[0090] Specifically, the query vector Q calculates the dot product with all key vectors K to obtain the attention scores for each pair of query and key. To avoid overly large dot products, the results are scaled and then converted into a probability distribution through Softmax, as shown in the formula: The attention weights for each position are obtained. These attention weights are applied to the value vector V, and the values at each position are weighted and summed to obtain the final output.
[0091] In this embodiment, a self-attention mechanism structure is used to construct the cognitive load classification model, which can automatically focus on the most critical features in the input data and improve the model's ability to pay attention to important information. By dynamically calculating attention scores, dependencies are effectively captured, enhancing the accuracy of the model in classifying cognitive load results.
[0092] In some alternative embodiments, the method further includes:
[0093] Sending the cognitive load classification result predicted by the cognitive load classification model to the client;
[0094] Receiving the feedback result sent by the client and performing real-time optimization of the cognitive load classification model according to the feedback result.
[0095] In this embodiment, the classification result of the cognitive load (e.g., low, medium, high load levels) generated after being processed by the model is transmitted to the client through the network. The client sends feedback data (e.g., the cognitive load result self-evaluated by the user, eye movement data, etc.) to the model according to the actual situation or the feedback classification result input by the user. Based on the received feedback, the model can be fine-tuned or perform online learning. By adjusting the parameters of the model, the classification accuracy is gradually improved to achieve real-time optimization of the model.
[0096] In some alternative embodiments, the cognitive load classification model is constructed through the following steps:
[0097] Step a1, collect eye movement data, which includes the cognitive information of the subject and data such as fixation data, fixation movement data, pupil data, and pupil change data. The cognitive information is the corresponding individual's cognitive load.
[0098] Step a2, perform data cleaning on the eye movement data to obtain the initial eye movement data. Among them, the dimension of the processed eye movement data can be [Batch_size, C = 8, T = 500].
[0099] Step a3, perform feature extraction on the initial eye movement data to obtain the initial eye movement feature data.
[0100] Step a4, perform normalization processing on the initial eye movement feature data to obtain the eye movement data set.
[0101] The initial eye movement feature data can be represented by X, where X ∈ R C×T , where C is the number of signal channels and T is the length of the signal sequence. Normalization can be expressed by the following formula:
[0102] where X is the initial eye movement feature data, is the normalized eye movement data, and μ and σ 2 represent the signal mean and variance respectively.
[0103] Normalizing the extracted features and scaling the data to a specific range (such as [0, 1] or [-1, 1]) can eliminate the dimensional differences between different features and improve the training effect and convergence speed of the model.
[0104] Step a5, construct the initial classification learning model architecture. The initial classification learning model architecture can refer to the above-mentioned cognitive load classification model architecture.
[0105] Step a6, train the initial classification learning model architecture based on the eye movement data set to obtain the cognitive load classification model after training is completed.
[0106] During training, an appropriate loss function can be selected to measure the difference between the predicted results of the model and the true labels. For the cognitive load classification task, cross-entropy loss can be used as the loss function, which can effectively measure the difference between two probability distributions and alleviate the problem of class imbalance. Further, an optimizer is selected to update the parameters of the model to minimize the loss function. Optimizers such as stochastic gradient descent (SGD) and Adam can be used. In this embodiment, the Adam optimizer is preferably used, with an initial learning rate of 3e-4, in conjunction with a cosine annealing scheduler. The Adam optimizer combines the advantages of SGD, can adaptively adjust the learning rate, and has a faster convergence speed and better stability. Specifically, the training set data is input into the model, the loss function value is calculated, and the parameters of the model are updated through the backpropagation algorithm. During the training process, the validation set data is used to evaluate the performance of the model, and the hyperparameters of the model, such as the learning rate and regularization parameter, are adjusted according to the loss value on the validation set. When the loss value on the validation set no longer decreases significantly or overfitting occurs, stop training to obtain the cognitive load classification model after training is completed.
[0107] The model constructed in this embodiment has a certain generalization ability through training and testing on multi-tasks and multi-subject groups and can adapt to different tasks and subjects.
[0108] In the model provided by the present invention, a bidirectional spatio-temporal convolutional network is used to extract temporal correlations, and a self-attention mechanism is introduced to calculate the correlation of eye movement channels after extracting temporal correlations, thereby highlighting important features and improving the effect of the model. This model is a brand-new model. Compared with ordinary deep learning networks, the accuracy and generalization ability of this model have been improved. Specifically, the present invention has been experimentally verified on the CL-Drive public dataset and the self-collected dataset. Among them, the three-classification and two-classification of the CL-Drive public dataset reach 63% and 71% respectively. On the self-collected dataset, the three-classification reaches 82%, which is relatively high in accuracy compared with other algorithms.
[0109] Referring to Figure 6 As shown, in this embodiment, after the original eye movement signal is input into the system, it is first simply preprocessed, and then sent to the cognitive load classification model, that is, a deep learning model for training. The cognitive load classification model consists of two bidirectional spatio-temporal convolutional networks. The preprocessed signal first enters the convolutional network for embedding encoding, and then the attention mechanism module highlights the key features and extracts features from the time dimension. After the network is trained, the eye movement signal is input into the network to output the final cognitive load classification result. The cognitive load classification model combining the bidirectional spatio-temporal convolutional network and the self-attention mechanism structure effectively improves the accuracy and efficiency of the cognitive load classification result.
[0110] In this embodiment, a cognitive load classification system based on eye movement tracking is also provided. This system is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0111] This embodiment provides a cognitive load classification system based on eye movement tracking, as Figure 7 shown, the system includes:
[0112] An acquisition module 201, configured to acquire original eye movement data to be analyzed, and the original eye movement data to be analyzed includes fixation data, fixation movement data, pupil data, and pupil change data;
[0113] A processing module 202, configured to perform data preprocessing on the original eye movement data to be analyzed to obtain eye movement feature data;
[0114] A prediction module 203, configured to input the eye movement feature data into a pre-constructed cognitive load classification model to obtain a cognitive load classification result. The cognitive load classification model includes a bidirectional spatio-temporal convolutional network structure, and the bidirectional spatio-temporal convolutional network structure includes a self-attention mechanism structure. The bidirectional spatio-temporal convolutional network structure is used to extract spatio-temporal features from both the forward time series and the backward time series simultaneously, and the self-attention mechanism structure is used to highlight key features based on the extracted spatio-temporal features. Among them, the bidirectional spatio-temporal convolutional network structure includes: at least two convolutional kernels, each convolutional kernel includes a forward convolution branch for the forward time series and a backward convolution branch for the backward time series. The forward convolution branch is used to extract the eye movement feature data of the forward time series, and the backward convolution branch is used to extract the eye movement feature data of the backward time series; each group of convolution branches includes convolution blocks with different dilation rates connected in sequence; among them, the first convolution block includes a convolution layer and a self-attention mechanism structure connected in sequence; other convolution blocks include a normalization layer, an activation layer, a convolution layer, and a self-attention mechanism structure connected in sequence. The cognitive load classification model further includes: an eye movement data input layer, configured to obtain the eye movement feature data and input the eye movement feature data into the forward convolution branch; a time series inversion structure, connected to the eye movement data input layer, configured to invert the time series of the eye movement feature data to obtain backward eye movement feature data; the time series inversion structure is further configured to input the backward eye movement feature data into the backward convolution branch; a pooling layer, connected to the convolutional kernel, configured to splice the features extracted by the convolutional kernel and perform average pooling; an embedding layer, connected to the pooling layer, configured to map the data after average pooling to a low-dimensional space; a classification layer, connected to the embedding layer, configured to make a classification decision based on the features output by the embedding layer. A linear layer is configured to create key, query, and value vectors; a scaled dot-product attention mechanism layer is configured to generate attention scores based on the query and the key, and use the attention scores as the weights of the corresponding values to obtain a weighted sum value for each position.
[0115] In some alternative embodiments, the system further includes:
[0116] A feedback module, configured to send the cognitive load classification result predicted by the cognitive load classification model to the client; receive the feedback result sent by the client, and perform real-time optimization on the cognitive load classification model according to the feedback result.
[0117] A construction module, configured to collect eye movement data, where the eye movement data includes the cognitive information of the subject; perform data cleaning on the eye movement data to obtain initial eye movement data; perform feature extraction on the initial eye movement data to obtain initial eye movement feature data; perform normalization processing on the initial eye movement feature data to obtain an eye movement data set; construct an initial classification learning model architecture; and train the initial classification learning model architecture based on the eye movement data set to obtain a trained cognitive load classification model.
[0118] The cognitive load classification system based on eye movement tracking in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0119] The further functional descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be repeated here.
[0120] The embodiment of the present invention also provides a computer device having the above-mentioned Figure 7 cognitive load classification system based on eye movement tracking as shown.
[0121] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As shown in Figure 8 , the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 8 In
[0122] , a single processor 10 is taken as an example.
[0123] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0124] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0125] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.
[0126] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0127] Embodiments of the present invention also provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0128] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways for a computer to execute computer program instructions include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0129] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A cognitive load classification method based on eye tracking, characterized in that: The method comprises: Acquiring raw eye movement data to be analyzed, wherein the raw eye movement data to be analyzed includes gaze data, gaze movement data, pupil data, and pupil change data; Preprocessing the raw eye movement data to be analyzed to obtain eye movement feature data; The eye movement feature data is input into a pre-built cognitive load classification model to obtain a cognitive load classification result, wherein the cognitive load classification model includes a bidirectional spatiotemporal convolutional network structure, the bidirectional spatiotemporal convolutional network structure includes a self-attention mechanism structure, the bidirectional spatiotemporal convolutional network structure is used to simultaneously extract spatiotemporal features from forward time sequence and reverse time sequence, and the self-attention mechanism structure is used to highlight key features based on the extracted spatiotemporal features.
2. The method according to claim 1, characterized in that The bidirectional spatiotemporal convolutional network structure includes: At least two convolution kernels, each convolution kernel includes a group of forward convolution branches in forward time sequence and a group of reverse convolution branches in reverse time sequence, wherein the forward convolution branches are used to extract eye movement feature data in forward time sequence, and the reverse convolution branches are used to extract eye movement feature data in reverse time sequence; each group of convolution branches includes convolution blocks with different expansion rates connected in sequence; wherein the first convolution block includes convolution layers connected in sequence and the self-attention mechanism structure; other convolution blocks include normalization layers, activation layers, convolution layers and the self-attention mechanism structure connected in sequence.
3. The method according to claim 2, characterized in that The cognitive load classification model also includes: An eye movement data input layer, used for acquiring the eye movement feature data and inputting the eye movement feature data into the forward convolution branch; A timing reversal structure, connected to the eye movement data input layer, for reversing the timing of the eye movement feature data to obtain reverse eye movement feature data; the timing reversal structure is also used to input the reverse eye movement feature data into the reverse convolution branch; A pooling layer, connected to the convolution kernel, for concatenating the features extracted by the convolution kernel and performing average pooling; An embedding layer, connected to the pooling layer, for mapping the average pooled data to a low-dimensional space; The classification layer is connected to the embedding layer and is used to make classification decisions based on the features output by the embedding layer.
4. The method according to claim 1, characterized in that: The self-attention mechanism structure includes: Linear layers to create key, query, and value vectors; A scaled dot product attention mechanism layer is used to generate attention scores based on the query and key, and use the attention scores as weights for the corresponding values to obtain the weighted sum value for each position.
5. The method according to claim 1, characterized in that The method further comprises: Sending the cognitive load classification result predicted by the cognitive load classification model to the client; Receive feedback results sent by the client, and optimize the cognitive load classification model in real time according to the feedback results.
6. The method according to claim 1, characterized in that The cognitive load classification model is constructed through the following steps: collecting eye movement data, wherein the eye movement data includes cognitive information of the subject; Performing data cleaning on the eye movement data to obtain initial eye movement data; Performing feature extraction on the initial eye movement data to obtain initial eye movement feature data; Normalizing the initial eye movement feature data to obtain an eye movement data set; Build the initial classification learning model architecture; The initial classification learning model architecture is trained based on the eye movement data set to obtain the cognitive load classification model after training.
7. A cognitive load classification system based on eye tracking, characterized in that: The system comprises: An acquisition module, used for acquiring raw eye movement data to be analyzed, wherein the raw eye movement data to be analyzed includes gaze data, gaze movement data, pupil data, and pupil change data; A processing module, used for performing data preprocessing on the raw eye movement data to be analyzed to obtain eye movement feature data; A prediction module is used to input the eye movement feature data into a pre-built cognitive load classification model to obtain a cognitive load classification result, wherein the cognitive load classification model includes a bidirectional spatiotemporal convolutional network structure, the bidirectional spatiotemporal convolutional network structure includes a self-attention mechanism structure, the bidirectional spatiotemporal convolutional network structure is used to extract spatiotemporal features from forward time sequence and reverse time sequence at the same time, and the self-attention mechanism structure is used to highlight key features based on the extracted spatiotemporal features.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the eye tracking-based cognitive load classification method according to any one of claims 1 to 6 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the cognitive load classification method based on eye tracking according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the cognitive load classification method based on eye tracking according to any one of claims 1 to 6.