A multimodal reading eye movement representation method based on graph neural network
Through the multimodal reading eye movement representation method based on graph neural network, the problem of difficulty in integrating multimodal data in the prior art is solved, efficient representation and analysis of eye movement data is realized, and a more comprehensive perspective is provided to understand eye movement behavior and cognitive processes.
Patent Information
- Application Number
- CN202411507526.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The prior art is difficult to effectively integrate multimodal data in eye movement data analysis, especially ignoring stimulus information related to eye movement tracking characteristics, making it difficult to provide a more comprehensive perspective to understand eye movement behavior and cognitive processes.
The multimodal reading eye movement representation method based on graph neural network is adopted. By acquiring and preprocessing eye movement data, iteratively and multi-dimensional attention processing is performed. Combining the multi-dimensional modeling method of graph attention network, node features and multi-dimensional edge features are fused to generate the final reading eye movement representation output.
It realizes the full utilization of eye movement tracking data of complex structures, accurately captures and analyzes subtle changes in eye movement data, integrates eye movement data and text information, provides a more comprehensive perspective to understand eye movement behavior and cognitive processes, and improves the universality and accuracy of the representation results.
Smart Images

Figure CN119357900B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent learning technology, and in particular relates to a multimodal reading eye movement representation method based on graph neural network. Background Art
[0002] The movement of the eyes during reading is affected by many factors, which can be captured by the machine and calculated into a variety of indicators. The influencing factors mainly include the individual attributes of the reader and the attributes of the reading material. The individual attributes of the reader include age, cognitive background, etc. The attributes of the reading material include vocabulary difficulty, meaning difficulty, etc. The instrument can collect the focus movement trajectory of the subject when reading the material. The trajectory data is a sequence of tuples (i.e., x, y coordinates). According to the research on reading eye movements, the movement of the human eye includes two basic movement phenomena, namely saccades (movement of the eyes) and fixations (relative stillness of the eyes). These two basic movements can be used to calculate a series of indicators through the reading trajectory. The saccade-related indicators include saccade direction, saccade angle, saccade amplitude, saccade average velocity, blink (sac contains blink), and saccade duration. The fixation-related indicators include fixation duration, pupil diameter, and interest area dwell time.
[0003] Eye movement data plays a vital role in the human cognitive process. It not only reflects the dynamic process of visual information processing, but also is the key to understanding fields such as human-computer interaction, mental health assessment, and visual impairment diagnosis. The representation and analysis of eye movement data has irreplaceable value in revealing the human visual cognitive mechanism. In the field of human-computer interaction, eye movement data can help design more natural interactive interfaces and improve user experience, such as virtual reality and online education; in mental health assessment, changes in eye movement patterns can be used as an auxiliary tool for diagnosing psychological states, such as depression diagnosis and Alzheimer's diagnosis; and in visual impairment diagnosis, the analysis of eye movement data helps to identify and assess the degree of visual impairment, such as dyslexia classification.
[0004] In the prior art, the EZ reader model predicts the reader's eye movement by simulating the perception, cognition and movement process during reading. Basic machine learning algorithms such as SVM (support vector machine), Naive Bayes, Logistic regression, Neural network, K-NN (K-Nearest Neighbor) and LDA (Linear discriminant analysis) are used in the diagnosis of depression. With the development of deep learning technology, Nora et al. used the transformer pre-training model to predict eye movement features, and used BERT and XLM to predict single features respectively; David et al. used a bidirectional long short-term memory network and fixation time and pupil size to predict the degree of human understanding when reading, ignoring the importance of stimulus information.
[0005] However, in the existing field of eye movement data analysis, traditional methods mainly rely on statistical models and basic machine learning algorithms. Eye movement data contains rich dynamic changes and complex patterns. Traditional methods are limited by the high dimensionality and complexity of eye movement data and find it difficult to capture and analyze subtle changes in eye movement data. These methods usually rely on a single eye tracking feature, such as saccades, fixation points, and pupil diameter, and do not fully utilize all the rich eye movement features. In addition, these methods often ignore stimulus information related to eye tracking features, such as text stimulus information in reading tasks. The integration of eye tracking data and visual stimulus information is crucial for a comprehensive understanding of eye movement behavior and cognitive processes. Existing methods lack effective strategies on how to effectively integrate these multimodal data, making it difficult to provide a more comprehensive perspective to understand eye movement behavior and cognitive processes. Eye tracking data provides rich information, but differences in individual behavioral habits challenge the universality and accuracy of the analysis results.
[0006] Therefore, the present invention proposes a multimodal reading eye movement representation method based on graph neural network. Summary of the invention
[0007] In order to solve the above technical problems, the present invention proposes a multimodal reading eye movement representation method based on graph neural network to solve the problems existing in the above-mentioned prior art.
[0008] To achieve the above objectives, the present invention provides a multimodal reading eye movement representation method based on graph neural network, comprising:
[0009] Acquiring eye movement data of the subject, and preprocessing the eye movement data of the subject to obtain preprocessed eye movement data;
[0010] Converting the pre-processed eye movement data into a topological structure graph;
[0011] Iterating and multi-dimensional attention processing the topological structure graph to obtain node features and multi-dimensional edge features;
[0012] The multi-dimensional modeling method based on the graph attention network interactively fuses the node features and the multi-dimensional edge features to obtain the final reading eye movement representation output.
[0013] Optionally, the process of preprocessing the subject's eye movement data includes:
[0014] Dividing the eye movement data of the subject according to the subject ID to obtain a plurality of eye movement data of the same subject;
[0015] The eye movement data of the same subject are normalized one by one according to the eye movement characteristics to obtain preprocessed eye movement data.
[0016] Optionally, the pre-processed eye movement data is converted into a topological structure graph based on fixation points and saccades;
[0017] The process of converting into a topological structure diagram includes:
[0018] Divide the word area where the fixation point is located into a node;
[0019] The subject's saccade is considered as a one-way edge between two nodes;
[0020] The pre-processed eye movement data is converted into a topological structure diagram including text and eye saccade relationship based on the nodes divided by the gaze points and the unidirectional edges formed by the eye saccades.
[0021] Optionally, the process of obtaining the node feature includes:
[0022] Use the pre-trained GloVe model to encode the text words in the topological structure graph to obtain word vectors;
[0023] Extract the words at the fixation point to obtain several fixation features;
[0024] Initialize the features of the node using the gaze feature and the word vector;
[0025] The node features are obtained by iterating the initialized nodes based on the gated graph neural network.
[0026] Optionally, the expression for iterating the initialized nodes based on the gated graph neural network is:
[0027]
[0028] Where σ is the sigmoid function, W, U, and b are trainable weights and bias terms, and functions z and r are the update gate and reset gate, respectively. It represents the node representation that integrates the text features and gaze movement features of the node after multiple iterations. The node v∈v, ν is the node set, and t represents the number of iterations.
[0029] Optionally, the expression for obtaining the multidimensional edge feature is:
[0030]
[0031] In the formula, E is the normalized edge feature, i represents the starting node subscript of the edge, j represents the ending node subscript of the edge, and p is the type of edge feature. is the original edge feature.
[0032] Optionally, the expression for interactively fusing the node feature and the multi-dimensional edge feature is:
[0033]
[0034] In the formula, l represents the number of EGNN Layers, represents the connection operation of the node representations of p channels, g l Represents the transformation of the shape of the node feature from the input space to the output space, represents the attention correlation coefficient of a certain channel, X l-1 Indicates the edge and node fusion results of the previous layer in the EGNN model. Represents the multi-dimensional edge fusion result of the previous layer in the EGNN model.
[0035] The present invention also provides a computer terminal device, comprising:
[0036] one or more processors;
[0037] A memory, coupled to the processor, for storing one or more programs;
[0038] When the one or more programs are executed by the one or more processors, the one or more processors implement a multimodal reading eye movement representation method based on graph neural network.
[0039] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, a multimodal reading eye movement representation method based on a graph neural network is implemented.
[0040] Compared with the prior art, the present invention has the following advantages and technical effects:
[0041] The present invention utilizes multi-dimensional eye movement features through the edge feature graph attention network, and integrates gaze features and text stimulus information through iterative updates of node features in the gated neural network. It fully utilizes complex structured eye movement tracking data and more accurately captures and analyzes subtle changes in eye movement data.
[0042] The present invention takes text stimulus information as the main component of node features, and integrates the stimulus information into the final representation in the iterative update of the gated neural network. By integrating eye movement data and text information, the present method provides a more comprehensive perspective to understand eye movement behavior and cognitive processes.
[0043] The present invention converts eye tracking data from a numerical sequence structure to a topological structure, thereby weakening the influence of individual differences on eye movement characteristics and improving the universality and accuracy of the characterization results. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0045] Figure 1 This is a flow chart of a multimodal reading eye movement representation method based on a graph neural network according to an embodiment of the present invention;
[0046] Figure 2 Schematic diagram of topology conversion according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0048] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0049] Embodiment 1
[0050] like Figure 1 As shown, this embodiment provides a multimodal reading eye movement representation method based on a graph neural network, including the following steps:
[0051] Step 1: Normalize the eye tracking data by grouping them according to the subjects to initially eliminate the dimensionality effect brought by the individuals and ensure the consistency of the data in training and testing.
[0052] Step 2: Use the pre-trained GloVe vector results to embed the text stimulus information.
[0053] Step 3: Complete the transformation of eye movement data from numerical sequence structure to topological structure according to the eye saccade sequence. Use text vectors and gaze features to initialize node features in the graph.
[0054] Step 4: According to the principle of gated neural network, iterate the node features until the parameters are stable.
[0055] Step 5: Based on the principle of multi-dimensional edge graph attention network, use the multi-dimensional edge features related to eye saccades and the iterated node features to complete the complete fusion of eye tracking data and text data.
[0056] Step 6: Merge and concatenate the channels related to each edge feature in the multi-dimensional edge graph attention network to complete the multimodal eye movement representation output.
[0057] As a specific implementation method of this embodiment, the following steps are included: obtaining the eye movement data of the subject, preprocessing the eye movement data of the subject of the present invention to obtain preprocessed eye movement data; converting the preprocessed eye movement data of the present invention into a topological structure graph; iterating and multi-dimensionally processing the topological structure graph of the present invention to obtain node features and multi-dimensional edge features; and interactively fusing the node features of the present invention and the multi-dimensional edge features of the present invention using a multi-dimensional modeling method based on a graph attention network to obtain the final reading eye movement representation output.
[0058] Furthermore, data preprocessing: The process of preprocessing the eye movement data of the subject of the present invention includes: dividing the eye movement data of the subject of the present invention according to the subject ID to obtain a number of eye movement data of the same subject; normalizing the eye movement data of the same subject of the present invention one by one according to the eye movement characteristics to obtain preprocessed eye movement data.
[0059] Furthermore, different subjects have different reading habits. Reading habits will cause individual differences in the data scale of eye movement features. In order to eliminate this difference as much as possible, for each eye movement feature, data normalization based on the subject is required. Specifically, the data is divided into several parts according to the subject ID. For the data of the same subject, Min-Max normalization is performed on each eye movement feature one by one.
[0060] Furthermore, topological structure conversion: based on the fixation point and saccades, the preprocessed eye movement data of the present invention is converted into a topological structure diagram; wherein, the process of converting into the topological structure diagram includes: dividing the word area where the fixation point of the present invention is located into a node; the subject's saccade of the present invention is used as a unidirectional edge between two nodes; based on the nodes divided by the fixation point of the present invention and the unidirectional edges formed by the saccades of the present invention, the preprocessed eye movement data is converted into a topological structure diagram including the relationship between text and saccades.
[0061] Furthermore, although the eye movement data is recorded in a time-ordered sequence format, eye movement can be abstracted as the movement of the focus on a plane. Specifically, the surface of the display can be compared to a two-dimensional plane. When reading, the focus of the eye moves on this two-dimensional plane, showing a certain topological relationship. The text information that corresponds to the eye movement can also be converted into a topological structure based on the two-dimensional movement of the eye movement.
[0062] like Figure 2 As shown in Figure 1, the two most important explicit activities in eye tracking are fixation and saccade. Fixation is when the focus stays on a certain point in the material, which is related to the subject's ability to understand the material. Saccade is when the subject's focus quickly jumps on the stimulus material, which is related to the subject's reading speed.
[0063] In the process of converting eye movements to graphs, each word is also regarded as a node. Note that the same word will have different nodes in the short text. The word area where the fixation point is located is regarded as a node. The subject's eye saccade process is regarded as a one-way edge between two nodes. Through this connection method, the complete eye movement process is converted into a topological structure graph containing the relationship between text and eye saccades.
[0064] Furthermore, the process of obtaining the node features of the present invention includes: using a pre-trained GloVe model to encode text words in a topological structure diagram to obtain word vectors; extracting words at the gaze point position to obtain a number of gaze features; initializing the nodes with the gaze features of the present invention and the word vectors of the present invention; and iterating the initialized nodes based on a gated graph neural network to obtain the node features of the present invention.
[0065] Furthermore, multimodal fusion strategy: Eye movement data is generated by subjects viewing stimulus information, and its influencing factors include individual attributes of subjects and text stimulus information. In order to effectively utilize stimulus information, a multimodal data fusion strategy is designed.
[0066] The present invention focuses on eye movement during reading, where the text content is the reading stimulus. When reading, the human eye pattern is related to difficult words, long words, etc. in the material, and the subject's comprehension process will affect their reading behavior. The words at the fixation point are extracted and used together with multiple fixation features to initialize the node representation (Step 3). The words at the fixation point are encoded using the pre-trained GloVe text (Step 2) and vectorized. Words not included in the vocabulary (out-of vocabulary, OOV) are randomly sampled from a uniform distribution [-0.01, 0.01].
[0067] In each graph composed of reading texts, a gated graph neural network (GGNN) is used to iteratively optimize the node representation. By modifying the unit structure of GRU, each node uses the information of neighboring nodes and its own node information for iterative update (Step 4). Word iteration can only make the information in the graph interact between first-order neighbors, and multiple iterations can realize the interaction of node information in the entire graph. The specific calculation formula is as follows:
[0068]
[0069] Where t represents the number of iterations. Node v∈v, ν is a node set, and |ν| represents the node subscript. It is the initial node representation composed of text vector and gaze feature. D|ν|×2D|ν| is the positive and negative adjacency matrix of two concatenations (A (in) and A (out) ). v: ∈R D|ν|×2D is the matrix A (in) and A (out) The two lines about node v in .
[0070]
[0071]
[0072] Where σ is the sigmoid function, W, U and b are trainable weights and bias terms. Functions z and r are the update gate and reset gate, respectively, which determine the degree of information interaction between neighbor nodes and the current node. Node representation after multiple iterations It combines the text features and eye movement features of nodes, and through information interaction with neighboring nodes, it can learn the overall topological information of the graph.
[0073] Furthermore, multi-dimensional edge feature modeling: There are multiple features in eye movement behavior. In the traditional graph neural network structure, the edge feature only contains a single attribute related to the connectivity state. In order to effectively utilize all eye movement features, a multi-dimensional edge processing method is introduced.
[0074] In graph operations, the edge feature matrix will be used as a filter and multiplied with the node feature matrix. In order to avoid multiplication increasing the scale of the output dimension, double random normalization is used to uniformly process the edge features, as shown in formulas (6)-(7).
[0075]
[0076] in is the original edge feature, E is the normalized edge feature. i represents the starting node subscript of the edge, j represents the ending node subscript of the edge, and p is a certain edge feature, which refers to a certain eye saccade feature in this patent. The original edge feature matrix is projected onto the doubly random matrix space so that the sum of each row and column of the matrix is equal to 1.
[0077] In order to interactively fuse the node features and multi-dimensional edge features after iteration, a two-layer EGNN layer is introduced, which is a multi-dimensional edge modeling method based on Graph Attention Networks (GAT). This method regards multi-dimensional edge features as multi-channel signals, each channel guides an independent attention operation (Step 5), and finally concatenates the results of each channel as the final output (Step 6). The calculation process is shown in formulas (8)-(13).
[0078]
[0079] Iteratively integrate the node representation x of text features and gaze features v ∈X, l represents the number of EGNN Layer. σ represents the nonlinear activation function. represents the concatenation operation on the node representations of p channels. l-1 Indicates the edge and node fusion results of the previous layer in the EGNN model. Indicates the multi-dimensional edge fusion result of the previous layer in the EGNN model. l Used to transform the shape of a node feature from the input space to the output space.
[0080] g l =X l-1 W l #9)
[0081] It represents the attention coefficients of a channel. Through the connection operation, p Splice to α l The edge feature E of the next layer l Also from α l .
[0082] E l =α l #10)
[0083] Specifically, About and is a function of , as shown in formulas (11)-(12).
[0084]
[0085] Where DS represents the double random normalization operation, as shown in formulas (6)-(7). l is an attention function that can generate a scalar value based on two vector inputs, as shown in formula (13).
[0086]
[0087] Where L represents the LeakyReLU activation function, W is the same mapping weight as in formula (9), and || is the connection operation.
[0088] After two layers of EGNN Layer, the node features and multi-dimensional edge features are fully integrated. l The channels in the are connected in series as the final reading eye movement representation output.
[0089] The present invention proposes a multimodal eye movement representation method which combines a gated neural network and an edge feature map attention network, which can effectively fuse eye movement features and text stimulus information to achieve efficient representation of eye movement tracking data.
[0090] The present invention can fully utilize multi-dimensional eye movement characteristics and realize efficient utilization of eye movement tracking data of complex structures.
[0091] The present invention, by converting the numerical sequence into a topological structure, weakens the influence of individual differences in numerical values on the characterization result to a certain extent.
[0092] Embodiment 2
[0093] This embodiment also provides a computer terminal device, including:
[0094] one or more processors;
[0095] A memory, coupled to the processor of the present invention, for storing one or more programs;
[0096] When the one or more programs are executed by the one or more processors, the one or more processors implement a multimodal reading eye movement representation method based on graph neural network.
[0097] This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, a multimodal reading eye movement representation method based on a graph neural network is implemented.
[0098] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A multimodal reading eye movement representation method based on graph neural network, characterized in that: The following steps are involved: Acquiring eye movement data of the subject, and preprocessing the eye movement data of the subject to obtain preprocessed eye movement data; Converting the pre-processed eye movement data into a topological structure graph; Converting the pre-processed eye movement data into a topological structure graph based on fixation points and saccades; The process of converting into a topological structure diagram includes: dividing the word area where the fixation point is located into a node; the eye saccade of the subject is used as a unidirectional edge between two nodes; based on the nodes divided by the fixation point and the unidirectional edge formed by the eye saccade, converting the pre-processed eye movement data into a topological structure diagram including the relationship between text and eye saccade; Iterating and multi-dimensional attention processing the topological structure graph to obtain node features and multi-dimensional edge features; The process of obtaining the node features includes: using a pre-trained GloVe model to encode text words in a topological structure diagram to obtain word vectors; extracting words at the gaze point position to obtain a number of gaze features; initializing the features of the node using the gaze features and the word vectors; iterating the initialized nodes based on a gated graph neural network to obtain the node features; The multi-dimensional modeling method based on the graph attention network interactively fuses the node features and the multi-dimensional edge features to obtain the final reading eye movement representation output.
2. According to claim 1, the multimodal reading eye movement representation method based on graph neural network is characterized in that: The process of preprocessing the subject's eye movement data includes: Dividing the eye movement data of the subject according to the subject ID to obtain a plurality of eye movement data of the same subject; The eye movement data of the same subject are normalized one by one according to the eye movement characteristics to obtain preprocessed eye movement data.
3. The multimodal reading eye movement representation method based on graph neural network according to claim 1 is characterized in that: The expression for iterating the initialized nodes based on the gated graph neural network is: Where σ is the sigmoid function, W, U, and b are trainable weights and bias terms, and functions z and r are the update gate and reset gate, respectively. It represents the node representation that integrates the text features and gaze movement features of the node after multiple iterations. The node v∈v, v is a node set, and t represents the number of iterations.
4. The multimodal reading eye movement representation method based on graph neural network according to claim 3 is characterized in that: The expression for obtaining the multi-dimensional edge feature is: Where E is the normalized edge feature, i represents the starting node subscript of the edge, j represents the ending node subscript of the edge, and p is the type of edge feature. is the original edge feature.
5. The multimodal reading eye movement representation method based on graph neural network according to claim 4 is characterized in that: The expression for interactively fusing the node features and the multi-dimensional edge features is: In the formula, l represents the number of EGNN Layers, represents the connection operation of the node representations of p channels, g l Represents the transformation of the shape of the node feature from the input space to the output space, represents the attention correlation coefficient of a certain channel, X l-1 Indicates the edge and node fusion results of the previous layer in the EGNN model. Represents the multi-dimensional edge fusion result of the previous layer in the EGNN model.
6. A computer terminal device, characterized in that: include: one or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multimodal reading eye movement representation method based on graph neural network as described in any one of claims 1-5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the multimodal reading eye movement representation method based on graph neural network as described in any one of claims 1-5.
Citation Information
Patent Citations
Image preference prediction method based on eye movement image reasoning
CN115439921A
Method for pre-judging model establishment of information awareness degree based on classification model
CN117390517A