Image-based multi-modal eye movement characterization method and system

By transforming, feature extraction, encoding and multimodal fusion of the original eye movement data, high-quality eye movement representation is generated, which solves the problems of data discretization and information loss in traditional methods, and achieves more accurate eye movement dynamic changes and gaze order reflection.

CN120220219APending Publication Date: 2025-06-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510305690.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional eye movement data feature extraction methods lead to data discretization and information loss, resulting in poor eye movement characterization quality.

Method used

Using an image-based multimodal eye movement representation method, by transforming, feature extraction, encoding and multimodal fusion of the original eye movement data, a comprehensive hidden vector is generated and input into a self-attention network to obtain high-quality eye movement representation.

Benefits of technology

It effectively retains the dynamic change information of eye movement data, improves the quality of eye movement representation, can clearly show the dynamic changes of eye movement over time, and reflect the order of gaze during reading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220219A_ABST
    Figure CN120220219A_ABST
Patent Text Reader

Abstract

The invention provides an image-based multi-modal eye movement characterization method and system, and belongs to the technical field of eye movement characterization, and the method comprises the steps: converting original eye movement data, and obtaining an eye movement linear graph corresponding to the original eye movement data; obtaining a glancing path based on an eye movement feature, wherein the eye movement feature is a feature obtained after feature extraction is performed on the original eye movement data; encoding the eye movement line graph to obtain a target image implicit vector, and encoding based on text information corresponding to the glancing path to obtain a target text implicit vector; fusing the target image implicit vector and the target text implicit vector to obtain a comprehensive implicit vector; and inputting the comprehensive implicit vector into a self-attention network for transmission to obtain eye movement characterization. According to the image-based multi-modal eye movement characterization method and system provided by the invention, the quality of eye movement characterization can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of eye movement representation, and more specifically, relates to an image-based multi-modal eye movement representation method and system. Background Art

[0002] Eye movement data plays a crucial role in the cognitive process of humans. The representation and analysis of eye movement data have irreplaceable value for revealing the human visual cognitive mechanism.

[0003] Eye movement data contains rich dynamic changes and complex patterns. Traditional feature extraction discretizes and averages the coherent raw data, losing valuable information contained in the dynamic changes, thus resulting in poor quality of eye movement representation. Summary of the Invention

[0004] The purpose of the present disclosure is to provide an image-based multi-modal eye movement representation method and system to improve the quality of eye movement representation.

[0005] In the first aspect of the embodiments of the present disclosure, an image-based multi-modal eye movement representation method is provided, including: Converting the original eye movement data to obtain an eye movement line graph corresponding to the original eye movement data. The original eye movement data is a multi-variable time series, the time series includes multiple timestamps, and the multi-variables are respectively the abscissa, ordinate, and pupil diameter of the line-of-sight landing point at each timestamp; Obtaining a saccade path based on eye movement features, where the eye movement features are features obtained by performing feature extraction on the original eye movement data; Encoding the eye movement line graph to obtain a target image hidden vector; encoding the text information corresponding to the saccade path to obtain a target text hidden vector; Performing multi-modal fusion on the target image hidden vector and the target text hidden vector to obtain a comprehensive hidden vector; Inputting the comprehensive hidden vector into a self-attention network to obtain an eye movement representation.

[0006] In the second aspect of the embodiments of the present disclosure, an image-based multi-modal eye movement representation system is provided, including a line graph generation module for converting the original eye movement data to obtain an eye movement line graph corresponding to the original eye movement data. The original eye movement data is a multi-variable time series, the time series includes multiple timestamps, and the multi-variables are respectively the abscissa, ordinate, and pupil diameter of the line-of-sight landing point at each timestamp; A saccade path generation module for obtaining a saccade path based on eye movement features, where the eye movement features are features obtained by performing feature extraction on the original eye movement data; An encoding module for encoding the eye movement line graph to obtain a target image hidden vector; encoding the text information corresponding to the saccade path to obtain a target text hidden vector; A fusion module for performing multimodal fusion on a target image latent vector and a target text latent vector to obtain a comprehensive latent vector; An eye movement representation module for inputting the comprehensive latent vector into a self-attention network to obtain an eye movement representation.

[0007] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned method for image-based multimodal eye movement representation are implemented.

[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for image-based multimodal eye movement representation are implemented.

[0009] The beneficial effects of the method and system for image-based multimodal eye movement representation provided by the embodiments of the present disclosure are as follows: In the embodiments of the present disclosure, the original eye movement data is converted into an eye movement line graph, avoiding the discretization problem of traditional feature extraction, completely retaining data information, and clearly showing the dynamic changes of eye movement over time. By extracting eye movement features to obtain a saccade path, the fixation order during the reading process can be effectively reflected. The eye movement line graph and the saccade path text information are respectively encoded, and then the obtained image latent vector and text latent vector are fused. By using the complementarity of different modal information, a comprehensive latent vector containing rich information is obtained, improving the quality of eye movement representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a schematic flowchart of a method for image-based multimodal eye movement representation provided by an embodiment of the present disclosure; Figure 2 It is a schematic diagram of eye movement visualization provided by an embodiment of the present disclosure; Figure 3 It is a schematic diagram of an interest area provided by an embodiment of the present disclosure; Figure 4 It is a schematic flowchart of another method for image-based multimodal eye movement representation provided by an embodiment of the present disclosure; Figure 5Schematic flowchart of another image-based multi-modal eye movement representation method provided by an embodiment of the present disclosure; Figure 6 Block diagram of a structure of an image-based multi-modal eye movement representation system provided by an embodiment of the present disclosure; Figure 7 Schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0012] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0013] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.

[0014] Please refer to Figure 1 , Figure 1 Schematic flowchart of an image-based multi-modal eye movement representation method provided by an embodiment of the present disclosure. The method includes: S101: Convert the original eye movement data to obtain an eye movement line graph corresponding to the original eye movement data. The original eye movement data is a multi-variable time series, the time series includes multiple timestamps, and the multi-variables are respectively the abscissa, ordinate, and pupil diameter of the line-of-sight landing point at each timestamp.

[0015] In this embodiment, the eye movement data can be collected by an eye movement tracking instrument to record the landing point of the line of sight on the screen. The original eye movement data is a multi-variable time series, and each timestamp includes the two-dimensional coordinates (i.e., x and y coordinates) of the line-of-sight landing point and the pupil diameter. According to the change of the landing point, the original eye movement can be divided into fixation (the eyes are relatively stationary) and saccade (the eyes move significantly). The eye movement can be regarded as a process of fixation and saccade cycling in turn, as Figure 2 shown.

[0016] In this embodiment, the original eye movement data is converted into an eye movement line graph. With time as the horizontal axis, the abscissa, ordinate, and pupil diameter are respectively used as the vertical axes to draw individual line graphs. Thereby, the complete information of the data is retained, the dynamic changes are displayed, the discretization and information loss caused by traditional feature extraction are avoided, and the change of eye movement over time can be intuitively presented.

[0017] S102: Obtain a saccade path based on the eye movement features. The eye movement features are the features obtained by performing feature extraction on the original eye movement data.

[0018] In this embodiment, there is text information on the screen, and the text information can be divided into independent regions, namely regions of interest, as Figure 3 shown. Based on the regions of interest, a series of eye movement features can be calculated, such as the number of fixations (the number of fixation points within the region of interest), the average fixation duration, the saccade amplitude, etc.

[0019] The fixation points can be determined based on three thresholds of eye movement displacement, velocity, and acceleration. By sorting the time series of the fixation points, the saccade path can be obtained.

[0020] The saccade path is a sequence of fixation indices arranged in chronological order of the fixation points, marking the word numbers where the fixation points are located, and reflecting the fixation order during reading.

[0021] S103: Encode the eye movement line graph to obtain a target image hidden vector; encode the text information corresponding to the saccade path to obtain a target text hidden vector.

[0022] In this embodiment, a Vision Transformer (ViT) can be used to encode the eye movement line graph. ViT has a powerful ability to capture local and global dependencies, and the columnar structure can maintain the correspondence between the output features and time, which helps to better capture the local subtle movements and overall global movements involved in eye movement, and finally obtain a target image hidden vector.

[0023] In this embodiment, according to the text information corresponding to the saccade path, a Bidirectional Encoder Representations from Transformers (BERT) is used to encode the text information. BERT is a model based on the Transformers architecture and can obtain text features, thereby obtaining a target text hidden vector.

[0024] S104: Perform multimodal fusion on the target image hidden vector and the target text hidden vector to obtain a comprehensive hidden vector.

[0025] In this embodiment, the target image hidden vector and the target text hidden vector are fused to obtain a comprehensive hidden vector. Among them, the target image hidden vector represents eye movement features, and the target text hidden vector represents text features.

[0026] During the fusion process, Canonical Correlation Analysis (CCA) can be used to maximize the correlation between the eye movement features and the text feature vector set, coordinating the two modalities. At the same time, the Wasserstein Distance (WD) is used to measure the minimum cost of aligning the distributions of the two modalities, solving the cross-modal difference problem, so that the fused comprehensive latent vector contains richer cross-domain complementary information.

[0027] S105: Input the comprehensive latent vector into the self-attention network to obtain the eye movement representation.

[0028] In this embodiment, the comprehensive latent vector can be input into the self-attention network to transfer information. The self-attention mechanism can capture time-dependent relationships and transfer motion states between different patch time periods. After being processed by the self-attention network, the eye movement representation is obtained, which can be used for various downstream tasks. For example, in the field of human-computer interaction, the eye movement representation can help designers design more natural interaction interfaces and improve the user experience, such as virtual reality and online education.

[0029] It can be concluded from the above that in this embodiment, the original eye movement data is converted into an eye movement line graph, avoiding the discretization problem of traditional feature extraction, completely retaining the data information, and clearly showing the dynamic changes of eye movement over time. By extracting the eye movement features to obtain the saccade path, the fixation order during the reading process can be effectively reflected. The eye movement line graph and the saccade path text information are encoded respectively, and then the obtained latent vectors are fused. Utilizing the complementarity of different modality information, a comprehensive latent vector containing rich information is obtained, improving the quality of the eye movement representation.

[0030] In an embodiment of the present disclosure, obtaining the saccade path based on the eye movement features includes: Determine the fixation points within the region of interest based on the displacement, velocity, and acceleration of the eye movement within the region of interest; Sort based on the time series of the fixation points to obtain the saccade path.

[0031] In this embodiment, the text information on the screen can be divided into independent regions, and these regions are defined as regions of interest. By analyzing the eye movement situation within the regions of interest, relevant information can be obtained.

[0032] After the eye-tracking instrument collects the original eye movement data, it contains multivariate information such as the abscissa, ordinate, and pupil diameter of the line-of-sight landing point at each timestamp. According to the relative movement size of the landing point, the fixation points in each region of interest can be calculated. Fixations are detected based on three thresholds of displacement, velocity, and acceleration. When the original data is less than the threshold, it can be considered a fixation behavior. These three thresholds can be set according to the purpose of the eye-tracking task, and one or more sampling points in the original data that are less than the threshold are aggregated into a fixation.

[0033] For example, it can be set that the displacement between the front and back landing points is less than , the velocity is lower than , and the acceleration is less than . Then, the central position of the line-of-sight landing points within the above thresholds is determined as a fixation point within the region of interest.

[0034] By sorting the fixation points according to their corresponding timestamps, an ordered sequence of fixation points can be obtained, and this sequence is the saccade path.

[0035] For example, when reading a text, in chronological order, the line of sight first lands on the region of interest where the first word is located (generating the first fixation point), then moves to the region of interest where the second word is located (generating the second fixation point), and so on. After sorting these fixation points by time, the saccade path reflecting the fixation order during reading is formed.

[0036] From the above, it can be concluded that in this embodiment, the fixation points are determined based on the displacement, velocity, and acceleration of eye movements within the region of interest, which can accurately capture the attention of the human eye in the key region. Then, by sorting the time series of fixation points to obtain the saccade path, the movement trajectory and attention order of the human eye can be clearly presented.

[0037] As Figure 4 shown, in an embodiment of the present disclosure, the eye movement line graph is encoded to obtain a target image hidden vector, including: Segment the eye movement line graph to obtain multiple image patches; Aggregate the multiple image patches based on the patch time periods corresponding to each image patch to obtain multiple image hidden vectors; Input the multiple image hidden vectors into a recurrent neural network to obtain the target image hidden vector.

[0038] In this embodiment, the eye movement line graph can be input into ViT for segmentation to obtain multiple image patches.

[0039] The eye movement line graph records eye movement information over time, including the coordinates of the line of sight landing point and the pupil diameter, etc. To align eye movements with text at the same unit time, ViT can divide the eye movement line graph into multiple image patches, each patch corresponding to a specific duration, i.e., patch time.

[0040] In this embodiment, according to the patch time corresponding to each image patch, each fixation point in the saccade path is aggregated according to its occurrence time, and the fixation points within the same patch time are grouped into one category. This can make the eye movement information and text information better correspond in the time dimension.

[0041] Subsequently, ViT is used to encode each image patch. Since eye movements include local fine movements in a single direction and global movements between different directions, and ViT has a stronger ability to capture local and global dependency relationships compared to mainstream CNNs, and its columnar structure can also maintain the correspondence between the output features and time, which is beneficial for subsequent multimodal fusion. Through ViT encoding, each image patch is transformed into a corresponding image hidden vector, and the image hidden vector contains the eye movement features within the corresponding time period.

[0042] In this embodiment, the recurrent neural network can be a Gated Recurrent Unit (GRU) network. GRU can update the hidden state at the current moment according to the input at the current moment and the hidden state at the previous moment. In this way, GRU can learn the temporal order and dependency relationships between the image hidden vectors, so as to integrate the information of the entire eye movement line graph in different time periods. Finally, the hidden state output by GRU is the target image hidden vector, and the target image hidden vector contains the global feature information of the entire eye movement process, and can represent the eye movement features contained in the eye movement line graph more comprehensively and accurately.

[0043] It can be concluded from the above that in this embodiment, the image patches are first segmented to facilitate focusing on local eye movement features; multiple image hidden vectors are obtained by classifying according to the patch time period to achieve effective representation of eye movement data; and then through integration by the recurrent neural network, the obtained target image hidden vector can comprehensively reflect the global features of eye movements, thereby improving the accuracy of eye movement analysis.

[0044] As Figure 4 shown, in an embodiment of the present disclosure, encoding the text information corresponding to the saccade path to obtain a target text hidden vector includes: Performing feature extraction on the text information to obtain text features; Sorting the text features based on the saccade path to obtain multiple text hidden vectors; Input multiple text hidden vectors into a recurrent neural network to obtain a target text hidden vector.

[0045] In this embodiment, BERT is used to process text information. As a model based on the Transformer architecture, BERT can effectively extract rich features such as semantics and grammar in the text. Through the encoding mechanism of BERT, the text information is transformed into a text feature representation that contains its inherent semantics and structure.

[0046] The saccade path records the order of fixations during reading. According to this order, retrieve the features at the corresponding positions from the text features output by BERT, thereby obtaining multiple text hidden vectors. Combine the saccade path with the text features to associate the order of the text hidden vectors with the eye movement order, reflecting the features corresponding to the text content focused on at different times during reading, and embodying the attention order during reading.

[0047] Multiple text hidden vectors sorted by the saccade path are input into the GRU in chronological order. The GRU can integrate the text feature information at different times during the processing. Extract the output of the last time step from the same patch time as the text feature of that patch time because the output of the last time step synthesizes the text information within that period. After the GRU integrates the text features corresponding to all patch times, finally output the target text hidden vector, which fuses the text information and the reading order information related to eye movements.

[0048] It can be concluded from the above that the extraction of text information features in this embodiment can mine its inherent semantics. Sorting the text features based on the saccade path can combine eye movement information and reflect the attention order to different text contents during reading. After being processed by the recurrent neural network and integrating the information, the target text hidden vector becomes more comprehensive.

[0049] In an embodiment of the present disclosure, fuse the target image hidden vector and the target text hidden vector to obtain a comprehensive hidden vector, including: Perform time dimension alignment on the target image hidden vector and the target text hidden vector based on canonical correlation analysis and optimal transport distance to obtain a comprehensive hidden vector.

[0050] Canonical correlation analysis is a statistical method used to analyze the correlation between two sets of variables. In this embodiment, canonical correlation analysis is used to maximize the correlation between two sets of vectors, namely the target image latent vectors and the target text latent vectors. Through CCA, linear combinations of the two sets of variables can be found such that the correlation between these two linear combinations is maximized. CCA can separately find suitable projection directions in the target image latent vector space and the target text latent vector space, and project the original latent vectors into a new low-dimensional space. In the low-dimensional space, the correlation between the two sets of projected vectors is maximized. Thus, the data of the image and text modalities are correlated with each other.

[0051] In this embodiment, since the target image latent vectors and the target text latent vectors come from different modalities, there may be significant differences in their data distributions. The optimal transport distance can be used to measure the minimum cost of aligning the distributions of the two modalities.

[0052] By calculating the optimal transport distance, an optimal way can be found to align the distribution of the target image latent vectors and the distribution of the target text latent vectors. This means that during the fusion process, the differences between the data of the two modalities can be minimized as much as possible, enabling them to better complement each other in the fused comprehensive latent vectors, thereby achieving a better fusion effect.

[0053] After maximizing the correlation between the two sets of vectors through canonical correlation analysis and aligning the distributions of the two modalities using the optimal transport distance, the processed target image latent vectors and target text latent vectors are fused to obtain comprehensive latent vectors, which contain information from both the image and text modalities.

[0054] It can be concluded from the above that in this embodiment, through canonical correlation analysis and the optimal transport distance, the effective alignment and fusion of the target image latent vectors and the target text latent vectors are achieved, and comprehensive latent vectors containing cross-domain complementary information are obtained.

[0055] In one embodiment of the present disclosure, canonical correlation analysis is expressed as:

[0056] where represents the optimal linear projection vector group of the target image latent vectors and the target text latent vectors, represents the optimal linear projection vector of the target image latent vectors, represents the optimal linear projection vector of the target text latent vectors, represents the linear projection vector of the target image latent vectors, represents the linear projection vector of the target text latent vectors, represents the cross-covariance matrix of the target image latent vectors and the target text latent vectors, represents the covariance matrix of the target image hidden vector, represents the covariance matrix of the target text hidden vector, represents the transpose of the linear projection vector of the target image hidden vector, represents the transpose of the linear projection vector of the target text hidden vector.

[0057] In this embodiment, let the matrix formed by the target image hidden vectors (eye movement features) be , and the matrix formed by the target text hidden vectors (text features) be . Among them, c represents the number of patch times, and d represents the feature dimension. The canonical correlation analysis is expressed as:

[0058] Calculate their covariance matrices , and the cross-covariance matrix based on the target image hidden vector and the target text hidden vector. By solving the formula, find the optimal linear projection vectors and . Finally, use these two projection vectors to perform linear projection on the target image hidden vector and the target text hidden vector to obtain the projected vectors, and the correlation between them is the maximum correlation of the original vector set.

[0059] In an embodiment of the present disclosure, the optimal transport distance is expressed as:

[0060] where, represents the distance between the target image hidden vector and the target text hidden vector, represents the element of the nth target image hidden vector, represents the element of the mth target text hidden vector, represents the transport plan matrix, represents the joint distribution of the target image hidden vector and the target text hidden vector, represents the discrete distribution of the target image hidden vector, represents the discrete distribution of the target text hidden vector, represents the cost function between the element of the nth target image hidden vector and the element of the mth target text hidden vector.

[0061] In multimodal data processing, the target image hidden vector and the target text hidden vector belong to different modalities, and there are differences in data distributions. The goal of optimal transport is to find a transport mapping to clarify how much one distribution needs to move towards the other distribution, so as to achieve the alignment of the two distributions.

[0062] In this embodiment, the target image latent vector matrix X and the target text latent vector matrix Y are converted into one-dimensional vectors x and y, both of which have dimensions of , where c is the number of patch times and d is the feature dimension. Define the discrete distributions of eye movement features and text features and , which are respectively represented by the weighted sum of Dirac functions, as follows: and , where is the Dirac function centered at x.

[0063] The weight vectors and belong to the b-dimensional simplex and satisfy , then the WD between and is defined as:

[0064] where, represents all possible joint distributions , , defining the set of transfer matrices T. Among them, is a b-dimensional vector all ones. Each element of the transfer matrix T represents the transfer amount from to , and satisfies that the row sum is equal to the corresponding element of , and the column sum is equal to the corresponding element of , ensuring the conservation of probability mass during the transfer process.

[0065] The optimal transport distance formula is to find an optimal T among all possible transfer matrices such that the total cost of the transfer is minimized. By minimizing this cost, an optimal way is found to align the distribution of the target image latent vectors to the distribution of the target text latent vectors. The minimum cost is the WD between the two distributions, that is, the distance between the target image latent vectors and the target text latent vectors.

[0066] It can be concluded from the above that in this embodiment, by finding the optimal transport plan matrix and considering the cost function between elements, the minimum-cost transfer of distribution alignment is achieved, and multi-modal information is retained and fused to the greatest extent.

[0067] Please refer to Figure 5 , Figure 5Schematic flowchart of another image - based multi - modal eye movement representation method provided by an embodiment of the present disclosure. The original eye movement data is preliminarily processed by being converted into a line graph and determining the saccade path; the ViT network is used to encode the eye movement line graph to extract eye movement features; the BERT network is used to process the text information related to the saccade path to obtain text features. The two features are aligned in the time dimension, and then multi - modal fusion is performed to integrate information of different modalities, and finally a representation vector containing rich information is output.

[0068] A method for image - based multi - modal eye movement representation corresponding to the above - mentioned embodiment Figure 6 Block diagram of the structure of an image - based multi - modal eye movement representation system provided by an embodiment of the present disclosure. For the sake of illustration, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 6 The image - based multi - modal eye movement representation system 20 includes: a line graph generation module 21, a saccade path generation module 22, an encoding module 23, a fusion module 24, and an eye movement representation module 25.

[0069] Among them, the line graph generation module 21 is used to convert the original eye movement data to obtain an eye movement line graph corresponding to the original eye movement data. The original eye movement data is a multi - variable time series, the time series includes multiple timestamps, and the multi - variables are respectively the abscissa, ordinate, and pupil diameter of the line - of - sight landing point at each timestamp. The saccade path generation module 22 is used to obtain the saccade path based on the eye movement features, and the eye movement features are the features obtained by extracting features from the original eye movement data. The encoding module 23 is used to encode the eye movement line graph to obtain a target image hidden vector; encode the text information corresponding to the saccade path to obtain a target text hidden vector. The fusion module 24 is used to perform multi - modal fusion on the target image hidden vector and the target text hidden vector to obtain a comprehensive hidden vector. The eye movement representation module 25 is used to input the comprehensive hidden vector into the self - attention network to obtain the eye movement representation.

[0070] In an embodiment of the present disclosure, the saccade path generation module 22 is specifically used for: Determining the fixation points within the region of interest based on the displacement, velocity, and acceleration of eye movement within the region of interest; Sorting the time series of the fixation points to obtain the saccade path.

[0071] In an embodiment of the present disclosure, the encoding module 23 is specifically used for: Segmenting the eye movement line graph to obtain a plurality of image patches; Aggregating the plurality of image patches based on the patch time periods corresponding to each image patch to obtain a plurality of image hidden vectors. Input multiple image latent vectors into a recurrent neural network to obtain a target image latent vector.

[0072] In an embodiment of the present disclosure, the encoding module 23 is further specifically configured to: Encode the text information corresponding to the saccade path to obtain a target text latent vector, including: Extract features from the text information to obtain text features; Sort the text features based on the saccade path to obtain multiple text latent vectors; Input the multiple text latent vectors into a recurrent neural network to obtain a target text latent vector.

[0073] In an embodiment of the present disclosure, the fusion module 24 is specifically configured to: Align the target image latent vector and the target text latent vector in the time dimension based on canonical correlation analysis and optimal transport distance to obtain a comprehensive latent vector.

[0074] In an embodiment of the present disclosure, the canonical correlation analysis is expressed as:

[0075] Among them, represents the optimal linear projection vector group of the target image latent vector and the target text latent vector, represents the optimal linear projection vector of the target image latent vector, represents the optimal linear projection vector of the target text latent vector, represents the linear projection vector of the target image latent vector, represents the linear projection vector of the target text latent vector, represents the cross-covariance matrix of the target image latent vector and the target text latent vector, represents the covariance matrix of the target image latent vector, represents the covariance matrix of the target text latent vector, represents the transpose of the linear projection vector of the target image latent vector, represents the transpose of the linear projection vector of the target text latent vector.

[0076] In an embodiment of the present disclosure, the optimal transport distance is expressed as:

[0077] Among them, represents the distance between the target image latent vector and the target text latent vector, represents the element of the nth target image latent vector, represents the element of the mth target text latent vector, represents the transport plan matrix, represents the joint distribution of the target image latent vector and the target text latent vector, represents the discrete distribution of the target image latent vector, represents the discrete distribution of the target text latent vector, represents the cost function between the elements of the nth target image latent vector and the elements of the mth target text latent vector.

[0078] See Figure 7 , Figure 7 which is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 7 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above system embodiments, such as Figure 6 the functions of the modules 21 to 25 shown.

[0079] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or this processor may also be any conventional processor, etc.

[0080] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0081] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0082] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first and second embodiments of a multi-modal eye movement representation method based on images provided by the embodiments of the present disclosure, and may also execute the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0083] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or system capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0084] The computer-readable storage medium may be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0085] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0086] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0087] In several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, or can also be electrical, mechanical, or other forms of connection.

[0088] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.

[0089] In addition, the functional units in various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0090] The above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by this disclosure, and these modifications or substitutions should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. An image-based multimodal eye movement representation method, characterized in that: include: Converting the original eye movement data to obtain an eye movement line graph corresponding to the original eye movement data, wherein the original eye movement data is a multivariate time series, the time series includes multiple timestamps, and the multiple variables are the horizontal coordinate, the vertical coordinate, and the pupil diameter of the sight point at each timestamp; Obtaining a scanning path based on eye movement features, wherein the eye movement features are features extracted from the original eye movement data; Encoding the eye movement line graph to obtain a target image latent vector; encoding the text information corresponding to the scanning path to obtain a target text latent vector; Performing multimodal fusion on the target image latent vector and the target text latent vector to obtain a comprehensive latent vector; The comprehensive latent vector is input into the self-attention network to obtain eye movement representation.

2. The image-based multimodal eye movement representation method according to claim 1, characterized in that: The obtaining of the scanning path based on the eye movement characteristics comprises: determining a fixation point within the region of interest based on displacement, velocity, and acceleration of eye movements within the region of interest; The time series of fixation points are sorted to obtain the scanning path.

3. The image-based multimodal eye movement representation method according to claim 1, characterized in that: The step of encoding the eye movement line graph to obtain a target image latent vector comprises: Segmenting the eye movement line graph to obtain a plurality of image patches; Aggregating the plurality of image patches based on a patch period corresponding to each image patch to obtain a plurality of image latent vectors; The multiple image latent vectors are input into a recurrent neural network to obtain a target image latent vector.

4. The image-based multimodal eye movement representation method according to claim 1, characterized in that: The encoding of the text information corresponding to the scanning path to obtain a target text latent vector includes: Performing feature extraction on the text information to obtain text features; Sorting the text features based on the scanning path to obtain a plurality of text latent vectors; The multiple text latent vectors are input into a recurrent neural network to obtain a target text latent vector.

5. The image-based multimodal eye movement representation method according to claim 1, characterized in that: The step of fusing the target image latent vector and the target text latent vector to obtain a comprehensive latent vector includes: Based on canonical correlation analysis and optimal transmission distance, the target image latent vector and the target text latent vector are aligned in time dimension to obtain a comprehensive latent vector.

6. The image-based multimodal eye movement representation method according to claim 5, characterized in that: The canonical correlation analysis is expressed as: in, Represents the optimal linear projection vector group of the target image latent vector and the target text latent vector, represents the optimal linear projection vector of the target image latent vector, represents the optimal linear projection vector of the target text latent vector, represents the linear projection vector of the target image latent vector, Represents the linear projection vector of the target text latent vector, represents the cross variance matrix between the target image latent vector and the target text latent vector, represents the covariance matrix of the target image latent vector, represents the covariance matrix of the target text latent vector, represents the transpose of the linear projection vector of the target image latent vector, Represents the transpose of the linear projection vector of the target text latent vector.

7. The image-based multimodal eye movement representation method according to claim 5, characterized in that: The optimal transmission distance is expressed as: in, represents the distance between the target image latent vector and the target text latent vector, represents the element of the nth target image latent vector, represents the element of the mth target text latent vector, represents the transmission plan matrix, represents the joint distribution of the target image latent vector and the target text latent vector, represents the discrete distribution of the latent vector of the target image, represents the discrete distribution of the target text latent vector, Represents the cost function between the elements of the nth target image latent vector and the elements of the mth target text latent vector.

8. An image-based multimodal eye movement representation system, characterized in that: include: A line graph generation module is used to convert the original eye movement data to obtain an eye movement line graph corresponding to the original eye movement data, wherein the original eye movement data is a multivariate time series, the time series includes multiple timestamps, and the multiple variables are the horizontal coordinate, the vertical coordinate and the pupil diameter of the sight point at each timestamp; A scanning path generation module, used for obtaining a scanning path based on eye movement features, wherein the eye movement features are features extracted from the original eye movement data; An encoding module is used to encode the eye movement line graph to obtain a target image latent vector; encode the text information corresponding to the scanning path to obtain a target text latent vector; A fusion module, used for performing multimodal fusion on the target image latent vector and the target text latent vector to obtain a comprehensive latent vector; The eye movement representation module is used to input the comprehensive latent vector into the self-attention network to obtain the eye movement representation.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.