A multimodal three-dimensional visual attention prediction method and its application
Through the multimodal three-dimensional visual attention prediction method, combined with eye movement and head motion data, the spherical convolution model and attention long short-term memory artificial module are used to realize high-precision visual attention prediction in three-dimensional space, solving the problems of noise interference and low accuracy in traditional technology, and providing efficient and low-cost three-dimensional space design evaluation support.
Patent Information
- Application Number
- CN202111465974.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-03
AI Technical Summary
Traditional eye tracking technology is difficult to achieve high-precision visual attention detection in three-dimensional space, and the lack of multimodal data fusion makes noise interference difficult to remove, affecting the accuracy of prediction results.
The multimodal three-dimensional visual attention prediction method is adopted to comprehensively utilize the data of multiple modes of eye movement and head movement, and two-dimensional features are extracted using the pre-trained spherical convolution model, and behavioral and visual features are extracted through the attention length short-term memory artificial module and the residual fully connected convolution network module to perform fusion prediction.
High-precision visual attention prediction in three-dimensional space is realized, which reduces data noise, improves the accuracy of prediction results, and can be used to locate visual areas of interest and visual search paths to assist in three-dimensional spatial design evaluation.
Smart Images

Figure CN114170537B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of eye tracking, and in particular to a multimodal three-dimensional visual attention prediction method and application thereof. Background Art
[0002] Eye tracking technology obtains gaze point data by tracking eye features and mapping them to the real world or virtual screen. Multimodal fusion technology integrates information from two or more modalities to achieve information supplementation, thereby improving the accuracy of prediction results and the robustness of prediction models. Traditional eye tracking technology performs visual attention detection based on two-dimensional images and video sequences. For example, the patent applications with publication numbers CN111309138A and CN113040700A only improve the accuracy and efficiency of eye tracking based on two-dimensional images, and cannot be used for visual attention detection in three-dimensional space. Traditional eye tracking technology only performs eye tracking based on eyes or eye features. For example, the patent applications with application numbers CN111625090A and CN111417335A only focus on the processing of eye images, without multimodal data fusion. It is difficult to remove errors when there is noise interference, which affects the accuracy of prediction results.
[0003] Gaze point data can reflect the user's attention and cognitive state, and can thus be used for evaluation. Traditional three-dimensional space design evaluation methods usually use questionnaires, interviews, behavioral observations, and expert evaluation methods. These methods require recruiting a large number of subjects to obtain reliable data, which often consumes a lot of money and time, and the conclusions lack objective data support. Using multimodal visual attention to predict the visual interest area and visual search path provides information such as the user's gaze pattern and gaze focus in three-dimensional space, assisting designers in evaluating interference items and visual blind spots in three-dimensional space, which can not only improve efficiency and save costs, but also provide strong support for three-dimensional space design evaluation with objective data.
[0004] The Chinese patent document with publication number CN113177515A discloses an image-based eye tracking method, including performing face detection on the image to be detected to obtain a face detection frame; using a face key point positioning network to locate the eye area of interest and perform pupil key point positioning; and calculating the horizontal offset ratio based on the pupil center and the eye area center to determine the eye direction. This method can effectively locate the face and pupil center and perform eye tracking under conditions such as unsatisfactory ambient lighting conditions and complex backgrounds. This method also only focuses on the processing of eye images. Summary of the invention
[0005] The present invention provides a multimodal three-dimensional visual attention prediction method, which uses multimodal fusion technology to comprehensively utilize data from multiple modes of eye movement and head movement to predict visual attention, thereby improving the prediction accuracy.
[0006] The specific technical solutions adopted are as follows:
[0007] 1. A multimodal three-dimensional visual attention prediction method, comprising the following steps:
[0008] (1) Collecting the user's browsing screen and recording the user's head turning speed, head turning direction and visual gaze point when browsing the screen, wherein the browsing screen, the user's head turning speed and head turning direction are used as sample data, and the visual gaze point is used as a sample label;
[0009] (2) preprocessing the sample data, wherein the preprocessing steps are: extracting two-dimensional features of the sample data using a pre-trained spherical convolution model, and then sequentially performing timestamp alignment, data leak filling, noise cleaning, and normalization processing on the two-dimensional features to obtain preprocessed sample data; the preprocessed sample data includes head motion sample data and picture sample data;
[0010] (3) constructing a multimodal visual attention model including an attention long-term and short-term memory artificial module, a residual fully connected convolutional network module, and a fusion module; wherein the head movement sample data is input into the attention long-term and short-term memory artificial module to extract the behavioral features, and the picture sample data is input into the residual fully connected convolutional network module to extract the visual features, and the behavioral features and the visual features are fused by the fusion module to predict the attention position;
[0011] (4) Using the preprocessed sample data to train the multimodal visual attention model under the supervision of the sample labels to optimize the parameters of the multimodal visual attention model;
[0012] (5) Use a multimodal visual attention model with optimized parameters to predict and display the user's attention when browsing the screen.
[0013] Preferably, in step (1), a VR device is used to simulate a three-dimensional space, and the VR device has a sensor and a built-in eye tracker, the sensor is used to collect browsing images and record the user's head turning speed and user head turning direction when browsing the images; the built-in eye tracker is used to record the user's visual gaze point when browsing the images.
[0014] Preferably, in step (2), the spherical convolution model uses a generalized Fourier transform to project the sample data into the spectral domain, and after convolution, the two-dimensional features of the sample data are obtained by projection through an inverse Fourier transform.
[0015] Preferably, in step (2), linear interpolation is used to fill in data gaps; maximum and minimum filtering is used to clean noise; and all two-dimensional features of the sample data are normalized.
[0016] Preferably, in step (3), the residual fully connected convolutional network module includes a feature extraction module, a maximum pooling module and an average pooling module; after the image sample data is subjected to feature extraction by the feature extraction module, the obtained features are respectively input into the maximum pooling module and the average pooling module, and after the maximum pooling operation, the first visual feature is output, and after the average pooling operation, the second visual feature is output, and the first visual feature and the second visual feature are concatenated to obtain the visual feature.
[0017] Further preferably, the feature extraction module includes multiple block modules and spherical convolution layers, the block module is used to extract features of picture sample data, and the spherical convolution layer is used to process the features obtained by the block module to reduce the impact of panoramic distortion and capture deeper features through jump connections.
[0018] Preferably, in step (5), the user's browsing screen, and the user's head turning speed and head turning direction when browsing the screen are collected as test data, and the test data are pre-processed and input into a parameter-optimized multimodal visual attention model to predict and display the user's attention when browsing the screen.
[0019] The present invention also provides a method for locating a visual interest area and a visual search path, comprising the following steps:
[0020] Upload pictures from the front, back, left, right, top and bottom of the space to synthesize a panoramic image;
[0021] Capture panoramic images, and record the user's head turning speed and direction when browsing the panoramic images as test data;
[0022] The test data is preprocessed and input into the multimodal visual attention model. The coordinates of the user's attention position when browsing the panoramic image are calculated to form an attention position set. The attention position set is clustered to obtain the visual interest area, and the attention position set is timestamp-sorted to obtain the visual search path.
[0023] The present invention also provides a method for evaluating spatial information layout, comprising the following steps:
[0024] Collect the user's browsing screen, the user's head turning speed and the user's head turning direction when browsing the screen as the test data;
[0025] The test data is preprocessed and input into the multimodal visual attention model. The coordinates of the user's attention position when browsing the panoramic image are calculated to form an attention position set. The attention position set is clustered to obtain the visual interest area. The attention position set is sorted by timestamp to obtain the visual search path.
[0026] The visual search path and visual interest area are combined with the spatial design requirements to evaluate the current spatial information layout, including: when unimportant information is left in the visual interest area, it can be judged as interference information, and the interference information is moved out of the visual interest area; when important information is excluded from the visual interest area, it can be judged as easily ignored information, and the important information is moved to the visual interest area.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) The multimodal three-dimensional visual attention prediction method provided by the present invention can achieve high-precision visual attention prediction in three-dimensional space, and combine multimodal data to remove data noise, thereby further improving the accuracy of the prediction results.
[0029] (2) The multimodal three-dimensional visual attention prediction method provided by the present invention can be used to locate visual interest areas and visual search paths, and can combine the visual search paths and visual interest areas with spatial design requirements to evaluate the current spatial information layout, which can improve evaluation efficiency, save evaluation costs, and provide strong support for three-dimensional spatial design evaluation with objective data. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Flowchart of the multimodal 3D visual attention prediction method.
[0031] Figure 2 A technical roadmap for multimodal 3D visual attention prediction methods.
[0032] Figure 3 Diagram of the framework for building a multimodal visual attention model. DETAILED DESCRIPTION
[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.
[0034] like Figure 1 and Figure 2 As shown, this embodiment provides a multimodal three-dimensional visual attention prediction method, including the following steps: (1) sample data and sample label collection, (2) sample data preprocessing, (3) multimodal visual attention model construction, (4) training the multimodal visual attention model, and (5) predicting the user's attention when browsing the screen and displaying it.
[0035] (1) Sample data and sample label collection
[0036] Use VR equipment to simulate three-dimensional space, collect the user's browsing screen, and record the user's head turning speed, user's head turning direction and visual gaze point when browsing the screen. Among them, the browsing screen, user's head turning speed and user's head turning direction are used as sample data, and the visual gaze point is used as a sample label.
[0037] The VR device selected is Oculus Rift DK2, which has a sensor and a built-in Pupil Lab eye tracker. The sensor is used to collect browsing images and record the user's head turning speed and user head turning direction when the user browses the virtual reality image; the Pupil Lab built-in eye tracker is used to record the user's visual gaze point when browsing the virtual reality image.
[0038] (2) Sample data preprocessing
[0039] The sample data obtained in step (1) is preprocessed, and the preprocessing step is as follows: after extracting the two-dimensional features of the sample data using a pre-trained spherical convolution model, the two-dimensional features are sequentially timestamped, data leak-filled, noise cleaned, and normalized to obtain preprocessed sample data, wherein the preprocessed sample data includes head movement sample data (preprocessed user head turning speed and user head turning direction) and screen sample data (preprocessed browsing screen).
[0040] A pre-trained spherical convolution model is used to extract the two-dimensional features of the sample data. The spherical convolution model uses generalized Fourier transform to project the sample data into the spectral domain. After convolution, the two-dimensional features of the sample data are obtained by inverse Fourier transform projection.
[0041] The timestamps of the two-dimensional features are aligned to obtain the time series [(0, x0), (t1-t0, x1), ..., (t N -t0,x N )], where t0 is the start time, x N is time t N The corresponding eigenvalues.
[0042] Then use linear interpolation to fill in the gaps in the time series data, and use x n ,x n+2 Predict x n+1 :x n+1 =(x n +x n+1 ) / 2, n=1,2,3,…,N.
[0043] Use maximum and minimum filtering to clean noise, that is, for any x n , if x n >max,x n =max; if x n <min,xn =min; otherwise x n The max and min values are set manually.
[0044] Normalize all two-dimensional features of the sample data, and for any x n ,x n =x n / max0, max0 is all x n Then all normalized two-dimensional features are concatenated into a feature vector as the input of the multimodal visual attention model in step (3).
[0045] (3) Construction of multimodal visual attention model
[0046] A multimodal visual attention model is constructed, which includes an artificial module of long-short-term memory of attention, a residual fully connected convolutional network module and a fusion module. The head movement sample data is input into the artificial module of long-short-term memory of attention to extract the behavioral features, and the picture sample data is input into the residual fully connected convolutional network module to extract the visual features. The behavioral features and the visual features are fused by the fusion module to predict the attention position.
[0047] The Attention Long Short-Term Memory artificial module integrates the attention mechanism - calculating the match between the current input sequence and the fixation point coordinates, thereby selectively focusing on the corresponding information in the input - to capture long-range related dependency features.
[0048] In the attention long-term and short-term memory artificial module, the head movement sample data is calculated to obtain the hidden layer variable h j , hidden layer variable h j The corresponding total weight C t for: Among them, Tx is the total duration of each sample data, α tj is the hidden layer variable h j The corresponding weight, α tj The calculation formula is:
[0049]
[0050] e tj is the match between the output at time t and the input at time j, e tj =g(S t-1 ,h j ), g can be regarded as a fully connected sub-network used to learn new representations of features in the model, S t-1 is the output of the attention long short-term memory artificial module at time t-1.
[0051] In addition, multiple representations of the output of the long short-term memory artificial module are introduced into the dropout layer to improve the efficiency of model training. The dropout layer will randomly discard nodes from the network with a given probability during training, which will also reduce the generalization error of the model. Finally, the output of the residual layer will be used as the input of the residual network.
[0052] like Figure 3 As shown, in the residual fully connected convolutional network module, after the features of the picture sample data are extracted by the feature extraction module, the obtained features are respectively input into the maximum pooling module (Max pooling) and the average pooling module (Average pooling), and the first visual feature is output after the maximum pooling operation, and the second visual feature is output after the average pooling operation, and the first visual feature and the second visual feature are spliced to obtain the visual feature.
[0053] Each feature extraction module includes multiple block modules and ball convolution layers. The block module is used to extract the features of the picture sample data. The ball convolution layer is used to process the features obtained by the block module to reduce the impact of panoramic distortion and capture deeper features through jump connections.
[0054] Each block module has a residual structure formed by a spherical convolution layer and a batch normalization layer (BN), which helps to transmit features deeper in the network. It solves the network degradation problem and speeds up the network convergence; secondly, after the last layer of spherical convolution, the residual fully connected convolutional network module adds a maximum pooling layer and an average pooling layer, which helps the network learn semantic information from the input.
[0055] The residual fully connected convolutional network module is improved on the basis of the classic fully connected convolutional network. Compared with the classic fully connected convolutional network, the residual fully connected convolutional network module constructed in the present invention, which includes a feature extraction module, a maximum pooling module and an average pooling module, can better learn three-dimensional attention information, and also has better ability to recognize rotation and deformation. The residual structure directly connects the input of the previous layer to the output of the next layer using a jump. This structure reduces the risk of overfitting caused by the increase in model depth, so the entire network can try a greater depth and can process more information from lower layers. The residual fully connected convolutional network module combines the maximum pooling module and the average pooling module to improve the robustness of the model. The residual fully connected convolutional network module uses a maximum pooling module to reduce the parameters of the full connection and extract these parameters at the semantic level, reducing the variance of the estimated value and the feature extraction error caused by the limited neighborhood size. The average pooling module is used to extract more fuzzy global abstract features and reduce the estimated mean deviation caused by the convolution layer parameter error.
[0056] (4) Training a multimodal visual attention model
[0057] The processed sample data is used to train the multimodal visual attention model under the supervision of sample labels to optimize the parameters of the multimodal visual attention model.
[0058] The head movement sample data obtained in step (2) is used as the input of the attention long short-term memory artificial module, which is equipped with 640 neurons; the picture sample data is used as the input of the residual fully connected convolutional network module, which is stacked with 128, 256 and 640 filter time convolution layers. The outputs of the attention long short-term memory artificial module and the residual fully connected convolutional network module are input to the fusion module, that is, they are fused and connected through the concatenate layer of the fusion module, and the coordinates of the gaze point at the current moment are obtained through sigmoid regression.
[0059] The method of the present invention introduces regularization into the loss function of the residual fully connected convolutional network module to accelerate model training and improve the generalization ability of the model to eliminate overfitting during training.
[0060] The prediction of the user's future gaze area is defined as a classification problem, and the parameters of the multimodal visual attention model are continuously optimized until the loss converges during model training, so as to learn the mapping relationship between input and output from the training data and regress the gaze point coordinates.
[0061] The multimodal visual attention model is trained by the Adam optimizer with an initial learning rate of 1e-3, a final learning rate of 1e-4, and a batch size of 128. The learning rate is reduced by half In every 50 epochs, the validation score does not improve until the preset final learning rate is reached. The loss function is defined as:
[0062]
[0063] Among them, y i and f(x i ) represent the true value and predicted value of the i-th sample data, and m is the number of samples. Finally, training is performed on the training set and cross validation is performed to optimize the parameters of the multimodal visual attention model.
[0064] (5) Predicting the user’s attention when browsing the screen
[0065] The user's browsing screen, the user's head turning speed and the user's head turning direction when browsing the screen are collected as the test data, and the test data are preprocessed and input into the parameter-optimized multimodal visual attention model. The parameter-optimized multimodal visual attention model is used to predict the user's attention when browsing the screen and display it.
[0066] The visual interest region is generated by the parameter-optimized multimodal visual attention model, and the visual search path is obtained by connecting the visual interest regions according to the head movement direction. Using the visual interest region and the visual search path. Based on these outputs, the embodiment can analyze the following two points: (1) the browsing order of the user in processing information in the three-dimensional space, and the movement trajectory of the line of sight; (2) the user's browsing focus in the three-dimensional space, and the area where the line of sight stays for a long time; the designer can evaluate and judge whether there is interference information in the three-dimensional space, whether important information in the three-dimensional space is ignored, etc. based on the information provided.
[0067] The visual attention prediction in three-dimensional space takes panoramic images as input. Panoramic images contain all-round angles of three-dimensional space and are displayed in a spherical shape. Therefore, extracting the global and local information of the image can better capture the coarse-grained and fine-grained features of the image.
[0068] After uploading six orientation pictures of the space, namely the front, back, left, right, top and bottom, to the system, the pictures are synthesized into a panoramic image through the PTGui model, and the panoramic image is collected. The user's head turning speed and direction when browsing the panoramic image are recorded as the test data; the test data is preprocessed and input into the multimodal visual attention model. In the multimodal visual attention model constructed by the multimodal three-dimensional visual attention prediction method, the coordinates of the user's attention position when browsing the panoramic image are calculated to form an attention position set. The attention position set is clustered to obtain the visual interest area, and the attention position set is timestamped to obtain the visual search path.
[0069] After obtaining the predicted visual interest area and visual search path, the visual search path can be output as the browsing order (visual movement trajectory) of the user processing information in the three-dimensional space, and the visual interest area can be output as the browsing focus area (visual center of gravity area) where the user processes information in the three-dimensional space. Then, the visual movement trajectory and the visual center of gravity are combined with the space design requirements to evaluate the information layout of the space. When unimportant information is left in the browsing focus area, it can be judged as interference information and the interference information is moved out of the browsing focus area; when important information is excluded from the browsing area, it can be judged as easy to ignore information and the important information is moved to the visual center of gravity area.
[0070] The multimodal three-dimensional visual attention prediction method and application provided by the present invention are based on the visual attention model of the user's head turning speed, head turning direction and three-dimensional scene browsing screen, and realize the joint acquisition of multimodal user data of the built-in sensor and eye tracker of the VR helmet in a way of simulating three-dimensional space with virtual reality, and obtain a usable three-dimensional visual attention model through training of the multimodal visual attention data set to realize the prediction and evaluation of visual attention in three-dimensional space.
[0071] The present invention realizes the separate learning of head movement sample data and picture sample data through dual tributaries. The attention long short-term memory artificial module extracts local temporal features of the head movement sample data and has strong contextual text learning ability; the residual fully connected convolutional network module extracts visual features of the picture sample data, reduces the impact of panoramic distortion through splicing, and captures deeper features through jump connections.
[0072] The present invention combines multimodal data to reduce data noise and achieves high-precision three-dimensional visual attention prediction; the present invention provides a visual interest area and a visual search path for attention prediction, thereby achieving efficient, low-cost, and objective data-supported three-dimensional space design evaluation.
[0073] The embodiments described above provide a detailed description of the technical solutions of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements or similar substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multimodal three-dimensional visual attention prediction method, comprising the following steps: (1) Collecting the user's browsing screen and recording the user's head turning speed, head turning direction, and visual gaze point during browsing. The browsing screen, head turning speed, and head turning direction are used as sample data, and the visual gaze point is used as a sample label. (2) Preprocessing the sample data, wherein the preprocessing steps are as follows: extracting two-dimensional features of the sample data using a pre-trained spherical convolution model, and then sequentially performing timestamp alignment, data leak filling, noise cleaning, and normalization on the two-dimensional features to obtain preprocessed sample data; the preprocessed sample data includes head motion sample data and picture sample data; the spherical convolution model uses a generalized Fourier transform to project the sample data into a spectral domain, and after convolution, the two-dimensional features of the sample data are obtained by projection through an inverse Fourier transform; (3) constructing a multimodal visual attention model including an artificial module for long-term and short-term memory of attention, a residual fully connected convolutional network module and a fusion module; wherein the head movement sample data is input into the artificial module for long-term and short-term memory of attention to extract behavioral features, the picture sample data is input into the residual fully connected convolutional network module to extract visual features, and the behavioral features and the visual features are fused by the fusion module to predict the attention position; the residual fully connected convolutional network module includes a feature extraction module, a maximum pooling module and an average pooling module; after the picture sample data is extracted by the feature extraction module, the features obtained are respectively input into the maximum pooling module and the average pooling module, and the first visual features are output after the maximum pooling operation, and the second visual features are output after the average pooling operation, and the first visual features and the second visual features are spliced to obtain visual features; the feature extraction module includes a plurality of block modules and a ball convolution layer, the block module is used to extract the features of the picture sample data, and the ball convolution layer is used to process the features obtained by the block module; The attention long-term short-term memory artificial module integrates the attention mechanism to calculate the matching degree between the current input sequence and the fixation point coordinates; in the attention long-term short-term memory artificial module, the preprocessed head movement sample data is calculated to obtain the hidden layer variables h j , hidden layer variables h j The corresponding total weight C t for: ;in, Tx is the total duration of each sample data, α tj is the hidden layer variable h j The corresponding weight of α tj The calculation formula is: ; e tj It's time t Output and time j The matching degree between the input ,g It can be viewed as a fully connected subnetwork that is used to learn new representations in the feature remodel. S t-1 It's time t-1 The output of the artificial module of long-term short-term memory; (4) Using the preprocessed sample data to train the multimodal visual attention model under the supervision of sample labels to optimize the parameters of the multimodal visual attention model; (5) Use a multimodal visual attention model with optimized parameters to predict and display the user’s attention when browsing the screen.
2. The multimodal three-dimensional visual attention prediction method according to claim 1, characterized in that: In step (1), a three-dimensional space is simulated using a VR device, wherein the VR device has a sensor and a built-in eye tracker, wherein the sensor is used to collect browsing images and record the user's head turning speed and the user's head turning direction when browsing the images; and the built-in eye tracker is used to record the user's visual gaze point when browsing the images.
3. The multimodal three-dimensional visual attention prediction method according to claim 1, characterized in that: In step (2), linear interpolation is used to fill in data gaps; maximum and minimum filtering is used to clean noise; and all two-dimensional features of the sample data are normalized.
4. The multimodal three-dimensional visual attention prediction method according to claim 1, characterized in that: In step (5), the user's browsing screen, the user's head turning speed and the user's head turning direction when browsing the screen are collected as test data, and the test data is pre-processed and input into the parameter-optimized multimodal visual attention model to predict the user's attention when browsing the screen and display it.
5. A method for locating a visual interest region and a visual search path, characterized in that: The following steps are involved: Upload pictures from the front, back, left, right, top and bottom of the space to synthesize a panoramic image; Capture panoramic images, and record the user's head turning speed and direction when browsing the panoramic images as test data; The data to be tested are preprocessed and input into a multimodal visual attention model constructed according to the multimodal three-dimensional visual attention prediction method according to any one of claims 1-4, and the coordinates of the user's attention position when browsing the panoramic image are obtained by calculation to form an attention position set. The attention position set is clustered to obtain the visual interest area, and the attention position set is timestamp-sorted to obtain the visual search path.
6. A method for evaluating spatial information layout, characterized in that The following steps are involved: Collect the user's browsing screen, the user's head turning speed and the user's head turning direction when browsing the screen as the test data; The test data is pre-processed and input into a multimodal visual attention model constructed according to the multimodal three-dimensional visual attention prediction method described in any one of claims 1 to 4, and the coordinates of the user's attention position when browsing the panoramic image are calculated to form an attention position set, the attention position set is clustered to obtain a visual interest area, and the attention position set is timestamped to obtain a visual search path; The visual search path and visual interest area are combined with the spatial design requirements to evaluate the current spatial information layout, including: when unimportant information is left in the visual interest area, it can be judged as interference information, and the interference information is moved out of the visual interest area; When important information is excluded from the visual interest area, it can be judged as easily ignored information, and the important information is moved to the visual interest area.
Citation Information
Patent Citations
Method for reducing eyeball tracking operation and eye movement tracking device thereof
CN111309138A
Tracking movement of an eye within a tracking range
CN111417335A
Large-range eye movement tracking and sight line estimation algorithm comprehensive test platform
CN111625090A
Eye movement tracking system and tracking method thereof
CN113040700A
Image-based eye movement tracking method and system
CN113177515A