A driving ability evaluation method based on eye movement signals

By producing driving situation video materials and recording the driver's eye movement data using eye trackers, and establishing a driving capability evaluation model in combination with machine learning algorithms, the problem of insufficient training for novice drivers in the existing technology is solved, and efficient evaluation and accurate distinction of drivers' risk perception capabilities is achieved.

CN116383711BActive Publication Date: 2025-08-29NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211667555.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2025-08-29
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

The existing driver assessment system mainly focuses on theoretical knowledge and skills examinations, and lacks training on the risk perception of novice drivers in actual driving, resulting in a high proportion of traffic accidents. A more intuitive driving ability evaluation method is needed to improve drivers' safety awareness.

Method used

By producing driving situation video materials, using eye trackers to record the driver's eye movement data, combining machine learning algorithms to establish a driving ability evaluation model, evaluate the driver's risk perception ability, design visual stimulation sources and conduct feedback experiments, screen eye movement characteristics with high correlation, and construct a CNN-LSTM feature fusion network based on Attention for evaluation.

Benefits of technology

It has achieved efficient evaluation of drivers' hazard perception ability, can better distinguish between experienced and inexperienced drivers, improve the accuracy of evaluation and drivers' safety awareness, and has good application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116383711B_ABST
    Figure CN116383711B_ABST
Patent Text Reader

Abstract

This invention discloses a driving ability assessment method based on eye movement signals. First, a driving scenario video is produced, and subjects immersively watch the video to distinguish between two situations. A visual stimulus presentation and feedback experiment is designed, and the data is synchronized with the eye movement data recorded by an eye tracker. After data preprocessing, features are extracted and statistically analyzed, and correlation indicators are further screened. A machine learning algorithm is used for modeling and training to determine a benchmark. A targeted attention-based CNN-LSTM feature fusion network is designed to conduct a cross-subject assessment of driving ability, which can effectively distinguish between inexperienced and experienced drivers. Based on the theories of cognitive psychology and combining subjective and objective analysis, this invention has low investment costs and is easy to implement. The subjects have a good experience and the assessment accuracy is high, which has good application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation technology, and in particular relates to a driving ability evaluation method based on eye movement signals. Background Art

[0002] In modern life, cars have become an essential means of transportation. While these conveniences come with a host of traffic safety issues, they also raise a host of safety concerns. Among the three factors of road traffic—people, vehicles, and the environment—traffic accidents are generally caused by both subjective and objective factors. Subjective factors primarily involve the driver, while objective factors include the vehicle itself and the road environment. Extensive traffic accident data analysis and literature research indicate that drivers themselves are responsible for as much as 90% of accidents. In a complex and ever-changing traffic system, external factors are uncontrollable; only the driver's own behavior can be controlled. Therefore, driving ability plays a crucial role in overall system safety. Driver errors in any of these areas, including perception, judgment, and operation, can lead to accidents. Hazard perception is a hot topic in the current field of safe driving. A growing body of research is dedicated to studying human hazard perception in accident scenarios. This not only helps monitor driver behavior but also contributes to the improvement of assisted or automated driving systems.

[0003] Eye tracking is one of the primary ways humans perceive external information. Eye tracking technology uses auxiliary equipment such as eye trackers to record information related to visual attention. Due to its convenient acquisition method and strong signal immunity to interference, it is widely used in cognitive interaction research. Driving behavior itself is a complex cognitive process. During driving, the driver obtains information about the traffic environment through the visual system, and the brain processes this information cognitively, making judgments based on driving experience, the current road environment, and vehicle status information. Eye movement signals can more intuitively reveal the driver's visual attention. They can provide information such as gaze area, scanning trajectory, blink frequency, and pupil diameter. Using deep learning methods to study drivers' eye movement data in driving scenarios can reveal deeper connections between eye movement characteristics and driving behavior, which will greatly promote the development of safe driving.

[0004] my country's current driving test focuses more on theoretical knowledge and driving skills, lacking training for novice drivers on real-world hazard perception. Incorporating this driving ability assessment system into the driver's license exam, or conducting periodic assessments of driver ability, could significantly improve safety awareness among drivers, especially novice drivers. Through intuitive testing methods, drivers can firsthand experience potential dangers in the road environment, understand which clues to focus on, and learn to discern how potential dangers can translate into actual dangers, thereby improving their driving skills accordingly. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the present invention provides a driving ability assessment method based on eye movement signals. First, a driving scenario video is created, and subjects immersively watch the video to distinguish between two scenarios. A visual stimulus presentation and feedback experiment is designed, synchronized with the eye movement data recorded by an eye tracker. After data preprocessing, features are extracted and statistically analyzed, and correlation indicators are further screened. A machine learning algorithm is used for modeling and training to determine a benchmark. A targeted attention-based CNN-LSTM feature fusion network is designed to conduct a cross-subject assessment of driving ability, effectively distinguishing between inexperienced and experienced drivers. Based on the theories of cognitive psychology and combining subjective and objective analysis, the present invention is low-cost and easy to implement. Subjects experience a positive experience, and the assessment accuracy is high, suggesting promising application prospects.

[0006] The technical solution adopted by the present invention to solve the technical problem includes the following steps:

[0007] Step 1: Create a driving video with a specific scenario as the visual stimulus material, design a driving ability assessment simulation experiment, obtain real-time feedback from the subjects' perception results, and use an eye tracker to collect the subjects' eye movement data during the experiment to construct the subject's eye movement dataset;

[0008] Step 2: Label the subject's eye movement data, clean and filter the valid data, and perform preprocessing operations;

[0009] Step 3: Manually extract features from the preprocessed data, calculate features of pupil diameter, gaze, saccade, and blink data from the three dimensions of time, space, and frequency, and extract frequency domain features in different bands;

[0010] Step 4: Use one-way ANOVA test to perform statistical significance analysis on the frequency domain features extracted in step 3, and screen out indicators with correlations higher than the set threshold;

[0011] Step 5: Use machine learning algorithms to build models and train them, determine the optimal time window size, the training set and test set division method, and the indicator evaluation method, and establish a cross-subject classification framework;

[0012] Step 6: Establish a driving ability evaluation network model. First, build a three-layer CNN network for feature extraction and dimensionality reduction. Then, use the attention mechanism module to assign weights to the channels and use the LSTM part to process the time-related series data. Finally, classify it through the fully connected layer.

[0013] Furthermore, the visual stimulus material in step 1 is a first-person perspective video captured by a driving recorder in a vehicle following scenario, including videos of accidents and videos without accidents, and the research object is a car.

[0014] Furthermore, the driving ability evaluation simulation experiment is specifically as follows:

[0015] Step 1-1: Experimental equipment environment:

[0016] A desktop computer was used; the eye tracker was a Tobii Pro Nano with a sampling rate of 60 Hz / s. The experimental environment was a laboratory, with only the experimenter and the subject present during the experiment, with no other distracting factors. The visual stimuli were displayed on a computer screen 55–65 cm from the subject's eyes.

[0017] Step 1-2: Experimenter;

[0018] The participants had normal or corrected-to-normal vision, with no color weakness, color blindness, astigmatism, or strabismus. Those without a driver's license or with a license but no driving experience were classified as inexperienced drivers, while those with a license and previous driving experience were classified as experienced drivers, with an average driving experience of three years. The participants viewed 40 visual stimuli while connected to an eye tracker and were asked to identify dangerous situations.

[0019] Steps 1-3: Experimental steps;

[0020] 1) Before the experiment begins, the experimental procedures and requirements are introduced to the subjects, and they sign the informed consent form;

[0021] 2) The experimenter calibrates the eye tracker and performs eye calibration on the subject to confirm whether the subject is qualified as a subject;

[0022] 3) At the start of the experiment, the experimental instructions and video stimulus materials were presented through the Psychopy software, serving as the external stimulus source for the Tobii Pro Lab software. The subject selected the corresponding response button according to the prompt, and the Tobii Pro Lab software monitored the eye tracker data in real time.

[0023] 4) After the experiment, fill in the questionnaire, including the driving experience questionnaire and the driving style scale.

[0024] Furthermore, the driving eye movement data in step 1 is in a two-dimensional array format, where the first dimension of the array represents time information, and the second dimension represents eye movement information, including sampling timestamp t, gaze coordinates (x, y), and eye movement type. Each row of data in the array is called a sampling point.

[0025] Furthermore, the categories in step 2 are labeled as the driving experience of the subject, i.e., experienced = 0 and inexperience = 1, and the prediction result of the subject for each driving video is correct = 0 and wrong = 1.

[0026] Furthermore, the cleaning and screening of valid data in step 2 and the pre-processing operations include:

[0027] Step 2-1: Eliminate abnormal subject data, i.e., data where the subject's head moved during the experiment, causing the eye tracker to be unable to effectively capture pupil loss, or data where the recording was missing due to unknown software issues;

[0028] Step 2-2: Remove the data lost during the experiment, that is, the data with Eye movement type as Unclassified;

[0029] Step 2-3: For data lost due to blinking, the blink segment is filled using linear interpolation, i.e. data with Eye movement type of EyesNotFound;

[0030] Step 2-4: For missing values ​​caused by improper data collection, use the average value of the data before and after the missing value to fill it, that is, the data with Validity left or Validity right as Invalid;

[0031] Step 2-5: Normalize each dimension of the feature on the time axis, and the maximum-minimum value is normalized to: normalized value = data value - minimum value / maximum value - minimum value.

[0032] Furthermore, the frequency domain feature is a differential entropy feature DE or a power spectrum density feature PSD.

[0033] Furthermore, the training set and test set division method in step 5 is: 70% of the subjects are divided into the training set and 30% of the subjects are divided into the test set; the indicator evaluation method adopts the accuracy, precision, recall and F1-score indicators.

[0034] Furthermore, the step 6 is specifically as follows:

[0035] The first module is a CNN network, which uses three convolutional layers and two pooling layers. The convolution is one-dimensional, and the nonlinear activation function ReLU is used between the convolutional layers. Maxpooling is used to achieve downsampling processing, and the dropout function is applied to reduce overfitting. The back propagation of the convolutional neural network is used to optimize the overall network parameters.

[0036] The second module is the SE attention mechanism, which adds an attention mechanism to the eye movement feature channel dimension, assigns a weight value to each feature based on the importance of each channel, and finally outputs a sequence as the input of the subsequent LSTM;

[0037] The third module is an LSTM network, which uses LSTM to perform temporal modeling on the obtained feature map from the dimension of time steps. The number of hidden neurons is equal to the input frame length. Finally, feature fusion is completed in the fully connected layer. The loss function of the model is selected as the mean square error, and the optimization method adopts the Adam optimizer.

[0038] The beneficial effects of the present invention are as follows:

[0039] In terms of evaluating the subjects' hazard perception ability, the present invention not only combines subjective and objective aspects using driving experience questionnaires and driving video feedback for qualitative analysis, but also introduces eye tracking technology to record the subjects' eye movements during the experiment, and uses a deep learning model for in-depth analysis. It performs well in evaluating the accuracy of driving ability across subjects and has strong application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is the video material editing standard used in the embodiments of the present invention.

[0041] Figure 2 Experimental paradigm designed for embodiments of the present invention.

[0042] Figure 3 This is a network model for identifying driving experience based on driver eye movement data of the present invention.

[0043] Figure 4 These are the hazard perception ability results for different driving experiences in the embodiment of the present invention.

[0044] Figure 5 The eye movement trajectory diagram and eye movement heat map of inexperienced and experienced subjects when watching the same driving video according to an embodiment of the present invention, (a) eye movement trajectory diagram, (b) eye movement heat map. DETAILED DESCRIPTION

[0045] The present invention will be further described below with reference to the accompanying drawings and examples.

[0046] A driving ability evaluation method based on eye movement signals comprises the following steps:

[0047] Step 1: Create driving video visual stimulation materials for specific scenarios, design a driving ability assessment simulation experiment, obtain real-time feedback from the subjects' perception results, and use an eye tracker to collect the subjects' driving eye movement data during the experiment to construct the subject's eye movement dataset;

[0048] The stimulus material in step 1 consists of first-person perspective videos captured by a dashcam in a car-following scenario, including both those in which an accident occurred and those in which no accident occurred. The subject is the car, and the subject's feedback is based on the prediction of whether an accident occurred. Key interaction and eye tracking device synchronization were achieved using Psychopy. Rear-end collisions are a progressive type of hazard, involving a dynamic process of potential danger transforming into actual danger, and are a highly challenging driver cognition tool. The passive elicitation method of distinguishing between car-following and rear-end collision videos enabled the subjects to immerse themselves in the experiment.

[0049] The step 1 experiment is specifically as follows:

[0050] 1) Experimental equipment environment:

[0051] A HP desktop computer with a 23-inch monitor and a resolution of 1080*1920 was used. The eye tracker was a Tobii ProNano with a sampling rate of 60 Hz / s. The experimental environment was a laboratory. Only the experimenter and the subject were present during the experiment, with no other distractions. The temperature and lighting were appropriate, and participants were instructed to mute their electronic communication devices. Participants sat in a relatively relaxed position throughout the experiment. Video stimuli were displayed on a computer screen approximately 55–65 cm from the subject's eyes. Throughout the experiment, participants were instructed to avoid significant head movements, except when completing the questionnaire, to minimize the loss of eye movement data.

[0052] 2) Experimenters:

[0053] The participants were healthy and mentally healthy university students with normal or corrected-to-normal vision and no color deficiency, color blindness, astigmatism, or strabismus. All participants were approximately 25 years old. Those without a driver's license or with a license but no driving experience were classified as inexperienced drivers, while those with a license and previous driving experience were classified as experienced drivers. The average driving experience was three years. The participants watched 40 hazard perception videos while connected to an eye tracker and asked to identify hazardous situations. The driving abilities of young, experienced, and inexperienced participants of the same age group were compared.

[0054] 4) Experimental steps:

[0055] 4-1) Before the experiment begins, the experimental procedures and requirements are introduced to the subjects, and they sign the informed consent form;

[0056] 4-2) The experimenter calibrates the eye tracker and other equipment, and performs eye calibration on the subject to confirm whether the subject is qualified as a subject;

[0057] 4-3) At the start of the experiment, the experimental instructions and video stimulus materials were presented via Psychopy software, serving as the external stimulus source for the Tobii Pro Lab. The subject selected the corresponding response button according to the prompts, and the Tobii Pro Lab monitored the eye tracker data in real time.

[0058] 4-4) After the experiment, participants will complete the questionnaires, including the driving experience questionnaire and the driving style scale, and will be paid the experimental fee.

[0059] Furthermore, the eye movement data in step 1 is in a two-dimensional array format, where the first dimension of the array represents time information and the second dimension represents eye movement information, mainly including sampling timestamp (t), gaze coordinates (x, y), and eye movement type. Each row of data in the array is called a sampling point.

[0060] Step 2: Label the original experimental data, clean and filter the valid data, and perform preprocessing operations;

[0061] The label category is annotated as the subject's driving experience (experienced - 0) and inexperienced - 1, and the subject's prediction result for each driving video is correct - 0 and wrong - 1.

[0062] Step 3: Manually extract features from the processed data. Calculate features of pupil diameter, fixation, saccade, and blink data from the three dimensions of time, space, and frequency. Also, extract frequency domain features such as differential entropy (DE) and power spectral density (PSD) in different bands.

[0063] Step 4: Use one-way ANOVA to conduct statistical significance analysis on the initially extracted eye movement features and screen out indicators with high correlation;

[0064] In theory, data must pass normality test and homogeneity of variance test before parametric experience can be performed, otherwise non-parametric test is used. However, in actual application scenarios, this may not be necessary.

[0065] Step 5: Use machine learning algorithms to build models and train, determine the optimal time window size, training set and test set division method, and evaluate indicators, etc., to establish a cross-subject classification framework;

[0066] Step 6: Establish a driving ability evaluation network model. First, build a three-layer CNN network for feature extraction and dimensionality reduction. Then, use the attention mechanism module to assign weights to channels. Finally, use the LSTM part to process time-related series data, and finally classify it through the fully connected layer. Specific embodiment:

[0068] The present invention provides a driving ability assessment method based on eye movement signals. By having young driving experience and inexperienced subjects watch accident and non-accident videos of specific driving scenarios, a sample driving experience database and driving eye movement dataset are established. By building a deep learning model, the connection between eye movement behavior and driving experience, as well as the impact on driving hazard perception ability, is explored. Eye movement signal characteristics are used to achieve cross-subject driving ability assessment, providing a new idea and method for research in the field of safe traffic driving.

[0069] 1. Design a driving ability evaluation experiment to collect eye movement data from drivers with different experience levels. The details are as follows:

[0070] 1-1) Prepare the experimental visual stimuli. The experimental design consisted of a within-group factor (video clips of accidents and non-accidents) and a between-group factor (participants with driving experience and those without). The dependent variable was the participants' accuracy in identifying potential and actual hazards. To collect as many real-world driving accident videos as possible, we searched nearly all public datasets and mainstream video websites, such as AcFUNs, YouTube, Youku, and Bilibili. Because first-person accident scenes are relatively rare, we set the driving scenario to car-following scenarios to facilitate a control group. Furthermore, we selected 20 video clips of typical car-following accident types. These served as an experimental group. These videos were all first-person footage, captured by a dashcam from the driver's perspective, rather than from a third-person perspective. Correspondingly, 20 video clips of non-car-following accidents were selected as a control group. The video resolution was 1920*1080, and the frame rate for all videos was 30 frames per second.

[0071] Video editing process such as Figure 1For an accident video, we need to record the frames at the moment of the accident and divide the video into three parts: the pre-accident frame, the accident frame, and the post-accident frame. The video used in the experiment is the accident prediction portion. The editing criteria are based on the analysis of driver reaction characteristics in the "Traffic Engineering": Under high concentration, the minimum reaction time required before the driver begins braking is 0.4 seconds, and the braking effect requires 0.3 seconds. In total, the minimum total time for the braking effect is 0.7 seconds. Because the subjects of this experiment are inexperienced and young drivers, several volunteers were tested using the original unedited video. They were instructed to immediately press the pause button when they perceived danger. The reaction time (i.e., the impact time minus the key press time) was recorded and the average reaction time was calculated to be approximately 1 second. Therefore, the accident prediction portion is the portion before the accident frame but does not include the reaction time. The length is controlled at around 10 seconds, and the clarity is suitable. The non-accident control group experimental video clips serve as filler to prevent the experimenters from always perceiving the driving video clips as dangerous. The editing criteria for the non-accident video clips are that they do not have obvious characteristics similar to the accident video clips.

[0072] The videos were stripped of distracting watermarks and time information, effectively avoiding irrelevant interference. The order of the videos, including those showing accidents and those showing no accidents, was randomly shuffled. The videos covered various scenarios (highways, cities, rural areas, and tunnels), weather conditions (sunny, rainy, and snowy), and lighting conditions (daytime and nighttime). The videos also showed signs of danger, including slowing down the preceding vehicle, activating brake lights, emergency stops, lane changes, steering, and acceleration. If the subject noticed the correct clues before the video ended, they could correctly predict what would happen next.

[0073] 1-2) Design experimental paradigm, such as Figure 2 Before the experiment officially began, each participant was informed of the experimental procedures in detail. After the participant understood all the details of the experiment and had their questions answered, they voluntarily signed a written informed consent form and confidentiality agreement.

[0074] The experimental environment was a laboratory. During the experiment, only the experimenter and the subjects were present. Participants were required to mute their electronic communication devices and reduce the lighting in the laboratory to reduce the influence of the surrounding environment so that they could focus only on the computer screen. The subjects participated in the entire experiment in a relatively relaxed sitting position. The experimenter sat 60 cm away from the screen and conducted

[0075] Eye tracker equipment calibration: After screening for visual impairments, the experimenter calibrated the Tobii eye tracker using a 9-point calibration procedure. The calibration procedure was used between each video clip to recalibrate the central fixation point.

[0076] During the test, try to avoid large head movements except when completing the questionnaire.

[0077] The experimenter prepared two practice clips to watch. If the participant did not meet the expected performance, the requirements would be further explained. When the experimenter was in the formal experiment, the screen would first show a relevant introduction to the experiment.

[0078] After the introduction, the first stage of the experiment will begin. Three landscape pictures that are not related to this experiment will be randomly displayed on the screen, and each display will be 10 seconds. Before each picture is displayed, a fixation point will appear on the screen to remind the user that the image is about to be displayed.

[0079] The data recorded in this stage is marked as relax and will not be used as experimental data. Its main function is to allow the experimenters to relax and enter the experimental state in advance. After the landscape image is finished, the second stage of the experiment will begin.

[0080] In the second stage, the screen will randomly play real driving videos shot by the dashcam from the driver's perspective.

[0081] There are some videos where accidents have occurred and some where no accidents have occurred. The subjects need to concentrate on watching the videos. Each video is presented for about 10 seconds. The video will stop in advance at an appropriate point. After each video is played, the subjects will speculate whether it is related to the accident in front of the car.

[0082] The subjects were asked to press the corresponding keyboard letter to select the correct answer. After the subjects finished answering, the next video was played. Markers were set to indicate the start of the experiment, the presentation of the landscape image, each fixation point, the video playback, and the end of the experiment.

[0083] After the experiment, the participants were required to fill out a driving experience questionnaire, which included details about their driving experience and accident history, and were given experimental compensation.

[0084] 2. Pre-process the driving eye movement data collected in the experiment and filter out the effective eye movement data. Since the original eye movement data itself contains a large number of abnormal values ​​and default values, it is necessary to perform targeted feature extraction before formal feature extraction.

[0085] Data cleaning is carried out in a targeted manner and is divided into the following categories according to different categories:

[0086] a. Eliminate abnormal subject data, i.e., data where the subject's head moved significantly during the test, causing the eye tracker to be unable to effectively capture pupil loss data, or data where missing data was recorded due to unknown software issues;

[0087] b. Remove the data that was lost during the experiment, that is, the data with Eye movement type as Unclassified;

[0088] c. Data lost due to blinking: the blink segment is filled using linear interpolation, i.e. data with Eye movement type set to EyesNotFound;

[0089] 4) Missing values ​​due to improper lab collection are filled with the average value of the data before and after the missing value, that is, the data with Validity left or Validity right as Invalid.

[0090] Since the original eye movement data has large numerical variations, in order to eliminate the differences, each dimension of the feature is normalized on the time axis, and the maximum-minimum value is normalized to: normalized value = data value - minimum value / maximum value - minimum value.

[0091] 3. Feature extraction was performed on the preprocessed eye movement data. The raw eye tracking data was directly sampled using an eye tracker to obtain time-stamped gaze data. This data was large and computationally intensive, making its significance unclear. The experiment employed a manual feature extraction method, calculating features for four types of data: pupil diameter, gaze, saccade, and blink, along the three dimensions of time, space, and frequency. Frequency domain features such as differential entropy (DE) and power spectral density (PSD) were extracted from different bands, totaling 62 dimensions of eye movement features. Secondly, to address the issues of small sample sizes and data alignment, a sliding window truncation data augmentation method was employed, splitting the raw eye movement data into several sub-data sequences as input. It has been verified that eye movement behavior is a process with fluctuations, and the truncation method can better capture eye movement features over a certain period of time.

[0092] 4. Statistical methods were used to perform a significance analysis on the initially extracted eye movement features to identify highly correlated indicators. Since many different eye tracking signal indicators were initially extracted, it was unclear whether all of them were cognition-related. Therefore, a significance analysis was performed on these eye movement signals to verify whether any of them differed significantly between different driving experience categories. One-way analysis of variance (ANOVA) tests the null hypothesis that the samples in two or more groups are drawn from distributions with the same mean. ANOVA calculates an F-statistic, which is the ratio of the between-group sample variance to the within-group sample variance. A larger ratio indicates that the samples in different groups are drawn from distributions with different means. The F-statistic can be used to generate a P value. If P < 0.05, the null hypothesis is rejected, indicating that the samples in two or more groups are drawn from distributions with different means, indicating that the differences between the means of the multiple groups are significant. Otherwise, the null hypothesis is accepted, indicating that the differences in the means are not significant.

[0093] A one-way ANOVA of all the aforementioned eye movement signal features revealed that the P values ​​for all 39 features, including pupil diameter, fixation duration, fixation deviation, blink frequency, saccade duration, saccade amplitude, and blink duration, were all less than 0.05, indicating that these eye movement signal features showed significant differences across different experience states. The accuracy of driving experience classification using these 39 features was actually higher than that using the 62 features, indicating that the 62 eye movement features contain features unrelated to cognitive state. Therefore, these 39 statistical event features were selected for the subsequent classification task.

[0094] 5. Model and train the eye movement features selected after feature correlation analysis using machine learning algorithms. The optimal time window size, training set / test set division method, and metric evaluation were determined to establish a cross-subject classification framework. Regarding sample partitioning, to test the model's generalization ability for unknown sample classification, 70% of the subjects were assigned to the training set and 30% to the test set. Accuracy, precision, recall, and F1-score were used to measure model performance.

[0095] When selecting the data range for normalization, if the entire dataset or all subjects are used as the data range for single normalization, the large differences in feature values ​​between subjects will make the influence of driving experience less prominent, and the classification results will be less than ideal. If the feature values ​​corresponding to each feature attribute of a single subject in a single experiment are used as the data for single normalization, there will be no mutual influence between subjects and feature values, and the resulting normalization results will better preserve the internal characteristics of the original data.

[0096] 6. Establish a driving ability evaluation network model. First, build a three-layer CNN network for feature extraction and dimensionality reduction. Then, use the attention mechanism module to assign weights to the channels. Finally, use the LSTM part to process the time-related sequence data. Finally, use the fully connected layer for classification. Figure 3 As shown in the figure, the raw eye movement data obtained from preprocessing is concatenated and fused with the eye movement features. The fused features are then divided into training and test sets according to an appropriate ratio across subjects. These serve as the input data for the network model. The model is constructed by analyzing the characteristics of the existing dataset and the specific application scenario. Each input data is in the format of (chanel, size), where chanel represents the dimension of the eye movement feature and size represents the 3s*60hz input data at once, i.e., the window size.

[0097] The first module is a CNN network with three convolutional layers and two pooling layers. The convolutions are one-dimensional, and the nonlinear activation function ReLU is used between convolutional layers to better normalize the dataset and reduce exploding or vanishing gradients. Maxpooling is used for downsampling, and dropout is applied to reduce overfitting. The convolutional neural network's backpropagation is used to optimize the overall network parameters. The second module is the SE attention mechanism, which adds attention to the eye movement feature channel dimension. The key operations are squeeze and excitation. Each feature is assigned a weight based on its importance, allowing the network to focus on certain feature channels. The final output is a sequence that serves as the input for the subsequent LSTM. The third module is an LSTM network. The feature map is modeled in the time step dimension using the LSTM. The number of hidden neurons is equal to the input frame length. Finally, the fully connected layer completes feature fusion. The model's loss function is mean squared error, and the Adam optimizer is used for optimization.

[0098] The results show that the model uses eye movement behavior characteristics to distinguish between experienced and inexperienced drivers in the hazard perception ability assessment with an accuracy of 83.29%, a precision of 77.02%, a recall rate of 84.87%, and an F1-score of 70.50%, proving that this method can effectively distinguish between inexperienced and experienced drivers.

[0099] 7. The total prediction error rate of experienced drivers in all videos is significantly lower than that of inexperienced drivers, e.g. Figure 4 As shown in Figure 2, it shows that experienced drivers have better hazard perception ability than inexperienced drivers, and can detect the process of potential danger turning into actual danger earlier and more accurately. Figure 5 As shown in the figure, the eye movement heat map and gaze map of drivers with different experience levels when watching the same driving video show that experienced drivers tend to keep their eyes farther away from the road rather than in front of the car, and their eye movement search range is wider, and they can assist in judgment by observing clues in the surrounding environment. This overall perception can help experienced drivers prevent possible accidents.

Claims

1. A driving ability evaluation method based on eye movement signals, characterized in that: The following steps are involved: Step 1: Create a driving video with a specific scenario as the visual stimulus material, design a driving ability assessment simulation experiment, obtain real-time feedback from the subjects' perception results, and use an eye tracker to collect the subjects' eye movement data during the experiment to construct the subject's eye movement dataset; Step 2: Label the subject's eye movement data, clean and filter the valid data, and perform preprocessing operations; Step 3: Manually extract features from the preprocessed data, calculate features of pupil diameter, gaze, saccade, and blink data from the three dimensions of time, space, and frequency, and extract frequency domain features in different bands; Step 4: Use one-way ANOVA test to perform statistical significance analysis on the frequency domain features extracted in step 3, and screen out indicators with correlations higher than the set threshold; Step 5: Use machine learning algorithms to build models and train them, determine the optimal time window size, the training set and test set division method, and the indicator evaluation method, and establish a cross-subject classification framework; Step 6: Build a driving ability evaluation network model. First, construct a three-layer CNN network for feature extraction and dimensionality reduction. Then, use the attention mechanism module to assign weights to channels and use the LSTM part to process time-related series data. Finally, use the fully connected layer for classification. The first module is a CNN network, which uses three convolutional layers and two pooling layers. The convolution is one-dimensional, and the nonlinear activation function ReLU is used between the convolutional layers. Maxpooling is used to achieve downsampling processing, and the dropout function is applied to reduce overfitting. The back propagation of the convolutional neural network is used to optimize the overall network parameters. The second module is the SE attention mechanism, which adds an attention mechanism to the eye movement feature channel dimension, assigns a weight value to each feature based on the importance of each channel, and finally outputs a sequence as the input of the subsequent LSTM; The third module is an LSTM network, which uses LSTM to perform temporal modeling on the obtained feature map from the dimension of time steps. The number of hidden neurons is equal to the input frame length. Finally, feature fusion is completed in the fully connected layer. The loss function of the model is selected as the mean square error, and the optimization method adopts the Adam optimizer.

2. The driving ability evaluation method based on eye movement signals according to claim 1, characterized in that: The visual stimulus material in step 1 is a first-person perspective video captured by a driving recorder in a vehicle-following scenario, including videos of accidents and videos without accidents, and the research object is a car.

3. The driving ability evaluation method based on eye movement signals according to claim 1, characterized in that: The driving ability evaluation simulation experiment is as follows: Step 1-1: Experimental equipment environment: A desktop computer was used; the eye tracker was a Tobii Pro Nano with a sampling rate of 60 Hz / s. The experimental environment was a laboratory, with only the experimenter and the subject present during the experiment, with no other distracting factors. The visual stimuli were displayed on a computer screen 55–65 cm from the subject's eyes. Step 1-2: Experimenter; The participants had normal or corrected-to-normal vision, with no color weakness, color blindness, astigmatism, or strabismus. Those without a driver's license or with a license but no driving experience were classified as inexperienced drivers, while those with a license and previous driving experience were classified as experienced drivers, with an average driving experience of three years. The participants viewed 40 visual stimuli while connected to an eye tracker and were asked to identify dangerous situations. Steps 1-3: Experimental steps; 1) Before the experiment begins, the experimental procedures and requirements are introduced to the subjects, and they sign the informed consent form; 2) The experimenter calibrates the eye tracker and performs eye calibration on the subject to confirm whether the subject is qualified as a subject; 3) At the start of the experiment, the experimental instructions and video stimulus materials were presented through the Psychopy software, serving as the external stimulus source for the Tobii ProLab software. The subject selected the corresponding response button according to the prompt, and the Tobii Pro Lab software monitored the eye tracker data in real time. 4) After the experiment, fill in the questionnaire, including the driving experience questionnaire and the driving style scale.

4. The driving ability evaluation method based on eye movement signals according to claim 1, characterized in that: The driving eye movement data in step 1 is in a two-dimensional array format, where the first dimension of the array represents time information, and the second dimension represents eye movement information, including sampling timestamp t, gaze coordinates (x, y), and eye movement type. Each row of data in the array is called a sampling point.

5. The driving ability evaluation method based on eye movement signals according to claim 1, characterized in that: In step 2, the categories are labeled as experienced (0) and inexperienced (1), and the subject's prediction results for each driving video are correct (0) and incorrect (1).

6. The driving ability evaluation method based on eye movement signals according to claim 1, characterized in that: In step 2, the cleaning and screening of valid data and the pre-processing operations include: Step 2-1: Eliminate abnormal subject data, i.e., data where the subject's head moved during the experiment, causing the eye tracker to be unable to effectively capture pupil loss, or data where the recording was missing due to unknown software issues; Step 2-2: Remove the data lost during the experiment, that is, the data with Eye movement type as Unclassified; Step 2-3: For data lost due to blinking, the blink segment is filled using linear interpolation, i.e. data with Eye movement type of EyesNotFound; Step 2-4: For missing values ​​caused by improper data collection, use the average value of the data before and after the missing value to fill it, that is, the data with Validity left or Validity right as Invalid; Step 2-5: Normalize each dimension of the feature on the time axis, and the maximum-minimum value is normalized to: normalized value = data value – minimum value / maximum value – minimum value.

7. The driving ability evaluation method based on eye movement signals according to claim 1, characterized in that: The frequency domain feature is a differential entropy feature DE or a power spectrum density feature PSD.

8. The driving ability evaluation method based on eye movement signals according to claim 1, characterized in that: The training set and test set division method in step 5 is: 70% of the subjects are divided into the training set and 30% of the subjects are divided into the test set; the indicator evaluation method adopts the accuracy, precision, recall and F1-score indicators.

Citation Information

Patent Citations

  • Driver eye movement glancing-based safe driving recommendation method in urban environment

    CN113569733A

  • Driving evaluation method and system

    CN113743471A