A method for constructing Ti-SA model for pilot situational awareness assessment
By building a Ti-SA model that integrates multi-head attention mechanism and heuristic modules, using the pilot's eye tracking and AOI gaze data, the pilot's situational awareness is evaluated in real time, and the problem of inability to evaluate and adapt to complex environments in real time is solved in the existing technology, and the accuracy and efficiency of the evaluation are improved.
Patent Information
- Application Number
- CN202411229154.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-09-03
AI Technical Summary
The prior art cannot evaluate the pilot's situational awareness ability in real time, and the effectiveness of selecting data in complex environments is poor.
A Ti-SA model integrating multi-head attention mechanism and heuristic module was constructed. By obtaining pilot's eye track timing data and AOI gaze timing data, the pilot's situational awareness level is evaluated in real time.
Real-time assessment of pilot situation awareness abilities is achieved, the accuracy and efficiency of evaluation results are improved, the risk of flight accidents is reduced, and personalized evaluation results are provided.
Smart Images

Figure CN119202587B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aviation safety and situational awareness, and in particular to a method for constructing a Ti-SA model for pilot situational awareness assessment. Background Art
[0002] Aviation technology has made great progress, and the performance, function and complexity of aircraft have been significantly improved. However, flight safety is still an urgent problem to be solved. The pilot's situational awareness (SA) ability is the pilot's attitude perception and motion perception of time, space and direction in a high-altitude closed cockpit environment. It is an important factor affecting the pilot's flight decision-making and flight safety. The higher the situational awareness ability, the more timely, accurate and comprehensive the pilot can grasp the flight situation, so as to make better flight decisions and operations. On the contrary, low or loss of situational awareness will lead to cognitive bias, misjudgment or ignorance of the flight situation by the pilot, thereby increasing the risk of flight accidents. The pilot's situational awareness ability is crucial to flight safety. Studies have shown that 15%-20% of accidents are related to the pilot's lack of situational awareness. Spatial Disorientation (SD) is closely related to situational awareness SA. Pilots need to maintain good spatial orientation ability to ensure flight safety. SD is the inability or failure of directional judgment relative to the coordinate system composed of the ground and the vertical line of gravity. SD is closely related to SA ability. Specifically, the loss of SA ability by the pilot is inevitably accompanied by SD. Therefore, it is very important to build a system that can objectively evaluate the pilot's situational awareness during flight.
[0003] Pilot situational awareness assessment models constructed by existing technologies usually rely on traditional data analysis methods and subjective evaluation systems. These models are constructed based on questionnaire surveys, expert observations or simple behavioral indicator analysis methods. The sample data is single and relies on the subjective judgment of experts, rather than on the real-time attitude and motion perception of pilots during flight training. The assessment is also conducted after the training is completed, which is lagging and inefficient. It is not suitable for large-scale motion capture and attitude capture, as well as real-time situational awareness assessment. Summary of the invention
[0004] In view of the above analysis, an embodiment of the present invention aims to provide a method for constructing a Ti-SA model for pilot situational awareness assessment, so as to solve the technical problems that the existing perception model cannot perform real-time assessment of the pilot's situational awareness ability and the selected data cannot adapt to complex environments, resulting in poor effectiveness.
[0005] The purpose of the present invention is mainly achieved through the following technical solutions:
[0006] The present invention provides a method for constructing a Ti-SA model for pilot situational awareness assessment, comprising the following steps:
[0007] Step S1, obtaining eye movement tracking time series data and AOI gaze time series data of multiple pilots during simulated flight as sample data, wherein the sample data and the corresponding generated category labels constitute a training data set; wherein the category label is a level label of the pilot's situational awareness;
[0008] Step S2, constructing a Ti-SA model for pilot situational awareness assessment; using the training data set to train the Ti-SA model until a trained Ti-SA model is obtained when the loss function converges.
[0009] Further, the eye movement tracking time series data includes two-dimensional coordinates of gaze focus, three-dimensional coordinates of gaze focus, three-dimensional coordinates of left eye original position, three-dimensional coordinates of left eye moving direction, left eye pupil diameter, three-dimensional coordinates of right eye original position, three-dimensional coordinates of right eye moving direction and right eye pupil diameter;
[0010] The AOI gaze timing data includes a number sequence of AOI subdivision areas that the pilot gazes at in sequence based on time during flight; the AOI subdivision areas include A sub-areas divided based on instrument display areas and screen display areas, and exterior scenery areas.
[0011] Furthermore, the eye movement tracking time series data and AOI gaze time series data of the pilot during the simulated flight are obtained as sample data, including:
[0012] Obtaining the eye movement tracking time series data and the flight training experiment video;
[0013] Based on the flight experiment video detection, all AOI subdivision areas in the flight experiment video in time sequence are obtained;
[0014] Based on the two-dimensional and three-dimensional coordinates of the gaze focus in the eye tracking time series data, it is determined whether the coordinates are located in a certain AOI subdivision area. According to the matching results, the AOI subdivision area numbers of the pilot's gaze are recorded in chronological order to obtain the AOI gaze time series data.
[0015] Furthermore, the following process is performed to generate the corresponding category labels:
[0016] Based on the SA assessment items, measurement indicators, recovery results, scheduled results, and the full score of the SA assessment items, the assessment result Score of each pilot's SA assessment item is calculated as follows:
[0017]
[0018] Where n is the total number of indicators of SA evaluation items, D i M is the deviation between the pilot's recovery result and the expected result. i is the maximum deviation value in the indicator, fullscore i It is the full score obtained by the pilot when the recovery result of the indicator corresponding to the SA assessment item is completely consistent with the expected result;
[0019] The level corresponding to the evaluation result is used as the corresponding category label of the eye tracking time series data and AOI gaze time series data obtained during the flight simulation.
[0020] Furthermore, the Ti-SA model for pilot situational awareness assessment includes an eye tracking multi-dimensional attention enhancement module, an AOI multi-scale comprehensive feature module, a fusion module, and a SA level output module;
[0021] The eye movement trace multi-dimensional attention enhancement module is used to receive the eye movement trace time series data, perform multi-dimensional attention feature extraction, and output the category probability value corresponding to the eye movement trace time series data;
[0022] The AOI multi-scale comprehensive feature module is used to receive the AOI gaze time series data, perform multi-scale feature fusion, and output the category probability value corresponding to the AOI fusion feature;
[0023] The fusion module fuses the eye tracking depth feature and the AOI fusion feature to obtain a predicted situational awareness level result;
[0024] The SA level output module is used to output the predicted situational awareness level result.
[0025] Furthermore, the eye movement tracking multi-dimensional attention enhancement module includes:
[0026] A first linear projection layer, used to map the eye movement tracking time series data features to a high dimension;
[0027] The first random deactivation layer is used to perform random deactivation processing on the mapped high-dimensional time series data to enhance the generalization of the eye movement trace multi-dimensional attention enhancement module and prevent overfitting;
[0028] A first, a second and a third multi-layer attention structure connected in sequence, for obtaining an encoded representation of the eye tracking temporal data;
[0029] Bi-LSTM (Bi-directional Long Short-Term Memory), which uses the bidirectional capability of LSTM to enhance the understanding of the time series of the encoded representation output after three multi-layer attention structures;
[0030] Adaptive average pooling layer to reduce the dimensionality of Bi-LSTM output features while retaining important information;
[0031] The second linear projection layer is used to match the output dimension and obtain three category probability values corresponding to the eye movement trace.
[0032] Further, the AOI multi-scale comprehensive feature module includes six sequentially connected first to sixth heuristic modules, a global average pooling layer, a third random inactivation layer and a third linear projection layer;
[0033] The first to sixth heuristic modules are used to perform multi-scale feature extraction on the AOI gaze time series data; wherein the AOI gaze time series data is connected to the output of the second heuristic module through a first residual; the output of the third heuristic module is connected to the output of the fifth heuristic module through a second residual;
[0034] The global average pooling layer is used to reduce the dimension of the output features of the first heuristic module and extract global features;
[0035] The third random deactivation layer is used to perform random deactivation processing on the extracted global features to enhance the generalization ability of the AOI multi-scale comprehensive feature module;
[0036] The third linear projection layer is used to reduce the dimension of the matching output dimension to obtain three category probability values corresponding to the AOI fusion feature.
[0037] Furthermore, the multi-head attention module in the three-layer attention structure splits the input features of the module into multiple attention heads. Each head processes the feature information independently and understands the features from different angles. Finally, the outputs of multiple attention heads are merged to obtain a comprehensive feature representation.
[0038] Further, the first to sixth heuristic modules each include a first component, a second component and an output layer, as follows:
[0039] The first component is the first bottleneck layer, which uses a filter with a length of 1 to reduce the dimension of the input multi-dimensional AOI gaze time series data;
[0040] The second component includes three parallel convolutions, a max pooling, and a second bottleneck layer;
[0041] The three parallel convolutions are used to perform convolution operations of three lengths on the output of the first bottleneck layer to capture temporal features of different scales; wherein the three convolutions of different lengths are one-dimensional convolutions of different lengths, and the convolution kernels are 15, 25 and 50 respectively;
[0042] The maximum pooling is used to perform a maximum pooling operation on the input multi-dimensional AOI gaze time series data;
[0043] The second bottleneck layer is used to reduce the dimension of the maximum pooling output;
[0044] The output layer is used to concatenate the output features of each parallel convolution and maximum pooling to form a final output multi-dimensional time series.
[0045] Further, the fusion module averages the probability values of the three categories corresponding to the eye tracking time series data and the three category probability values corresponding to the AOI fusion feature to obtain an averaged probability value;
[0046] Comparing the averaged probability values of the three categories to determine the category with the highest probability;
[0047] The category with the largest average probability value is taken as the pilot's SA level status.
[0048] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0049] 1. The present invention constructs a Ti-SA model that integrates a multi-head attention mechanism and a heuristic module. During the model construction process, the long-term dependency relationship in the eye tracking time series data and the local short-term dependency features in the AOI gaze sequence data are simultaneously learned. Compared with the existing method of evaluating after the experiment, the Ti-SA model is designed for real-time or near real-time evaluation, reducing the delay time of the evaluation. The model evaluates the pilot's situational awareness ability in real time, so that the evaluation results are fed back to the pilot or the training system in real time, so as to adjust the training or operation strategy in time;
[0050] 2. The present invention solves the technical problem that the selected data in the existing model cannot adapt to the complex environment and leads to poor effectiveness by combining the pilot's eye movement tracking time series data and AOI gaze time series data, and provides a more comprehensive pilot situational awareness ability assessment, which helps to more accurately capture the pilot's cognitive state of the flight environment, thereby improving the accuracy of the assessment results;
[0051] 3. The model building method of the present invention includes a complete process from data collection, feature extraction, model training to evaluation level output. The optimization of this process helps to simplify the actual operation of pilot situational awareness evaluation, improve the efficiency and practicality of the evaluation process, and thus optimize the pilot training and situational awareness evaluation process;
[0052] 4. The Ti-SA model of the present invention uses a multi-attention mechanism, adaptive average pooling, and multi-scale feature extraction of a heuristic module. The Ti-SA model can adapt to different flight phases and the cognitive characteristics of different pilots and provide personalized evaluation results;
[0053] 5. Through more accurate and real-time situational awareness assessment of pilots during flight, the present invention helps to identify possible cognitive deficiencies or errors of pilots at an early stage, so as to take preventive measures, reduce flight accidents caused by insufficient situational awareness, and improve overall flight safety; the Ti-SA model can not only be used to assess the SA ability level of pilots, but also assist in pilot training. By analyzing the SA ability level during the training process, customized training suggestions can be provided to pilots to help them improve their flying skills and ability to cope with complex situations.
[0054] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like components throughout the drawings.
[0056] Figure 1 is a flow chart of a method for constructing a Ti-SA model for pilot situational awareness assessment in an embodiment of the present invention;
[0057] Figure 2 Schematic diagram of the overall structure of the Ti-SA model in an embodiment of the present invention;
[0058] Figure 3 is a flow chart of a method for acquiring eye movement tracking time series data and AOI gaze time series data in an embodiment of the present invention;
[0059] Figure 4 This is a schematic diagram of the pilot's eye movement traces in a spatial directional flight simulation state in an embodiment of the present invention - a schematic diagram of the pilot's gaze focus above the runway;
[0060] Figure 5 This is a schematic diagram of the pilot's eye movement traces in a spatial directional flight simulation state in an embodiment of the present invention - a schematic diagram of the pilot's gaze focus on the instrument panel area;
[0061] Figure 6 Schematic diagram of AOI area division of real scene in the cabin in an embodiment of the present invention;
[0062] Figure 7 Schematic diagram of the heuristic module structure in an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0064] In view of the problems of the current existing technology, this paper constructs a data set based on deep learning, and uses deep learning methods to perform SA in the pilot's spatial orientation flight simulation state, proposes a highly accurate Ti-SA model, and improves the evaluation accuracy by fusing eye tracking time series data and AOI (Area of Interest) gaze time series data. It is used for SA evaluation in the pilot's spatial orientation flight simulation state as a preliminary exploration in this field.
[0065] A specific embodiment of the present invention discloses a method for constructing a Ti-SA model for pilot situational awareness assessment, such as Figure 1 As shown, the following steps are included:
[0066] Step S1, obtaining eye movement tracking time series data and AOI gaze time series data of multiple pilots during simulated flight as sample data, wherein the sample data and the corresponding generated category labels constitute a training data set; wherein the category label is a level label of the pilot's situational awareness;
[0067] Step S2, constructing a Ti-SA model for pilot situational awareness assessment; using the training data set to train the Ti-SA model until a trained Ti-SA model is obtained when the loss function converges.
[0068] The test equipment involved in the present invention is as follows:
[0069] (1) Flight simulator: The DISO space orientation flight simulator produced by AMST of Austria is used. It uses a three-dimensional 6-DOF hydraulic platform and a highly simulated cockpit vision system. It can simulate more than 10 kinds of vestibular and visual illusions to intuitively experience the various SDs generated by the movement of the human body's orientation organs when facing an abnormal three-dimensional space environment.
[0070] (2) Wearable eye tracker: The Tobii Pro Glasses 3 eye tracker produced by Swedish company Tobii is used. It has a sampling rate of 100 Hz, integrates 16 light sources and 4 eye tracking sensors, and is embedded in scratch-resistant lenses. The eye tracker has a 106° wide-angle scene camera that can record behavioral details from a first-person perspective.
[0071] The test conditions involved in the present invention are as follows: the test sequentially conducts the following simulated flight links: at an altitude of 914m and a speed of 400 knots, the pitch illusion is experienced at an elevation angle of 10°, the Coriolis illusion is experienced, and the tilt illusion is experienced at a slope of 45°; at a speed of 400 knots, the cloud base is 914m, and the cloud top is 3657m to experience the false horizon illusion; at an altitude of 1219m, a speed of 250 knots, and a heading of 163° to experience the black hole approach illusion.
[0072] Step S1 includes steps S11-S13.
[0073] Step S11, obtaining the eye movement tracking time series data and AOI gaze time series data of each pilot during the simulated flight as sample data. Figure 3 shown.
[0074] The eye movement tracking time series data includes the two-dimensional coordinates of the gaze focus, the three-dimensional coordinates of the gaze focus, the three-dimensional coordinates of the original position of the left eye, the three-dimensional coordinates of the moving direction of the left eye, the pupil diameter of the left eye, the three-dimensional coordinates of the original position of the right eye, the three-dimensional coordinates of the moving direction of the right eye and the pupil diameter of the right eye;
[0075] The AOI gaze timing data includes a number sequence of AOI subdivision areas that the pilot gazes at in sequence based on time during flight; the AOI subdivision areas include A sub-areas divided based on instrument display areas and screen display areas, and exterior scenery areas.
[0076] The eye tracking time series data and AOI gaze time series data of the pilot during the simulated flight are obtained as sample data, including:
[0077] Obtaining the eye movement tracking time series data and the flight training experiment video;
[0078] Based on the flight experiment video detection, all AOI subdivision areas in the flight experiment video in time sequence are obtained;
[0079] Based on the two-dimensional and three-dimensional coordinates of the gaze focus in the eye tracking time series data, it is determined whether the coordinates are located in a certain AOI subdivision area. According to the matching results, the AOI subdivision area numbers of the pilot's gaze are recorded in chronological order to obtain the AOI gaze time series data.
[0080] (1) Obtaining eye movement tracking time series data and flight training experiment videos
[0081] The pilot's eye movement time series data and flight training experiment video were obtained through a spatial orientation flight simulator and a wearable eye tracker. When the pilot's eyes looked at the spatial orientation flight simulator screen through the eye tracker, the wearable eye tracker recorded his eye movement traces.
[0082] The basic working principle of the eye tracker: First, use infrared or other light sources to illuminate the eyes to produce two bright spots: corneal reflection and pupil reflection. Then use a camera or other image sensor to capture the image of the eye and extract the position information of the corneal reflection and pupil reflection. Then use advanced image processing algorithms and 3D eyeball models to calculate the direction of the eye and the coordinate position of the gaze point relative to the screen. Finally, there is the calibration process, which maps the eye's gaze point coordinates to the visual gaze coordinates on the screen, thereby obtaining the eye movement tracking time series data and flight training experiment video of the test pilot.
[0083] Figure 4 The white circular outline area is the pilot's visual focus, and the red curve is the historical trajectory of the pilot's previous gaze movement. Figure 4 to Figure 5 It shows that the pilot's gaze focus shifts from the runway to the instrument panel. The characteristics of the acquired eye tracking time series data are shown in Table 1.
[0084] Table 1: Eye tracking time series data characteristics
[0085]
[0086] The two-dimensional coordinates of visual gaze are split into two-dimensional features of X and Y coordinates, and the three-dimensional coordinates are split into three-dimensional features of X, Y and Z coordinates.
[0087] The eye movement coordinate position is an important feature that reflects the pilot's visual information. In addition, the pupil diameter feature also plays a pivotal role. The pilot's left and right pupil diameter feature is an indicator that reflects the pilot's physiological and psychological state. The pilot's pupil diameter will change with factors such as lighting, visual stimulation, cognitive load, emotions, and fatigue. Generally speaking, the larger the pupil diameter, the higher the pilot's cognitive load; conversely, the smaller the pupil diameter, the lower the pilot's cognitive load. In addition, the difference in the pilot's left and right pupil diameters can also reflect the pilot's spatial orientation ability. If the difference in the left and right pupil diameters is large, it means that the pilot's spatial orientation ability is poor and is prone to spatial orientation disorders.
[0088] (2) Obtaining AOI gaze sequence data
[0089] AOI is an important source of information for pilots to acquire and maintain situational awareness. Pilots form a comprehensive understanding and prediction of current and future situations by observing and analyzing information on aircraft instruments, radar, navigation, communication and other equipment, as well as external weather, terrain and other information. Therefore, pilots need to reasonably allocate and adjust their attention to AOI according to mission requirements and environmental changes. Studies have shown that the longer the time spent paying attention to the AOI subdivision area, the more attention they pay to SA information.
[0090] There are three main areas that the pilot observes during flight: in the figure, number 1 is the instrument panel, number 2 is the screen display, and number 3 is the exterior view. For example, in the present invention, the instrument panel, screen display, and exterior view are divided into 13 AOI subdivision areas, such as Figure 6 And as shown in Table 2.
[0091] The main basis for the division of AOI subdivision areas is the type of instrument. Pilots will use and pay attention to different instruments at different stages. The specific meanings of AOI subdivision areas are shown in Table 2.
[0092] Table 2: Meaning of AOI subdivision areas in the real scene inside the space directional flight simulator and the outdoor location
[0093]
[0094]
[0095] Among them, the steps for determining the AOI subdivision area are as follows:
[0096] Step 1: Identify the types of instrument panels and screen displays in the flight simulator:
[0097] Step 2: Based on the instrument panel function module and the screen information display area, the instrument panel and the screen display are subdivided into a plurality of logical AOI subdivision areas; the exterior scene is one AOI subdivision area;
[0098] Step 3: Clarify the scope and meaning of each AOI segment to ensure that each AOI segment covers the information that pilots need to pay attention to in a specific situation.
[0099] Step 4: number the AOI subdivision areas to obtain m AOI subdivision areas.
[0100] Exemplarily, based on the instrument categories of the space-oriented simulated aircraft in the present invention, 13 AOI subdivision areas are determined.
[0101] Among them, the exterior scene is counted as one AOI subdivision area.
[0102] Since the 6-DOF spatial orientation flight simulator has a non-fixed viewing angle, direct Figure 6 Transforming to region templates for matching does not scale well to all viewpoints.
[0103] If you directly Figure 6 The displayed flight simulator instrument panel is treated as a regional template and then matched with the pilot's eye movement data, resulting in the following problems:
[0104] When the pilot's view rotates or tilts, Figure 6The regional template also needs to be rotated or tilted accordingly to align with the pilot's eye movement data, which increases the complexity and error of the calculation; when the pilot's perspective is translated or scaled, Figure 6 The area template also needs to be translated or scaled accordingly to align with the pilot's eye movement data, which will cause the position and shape of the AOI to change, thereby affecting the accuracy of AOI recognition and matching; when the pilot's perspective is blocked or distorted, Figure 6 The area template may not completely cover the pilot's eye movement data, or may be inconsistent with the pilot's eye movement data, which will result in loss or error of AOI.
[0105] In order to solve the above problems, the present invention uses a method of AOI recognition and matching based on deep learning to train the YOLOv8n model to recognize the AOI area, as follows:
[0106] The first step is to use the lightweight YOLOv8n model, input the image data of 13 AOI subdivision areas for training, and obtain a neural network model that can detect AOI. The YOLOv8n model is an improved target detection model based on the YOLO series. YOLOv8n replaces the C3 structure of YOLOv8 with a C2f structure with richer gradient flow, and adjusts the number of channels for models of different scales, thereby improving the performance and efficiency of the model; uses a new detection head to separate the classification and regression branches, and changes from Anchor-Based to Anchor-Free, thereby improving the accuracy and robustness of detection; uses a new loss function to balance the contributions of different targets, thereby improving the stability and generalization of detection;
[0107] The second step is to input the flight training experiment video recorded by the eye tracker into the trained network for inference, which can detect the AOI subdivision area in the video, as well as the position and shape of the AOI subdivision area. No matter how the pilot's perspective changes, the pilot's visual gaze area can be identified in real time, thus adapting to all perspectives;
[0108] The third step is to use the focus coordinates in the eye tracking time series data to calculate whether it is located in a certain AOI subdivision area, and then get the AOI gaze sequence. The AOI gaze sequence refers to the number of AOIs that the pilot gazes in chronological order during the flight, which can reflect the pilot's visual attention and cognitive process, as well as the pilot's situational awareness level.
[0109] Since the AOI area features are relatively simple, we use the lightweight YOLOv8n model to input 13 ( Figure 6The AOI subdivision areas are trained, and then the flight training test video recorded by the eye tracker is input into the trained network for inference, so that the AOI subdivision areas can be detected. The AOI gaze sequence data can be obtained by calculating whether the focus coordinates in the eye movement tracking time series data are located in a certain AOI subdivision area.
[0110] Step S12: Generate category labels for the eye movement tracking time series data and the AOI gaze time series data.
[0111] Exemplarily, the simulation experiment training SA assessment items include: pitch illusion, Coriolis illusion, tilt illusion, false horizon illusion, runway width and depth illusion, and black hole approach illusion, as described in Table 3.
[0112] Table 3: SA assessment items
[0113]
[0114] If the correction result of the indicator corresponding to each SA assessment item is completely consistent with the expected result, the full score can be obtained. The specific SA assessment item and the corresponding correction success full score fullscore i Definition, set based on the specific needs of each assessment.
[0115] Execute the following process to generate the corresponding category labels:
[0116] Based on the SA assessment items, measurement indicators, recovery results, expected results and the full score of the SA assessment items, the evaluation result Score of each pilot's SA assessment item is calculated, as shown in formula (1):
[0117]
[0118] Where n is the total number of indicators of SA evaluation items, D i M is the deviation between the pilot's recovery result and the expected result. i is the maximum deviation value in the indicator, fullscore i It is the full score obtained by the pilot when the recovery result of the indicator corresponding to the SA assessment item is completely consistent with the expected result;
[0119] The level corresponding to the evaluation result is used as the corresponding category label of the eye tracking time series data and AOI gaze time series data obtained during the flight simulation.
[0120] For each metric of the SA assessment item, illustratively, as shown in Table 4, for a pitch illusion dive situation, the pilot may recover successfully, or fail to recover, or be in the intermediate stage between successful recovery and failure to recover.
[0121] Table 4: Complexity of SA Assessment Items
[0122]
[0123] The eye tracking time series data and AOI gaze sequence data in the sample data are both 1 matrix, that is, the data itself is a three-dimensional tensor of multi-dimensional time series data, which is expanded and spliced into two dimensions according to the samples. Each sample has only one category label.
[0124] During the 60-minute training course, the pilots were given the background knowledge of the SA problem and were instructed by the test personnel on how to use the space orientation flight simulator to complete the SA assessment project. Each pilot had a 30-minute space orientation flight simulator operation experience to ensure that the pilots could successfully operate the space orientation flight simulator to collect data.
[0125] During flight, pilots face various complex situations. Complex situation recovery for pilots refers to the process of the pilot restoring the aircraft from an abnormal flight state to a normal controllable state, which is an important part of pilot training.
[0126] The level corresponding to the scoring result is as shown in formula (2):
[0127]
[0128] The corresponding level of the scoring result is used as the corresponding category label of the eye tracking time series data and AOI gaze time series data obtained during the flight simulation. If the scoring result is less than 60 points, it is considered a simulated flight failure, and it is not counted as a sample, and the spatial orientation flight simulation test is repeated.
[0129] Step S2 includes steps S21-S22.
[0130] Step S21, constructing a Ti-SA model for pilot situational awareness assessment;
[0131] like Figure 2 As shown in the overall structure diagram of the Ti-SA model, the Ti-SA model for pilot situational awareness assessment includes an eye tracking multi-dimensional attention enhancement module, an AOI multi-scale comprehensive feature module, a fusion module and a SA level output module;
[0132] The eye movement trace multi-dimensional attention enhancement module is used to receive the eye movement trace time series data, perform multi-dimensional attention feature extraction, and output the category probability value corresponding to the eye movement trace time series data;
[0133] The AOI multi-scale comprehensive feature module is used to receive the AOI gaze time series data, perform multi-scale feature fusion, and output the category probability value corresponding to the AOI fusion feature;
[0134] The fusion module fuses the eye tracking depth feature and the AOI fusion feature to obtain a predicted situational awareness level result;
[0135] The SA level output module is used to output the predicted situational awareness level result.
[0136] The present invention transforms the pilot SA evaluation problem into a multi-dimensional time series data classification problem. Since the amount of experimental data is huge and the sequence size is particularly long, it is particularly necessary to consider the model selection in this case. The overall structure of the proposed Ti-SA model is as follows: Figure 2 shown.
[0137] The eye tracking multi-dimensional attention enhancement module includes:
[0138] A first linear projection layer, used to map the eye movement tracking time series data features to a high dimension;
[0139] The first random deactivation layer is used to perform random deactivation processing on the mapped high-dimensional time series data to enhance the generalization of the eye movement trace multi-dimensional attention enhancement module and prevent overfitting;
[0140] A first, a second and a third multi-layer attention structure connected in sequence, for obtaining an encoded representation of the eye tracking temporal data;
[0141] Bi-LSTM, which uses the bidirectional capability of LSTM to enhance the understanding of the time series of the encoded representation output after three multi-layer attention structures;
[0142] Adaptive average pooling layer to reduce the dimensionality of Bi-LSTM output features while retaining important information;
[0143] The second linear projection layer is used to match the output dimension and obtain three category probability values corresponding to the eye movement trace.
[0144] Wherein, the multi-layer attention structure includes:
[0145] The multi-head attention module in the first multi-layer attention structure is used to obtain the output features of the first random inactivation layer, and use multiple attention heads to obtain comprehensive features; the multi-head attention module in the second multi-layer attention structure obtains the output features of the first multi-layer attention structure; the multi-head attention module in the third multi-layer attention structure obtains the output features of the second multi-layer attention structure;
[0146] The first residual connection and normalization layer in the first multi-layer attention structure is used to connect the data residual of the first random inactivation layer to the comprehensive features output by the multi-head attention module and then normalize them to obtain a normalized result 1A;
[0147] The first residual connection and normalization layer in the second multi-layer attention structure are used to connect the output feature residuals of the second multi-layer attention structure and then normalize them to obtain a normalized result 1B;
[0148] The first residual connection and normalization layer in the third multi-layer attention structure are used to residually connect the output features of the third multi-layer attention structure and then normalize them to obtain a normalized result 1C;
[0149] The feedforward network in each multi-layer attention structure is used to learn the nonlinear representation of the normalized results 1A, 1B and 1C respectively;
[0150] The second residual connection and normalization layer in each multi-layer attention structure are respectively used to perform residual connection on the normalized results 1A, 1B and 1C and the output of the feedforward network and then normalize them to obtain normalized results 2A, 2B and 2C;
[0151] The second random deactivation layer in each multi-layer attention structure is used to perform random deactivation processing on the normalized results 2A, 2B and 2C respectively to enhance the generalization ability of the multi-layer attention structure.
[0152] Among them, the multi-head attention module in the three-layer attention structure divides the input features of the module into multiple attention heads. Each head processes the feature information independently and understands the features from different angles. Finally, the outputs of multiple attention heads are merged to obtain a comprehensive feature representation.
[0153] The core idea of the attention mechanism is to use query vectors (Query), key vectors (Key) and value vectors (Value) to represent different parts of the input sequence, and obtain the part most relevant to the query vector by calculating the similarity between them. Then, by normalizing and weighted summing the similarity, an output vector is obtained, which is the output result of the multi-head attention module.
[0154] The attention mechanism imitates the principle of human attention, allowing the model to focus on key information in the input data. This mechanism can assign weights to the input data by calculating the importance of different parts, thereby more effectively capturing the key features of the data. Therefore, for eye tracking time series data, the attention mechanism is used to construct a network for training. Adding an LSTM module can use the implicit state in the LSTM to alleviate errors or noise in the past sequence and improve the robustness of the model.
[0155] The input of the eye movement track multi-dimensional attention enhancement module is the eye movement track time series data in the training data set, and the output is the category probability value corresponding to the eye movement track time series data.
[0156] The AOI multi-scale comprehensive feature module includes six sequentially connected first to sixth heuristic modules, a global average pooling layer, a third random inactivation layer and a third linear projection layer;
[0157] The first to sixth heuristic modules are used to perform multi-scale feature extraction on the AOI gaze time series data; wherein the AOI gaze time series data is connected to the output of the second heuristic module through a first residual; the output of the third heuristic module is connected to the output of the fifth heuristic module through a second residual;
[0158] The global average pooling layer is used to reduce the dimension of the output features of the first heuristic module and extract global features;
[0159] The third random deactivation layer is used to perform random deactivation processing on the extracted global features to enhance the generalization ability of the AOI multi-scale comprehensive feature module;
[0160] The third linear projection layer is used to reduce the dimension of the matching output dimension to obtain three category probability values corresponding to the AOI fusion feature.
[0161] like Figure 6 As shown, the heuristic module is not a traditional convolutional layer.
[0162] The first to sixth heuristic modules each include a first component, a second component and an output layer, as follows:
[0163] The first component is the first bottleneck layer, which uses a filter with a length of 1 to reduce the dimension of the input multi-dimensional AOI gaze time series data;
[0164] The second component includes three parallel convolutions, a max pooling, and a second bottleneck layer;
[0165] The three parallel convolutions are used to perform convolution operations of three lengths on the output of the first bottleneck layer to capture temporal features of different scales;
[0166] The maximum pooling is used to perform a maximum pooling operation on the input multi-dimensional AOI gaze time series data;
[0167] The second bottleneck layer is used to reduce the dimension of the maximum pooling output;
[0168] The output layer is used to concatenate the output features of each parallel convolution and maximum pooling to form a final output multi-dimensional time series;
[0169] Among them, the three convolution lengths are one-dimensional convolutions of different lengths, and the convolution kernels are 15, 25, and 50 respectively.
[0170] The first bottleneck layer performs dimensionality reduction, thereby reducing the complexity of the model and the risk of overfitting. The heuristic module is a structure that can simultaneously apply convolution kernels and pooling layers of different sizes, thereby improving the expressiveness and efficiency of the model.
[0171] The eye movement data category probability value output by the eye movement multi-dimensional attention enhancement module and the AOI gaze time series data category probability value output by the AOI multi-scale comprehensive feature module are fused through the fusion module to obtain the SA level state, which is output through the SA level output module.
[0172] The fusion module averages the probability values of the three categories corresponding to the eye tracking time series data and the three category probability values corresponding to the AOI fusion features to obtain an averaged probability value;
[0173] Comparing the averaged probability values of the three categories to determine the category with the highest probability;
[0174] The category with the largest average probability value is taken as the pilot's SA level status.
[0175] Step S22: train the Ti-SA model using the training data set until a trained Ti-SA model is obtained when the loss function converges.
[0176] For the training feature data input into the Ti-SA model (including eye tracking time series data and AOI gaze time series data), Spearman correlation coefficient analysis was used, p < 0.05, which was significantly correlated, so the input feature data did not need to be screened again.
[0177] The loss function of the Ti-SA model training is the cross entropy function. Until the loss function converges, the trained Ti-SA model is obtained. The cross entropy function E is shown in formula (3):
[0178]
[0179] Among them, n is the number of category labels, yi and p i are the true label and predicted category probability value of the i-th category label respectively.
[0180] The pilot data of eye tracking time series data and AOI gaze time series data obtained in the present invention specially selected professional pilots as the subjects of this test group. Through simple random sampling, fighter pilots who met the standards were randomly selected from the professional pilot group participating in flight training by lottery. Inclusion criteria: The pilots were in good physical condition, the Minnesota Multiphasic Personality Inventory (MMPI) met the standards, all had the necessary flying skills and experience, and the pilots' physical and cognitive abilities were normal during the test. All professional pilots participating in the experiment understood the test process and precautions, and voluntarily signed a written informed consent. A total of 30 pilots were included in the simulated flight process. They were in good physical and mental condition during the experiment. They were all male, aged 23-38 years old, with an average of (25.15±4.52) years old, with a flight time of 300-2200 hours, and an average flight time of 350.00 (340.75,912.50) hours.
[0181] The performance comparison of the Ti-SA model is as follows:
[0182] For example, the Ti-SA model training is performed under the experimental conditions of Python 3.8 and the deep learning framework Pytorch 2.1.2. In order to evaluate the performance of the Ti-SA model, the accuracy, precision, recall, and F1-score are used for measurement, as shown in formulas (4)-(7):
[0183]
[0184]
[0185] Where Acc, Pre, Rec and F1 represent accuracy, precision, recall and F1 score respectively. TP represents the number of samples that are actually positive and correctly predicted as positive by the Ti-SA model. TN represents the number of samples that are actually negative and correctly predicted as negative by the Ti-SA model. FP represents the number of samples that are actually negative but incorrectly predicted as positive by the Ti-SA model. FN represents the number of samples that are actually positive but incorrectly predicted as negative by the Ti-SA model.
[0186] The performance comparison of the Ti-SA model of the present invention and other multi-dimensional time series classification domain models is shown in Table 5:
[0187] Table 5: Performance comparison between each model and the proposed model
[0188] Model Accuracy (%) Accuracy (%) Recall rate (%) F1-score (%) MLSTM-FCN 83.27 84.44 90.46 87.35 Res-CNN 83.36 85.24 89.90 87.51 MiniRocket 88.36 89.10 93.58 91.28 TST 89.09 90.43 93.41 91.89 Ti-SA 92.18 92.95 95.49 94.20
[0189] (1) MLSTM-FCN: Convert the long short-term memory fully convolutional network (LSTM-FCN) to a multivariate model and expand the fully convolutional block;
[0190] (2) Res-CNN: Integrates residual networks and convolutional neural networks for time series classification;
[0191] (3) MiniRocket: A fast implementation that uses random convolution kernels to transform input time series and uses the transformed features to train the model;
[0192] (4) TST: A multi-dimensional temporal learning framework based on Transformer encoder.
[0193] (5) The Ti-SA model proposed in this invention: a multi-dimensional SA data learning model that integrates multi-head attention mechanism and heuristic module.
[0194] The performance of different models on the training dataset is shown in Table 4. The Ti-SA model achieves the best performance when used on the obtained dataset. The accuracy rate can reach 92.18%, which is better than other commonly used models in the field of multidimensional time series classification.
[0195] SA affects the pilot's decision-making process and information acquisition, and is related to flight safety, efficiency, and the smooth operation of the entire aviation system. Therefore, evaluating the pilot's SA level has always been a focus of concern in the aviation industry. The present invention constructs a training data set under spatial directional flight simulation conditions and evaluates the pilot's SA level through a deep learning method, obtaining a Ti-SA model with high evaluation accuracy.
[0196] During the flight simulation test, each pilot was randomly frozen in the simulation scene at least 5 times. The purpose of this is to simulate the emergencies and challenges in real flight. Freezing the simulation scene helps to assess their SA level. Such comprehensive training can enhance the pilot's physiological adaptability and coping ability in flight missions, thereby improving flight safety and operational effectiveness.
[0197] The Ti-SA model in the present invention integrates the multi-head attention mechanism and the heuristic module to extract and learn the features of the eye tracking time series data and the AOI gaze time series data. Compared with the model that only learns unilateral features, the Ti-SA model of the present invention can comprehensively consider the features of the eye tracking time series data and the AOI gaze sequence, and realize a more comprehensive pilot SA evaluation.
[0198] The Ti-SA model uses a multi-head attention mechanism to capture long-term dependencies in eye movement time series data and mitigates errors or noise in past sequences through a Bi-LSTM model.
[0199] At the same time, the local short-term dependency features in the AOI gaze sequence are captured through the heuristic module to adapt to the situation where the pilot focuses on a specific AOI segment within a specific time.
[0200] The results in Table 5 show that the proposed situational awareness assessment method that integrates eye tracking time series data and AOI gaze sequence is effective and can accurately judge the pilot's situational awareness level. The model in this study has a comprehensive feature learning mechanism, robustness, and a specific design for pilot SA tasks, which has obvious advantages in pilot situational awareness assessment.
[0201] In summary, the method for constructing a Ti-SA model for pilot situational awareness assessment according to an embodiment of the present invention has the following beneficial effects:
[0202] 1. The present invention constructs a Ti-SA model that integrates a multi-head attention mechanism and a heuristic module, and simultaneously learns the long-term dependency relationship in the eye tracking time series data and the local short-term dependency features in the AOI gaze sequence data. Compared with the existing method of evaluating after the experiment, the Ti-SA model is designed for real-time or near real-time evaluation, reducing the delay time of the evaluation. The model evaluates the pilot's situational awareness ability in real time, so that the evaluation results are fed back to the pilot or the training system in real time, so as to adjust the training or operation strategy in time;
[0203] 2. The present invention solves the technical problem that the selected data in the existing model cannot adapt to the complex environment and leads to poor effectiveness by combining the pilot's eye movement tracking time series data and AOI gaze time series data, and provides a more comprehensive pilot situational awareness ability assessment, which helps to more accurately capture the pilot's cognitive state of the flight environment, thereby improving the accuracy of the assessment results;
[0204] 3. The model building method of the present invention includes a complete process from data collection, feature extraction, model training to evaluation level output. The optimization of this process helps to simplify the actual operation of pilot situational awareness evaluation, improve the efficiency and practicality of the evaluation process, and thus optimize the pilot training and situational awareness evaluation process;
[0205] 4. The Ti-SA model of the present invention uses a multi-attention mechanism, adaptive average pooling, and multi-scale feature extraction of a heuristic module. The Ti-SA model can adapt to different flight phases and the cognitive characteristics of different pilots and provide personalized evaluation results;
[0206] 5. Through more accurate and real-time situational awareness assessment of pilots during flight, the present invention helps to identify possible cognitive deficiencies or errors of pilots at an early stage, so as to take preventive measures, reduce flight accidents caused by insufficient situational awareness, and improve overall flight safety; the Ti-SA model can not only be used to assess the SA ability level of pilots, but also assist in pilot training. By analyzing the SA ability level during the training process, customized training suggestions can be provided to pilots to help them improve their flying skills and ability to cope with complex situations.
[0207] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for constructing a Ti-SA model for pilot situational awareness assessment, characterized in that: The steps include: Step S1, obtaining eye movement tracking time series data and AOI gaze time series data of multiple pilots during simulated flight as sample data, wherein the sample data and the corresponding generated category labels constitute a training data set; wherein the category label is a level label of the pilot's situational awareness; Step S2, constructing a Ti-SA model for pilot situational awareness assessment; training the Ti-SA model using the training data set until a trained Ti-SA model is obtained when the loss function converges; Execute the following process to generate the corresponding category labels: Based on the SA assessment items, measurement indicators, recovery results, scheduled results, and the full score of the SA assessment items, the assessment result Score of each pilot's SA assessment item is calculated as follows: Where n is the total number of indicators of SA evaluation items, D i M is the deviation between the pilot's recovery result and the expected result. i is the maximum deviation value in the indicator, fullscore i It is the full score obtained by the pilot when the recovery result of the indicator corresponding to the SA assessment item is completely consistent with the expected result; The level corresponding to the evaluation result is used as the corresponding category label of the eye tracking time series data and AOI gaze time series data obtained during the flight simulation.
2. The method according to claim 1, characterized in that: The eye movement tracking time series data includes the two-dimensional coordinates of the gaze focus, the three-dimensional coordinates of the gaze focus, the three-dimensional coordinates of the original position of the left eye, the three-dimensional coordinates of the moving direction of the left eye, the pupil diameter of the left eye, the three-dimensional coordinates of the original position of the right eye, the three-dimensional coordinates of the moving direction of the right eye and the pupil diameter of the right eye; The AOI gaze timing data includes a number sequence of AOI subdivision areas that the pilot gazes at in sequence based on time during flight; the AOI subdivision areas include A sub-areas divided based on instrument display areas and screen display areas, and exterior scenery areas.
3. The method according to claim 2, characterized in that: The eye tracking time series data and AOI gaze time series data of the pilot during the simulated flight are obtained as sample data, including: Obtaining the eye movement tracking time series data and the flight training experiment video; Based on the flight experiment video detection, all AOI subdivision areas in the flight experiment video in time sequence are obtained; Based on the two-dimensional and three-dimensional coordinates of the gaze focus in the eye tracking time series data, it is determined whether the coordinates are located in a certain AOI subdivision area. According to the matching results, the AOI subdivision area numbers of the pilot's gaze are recorded in chronological order to obtain the AOI gaze time series data.
4. The method according to claim 1, characterized in that: The Ti-SA model for pilot situational awareness assessment includes an eye tracking multi-dimensional attention enhancement module, an AOI multi-scale comprehensive feature module, a fusion module and an SA level output module; The eye movement trace multi-dimensional attention enhancement module is used to receive the eye movement trace time series data, perform multi-dimensional attention feature extraction, and output the category probability value corresponding to the eye movement trace time series data; The AOI multi-scale comprehensive feature module is used to receive the AOI gaze time series data, perform multi-scale feature fusion, and output the category probability value corresponding to the AOI fusion feature; The fusion module fuses the eye tracking depth feature and the AOI fusion feature to obtain a predicted situational awareness level result; The SA level output module is used to output the predicted situational awareness level result.
5. The method according to claim 4, characterized in that: The eye tracking multi-dimensional attention enhancement module includes: A first linear projection layer, used to map the eye movement tracking time series data features to a high dimension; The first random deactivation layer is used to perform random deactivation processing on the mapped high-dimensional time series data to enhance the generalization of the eye movement trace multi-dimensional attention enhancement module and prevent overfitting; A first, a second and a third multi-layer attention structure connected in sequence, for obtaining an encoded representation of the eye tracking temporal data; Bi-LSTM, which uses the bidirectional capability of LSTM to enhance the understanding of the time series of the encoded representation output after three multi-layer attention structures; Adaptive average pooling layer to reduce the dimensionality of Bi-LSTM output features while retaining important information; The second linear projection layer is used to match the output dimension and obtain three category probability values corresponding to the eye movement trace.
6. The method according to claim 5, characterized in that: The AOI multi-scale comprehensive feature module includes six sequentially connected first to sixth heuristic modules, a global average pooling layer, a third random inactivation layer and a third linear projection layer; The first to sixth heuristic modules are used to perform multi-scale feature extraction on the AOI gaze time series data; wherein the AOI gaze time series data is connected to the output of the second heuristic module through a first residual; the output of the third heuristic module is connected to the output of the fifth heuristic module through a second residual; The global average pooling layer is used to reduce the dimension of the output features of the first heuristic module and extract global features; The third random deactivation layer is used to perform random deactivation processing on the extracted global features to enhance the generalization ability of the AOI multi-scale comprehensive feature module; The third linear projection layer is used to reduce the dimension of the matching output dimension to obtain three category probability values corresponding to the AOI fusion feature.
7. The method according to claim 5, characterized in that: The multi-head attention module in the three-layer attention structure splits the input features of the module into multiple attention heads. Each head processes the feature information independently and understands the features from different angles. Finally, the outputs of multiple attention heads are merged to obtain a comprehensive feature representation.
8. The method according to claim 6, characterized in that: The first to sixth heuristic modules each include a first component, a second component and an output layer, as follows: The first component is the first bottleneck layer, which uses a filter with a length of 1 to reduce the dimension of the input multi-dimensional AOI gaze time series data; The second component includes three parallel convolutions, a max pooling, and a second bottleneck layer; The three parallel convolutions are used to perform convolution operations of three lengths on the output of the first bottleneck layer to capture temporal features of different scales; wherein the three convolutions of different lengths are one-dimensional convolutions of different lengths, and the convolution kernels are 15, 25 and 50 respectively; The maximum pooling is used to perform a maximum pooling operation on the input multi-dimensional AOI gaze time series data; The second bottleneck layer is used to reduce the dimension of the maximum pooling output; The output layer is used to concatenate the time series features of different scales captured by the three parallel convolutions and the output features of the maximum pooling to form the final output multi-dimensional time series.
9. The method according to claim 4, characterized in that: The fusion module averages the probability values of the three categories corresponding to the eye tracking time series data and the three category probability values corresponding to the AOI fusion features to obtain an averaged probability value; Comparing the averaged probability values of the three categories to determine the category with the highest probability; The category with the largest average probability value is taken as the pilot's SA level status.
Citation Information
Patent Citations
Pilot real-time situation awareness evaluation method based on eye movement and visual fixation data
CN116993202A