Attention state recognition method for air traffic control personnel

By using attention classification models in the aviation control system, the attention status of aviation control personnel is automatically identified, which solves the problem of poor timeliness of attention allocation errors in the prior art, and improves identification efficiency and safety.

CN120220220APending Publication Date: 2025-06-27NAVAL AVIATION UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343717.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-22
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Aviation control personnel may have errors or loopholes in their work. The existing technology relies on manual observations to be inefficient and prone to errors, and cannot detect and correct attention allocation errors in a timely manner.

Method used

An air control personnel attention state recognition method is adopted. By determining the difference between the expected attention distribution label and the actual attention distribution label of the current aircraft video segment, the attention classification model is used to automatically identify the attention state of the controlled personnel and issue corresponding prompts.

Benefits of technology

It improves the efficiency of attention state recognition, reduces the error rate, realizes timely attention state monitoring, improves the level of tower automation assisted decision-making, and ensures flight safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220220A_ABST
    Figure CN120220220A_ABST
Patent Text Reader

Abstract

The invention discloses a method for identifying the attention state of air traffic control personnel. The method comprises the following steps: step 1, determining an expected attention distribution label corresponding to a current aircraft video segment; step 2, inputting the eye movement video segment of the current controller into the attention classification model to obtain a current actual attention distribution label of the controller; and step 3, judging whether the attention of the control personnel is proper or not according to the difference between the current actual attention distribution label and the expected attention distribution label in the corresponding time period, and if not, giving out a prompt. According to the method, the attention state is automatically and intelligently recognized, the recognition efficiency is improved, the error rate is reduced, meanwhile, the recognition timeliness is guaranteed, the automatic auxiliary decision-making level of the control tower is improved, and the flight safety is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of behavior type recognition, and particularly relates to a method for recognizing the attention state of air traffic controllers. Background Art

[0002] In air traffic control tasks, the attention allocation of controllers is of crucial value for timely detecting potential dangers and ensuring flight safety. According to the state and position of aircraft, the eye movement attention patterns of air traffic controllers can be mainly divided into the following three categories: 1. Saccadic movement: When one aircraft flies out of the airspace and another aircraft enters the airspace to prepare for landing, the controller needs to make a rapid eye movement, that is, a rapid eye movement from one fixation point to another fixation point.

[0003] 2. Vergence movement: When the controller discovers an aircraft and needs to focus on tracking its takeoff and landing process, binocular simultaneous vergence eye movement needs to be made according to the depth change of the aircraft being stared at. This adjustment enables the controller to quickly concentrate attention on a specific aircraft after saccadic discovery of the target.

[0004] 3. Smooth pursuit: For the process of aircraft takeoff or landing, the controller needs to continuously track the aircraft in the field of view through the smooth rotation of the eyes to ensure continuous monitoring of its flight trajectory, thereby ensuring flight safety.

[0005] In addition, among the above three eye movement patterns, the vestibulo-ocular reflex always plays an important physiological role. This reflex can automatically compensate for the eye position when the controller's head or body moves, ensuring that the visual image remains stable on the retina, thereby maintaining continuous and effective attention to key flight targets.

[0006] Due to reasons such as fatigue and negligence, controllers may have attention allocation mistakes or loopholes during work, and trainees are also prone to unreasonable attention allocation problems during the learning process. To solve the above problems, it is necessary to rely on manual observation of the eye movements and expressions of controllers to judge their attention state. However, this method is inefficient and error-prone, and it is impossible to detect and correct attention allocation mistakes in a timely manner. Summary of the Invention

[0007] The present invention proposes a method for recognizing the attention state of air traffic controllers, and its purpose is to solve the problems of low efficiency, easy error and poor timeliness in attention state recognition.

[0008] The technical solution of the present invention is as follows: A method for recognizing the attention state of air traffic controllers, the steps include: Step 1, determine the expected attention allocation label corresponding to the current aircraft video segment; Step 2: Input the eye movement video segment of the current controller into the attention classification model to obtain the actual attention allocation label of the controller at present; The eye movement video segment is intercepted from the controller's eye movement video in a sliding window manner; The input of the attention classification model is the eye movement video segment, and the output is the actual attention allocation label; Step 3: Determine whether the controller's attention is appropriate according to the difference between the actual attention allocation label at present and the expected attention allocation label in the corresponding time period. If it is not appropriate, a prompt is issued.

[0009] As a further improvement of the above method for identifying the attention state of air traffic controllers, Step 1 specifically includes: Step 1-1: Use the camera device on the tower to obtain the aircraft video near the tower in real time, and intercept the current aircraft video segment from the aircraft video; Step 1-2: Use the aircraft target detection model to determine the position of the aircraft in each video frame of the aircraft video segment; The input of the aircraft target detection model is the aircraft video frame, and the output is the position of the aircraft in this video frame; Step 1-3: Use the following rules to determine the attention type that the controller should adopt in each aircraft video frame: A. If the aircraft is in the position of taking off or landing, the expected attention type is smooth tracking; B. If two aircraft are detected, and one is in the position of entering the airspace and the other is in the position of flying out of the airspace, the expected attention type is saccade movement; C. If the aircraft is in the position of about to take off or about to land, the expected attention type is accommodation movement; D. If the aircraft position does not belong to any of the situations in A, B, and C, the expected attention type is free action, that is, any attention action is acceptable; Step 1-4: Construct the expected attention allocation label according to the expected attention type corresponding to each frame in the aircraft video segment.

[0010] As a further improvement of the above method for identifying the attention state of air traffic controllers: The length of the expected attention allocation label is 4, each element corresponds to a different attention type respectively, and the value of each element represents the time percentage of the corresponding attention type in the video segment.

[0011] As a further improvement of the method for identifying the attention state of the air traffic controller: The actual attention allocation label is a label vector of length 4. Each element in the vector corresponds to an attention type, and the value of each element represents the percentage of time occupied by the corresponding attention type in the eye movement video segment.

[0012] As a further improvement of the method for identifying the attention state of the air traffic controller: The attention classification model includes a number of DW3DPM modules connected in sequence, a global average pooling layer, and a fully connected layer.

[0013] As a further improvement of the method for identifying the attention state of the air traffic controller: The DW3DPM module includes a convolutional module, a depth convolutional module, a pointwise convolutional module, and a three-dimensional projection attention module connected in sequence.

[0014] As a further improvement of the method for identifying the attention state of the air traffic controller: The three-dimensional projection attention module includes a first projection pooling module, a second projection pooling module, and a third projection pooling module; Suppose the size of the input feature map of the three-dimensional projection attention module is m×m×n. The input feature map passes through the first projection pooling module to obtain the first projection pooling feature map, and then the input feature map is multiplied by the first projection pooling feature map to obtain the first intermediate feature map; The first intermediate feature map passes through the second projection pooling module to obtain the second projection pooling feature map, and then the first intermediate feature map is multiplied by the second projection pooling feature map to obtain the second intermediate feature map; The second intermediate feature map passes through the third projection pooling module to obtain the third projection pooling feature map, and then the second intermediate feature map is multiplied by the third projection pooling feature map to obtain the output feature map of the three-dimensional projection attention module, and the size of the output feature map is m×m×n.

[0015] As a further improvement of the method for identifying the attention state of the air traffic controller: The working process of the first projection pooling module is as follows: First, perform max pooling on the input feature map in the width direction and then perform max pooling in the height direction to obtain the feature map M1. At the same time, perform average pooling on the input feature map in the width direction and then perform average pooling in the height direction to obtain the feature map A1. Then, the transformation results obtained by inputting the feature map M1 and the feature map A1 into a 3-layer fully connected layer are superimposed, and the superimposed result is input into the sigmoid function to obtain the first projection pooling feature map; The working process of the second projection pooling module is as follows: perform max pooling on the first intermediate feature map in the channel direction first, and then perform max pooling in the height direction to obtain the feature map M2. At the same time, perform average pooling on the first intermediate feature map in the channel direction first, and then perform average pooling in the height direction to obtain the feature map A2. Then, stack the transformation results obtained by inputting the feature map M2 and the feature map A2 into a 3-layer fully connected layer respectively, and input the stacked result into the sigmoid function to obtain the second projection pooling feature map; The working process of the third projection pooling module is as follows: perform max pooling on the second intermediate feature map in the channel direction to obtain the feature map M3, and perform average pooling in the channel direction at the same time to obtain the feature map A3. Then, for the stacked result of the feature map M3 and the feature map A3, after passing through convolution and the sigmoid activation layer in sequence, the third projection pooling feature map is obtained.

[0016] As a further improvement of the above air traffic controller attention state recognition method, the training process of the attention classification model is as follows: Step T-1: Take e <eye movement video segment, label> data pairs, input the eye movement video segments into the constructed attention classification model, and obtain the prediction results of these e data pairs; The label in the data pair is the true label manually annotated, with a length of 4, corresponding to 4 attention types respectively. The value of each element in the label represents the time percentage of the corresponding attention type in the video segment; Step T-2: For each data pair, respectively according to its prediction result and label calculate the cross-entropy loss function: , is the value of the th element in the label , is the value of the th element in the prediction result; Backpropagate the average value of the cross-entropy loss functions of the e data pairs to optimize the entire network;

[0017] As a further improvement of the above air traffic controller attention state recognition method: in step 3, subtract the actual attention allocation label from the expected attention allocation label element by element, and then sum the squares of each element of the subtraction result as the difference value; if the difference value exceeds the preset threshold, it is determined that the attention is improper.

[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. Based on the attention classification model, the present invention converts the attention allocation problem into a classification problem, and then judges whether the current controller's attention allocation meets the requirements based on the difference between the actual attention allocation label and the expected attention allocation label, realizing the automatic and intelligent recognition of the attention state, improving the recognition efficiency, reducing the error rate, ensuring the timeliness of recognition, improving the level of automated auxiliary decision-making in the tower, and ensuring flight safety.

[0019] 2. The three-dimensional projection attention module in the present invention obtains the attention situation at each angle by projecting the feature map in three directions, fully considering the information in the three-dimensional space, and more accurately capturing the movement pattern of the eyeball in the three-dimensional space. By calculating the attention coefficients in different directions, the three-dimensional projection attention module can more effectively utilize the information in the feature map, thereby improving the expression ability of the features. At the same time, as a plug-and-play module, the three-dimensional projection attention module can be integrated into the existing network structure to fuse the features from different perspectives through the cross-view spatial attention mechanism, enhancing the adaptability of the model to different perspectives and scenarios.

[0020] 3. The attention classification model uses depthwise separable convolution to replace the traditional two-dimensional convolution, significantly reducing the number of model parameters, reducing the computational complexity of the model, reducing the computational resources required for training and inference, and accelerating the training speed and inference speed of the model. At the same time, depthwise convolution can independently learn spatial features on each input channel, while pointwise convolution can effectively combine these features, which helps the model learn richer and more discriminative feature representations. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the processing process of the three-dimensional projection attention module. DETAILED DESCRIPTION OF THE INVENTION

[0022] The technical solution of the present invention will be described in detail below with reference to the drawings: A method for recognizing the attention state of an air traffic controller, the steps including: Step 1. Determine the expected attention allocation label corresponding to the current aircraft video segment.

[0023] The specific method is: Step 1-1. Use the camera device on the tower to obtain the aircraft video near the tower in real time, and intercept the current aircraft video segment from the aircraft video.

[0024] Step 1-2. Use the aircraft target detection model (such as YOLOv10) to determine the position of the aircraft in each video frame of the aircraft video segment.

[0025] The input of the aircraft target detection model is the aircraft video frame, and the output is the position of the aircraft in the video frame.

[0026] The YOLOv10 model can be trained and predicted in the following ways in advance: Using the camera device on the tower, obtain the videos of the takeoff and landing airspaces near the tower, and manually label the position tags [x1, y1, x2, y2] of the aircraft targets in each frame image of the video in the form of positive frames, which are the x and y coordinates of the upper left corner and the lower right corner of the target respectively, thereby constructing an aircraft target detection task dataset. Assume the imaging frequency is 20 frames per second. Randomly divide the aircraft target detection task dataset into training and test samples according to a ratio of 7:3 to construct a training set and a test set. Train the YOLOv10 model based on the training set and test the training effect through the test set.

[0027] Steps 1-3: Determine the type of attention that the air traffic controller should take in each aircraft video frame using the following rules: A. If the aircraft is in the takeoff or landing position, the expected type of attention is smooth tracking.

[0028] B. If two aircraft are detected, and one is in the entering airspace position and the other is in the leaving airspace position, the expected type of attention is saccadic movement.

[0029] C. If the aircraft is in the about-to-takeoff or about-to-land position, the expected type of attention is accommodation movement.

[0030] D. If the aircraft position does not belong to any of the situations in A, B, or C, the expected type of attention is free action, that is, any attention action is acceptable.

[0031] Steps 1-4: Construct an expected attention allocation label according to the expected type of attention corresponding to each frame in the aircraft video segment: The length of the expected attention allocation label is 4, each element corresponds to a different type of attention respectively, and the value of each element represents the percentage of time that the corresponding type of attention occupies in the video segment.

[0032] Step 2: Input the current air traffic controller's eye movement video segment into the attention classification model to obtain the current actual attention allocation label of the air traffic controller.

[0033] The eye movement video segment is intercepted from the air traffic controller's eye movement video in a sliding window manner.

[0034] In this embodiment, the eye movement video is 20 frames per second. The video segment is intercepted with a length of 0.5 seconds (window size) and a step size of 0.1 seconds to segment the eye movement video. Each video segment contains 10 frames of images. The image size is 512×512×3, so the size of the eye movement video segment is 512×512×30.

[0035] The input of the attention classification model is an eye movement video segment, and the output is the actual attention allocation label.

[0036] The actual attention allocation label is a label vector of length 4. Each element in the vector corresponds to an attention type respectively, and the value of each element represents the percentage of time that the corresponding attention type occupies in the eye movement video segment.

[0037] The attention classification model includes a number of DW3DPM modules connected in sequence, a global average pooling layer, and a fully connected layer.

[0038] The DW3DPM module includes a convolutional module (Conv2D), a depthwise convolutional module (DepthwiseConv), a pointwise convolutional module (Pointwise Conv), and a three-dimensional projection attention module (3DPM) connected in sequence.

[0039] In order to reduce the number of parameters in the present invention, depthwise separable convolution is introduced. By replacing the traditional Conv2D with two layers, namely Depthwise Conv and PointwiseConv, the number of parameters is reduced, the model complexity is lowered, and the inference speed is improved. The structure of the attention classification model is as follows: Table 1: Structure of the attention classification model.

[0040] Layer Type Output Size Description Input 512×512×30 Input video segment 512×512×30 Conv2D 256×256×32 3x3 convolution, 32 convolutional kernels, stride of 2, activation function ReLU Depthwise Conv 256×256×32 3x3 depthwise convolution, stride of 1, activation function ReLU Pointwise Conv 256×256×64 1x1 convolution, 64 convolutional kernels, activation function ReLU 3DPM 256×256×64 Channel attention mechanism to enhance feature representation Conv2D 128×128×64 3x3 convolution, 64 convolutional kernels, stride of 2, activation function ReLU Depthwise Conv 128×128×64 3x3 depthwise convolution, stride of 1, activation function ReLU Pointwise Conv 128×128×128 1x1 convolution, 128 convolutional kernels, activation function ReLU 3DPM 128×128×128 Channel attention mechanism to enhance feature representation Conv2D 64×64×128 3x3 convolution, 128 convolutional kernels, stride of 2, activation function ReLU Depthwise Conv 64×64×128 3x3 depthwise convolution, stride of 1, activation function ReLU Pointwise Conv 64×64×256 1x1 convolution, 256 convolutional kernels, activation function ReLU 3DPM 64×64×256 Channel attention mechanism to enhance feature representation Conv2D 32×32×256 3x3 convolution, 256 convolutional kernels, stride of 2, activation function ReLU Depthwise Conv 32×32×256 3x3 depthwise convolution, stride of 1, activation function ReLU Pointwise Conv 32×32×512 1x1 convolution, 512 convolutional kernels, activation function ReLU 3DPM 32×32×512 Channel attention mechanism to enhance feature representation Global Average Pooling 512 Global average pooling to reduce the feature map size Dense 4 Fully connected layer, activation function Softmax, output the final classification result As Figure 1 , the three-dimensional projection attention module includes a first projection pooling module, a second projection pooling module, and a third projection pooling module. Assume that the size of the input feature map of the three-dimensional projection attention module is m×m×n. The input feature map passes through the first projection pooling module to obtain the first projection pooling feature map, and then the input feature map is multiplied by the first projection pooling feature map to obtain the first intermediate feature map; the first intermediate feature map passes through the second projection pooling module to obtain the second projection pooling feature map, and then the first intermediate feature map is multiplied by the second projection pooling feature map to obtain the second intermediate feature map; the second intermediate feature map passes through the third projection pooling module to obtain the third projection pooling feature map, and then the second intermediate feature map is multiplied by the third projection pooling feature map to obtain the output feature map of the three-dimensional projection attention module, and the size of the output feature map is m×m×n.

[0041] The working process of the first projection pooling module is as follows: First, perform max pooling on the input feature map in the width direction, and then perform max pooling in the height direction to obtain the feature map M1. At the same time, perform average pooling on the input feature map in the width direction, and then perform average pooling in the height direction to obtain the feature map A1. Then, stack the transformation results obtained by respectively inputting the feature map M1 and the feature map A1 into a 3-layer fully connected layer (the number of nodes are: n, 2 / n, n), and then input the stacked result into the sigmoid function to obtain the first projection pooling feature map (the channel attention results of the projections in the width direction and the height direction).

[0042] The working process of the second projection pooling module is as follows: First, perform max pooling on the first intermediate feature map in the channel direction, and then perform max pooling in the height direction to obtain the feature map M2. At the same time, perform average pooling on the first intermediate feature map in the channel direction, and then perform average pooling in the height direction to obtain the feature map A2. Then, stack the transformation results obtained by respectively inputting the feature map M2 and the feature map A2 into a 3-layer fully connected layer (the number of nodes are: n, 2 / n, n), and then input the stacked result into the sigmoid function to obtain the second projection pooling feature map (the channel attention results of the projections in the channel direction and the height direction).

[0043] The working process of the third projection pooling module is as follows: Perform max pooling on the second intermediate feature map in the channel direction to obtain the feature map M3, and perform average pooling in the channel direction at the same time to obtain the feature map A3. Then, for the stacked result of the feature map M3 and the feature map A3, after passing through convolution (1×1, stride is 1) and the sigmoid activation layer in sequence, obtain the third projection pooling feature map.

[0044] The training process of the attention classification model is as follows: Set the learning rate to 0.0002, the number of training times to 500, and the batch size to e = 128. Train the model constructed in step 5 in the following order: Step T-1: Take e pairs of <eye movement video segment, label> data, input the eye movement video segment into the constructed attention classification model, and obtain the prediction results of these e data pairs.

[0045] The label in the data pair is the manually labeled true label, with a length of 4, corresponding to 4 attention types respectively. The value of each element in the label represents the time percentage of the corresponding attention type in the video segment.

[0046] Step T-2: For each data pair, respectively according to its prediction result and label calculate the cross-entropy loss function: , is the label in the The value of an element is the prediction result in the th element value. The average value of the cross-entropy loss function of e data pairs is backpropagated to optimize the entire network.

[0047] Step T-3: Repeat steps T-1 to T-2 for 500 times, and finally obtain a trained model.

[0048] Step 3: Determine whether the controller's attention is appropriate based on the difference between the current actual attention allocation label and the expected attention allocation label for the corresponding time period. If it is not appropriate, a prompt is issued.

[0049] Specifically, the difference between the actual attention allocation label and the expected attention allocation label is calculated element by element, and then the sum of the squares of the elements of the difference result is used as the difference value. If the difference value exceeds a preset threshold (0.2 in this embodiment), it is determined that the attention is not appropriate.

[0050] It should be noted that for those skilled in the art, obviously, the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. The scope of the present invention is defined by the claims rather than the above description.

Claims

1. A method for identifying the attention state of an air traffic controller, characterized in that the steps include: Step 1: Determine the expected attention allocation label corresponding to the current aircraft video segment; Step 2: Input the eye movement video segment of the current controller into the attention classification model to obtain the current actual attention allocation label of the controller; The eye-movement video segments were extracted from the controller’s eye-movement videos using a sliding window approach; The input of the attention classification model is the eye movement video segment, and the output is the actual attention allocation label; Step 3: Determine whether the controller's attention is appropriate based on the difference between the current actual attention allocation label and the expected attention allocation label for the corresponding time period, and issue a prompt if it is not appropriate.

2. The method for identifying the attention state of air traffic controllers according to claim 1, characterized in that: Step 1 specifically includes: Step 1-1, using the camera device on the tower to obtain the aircraft video near the tower in real time, and intercept the current aircraft video segment from the aircraft video; Step 1-2, using the aircraft target detection model to determine the position of the aircraft in each video frame in the aircraft video segment; The input of the aircraft target detection model is an aircraft video frame, and the output is the position of the aircraft in the video frame; Step 1-3: Determine the type of attention the controller should take in each aircraft video frame using the following rules: A. If the aircraft is in a take-off or landing position, the desired attention type is smooth pursuit; B. If two aircraft are detected, one at a position to enter the airspace and the other at a position to leave the airspace, the expected attention type is a saccadic movement; C. If the aircraft is in a position to take off or land, the expected attention type is accommodative movement; D. If the aircraft position does not belong to any of the cases A, B, and C, the desired attention type is free action, that is, any attention action is acceptable; Step 1-4: Construct the expected attention allocation label according to the expected attention type corresponding to each frame in the airplane video segment.

3. The method for identifying the attention state of air traffic controllers according to claim 2, characterized in that: The expected attention allocation label length is 4, each element corresponds to a different attention type, and the value of each element represents the percentage of time occupied by the corresponding attention type in the video segment.

4. The method for identifying the attention state of air traffic controllers according to claim 2, wherein: The actual attention allocation label is a label vector with a length of 4, each element in the vector corresponds to an attention type, and the value of each element represents the percentage of time occupied by the corresponding attention type in the eye movement video segment.

5. The method for identifying the attention state of air traffic controllers according to claim 1, characterized in that: The attention classification model consists of several DW3DPM modules connected sequentially, a global average pooling layer, and a fully connected layer.

6. The method for identifying the attention state of air traffic controllers according to claim 5, characterized in that: The DW3DPM module includes a convolution module, a depth convolution module, a point-by-point convolution module and a three-dimensional projection attention module which are connected in sequence.

7. The method for identifying the attention state of air traffic controllers according to claim 6, characterized in that: The three-dimensional projection attention module includes a first projection pooling module, a second projection pooling module and a third projection pooling module; Assume that the input feature map size of the three-dimensional projection attention module is m×m×n. The input feature map passes through the first projection pooling module to obtain the first projection pooling feature map, and then the input feature map is multiplied by the first projection pooling feature map to obtain the first intermediate feature map; the first intermediate feature map passes through the second projection pooling module to obtain the second projection pooling feature map, and then the first intermediate feature map is multiplied by the second projection pooling feature map to obtain the second intermediate feature map; the second intermediate feature map passes through the third projection pooling module to obtain the third projection pooling feature map, and then the second intermediate feature map is multiplied by the third projection pooling feature map to obtain the output feature map of the three-dimensional projection attention module, and the size of the output feature map is m×m×n.

8. The method for identifying the attention state of air traffic controllers according to claim 7, characterized in that: The working process of the first projection pooling module is as follows: the input feature map is firstly subjected to maximum pooling in the width direction and then to maximum pooling in the height direction to obtain the feature map M1, and at the same time, the input feature map is first subjected to average pooling in the width direction and then to average pooling in the height direction to obtain the feature map A1, and then the feature map M1 and the feature map A1 are respectively input into the three layers of fully connected layers to obtain the transformation results to be superimposed, and then the superposition result is input into the sigmoid function to obtain the first projection pooling feature map; The working process of the second projection pooling module is as follows: the first intermediate feature map is firstly subjected to maximum pooling in the channel direction and then to maximum pooling in the height direction to obtain feature map M2; at the same time, the first intermediate feature map is first subjected to average pooling in the channel direction and then to average pooling in the height direction to obtain feature map A2; then, feature map M2 and feature map A2 are respectively input into the three-layer fully connected layer to obtain the transformation results to be superimposed; and then, the superimposed results are input into the sigmoid function to obtain the second projection pooling feature map; The working process of the third projection pooling module is as follows: perform maximum pooling on the second intermediate feature map in the channel direction to obtain feature map M3, and perform average pooling in the channel direction to obtain feature map A3, and then superimpose the feature map M3 and the feature map A3. After the superimposed results pass through the convolution and sogmoid activation layers in turn, the third projection pooling feature map is obtained.

9. The method for identifying the attention state of air traffic controllers according to claim 2, characterized in that: The attention classification model training process is as follows: Step T-1, taking e <eye movement video segment, label> data pairs, inputting the eye movement video segment into the constructed attention classification model, and obtaining the prediction results of the e data pairs; The label in the data pair is a manually annotated real label, and its length is 4, corresponding to 4 types of attention respectively. The value of each element in the label represents the time percentage of the corresponding attention type in the video segment; Step T-2: For each data pair, according to its prediction results and tags Calculate the cross entropy loss function: , For label Middle element values, To predict the results Middle element values; the average value of the cross entropy loss function of e data pairs is returned to optimize the entire network; Step T-3, repeat steps T-1 to T-2 several times.

10. The method for identifying the attention state of air traffic controllers according to claim 1, characterized in that: In step 3, the actual attention allocation label is subtracted from the expected attention allocation label element by element, and then the square sum of each element of the difference result is taken as the difference value; if the difference value exceeds the preset threshold, it is judged that the attention is inappropriate.