Remote control tower intelligent man-machine interaction system driven by watching

The gaze-driven remote control intelligent human-machine interaction system analyzes the controller's gaze behavior in real time and dynamically adjusts the information display, solving the problems of low information acquisition efficiency and high safety risks in complex environments for remote control systems, and realizing efficient and safe multi-target monitoring and interaction.

CN120949935APending Publication Date: 2025-11-14CIVIL AVIATION FLIGHT UNIV OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511069934.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Remote control tower systems are inefficient in terms of information presentation and human-computer interaction, especially in complex traffic situations or multi-target tracking scenarios. Controllers need to frequently move their eyes, the information acquisition path is lengthy, the operation process is fragmented, and there is a lack of monitoring of abnormal areas and systematic attention management, which leads to the dispersion of attention resources and the risk of missing or misjudging.

Method used

The gaze-driven remote control intelligent human-computer interaction system analyzes the controller's gaze behavior in real time through a target mapping module, an eye-tracking data processing module, and an information enhancement and interaction support module. It dynamically adjusts the information display strategy, evaluates the rationality of attention allocation by combining a lightweight convolutional neural network model, and introduces a dual-channel early warning mechanism to achieve dynamic interaction and human-factor adaptation.

Benefits of technology

It improves information acquisition efficiency and interaction flexibility, reduces visual interference, shortens operation paths, enhances the system's real-time response capability and security level, and ensures control efficiency and security in high-density operating environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949935A_ABST
    Figure CN120949935A_ABST
Patent Text Reader

Abstract

The invention relates to the field of electric data processing, in particular to a fixation-driven remote control tower intelligent man-machine interaction system which comprises a target mapping module, an eye movement data processing module and an information enhancement and interaction support module. The target mapping module is used for fusing the video data and the radar data and constructing a region-of-interest mapping sequence; the eye movement data processing module is used for processing user eye movement data and generating a gazing sequence; and the information enhancement and interaction support module is used for matching the region-of-interest mapping sequence and the watching sequence in a spatial dimension and driving dynamic enhancement display of the sign information. According to the invention, a man-machine cooperation mechanism taking the watching behavior of the controller as the core drive is realized, so that the real-time response capability, the information presentation efficiency and the operation safety guarantee level of the system are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of air traffic control technology, and in particular to a gaze-driven remote control intelligent human-machine interaction system. Background Technology

[0002] With the development of the civil aviation industry and the continuous improvement of air traffic control automation, remote control tower technology, as an emerging direction in the field of air traffic control, is gradually becoming a feasible solution to replace traditional towers. This technology, through the deployment of high-resolution cameras, panoramic imaging systems, radar, and various sensors, transmits real-time images and related information from the airport to a remote control center, enabling controllers to achieve comprehensive control over airport operations such as takeoffs and landings and ground taxiing from a remote location. Compared with traditional towers, remote towers not only reduce manpower deployment and infrastructure construction costs but also improve operational flexibility and multi-airport collaboration capabilities. However, current remote tower systems still have significant shortcomings in information presentation and human-computer interaction. Especially when facing complex traffic situations or multi-target tracking scenarios, the system often fails to effectively support controllers in making rapid identification and accurate judgments. Controllers need to frequently move their gaze between multiple screens and manually interpret target and flight information, resulting in lengthy information acquisition paths, fragmented operational processes, and low overall interaction efficiency. Furthermore, in the absence of abnormal area monitoring and systematic attention management, controllers need to actively monitor multiple operational objects, leading to a dispersed attentional resource and a heavy workload, increasing the risk of missed observations and misjudgments.

[0003] In traditional air traffic control tower systems, research and engineering practices have incorporated Head-Up Display (HUD) technology. By overlaying key information such as flight identification, call signs, and flight numbers onto the tower's viewing glass as a semi-transparent layer, HUD reduces the frequency of controllers switching between looking down at radar or information terminals and looking up at the airport, improving their ability to maintain continuous focus on key targets and enhancing operational smoothness. For example, Chinese invention publication CN113990113 A proposes a HUD-enhanced air traffic control tower visual display method. After identifying an aircraft using target detection technology, simultaneously displaying the aircraft's key information can improve control efficiency and reduce the workload of controllers.

[0004] However, these methods are merely static displays of fixed information, serving only as a one-way information output mechanism for the control tower system. They lack effective linkage with the dynamic input behavior of controllers, resulting in numerous defects and limitations. When the system faces complex scenes such as densely packed multiple targets, overlapping targets, or crowded screen edges, static rules often lead to information redundancy, severe signage obstruction, and even visual confusion or misunderstanding, thereby increasing the cognitive burden on controllers and reducing information acquisition efficiency. Summary of the Invention

[0005] The purpose of this invention is to provide a gaze-driven remote control intelligent human-machine interaction system, which is particularly suitable for rapid and accurate information acquisition in complex environments, thereby improving control efficiency.

[0006] Therefore, the embodiments of the present invention provide the following technical solutions:

[0007] A gaze-driven remote control intelligent human-computer interaction system includes a target mapping module, an eye-tracking data processing module, and an information enhancement and interaction support module. The target mapping module is used to fuse video data and radar data to construct a region of interest mapping sequence. The eye-tracking data processing module is used to process user eye-tracking data to generate a gaze sequence. The information enhancement and interaction support module is used to match the region of interest mapping sequence and the gaze sequence in a spatial dimension to drive the dynamic enhancement display of signage information.

[0008] The above solution not only performs target detection but also collects eye-tracking data from users, i.e., controllers. Based on the eye-tracking data, it analyzes the user's current focus in real time and dynamically adjusts the target display information accordingly, either enhancing or simplifying the display. In other words, by analyzing the gaze state in the multi-target monitoring screen in real time and dynamically adjusting the information display strategy, it achieves human factor enhancement and intelligent assistance in the multi-target, high-load, and high-density operating environment under the remote control tower, thereby comprehensively improving control efficiency and operational safety.

[0009] The target mapping module first performs target detection based on video data to obtain the bounding boxes of the detected targets, and obtains the position coordinates of the radar targets based on radar data. Then, when the overlapping area of ​​the bounding boxes of the detected targets and the position coordinates of the radar targets is determined to be the same target, the target detection results and the radar data results are fused. Based on the fusion results, region of interest boundary data is constructed, forming a temporally continuous region of interest mapping sequence. This represents the k-th region of interest. This represents the bounding box coordinates and dimensions of the region of interest in the current video frame image, x k y k w represents the center coordinates of the k-th region of interest. k h k These are the width and height of the k-th region of interest, respectively, and ID. k N is the unique identifier for this region of interest. t This represents the number of targets detected in the current video frame t.

[0010] The eye-tracking data processing module is configured with an input buffer, an output buffer, a denoising thread, a gaze behavior recognition thread, and a merging output thread. The denoising thread denoises the real-time acquired eye-tracking data and writes the resulting gaze position sequence into the input buffer. The gaze behavior recognition thread reads the gaze position sequence data segment from the input buffer, identifies whether a gaze behavior has occurred, and writes the gaze point of the gaze behavior into the output buffer when a gaze behavior is identified. The merging output thread reads the gaze point from the output buffer and determines whether two adjacent gaze points meet the merging condition. If so, it merges the two adjacent gaze points, and all the retained gaze points form the gaze sequence.

[0011] In the above scheme, a double-buffered asynchronous acquisition mechanism and an improved Fast-I-DT algorithm are used to achieve high-precision extraction and recognition of gaze points, ensuring the real-time performance and stability of the system under high-frequency sampling conditions.

[0012] The information enhancement and interaction support module first reads the region of interest mapping sequence and overlays a simplified sign on the identified target region. Then, it determines whether the gaze behavior is a valid gaze based on a second time threshold. If so, it expands the information of the simplified sign to obtain a new sign and highlights it. It also determines whether the overlap rate between the proposed display position of the new sign and the display position of the existing sign exceeds a preset threshold. If so, it selects a candidate display position to display the new sign.

[0013] The above scheme introduces a gaze-driven information enhancement and interaction mechanism. After detecting a valid gaze from the controller, the system can dynamically adjust the display mode of the corresponding target sign, including highlighting, smooth zooming and loading additional information, and combine it with an avoidance and rearrangement strategy to prevent sign overlap and occlusion, thereby improving information recognition efficiency.

[0014] The information enhancement and interaction support module is also used to detect whether a user operation instruction is received after the duration of the gaze behavior exceeds a third time threshold. If so, the current gaze object is selected as the target and the user operation instruction is executed.

[0015] The above solution supports a human-computer interaction method of "gaze selection + quick operation", which allows controllers to perform operations such as information locking, screen focusing, and historical trajectory retrieval by looking at the target and inputting data on the console. This reduces visual search and click paths, and improves interaction efficiency and user-friendliness.

[0016] It also includes an attention deviation intervention module, which is used to determine whether there is a deviation in the user's attention based on the gaze sequence and the region of interest mapping sequence, and to intervene when a deviation exists.

[0017] The above approach can effectively improve safety control by detecting attention deviations and intervening when deviations occur.

[0018] The attention bias intervention module performs the following operations:

[0019] For each region of interest in the region of interest mapping sequence, features are extracted based on the gaze sequence. These features include the most recent gaze interval, the duration of the most recent gaze, the number of gazes, the total gaze duration, the coordinates of the region center, and the region size ratio.

[0020] The features of all targets are stacked sequentially into an input matrix to obtain the target feature matrix. This represents the feature extracted from the m-th target at the current time t. Represents a matrix, where the number of rows is the number of targets N at the current time t. t The number of columns is the feature dimension d of each target;

[0021] Perform a one-dimensional convolution along the target dimension on the target feature matrix to obtain the hidden feature vector for each target. Conv1D is a one-dimensional convolution operation; These are the convolution kernel parameters, with a total of h channels, and the kernel size. ; ReLU is the activation function;

[0022] A fully connected mapping is performed on the hidden feature vector of each target to obtain the score of that target. This represents the score of objective m. It is a weight matrix. b1 is the hidden feature vector of the m-th target; b2 is the bias term;

[0023] Normalize the scores of all targets to obtain the probability that attention should be allocated to each target at this time. Let m represent the probability of the target m.

[0024] Sort the probabilities from largest to smallest, and select the target regions corresponding to the top L probabilities as candidate target regions;

[0025] Determine the coverage rate of the candidate target area that has been focused on. If the coverage rate is lower than a set threshold, it is determined that there is a deviation in user attention. Select the target area corresponding to the maximum probability from the candidate target areas and provide a focus prompt.

[0026] In the above scheme, a lightweight attention modeling and bias intervention mechanism based on CNN is used to predict the attention weight of each target in real time by integrating the spatial features of the region of interest and historical gaze behavior, and to judge whether the controller's attention allocation is reasonable. When attention bias is identified, the system will select suggested attention objects from the high-priority targets that have not been gazed at, and perform gentle intervention to guide the controller to reasonably adjust the allocation of visual resources.

[0027] It also includes a monitoring and early warning module, which identifies abnormal events through target detection results and, in conjunction with gaze sequences, determines in real time whether potential risks have received visual attention and triggers an early warning prompt if they have not received visual attention.

[0028] The monitoring and early warning module first determines whether an abnormal event exists by the following definition: Indicates a collection of legal facilities. Indicate target Spatial location coordinates, Indicate target If the area is a legitimate activity zone, then a warning will be displayed showing the target corresponding to the abnormal event. Then, the target is determined based on the gaze sequence. If the unattended time is greater than or equal to the fourth time threshold, a multimodal alert response is triggered.

[0029] The monitoring and early warning module also detects target areas whose gaze time is greater than the fifth time threshold based on the gaze sequence, and marks the target area for further verification and confirmation based on the target detection results to determine whether an abnormal event has occurred in the target area.

[0030] The above scheme constructs a dual-channel early warning and reverse compensation mechanism for abnormal targets. On the one hand, after the system detects abnormal objects such as drones or illegal vehicles, it continuously monitors whether they are covered by gaze. If they are not paid attention to for a long time, an early warning is automatically triggered to ensure that the warning information is perceived in a timely manner. On the other hand, if the system recognizes that the controller is continuously watching a certain area but the detection module does not find any abnormalities, it marks the area as a "area to be verified" and calls other modal sensing devices for auxiliary confirmation, thereby improving the system's risk detection capability and safety fault tolerance level.

[0031] Compared with existing technologies, this invention constructs a closed-loop structure of "gaze perception—information reconstruction—intelligent prompting—feedback intervention," realizing a human-machine collaboration mechanism driven by controller gaze behavior, thereby significantly improving the system's real-time response capability, information presentation efficiency, and operational safety assurance level. It has the following technical effects:

[0032] 1) Constructing a real-time perception mechanism based on eye-tracking behavior to achieve dynamic interaction and human-computer adaptation: By establishing a data processing flow for eye-tracking data and an asynchronous parallel mechanism with multiple threads and buffers, the system acquires the controller's gaze target, gaze duration, and transition patterns in real time. Combined with the region of interest mapping results, it establishes a spatial matching relationship between the gaze sequence and the target object. Based on this, the system can dynamically identify the user's current focus object and use this as a driving force to adjust interface elements. This breaks the original static interaction mode of "system output—manual reading" and realizes a closed-loop human-computer interaction process of "perception—judgment—response," improving information acquisition efficiency and interaction flexibility.

[0033] 2) Optimized Information Display and Gazing Interaction Support Based on Gaze Behavior: A gaze-driven information enhancement and rearrangement mechanism was designed. After identifying the target being gazed at by the controller, the system can highlight, enlarge, and expand the information of the target sign only. Simultaneously, a sign avoidance algorithm is introduced to automatically adjust its display position to avoid overlapping or obstructing other targets, thereby improving the salience and recognition efficiency of key target information and reducing visual interference caused by information redundancy. Furthermore, the system supports a "gazing selection + quick operation" mode. Controllers can perform interactive functions such as target locking, interface focusing, and historical trajectory retrieval by gazing at the target and using console shortcut keys, further shortening the operation path, improving response speed, and enhancing the system's user-friendliness and operational efficiency.

[0034] 3) Construct an attention rationality assessment mechanism based on gaze behavior modeling to assist in judging the attention allocation status: Introduce a lightweight convolutional neural network model to construct a temporal feature matrix using the spatial features of all targets in the video frame and the gaze history. Infer the target attention allocation probability for each frame to determine whether the controller has missed gazing on high-priority areas. If the attention allocation is unreasonable, the system will proactively suggest key targets to focus on, enabling timely intervention in attention deviations and effectively reducing safety risks caused by missed observations.

[0035] 4) A gaze-driven anomaly alert mechanism enhances the system's responsiveness to high-risk situations: For anomaly target identification, this invention introduces a dual-trigger mechanism: upon detecting an anomaly target, the system not only highlights it on the screen but also continuously monitors whether the target receives a gaze response. If it remains unattended for an extended period, the system escalates the alert through multimodal channels, including visual and audio, ensuring timely reception of the alert. Furthermore, a reverse attention compensation mechanism is introduced, which analyzes the areas where controllers have gazed for extended periods to infer potential risks, further enhancing the system's safety tolerance. Attached Figure Description

[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a block diagram of the result of the gaze-driven remote tower intelligent human-computer interaction system in the embodiment.

[0038] Figure 2 This is a flowchart of the eye-tracking data processing module.

[0039] Figure 3 This is a simplified illustration of the signage's display effect.

[0040] Figure 4 This is a rendering of the new signage.

[0041] Figure 1 The following modules are marked: 10 - Target mapping module; 20 - Eye-tracking data processing module; 30 - Information enhancement and interaction support module; 40 - Attention deviation intervention module; 50 - Monitoring and early warning module. Detailed Implementation

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0043] In the field of air traffic control, maintaining efficient and accurate situational awareness and target identification capabilities is a core requirement for ensuring flight safety. With the widespread application of remote control tower technology, controllers need to achieve real-time monitoring and command of airport operations from remote environments through multi-screen, high-definition video surveillance. Existing remote control tower auxiliary systems generally adopt fixed information display and static interaction modes, which cannot dynamically adjust interface information according to the controller's line of sight behavior, nor can they proactively identify the controller's attention distribution and perception gaps. This can lead to high-priority targets being overlooked and delayed responses to abnormal events, posing significant safety hazards. Especially in complex operating environments with multiple concurrent targets and frequent anomalies, traditional systems lack a coupling mechanism between "human state and interface response," making it difficult to achieve optimal intelligent assistance and human-machine collaboration.

[0044] To address the aforementioned issues, this embodiment provides a gaze-driven remote control intelligent human-computer interaction system. It introduces gaze-driven attention guidance, anomaly warning, information enhancement, and interaction enhancement mechanisms to construct a human-computer interaction closed loop centered on gaze behavior. This system can identify the controller's gaze targets and distribution patterns in real time, dynamically adjust the display of target information, and achieve visual enhancement based on the controller's visual intent. Simultaneously, it possesses real-time detection and intervention capabilities for attention deviations and integrates video detection and gaze feedback to construct a dual-channel anomaly monitoring and warning mechanism, thereby improving cognitive reliability and operational safety in multi-target environments.

[0045] Specifically, please refer to Figure 1 This embodiment provides a gaze-driven remote control intelligent human-computer interaction system, including a target mapping module, an eye-tracking data processing module, an information enhancement and interaction support module, an attention deviation intervention module, and a monitoring and early warning module, forming a closed-loop processing flow of "perception-matching-judgment-response". By fusing radar and video data, a region of interest mapping sequence is constructed; simultaneously, eye-tracking data is collected and processed to generate a gaze sequence. The two types of data are matched in the spatial dimension to drive dynamic highlighting of signage information, gaze selection interaction, and real-time evaluation of the rationality of attention allocation. If an attention deviation is detected, the system actively outputs suggested targets to focus on based on a pre-trained lightweight convolutional neural network model; in the event of an abnormal event, the system combines the gaze state to determine whether to trigger a multimodal warning, while also supporting a reverse attention compensation mechanism to achieve higher response priority and reliability of risk identification. The entire interaction system improves the information perception efficiency, interaction convenience, and safety assurance capabilities in multi-target dense scenes under a remote control tower.

[0046] More specifically, the target mapping module first utilizes a neural network model based on the YOLO (You Only Look Once) architecture as the target detection model to perform real-time detection of aircraft in the tower's external video footage. To meet the deployment requirements of the remote tower system under conditions of limited computing resources, this invention adopts the YOLOv7-tiny lightweight model architecture, reducing computational complexity and model size while maintaining detection accuracy and improving inference speed. The output of the target detection model is multiple detection boxes (boundary boxes of detected targets) and their corresponding class probabilities D, denoted as:

[0047]

[0048] Where, x i y i w represents the center coordinates of the detection box for the i-th target; i h i These are the width and height of the bounding box, respectively; For category confidence scores; N t This represents the number of targets detected in the current video frame t (i.e., at the current time t).

[0049] Simultaneously, the target mapping module synchronously accesses airport surface surveillance radar (SMR) data to supplement the motion status and position information of ground targets. The target positions provided by surface surveillance radar data are typically given in Cartesian or polar coordinate systems. The radar observation result A for the j-th radar target is defined as:

[0050]

[0051] Where, x j y j Let v be the position coordinates of the j-th radar target; j ID is the instantaneous velocity vector of the j-th radar target. j M is the unique radar identifier for the j-th radar target; t This represents the number of radar targets identified by the radar at the current time t.

[0052] To achieve spatial matching and fusion between the detection results of the YOLO target detection model and radar data, a homomorphic registration function based on perspective projection and coordinate mapping is established here, denoted as:

[0053]

[0054] Wherein, mapping function The perspective transformation function, which maps from the radar coordinate system (geographic coordinate system) to the image coordinate system (video pixel space), can be obtained using the camera's extrinsic matrix and the radar system calibration parameters. The mapped radar target point... Projected onto the video image, and compared with the detection bounding box output by YOLO. Perform spatial overlap judgment. If the overlapping regions satisfy... If they are identified as the same target, information fusion is performed, which means integrating the information obtained from target detection with the information obtained from radar data. Among them, BB i BB is the detection bounding box for the i-th target. j Let θ be the projection area of ​​the j-th radar target; θ is the matching threshold, which can be adjusted according to the actual application scenario, with a default value of 0.5. If it is determined that it is not the same target, for example, if the target detection result is an uncontrolled intrusion target such as a bird (not registered in the airport database), then the detection result of the target detection model is retained.

[0055] The radar target is a target in the set of legitimate facilities registered in the airport database. If the target detection result matches the radar data successfully, it means that the detected target is a target in the set of legitimate facilities registered in the airport database. Otherwise, it does not belong to the set of legitimate facilities and is called an uncontrolled intrusion target.

[0056] Then, based on the fusion results of target detection and radar data, the boundary data of Areas of Interest (AOIs) are continuously constructed, forming a temporally continuous sequence of AOI mappings. In this paper, an AOI refers to an airborne or ground target area identified through target detection and radar data matching, where the radar data contains a corresponding target. Each AOI typically corresponds to a target area of ​​a target such as an aircraft or vehicle registered in a legitimate facility set in an airport database. Its spatial boundary is the detection box of that target in the video image, which can be dynamically updated over time.

[0057] Define the region of interest mapping sequence in the current video frame t. for:

[0058]

[0059] Among them, AOI k This represents the k-th region of interest. Indicates the bounding box coordinates and dimensions of the region of interest in the current frame image, ID k N is the unique identifier for this region of interest. t This represents the number of targets detected in the current video frame t.

[0060] The region of interest mapping sequence provides a set of spatial target locations and their unique identifiers that are dynamically updated over time, providing spatial benchmarks and matching support for subsequent processing steps such as the attention deviation intervention module and the monitoring and early warning module.

[0061] To meet the real-time and high-frequency data stream processing requirements of the interactive system, the eye-tracking data processing module adopts a dual-buffer and asynchronous thread mechanism to decouple the acquisition, processing, and output tasks, thereby achieving concurrent non-blocking processing. (See also...) Figure 2 The eye-tracking data processing module is configured with input and output buffers. The input buffer stores the denoised eye-tracking data, while the output buffer stores the recognized gaze point data stream. Through thread partitioning, the eye-tracking data processing module internally uses three parallel asynchronous threads.

[0062] 1) Denoising thread T1

[0063] First, the thread continuously receives eye-tracking data from the eye tracker device, which is the raw eye-tracking coordinate time series data.

[0064] Then, to address the common issues in eye-tracking data such as high-frequency noise, momentary frame drops, and blink interference, a weighted moving average filter is used for noise reduction. The original eye-tracking coordinate time series of the input is defined as follows: ,in Let time t o The coordinates of the line of sight position are recorded below.

[0065] A Gaussian weighted filter with a window radius of R is introduced to perform noise reduction calculations on the gaze position coordinates of each frame:

[0066]

[0067] in, The weights are Gaussian kernels, satisfying... t is used to enhance the consistency of data within the time neighborhood. t This represents the timestamp corresponding to the coordinates of the gaze position. This method can effectively suppress high-frequency noise and preserve the changing trend of the gaze during the switching process, preventing delays caused by excessive smoothing of information.

[0068] Finally, the smoothed eye-tracking time series is obtained. The data is written to the input buffer in timestamp order. For ease of distinction and description, the smoothed eye-tracking time series can be referred to as the gaze position series.

[0069] 2) Gaze behavior recognition thread T2

[0070] First, the thread reads line-of-sight position sequence data segments from the input buffer in real time.

[0071] Then, to avoid triggering subsequent module errors (such as error highlighting or frequent switching) due to unstable behaviors such as saccades and jumps, a lightweight fast gaze detection logic is introduced here—a low-latency improved scatter-duration algorithm (Fast-I-DT)—to determine gaze behavior in real time within the sliding window.

[0072] Specifically, define a duration of Within the first time window, determine whether the maximum Euclidean distance between all fixation points (the geometric centers of all video frames within the first time window) is less than the spatial dispersion threshold.

[0073]

[0074] Among them, D th The spatial dispersion threshold is determined based on the pixel distance of the tower exterior display in the actual application scenario; t n It is the starting time point of the first time window; p and q represent the line-of-sight points at any two moments within the first time window; These represent the denoised coordinates of the line-of-sight points p and q, respectively.

[0075] At the same time, the start and end times of the first time window must satisfy the first time threshold T. th :

[0076] ;

[0077] Among them, T th It is the first time threshold, which is the minimum duration required for the system to determine a segment of eye movement as "fixation", and is generally set to 100ms.

[0078] If both conditions are met, it is considered a gaze action, the gaze point is output, and a gaze point sequence F is constructed. Q :

[0079]

[0080] in, For the gaze center position, t q Q is the gaze start timestamp, and Q is the number of gaze points identified within the first time window.

[0081] Each gaze segment is continuously recorded as a data point and appended to the output buffer, thus realizing the time serialization of gaze points.

[0082] After obtaining the gaze point, it can be spatially matched with the region of interest (ROI) for subsequent use. For each gaze point identified at the current moment, it is checked from the ROI mapping sequence of the current frame to see if it falls within the bounding box of a certain ROI. If the gaze point coordinates are inside the bounding box of a ROI or the distance from its center is less than a set threshold, the gaze action is considered to have hit the ROI, and a binding relationship is established between the gaze point and the ROI, thus completing the matching.

[0083] 3) Merge output thread T3

[0084] First, the thread reads the foveation stream from the output buffer in real time.

[0085] Then, to further improve the stability of gaze point output and avoid frequent switching and redundant output caused by slight fluctuations in eye movement or algorithm errors, a gaze point merging strategy is introduced here. This strategy determines whether a newly identified gaze point can be merged with the previous output gaze point based on spatial proximity and temporal continuity, and only executes the output action when it is determined to be a completely new gaze behavior.

[0086] Specifically, whenever the gaze behavior recognition thread T2 identifies a new gaze point f new =(x new ,y new ,t new When the new gaze point is reached, it is written to the output buffer. The merging output thread T3 reads the new gaze point from the output buffer and merges it with the previous output gaze point. Compare the results to determine if the spatial proximity condition is met:

[0087]

[0088] in, The spatial proximity threshold is determined based on the pixel distance of the tower exterior display in the actual application scenario.

[0089] At the same time, determine whether the time continuity condition is met:

[0090]

[0091] in It is the time interval threshold, usually set to 100ms.

[0092] If both of the above conditions are met, then the two fixation points are merged into a new fixation point f. new And suppress intermediate redundant output. Otherwise, retain it as an independent gaze record.

[0093] The final output is the continuous gaze sequence F: .

[0094] The three threads mentioned above coordinate access to the buffer using shared memory to ensure data consistency and synchronization stability. The entire double-buffered + three-threaded structure supports stable operation at high-frequency (≥60Hz) eye-tracking acquisition rates, with an average processing latency controlled within 100ms.

[0095] The Information Enhancement and Interaction Support module aims to implement a "what you see is what you get" information presentation strategy based on gaze behavior, thereby improving the accessibility of target information and the efficiency of interaction in remote control scenarios. The input consists of a real-time updated region of interest mapping sequence and a sequence of gaze points that have been matched (spatial location matching between gaze points and regions of interest), where each gaze point is bound to a unique target identifier.

[0096] First, by reading the real-time updated region of interest (ROI) mapping sequence, simplified labels are overlaid on the identified target regions to highlight key information about the target, such as flight numbers, while reducing information overload. The simplified labels can be drawn as semi-transparent text at preset positions above or to the side of the detection box, and are updated in real-time based on the target's motion trajectory. The display effect of the simplified labels is as follows: Figure 3 ( Figure 3 The flight number shown is an arbitrary and fabricated number, and is for illustrative purposes only.

[0097] Then, a second time threshold is set, and based on the sequence of matched gaze points, a gaze is considered valid if its duration is greater than or equal to the second time threshold. The second time threshold is greater than the first time threshold.

[0098] Once the gaze target is identified, i.e., for the target deemed to be in a valid gaze state, the system will perform information enhancement operations on its corresponding simplified sign. Specifically, the simplified sign is first augmented to obtain a new sign, and then the information in the new sign is highlighted (i.e., the sign is in an information-enhanced state). For example, the information in the new sign will be highlighted and presented through a smooth zoom-in method to display several configurable additional information items. The types of information presented can be flexibly set and customized according to the user's actual usage scenario and operating habits. For example, Table 1 below provides an example.

[0099]

[0100] Enhanced display of new signage containing additional information, such as Figure 4 ( Figure 4 All information in this document is fabricated and is for illustrative purposes only. Figure 4In the signage, CCA2705 represents the flight number, consisting of the airline code and numbers; A6344 represents the transponder code, consisting of A plus a four-digit octal number; TA indicates the current status is taxiing; TM-01D represents the planned arrival / departure route number; ZHHH-ZUUU represents the departure airport code to the arrival airport code; 0712-0915 represents the estimated departure time to the estimated arrival time; S102 represents the parking stand, consisting of S plus the parking stand number; 089 represents the cruising altitude layer, in 100 meters; 040 represents the ground taxiing speed, in knots; and A320 M represents the aircraft type, consisting of the aircraft type code and weight class. The signage may also display... This indicates that the identifier is manually fixed and will be displayed when manually fixed.

[0101] To prevent bright signs from obscuring other targets, this module introduces a sign avoidance and dynamic arrangement mechanism. The interactive system maintains a set of candidate display positions for each target sign (such as upper left, upper top, upper right, etc.), and automatically selects the optimal non-overlapping position for display based on the distribution of other signs in the current frame, target density, and screen boundaries. This ensures that the focus target information is always in a clear and prominent area, improving the visual priority of key task targets.

[0102] Specifically, the interactive system maintains a set of coordinates for all displayed signs in the current frame in real time. New signs are displayed by default in a standard relative position (e.g., top left corner). If the overlap with an existing sign exceeds a preset threshold (e.g., 30%), other candidate display positions are tried sequentially until the optimal solution without occlusion is found. This avoidance mechanism effectively ensures the sign's recognizability, positional uniqueness, and visual neatness when multiple target information is presented simultaneously.

[0103] In addition, this module introduces an interaction command mechanism based on gaze behavior, which allows users to trigger a shortcut input device after gazing at a target, enabling various interactive functions for that target.

[0104] For example, a specific approach could be as follows: When an controller focuses on a target and the duration of the focus exceeds a third time threshold (i.e., the duration of the focus is greater than or equal to the third time threshold), pressing the corresponding key on the console (e.g., F1-F10 or a light touch on a physical key) will cause the system to select the currently focused object and trigger relevant operation commands, including but not limited to: maintaining the information enhancement state (without this operation, the information enhancement state will continue until the controller is detected focusing on another target, triggering information enhancement for that target. However, if the information enhancement state of the current object is fixed using this interactive command mechanism, it will remain indefinitely); focusing and zooming the video feed to the target area; and retrieving historical data of the target (dispatch trajectories, command records, delay logs, etc., for situation review).

[0105] This "gaze selection + quick operation" interaction mode has advantages such as high efficiency, low load and strong intuitiveness. Users can complete the input of instructions by looking at the target and pressing the button, without the need for mouse selection. This significantly shortens the operation path, reduces cognitive burden, and improves the response speed and accuracy in multi-target and high-density scenarios. It is particularly suitable for high-intensity air traffic control environments such as remote towers, and helps to improve the human-centeredness and operational reliability of the system.

[0106] The attention deviation intervention module, as an optional configuration, aims to provide timely warnings when a target that should be focused on is not being focused on, so as to guide controllers to make reasonable adjustments to the allocation of visual resources and improve the level of safety assurance.

[0107] In this embodiment, the attention deviation intervention module provides a lightweight convolutional neural network (CNN) attention model based on gaze behavior and target spatial features. This model is used to determine in real-time whether the air traffic controller's gaze allocation is reasonable during remote tower operations, and outputs suggested key targets to focus on when attention is inappropriate. Finally, the attention deviation intervention module can provide gentle prompts when it detects a deviation in the controller's attention. This attention model does not require manual setting of target priorities and does not rely on specific flight information. Instead, it is trained using historical safe operation data, allowing it to automatically learn a reasonable matching pattern between gaze behavior and spatial target distribution, forming universal attention assessment rules.

[0108] The attention model was obtained through supervised training using a convolutional neural network (CNN) based on training samples. The training data was derived from historical accident-free periods in the remote control tower system, and the training samples were constructed as follows:

[0109] 1) Use the gaze trajectory and target area matching relationship during normal operation as a reasonable label for attention;

[0110] 2) By randomly shuffling the gaze targets, “pseudo-irrational” samples are constructed to improve the attention model’s ability to identify attention biases.

[0111] In practical applications, the trained attention model processes the input data according to the following steps:

[0112] (1) Feature extraction

[0113] With a fixed frequency of 500ms as the period, in the second time window T w =Attention assessment calculations are performed every 10 seconds, that is, the gaze distribution in the most recent 10 seconds is analyzed every half second to ensure a balance between real-time performance and robustness.

[0114] First, the gaze sequence F={f1,f2,…,f1} output by the eye-tracking data processing module is received synchronously.new} and the region of interest mapping sequence output by the target mapping module Among them, B k This represents the center coordinates and size of the bounding box of the k-th region of interest (AOI) in the current frame image, independent of specific flight information or identifiers. k The features shown in Table 2 below were extracted.

[0115]

[0116] Then, the features of all targets are stacked sequentially into an input matrix to obtain the target feature matrix:

[0117]

[0118] Among them, X m (t) represents the m-th (m=1,2,…,N)-th node. t The features extracted from the targets at the current time t (each feature is a vector in the matrix) are of length d dimensions, as shown in Table 2, where d=6; This means that the final result is a matrix, with the number of rows being the number of targets N at the current time t. t The number of columns is the feature dimension d of each target.

[0119] (2) Local convolutional coding

[0120] Next, a one-dimensional convolution operation is performed on the feature matrix formed by all targets in the current frame along the direction of target arrangement (i.e., the row direction of the matrix) to extract the relative feature relationships between targets:

[0121]

[0122] Conv1D is a one-dimensional convolution operation; These are the convolution kernel parameters, with a total of h channels, and the kernel size. ; It is the hidden feature vector of each target in the output; ReLU is the activation function.

[0123] (3) Calculation of target priority score

[0124] Then, a fully connected mapping is performed on the hidden feature vector H1(t) of each target to obtain the original score of that target:

[0125]

[0126] Among them, Z m (t) is the raw score of the attention model for the "attention priority" of target m at the current time t; It is a weight matrix used to map hidden features to a score; b1 is the hidden feature vector of the m-th target; b2 is the bias term used to adjust the overall score offset.

[0127] (4) Attention probability normalization

[0128] Using Softmax to normalize the scores of all targets, we obtain the probability that attention should be assigned to each target at this time:

[0129]

[0130] in, It is the "relative probability" or "importance" of target m that it should receive attention, and all targets The sum is 1, that is: .

[0131] (5) Reasonableness assessment and key target output

[0132] To improve the system's robustness in detecting controller attention allocation biases, this module selects the top 30% of regions with high attention probabilities as key monitoring targets based on the attention probabilities of all regions of interest in the current frame, and constructs a multi-target attention coverage judgment mechanism.

[0133] The total number of targets in the current video frames is N. t Based on the attention probability after Softmax normalization Before selection The target regions are selected as candidate target regions to form a high-priority region set:

[0134]

[0135] The system in the third time window T L Within 10 seconds, check whether each candidate target area has been observed by the controller.

[0136] Define gaze coverage For the third time window T L The percentage of regions that have been gazed upon among the top L high-priority regions:

[0137] ;

[0138] When the coverage rate is lower than the preset threshold (Generally taken as 0.35), indicating that the fixation does not fully cover the important regions inferred by the model, and is judged as unreasonable attention allocation. The number of fixated regions is the number of target regions judged as valid fixations among the calculated high-priority regions. Define the reasonableness judgment function A(t):

[0139]

[0140] When A(t)=0, the system, based on the judgment that the attention allocation is unreasonable, selects the target with the highest attention probability from the high-priority targets that have not been observed as the current suggested target for attention and outputs its bounding box. .

[0141] (6) It is recommended to pay attention to the target's gentle prompts.

[0142] When the attention rationality judgment function outputs A(t)=0, meaning that there is a high-priority region with insufficient attention coverage within the time window, the system will consider the inferred suggested attention target region. Provide attention prompts, for example, within the target bounding box. It is surrounded by a semi-transparent blue soft light border, which flashes at a frequency of 1Hz and lasts for 5-10 seconds or fades away after being looked at.

[0143] This notification mechanism ensures a balance between the effectiveness of the notifications and the user-friendliness of the interface, respects the controller's subjective judgment, and avoids excessive interference.

[0144] This module employs a shallow network structure, resulting in a small number of model parameters and low inference latency (<100ms) after training. This meets the real-time requirements of remote control tower systems for high-frequency attention assessment tasks. By continuously monitoring controllers' gaze behavior and quickly identifying attention deviations, this module can effectively assist controllers in timely detection and attention to key targets in complex traffic situations, reducing the risk of omissions and improving operational safety and monitoring efficiency.

[0145] The monitoring and early warning module is an optional configuration. Its purpose is to improve the system's risk detection capability and safety tolerance level through a dual-channel early warning and reverse compensation mechanism for abnormal targets. That is, when the target detection identifies an abnormal object but it is not observed, or when the controller continuously observes a certain area but the target detection does not find any abnormality, an early warning is issued.

[0146] In this embodiment, to address abnormal events in non-routine operational scenarios, the monitoring and early warning module introduces a safety monitoring and early warning method that integrates video analysis and controller gaze behavior. This method identifies abnormal events using target detection results from the target mapping module, combines this with gaze trajectories generated from eye-tracking data, and determines in real-time whether potential risks have garnered visual attention, triggering corresponding early warning prompts accordingly. Compared to attention deviation intervention modules used in conventional scenarios, this module has a higher response priority and is suitable for rapid identification and intervention in emergency safety situations.

[0147] First, targets identified by the target detection model but not matched in the surface radar data (i.e., not in the set of legitimate facilities registered in the airport database) are marked as uncontrolled intrusion targets. Uncontrolled intrusion targets include, for example, birds and small drones, which are not in the set of legitimate facilities registered in the airport database. The goal in the process.

[0148] Then, by combining target detection and rule constraints, the system identifies violations by controlled objects, such as vehicles or personnel mistakenly entering restricted areas like runways or taxiways. The identification logic is as follows: if a target... It belongs to a set of legal facilities, but the target's current spatial location... If it falls within the area where the activity is legal, then it is considered abnormal behavior.

[0149] The definition of an abnormal event is as follows:

[0150]

[0151] in, This is a collection of legally registered facilities in the scene database. For the goal The legal activity area.

[0152] For each detected anomalous target The monitoring and early warning module highlights warning prompts in the video footage, such as highlighting the target area with a red frame and flashing at a frequency of 4Hz, forming a highly conspicuous visual marker to attract the attention of controllers.

[0153] At the same time, a fourth time threshold is set. Continuously monitor the target In the coverage state of the gaze sequence, if an abnormal target From the time of being monitored to the present moment, the unattended time is greater than [missing information]. If an event is detected as an unattended anomaly, the monitoring and early warning module will immediately trigger a multimodal alert response, such as issuing an audio warning through the console speaker or headphone channel, ensuring that controllers can quickly locate the source of the anomaly and take appropriate control measures. If necessary, the anomaly information can also be synchronously uploaded to the superior monitoring system or dispatch center to achieve cross-position risk sharing and response coordination. Unattended anomalies are considered more urgent than attention-related inappropriate events, meaning they have a higher response priority. Therefore, to prevent less urgent attention-related intervention modules from affecting the functionality of this module, the function of the attention-related intervention module can be temporarily suppressed after an unattended anomaly is identified.

[0154] In addition, this module introduces a reverse attention compensation mechanism: if the eye-tracking data processing module shows that a target area has been gazed at for a long time (gaze time greater than the set fifth time threshold), but no abnormal event is detected, it may mean that there is a potential risk in the area that has not been captured by the algorithm. At this time, the system automatically marks the area as a region to be reviewed and calls other available modal pavement auxiliary sensing devices (such as infrared thermal imaging) to confirm the area.

[0155] This module constructs a collaborative monitoring loop of "algorithm-human eye" by fusing and compensating information from the two channels of "visual perception" and "video recognition," effectively improving the remote control tower's response capability and fault tolerance redundancy to sudden anomalies and potential conflicts.

[0156] To verify the effectiveness and accuracy of the interactive system of this invention, an experimental scenario simulating a remote control tower environment was designed. The following is a detailed description of the experimental conditions, steps, and conclusions.

[0157] Experimental conditions: The experiment used the Tower Client simulator to build a standard remote control tower visual environment and generate realistic airport flight operation video footage. The experiment was equipped with the remote control intelligent human-computer interaction system proposed in this invention, which supports gaze-driven eye movement behavior recognition, video object detection, attention modeling, anomaly detection and warning, information enhancement, and interaction enhancement functions.

[0158] Computer configuration: CPU: Intel Core i5-10400 @ 2.90GHz; GPU: NVIDIA GeForce RTX3070; RAM: 16GB.

[0159] The eye-tracking device used for real-time recording of subjects' gaze behavior is the TOBII PRO FUSION. (Sampling frequency 250 Hz, error ≤0.5°)

[0160] The subjects of the experiment were 12 participants with an air traffic control background.

[0161] To evaluate the actual effectiveness of the present invention, three experimental conditions were set, as shown in Table 3 below:

[0162]

[0163] Experimental Procedure: Participants performed simulated air traffic control tasks under three different experimental conditions, each lasting one hour to fully simulate the real-world working environment of long-duration, multi-objective, and high-information-load conditions. Each participant completed three rounds of experiments, with the order randomly shuffled to avoid sequence effects. The tasks covered typical air traffic control operations such as flight taxiing guidance, takeoff and landing sequence management, and anomaly response. In each round, the system randomly generated and scheduled multiple flight objects, dynamically setting different task objectives.

[0164] The evaluation indicators for the experimental results are shown in Table 4 below:

[0165] Table 4: Explanation of Evaluation Indicators

[0166]

[0167] The experimental results are shown in Table 5 below:

[0168]

[0169] As shown in Tables 4 and 5, the remote tower intelligent human-computer interaction system based on the fusion of eye-tracking behavior and video detection proposed in this invention can effectively improve the information acquisition efficiency, attention allocation rationality, and response speed of controllers in complex and high-density operation scenarios. It is significantly better than existing fixed display and traditional early warning technologies and has clear practical value and promotion prospects.

[0170] Based on the basic monitoring functions of remote control towers, this invention further constructs an information enhancement and security framework dominated by "attention perception". This not only improves the information processing efficiency and human-computer interaction intelligence of the system, but also significantly enhances the sensitivity and accuracy of abnormal situation perception and response. It solves key problems in the prior art such as low interaction efficiency, rigid information display, inability to recognize attention, and lack of feedback mechanism for abnormal prompts.

[0171] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A gaze-driven remote control intelligent human-machine interaction system for air traffic control, characterized in that, It includes a target mapping module, an eye-tracking data processing module, and an information enhancement and interaction support module. The target mapping module is used to fuse video data and radar data to construct a region of interest mapping sequence. The eye-tracking data processing module is used to process user eye-tracking data to generate a gaze sequence. The information enhancement and interaction support module is used to match the region of interest mapping sequence and the gaze sequence in the spatial dimension to drive the dynamic enhancement display of signage information.

2. The gaze-driven remote control intelligent human-computer interaction system for towers according to claim 1, characterized in that, The target mapping module first performs target detection based on video data to obtain the bounding boxes of the detected targets, and obtains the position coordinates of the radar targets based on radar data. Then, when the overlapping area of ​​the bounding boxes of the detected targets and the position coordinates of the radar targets is determined to be the same target, the target detection results and the radar data results are fused. Based on the fusion results, region of interest boundary data is constructed, forming a temporally continuous region of interest mapping sequence. This represents the k-th region of interest. This represents the bounding box coordinates and dimensions of the region of interest in the current video frame image, x k y k w represents the center coordinates of the k-th region of interest. k h k These are the width and height of the k-th region of interest, respectively, and ID. k N is the unique identifier for this region of interest. t This represents the number of targets detected in the current video frame t.

3. The gaze-driven remote control intelligent human-computer interaction system for towers according to claim 1, characterized in that, The eye-tracking data processing module is configured with an input buffer, an output buffer, a denoising thread, a gaze behavior recognition thread, and a merging output thread. The denoising thread denoises the real-time acquired eye-tracking data and writes the resulting gaze position sequence into the input buffer. The gaze behavior recognition thread reads the gaze position sequence data segment from the input buffer, identifies whether a gaze behavior has occurred, and writes the gaze point of the gaze behavior into the output buffer when a gaze behavior is identified. The merging output thread reads the gaze point from the output buffer and determines whether two adjacent gaze points meet the merging condition. If so, it merges the two adjacent gaze points, and all the retained gaze points form the gaze sequence.

4. The gaze-driven remote control intelligent human-computer interaction system for towers according to claim 3, characterized in that, The information enhancement and interaction support module first reads the region of interest mapping sequence and overlays a simplified sign on the identified target region. Then, it determines whether the gaze behavior is a valid gaze based on a second time threshold. If so, it expands the information of the simplified sign to obtain a new sign and highlights it. It also determines whether the overlap rate between the proposed display position of the new sign and the display position of the existing sign exceeds a preset threshold. If so, it selects a candidate display position to display the new sign.

5. The gaze-driven remote control intelligent human-computer interaction system for towers according to claim 4, characterized in that, The information enhancement and interaction support module is also used to detect whether a user operation instruction is received after the duration of the gaze behavior exceeds a third time threshold. If so, the current gaze object is selected as the target and the user operation instruction is executed.

6. The gaze-driven remote control intelligent human-machine interaction system for any one of claims 1-5, characterized in that, It also includes an attention deviation intervention module, which is used to determine whether there is a deviation in the user's attention based on the gaze sequence and the region of interest mapping sequence, and to intervene when a deviation exists.

7. The gaze-driven remote control intelligent human-machine interaction system for towers according to claim 6, characterized in that, The attention bias intervention module performs the following operations: For each region of interest in the region of interest mapping sequence, features are extracted based on the gaze sequence. These features include the most recent gaze interval, the duration of the most recent gaze, the number of gazes, the total gaze duration, the coordinates of the region center, and the region size ratio. The features of all targets are stacked sequentially into an input matrix to obtain the target feature matrix. This represents the feature extracted from the m-th target at the current time t. Represents a matrix, where the number of rows is the number of targets N at the current time t. t The number of columns is the feature dimension d of each target; Perform a one-dimensional convolution along the target dimension on the target feature matrix to obtain the hidden feature vector for each target. Conv1D is a one-dimensional convolution operation; These are the convolution kernel parameters, with a total of h channels, and the kernel size. ; ReLU is the activation function; A fully connected mapping is performed on the hidden feature vector of each target to obtain the score of that target. This represents the score of objective m. It is a weight matrix. It is the hidden feature vector of the m-th target; b2 is the bias term; Normalize the scores of all targets to obtain the probability that attention should be allocated to each target at this time. Let m represent the probability of the target m. Sort the probabilities from largest to smallest, and select the target regions corresponding to the top L probabilities as candidate target regions; Determine the coverage rate of the candidate target area that has been focused on. If the coverage rate is lower than a set threshold, it is determined that there is a deviation in user attention. Select the target area corresponding to the maximum probability from the candidate target areas and provide a focus prompt.

8. The gaze-driven remote control intelligent human-machine interaction system for towers according to claim 6, characterized in that, It also includes a monitoring and early warning module, which identifies abnormal events through target detection results and, in conjunction with gaze sequences, determines in real time whether potential risks have received visual attention and triggers an early warning prompt if they have not received visual attention.

9. The gaze-driven remote control intelligent human-computer interaction system for towers according to claim 8, characterized in that, The monitoring and early warning module first determines whether an abnormal event exists by the following definition: Indicates a collection of legal facilities. Indicate target Spatial location coordinates, Indicate target If the area is a legitimate activity zone, then a warning will be displayed showing the target corresponding to the abnormal event. Then, the target is determined based on the gaze sequence. If the unattended time is greater than or equal to the fourth time threshold, a multimodal alert response is triggered.

10. The gaze-driven remote control intelligent human-computer interaction system for towers according to claim 9, characterized in that, The monitoring and early warning module also detects target areas whose gaze time is greater than the fifth time threshold based on the gaze sequence, and marks the target area for further verification and confirmation based on the target detection results to determine whether an abnormal event has occurred in the target area.

Citation Information

Patent Citations

  • Air traffic control tower visual display method based on HUD enhancement

    CN113990113A