Eye movement interaction air-ground cooperative command system and method
Patent Information
- Application Number
- CN202511375455.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-09-29
AI Technical Summary
目前,传统的指挥方式是指挥员在直升机内通过手动操作(如通过平板、键盘、触控屏等)方式向地面传达指令,从实际实施可以发现,这种方式存在操作繁琐、指挥效率较低的问题,尤其是在高动态的复杂作战环境中,指挥员需要在短时间内处理大量信息,上述传统的指挥方式显然无法满足高效决策的需求
[0020]本发明在空地协同作战中显著提升了空中指挥员的指挥效率和执行准确性,使指挥变得直观、便捷与高效,适用于指挥员在直升机等空中飞行设备内的指挥,具有如下有益效果:
Smart Images

Figure CN122837618A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a command system for air-ground coordinated operations, and more particularly to an air-ground coordinated command system and method based on eye-tracking interaction. Background Technology
[0002] With the evolution of modern warfare and technological advancements, helicopters are playing an increasingly important role in air-ground coordinated operations. To improve operational efficiency, commanders need to maintain real-time and effective control and command of ground forces from within the helicopter, enabling better coordination between ground and airborne helicopters. Currently, the traditional command method involves commanders manually transmitting instructions to the ground from within the helicopter (e.g., via tablets, keyboards, touchscreens). However, practical experience shows that this method is cumbersome and inefficient, especially in highly dynamic and complex combat environments where commanders need to process large amounts of information in a short time. The aforementioned traditional command method clearly cannot meet the demands of efficient decision-making.
[0003] In recent years, eye-tracking technology has matured and has been widely applied in various human-computer interaction scenarios. Its high efficiency and intuitiveness can effectively free up hands and are suitable for operation in high-pressure and complex environments. However, the technology for enabling air-to-ground coordinated operations through airborne command is still in its early stages. How to make commanders' commands more intuitive, convenient, and efficient, and improve the command efficiency and execution accuracy of air-to-ground coordinated operations, is an urgent problem to be solved. Summary of the Invention
[0004] The purpose of this invention is to provide an eye-tracking interactive air-to-ground collaborative command system and method, which makes command by air commanders more intuitive, convenient and efficient, and significantly improves the command efficiency and execution accuracy of air-to-ground collaborative operations.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] An eye-tracking interactive air-to-ground collaborative command system includes a head-mounted device. The head-mounted device is equipped with a ground information display module, an eye-tracking infrared camera, a command interaction module, and an air-to-ground communication module, wherein:
[0007] The ground information display module is used to receive ground data transmitted by the data acquisition device, obtain ground information based on the ground data, and display the ground information in the form of a map.
[0008] The eye-tracking infrared camera is used to acquire images of the commander's eyes;
[0009] The command interaction module is used to lock onto targets by tracking the commander's gaze point based on the eye images and to obtain instructions by recognizing eye movements based on the eye images.
[0010] The air-to-ground communication module is used to send the instructions to the locked target to complete the command operation, and to receive feedback information from the target so that the ground information display module can update and display the ground information based on the feedback information.
[0011] An eye-tracking interactive air-to-ground collaborative command method, implemented based on the aforementioned eye-tracking interactive air-to-ground collaborative command system, includes the following steps:
[0012] 1) The commander puts on and activates the aforementioned head-mounted device, entering combat readiness;
[0013] 2) The ground information display module receives ground data transmitted by each of the data acquisition devices and displays the ground information obtained based on the ground data in a map form.
[0014] 3) The eye-tracking infrared camera acquires images of the commander's eyes and transmits these images to the command interaction module;
[0015] 4) The command interaction module determines whether the commander is looking at the eye image: if the commander is not looking at the eye, proceed to step 5); if the commander is looking at the eye, calculate the coordinates of the gaze point and determine the object of gaze. If the object of gaze is friendly troops, lock the object of gaze as the target and proceed to the next step; otherwise, return to step 3.
[0016] 5) The command interaction module determines whether the commander has issued an instruction based on the eye image: if it determines that the commander has not issued an instruction, it returns to step 3); if it determines that the commander has issued an instruction, it identifies the commander's eye movement behavior, obtains the instruction corresponding to this eye movement behavior, and proceeds to the next step.
[0017] 6) The air-to-ground communication module sends the command to the locked target and determines whether the target returns feedback: if there is feedback from the ground, proceed to the next step; otherwise, return to step 3).
[0018] 7) The air-to-ground communication module sends the returned information to the ground information display module to update and display the map based on the returned information, and then returns to step 3).
[0019] The advantages of this invention are:
[0020] This invention significantly improves the command efficiency and execution accuracy of air commanders in air-to-ground coordinated operations, making command more intuitive, convenient, and efficient. It is applicable to command from within helicopters and other aerial flight equipment, and has the following beneficial effects:
[0021] 1) Easy to operate. This invention utilizes eye-tracking technology, enabling commanders to select and confirm targets and instructions through eye movements (such as rapid double blinks) without using their hands or other complex operating devices. This significantly simplifies the command process and ensures accurate and timely system responses.
[0022] 2) High real-time performance and accuracy. The head-mounted device can display ground information in real time, quickly track and identify targets that the commander is looking at and lock onto them, and quickly transmit instructions to targets and receive feedback information to achieve real-time updates of ground information. This significantly improves the reaction speed of air-ground coordinated operations, enhances the accuracy of command execution, and ensures the command efficiency of air-ground coordination.
[0023] 3) Visualized and comprehensive information. The map interface presents the ground information of the area to be commanded in a visual way, which helps commanders quickly identify enemy and friendly forces and have a comprehensive understanding of the ground battlefield situation, thereby improving the accuracy of decision-making.
[0024] 4) Freeing up the commander's hands. During combat, freeing up the commander's hands allows them to perform other high-priority operations and improves their ability to respond to emergencies. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the composition of the eye-tracking interactive air-ground collaborative command system of the present invention.
[0026] Figure 2 This is a schematic diagram illustrating the implementation of the eye-tracking interactive air-ground collaborative command method of the present invention.
[0027] Figure 3 This is a schematic diagram illustrating the operation of the ground information display module.
[0028] Figure 4 This is a schematic diagram illustrating the operation of the command and interaction module.
[0029] Figure 5 This is a schematic diagram illustrating the operation of the air communication module. Detailed Implementation
[0030] like Figures 1 to 5 As shown, this invention proposes an eye-tracking interactive air-to-ground collaborative command system, which includes a head-mounted device 10. The head-mounted device 10 is equipped with a ground information display module 20, an eye-tracking infrared camera 30, a command interaction module 40, and an air-to-ground communication module 50, wherein:
[0031] The ground information display module 20 is used to receive ground data transmitted by the data acquisition equipment, obtain ground information based on the ground data, and display the ground information intuitively in the form of a map (two-dimensional or three-dimensional view);
[0032] An eye-tracking infrared camera 30 is used to acquire images of the commander's eyes;
[0033] The command interaction module 40 is used to lock onto targets by tracking the commander's gaze point based on eye images and to obtain instructions by recognizing eye movements based on eye images.
[0034] The air-to-ground communication module 50 is used to send commands to the locked target, complete command operations, and receive feedback information from the target so that the ground information display module 20 can update and display the ground information based on the feedback information (i.e., update the map).
[0035] In the actual design, the head-mounted device 10 includes a helmet worn on the commander's head, and a ground information display module 20 is installed inside the helmet. The ground information display module 20 includes a screen that displays ground information in the form of a map. When the head-mounted device 10 is worn on the commander's head, the screen is in front of his eyes. An eye-tracking infrared camera 30 installed inside the helmet is positioned opposite to the eyes. The screen is semi-transparent so that the commander can see the scene in front of him through the screen.
[0036] In practice, eye-tracking infrared cameras 30 are installed on both the left and right sides inside the helmet, and the eye-tracking infrared cameras 30 on both sides capture images of the left and right eyes respectively.
[0037] In this invention, the data acquisition equipment includes any one or a combination of several of the following: reconnaissance satellites, drones, or ground force equipment (such as command center equipment). Ground data includes terrain information, the location and distribution of friendly forces, the location and distribution of enemy forces, and the weapon and ammunition status of each side. Of course, the aforementioned data acquisition equipment and the ground data it collects are not limited to these. The ground information system processes and correlates the collected ground data, presenting it dynamically on a map interface to the commander's field of vision.
[0038] In this invention, the ground information display module 20 displays ground information in map form, which is the ground information of the area to be commanded below the commander's aerial position. The area to be commanded is either a rectangular area with a predetermined area or a circular area with a predetermined radius. In practical applications, the commander can be seated inside an aerial flight device such as a helicopter.
[0039] Based on the above-described eye-tracking interactive air-to-ground collaborative command system, this invention also proposes an eye-tracking interactive air-to-ground collaborative command method, comprising the following steps:
[0040] 1) The commander puts on and activates the head-mounted device, entering combat readiness;
[0041] 2) The ground information display module 20 receives ground data transmitted from each data acquisition device and displays the ground information obtained based on the ground data in the form of a map, presenting it in the commander's field of vision;
[0042] 3) The eye-tracking infrared camera 30 acquires images of the commander's eyes and transmits these images to the command interaction module 40;
[0043] 4) The command interaction module 40 determines whether the commander is looking at the eye image: if the commander is not looking at the eye, proceed to step 5); if the commander is looking at the eye, calculate the coordinates of the gaze point of the commander, and then determine the gaze object based on the position of the gaze point coordinates on the ground information display module 20, i.e. the screen. If the gaze object is friendly troops, lock the gaze object as the target and proceed to the next step; otherwise, if the gaze object is friendly or enemy troops or other positions, do not react and return to step 3).
[0044] 5) The command interaction module 40 determines whether the commander has issued an instruction based on the eye image: if it determines that the commander has not issued an instruction, it returns to step 3); if it determines that the commander has issued an instruction, it identifies the commander's eye movement behavior, obtains the instruction corresponding to this eye movement behavior from the behavior-instruction mapping table (such as blinking twice quickly to correspond to an order for the troops to advance, turning the eyeballs from left to right to correspond to a retreat order, and closing the eyes to correspond to ending the command), and proceeds to the next step;
[0045] 6) The air-to-ground communication module 50 sends the command to the locked target, i.e., to its own troops, and determines whether the target has returned information: if there is feedback from the ground, i.e. the target has sent a return message, then proceed to the next step; otherwise, return to step 3).
[0046] 7) The air-to-ground communication module 50 sends the returned information to the ground information display module 20 to update the map (or update the original ground information) and display it, then returns to step 3).
[0047] In actual design, in step 4), determining whether the commander is in a gaze state includes the following steps: the command interaction module 40 obtains a monocular image of the left or right eye from the eye image, detects the pupil center position from the monocular image, and if the pupil center position does not change for a set time (e.g., 2-3 seconds), it is considered that the commander is gazing at a point on the map, and then the coordinates of the gaze point are calculated; otherwise, it is considered that the commander is not gazing at the map.
[0048] The calculation of the gaze point coordinates includes the following steps:
[0049] The command interaction module 40 obtains left and right eye images from the eye images. Through pupil detection, it extracts the coordinates of the left pupil center from the left eye image and the coordinates of the right pupil center from the right eye image. It then obtains the gaze point coordinates based on the polynomial fitting curve between the centers of the left and right pupils and the screen position. The polynomial fitting curve between the centers of the left and right pupils and the screen position is obtained through the following steps:
[0050] Test participants wore a head-mounted device 10 and successively looked at the ground information display module 20, i.e., multiple locations on the screen (e.g., the four corners, the midpoints of the four sides, and the center of the screen, for a total of nine locations), obtaining multiple pairs of left and right eye images. For each pair of left and right eye images, the coordinates of the left eye pupil center (x1, y1) and the right eye pupil center (x2, y2) were extracted through pupil detection. The mapping relationship between the coordinates of the left and right eye pupil centers and the coordinates (x, y) of the positions being looked at on the screen was established according to the following formula:
[0051] x = a1X1 + a2y1 + a3x2 + a4y2 + a5,
[0052] y = b1x1 + b2y1 + b3x2 + b4y2 + b5,
[0053] In the formula, a1-a5 and b1-b5 are coefficients (constants);
[0054] In practice, the pupil center coordinates (x1, y1) and (x2, y2) are detected in the left and right eye images using image processing algorithms (such as image thresholding and edge detection). In addition, the accuracy of pupil detection can be improved by preprocessing such as filtering and denoising.
[0055] Based on the mapping relationship between the coordinates of the left and right pupil centers and the screen position coordinates (x, y) established by each pair of left and right eye images, a polynomial fitting curve is obtained. The coefficients of the polynomial fitting curve are determined by the least squares method, thereby establishing the polynomial fitting curve between the center of the left and right pupils and the screen position.
[0056] In practical applications, one or more friendly forces can be deployed. When there are multiple friendly forces, the commander needs to issue different instructions to each force in sequence to coordinate their joint operations. Similarly, one or more friendly forces can be deployed, and one or more enemy forces can be deployed.
[0057] In practical design, step 5) of eye movement behavior recognition includes the following steps: the command interaction module 40 obtains a sequence of monocular images within a certain time period from the eye images, i.e., a series of consecutive monocular images, and inputs the monocular image sequence into the deep learning model to obtain the eye movement behavior results. The deep learning model is obtained through the following steps:
[0058] Constructing a deep learning model: The deep learning model includes a three-dimensional convolutional neural network (3D-CNN) and a classification head. The 3D-CNN consists of multiple stacked 3D convolutional and pooling layers, performing convolution operations in the height, width, and time dimensions. It is used to extract temporal and spatial features based on the input monocular image sequence. The temporal and spatial features correspond to the temporal and spatial changes that occur during eye movement, respectively (e.g., the time of eye movement is greater than that of blinking, but the spatial change of blinking is greater than that of eye movement). The classification head includes fully connected layers and softmax layers, which are used to classify eye movement behavior based on the temporal and spatial features output by the 3D-CNN. It establishes a mapping relationship between spatiotemporal features and eye movement behavior (e.g., closing the eyes, blinking, eye turning left, turning right, etc.) and outputs the eye movement behavior results.
[0059] Training the deep learning model: The deep learning model is trained by inputting multiple monocular image sequences (training data), and the weights of the deep learning model are determined by the backpropagation algorithm. The difference between the predicted results and the true labels is reduced by calculating the cross-entropy loss, and gradient updates are performed based on the difference to complete the training of the deep learning model.
[0060] In practice, the trained deep learning model should be further optimized or pruned to improve recognition speed.
[0061] In practice, commanders can make different eye movements in response to the locked target, thereby issuing different instructions in succession.
[0062] In addition, the ground information display module 20 can display all the instructions in the behavior-instruction mapping table on the screen so that commanders can make decisions more efficiently.
[0063] In actual design, step 6) includes: the air-to-ground communication module 50 extracts ground change information from the returned information so that the ground information display module 20 updates the map according to the ground change information. The returned information includes command execution results, battlefield situation changes and ground change information, and is not limited to these. Ground change information may include, for example, changes in the positions of various forces.
[0064] In this invention, the air-to-ground communication module 50 is responsible for transmitting the commander's instructions to relevant equipment of friendly forces on the ground via a relevant communication channel, enabling the forces to perform corresponding tasks such as adjusting their positions according to the instructions, thereby achieving effective coordination between ground and air deployment (such as helicopter formation layout). Furthermore, the air-to-ground communication module 50 transmits the feedback information to the ground information display module 20 of the reflex headset 10 in real time. Through the feedback information from the ground (such as data on changes in the positions of various forces), the commander can adjust decisions according to the battlefield situation at any time, and update map data based on the feedback information so that the commander can see the latest battlefield situation.
[0065] In the actual design, the ground information display module 20 includes a screen, the command and interaction module 40 includes a microprocessor, and the air-to-ground communication module 50 includes communication equipment. The eye-tracking infrared camera 30 is an existing camera in the field, used to capture infrared radiation emitted by the eye and convert it into an image. The eye-tracking infrared camera 30 includes an infrared transmitter and an infrared receiver.
[0066] To adapt to different combat scenarios and environments, this invention can be combined with other operating methods to further improve combat command efficiency.
[0067] 1) Eye movement combined with gesture operation: In addition to eye movement interaction, gesture recognition function can be combined to assist in the selection and confirmation of targets and instructions.
[0068] 2) Eye movement combined with voice: In some situations, commanders can use voice commands to assist in confirming certain operations, which is applicable to situations where visual means are not sufficient.
[0069] In practice, commanders can switch between multiple modes to improve the system's adaptability, depending on different mission requirements or environmental changes.
[0070] In addition, the ground information display module 20 in this invention can be integrated with AR or VR display modes to enhance intuitiveness.
[0071] The advantages of this invention are:
[0072] This invention significantly improves the command efficiency and execution accuracy of air commanders in air-to-ground coordinated operations, making command more intuitive, convenient, and efficient. It is applicable to command from within helicopters and other aerial flight equipment, and has the following beneficial effects:
[0073] 1) Easy to operate. This invention utilizes eye-tracking technology, enabling commanders to select and confirm targets and instructions through eye movements (such as rapid double blinks) without using their hands or other complex operating devices. This significantly simplifies the command process and ensures accurate and timely system responses.
[0074] 2) High real-time performance and accuracy. The head-mounted device can display ground information in real time, quickly track and identify targets that the commander is looking at and lock onto them, and quickly transmit instructions to targets and receive feedback information to achieve real-time updates of ground information. This significantly improves the reaction speed of air-ground coordinated operations, enhances the accuracy of command execution, and ensures the command efficiency of air-ground coordination.
[0075] 3) Visualized and comprehensive information. The map interface presents the ground information of the area to be commanded in a visual way, which helps commanders quickly identify enemy and friendly forces and have a comprehensive understanding of the ground battlefield situation, thereby improving the accuracy of decision-making.
[0076] 4) Freeing up the commander's hands. During combat, freeing up the commander's hands allows them to perform other high-priority operations and improves their ability to respond to emergencies.
[0077] The above description describes the preferred embodiments of the present invention and the technical principles applied thereto. For those skilled in the art, any obvious changes such as equivalent transformations or simple substitutions based on the technical solutions of the present invention, without departing from the spirit and scope of the present invention, shall fall within the protection scope of the present invention.
Claims
1. An eye-tracking interactive air-to-ground collaborative command system, characterized in that, The device includes a head-mounted display unit, which is equipped with a ground information display module, an eye-tracking infrared camera, a command and interaction module, and an air-to-ground communication module, wherein: The ground information display module is used to receive ground data transmitted by the data acquisition device, obtain ground information based on the ground data, and display the ground information in the form of a map. The eye-tracking infrared camera is used to acquire images of the commander's eyes; The command interaction module is used to lock onto targets by tracking the commander's gaze point based on the eye images and to obtain instructions by recognizing eye movements based on the eye images. The air-to-ground communication module is used to send the instructions to the locked target to complete the command operation, and to receive feedback information from the target so that the ground information display module can update and display the ground information based on the feedback information.
2. The eye-tracking interactive air-to-ground collaborative command system as described in claim 1, characterized in that, The head-mounted device includes a helmet worn on the commander's head, and the ground information display module is installed inside the helmet. The ground information display module includes a screen that displays the ground information in the form of a map. When the head-mounted device is worn on the commander's head, the screen is in front of his eyes. The eye-tracking infrared camera installed inside the helmet is positioned opposite to the eyes. The screen is semi-transparent so that the commander can see the scene in front of him through the screen.
3. The eye-tracking interactive air-to-ground collaborative command system as described in claim 1, characterized in that, The data acquisition equipment includes any one or a combination of reconnaissance satellites, drones, or ground force equipment. The ground data includes terrain information, the location and distribution of friendly forces, the location and distribution of enemy forces, and the weapon and ammunition status of each side. The ground information will associate and process the collected ground data and dynamically present it on a unified map interface in the commander's field of vision.
4. The eye-tracking interactive air-to-ground collaborative command system as described in claim 1, characterized in that, The ground information display module displays the ground information in map form as the ground information of the area to be commanded below the commander's air position, wherein the area to be commanded is a rectangular area with a set area or a circular area with a set radius.
5. An eye-tracking interactive air-to-ground collaborative command method, implemented based on the eye-tracking interactive air-to-ground collaborative command system according to any one of claims 1 to 4, characterized in that, Including the following steps: 1) The commander puts on and activates the aforementioned head-mounted device, entering combat readiness; 2) The ground information display module receives ground data transmitted by each of the data acquisition devices and displays the ground information obtained based on the ground data in a map form. 3) The eye-tracking infrared camera acquires images of the commander's eyes and transmits these images to the command interaction module; 4) The command interaction module determines whether the commander is looking at the eye image: if the commander is not looking at the eye, proceed to step 5); if the commander is looking at the eye, calculate the coordinates of the gaze point and determine the object of gaze. If the object of gaze is friendly troops, lock the object of gaze as the target and proceed to the next step; otherwise, return to step 3. 5) The command interaction module determines whether the commander has issued an instruction based on the eye image: if it determines that the commander has not issued an instruction, it returns to step 3); if it determines that the commander has issued an instruction, it identifies the commander's eye movement behavior, obtains the instruction corresponding to this eye movement behavior, and proceeds to the next step. 6) The air-to-ground communication module sends the command to the locked target and determines whether the target returns feedback: if there is feedback from the ground, proceed to the next step; otherwise, return to step 3). 7) The air-to-ground communication module sends the returned information to the ground information display module to update and display the map based on the returned information, and then returns to step 3).
6. The eye-tracking interactive air-to-ground collaborative command method as described in claim 5, characterized in that, In step 4), determining whether the commander is in a gaze state includes the following steps: The command interaction module obtains a monocular image of the left or right eye from the eye image, detects the pupil center position from the monocular image, and if the pupil center position does not change for a set time, the commander is considered to be in a gaze state, and then the gaze point coordinates are calculated; otherwise, the commander is considered not to be in a gaze state. The calculation of the gaze point coordinates includes the following steps: The command and interaction module obtains left and right eye images from the eye images. Through pupil detection, it extracts the coordinates of the left pupil center from the left eye image and the coordinates of the right pupil center from the right eye image. It then obtains the gaze point coordinates based on the polynomial fitting curve between the centers of the left and right pupils and the screen position. The polynomial fitting curve between the centers of the left and right pupils and the screen position is obtained through the following steps: The tester wore the head-mounted device and looked at multiple locations on the ground information display module in succession, obtaining multiple pairs of left and right eye images; For each pair of left and right eye images, the coordinates of the left eye pupil center (x1, y1) and the right eye pupil center (x2, y2) are extracted through pupil detection. The mapping relationship between the left and right eye pupil center coordinates and the screen position coordinates (x, y) is established according to the following formula: x = a1x1 + a2y1 + a3x2 + a4y2 + a5, y = b1x1 + b2y1 + b3x2 + b4y2 + b5, In the formula, a1-a5 and b1-b5 are coefficients; Based on the mapping relationship between the center coordinates of the left and right pupils and the screen position coordinates established by each pair of left and right eye images, a polynomial fitting curve is obtained. The coefficients of the polynomial fitting curve are determined by the least squares method, thereby establishing the polynomial fitting curve between the center of the left and right pupils and the screen position.
7. The eye-tracking interactive air-to-ground collaborative command method as described in claim 5, characterized in that, In step 5), eye movement recognition includes the following steps: the command interaction module obtains a monocular image sequence from the eye images, inputs the monocular image sequence into a deep learning model, and then obtains the eye movement result. The deep learning model is obtained through the following steps: Constructing a deep learning model: The deep learning model includes a three-dimensional convolutional neural network and a classification head. The three-dimensional convolutional neural network is used to extract temporal and spatial features based on the input monocular image sequence. The temporal and spatial features correspond to the temporal and spatial changes that occur during eye movement, respectively. The classification head is used to classify eye movement behavior based on the temporal and spatial features output by the three-dimensional convolutional neural network, establish the mapping relationship between spatiotemporal features and eye movement behavior, and output eye movement behavior results. Training the deep learning model: The deep learning model is trained by inputting multiple monocular image sequences.
8. The eye-tracking interactive air-to-ground collaborative command method as described in claim 5, characterized in that, Step 7) includes: the air-to-ground communication module extracts ground change information from the returned information so that the ground information display module updates the map according to the ground change information, wherein the returned information includes command execution results, battlefield situation changes and ground change information.