Efficiency evaluation method based on eye movement dominated multi-channel interaction data alignment
By synchronizing multi-channel interaction data with eye movement data, a high-precision performance evaluation model is constructed, which solves the problem of insufficient synchronization of multimodal data in traditional methods, achieves high-precision evaluation and interface optimization, and improves the efficiency of operators.
Patent Information
- Application Number
- CN202510701097.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional human-computer interaction effectiveness evaluation methods rely on a single data source and lack a spatiotemporal synchronization mechanism for multimodal data, resulting in insufficient comprehensiveness and accuracy of evaluation results.
Using eye movement data as the time reference, touch, voice, and vibration data are collected synchronously to build a performance evaluation model for multi-channel interaction data alignment. A timestamp compensation algorithm is used to achieve spatiotemporal synchronization of multimodal data. Performance indicators are calculated based on task phases to generate a comprehensive score and dynamically adjust the interface layout or training difficulty.
It improves the accuracy of multi-channel data integration and the precision of evaluation, can optimize the human-computer interaction interface in real time, and improve the operator's decision-making efficiency and execution accuracy, especially in military simulation and UAV command and control scenarios. It has important application value.
Smart Images

Figure CN120670796A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of human-computer interaction technology, specifically an effectiveness evaluation method based on eye-movement-led multi-channel interaction data alignment, which is suitable for high-precision human-computer collaboration scenarios such as military simulation training and unmanned aerial vehicle command and control systems. Background Art
[0002] Multi-channel human-computer interaction technology integrates eye movement, touch, voice and other interaction methods, and is widely used in high-precision operation scenarios such as military simulation and drone command and control. It aims to improve the decision-making efficiency and execution accuracy of operators in complex tasks.
[0003] However, traditional human-computer interaction effectiveness evaluation methods often rely on a single data source, such as only considering eye movement data or touch data, which limits the comprehensiveness and accuracy of the evaluation; in addition, existing methods often lack effective spatiotemporal synchronization mechanisms when processing multimodal data, resulting in errors in data integration and affecting the reliability of the evaluation results.
[0004] Therefore, a performance evaluation method is needed that can comprehensively consider multi-channel interaction data and achieve high-precision spatiotemporal synchronization. Summary of the Invention
[0005] The purpose of the present invention is: in order to solve the above problems, the present invention proposes an effectiveness evaluation method based on the alignment of eye-movement-dominated multi-channel interaction data. This method uses eye movement data as a time reference to perform spatiotemporal synchronization of multi-channel data such as touch, voice, and vibration, and constructs a dynamic effectiveness evaluation model based on the task phase division (reconnaissance, locking, and attack). By quantifying indicators such as target discovery time, fractal dimension of touch trajectory, and consistency of voice commands, a comprehensive effectiveness score is generated. When the score is lower than the threshold, the system can dynamically adjust the interface layout or training difficulty to improve the effectiveness of human-computer interaction.
[0006] In a first aspect, the present invention provides a performance evaluation method based on eye-movement-dominated multi-channel interaction data alignment, comprising:
[0007] S1. Synchronously collect interaction data from the eye tracker, touch screen, microphone, and vibration sensor, including eye gaze coordinates, touch trajectory, voice command timestamp, and vibration feedback intensity;
[0008] S2. Using the eye movement data sampling rate as the reference clock, resample and align the touch, voice, and vibration data to generate a spatiotemporally synchronized multimodal dataset with a synchronization error of ≤2ms.
[0009] S3, when it is detected that the eye movement fixation point stays on the interface control for more than 300ms, extract the touch, voice and vibration data segments associated with this period;
[0010] S4. Based on the mission phase, which includes reconnaissance, locking, and strike phases, calculate the effectiveness indicators of each phase, including target detection time, touch trajectory fractal dimension, voice command consistency, etc.
[0011] S5. Generate a comprehensive performance score based on the dynamic weight distribution model. When the score is lower than the threshold, output an interface layout optimization plan or a training difficulty adjustment instruction.
[0012] Preferably, the method for implementing the spatiotemporal synchronization in step S2 is:
[0013] Eye tracking data is based on a 120Hz sampling rate, and touch data is resampled to 120Hz through linear interpolation;
[0014] Voice commands are aligned to the eye movement timeline based on the start timestamp, and vibration feedback intensity is matched to the event triggering moment;
[0015] A timestamp compensation algorithm is used to perform linear interpolation resampling on the touch data, and combined with voice command start timestamp alignment and vibration event trigger matching to achieve spatiotemporal synchronization of multimodal data. The formula of the timestamp compensation algorithm includes:
[0016]
[0017] Where t′≤ET i ≤t″, t′ and t″ are the adjacent timestamps of touch data, Touch t″ and Touch t′ adjacent touch data;
[0018] Synchronization error verification:
[0019] Calculate the maximum time deviation (Max TimeSkew, MTS)
[0020]
[0021] If MTS>2ms, the hardware clock calibration protocol is triggered.
[0022] Preferably, in step S4, in the reconnaissance phase: target discovery time, false alarm rate; in the locking phase: number of eye movement gaze jumps, fractal dimension of touch trajectory; in the attack phase: consistency of voice commands, false vibration touch rate.
[0023] Preferably, the method for calculating the fractal dimension of the touch track in step 4 is:
[0024] The fractal dimension FD of the touch path is calculated using the box counting method and fitted using the least squares method. The formulas involved include:
[0025]
[0026] Among them, ∈ i The side length of the box in the i-th experiment, n is the total number of box size sequences;
[0027] When FD>1.2, it is judged as operation hesitation and trajectory optimization suggestions are generated.
[0028] In a second aspect, the present invention provides a multi-channel human-computer interaction performance evaluation module based on eye movement dominance, which applies the above-mentioned performance evaluation method based on eye movement dominance multi-channel interaction data alignment, including:
[0029] Data acquisition module: used to synchronously acquire interaction data from the eye tracker, touch screen, microphone, and vibration sensor;
[0030] Spatiotemporal alignment module: uses eye movement data as the time reference to resample and timestamp multi-channel data;
[0031] Dynamic Assessment Engine: Calculates performance indicators by military mission phase and generates real-time optimization instructions;
[0032] Feedback execution unit: adjusts the interface control layout or training scenario parameters based on the evaluation results.
[0033] In a third aspect, an embodiment of the present invention further provides a computer storage medium storing a computer program, which implements the above method when executed by a processor.
[0034] In a fourth aspect, an embodiment of the present invention further provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.
[0035] The present invention is based on the effectiveness evaluation module of the multi-channel interactive system, which includes a multi-channel human-computer interaction subsystem and a task scenario simulation subsystem; Figure 1As shown, the performance evaluation method based on eye movement-dominated multi-channel interaction data alignment is included in the interaction performance evaluation module. The evaluation method is based on the eye movement, touch, voice and vibration data collected by the multi-channel human-computer interaction subsystem. The eye movement sampling rate is used as a unified time reference, and the touch data is linearly interpolated and resampled by the timestamp compensation algorithm. The touch data is combined with the voice command start timestamp alignment and vibration event trigger matching to achieve multimodal data spatiotemporal synchronization; based on the reconnaissance, locking and strike phase division of the mission scenario simulation subsystem, a dynamic evaluation model is constructed: the target discovery time and false alarm rate are quantified in the reconnaissance phase, the fractal dimension of the touch trajectory is calculated by the box counting method in the locking phase, and the voice command consistency and vibration false touch rate are verified in the strike phase; further, through the dynamic weight distribution model or training difficulty, an "evaluation-optimization" closed loop is formed, which ultimately improves the human-computer collaboration efficiency and operation accuracy in the UAV command and control mission.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] The present invention improves the accuracy of multi-channel data integration and provides a more refined performance evaluation for different mission stages through a dynamic weight allocation model. In addition, the system can also provide optimization instructions in real time based on the evaluation results, further improving the operator's decision-making efficiency and execution accuracy. It has important application value in high-precision human-machine collaboration scenarios such as military simulation training and drone command and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a structural diagram of the multi-channel human-computer interaction system of the present invention;
[0039] Figure 2 This is a flow chart of the performance evaluation method based on eye-movement-led multi-channel interaction data alignment of the present invention. DETAILED DESCRIPTION
[0040] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0041] Embodiment: The present invention provides an effectiveness evaluation method based on eye movement-dominated multi-channel interaction data alignment, such as Figure 2 As shown, including:
[0042] Step 1: Multi-channel data acquisition, as follows:
[0043] 1.1 Gaze coordinates were collected using an eye tracker (120 Hz) with an accuracy of ±0.5°.
[0044] 1.2 Touch trajectory data is recorded via a capacitive touch screen (100 Hz) with a coordinate resolution of 1920 × 1080.
[0045] 1.3 Voice commands are captured by a microphone (16kHz sampling) and the vibration feedback intensity is measured by a tactile glove (200Hz).
[0046] Step 2: Time and space synchronization, as follows:
[0047] 2.1 Eye tracking data benchmark alignment:
[0048] Timestamp compensation algorithm:
[0049] Assume that the time series of eye movement data is The touch data time series is Resample the touch data to 120Hz using linear interpolation:
[0050]
[0051] Where t′≤ET i ≤t″, t′ and t″ are the adjacent timestamps of touch data, Touch t″ and Touch t′ It is the adjacent touch data.
[0052] Synchronization error verification:
[0053] Calculate the maximum time deviation (Max TimeSkew, MTS)
[0054]
[0055] If MTS>2ms, the hardware clock calibration protocol is triggered.
[0056] 2.2 Synchronous error control:
[0057] Calculate the mean time deviation of eye movement and touch data:
[0058]
[0059] If Δt>2ms, the hardware clock calibration procedure is triggered.
[0060] Voice command alignment algorithm
[0061] Step 3: Dynamic data slicing
[0062] 3.1 When it is detected that the gaze point stays on a certain control for ≥300ms, it is marked as a valid gaze event.
[0063] 3.2 Extract multi-channel data within the 50ms time window before and after the event and generate data fragments:
[0064] D slice ={ET t-50ms:t+50ms,Touch t-50ms:t+50ms ,Voice,Vib}
[0065] Step 4: Stage performance calculation
[0066] 4.1 Reconnaissance Phase Indicators:
[0067] Target detection time: the duration from the start of the task to the locking of the first target gaze point (≤3 seconds to meet the requirement).
[0068] T discover =min(t gaxe_end )-t task_start
[0069] False alarm rate: the proportion of times of incorrectly looking at non-target areas (≤5% to meet the standard).
[0070]
[0071] Among them, is T g The total number of fixations is P g Number of incorrect fixations.
[0072] The goal of this phase is to assess the operator's ability to quickly and accurately identify targets in complex environments. Target detection time provides insights into the operator's reaction speed and attention efficiency, while the false alarm rate reflects the operator's ability to accurately assess the environment and avoid operational errors.
[0073] 4.2 Locking Phase Indicators:
[0074] Number of eye movement jumps: the number of times the gaze point switches between the target and the weapon controls during the lock phase (≤2 times / target).
[0075] Fractal dimension of touch track:
[0076]
[0077] Among them, ∈ i The side length of the box in the i-th experiment, n is the total number of box size sequences;
[0078] When FD>1.2, it is judged as operation hesitation, and trajectory optimization suggestions are generated to generate trajectory optimization paths.
[0079] During the lock phase, the operator needs to stably track and lock onto the target. The number of eye gaze jumps reflects the operator's attention switching frequency and cognitive load, while the fractal dimension of the touch trajectory provides a deep understanding of the operator's hand movement patterns, helping to identify potential operational difficulties or optimization points.
[0080] 4.3 Indicators of the strike phase:
[0081] Voice command consistency: How well the voice command matches the preset command library.
[0082]
[0083] Among them, P v is the number of matching instructions, T v is the total number of instructions.
[0084] Vibration False Trigger Rate: The number of vibration feedbacks that falsely trigger weapon firing.
[0085] The strike phase is a critical step in mission execution, requiring the operator to accurately and promptly execute commands. Voice command consistency ensures effective communication between the operator and the system, while the vibration false touch rate reflects the accuracy and reliability of the system's tactile feedback. By evaluating these two metrics, we can further optimize the design of the human-machine interface and feedback mechanism, improving operator precision and efficiency.
[0086] Step 5: Dynamic Feedback Optimization
[0087] 5.1 Comprehensive scoring formula:
[0088] Score=0.4×S recon +0.5×S lock +0.1×S strike
[0089] In air combat mode, the weight is adjusted to: 0.3×S recon +0.6×S strike ;
[0090] 5.2 If the score of the performance evaluation method based on eye-movement-dominated multi-channel interaction data alignment is less than 80:
[0091] Interface optimization: The space in high-density gaze areas is enlarged by 20% and the spacing is increased by 15px.
[0092] Difficulty Adjustment: Increase the number of enemy interference targets.
[0093] As can be seen from the above, the present invention uses eye movement data as a time reference to synchronize the time and space of multi-channel data such as touch, voice, and vibration, and constructs a dynamic performance evaluation model based on task phase division. This method improves the accuracy of multi-channel data integration and, through a dynamic weight distribution model, provides a more refined performance evaluation for different task phases. In addition, the system can also provide real-time optimization instructions based on the evaluation results, further improving the operator's decision-making efficiency and execution accuracy. It has important application value in high-precision human-machine collaboration scenarios such as military simulation training and drone command and control.
[0094] Working Principle: A timestamp compensation algorithm is used to synchronize the temporal and spatial data of multi-channel interactions, such as touch, voice, and vibration, to generate a high-precision multimodal dataset. Then, based on the mission phases (reconnaissance, locking, and strike), a dynamic performance evaluation model is constructed to quantify evaluation indicators such as target detection time, fractal dimension of touch trajectory, and consistency of voice commands. When the overall performance score falls below the threshold, the system provides real-time optimization instructions for interface layout optimization or training difficulty adjustment, forming an "evaluation-optimization" closed loop to improve human-machine collaboration efficiency and operational precision.
[0095] The present invention provides an eye-movement-dominated multi-channel human-computer interaction performance evaluation module, which applies the above-mentioned eye-movement-dominated multi-channel interaction data alignment performance evaluation method, including:
[0096] Data acquisition module: used to synchronously acquire interaction data from the eye tracker, touch screen, microphone, and vibration sensor;
[0097] Spatiotemporal alignment module: uses eye movement data as the time reference to resample and timestamp multi-channel data;
[0098] Dynamic Assessment Engine: Calculates performance indicators by military mission phase and generates real-time optimization instructions;
[0099] Feedback execution unit: adjusts the interface control layout or training scenario parameters based on the evaluation results.
[0100] An embodiment of the present application provides an electronic device applicable to the above-mentioned performance evaluation method based on eye-movement-dominated multi-channel interaction data alignment, including:
[0101] Memory, used to protect computer programs and data;
[0102] Processor, used to run system programs.
[0103] An embodiment of the present application provides a computer storage medium, which is applicable to the above-mentioned performance evaluation method based on eye-movement-dominated multi-channel interaction data alignment, and performs hierarchical confidentiality management on the above-mentioned system and data in accordance with confidentiality management requirements.
[0104] Those skilled in the art will appreciate that the embodiments of the present application can be provided as a system or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0105] The present application is described with reference to the flowcharts and / or block diagrams of the devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0106] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0108] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0109] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0110] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0111] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, commodity, or apparatus comprising the element.
[0112] The embodiments of the present invention are provided for the purpose of illustration and description. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present invention. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. The performance evaluation method based on eye-movement-dominated multi-channel interaction data alignment is characterized by: include: S1. Synchronously collect interaction data from the eye tracker, touch screen, microphone, and vibration sensor, including eye gaze coordinates, touch trajectory, voice command timestamp, and vibration feedback intensity; S2. Using the eye movement data sampling rate as the reference clock, resample and align the touch, voice, and vibration data to generate a spatiotemporally synchronized multimodal dataset with a synchronization error of ≤2ms. S3, when it is detected that the eye movement fixation point stays on the interface control for more than 300ms, extract the touch, voice and vibration data segments associated with this period; S4. Based on the mission phase, which includes reconnaissance, locking, and strike phases, calculate the effectiveness indicators of each phase, including target detection time, touch trajectory fractal dimension, and voice command consistency; S5. Generate a comprehensive performance score based on the dynamic weight distribution model. When the score is lower than the threshold, output an interface layout optimization plan or a training difficulty adjustment instruction.
2. The performance evaluation method based on eye-movement-driven multi-channel interaction data alignment according to claim 1, characterized in that: The method for implementing the spatiotemporal synchronization in step S2 is: Eye tracking data is based on a 120Hz sampling rate, and touch data is resampled to 120Hz through linear interpolation; Voice commands are aligned to the eye movement timeline based on the start timestamp, and vibration feedback intensity is matched to the event triggering moment; A timestamp compensation algorithm is used to perform linear interpolation resampling on the touch data, and combined with voice command start timestamp alignment and vibration event trigger matching to achieve spatiotemporal synchronization of multimodal data. The formula of the timestamp compensation algorithm includes: Where t′≤ET i ≤t″, t′ and t″ are the adjacent timestamps of touch data, Touch t″ and Touch t′ adjacent touch data; Synchronization error verification: Calculate the maximum time deviation (Max TimeSkew, MTS) If MTS>2ms, the hardware clock calibration protocol is triggered.
3. The performance evaluation method based on eye-movement-driven multi-channel interaction data alignment according to claim 1, characterized in that: In step S4, during the reconnaissance phase, the following parameters are measured: target discovery time and false alarm rate; during the locking phase, the following parameters are measured: number of eye gaze jumps and fractal dimension of touch trajectory; and during the attack phase, the following parameters are measured: consistency of voice commands and false vibration touch rate.
4. The performance evaluation method based on eye-movement-driven multi-channel interaction data alignment according to claim 1, characterized in that: The calculation method of the touch track fractal dimension in step 4 is: The fractal dimension FD of the touch path is calculated using the box counting method and fitted using the least squares method; the formula involved is include: Among them, ∈ i The side length of the box in the i-th experiment, n is the total number of box size sequences; When FD>1.2, it is judged as operation hesitation and trajectory optimization suggestions are generated.
5. A multi-channel human-computer interaction effectiveness evaluation module based on eye movement, characterized by: The performance evaluation method based on eye-movement-led multi-channel interaction data alignment as described in any one of claims 1 to 4 comprises: Data acquisition module: used to synchronously acquire interaction data from the eye tracker, touch screen, microphone, and vibration sensor; Spatiotemporal alignment module: uses eye movement data as the time reference to resample and timestamp multi-channel data; Dynamic Assessment Engine: Calculates performance indicators by military mission phase and generates real-time optimization instructions; Feedback execution unit: adjusts the interface control layout or training scenario parameters based on the evaluation results.
6. A computer storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to claims 1 to 4 is implemented.
7. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to claims 1 to 4 is implemented.