A Safety Risk Early Warning System for Engineering Workers Based on Eye Tracking Monitoring

CN122537015APending Publication Date: 2026-08-11XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0002]在现代工业与工程领域,诸多高风险作业场景高度依赖作业人员的持续警觉性与动态视觉监控能力;例如,在化工厂中央控制室内,操作员需长时间面对由数十块仪表屏幕构成的监控界面,持续追踪多项工艺参数的细微变化;在电力系统的运维检修中,电工常需在结构复杂、元件密集的配电盘前,执行高重复性、高精度的线路连接与检测作业;此类场景具有视觉信息密度高、作业流程单调重复、且对注意力分配灵活性要求极严的共同特点;作业人员长期处于此种认知负荷下,极易因精神疲劳、任务压力或过度聚焦于局部细节,诱发一种内源性的认知功能退化现象,具体表现为“认知隧道效应”即注意力范围病理性窄化,或“注意盲视”即对视野内显著异常变化视而不见;在此状态下,人员的眼球仍可能进行规律性移动,甚至持续注视关键区域,但大脑的感知与信息处理功能已出现阻滞,对潜在风险丧失有效感知;这种“视而不见”的认知脱节,是导致误操作、漏检等重大安全事故的深层人为因素

Benefits of technology

[0015] The beneficial effects of this invention are as follows: This solution achieves a forward-looking early warning of intrinsic cognitive risks for engineering workers by performing multi-level and multi-scale in-depth analysis and fusion of eye movement signals; it constructs a personalized behavioral baseline by extracting microscopic eye movement dynamics features and macroscopic visual scanning strategy features, and uses a temporal reasoning model to comprehensively judge the co-evolution of multi-source evidence; ultimately, it can issue early warnings with high specificity and low false alarm rate before risks such as cognitive tunneling and attentional blindness actually lead to operational errors, thereby effectively improving the level of human safety in high-risk work scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122537015A_ABST
    Figure CN122537015A_ABST
Patent Text Reader

Abstract

This invention relates to a safety risk early warning system for engineering workers based on eye-tracking monitoring, specifically in the field of eye-tracking data analysis. This solution achieves proactive early warning of intrinsic cognitive risks for engineering workers through multi-level, multi-scale in-depth analysis and fusion of eye-tracking signals. It constructs a personalized behavioral baseline by extracting microscopic eye-tracking dynamics features and macroscopic visual scanning strategy features, and uses a temporal reasoning model to comprehensively assess the co-evolution of multi-source evidence. Ultimately, it can issue early warnings with high specificity and low false alarm rate before risks such as cognitive tunneling and attentional blindness actually lead to operational errors, thereby effectively improving human safety levels in high-risk work scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of eye-tracking data analysis, and more specifically, to a safety risk early warning system for engineering workers based on eye-tracking monitoring. Background Technology

[0002] In modern industry and engineering, many high-risk work scenarios heavily rely on operators' continuous vigilance and dynamic visual monitoring capabilities. For example, in the central control room of a chemical plant, operators must spend long periods of time facing a monitoring interface composed of dozens of instrument screens, continuously tracking subtle changes in multiple process parameters. In the operation and maintenance of power systems, electricians often need to perform highly repetitive and high-precision line connection and testing operations in front of complex and densely packed switchboards. These scenarios share common characteristics such as high visual information density, monotonous and repetitive work processes, and extremely stringent requirements for the flexibility of attention allocation. Under such cognitive load, mental fatigue, task pressure, or excessive focus on local details can easily induce an endogenous cognitive decline, specifically manifested as the "cognitive tunneling effect," which is a pathological narrowing of the attention span, or "attentional blindness," which is the inability to see significant abnormal changes within the visual field. In this state, the person's eyes may still move regularly, or even continuously focus on key areas, but the brain's perception and information processing functions have been blocked, resulting in a loss of effective perception of potential risks. This cognitive disconnect of "seeing but not seeing" is a deep-seated human factor leading to major safety accidents such as misoperation and missed detection.

[0003] Currently, safety monitoring technologies for workers, especially those utilizing eye tracking, primarily focus on describing overt visual behavior and providing simple threshold alarms. Existing solutions generally infer a person's fatigue level or whether their attention has deviated from a preset area by analyzing macroscopic eye-tracking indicators such as the spatial distribution of fixation points, saccade trajectory, average fixation duration, and pupil diameter changes. However, this approach, based on superficial statistical features, has inherent limitations. First, it struggles to effectively distinguish between "physiological fixation" and "cognitive perception." When a person is caught in a cognitive tunnel, their fixation point may remain on the relevant instruments, but the information processing... First, the cognitive process has already been interrupted, and existing technologies cannot identify this state of separation between cognition and vision. Second, the indicators relied upon by existing methods are not sensitive to the migration of higher-order cognitive states, such as the flexibility of attention resource allocation and the efficiency changes of visual information sampling strategies, and cannot capture the early subtle features of cognitive solidification. More importantly, there is a lack of a non-invasive predictive model that can dynamically model and warn of the decline in attentional flexibility and perceptual efficiency from continuous eye-tracking data streams. Therefore, there is an urgent need for an innovative technical solution that can go beyond traditional macroscopic eye-tracking analysis, deeply decode cognitive states from visual behavior patterns, and provide forward-looking warnings of cognitive risks. Summary of the Invention

[0004] This invention addresses the technical problems existing in the prior art by providing a safety risk early warning system for engineering workers based on eye-tracking monitoring. The system utilizes a feature extraction module, a strategy evaluation module, a baseline comparison module, and a decision-making early warning module to solve the problems mentioned in the background.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: Specifically, it includes a feature extraction module, a strategy evaluation module, a baseline comparison module, and a decision warning module that are sequentially connected in communication. Feature extraction module: preprocesses the acquired raw eye movement signal, performs micro-salivation dynamic analysis, gaze stability quantification and pupil oscillation spectrum analysis, extracts high-dimensional eye movement micro-behavioral feature vectors that characterize the micro-oscillation and dynamics of eye movement, and parses the real-time gaze coordinate sequence of the operator from the preprocessed raw eye movement signal; The strategy evaluation module receives high-dimensional eye-tracking micro-behavior feature vectors and real-time gaze coordinate sequences and obtains current task context information. It dynamically divides visual interest regions according to the semantics of the current task scenario, and calculates the weighted conditional entropy of visual scanning strategy randomness and task logic conformity based on the real-time gaze transfer sequence determined by the real-time gaze coordinate sequence. It also calculates the strategy deviation between the actual gaze transfer sequence and the optimal observer strategy generated offline through a pre-built reinforcement learning model, and generates a strategy effectiveness index. Baseline comparison module: Receives high-dimensional eye-tracking micro-behavioral feature vectors and strategy effectiveness indicators. When calling the pre-stored personalized behavior baseline of the monitored operator, it performs a multi-dimensional deviation measurement based on Mahalanobis distance between the feature vector (composed of high-dimensional eye-tracking micro-behavioral feature vectors, weighted conditional entropy, and strategy deviation degree) within the current time window and the personalized behavior baseline, and outputs a comprehensive deviation index. The decision-making early warning module receives the comprehensive deviation index, a subset of sensitive micro-features predefined and selected from the high-dimensional eye-tracking micro-behavioral feature vector, and the continuous policy deviation. It then uses the short-term trend of the comprehensive deviation index, the state of the sensitive micro-feature subset, and the state of the policy deviation as a multi-source evidence stream and inputs it into a pre-built temporal reasoning model for analysis. When the calculated cognitive risk probability exceeds a preset threshold and the multi-source evidence stream exhibits a collaborative degradation mode, an early warning is triggered.

[0006] In a preferred embodiment, the specific process of preprocessing the raw eye movement signal in the feature extraction module is as follows: First, the original eye movement signal containing the coordinates of the left pupil center, the coordinates of the right pupil center, the diameter signal of the left pupil, and the diameter signal of the right pupil is denoised, and the missing frames in the original eye movement signal are interpolated and spatially calibrated. Subsequently, gaze coordinate analysis and pupil signal analysis are performed simultaneously on the preprocessed raw eye movement signal; The specific process of gaze coordinate resolution is as follows: the calibrated coordinates of the left pupil center and the coordinates of the right pupil center are averaged to generate a real-time gaze coordinate sequence in time order. The specific process of pupil signal analysis is as follows: the validity of the left eye pupil diameter signal and the right eye pupil diameter signal are verified and fused and averaged to generate a single pupil diameter time series signal; The real-time gaze coordinate sequence and the pupil diameter time sequence signal together serve as the basic input data for subsequent microsaccade dynamic analysis, gaze stability quantification, and pupil oscillation spectrum analysis.

[0007] In a preferred embodiment, the specific process of performing micro-salivation dynamic analysis, fixation stability quantification, and pupillary oscillation spectrum analysis to construct a high-dimensional eye-movement micro-behavioral feature vector is as follows: Based on the real-time gaze coordinate sequence, a dual-threshold detection algorithm based on motion speed and acceleration is used to identify all micro-saccade events with an amplitude of less than one degree of visual angle. For a set of micro-saccade events identified within a set sliding time window, the micro-saccade direction consistency index is calculated. Based on the real-time gaze coordinate sequence, gaze point events with a duration exceeding a preset duration threshold are identified. During the duration of a single gaze point event, the micro-fluctuation sequence of the pupil center coordinates of both eyes relative to their average position is extracted, and the corrected sample entropy of the micro-fluctuation sequence is calculated. The calculation result is used as a quantitative indicator of gaze stability. Based on the pupil diameter time-series signal, its spectrum as a function of time is obtained through short-time Fourier transform. The power spectral density within a predetermined frequency range related to neural oscillation is extracted, and the normalized entropy value of the power spectral distribution within this frequency range is calculated as the pupil oscillation spectral entropy. Finally, the microsaccade direction consistency index, gaze stability quantification index, and pupil oscillation spectrum entropy are combined to form a high-dimensional eye movement microbehavioral feature vector.

[0008] In a preferred embodiment, the specific process of dynamically dividing visual interest regions and generating a gaze-dwelling region sequence based on the semantics of the current work scenario in the strategy evaluation module is as follows: First, the current task context information is loaded. The task context information defines the logical description and spatial layout of each functional unit in the scene in the form of structured data. Based on this, a set containing multiple visual interest regions is dynamically constructed. Subsequently, the received real-time gaze coordinate sequence is used in conjunction with the established set of visual interest regions to determine spatial affiliation. Each coordinate point in the real-time gaze coordinate sequence is assigned a visual interest region number. If the coordinate point is not in any preset visual interest region, it is marked as a background region. Next, random gazes are filtered out, and the continuous real-time gaze coordinate sequence is compressed into a discrete, time-ordered gaze dwell region sequence, where each element represents the visual interest region number where a valid gaze lingers. The gaze-dwelling region sequence serves as the basic input data for calculating the weighted conditional entropy and policy deviation.

[0009] In a preferred embodiment, the specific process of calculating the weighted conditional entropy and policy deviation to generate a policy performance index is as follows: First, based on the gaze dwelling region sequence, we calculate the empirical conditional probability of moving from each visual interest region to another, and the empirical probability of each visual interest region appearing in the gaze dwelling region sequence. A predefined task logic weight matrix is ​​introduced, in which each weight element reflects the necessity and rationality of the action of moving from one specified visual interest region to another in the current task context. Based on empirical conditional probability, empirical probability of occurrence of each visual interest region, and task logical weight matrix, the task-weighted visual scanning conditional entropy is obtained through the following calculation process; an optimal observer policy pre-trained in a simulated environment using a reinforcement learning model is used. This optimal observer policy aims to minimize risk and contains an optimal state-action value function. For the actual scanning trajectory represented by the gaze dwell area sequence, the policy deviation is obtained through the following calculation process based on the optimal state action value function; finally, the calculated task-weighted visual scanning conditional entropy and the policy deviation are combined to form a policy effectiveness index.

[0010] In a preferred embodiment, the specific process of constructing a feature vector that integrates micro-features and macro-strategies and calling the personalized behavioral baseline in the baseline comparison module is as follows: First, after receiving the high-dimensional eye-tracking microbehavioral feature vector and the strategy effectiveness index, ensure that the high-dimensional eye-tracking microbehavioral feature vector and the strategy effectiveness index correspond to the same preset time window for analysis, and complete the time alignment. Subsequently, the multiple micro-eye movement features contained in the high-dimensional eye movement micro-behavior feature vector are sequentially concatenated with the task-weighted visual scanning conditional entropy and policy deviation contained in the policy effectiveness index in the feature dimension to form a new, higher-dimensional fusion feature vector. This fusion feature vector is a feature vector that combines micro-features and macro-policies, which is composed of the high-dimensional eye movement micro-behavior feature vector, weighted conditional entropy and policy deviation. At the same time, based on the identity of the currently monitored worker, the personalized behavior baseline pre-established for the worker is retrieved and called from the pre-stored personalized behavior baseline. The personalized behavior baseline includes the baseline mean vector and baseline covariance matrix obtained by statistical analysis of a large amount of historical normal work data of the worker.

[0011] In a preferred embodiment, the specific process of performing a multidimensional deviation metric based on Mahalanobis distance to output a comprehensive deviation index is as follows: Based on the fusion feature vector constructed in the current time window, and the baseline mean vector and baseline covariance matrix of the corresponding personnel, the Mahalanobis distance is calculated as a comprehensive deviation index. The Mahalanobis distance calculation process is as follows: First, calculate the difference vector between the current fused feature vector and the baseline mean vector; Then, calculate the inverse of the baseline covariance matrix; transpose the difference vector to obtain the transposed difference vector; multiply the transposed difference vector with the inverse matrix, and then multiply the result with the original, untransposed difference vector to obtain a scalar value. Finally, the arithmetic square root of the scalar value is calculated, and the result is the Mahalanobis distance value, which is the output comprehensive deviation index. The overall deviation index is a non-negative scalar that quantifies the degree of deviation of the behavioral pattern represented by the current fused feature vector from the individual's historical normal baseline.

[0012] In a preferred embodiment, the process of generating a multi-source evidence stream in the decision warning module is as follows: First, the system receives the comprehensive deviation index, the sensitive micro-feature subset, and the persistent strategy deviation. The sensitive micro-feature subset is pre-selected and extracted from the high-dimensional eye-tracking micro-behavior feature vector, including the micro-salivation direction consistency index and the pupillary oscillation spectrum entropy. The comprehensive deviation index, the sensitive micro-feature subset, and the persistent strategy deviation are resampled at fixed time intervals to ensure that they have corresponding values ​​at the same time point in the same series, thus completing time synchronization. Next, for the comprehensive deviation index, the microsaccade direction consistency index and pupil oscillation spectrum entropy in the sensitive micro feature subset, and the continuous strategy deviation, the linear regression slope of each index within a recent sliding window of a preset time length is calculated. This linear regression slope is the short-term trend. Then, the current value and short-term trend of the comprehensive deviation index, the current value and short-term trend of the microsaccade direction consistency index, the current value and short-term trend of the pupil oscillation spectrum entropy, and the current value and short-term trend of the strategy deviation are standardized respectively to obtain the standardized values ​​of the current value and short-term trend of each index. Finally, the standardized current value of the comprehensive deviation index and its standardized short-term trend, the standardized current value of the microsaccade direction consistency index and its standardized short-term trend, the standardized current value of the pupil oscillation spectrum entropy and its standardized short-term trend, and the standardized current value of the strategy deviation and its standardized short-term trend are concatenated and combined in a predetermined order to form a feature vector containing multiple dimensions for representing multi-source evidence flow.

[0013] In a preferred embodiment, the specific process by which the temporal reasoning model analyzes and calculates the probability of cognitive risk is as follows: The pre-built temporal reasoning model contains a deep learning network with a bidirectional gated recurrent unit network layer and a multi-head self-attention mechanism layer; the temporal reasoning model receives a time series consisting of feature vectors representing multi-source evidence streams at a consecutive preset number of time points as input; First, the input time series is fed into a bidirectional gated recurrent unit network layer for processing. This bidirectional gated recurrent unit network layer can capture the dependency relationship between each time point in the time series and its context time points from both forward and backward directions, and output a hidden state vector sequence containing temporal context information. Subsequently, the hidden state vector sequence is input into a multi-head self-attention mechanism layer, which can simultaneously calculate the association weight between any two hidden states at any time point in the hidden state vector sequence and output a context vector sequence weighted by attention weights. An adaptive max pooling operation is performed on the context vector sequence to extract the most indicative feature information from the entire input time series, resulting in a pooled feature vector. Finally, the pooled feature vector is processed through a fully connected layer and combined with a non-linear activation function to map and output a real value between zero and one, which is the cognitive risk probability.

[0014] In a preferred embodiment, the specific process of identifying the collaborative degradation mode and triggering the early warning is as follows: First, based on the standardized short-term trends of each piece of evidence, a co-degradation index is calculated; this co-degradation index reflects the degree to which multiple pieces of evidence change synchronously in a predetermined deterioration direction. Triggering an early warning requires meeting two conditions simultaneously: the first condition is that the current cognitive risk probability calculated by the time-series reasoning model exceeds a preset probability threshold; the second condition is that the currently calculated collaborative degradation index exceeds a preset collaborative threshold. Only when the probability of perceived risk and the synergistic degradation index both exceed their respective preset probability thresholds and synergistic thresholds is the multi-source evidence stream determined to exhibit a synergistic degradation mode, and an early warning signal is triggered for the operators.

[0015] The beneficial effects of this invention are as follows: This solution achieves a forward-looking early warning of intrinsic cognitive risks for engineering workers by performing multi-level and multi-scale in-depth analysis and fusion of eye movement signals; it constructs a personalized behavioral baseline by extracting microscopic eye movement dynamics features and macroscopic visual scanning strategy features, and uses a temporal reasoning model to comprehensively judge the co-evolution of multi-source evidence; ultimately, it can issue early warnings with high specificity and low false alarm rate before risks such as cognitive tunneling and attentional blindness actually lead to operational errors, thereby effectively improving the level of human safety in high-risk work scenarios. Attached Figure Description

[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a block diagram of the system structure of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0019] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0020] Example 1 This embodiment provides, for example Figure 1-2 The system illustrates a safety risk early warning system for engineering workers based on eye-tracking monitoring, which specifically includes: a feature extraction module, a strategy evaluation module, a baseline comparison module, and a decision early warning module that are sequentially connected via communication. Feature extraction module: preprocesses the acquired raw eye movement signal, performs micro-salivation dynamic analysis, gaze stability quantification and pupil oscillation spectrum analysis, extracts high-dimensional eye movement micro-behavioral feature vectors that characterize the micro-oscillation and dynamics of eye movement, and parses the real-time gaze coordinate sequence of the operator from the preprocessed raw eye movement signal; The strategy evaluation module receives high-dimensional eye-tracking micro-behavior feature vectors and real-time gaze coordinate sequences and obtains current task context information. It dynamically divides visual interest regions according to the semantics of the current task scenario, and calculates the weighted conditional entropy of visual scanning strategy randomness and task logic conformity based on the real-time gaze transfer sequence determined by the real-time gaze coordinate sequence. It also calculates the strategy deviation between the actual gaze transfer sequence and the optimal observer strategy generated offline through a pre-built reinforcement learning model, and generates a strategy effectiveness index. Baseline comparison module: Receives high-dimensional eye-tracking micro-behavioral feature vectors and strategy effectiveness indicators. When calling the pre-stored personalized behavior baseline of the monitored operator, it performs a multi-dimensional deviation measurement based on Mahalanobis distance between the feature vector (composed of high-dimensional eye-tracking micro-behavioral feature vectors, weighted conditional entropy, and strategy deviation degree) within the current time window and the personalized behavior baseline, and outputs a comprehensive deviation index. The decision-making early warning module receives the comprehensive deviation index, a subset of sensitive micro-features predefined and selected from the high-dimensional eye-tracking micro-behavioral feature vector, and the continuous policy deviation. It then uses the short-term trend of the comprehensive deviation index, the state of the sensitive micro-feature subset, and the state of the policy deviation as a multi-source evidence stream and inputs it into a pre-built temporal reasoning model for analysis. When the calculated cognitive risk probability exceeds a preset threshold and the multi-source evidence stream exhibits a collaborative degradation mode, an early warning is triggered.

[0021] In this embodiment, the specific process of preprocessing the original eye movement signal in the feature extraction module is as follows: First, the acquired raw eye-tracking signal, containing the coordinates of the left and right pupil centers, the diameter of the left pupil, and the diameter of the right pupil, is subjected to denoising, interpolation compensation for lost frames, and spatial calibration to complete the preprocessing of the raw eye-tracking signal. The denoising operation includes filtering out high-frequency noise in the signal using methods based on moving average or Kalman filtering. The interpolation compensation operation fills in the signal loss caused by blinking or brief head occlusion using linear interpolation or spline interpolation methods. The spatial calibration operation uses pre-calibrated mapping parameters to transform the pupil center coordinates from the image coordinate system to a unified screen or world coordinate system. Subsequently, gaze coordinate analysis and pupil signal analysis are performed simultaneously on the preprocessed raw eye movement signal; The specific process of gaze coordinate resolution is as follows: the calibrated coordinates of the left and right pupil centers are averaged to generate a real-time gaze coordinate sequence in time order; the averaging process is as follows: at each sampling time, the abscissa of the left and right pupil centers is summed and divided by two to obtain the abscissa of the fused gaze; the ordinate of the left and right pupil centers is summed and divided by two to obtain the ordinate of the fused gaze; the fused gaze coordinates of all sampling times are arranged in time order to form a real-time gaze coordinate sequence. The specific process of pupil signal analysis is as follows: the validity of the left eye pupil diameter signal and the right eye pupil diameter signal are verified and fused and averaged to generate a single pupil diameter time series signal; the validity verification includes checking whether the pupil diameter is within the physiologically reasonable range and comparing the correlation of the changes in pupil diameter of both eyes, and removing sampling points with too low correlation or obvious abnormalities; the fusion and averaging process is as follows: the arithmetic mean of the left eye pupil diameter value and the right eye pupil diameter value at the same sampling time that have passed the verification is calculated, and it is used as the fused pupil diameter value at that time, and arranged in chronological order to form the pupil diameter time series signal; The real-time gaze coordinate sequence and the pupil diameter time sequence signal together serve as the basic input data for subsequent microsaccade dynamic analysis, gaze stability quantification and pupil oscillation spectrum analysis. The specific process of performing microsaccadic dynamic analysis, fixation stability quantification, and pupillary oscillation spectrum analysis to construct a high-dimensional eye movement microbehavioral feature vector is as follows: Based on real-time gaze coordinate sequences, a dual-threshold detection algorithm based on motion velocity and acceleration is used to identify all microsalivation events with amplitudes less than one degree of visual angle. For a set of microsalivation events identified within a set sliding time window, the ratio of the magnitude of the sum of the unit vectors of all microsalivation direction vectors in the set of microsalivation events to the total number of microsalivation events in the set is calculated to obtain the microsalivation direction consistency index, which is a component of the high-dimensional eye-tracking microbehavior feature vector. The specific process of the dual-threshold detection algorithm based on motion speed and acceleration is as follows: First, the real-time gaze coordinate sequence is differentially divided in the time dimension to calculate the instantaneous motion speed of the gaze; then, the instantaneous motion speed sequence is differentially divided to calculate the instantaneous motion acceleration; a speed threshold and an acceleration threshold are set. When the instantaneous speed and instantaneous acceleration of the gaze movement simultaneously exceed their respective thresholds, and the motion amplitude is less than one degree of visual angle, the motion segment is determined to be a micro-saccade event; the length range of the set sliding time window is two to five seconds. The process of calculating the micro-scanning direction consistency index is as follows: For all micro-scanning events identified within the sliding time window, the direction vector of each micro-scanning event is converted into a unit vector, that is, a direction vector with a length of 1 and an unchanged direction; then, all unit vectors are added together to obtain a composite vector; the length of the composite vector is calculated; finally, the length of the composite vector is divided by the total number of micro-scanning events within the window, and the resulting ratio is the micro-scanning direction consistency index. The closer the value is to 1, the more consistent the micro-scanning direction is within that time period. Based on the real-time gaze coordinate sequence, gaze point events with a duration exceeding a preset duration threshold are identified. During the duration of a single gaze point event, the micro-fluctuation sequence of the pupil center coordinates of both eyes relative to their average position is extracted, and the corrected sample entropy of the micro-fluctuation sequence is calculated. The calculation result is used as another component of the high-dimensional eye movement micro-behavior feature vector, namely the gaze stability quantification index. The preset duration threshold is between 80 and 150 milliseconds. The process of extracting the micro-fluctuation sequence is as follows: For a fixation event with a duration exceeding the threshold, extract the fused gaze coordinates of all sampling moments throughout its duration; calculate the average values ​​of these fused gaze coordinates in the horizontal and vertical directions, which are used as the average center position of the gaze during the fixation period; subtract the horizontal coordinate of the average center position from the horizontal coordinate of each sampling moment to obtain the horizontal micro-fluctuation component at that moment; similarly, subtract the vertical coordinate of the average center position from the vertical coordinate of each sampling moment to obtain the vertical micro-fluctuation component at that moment; the horizontal and vertical micro-fluctuation components together constitute a two-dimensional micro-fluctuation sequence. The process of calculating the corrected sample entropy of this micro-fluctuation sequence is as follows: First, a pattern length parameter and a similarity tolerance parameter are set. Then, in the two-dimensional micro-fluctuation sequence, all continuous subsequences with a length equal to the pattern length (called templates) are counted. For each template, the number of other templates whose distance from it is less than the similarity tolerance is counted in the entire sequence. Next, all continuous subsequences with a length equal to the pattern length plus one are counted. For each template whose length is equal to the pattern length plus one, the number of other templates whose distance from it is less than the similarity tolerance is counted. Finally, the negative natural logarithm of the ratio of the number of matches of the latter template to the number of matches of the former template is calculated as the corrected sample entropy value, which is used to quantify the complexity and irregularity of the micro-fluctuation sequence. Based on the pupil diameter time-series signal, its spectrum as a function of time is obtained by short-time Fourier transform. The power spectral density in a predetermined frequency range related to neural oscillation is extracted, and the normalized entropy value of the power spectral distribution in this frequency range is calculated to obtain the pupil oscillation spectrum entropy, which is another component of the high-dimensional eye movement microbehavior feature vector. The short-time Fourier transform process is as follows: the pupil diameter time-series signal is segmented using a sliding time window, and a Fourier transform is applied to each segment to obtain the spectrum of the signal frequency components as a function of time; the preset specified frequency range related to neural oscillation is 0.8 Hz to 1.2 Hz. The process of calculating the normalized entropy of the power spectrum distribution within this frequency range is as follows: First, within the frequency range of 0.8 Hz to 1.2 Hz, calculate the power spectral density value at each discrete frequency point; sum the power spectral density values ​​at all frequency points to obtain the total power of this frequency range; then divide the power spectral density value at each frequency point by the total power to obtain the normalized power value corresponding to that frequency point; finally, calculate the information entropy of all normalized power values ​​according to the information entropy formula. The result is the pupil oscillation spectral entropy, which is used to quantify the regularity of the pupil rhythmic oscillation pattern. Finally, the microsaccade direction consistency index, gaze stability quantification index, and pupil oscillation spectrum entropy are combined to form a high-dimensional eye movement microbehavioral feature vector.

[0022] In this embodiment, it is specifically necessary to explain the process in the strategy evaluation module of dynamically dividing the visual interest region based on the semantics of the current work scenario and generating a sequence of gaze-dwelling regions as follows: First, the task context information corresponding to the current work scenario is loaded. The task context information defines the logical description and spatial layout of each functional unit in the scenario in the form of structured data. Based on this, a set of multiple visual interest regions is dynamically constructed, with each visual interest region corresponding to an independent functional unit. The task context information may include the scenario layout file, equipment list, or preset configuration file, which clearly records the identifiers of key functional units such as dashboards, switches, indicator lights, and designated work areas, their boundary coordinates on the screen or in actual space (such as the coordinates of the upper left and lower right corners of a rectangular area), and their functional descriptions. The process of dynamically constructing the set of visual interest regions is to instantiate a visual interest region object with a unique number and spatial range for each functional unit in the system based on these boundary coordinates. Subsequently, using the received real-time gaze coordinate sequence and the established set of visual interest regions, spatial attribution is determined. Each coordinate point in the real-time gaze coordinate sequence at each sampling time is assigned a visual interest region number. If a coordinate point is not within any preset visual interest region, it is marked as a background region. The specific process of spatial attribution determination is as follows: For each coordinate point in the real-time gaze coordinate sequence, it is checked whether it falls within the spatial range defined by each visual interest region object. If it falls within a visual interest region, the coordinate point is assigned the region's number. If it does not fall within any visual interest region after traversing all visual interest regions, the coordinate point is marked as belonging to the background region. Next, incidental gazes lasting less than 50 milliseconds are filtered out, compressing the continuous real-time gaze coordinate sequence into a discrete, chronologically ordered sequence of gaze dwell areas. Each element in this sequence represents the visual interest region number or background marker where a valid gaze lingers. The process of filtering out incidental gazes and compressing the sequence is as follows: First, in the continuous coordinate point sequence marked with region numbers, segments with the same region number (including background markers) are identified. Then, the duration of each segment is calculated (i.e., the number of coordinate points within the segment multiplied by the sampling interval). Next, segments with a duration of less than 50 milliseconds are discarded, considered as incidental gaze jumps or noise, and not counted as valid gazes. Finally, the remaining segments with a duration of 50 milliseconds or more are represented by their uniform region number (or background marker) and arranged in chronological order to form the gaze dwell area sequence. The 50-millisecond threshold is set based on the physiological basis of the shortest time required for effective attention in human visual perception. The gaze-dwell region sequence serves as the basic input data for calculating the weighted conditional entropy and policy deviation. The specific process of calculating the weighted conditional entropy and policy deviation to generate a policy performance index is as follows: First, based on the gaze-retention region sequence, the empirical conditional probability of moving from each visual interest region to another, and the empirical probability of each visual interest region appearing in the gaze-retention region sequence are calculated. The process of calculating the empirical probability is as follows: traverse the gaze-retention region sequence, count the total number of times each visual interest region number (excluding background markers) appears in the entire sequence, divide the number of occurrences of each region by the total number of occurrences of all visual interest regions (excluding background markers) in the sequence to obtain the empirical probability of occurrence of each visual interest region; at the same time, count the number of occurrences of all transition pairs (from the previous region to the next region) formed by all adjacent elements in the sequence, ignoring transition pairs involving background markers; for each specific transition pair from the previous region i to the next region j, divide its occurrence count by the total number of transition pairs originating from the previous region i to obtain the empirical conditional probability of moving from region i to region j. A task logic weight matrix, predefined based on task flowcharts, operation manuals, or expert knowledge, is introduced. Each weight element in this matrix reflects the necessity and rationality of the action of transferring from one specified visual interest region to another within the current task context. The task logic weight matrix is ​​a square matrix with the number of rows and columns equal to the total number of visual interest regions. The assignment rules for weight elements are typically as follows: for two regions that are explicitly required to be checked continuously or have a strong logical sequence in the standard operating procedure or operation procedure, their corresponding transfer weights are set to higher values, such as close to or equal to one; for two logically unrelated regions or regions that should not be viewed continuously, their transfer weights are set to lower values, such as zero or close to zero; for transfers that are not explicitly specified, they can be set to intermediate values, such as 0.5. The task logic weight matrix is ​​usually configured according to the specific work scenario during system deployment. Based on empirical conditional probabilities, the empirical probabilities of occurrence of each visual interest region, and the task logic weight matrix, the task-weighted visual scanning conditional entropy is obtained through the following calculation process: multiply the empirical probability of occurrence of each visual interest region by the conditional empirical probability of transitioning from that region to another region, then multiply by the element of the task logic weight matrix corresponding to the transition, and finally multiply by the base-2 logarithm of the conditional empirical probability. Perform the above calculation on all possible combinations of visual interest regions and sum them up. The negative value of the summation result is taken as the task-weighted visual scanning conditional entropy. This task-weighted visual scanning conditional entropy integrates the statistical uncertainty of gaze transitions with the prior knowledge of task logic, and is used to quantify the randomness of the scanning strategy and its conformity with the task logic. "Multiply by the base-2 logarithm of the conditional empirical probability" refers to calculating the binary logarithm of the conditional empirical probability; "Perform the above calculation and sum for all possible combinations of visual interest regions" means traversing all visual interest regions as the previous region i and all visual interest regions as the next region j (including the case where i equals j, i.e., looking at the same region), and accumulating the products calculated for each pair (i,j); the lower the entropy value obtained, the more likely it is that the gaze shifting pattern is either too rigid and predictable or completely random and does not conform to the task logic; if the entropy value is maintained at a moderately high level, it indicates that the scanning strategy has both flexibility and task orientation; Simultaneously, an optimal observer policy pre-trained in a simulated environment using a reinforcement learning model is employed. This optimal observer policy aims to maximize information acquisition or minimize risk, and internally contains an optimal state-action value function. The reinforcement learning model is typically trained using Q-learning or deep Q-network algorithms. The simulated environment constructs a virtual task consistent with real-world work scenarios and defines the state (current gaze region), action (next gaze region), reward (positive reward for correctly identifying key information, negative reward for missing risk), and termination condition. Through extensive simulation training, the model learns an optimal state-action value function. This function outputs a real value for a given state (region) and action (next region), representing the cumulative long-term expected reward of taking this action in that state and subsequently following the optimal observer policy. For the actual scanning trajectory represented by the gaze dwell region sequence, the policy deviation is obtained through the following calculation process based on the optimal state action value function: For each position in the gaze dwell region sequence except the last position, the current gaze region number and the next actual gaze region number are obtained; the value score corresponding to the transfer action represented by the next actual gaze region number in the optimal state action value function is calculated, based on the state represented by the current gaze region number, and the value score is converted using an exponential function and a preset temperature parameter; the preset temperature parameter is a positive real number used to adjust the "strictness" of the policy evaluation, and its typical value range is between 0.1 and 1; the specific calculation for the conversion using the exponential function and the temperature parameter is as follows: first, the quotient obtained by dividing the value score by the temperature parameter is calculated, and then the natural constant e is raised to the power of the quotient; this step converts the value score into a non-negative "propensity score"; Simultaneously, the sum of the value scores of all possible transfer actions represented by other visual interest region numbers, under the state represented by the current gaze region number, is calculated after transformation by the same exponential function and temperature parameter; the transformed score corresponding to the actual transfer action is divided by this sum to obtain a ratio; the average of all such ratios in the entire gaze dwell region sequence is calculated; finally, the average is subtracted from the average to obtain the strategy deviation; the strategy deviation measures the difference between the actual scan trajectory and the optimal observer strategy. "All possible transfer actions represented by other visual interest region numbers" includes actions represented by all visual interest region numbers (usually excluding background markers); the calculation of the ratio essentially normalizes the "preferential score" of the actual action to the sum of the "preferential scores" of all possible actions, reflecting the relative superiority or inferiority of the actual action among all possible actions; the average of all such ratios in the entire gaze-holding region sequence is calculated to obtain the degree of conformity between the actual trajectory and the optimal observer strategy; subtracting this average from one makes the final strategy deviation range between zero and one, with a larger value indicating a more severe deviation; Finally, the calculated task-weighted visual scanning conditional entropy and policy deviation are combined to form a policy effectiveness index that characterizes the effectiveness of visual information acquisition strategies. The combination usually uses the task-weighted visual scanning conditional entropy and policy deviation as two independent numerical components to form a two-dimensional policy effectiveness index vector.

[0023] In this embodiment, it is specifically necessary to explain the process of constructing a feature vector that integrates micro-features and macro-strategies and calling the personalized behavior baseline in the baseline comparison module as follows: First, after receiving the high-dimensional eye-tracking micro-behavior feature vector from the feature extraction module and the policy effectiveness index from the policy evaluation module, ensure that the high-dimensional eye-tracking micro-behavior feature vector and the policy effectiveness index correspond to the same preset time window for analysis, and complete the time alignment; the typical value of the preset time window for analysis is two to five seconds, which is long enough to capture short-term patterns of behavior without introducing excessive delay. Subsequently, multiple micro-eye movement features contained in the high-dimensional eye-movement micro-behavior feature vector are sequentially concatenated with the task-weighted visual scan conditional entropy and policy deviation contained in the policy effectiveness index, forming a new, higher-dimensional fusion feature vector. This fusion feature vector is a feature vector that integrates micro-features and macro-strategies, composed of the high-dimensional eye-movement micro-behavior feature vector, weighted conditional entropy, and policy deviation. Sequential concatenation means arranging all elements in the high-dimensional eye-movement micro-behavior feature vector (such as the micro-sagging direction consistency index, gaze stability quantification index, pupillary oscillation spectrum entropy, etc.) in a fixed order, followed by the task-weighted visual scan conditional entropy, and finally the policy deviation, thus forming a longer one-dimensional array, i.e., the fusion feature vector. This operation mathematically realizes the physical integration of micro-physiological response features and macro-behavioral strategy features. Simultaneously, based on the identity of the currently monitored worker, the system retrieves and calls the personalized behavioral baseline pre-established for that worker from the pre-stored personalized behavioral baseline. The personalized behavioral baseline includes a baseline mean vector and a baseline covariance matrix obtained from a large amount of historical normal work data of that worker. The baseline mean vector describes the typical average value of each feature in the fused feature vector of that worker under healthy cognitive state, while the baseline covariance matrix describes the interrelationship and dispersion of these features within their normal fluctuation range. "A large amount of historical normal work data" refers to eye-tracking data continuously collected and stored during the initial calibration phase of the system or during a long period of work when the worker is in a known, good condition and has no accident record. The calculation process of the baseline mean vector is as follows: For each analysis time window in the historical data, a fusion feature vector is constructed according to the above method. Then, the arithmetic mean of all historical fusion feature vectors on each feature dimension is calculated. The vector formed by arranging these averages in the same order is the baseline mean vector. The calculation process of the baseline covariance matrix is ​​as follows: Based on the same batch of historical fused feature vector samples, first calculate the variance of each feature dimension itself (i.e., the average of the squares of the differences between the data of that dimension and its mean), then calculate the covariance between any two different feature dimensions (i.e., the average of the product of the differences between the data of the two dimensions and their means), and arrange all the variances and covariances into a square matrix according to the feature order. This square matrix is ​​the baseline covariance matrix, which quantitatively describes the amplitude of the normal fluctuation of each feature and the correlation trend between the changes of each pair. The specific process for performing a multidimensional deviation metric based on the Mahalanobis distance to output a comprehensive deviation index is as follows: Based on the fusion feature vector constructed in the current time window, and the baseline mean vector and baseline covariance matrix of the corresponding personnel, the Mahalanobis distance is calculated as a comprehensive deviation index. The Mahalanobis distance calculation process is as follows: First, calculate the difference vector between the current fused feature vector and the baseline mean vector; the difference vector is the new vector formed by subtracting the mean of the corresponding feature in the baseline mean vector from the value of each feature in the current fused feature vector. Next, the inverse matrix of the baseline covariance matrix is ​​calculated. Calculating the inverse matrix of the baseline covariance matrix involves using a standard matrix inversion algorithm (such as Gaussian elimination or LU decomposition) to obtain a new matrix with the same number of rows and columns as the original covariance matrix; this is called the inverse matrix. The difference vector is then mathematically transposed to obtain the transposed difference vector. The transposed difference vector is multiplied by the inverse matrix, and the result is multiplied by the original, untransposed difference vector to obtain a scalar value. This involves three consecutive matrix or vector multiplication operations. First, the row vector (the transposed difference vector) is multiplied by the square matrix (the inverse matrix), resulting in a new row vector. Then, this new row vector is multiplied by the column vector (the original difference vector), and according to the rules of matrix multiplication, the result is a single numerical value, i.e., a scalar value. Geometrically, this scalar value can be understood as the "squared generalized distance" from the current observation point to the center in the feature space after being "distorted" by the inverse matrix. Finally, the arithmetic square root of the scalar value is calculated, and the result is the Mahalanobis distance value, which is the output comprehensive deviation index. Calculating the arithmetic square root restores the "squared distance" to "distance", making the final index have an intuitive distance meaning. The Comprehensive Deviation Index is a non-negative scalar that quantifies the overall deviation of the behavioral pattern represented by the current fused feature vector from the individual's historical normal baseline. When the Comprehensive Deviation Index is zero, it indicates that the current behavioral pattern is exactly at the typical center of the individual's history. The larger the index value, the further away from the normality. Under the assumption of a multivariate normal distribution, the square of the Mahalanobis distance approximately follows a chi-square distribution. Therefore, the index value can be correlated with a certain probability confidence interval to determine the statistical significance of the deviation. The essence of Mahalanobis distance calculation is to measure the probabilistic statistical distance of the current observation point from the center of the distribution in a multivariate feature space centered on the baseline mean vector and described by the baseline covariance matrix. The core advantage of this method is that it performs a linear transformation on the original feature space through the inverse of the covariance matrix. In the transformed space, each feature dimension is "decoupled" (correlation is eliminated) and "standardized" (variance is normalized). Therefore, the calculated distance can equally and unbiasedly measure the joint deviation of the current observation point from its individual normality across all feature dimensions, overcoming the comparison difficulties caused by different feature dimensions and fluctuation amplitudes. Mahalanobis distance automatically considers the dimensional differences and intrinsic correlations between different features, and can sensitively integrate the synergistic changes of micro-eye movement features and macro-strategy indicators into a single overall deviation measure. For example, when two phenomena that may co-occur in cognitive decline, namely increased microsacrifice direction consistency index (micro-rigidity) and increased strategy deviation (macro-strategy inefficiency), occur, Mahalanobis distance calculation will identify this synergistic change that "conforms to historical correlation patterns" because they may have a positive correlation in the historical baseline (which is reflected in the covariance matrix), and give a larger distance value than simply looking at the changes of these two features separately, thus more sensitively indicating systematic behavioral abnormalities.

[0024] In this embodiment, it is necessary to specifically explain the process of generating multi-source evidence streams in the decision warning module as follows: First, the system receives the comprehensive deviation index, the sensitive micro-feature subset, and the persistent strategy deviation. The sensitive micro-feature subset is pre-selected and extracted from the high-dimensional eye-tracking micro-behavior feature vector, and includes at least the micro-salivation direction consistency index and the pupillary oscillation spectrum entropy. The comprehensive deviation index, the sensitive micro-feature subset, and the persistent strategy deviation are resampled at fixed time intervals to ensure that they have corresponding values ​​at the same time point in the same series, thus completing time synchronization. The fixed time interval is usually set to 0.1 seconds or 0.05 seconds to ensure the precision of evidence synchronization and match the input requirements of subsequent time series models. Next, for the comprehensive deviation index, the microsaccade direction consistency index and pupil oscillation spectrum entropy in the sensitive micro feature subset, and the continuous strategy deviation, the linear regression slope of each index within a recent sliding window of a preset time length is calculated. This linear regression slope is the short-term trend. The preset time length sliding window is usually three seconds. The process of calculating the linear regression slope is as follows: using the time of each sampling point within the sliding window as the independent variable X and the corresponding index value as the dependent variable Y, a straight line is fitted using the least squares method. The slope of this line is the short-term trend. Its positive or negative sign indicates the average change direction of the index within the window (increase or decrease), and the absolute value indicates the rate of change. Then, the current values ​​and short-term trends of the comprehensive deviation index, the microsaccade direction consistency index, the pupil oscillation spectrum entropy, and the strategy deviation are standardized. This involves subtracting the mean of the corresponding index from the pre-stored personal historical data of the operator, and then dividing by the standard deviation of the corresponding personal historical data to obtain the standardized values ​​of the current values ​​and short-term trends of each index. The pre-stored personal historical data of the operator refers to data recorded during the system calibration phase or during the operator's long-term normal work. This data is used to calculate the mean and standard deviation of each index (including its short-term trend), thereby establishing a personalized normal distribution reference. The standardization process transforms indicators of different dimensions and orders of magnitude and their changing trends into Z-scores with a mean of zero and a standard deviation of one, facilitating subsequent fusion and comparison. Finally, the standardized current value of the comprehensive deviation index and its standardized short-term trend, the standardized current value of the microsaccade direction consistency index and its standardized short-term trend, the standardized current value of the pupil oscillation spectrum entropy and its standardized short-term trend, and the standardized current value of the strategy deviation and its standardized short-term trend are concatenated and combined in a predetermined order to form a feature vector containing multiple dimensions to characterize the multi-source evidence flow. The predetermined order may be, for example,: [standardized current value of the comprehensive deviation index and its standardized trend, standardized current value of the microsaccade direction consistency index and its standardized trend, standardized current value of the pupil oscillation spectrum entropy and its standardized trend, standardized current value of the strategy deviation and its standardized trend]. Thus, the multi-source evidence flow at each moment is quantized into an eight-dimensional feature vector, which simultaneously encodes the immediate state and recent dynamics of four key pieces of evidence. The specific process of analyzing and calculating the probability of cognitive risk using the time-series reasoning model is as follows: The pre-built temporal reasoning model internally contains a deep learning network with a bidirectional gated recurrent unit network layer and a multi-head self-attention mechanism layer. The temporal reasoning model receives a time series as input, which consists of feature vectors representing a multi-source evidence stream at a consecutive preset number of time points. The consecutive preset number of time points usually correspond to a duration of five to ten seconds, for example, fifty to one hundred time points (intervals of 0.1 seconds). This time series captures the recent evolutionary history of the evidence stream. First, the input time series is fed into a bidirectional gated recurrent unit (BRN) network layer for processing. This BRN network layer can capture the dependencies between each time point in the time series and its context time points from both forward and backward directions, outputting a sequence of hidden state vectors containing temporal context information. The BRN network layer contains two independent recurrent neural networks, forward and backward. The forward network processes the time series in chronological order, while the backward network processes the time series in reverse chronological order. For each time point in the time series, the hidden state of the forward network at that point is concatenated with the hidden state of the backward network at that point to form the final hidden state vector. This allows the representation of each time point to incorporate its past and future context information, enabling a better understanding of the continuity of trends. Subsequently, the hidden state vector sequence is input into a multi-head self-attention mechanism layer. This layer can simultaneously calculate the association weights between any two hidden states at any given time point in the hidden state vector sequence, allowing the temporal reasoning model to focus on the historical moments and evidence interaction patterns most relevant to the current risk assessment, and output a context vector sequence weighted by attention weights. The multi-head self-attention mechanism layer simultaneously maps the hidden state vector sequence to multiple "representation subspaces." In each subspace, it calculates the attention score between all time point pairs in the hidden state vector sequence, which is based on the dot product of the query vector and the key vector and normalized to weights using the Softmax function. Then, the value vectors are weighted and summed using these weights to obtain the output of the subspace. The outputs of multiple subspaces are concatenated and linearly transformed to form the final context vector sequence. This enables the temporal reasoning model to focus on multiple cooperative or anomalous patterns at different time points in the evidence stream in parallel. Adaptive max pooling is performed on the context vector sequence to extract the most indicative feature information in the entire input time series, resulting in a pooled feature vector. Adaptive max pooling is not a fixed-size pooling window, but rather operates on each feature dimension of the entire context vector sequence, taking the maximum value of that dimension at all time points. This enables the time series inference model to extract the most significant (maximum value) activation signal in each feature dimension within the entire time window, and these activation signals often correspond to the moments when the risk pattern is most prominent. Finally, the pooled feature vector is processed through a fully connected layer and combined with a non-linear activation function to map and output a real value between zero and one. This real value is the cognitive risk probability. The closer the value is to one, the higher the risk that the worker is in a state of cognitive decline at the current moment. The fully connected layer maps the pooled feature vector to a scalar. The non-linear activation function usually uses the sigmoid function to compress the output to the (0,1) interval, which can be directly interpreted as a probability. This temporal reasoning model is trained under supervision using historical data containing known changes in cognitive state (such as fatigue experiments and simulated distraction tasks) to learn the complex mapping from multi-source evidence sequences to risk probabilities. The specific process of identifying collaborative degradation patterns and triggering early warnings is as follows: First, based on the standardized short-term trends of each piece of evidence, a co-degeneration index is calculated. For the four pieces of evidence—overall deviation index, microsacrifice direction consistency index, pupillary oscillation spectrum entropy, and strategy deviation—the deterioration direction is pre-defined. Specifically, the deterioration direction of the overall deviation index, microsacrifice direction consistency index, and strategy deviation is a positive increase, while the deterioration direction of the pupillary oscillation spectrum entropy is a negative decrease. The pre-defined deterioration direction is based on the physiological and behavioral manifestations of cognitive degeneration: an increase in overall deviation indicates abnormal overall behavior, an increase in microsacrifice direction consistency indicates rigidity in visual exploration, an increase in strategy deviation indicates inefficient strategy, and a decrease in pupillary oscillation spectrum entropy indicates reduced neural activity rhythm and increased cognitive load. When calculating the co-degradation index, it is checked whether the actual standardized short-term trend direction of each piece of evidence is consistent with its preset deterioration direction. If the direction is consistent, the absolute value of the standardized short-term trend of the evidence is mapped to a value between zero and one using a sigmoid function and included in the contribution. If the direction is inconsistent, the contribution is zero. The sigmoid function is specifically the sigmoid function. The calculation process is as follows: the absolute value of the standardized short-term trend is input into the sigmoid function, and a value between zero and one is output, representing the "intensity contribution" of the trend. The larger the absolute value of the trend, the closer the intensity contribution is to one. This step ensures that only those trends that change in the direction of deterioration and have a certain intensity are included. The average contribution of all four pieces of evidence is used to obtain a co-degradation index ranging from zero to one. This co-degradation index reflects how much key evidence is synchronously changing in a deteriorating direction and the strength of this change. A co-degradation index of one indicates that all four pieces of evidence are changing in a deteriorating direction with a strong trend; a value of zero indicates that no evidence is changing in a deteriorating direction. This index supplements the risk probability output by the time series model from the perspective of evidence "consistency" and "synchronicity". Triggering an alert requires meeting two conditions simultaneously: the first condition is that the current cognitive risk probability calculated by the time-series inference model exceeds a preset probability threshold; the second condition is that the currently calculated collaborative degradation index exceeds a preset collaborative threshold. The preset probability threshold is usually set between 0.6 and 0.8, for example, 0.7; the preset collaborative threshold is usually set between 0.4 and 0.6, for example, 0.5. The thresholds can be adjusted according to the tolerance for false alarms and false negatives in different application scenarios. Only when the perceived risk probability and the collaborative degradation index simultaneously exceed their respective preset probability thresholds and collaborative thresholds will the decision-making early warning module determine that the multi-source evidence stream exhibits a collaborative degradation mode and trigger an early warning signal for the operators. The "collaborative degradation mode" refers to a dynamic characteristic where not only is the overall risk probability high, but multiple underlying pieces of evidence also show a coordinated and consistent deterioration in the recent period. The dual-condition criterion greatly reduces false alarms caused by transient fluctuations of a single indicator or accidental misjudgments by the model, making the early warning signal more reliable and credible. The early warning signal can trigger audio-visual prompts, vibration feedback, or display warning information in the augmented reality interface.

[0025] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0026] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0027] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0028] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0029] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0030] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0031] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A safety risk early warning system for engineering workers based on eye-tracking monitoring, characterized in that, Specifically, it includes: The feature extraction module, strategy evaluation module, baseline comparison module, and decision warning module are sequentially connected via communication. Feature extraction module: preprocesses the acquired raw eye movement signal, performs micro-salivation dynamic analysis, gaze stability quantification and pupil oscillation spectrum analysis, extracts high-dimensional eye movement micro-behavioral feature vectors that characterize the micro-oscillation and dynamics of eye movement, and parses the real-time gaze coordinate sequence of the operator from the preprocessed raw eye movement signal; The strategy evaluation module receives high-dimensional eye-tracking micro-behavior feature vectors and real-time gaze coordinate sequences and obtains current task context information. It dynamically divides visual interest regions according to the semantics of the current task scenario, and calculates the weighted conditional entropy of visual scanning strategy randomness and task logic conformity based on the real-time gaze transfer sequence determined by the real-time gaze coordinate sequence. It also calculates the strategy deviation between the actual gaze transfer sequence and the optimal observer strategy generated offline through a pre-built reinforcement learning model, and generates a strategy effectiveness index. Baseline comparison module: Receives high-dimensional eye-tracking micro-behavioral feature vectors and strategy effectiveness indicators. When calling the pre-stored personalized behavior baseline of the monitored operator, it performs a multi-dimensional deviation measurement based on Mahalanobis distance between the feature vector (composed of high-dimensional eye-tracking micro-behavioral feature vectors, weighted conditional entropy, and strategy deviation degree) within the current time window and the personalized behavior baseline, and outputs a comprehensive deviation index. The decision-making early warning module receives the comprehensive deviation index, a subset of sensitive micro-features predefined and selected from the high-dimensional eye-tracking micro-behavioral feature vector, and the continuous policy deviation. It then uses the short-term trend of the comprehensive deviation index, the state of the sensitive micro-feature subset, and the state of the policy deviation as a multi-source evidence stream and inputs it into a pre-built temporal reasoning model for analysis. When the calculated cognitive risk probability exceeds a preset threshold and the multi-source evidence stream exhibits a collaborative degradation mode, an early warning is triggered.

2. The safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 1, characterized in that: The specific process of preprocessing the raw eye movement signal in the feature extraction module is as follows: First, the original eye movement signal containing the coordinates of the left pupil center, the coordinates of the right pupil center, the diameter signal of the left pupil, and the diameter signal of the right pupil is denoised, and the missing frames in the original eye movement signal are interpolated and spatially calibrated. Subsequently, gaze coordinate analysis and pupil signal analysis are performed simultaneously on the preprocessed raw eye movement signal; The specific process of gaze coordinate resolution is as follows: the calibrated coordinates of the left pupil center and the coordinates of the right pupil center are averaged to generate a real-time gaze coordinate sequence in time order. The specific process of pupil signal analysis is as follows: the validity of the left eye pupil diameter signal and the right eye pupil diameter signal are verified and fused and averaged to generate a single pupil diameter time series signal; The real-time gaze coordinate sequence and the pupil diameter time sequence signal together serve as the basic input data for subsequent microsaccade dynamic analysis, gaze stability quantification, and pupil oscillation spectrum analysis.

3. The safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 2, characterized in that: The specific process of performing micro-salivation dynamic analysis, fixation stability quantification, and pupillary oscillation spectrum analysis to construct a high-dimensional eye movement micro-behavioral feature vector is as follows: Based on the real-time gaze coordinate sequence, a dual-threshold detection algorithm based on motion speed and acceleration is used to identify all micro-saccade events with an amplitude of less than one degree of visual angle. For a set of micro-saccade events identified within a set sliding time window, the micro-saccade direction consistency index is calculated. Based on the real-time gaze coordinate sequence, gaze point events with a duration exceeding a preset duration threshold are identified. During the duration of a single gaze point event, the micro-fluctuation sequence of the pupil center coordinates of both eyes relative to their average position is extracted, and the corrected sample entropy of the micro-fluctuation sequence is calculated. The calculation result is used as a quantitative indicator of gaze stability. Based on the pupil diameter time-series signal, its spectrum as a function of time is obtained through short-time Fourier transform. The power spectral density within a predetermined frequency range related to neural oscillation is extracted, and the normalized entropy value of the power spectral distribution within this frequency range is calculated as the pupil oscillation spectral entropy. Finally, the microsaccade direction consistency index, gaze stability quantification index, and pupil oscillation spectrum entropy are combined to form a high-dimensional eye movement microbehavioral feature vector.

4. A safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 3, characterized in that: In the strategy evaluation module, the specific process of dynamically dividing visual interest regions and generating a gaze-dwelling region sequence based on the semantics of the current work scenario is as follows: First, the current task context information is loaded. The task context information defines the logical description and spatial layout of each functional unit in the scene in the form of structured data. Based on this, a set containing multiple visual interest regions is dynamically constructed. Subsequently, the received real-time gaze coordinate sequence is used in conjunction with the established set of visual interest regions to determine spatial affiliation. Each coordinate point in the real-time gaze coordinate sequence is assigned a visual interest region number. If the coordinate point is not in any preset visual interest region, it is marked as a background region. Next, random gazes are filtered out, and the continuous real-time gaze coordinate sequence is compressed into a discrete, time-ordered gaze dwell region sequence, where each element represents the visual interest region number where a valid gaze lingers. The gaze-dwelling region sequence serves as the basic input data for calculating the weighted conditional entropy and policy deviation.

5. A safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 4, characterized in that: The specific process for calculating the weighted conditional entropy and policy deviation to generate a policy performance index is as follows: First, based on the gaze dwelling region sequence, we calculate the empirical conditional probability of moving from each visual interest region to another, and the empirical probability of each visual interest region appearing in the gaze dwelling region sequence. A predefined task logic weight matrix is ​​introduced, in which each weight element reflects the necessity and rationality of the action of moving from one specified visual interest region to another in the current task context. Based on empirical conditional probability, empirical probability of occurrence of each visual interest region, and task logical weight matrix, the task-weighted visual scanning conditional entropy is obtained through the following calculation process. An optimal observer policy, pre-trained in a simulated environment using a reinforcement learning model, is employed. This optimal observer policy aims to minimize risk and contains an optimal state-action value function. For the actual scanning trajectory represented by the gaze dwell area sequence, the policy deviation is obtained through the following calculation process based on the optimal state action value function; finally, the calculated task-weighted visual scanning conditional entropy and the policy deviation are combined to form a policy effectiveness index.

6. A safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 5, characterized in that: In the baseline comparison module, the specific process of constructing a feature vector that integrates micro-features and macro-strategies and calling the personalized behavior baseline is as follows: First, after receiving the high-dimensional eye-tracking microbehavioral feature vector and the strategy effectiveness index, ensure that the high-dimensional eye-tracking microbehavioral feature vector and the strategy effectiveness index correspond to the same preset time window for analysis, and complete the time alignment. Subsequently, the multiple micro-eye movement features contained in the high-dimensional eye movement micro-behavior feature vector are sequentially concatenated with the task-weighted visual scanning conditional entropy and policy deviation contained in the policy effectiveness index in the feature dimension to form a new, higher-dimensional fusion feature vector. This fusion feature vector is a feature vector that combines micro-features and macro-policies, which is composed of the high-dimensional eye movement micro-behavior feature vector, weighted conditional entropy and policy deviation. At the same time, based on the identity of the currently monitored worker, the personalized behavior baseline pre-established for the worker is retrieved and called from the pre-stored personalized behavior baseline. The personalized behavior baseline includes the baseline mean vector and baseline covariance matrix obtained by statistical analysis of a large amount of historical normal work data of the worker.

7. A safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 6, characterized in that: The specific process for performing a multidimensional deviation measurement based on Mahalanobis distance to output a comprehensive deviation index is as follows: Based on the fusion feature vector constructed in the current time window, and the baseline mean vector and baseline covariance matrix of the corresponding personnel, the Mahalanobis distance is calculated as a comprehensive deviation index. The Mahalanobis distance calculation process is as follows: First, calculate the difference vector between the current fused feature vector and the baseline mean vector; Then, calculate the inverse of the baseline covariance matrix; transpose the difference vector to obtain the transposed difference vector; multiply the transposed difference vector with the inverse matrix, and then multiply the result with the original, untransposed difference vector to obtain a scalar value. Finally, the arithmetic square root of the scalar value is calculated, and the result is the Mahalanobis distance value, which is the output comprehensive deviation index. The overall deviation index is a non-negative scalar that quantifies the degree of deviation of the behavioral pattern represented by the current fused feature vector from the individual's historical normal baseline.

8. A safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 7, characterized in that: In the decision warning module, the process of generating multi-source evidence streams is as follows: First, the system receives the comprehensive deviation index, the sensitive micro-feature subset, and the persistent strategy deviation. The sensitive micro-feature subset is pre-selected and extracted from the high-dimensional eye-tracking micro-behavior feature vector, including the micro-salivation direction consistency index and the pupillary oscillation spectrum entropy. The comprehensive deviation index, the sensitive micro-feature subset, and the persistent strategy deviation are resampled at fixed time intervals to ensure that they have corresponding values ​​at the same time point in the same series, thus completing time synchronization. Next, for the comprehensive deviation index, the microsaccade direction consistency index and pupil oscillation spectrum entropy in the sensitive micro feature subset, and the continuous strategy deviation, the linear regression slope of each index within a recent sliding window of a preset time length is calculated. This linear regression slope is the short-term trend. Then, the current value and short-term trend of the comprehensive deviation index, the current value and short-term trend of the microsaccade direction consistency index, the current value and short-term trend of the pupil oscillation spectrum entropy, and the current value and short-term trend of the strategy deviation are standardized respectively to obtain the standardized values ​​of the current value and short-term trend of each index. Finally, the standardized current value of the comprehensive deviation index and its standardized short-term trend, the standardized current value of the microsaccade direction consistency index and its standardized short-term trend, the standardized current value of the pupil oscillation spectrum entropy and its standardized short-term trend, and the standardized current value of the strategy deviation and its standardized short-term trend are concatenated and combined in a predetermined order to form a feature vector containing multiple dimensions for representing multi-source evidence flow.

9. A safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 8, characterized in that: The specific process by which the temporal reasoning model analyzes and calculates the probability of cognitive risk is as follows: The pre-built temporal reasoning model contains a deep learning network with a bidirectional gated recurrent unit network layer and a multi-head self-attention mechanism layer; the temporal reasoning model receives a time series consisting of feature vectors representing multi-source evidence streams at a consecutive preset number of time points as input; First, the input time series is fed into a bidirectional gated recurrent unit network layer for processing. This bidirectional gated recurrent unit network layer can capture the dependency relationship between each time point in the time series and its context time points from both forward and backward directions, and output a hidden state vector sequence containing temporal context information. Subsequently, the hidden state vector sequence is input into a multi-head self-attention mechanism layer, which can simultaneously calculate the association weight between any two hidden states at any time point in the hidden state vector sequence and output a context vector sequence weighted by attention weights. An adaptive max pooling operation is performed on the context vector sequence to extract the most indicative feature information from the entire input time series, resulting in a pooled feature vector. Finally, the pooled feature vector is processed through a fully connected layer and combined with a non-linear activation function to map and output a real value between zero and one, which is the cognitive risk probability.

10. A safety risk early warning system for engineering workers based on eye-tracking monitoring according to claim 9, characterized in that: The specific process of identifying collaborative degradation patterns and triggering early warnings is as follows: First, based on the standardized short-term trends of each piece of evidence, a co-degradation index is calculated; this co-degradation index reflects the degree to which multiple pieces of evidence change synchronously in a predetermined deterioration direction. Triggering an alert requires two conditions to be met simultaneously: the first condition is that the current cognitive risk probability calculated by the time-series reasoning model exceeds a preset probability threshold; The second condition is that the currently calculated collaborative degradation index exceeds a preset collaborative threshold; Only when the probability of perceived risk and the synergistic degradation index both exceed their respective preset probability thresholds and synergistic thresholds is the multi-source evidence stream determined to exhibit a synergistic degradation mode, and an early warning signal is triggered for the operators.