A method and system for monitoring fatigue status of workers in high-risk industries based on EEG-eye tracking fusion.

CN122664682APending Publication Date: 2026-09-01GUANGDONG YUNNAO INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610842613.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0003]现有技术仍存在一些不足,在检测模态层面,单纯依赖脑电或眼动单一模态进行疲劳判别,易受个体差异、环境噪声、传感器接触状态等因素干扰,导致误报率高、鲁棒性差,脑电信号虽对疲劳敏感,却易混入肌电、眼电伪迹,而眼动特征亦受光照条件、头部姿态影响较大,单一模态的信息维度有限,难以全面表征疲劳状态的多层次、多维度演变规律,在融合策略层面,现有少量多模态融合方案多采用特征向量直接拼接或加权平均的方式,未充分考虑脑电与眼动两种异质生理信号在疲劳不同阶段的重要性差异,忽略了模态间的互补性与冗余性,导致融合后的特征判别力不足

Benefits of technology

通过脑电与眼动双模态融合显著提升疲劳检测精度与鲁棒性:脑电从神经中枢层面反映大脑警觉状态,眼动从行为层面反映注意力外在表征,二者信息互补、互为校验,有效降低单一模态受噪声干扰导致的误判风险,引入自注意力机制实现自适应跨模态权重分配,对不同疲劳阶段下脑电与眼动特征的重要性进行动态学习,使模型能够自适应关注当前最具判别力的模态,疲劳初期眼动变化更为敏感,充分挖掘模态间互补性,提升融合特征表征能力,采用庞加莱球双曲空间嵌入技术,将融合特征映射至负曲率双曲几何空间,利用其天然层次表征能力,通过测地线距离进行分类,相比欧氏空间分类器更自然刻画清醒至重度疲劳的递进层级关系,有效减少跨等级误判,提升分级准确性,采用频域映射与一维时序卷积结合的双分支轻量化特征提取网络,结构精简、参数少、计算量低,可直接部署于智能安全帽内置边缘计算单元,全链路端侧完成,无需云端通信,确保毫秒级响应与离线可用性,实现即戴即用的现场实时监测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122664682A_ABST
    Figure CN122664682A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for monitoring fatigue status of workers in high-risk industries based on EEG-eye movement fusion, belonging to the field of safety early warning technology. The method includes: real-time acquisition of workers' raw physiological signals using a smart safety helmet; performing bandpass filtering and baseline drift removal processing on the raw physiological signals to obtain multimodal physiological monitoring data containing denoised EEG signals and primary eye movement feature data including blink frequency and gaze duration; extracting blink interval temporal features and gaze point shift features from the primary eye movement feature data in the multimodal physiological monitoring data, and extracting the power spectral density of the EEG feature frequency band within a preset time window. This invention enables accurate classification and reliable early warning of worker fatigue status under complex and high-risk working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety early warning technology, and in particular to a method and system for monitoring the fatigue status of workers in high-risk industries based on EEG-eye fusion. Background Technology

[0002] Currently, fatigue detection technologies are mainly divided into three categories: subjective rating methods based on subjective scales, behavioral feature analysis methods based on video images, and objective detection methods based on physiological signals. Electroencephalogram (EEG) signals are hailed as the gold standard for fatigue detection because they can directly reflect changes in the alertness state of the central nervous system; eye movement behavior characteristics have also been proven to be correlated with the degree of fatigue.

[0003] Current technologies still have some shortcomings. At the detection modality level, relying solely on a single modality such as EEG or eye movement for fatigue discrimination is easily affected by individual differences, environmental noise, and sensor contact status, resulting in a high false alarm rate and poor robustness. Although EEG signals are sensitive to fatigue, they are easily mixed with EMG and EMG artifacts, while eye movement features are also greatly affected by lighting conditions and head posture. The information dimension of a single modality is limited, making it difficult to comprehensively represent the multi-level and multi-dimensional evolution of fatigue states. At the fusion strategy level, the few existing multimodal fusion schemes mostly adopt the method of directly splicing feature vectors or weighted averaging, without fully considering the difference in importance between EEG and eye movement, two heterogeneous physiological signals, at different stages of fatigue, and ignoring the complementarity and redundancy between modalities, resulting in insufficient discriminative power of the fused features.

[0004] At the classification and modeling level, fatigue itself is a progressive evolutionary process from wakefulness to early fatigue, moderate fatigue, and severe fatigue, with a natural hierarchical structure between each level. However, existing methods mostly use traditional classifiers in Euclidean space, which are difficult to effectively represent the hierarchy and progressiveness of fatigue levels, and are prone to cross-level misjudgments. At the deployment level, network conditions are often limited in high-risk industry work sites. Cloud-based fatigue detection solutions have prominent problems such as high communication latency and high privacy leakage risks. At the same time, existing deep learning models generally have a large number of parameters and high computational complexity, making it difficult to deploy and run on low-computing-power edge devices such as safety helmets, and failing to meet the needs of real-time early warning. At the early warning feedback level, existing early warning methods are mostly simple sound and light prompts, without differentiated alarms according to the severity of fatigue levels, and lack redundant feedback channels such as tactile feedback. Summary of the Invention

[0005] This invention provides a method and system for monitoring the fatigue status of workers in high-risk industries based on EEG-eye-tracking fusion, enabling accurate classification and reliable early warning of fatigue status of workers under complex and high-risk working conditions.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a method for monitoring fatigue status of workers in high-risk industries based on EEG-eye-tracking fusion, the method comprising: Step 1: Collect raw physiological signals of workers in real time using a smart safety helmet. Based on the raw physiological signals, perform bandpass filtering and baseline drift removal processing to obtain multimodal physiological monitoring data, which includes denoised EEG signals and primary eye movement feature data including blink frequency and gaze duration. Step 2: Based on the primary eye movement feature data in the multimodal physiological monitoring data, extract the blink interval temporal feature and gaze point shift feature, and extract the power spectral density of the EEG feature frequency band within a preset time window; input the power spectral density of the EEG feature frequency band, the blink interval temporal feature and the gaze point shift feature into a preset lightweight feature extraction network for parsing to obtain the EEG feature vector and the eye movement behavior feature vector; Step 3: Based on the EEG feature vector and eye movement behavior feature vector, perform cross-modal feature splicing and attention weight allocation calculation to obtain the multimodal fusion feature vector; Step 4: Input the multimodal fusion feature vector into the preset fatigue state classification model for probability mapping calculation to obtain the fatigue state confidence. Based on the maximum value in the fatigue state confidence, obtain the fatigue state classification result of the worker's current fatigue level. Step 5: When the fatigue state is determined to be early fatigue, moderate fatigue or severe fatigue, the edge computing early warning logic built into the smart safety helmet is triggered to obtain a fatigue early warning command containing the fatigue level and timestamp; according to the fatigue early warning command, the sound and light alarm component and vibration feedback component of the smart safety helmet are driven to obtain a real-time fatigue early warning signal.

[0007] Secondly, a fatigue monitoring system for workers in high-risk industries based on EEG-eye tracking fusion includes: The acquisition module is used to collect the raw physiological signals of workers in real time through the smart safety helmet. Based on the raw physiological signals, bandpass filtering and baseline drift removal processing are performed to obtain multimodal physiological monitoring data, which includes denoised EEG signals and primary eye movement feature data including blink frequency and gaze duration. The parsing module is used to extract blink interval temporal features and gaze point shift features from the primary eye movement feature data in the multimodal physiological monitoring data, and to extract the power spectral density of the EEG feature band within a preset time window; the power spectral density of the EEG feature band, blink interval temporal features and gaze point shift features are input into a preset lightweight feature extraction network for parsing to obtain EEG feature vectors and eye movement behavior feature vectors; The splicing module is used to perform cross-modal feature splicing and attention weight allocation calculation based on EEG feature vectors and eye-tracking behavior feature vectors to obtain multimodal fusion feature vectors; The calculation module is used to input the multimodal fusion feature vector into the preset fatigue state classification model for probability mapping calculation, obtain the fatigue state confidence score, and obtain the fatigue state classification result of the current fatigue level of the worker based on the maximum value of the fatigue state confidence score. The driving module is used to trigger the edge computing early warning logic built into the smart safety helmet when the fatigue state is determined to be early fatigue, moderate fatigue or severe fatigue, and obtain a fatigue warning command containing fatigue level and timestamp; according to the fatigue warning command, it drives the sound and light alarm component and vibration feedback component of the smart safety helmet to obtain a real-time fatigue warning signal.

[0008] Thirdly, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0009] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0010] The above-described solution of the present invention has at least the following beneficial effects: The fusion of EEG and eye-tracking modalities significantly improves the accuracy and robustness of fatigue detection: EEG reflects the brain's alertness at the central nervous system level, while eye-tracking reflects the external representation of attention at the behavioral level. The two modalities complement and verify each other, effectively reducing the risk of misjudgment caused by noise interference in a single modality. A self-attention mechanism is introduced to achieve adaptive cross-modal weight allocation, dynamically learning the importance of EEG and eye-tracking features at different fatigue stages. This allows the model to adaptively focus on the most discriminative modality at the current stage, with eye-tracking changes being more sensitive in the early stages of fatigue. By fully exploring the complementarity between modalities, the fusion feature representation capability is enhanced. A Poincaré sphere is employed. Hyperbolic space embedding technology maps fused features to a negative curvature hyperbolic geometric space. Utilizing its natural hierarchical representation capabilities, it classifies features by geodesic distance. Compared to Euclidean space classifiers, it more naturally depicts the progressive hierarchical relationship from alertness to severe fatigue, effectively reducing cross-level misjudgments and improving classification accuracy. It employs a dual-branch lightweight feature extraction network combining frequency domain mapping and one-dimensional temporal convolution. The structure is simplified, with fewer parameters and lower computational cost. It can be directly deployed in the built-in edge computing unit of smart safety helmets, completing the entire link on the edge without cloud communication, ensuring millisecond-level response and offline availability, and enabling on-site real-time monitoring that is ready to use immediately upon wearing. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating a method for monitoring fatigue status of workers in high-risk industries based on EEG-eye fusion, as provided in an embodiment of the present invention.

[0012] Figure 2 This is a schematic diagram of a fatigue monitoring system for workers in high-risk industries based on EEG-eye fusion, provided by an embodiment of the present invention.

[0013] Figure 3 This is a graph showing the trend of EEG thermistor band power over time.

[0014] Figure 4 It is a heatmap of cross-modal attention weights. Detailed Implementation

[0015] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0016] like Figure 1 As shown, embodiments of the present invention propose a method for monitoring fatigue status of workers in high-risk industries based on EEG-eye fusion, the method comprising the following steps: Step 1: Collect raw physiological signals of workers in real time using a smart safety helmet. Based on the raw physiological signals, perform bandpass filtering and baseline drift removal processing to obtain multimodal physiological monitoring data, which includes denoised EEG signals and primary eye movement feature data including blink frequency and gaze duration. Step 2: Based on the primary eye movement feature data in the multimodal physiological monitoring data, extract the blink interval temporal feature and gaze point shift feature, and extract the power spectral density of the EEG feature frequency band within a preset time window; input the power spectral density of the EEG feature frequency band, the blink interval temporal feature and the gaze point shift feature into a preset lightweight feature extraction network for parsing to obtain the EEG feature vector and the eye movement behavior feature vector; Step 3: Based on the EEG feature vector and eye movement behavior feature vector, perform cross-modal feature splicing and attention weight allocation calculation to obtain the multimodal fusion feature vector; Step 4: Input the multimodal fusion feature vector into the preset fatigue state classification model for probability mapping calculation to obtain the fatigue state confidence. Based on the maximum value in the fatigue state confidence, obtain the fatigue state classification result of the worker's current fatigue level. Step 5: When the fatigue state is determined to be early fatigue, moderate fatigue or severe fatigue, the edge computing early warning logic built into the smart safety helmet is triggered to obtain a fatigue early warning command containing the fatigue level and timestamp; according to the fatigue early warning command, the sound and light alarm component and vibration feedback component of the smart safety helmet are driven to obtain a real-time fatigue early warning signal.

[0017] In this embodiment of the invention, by simultaneously acquiring EEG and eye-tracking modal data and adaptively fusing self-attention across modalities, combined with Poincaré sphere hyperbolic space hierarchical classification, dual-branch lightweight feature extraction network edge-side inference, and fatigue level differentiated audio-visual vibration multi-channel redundant early warning, the defects of single modality being susceptible to interference and having a high false alarm rate, fixed fusion ignoring stage differences, serious cross-level misjudgment in Euclidean space, difficulty in deploying deep models at the edge, and single prompts being easily ignored are overcome. Thus, the effect of fatigue state classification, edge-side millisecond-level offline real-time monitoring, and reliable perception of multi-channel early warning signals is achieved under complex and high-risk working conditions, ensuring the personal safety of workers.

[0018] In a preferred embodiment of the present invention, step 1 above may include: Step 11: Simultaneously collect the worker's raw EEG waveform data and raw eye video stream data using the EEG sensor and eye-tracking sensor integrated in the smart safety helmet; based on the raw EEG waveform data, perform bandpass filtering to remove power frequency interference and EMG artifacts, and remove low-frequency noise according to baseline drift correction processing to obtain the denoised EEG signal. Specifically, this includes: collecting the worker's raw EEG waveform data at a preset sampling frequency using the EEG sensor integrated in the smart safety helmet. The smart safety helmet includes a sensing acquisition unit, an edge computing unit, a warning execution unit, and a power supply unit. The sensing acquisition unit includes an EEG sensor group and an eye-tracking sensor group. The EEG sensor group uses 4 to 8 Ag / AgCl dry electrodes distributed on the forehead and occipital region, with a reference electrode and a ground electrode. The signal is output after pre-differential amplification and 24-bit analog-to-digital conversion. The eye-tracking sensor group uses 2 to 4 850nm infrared LEDs in conjunction with an infrared camera to achieve occult eye tracking based on the pupil-corneal reflection method. The edge computing unit uses ARM... The processor is based on a Cortex-M7 or RISC-V architecture and integrates all executable code and weight parameters for the lightweight feature extraction network, self-attention mechanism network, fatigue state classification model, and edge computing early warning logic. It connects to various sensors via SPI / I2C bus and includes a built-in real-time clock module to provide a unified time reference. The early warning execution unit includes a red and yellow dual-color LED array, a buzzer module, and an eccentric rotary vibration motor, driven by control parameters output from the edge computing early warning logic. The power supply unit is a rechargeable lithium battery with a continuous working time of ≥8 hours. All the above units work together to complete a closed-loop workflow of synchronous acquisition, edge inference, and hierarchical early warning. Based on the original EEG waveform data, bandpass filtering is performed to remove power frequency interference and high-frequency electromyography artifacts, resulting in a bandpass-filtered EEG signal. Specifically, a Butterworth bandpass filter is used, whose transfer function amplitude-frequency response is: ; In the formula, For the filter at angular frequency The amplitude-frequency response gain at that point is dimensionless. Angular frequency of the input signal, expressed in radians per second; The center angular frequency of the passband is expressed in radians per second and is determined by the geometric mean of the lower and upper cutoff frequencies of the passband. The passband bandwidth is expressed in radians per second and is equal to the difference between the upper and lower cutoff angular frequencies. This is the order of the Butterworth filter, taken as a positive integer. A higher order results in a steeper transition from the passband to the stopband, and a wider passband coverage. Wave, Wave, Wave, Affecting low Effective frequency bands for EEG, such as waves.

[0019] Based on the bandpass-filtered EEG signal, baseline drift correction is performed to remove low-frequency drift noise caused by the slow change in electrode-skin contact impedance, resulting in a denoised EEG signal. Specifically, a wavelet transform-based baseline correction method is used to perform multi-level discrete wavelet decomposition on the bandpass-filtered signal. The decomposition expression is as follows: ; In the formula, For bandpass filtered EEG signals at time 10:00 The amplitude, in microvolts; For the first Layer Each approximation coefficient represents the lowest frequency approximation component in the signal, corresponding to low-frequency components such as baseline drift; For the first Layer The detail coefficients characterize the signal at the _ . Detailed components at the layer decomposition scale correspond to the effective components of EEG in each frequency band; For the first The scaling function of the layer at time The value of is used to reconstruct the low-frequency approximation signal; For the first The wavelet function of the layer at time The value is used to reconstruct the detail signals of each layer; This represents the total number of layers in the discrete wavelet decomposition. This is the decomposition layer index number, with values ​​ranging from 1 to... ; The translation parameter index controls the position of the wavelet function on the time axis, and all approximation coefficients are used. After setting the values ​​to zero, the signal is reconstructed using detail coefficients and wavelet functions, thus obtaining the denoised EEG signal after removing baseline drift.

[0020] Step 12: Based on the original eye video stream data, perform eye key point localization and pupil center tracking. Calculate the blink count and gaze duration per unit time based on the temporal sequence of pupil center position changes, obtaining primary eye movement feature data including blink frequency and gaze duration. Specifically, this includes: acquiring the worker's original eye video stream data at a preset frame rate using an eye-tracking sensor integrated into the smart safety helmet; performing eye key point localization on each frame of the original eye video stream data; and iteratively locating the eyelid contour and corner of the eye coordinates using an ensemble regression tree algorithm based on gradient boosting regression trees. The algorithm obtains an image sequence of the eye region. Starting with an initial shape estimate, a cascaded regressor calculates the shape increment and updates the shape estimate based on the current shape estimate and image texture features. After multiple iterations until convergence, the final keypoint coordinates are output. Based on the eye region image sequence, pupil center tracking is performed on each frame of the eye region image. An ellipse fitting algorithm is used to locate the pupil center coordinates. Adaptive threshold binarization segmentation is performed on the eye region image to extract the pupil contour region. Based on the boundary point set of this contour region, the least squares ellipse fitting method is used to solve for the pupil center coordinates. The optimization objective function is: ; In the formula, This represents the total number of sampling points in the set of pupil contour boundary points; This is the index of the contour sampling point, with values ​​ranging from 1 to... ; For the first The x-coordinate value of each contour sampling point, in pixels; For the first The ordinate value of each contour sampling point, in pixels; Here is the x-coordinate of the center of the ellipse to be solved, which is the x-coordinate of the pupil center, in pixels; The ordinate of the center of the ellipse to be solved is the ordinate of the pupil center, in pixels. The radius estimate of the ellipse to be fitted is the pupil radius, expressed in pixels. Solving this minimization problem yields temporal data of the pupil center coordinates. Based on this temporal data, the number of blinks and gaze duration per unit time are calculated, resulting in primary eye movement feature data containing blink frequency and gaze duration. The formula for calculating blink frequency is: ; In the formula, This is the blink frequency value, measured in blinks per second. The total number of blink events detected per unit time. The criteria for determining a blink are that the pupil is lost for several consecutive frames and the duration of the loss is between the preset minimum closing time and the preset maximum closing time. To calculate the length of the statistical time window, in seconds, the formula for the gaze duration is: ; In the formula, The duration of gazing is a dimensionless percentage within the statistical time window. To count the total number of frames that were identified as staring within the time window; This is the index of the gaze frame, with values ​​ranging from 1 to... ; For the first The duration of a gaze segment is measured in seconds. The criteria for determining a gaze segment are that the displacement of the pupil center coordinates between consecutive frames is less than a preset displacement threshold and the duration of this low displacement state exceeds a preset minimum gaze duration threshold. The length of the statistical time window is in seconds.

[0021] Step 13: Based on the denoised EEG signal and primary eye movement feature data, perform timestamp alignment processing to obtain multimodal physiological monitoring data. Specifically, this includes: extracting the acquisition timestamp sequences of the denoised EEG signal and primary eye movement feature data. The EEG data timestamp sequence is generated by the EEG sensor sampling at a preset sampling frequency at equal intervals, and the eye movement data timestamp sequence is generated by the eye movement tracking sensor sampling at a preset frame rate at equal intervals. Both use the same clock source as a unified time reference. Based on the EEG timestamp sequence and the eye movement timestamp sequence, and using the EEG sampling time point as a reference, perform linear interpolation resampling alignment on the eye movement data to obtain time-aligned multimodal physiological monitoring data. For each sampling moment in the EEG timestamp sequence, find the two adjacent eye movement recording moments in the eye movement timestamp sequence, and perform linear interpolation on the eye movement feature data using the following formula: ; In the formula, Alignment to the first value after interpolation resampling The dimensions of the eye movement feature values ​​at each EEG sampling time are consistent with the interpolated eye movement features. The first in the EEG timestamp sequence Each sampling time is measured in seconds. For the eye-tracking timestamp sequence located in The previous and most recent sampling time, in seconds; For the eye-tracking timestamp sequence located in The time immediately following and the time closest to it, in seconds; for The original eye movement feature value corresponding to the time moment; for The original eye movement feature value corresponding to the time moment; For eye-tracking timestamp sequences that satisfy The frame number index of the condition is used to perform interpolation point by point for all EEG sampling times according to the above formula, and finally outputs multimodal physiological monitoring data with unified timestamps. Each sampling point contains both the denoised EEG signal value and the interpolated eye movement feature value.

[0022] In this embodiment of the invention, by simultaneously acquiring EEG and eye-tracking modal data and adaptively fusing self-attention across modalities, combined with Poincaré sphere hyperbolic space hierarchical classification, dual-branch lightweight feature extraction network edge-side inference, and fatigue level differentiated audio-visual vibration multi-channel redundant early warning, the defects of single modality being susceptible to interference and having a high false alarm rate, fixed weight fusion ignoring stage differences, serious cross-level misjudgment in Euclidean space, difficulty in deploying deep models at the edge, and single prompts being easily ignored are overcome. Thus, the effect of fatigue state classification, edge-side millisecond-level offline real-time monitoring, and reliable perception of multi-channel early warning signals is achieved under complex and high-risk working conditions, ensuring the personal safety of workers.

[0023] In a preferred embodiment of the present invention, step 2 above may include: Step 21: Based on the blink frequency in the primary eye movement feature data, calculate the time interval sequence between two adjacent blinks to obtain the blink interval temporal feature. Specifically, this includes: extracting the occurrence time of each blink event based on the blink frequency in the primary eye movement feature data, arranging them in chronological order to obtain a blink time sequence, and defining the time interval between two adjacent blinks as the blink interval. Iterating through all blink times and calculating the adjacent differences sequentially yields the blink interval temporal feature. The specific calculation formula is as follows: ; In the formula, For the first The blink interval value is in seconds; For the first The time of each blink is measured in seconds; The next one The time of each blink is measured in seconds; The index is the sequence number of the blink interval, with values ​​ranging from 1 to... ; This represents the total number of blinking events detected within the statistical period.

[0024] Step 22: Based on the gaze duration in the primary eye-tracking feature data, calculate the spatial coordinate offset of the gaze focus within the preset work area to obtain the gaze point offset feature; based on the denoised EEG signal in the multimodal physiological monitoring data, extract the preset time window data corresponding to the current work cycle, specifically including: based on the gaze duration in the primary eye-tracking feature data, extract the spatial coordinate sequence of the pupil center in the preset work area coordinate system within each gaze period, calculate the spatial offset of the gaze point relative to the reference point of the work area, and take the average of the coordinate offsets of all sampling points within the gaze period to obtain the gaze point offset feature for that gaze period. The specific calculation formula is as follows: ; In the formula, The average distance of gaze point offset during a single gaze period, in pixels; This represents the total number of sampling points included within the current gaze period; This is the index of the sampling point within the gaze period, with values ​​ranging from 1 to... ; For the first The x-coordinate value of the pupil center at each sampling time, in pixels; For the first The vertical coordinate value of the pupil center at each sampling time, in pixels; The x-coordinate value of the reference point within the preset work area, in pixels; The ordinate value of the reference point within the preset work area is given in pixels. Based on the denoised EEG signal from the multimodal physiological monitoring data, a continuous sampling point sequence of a preset length is extracted, starting from the start time of the current work cycle, to obtain the preset time window data. The extraction range is the continuous EEG sampling segment from the start time of the current work cycle to a preset time extension after that time. Step 23: Based on the data within the preset time window, perform Fast Fourier Transform (FFT) processing to extract energy distribution parameters including the alpha, beta, and theta bands, obtaining the power spectral density of the EEG characteristic frequency bands. Specifically, this includes: based on the data within the preset time window, performing FFT processing to convert the time-domain EEG signal into a frequency-domain representation, obtaining the spectral distribution of the EEG signal. The calculation formula for FFT is: ; In the formula, For the first Each frequency index corresponds to a complex Fourier transform value, which includes amplitude and phase information; For the first time within the preset time window The amplitude of the EEG signal at each sampling point is expressed in microvolts. This is the index of the time-domain sampling point number, with values ​​ranging from 0 to... ; This represents the total number of sampling points for EEG signals within the time window. This is the frequency domain index number, with values ​​ranging from 0 to... Each Corresponding to a discrete frequency point; is the base of the natural logarithm; The imaginary unit; For pi, calculate the power spectral density at each frequency point based on the Fourier transform results. The formula for calculating the power spectral density is: ; In the formula, For the first The power spectral density value corresponding to each frequency point, in microvolt squared per hertz; For the first Fourier transform results at each frequency point The squared value of the modulus; This represents the total number of sampling points for EEG signals within the time window. Frequency resolution, or the frequency interval between two adjacent discrete frequency points, is measured in Hertz (Hz) and is equal to the sampling frequency divided by the total number of sampling points. Based on the power spectral density calculation results, the power spectral density in the alpha, beta, and theta bands is integrated to obtain the energy distribution parameters for each band. The formula for calculating the energy in each band is as follows: ; In the formula, The total energy value of a specified frequency band within a time window, expressed in microvolt squared. This is the frequency domain index number corresponding to the lower limit frequency of this frequency band; This is the frequency domain index number corresponding to the upper limit frequency of this frequency band; For the first The power spectral density value corresponding to each frequency point, in microvolt squared per hertz; The frequency resolution is measured in Hertz. The alpha band corresponds to a frequency range of 8 Hz to 13 Hz, the beta band corresponds to a frequency range of 13 Hz to 30 Hz, and the theta band corresponds to a frequency range of 4 Hz to 8 Hz. The final output contains the power spectral density of the EEG characteristic bands, which includes the energy values ​​of the three frequency bands.

[0025] like Figure 3 As shown: The figure shows the trend of EEG power in the Theta band over time. The horizontal axis represents time, and the vertical axis represents the power value of the Theta band. The Theta band corresponds to the 4-8Hz frequency component in the EEG signal. The activity in this frequency band is directly related to the degree of drowsiness and the state of distraction in the human body.

[0026] In this embodiment of the invention, by simultaneously acquiring EEG and eye-tracking modal data and adaptively fusing self-attention across modalities, combined with Poincaré sphere hyperbolic space hierarchical classification, dual-branch lightweight feature extraction network edge-side inference, and fatigue level differentiated audio-visual vibration multi-channel redundant early warning, the defects of single modality being susceptible to interference and having a high false alarm rate, fixed weight fusion ignoring stage differences, serious cross-level misjudgment in Euclidean space, difficulty in deploying deep models at the edge, and single prompts being easily ignored are overcome. Thus, the effect of fatigue state classification, edge-side millisecond-level offline real-time monitoring, and reliable perception of multi-channel early warning signals is achieved under complex and high-risk working conditions, ensuring the personal safety of workers.

[0027] In a preferred embodiment of the present invention, step 2 above may include: Step 24: Based on the blink interval temporal features and gaze point offset features, perform feature dimension concatenation processing to obtain an eye movement feature concatenation sequence. Specifically, this includes: based on the blink interval temporal feature sequence and gaze point offset feature sequence, using a preset time window as the basic analysis unit, extracting statistical features from the blink interval subsequence and gaze offset subsequence within each time window, and then concatenating the two types of statistical features along the vector dimension to obtain the eye movement feature concatenation sequence. For the preset time window, extract all blink interval values ​​falling within that window from the blink interval temporal features to form the blink interval subsequence for that window. Calculate the arithmetic mean of this subsequence using the formula: ; In the formula, For the first The arithmetic mean of blink intervals within a time window, in seconds; For the first The total number of blink intervals contained within a time window; For the first The index of the blink interval within each window. ; For the first Within the first time window Each blink interval is a value in seconds. Based on the mean blink interval within the window, the standard deviation of the blink intervals within that window is calculated using the following formula: In the formula, For the first The standard deviation of blink intervals within a time window, in seconds.

[0028] Based on the gaze point offset feature sequence, extract the points falling into the first... The total gaze offset values ​​within a given time window constitute a gaze offset subsequence for that window. The arithmetic mean of this subsequence is calculated using the following formula: ; In the formula, For the first The arithmetic mean of the gaze point offset distance within a time window, in pixels; For the first The total number of gaze events contained within a time window; For the first The sequence index of gaze events within a window. ; For the first Within the first time window The offset distance value of the second gaze, in pixels, is used to calculate the standard deviation of the gaze offset within the window based on the mean gaze offset of the window. The formula is as follows: ; In the formula, For the first The standard deviation of the gaze point offset distance within each time window, in pixels, is concatenated based on the four statistical characteristics mentioned above along the vector dimension to obtain the first... The formula for the eye-tracking feature concatenation vector corresponding to each time window is: ; In the formula, For the first The concatenated vector of eye-tracking features corresponding to each time window is a 4-dimensional column vector; superscript The transpose operation represents the operation on all vectors. The stitching operation is performed sequentially in each time window to obtain the eye-tracking feature stitching sequence.

[0029] Step 25: Input the eye-tracking feature concatenation sequence into the temporal convolution branch of the pre-set lightweight feature extraction network, perform one-dimensional convolution and pooling operations to obtain an eye-tracking behavior feature vector representing the temporal evolution of eye-tracking behavior. Specifically, this includes: arranging the eye-tracking feature concatenation sequence into a two-dimensional feature matrix according to the time window sequence, and inputting it into the temporal convolution branch of the pre-set lightweight feature extraction network. The pre-set lightweight feature extraction network is a compact neural network that has completed all parameter learning and has been frozen. It cannot be changed online during the inference and deployment phase and is directly embedded in the edge device. The network employs two core technologies—depth-separable convolution and channel pruning—to compress the number of parameters and computational load, enabling efficient real-time feature extraction of EEG and eye-tracking signals on edge chips where storage resources and computing power are limited. It sequentially performs one-dimensional convolution, nonlinear activation, one-dimensional pooling, and fully connected mapping operations to obtain eye-tracking behavior feature vectors representing the temporal evolution of eye-tracking behavior. The feature vectors of several adjacent time windows in the eye-tracking feature concatenation sequence are stacked along the time axis to form an input feature matrix. A one-dimensional convolution operation is then performed on this matrix, with the convolution kernel sliding along the time axis. The calculation formula is as follows: ; In the formula, For the first Each convolutional kernel at the start of time The convolution output value at that point is dimensionless. This is the starting index of the convolutional kernel on the time axis; This is the kernel index. , This represents the total number of convolutional kernels in the convolutional layer. This is the time offset index inside the convolution kernel. , The length of the one-dimensional convolutional kernel in the time dimension; The dimension index of the input feature vector. ; For the first Each convolutional kernel at time offset Feature Dimension Learnable weight parameters at the location; The first in the eye movement feature splicing sequence The feature vector corresponding to the time window is at the _ ... Component values ​​in a dimension; For the first The learnable bias parameters corresponding to each convolutional kernel are processed by batch normalization based on the convolution output, and then the ReLU nonlinear activation function is used for feature activation. The calculation formula is as follows: ; In the formula, The output of the convolution is the activation value after the ReLU activation function is applied; This represents the element-wise operation between zero and the larger of the input values. Based on the activated feature maps, a one-dimensional max pooling operation is performed to reduce the time dimension and extract significant features. The calculation formula is as follows: ; In the formula, For the first The feature map corresponding to the convolutional kernel is pooled and then... The value of each pooling output position; For pooling output position index; The length of the pooling window; The pooling step size; This indicates that the maximum value is taken within the pooling window. Based on the pooling output, the pooling outputs of all convolutional kernels are flattened into a one-dimensional vector, which is then mapped through a fully connected layer to obtain the eye-tracking behavior feature vector. The calculation formula is as follows: ; In the formula, The final output eye-tracking behavior feature vector is denoted as . ; The learnable weight matrix of the fully connected layer. ,in, It is the set of real numbers, meaning that all elements in the matrix are real numbers. The output dimension, i.e., the number of rows in the matrix, is equal to the length of the output vector of the fully connected layer. Input dimension, i.e., the number of columns in the matrix; This is a one-dimensional column vector resulting from the flattening operation of all pooled outputs. is the learnable bias vector for the fully connected layer.

[0030] Step 26: Input the power spectral density of the EEG feature frequency bands into the frequency domain mapping branch of the lightweight feature extraction network to perform preliminary feature mapping processing and obtain the frequency domain mapping feature matrix. Specifically, this includes: constructing the EEG frequency band energy input vector from the alpha, beta, and theta band energy values ​​based on the power spectral density of the EEG feature frequency bands, inputting it into the frequency domain mapping branch of the preset lightweight feature extraction network, and sequentially performing the first fully connected mapping, nonlinear activation, and second fully connected mapping to expand the low-dimensional frequency band energy to a higher-dimensional feature space, and obtaining the frequency domain mapping feature matrix after shape reshaping. The input vector is constructed based on the three frequency band energy values, using the following formula: ; In the formula, The brainwave frequency band energy input vector is a 3-dimensional column vector. This represents the energy value in the alpha band, expressed in microvolt squared (µV). This represents the beta band energy value, expressed in microvolt squared (µV). The energy value of the Theta band, in microvolt squares, is initially mapped to a higher dimension through the first fully connected layer based on the EEG band energy input vector. The calculation formula is as follows: ; In the formula, Let be the frequency domain mapping vector output by the first fully connected layer, with dimension denoted as . ; Here is the learnable weight matrix of the first fully connected layer, with a matrix size of [value missing]. Rows x 3 columns; Here is the learnable bias vector of the first fully connected layer, with dimension . Based on the output of the first fully connected layer, a ReLU function is used for nonlinear activation, and the calculation formula is as follows: ; In the formula, The output vector of the first fully connected layer after ReLU activation has a dimension of . ; This represents the operation between taking zeros element-wise and the larger of the input value. Based on the activated feature vector, the feature dimension is further expanded through a second fully connected layer. The calculation formula is: ; In the formula, Let be the frequency domain mapping vector output by the second fully connected layer, with dimension denoted as . ; Here is the learnable weight matrix for the second fully connected layer, with a size of [missing information]. Line × List; Here is the learnable bias vector for the second fully connected layer, with dimension . Based on the output vector of the second fully connected layer, a shape reshaping operation is performed to transform it into a two-dimensional matrix form, resulting in the frequency domain mapping feature matrix, as shown in the formula: ; In the formula, The final output frequency domain mapping feature matrix has a size of . Line × Column, satisfying ; This represents rearranging vectors in row-major order to form a matrix with a specified number of rows and columns; This represents the number of rows in the reshaped matrix. This represents the number of columns in the reshaped matrix.

[0031] Step 27: Based on the frequency domain mapping feature matrix, perform dimensionality reduction and nonlinear activation processing using a fully connected layer to extract the nonlinear mapping relationship between the EEG frequency bands, obtaining the EEG feature vector. Specifically, this includes: based on the frequency domain mapping feature matrix, sequentially performing matrix flattening, fully connected layer dimensionality reduction, and hyperbolic tangent nonlinear activation processing to extract the nonlinear mapping relationship between the three EEG feature frequency bands, obtaining the EEG feature vector. Based on the frequency domain mapping feature matrix, flatten it into a one-dimensional column vector, as shown in the formula: ; In the formula, This is a one-dimensional column vector of the frequency domain mapping feature matrix after flattening, with dimension 1. ; This represents the operation of expanding a matrix into column vectors in row-major order. Based on the flattened vectors, feature compression is performed through a dimensionality-reduction fully connected layer, mapping the high-dimensional frequency domain mapping vector to the same feature space dimension as the eye-tracking behavior feature vector. The calculation formula is as follows: ; In the formula, Let be the linear mapping vector output by the dimension-reduced fully connected layer, denoted as . ; The learnable weight matrix for the dimensionality-reduced fully connected layer has a size of . Line × List; The learnable bias vector for the dimension-reduced fully connected layer has a dimension of . Based on the output of the dimension-reduced fully connected layer, a hyperbolic tangent function is used for nonlinear activation to extract the nonlinear mapping relationship of the EEG frequency bands. The calculation formula is as follows: ; In the formula, The final output EEG feature vector has a dimension of Each component takes values ​​between -1 and 1; It is a hyperbolic tangent activation function that acts element-wise on each component of the input vector; Represented by natural constant An exponential function with base 0.

[0032] In this embodiment of the invention, by simultaneously acquiring EEG and eye-tracking modal data and adaptively fusing self-attention across modalities, combined with Poincaré sphere hyperbolic space hierarchical classification, dual-branch lightweight feature extraction network edge-side inference, and fatigue level differentiated audio-visual vibration multi-channel redundant early warning, the defects of single modality being susceptible to interference and having a high false alarm rate, fixed weight fusion ignoring stage differences, serious cross-level misjudgment in Euclidean space, difficulty in deploying deep models at the edge, and single prompts being easily ignored are overcome. Thus, the effect of fatigue state classification, edge-side millisecond-level offline real-time monitoring, and reliable perception of multi-channel early warning signals is achieved under complex and high-risk working conditions, ensuring the personal safety of workers.

[0033] In a preferred embodiment of the present invention, step 3 above may include: Step 31: Based on the EEG feature vector and the eye-tracking behavior feature vector, perform feature dimension concatenation processing to obtain a cross-modal concatenated feature matrix. Specifically, this includes: concatenating the two feature vectors end-to-end along their feature dimension to obtain a merged long column vector; then reorganizing this long column vector into a two-dimensional matrix with two rows according to a preset row-first arrangement to obtain the cross-modal concatenated feature matrix. The EEG feature vector and the eye-tracking behavior feature vector have the same dimension length. The formula for concatenating them end-to-end into a unified column vector is as follows: ; In the formula, The concatenated cross-modal feature column vector has a total length that is twice the length of the dimension of a single modal feature. For the EEG feature vector at the th Component values ​​in each dimension The value ranges from 1 to All integers; represents the dimensional length of the EEG feature vector and the eye-tracking behavior feature vector, respectively. The dynamic behavior feature vector at the th Component values ​​in each dimension; superscript The transpose operation represents the operation of a vector or matrix. Based on the concatenated cross-modal feature column vectors, it rearranges them into a two-dimensional matrix in row-major order. The recombination formula is as follows: ; In the formula, This is a cross-modal concatenated feature matrix, consisting of two rows. The first row corresponds to the original feature vectors of the EEG modality, and the second row corresponds to the original feature vectors of the eye-tracking modality. Each row has 10 columns. .

[0034] Step 32: Input the cross-modal splicing feature matrix into a pre-set self-attention mechanism network for weight calculation to obtain attention weight coefficients representing the importance of different physiological signals. Specifically, this includes: calculating the query matrix and key matrix by passing the cross-modal splicing feature matrix through two independent and learnable linear mapping layers; performing matrix multiplication on the transposes of the query matrix and key matrix, and scaling them with a scaling factor to obtain the attention score matrix; independently performing Softmax normalization on each row of this score matrix to obtain the attention weight coefficient matrix representing the degree of mutual attention between different physiological modal signals; and obtaining the query matrix by linear mapping through the first learnable weight matrix based on the cross-modal splicing feature matrix, using the following calculation formula: ; In the formula, For query matrix; To query the learnable weight matrix corresponding to the mapping, the key matrix is ​​obtained by linear mapping through the second learnable weight matrix based on the same cross-modal concatenated feature matrix. The calculation formula is as follows: In the formula, The key matrix; The learnable weight matrix corresponding to the key mapping; two mapping matrices and Project each row of the cross-modal concatenated feature matrix onto the same dimensional space, with dimensions of [size missing]. That is, 64 rows and 32 columns, of which Calculate the dimension for the mapped attention; query matrix AND key matrix All dimensions are That is, a matrix with 2 rows and 32 columns. Based on the query matrix and the key matrix, the unnormalized raw attention score is calculated, and a scaling factor is introduced to prevent the value from being too large. The calculation formula is as follows: ; In the formula, The scaled attention score matrix has a size of two rows and two columns. The elements in the first row and first column represent the original attention score of the EEG modality to itself, the elements in the first row and second column represent the original attention score of the EEG modality to the eye-tracking modality, the elements in the second row and first column represent the original attention score of the eye-tracking modality to the EEG modality, and the elements in the second row and second column represent the original attention score of the eye-tracking modality to itself. Key matrix Transpose of; The length of the feature dimension after the query is mapped to the key; The scaling factor is used to suppress the vanishing gradient problem of the Softmax operation caused by excessive expansion of the dot product value as the dimension increases. Based on the scaled attention score matrix, a Softmax normalization operation is performed independently on each row to obtain the attention weight coefficient matrix. The calculation formula is as follows: ; In the formula, This is the attention weight coefficient matrix, with a size of two rows and two columns, and the sum of the two weight values ​​in each row is always equal to one; For the first The modality pair of the first Attention weight coefficients for each modality A value of 1 represents an EEG modality. A value of two represents the eye-tracking modality; Attention score matrix The Middle Line 1 The element values ​​of the column; For the natural constant An exponential function with base 0; the denominator is the first exponential function. The sum of the exponent values ​​of the two elements in the row.

[0035] Step 33: Based on the attention weight coefficients and the cross-modal concatenation feature matrix, perform a weighted summation calculation to obtain the weighted multimodal feature representation. Specifically, this includes: taking each row of the attention weight coefficient matrix as a set of weighting coefficients, performing a weighted summation operation on each row vector of the cross-modal concatenation feature matrix to generate a weighted fusion representation vector corresponding to the EEG modality and a weighted fusion representation vector corresponding to the eye-tracking modality; stacking the two weighted fusion representation vectors row by row to obtain the weighted multimodal feature representation matrix; and then performing a weighted summation operation on the attention weight coefficient matrix. The two weight values ​​in each row are used as weighting coefficients for the EEG modality features and the eye-tracking modality features, respectively. The two rows of the cross-modal concatenated feature matrix are then weighted and summed using the following formula: ; In the formula, For the first The weighted fusion representation vector corresponding to each modality is a row vector; The attention weight coefficient matrix output in step 32 is the first... Line 1 The weight value of the column indicates the weight of the first column. The modality pair of the first The degree of attention paid to modal feature information; The first feature matrix to be concatenated across modalities Row vectors The value corresponds to the original features of the EEG modality at one time. When the value is two, it corresponds to the original features of the eye-tracking modality. The weighted fusion representation vector of the EEG modality and the weighted fusion representation vector of the eye-tracking modality are stacked in row-wise into a matrix form. The stacking formula is as follows: In the formula, This is the weighted multimodal feature representation matrix, which consists of two rows. The first row... The second row represents the feature representation of EEG modalities after cross-modal attention weighting. This represents the feature representation of the eye-tracking modality after cross-modal attention weighting.

[0036] Step 34: Based on the weighted multimodal feature representation, perform linear transformation and normalization to obtain the multimodal fusion feature vector. Specifically, this includes: flattening all elements of the weighted multimodal feature representation matrix into a one-dimensional column vector in row-major order; mapping this flattened vector to the target fusion feature dimension through a learnable linear transformation fully connected layer to obtain the linear fusion feature vector; then performing layer normalization on this fusion feature vector to eliminate scale differences between components, finally outputting the multimodal fusion feature vector. The weighted multimodal feature representation matrix is ​​flattened row by row into a one-dimensional column vector in row-major order. The flattening formula is: ; In the formula, The flattened one-dimensional column vector has a total length that is twice the length of the given vector. ; The first row of the weighted multimodal feature representation matrix is... The element values ​​of the column; The second row of the weighted multimodal feature representation matrix The element values ​​of the column are used for feature mapping through a linear transformation fully connected layer based on the flattened vector. The mapping formula is as follows: In the formula, The fused feature vector after linear transformation; The learnable weight matrix of the linearly transformed fully connected layer; To linearly transform the learnable bias vector of the fully connected layer, based on the linearly fused feature vector, layer normalization is performed to eliminate scale differences between feature components and stabilize the numerical distribution. The formula is as follows: ; ; ; In the formula, The final output multimodal fusion feature vector serves as the input for the subsequent fatigue level classification step; Linear fusion feature vector The arithmetic mean of all component values; for The length of the dimension; for The variance of each component value; To prevent extremely small positive constants with a denominator of zero; A learnable vector of scaling parameters; A learnable translation parameter vector; operators This represents element-wise multiplication; for In the Component values ​​in each dimension.

[0037] In this embodiment of the invention, by simultaneously acquiring EEG and eye-tracking modal data and adaptively fusing self-attention across modalities, combined with Poincaré sphere hyperbolic space hierarchical classification, dual-branch lightweight feature extraction network edge-side inference, and fatigue level differentiated audio-visual vibration multi-channel redundant early warning, the defects of single modality being susceptible to interference and having a high false alarm rate, fixed weight fusion ignoring stage differences, serious cross-level misjudgment in Euclidean space, difficulty in deploying deep models at the edge, and single prompts being easily ignored are overcome. Thus, the effect of fatigue state classification, edge-side millisecond-level offline real-time monitoring, and reliable perception of multi-channel early warning signals is achieved under complex and high-risk working conditions, ensuring the personal safety of workers.

[0038] like Figure 4 As shown: This figure presents the distribution of attention weights automatically assigned to the EEG and eye-tracking modalities during the early fatigue stage in the form of a heatmap. The horizontal axis represents time, the vertical axis represents the two modalities involved in feature fusion, and the color bars on the right represent the numerical range of attention weights. The darker the color, the higher the contribution of that modality to the final fused features at the corresponding time.

[0039] In a preferred embodiment of the present invention, step 4 above may include: Step 41: Input the multimodal fusion feature vector into a pre-set fatigue state classification model to perform hyperbolic space embedding mapping processing, projecting the multimodal fusion feature vector into the hyperbolic geometric space of a Poincaré sphere to obtain a hyperbolic space feature representation. Specifically, this includes: inputting the multimodal fusion feature vector into a pre-set fatigue state classification model; calling the learnable linear projection layer embedded within the model, and using a pre-set weight matrix... With bias vector The fused feature vector is mapped to an intermediate transition vector in the tangent space. Based on this intermediate transition vector, an exponential mapping operation is performed. With the help of the model's pre-set curvature parameters, the tangent space vector is projected onto the interior of the hyperbolic geometric space of the Poincaré sphere, resulting in a hyperbolic space feature representation with a modulus strictly smaller than a unit length. Based on the multimodal fused feature vector, mapping is performed through a pre-set linear projection layer. The linear projection formula is: In the formula, It is the intermediate transition vector in the tangent space; The pre-defined learnable weight matrix for the linear projection layer is optimized, determined, and fixed during the model training phase. In the formula, It is the set of real numbers, meaning that all elements in the matrix are real numbers; The dimension of the Poincaré sphere tangent space is equal to the dimension of the hyperbolic space feature representation h. Multimodal fusion feature vector The dimensions of the matrix are such that each row of the matrix corresponds to one dimension of the tangent space, and each column corresponds to one dimension of the input fused features. This is a multimodal fusion feature vector; The pre-defined learnable bias vectors for the linear projection layer are optimized, determined, and fixed during model training. Based on the intermediate transition vectors, their Euclidean norm is calculated. The norm formula is: ; In the formula, The Euclidean norm 2 of the intermediate transition vector; The dimension of the intermediate transition vector; For the intermediate transition vector at the th The component values ​​in each dimension are projected onto the interior of the Poincaré sphere using an exponential mapping operation based on the intermediate transition vector and its norm. The mapping formula is as follows: ; In the formula, To represent the hyperbolic space characteristics inside the Poincaré sphere, ; The pre-defined curvature parameters of the Poincaré sphere are optimized, determined, and fixed during the model training phase. It is the hyperbolic tangent function; It is a very small positive constant to prevent division by zero; It is a vector of all zeros.

[0040] Step 42: Based on the hyperbolic space feature representation, perform spatial metric processing through a pre-set fatigue state classification model. Calculate the geodesic distance in hyperbolic geometric space between the hyperbolic space feature representation and the pre-set standard prototype points for four levels of fatigue: alertness, early fatigue, moderate fatigue, and severe fatigue. This yields a multi-dimensional geometric distance vector. Specifically, this includes: based on the hyperbolic space feature representation, sequentially reading the pre-set alertness level prototype points within the fatigue state classification model. Early fatigue level prototype point Moderate fatigue level prototype point and the prototype point of severe fatigue level The coordinate vector; calculate the hyperbolic geodesic distance between the hyperbolic space feature representation and the four preset prototype points one by one; concatenate the four geodesic distance values ​​in order from clear to severe to obtain a four-dimensional geometric distance vector, based on the hyperbolic space feature representation and the first... For each pre-set prototype point, the intermediate ratio parameter is calculated using the following formula: ; In the formula, For the first The intermediate ratio parameter corresponding to each level; The hyperbolic space characteristics inside the Poincaré sphere are represented; For the first The pre-defined standard prototype point coordinate vectors corresponding to each level are optimized, determined, and fixed during the model training phase. Corresponding to the states of wakefulness, early fatigue, moderate fatigue, and severe fatigue, respectively, the geodesic distance is calculated based on the intermediate ratio parameter using the following formula: ; In the formula, Representation of hyperbolic space features up to the th The geodesic distance between pre-set prototype points; The pre-defined curvature parameters of the Poincaré sphere are optimized, determined, and fixed during the model training phase. The inverse hyperbolic cosine function is used to concatenate the distances of the four geodesic lines into a geometric distance vector, as shown in the formula: ; In the formula, It is a four-dimensional geometric distance vector; to These represent the geodesic distances to the prototype points for awakening, early fatigue, moderate fatigue, and severe fatigue, respectively.

[0041] Step 43: Based on the multidimensional geometric distance vector, perform nonlinear transformation processing through a preset fatigue state classification model, and perform nonlinear probability normalization calculation based on the reciprocal of the distance to obtain the fatigue state confidence of the worker under four levels: alertness, early fatigue, moderate fatigue, and severe fatigue. Specifically, this includes: based on the four-dimensional geometric distance vector, calling a preset smoothing constant, and performing a reciprocal operation on each distance component to obtain the original correlation strength score; based on all four original correlation strength scores, calling a preset temperature adjustment factor... Perform a temperature-controlled flexible maximum normalization operation to convert the geometric distance information into a probability distribution, obtaining the fatigue state confidence vector of the worker under four fatigue levels. Based on the geometric distance vector and a preset smoothing constant, calculate the original correlation strength score using the following formula: ; In the formula, For the first The original association strength score corresponding to each level; For the first Geodesic distances at various levels; A pre-defined positive smoothing constant is determined and fixed during the model training phase to ensure that the denominator is always greater than zero. Normalization is performed based on the original correlation strength score and a pre-defined temperature adjustment factor, using the following formula: ; In the formula, For the first The confidence level of fatigue state is 100%. ; The preset temperature adjustment factor is determined and fixed during the model training phase, and its value is greater than zero. The four confidence scores are concatenated into a confidence vector, as shown in the formula: ; In the formula, This represents the confidence vector for the fatigue state. to The confidence levels are respectively for alertness, early fatigue, moderate fatigue, and severe fatigue.

[0042] Step 44: Based on the fatigue state confidence score, extreme value determination processing is performed using a pre-set fatigue state classification model to determine the target fatigue level with the highest fatigue state confidence score, thus obtaining the fatigue state classification result of the worker's current fatigue level. Specifically, this includes: based on the fatigue state confidence score vector, traversing all four components and performing element-by-element comparison operations to determine the index position corresponding to the component with the highest confidence score; according to the index-level label mapping relationship pre-set in the fatigue state classification model, converting the index into the corresponding fatigue level label text to obtain the final classification result of the worker's current fatigue state; and determining the final level position through maximum value index calculation, using the following formula: ; In the formula, The rank index corresponding to the highest confidence level; This is an index operation to retrieve the maximum value; For the first The fatigue level label is output based on the confidence level values ​​of each level, according to the maximum confidence index and a preset mapping rule. The formula is as follows: ; In the formula, This is the final output of the fatigue state classification result; to These are four pre-set level labels in the model: alertness, early fatigue, moderate fatigue, and severe fatigue.

[0043] In this embodiment of the invention, by simultaneously acquiring EEG and eye-tracking modal data and adaptively fusing self-attention across modalities, combined with Poincaré sphere hyperbolic space hierarchical classification, dual-branch lightweight feature extraction network edge-side inference, and fatigue level-differentiated audio-visual vibration multi-channel redundant early warning, the defects of single modality being susceptible to interference and having a high false alarm rate, fixed weight fusion ignoring stage differences, serious cross-level misjudgment in Euclidean space, difficulty in deploying deep models at the edge, and single prompts being easily ignored are overcome. Thus, the invention achieves the effect of accurate fatigue state classification, millisecond-level offline real-time monitoring at the edge, and reliable perception of multi-channel early warning signals under complex and high-risk working conditions, ensuring the personal safety of workers.

[0044] In a preferred embodiment of the present invention, step 5 above may include: Step 51: Based on the fatigue state classification results, perform fatigue level comparison processing. When the fatigue state classification result is determined to be early fatigue, moderate fatigue, or severe fatigue, trigger the edge computing early warning logic built into the smart safety helmet. Specifically, this includes: comparing the fatigue state classification result with the four preset fatigue level labels one by one; when the comparison result determines that the current fatigue level belongs to one of the three categories—early fatigue, moderate fatigue, or severe fatigue—a warning trigger enable signal is generated, activating the edge computing early warning logic module built into the smart safety helmet; when the comparison result determines that the current fatigue level belongs to the conscious level, no warning trigger enable signal is generated, the edge computing early warning logic module remains silent, and performs a level label equivalence comparison based on the fatigue state classification result and outputs a warning trigger enable signal. The trigger determination formula is: ; In the formula, The warning trigger enable signal has a value of 1 indicating that the edge computing warning logic module is activated, and a value of 0 indicates that the module remains silent. This is the final output of the fatigue state classification result; A level of alertness is indicated by a label. Early fatigue level label; Labeled as moderate fatigue level; Labels for severe fatigue levels; operators The determination that an element belongs to a set is based on the alert trigger enable signal, which triggers a state switch in the edge computing alert logic module. The state switch formula is as follows: ; In the formula, This indicates the working status of the edge computing early warning logic module. This indicates that the module has entered the active working state and has begun to execute the early warning process; This indicates that the module remains idle and silent, and does not perform any warning operations.

[0045] Step 52: Based on the triggered edge computing warning logic, obtain the current timestamp and fatigue state classification result, perform data encapsulation processing to obtain a fatigue warning instruction containing fatigue level and timestamp. Specifically, this includes: reading the current system timestamp from the real-time clock module built into the smart helmet when the warning trigger enable signal is active; performing structured data encapsulation processing on the timestamp and fatigue state classification result, combining the fatigue level label and timestamp into a structured data unit to obtain a fatigue warning instruction containing fatigue level parameters and timestamp parameters; and reading the current system timestamp from the real-time clock module after activation by the warning trigger enable signal. The timestamp reading formula is: In the formula, The current system timestamp is an integer value in milliseconds, representing the number of milliseconds that have elapsed since the preset base time. The current count value output by the real-time clock module built into the smart safety helmet is combined with the current timestamp and fatigue state classification result through a data structuring encapsulation operation to form a fatigue warning instruction. The encapsulation formula is as follows: In the formula, The fatigue warning instruction is a structured data unit containing two fields; The results of fatigue state classification are used as fatigue level parameter fields; The current system timestamp is used as a timestamp parameter field.

[0046] Step 53: Parse and extract the fatigue level parameter from the fatigue warning command. Match the corresponding alarm intensity level according to the fatigue level parameter to obtain the alarm control parameter. According to the alarm control parameter, drive the audible and visual alarm component and vibration feedback component of the smart safety helmet to execute the audible and visual prompts and vibration feedback of the corresponding intensity to obtain a real-time fatigue warning signal. Specifically, this includes: parsing and extracting the fatigue level parameter field from the fatigue warning command to obtain the current fatigue level value to be responded to; inputting the fatigue level value into a preset level intensity mapping function to match and output the corresponding alarm intensity level parameter; calculating the control drive parameters of the audible and visual alarm component and vibration feedback component respectively according to the alarm intensity level parameter; and driving the LED array, buzzer module, and vibration motor module of the smart safety helmet to synchronously execute the audible and visual prompts and vibration feedback of the corresponding intensity according to the control drive parameters to obtain a fatigue warning signal transmitted to the worker in real time. The fatigue level parameter field is parsed and extracted according to the fatigue warning command. The parsing formula is: ; In the formula, The fatigue level value is extracted from the early warning command. Parsing operations to extract the fatigue level field from structured data units; For fatigue warning commands, based on the fatigue level value, the corresponding alarm intensity level is matched using a preset intensity mapping table. The mapping formula is as follows: ; In the formula, This is an alarm intensity level parameter used to uniformly control the output intensity of the three alarm channels: sound, light, and vibration. This is a preset level intensity mapping function; The minimum value is selected as the alarm intensity value corresponding to the early fatigue level; The alarm intensity value corresponding to the moderate fatigue level is set to the middle value; The alarm intensity value corresponding to the severe fatigue level is taken as the maximum value; the three intensity values ​​satisfy... Based on the strictly increasing relationship of the alarm intensity level parameters, the control drive parameters for each alarm channel are calculated separately. The formula for calculating the control parameters is as follows: ; ; ; In the formula, The pulse width modulation duty cycle parameter for driving the LED array; The audio frequency parameters for driving the buzzer module, in Hertz; The power percentage parameter for driving the vibration motor; This is the preset maximum alarm intensity reference value; This is the maximum duty cycle limit for the LED array; This is the highest operating frequency of the buzzer module; This is the minimum operating frequency of the buzzer module; Based on the maximum output power ratio of the vibration motor and the control drive parameters of each channel, the three types of alarm actuators are synchronously driven to generate early warning signals. The output signal formula is: ; In the formula, The final output real-time fatigue warning signal is a composite warning signal that includes three modes: light, sound, and vibration. In order to use duty cycle The driving light alarm signal controls the LED array to flash at a preset frequency; For frequency The driving sound alarm signal controls the buzzer to emit a buzzing sound of the corresponding frequency; Based on power ratio The vibration alarm signal driven by the motor controls the vibration motor to generate intermittent vibration feedback of corresponding intensity.

[0047] In this embodiment of the invention, by simultaneously acquiring EEG and eye-tracking modal data and adaptively fusing self-attention across modalities, combined with Poincaré sphere hyperbolic space hierarchical classification, dual-branch lightweight feature extraction network edge-side inference, and fatigue level-differentiated audio-visual vibration multi-channel redundant early warning, the defects of single modality being susceptible to interference and having a high false alarm rate, fixed weight fusion ignoring stage differences, serious cross-level misjudgment in Euclidean space, difficulty in deploying deep models at the edge, and single prompts being easily ignored are overcome. Thus, the invention achieves the effect of accurate fatigue state classification, millisecond-level offline real-time monitoring at the edge, and reliable perception of multi-channel early warning signals under complex and high-risk working conditions, ensuring the personal safety of workers.

[0048] like Figure 2 As shown, embodiments of the present invention also provide a fatigue monitoring system for workers in high-risk industries based on EEG-eye fusion, comprising: The acquisition module is used to collect the raw physiological signals of workers in real time through the smart safety helmet. Based on the raw physiological signals, bandpass filtering and baseline drift removal processing are performed to obtain multimodal physiological monitoring data, which includes denoised EEG signals and primary eye movement feature data including blink frequency and gaze duration. The parsing module is used to extract blink interval temporal features and gaze point shift features from the primary eye movement feature data in the multimodal physiological monitoring data, and to extract the power spectral density of the EEG feature band within a preset time window; the power spectral density of the EEG feature band, blink interval temporal features and gaze point shift features are input into a preset lightweight feature extraction network for parsing to obtain EEG feature vectors and eye movement behavior feature vectors; The splicing module is used to perform cross-modal feature splicing and attention weight allocation calculation based on EEG feature vectors and eye-tracking behavior feature vectors to obtain multimodal fusion feature vectors; The calculation module is used to input the multimodal fusion feature vector into the preset fatigue state classification model for probability mapping calculation, obtain the fatigue state confidence score, and obtain the fatigue state classification result of the current fatigue level of the worker based on the maximum value of the fatigue state confidence score. The driving module is used to trigger the edge computing early warning logic built into the smart safety helmet when the fatigue state is determined to be early fatigue, moderate fatigue or severe fatigue, and obtain a fatigue warning command containing fatigue level and timestamp; according to the fatigue warning command, it drives the sound and light alarm component and vibration feedback component of the smart safety helmet to obtain a real-time fatigue warning signal.

[0049] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0050] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0051] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0052] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for monitoring fatigue status of workers in high-risk industries based on EEG-eye tracking fusion, characterized in that, The method includes: Step 1: Collect raw physiological signals of workers in real time using a smart safety helmet. Based on the raw physiological signals, perform bandpass filtering and baseline drift removal processing to obtain multimodal physiological monitoring data, which includes denoised EEG signals and primary eye movement feature data including blink frequency and gaze duration. Step 2: Based on the primary eye movement feature data in the multimodal physiological monitoring data, extract the blink interval temporal feature and gaze point shift feature, and extract the power spectral density of the EEG feature frequency band within a preset time window; input the power spectral density of the EEG feature frequency band, the blink interval temporal feature and the gaze point shift feature into a preset lightweight feature extraction network for parsing to obtain the EEG feature vector and the eye movement behavior feature vector; Step 3: Based on the EEG feature vector and eye movement behavior feature vector, perform cross-modal feature splicing and attention weight allocation calculation to obtain the multimodal fusion feature vector; Step 4: Input the multimodal fusion feature vector into the preset fatigue state classification model for probability mapping calculation to obtain the fatigue state confidence. Based on the maximum value in the fatigue state confidence, obtain the fatigue state classification result of the worker's current fatigue level. Step 5: When the fatigue state is determined to be early fatigue, moderate fatigue or severe fatigue, the edge computing early warning logic built into the smart safety helmet is triggered to obtain a fatigue early warning command containing the fatigue level and timestamp; according to the fatigue early warning command, the sound and light alarm component and vibration feedback component of the smart safety helmet are driven to obtain a real-time fatigue early warning signal.

2. The method for monitoring fatigue status of workers in high-risk industries based on EEG-eye tracking fusion according to claim 1, characterized in that, Step 1: Real-time acquisition of raw physiological signals from workers using a smart safety helmet. Based on these raw physiological signals, bandpass filtering and baseline drift removal are performed to obtain multimodal physiological monitoring data, including denoised EEG signals and primary eye movement feature data containing blink frequency and gaze duration. The system uses an integrated EEG sensor and eye-tracking sensor in a smart safety helmet to simultaneously collect raw EEG waveform data and raw eye video stream data from workers. Based on the raw EEG waveform data, bandpass filtering is performed to filter out power frequency interference and electromyography artifacts, and low-frequency noise is removed by baseline drift correction to obtain a denoised EEG signal. Based on the original eye video stream data, the key points of the human eye are located and the pupil center is tracked. Based on the temporal sequence of changes in the position of the pupil center, the number of blinks and the duration of gaze are calculated per unit time to obtain primary eye movement feature data containing blink frequency and gaze duration. Based on the denoised EEG signals and primary eye movement feature data, timestamp alignment processing was performed to obtain multimodal physiological monitoring data.

3. The method for monitoring fatigue status of workers in high-risk industries based on EEG-eye fusion according to claim 2, characterized in that, Step 2: Based on the primary eye movement feature data from the multimodal physiological monitoring data, extract the blink interval temporal features and gaze point shift features, and extract the power spectral density of the EEG feature frequency bands within a preset time window, including: Based on the blink frequency in the primary eye movement feature data, the time interval sequence between two adjacent blink actions is calculated to obtain the blink interval temporal feature. Based on the gaze duration in the primary eye movement feature data, the spatial coordinate offset of the gaze focus in the preset work area is calculated to obtain the gaze point offset feature; based on the denoised EEG signal in the multimodal physiological monitoring data, the preset time window data corresponding to the current work cycle is extracted. Based on the data within a preset time window, a fast Fourier transform is performed to extract energy distribution parameters including the alpha, beta, and theta bands, thus obtaining the power spectral density of the EEG characteristic frequency bands.

4. The method for monitoring fatigue status of workers in high-risk industries based on EEG-eye fusion according to claim 3, characterized in that, The power spectral density of the EEG characteristic frequency bands, the temporal features of blink intervals, and the gaze point shift features are input into a pre-built lightweight feature extraction network for analysis, resulting in EEG feature vectors and eye movement behavior feature vectors, including: Based on the blink interval temporal features and gaze point offset features, feature dimension splicing processing is performed to obtain the eye movement feature splicing sequence; The eye-tracking feature concatenation sequence is input into the temporal convolution branch of a pre-built lightweight feature extraction network, and one-dimensional convolution and pooling operations are performed to obtain an eye-tracking behavior feature vector that represents the temporal evolution of eye-tracking behavior. The power spectral density of the EEG feature frequency band is input into the frequency domain mapping branch of the lightweight feature extraction network to perform preliminary feature mapping processing and obtain the frequency domain mapping feature matrix. Based on the frequency domain mapping feature matrix, dimensionality reduction and nonlinear activation processing are performed using a fully connected layer to extract the nonlinear mapping relationship of the EEG frequency bands, thus obtaining the EEG feature vector.

5. The method for monitoring fatigue status of workers in high-risk industries based on EEG-eye fusion according to claim 4, characterized in that, Step 3: Based on the EEG feature vector and eye-tracking behavior feature vector, perform cross-modal feature concatenation and attention weight allocation calculation to obtain a multimodal fusion feature vector, including: Based on the EEG feature vector and the eye movement behavior feature vector, feature dimension splicing processing is performed to obtain a cross-modal spliced ​​feature matrix; The cross-modal splicing feature matrix is ​​input into a pre-set self-attention mechanism network for weight calculation to obtain attention weight coefficients that characterize the importance of different physiological signals; Based on the attention weight coefficients and the cross-modal concatenated feature matrix, a weighted summation calculation is performed to obtain the weighted multimodal feature representation; Based on the weighted multimodal feature representation, a linear transformation and normalization process are performed to obtain the multimodal fusion feature vector.

6. The method for monitoring fatigue status of workers in high-risk industries based on EEG-eye fusion according to claim 5, characterized in that, Step 4: Input the multimodal fusion feature vector into the preset fatigue state classification model for probability mapping calculation to obtain the fatigue state confidence score. Based on the maximum value in the fatigue state confidence score, obtain the fatigue state classification result of the worker's current fatigue level, including: The multimodal fusion feature vector is input into a pre-set fatigue state classification model to perform hyperbolic space embedding mapping processing. The multimodal fusion feature vector is then projected into the hyperbolic geometric space of the Poincaré sphere to obtain the hyperbolic space feature representation. Based on the hyperbolic space feature representation, spatial metric processing is performed through a pre-set fatigue state classification model to calculate the geodesic distance between the hyperbolic space feature representation and the pre-set standard prototype points of four levels of wakefulness, early fatigue, moderate fatigue and severe fatigue in the hyperbolic geometric space, thus obtaining a multi-dimensional geometric distance vector. Based on the multidimensional geometric distance vector, a nonlinear transformation process is performed through a pre-set fatigue state classification model, and a nonlinear probability normalization calculation based on the inverse of the distance is performed to obtain the fatigue state confidence of the worker under four levels: awake, early fatigue, moderate fatigue and severe fatigue. Based on the fatigue state confidence level, extreme value determination processing is performed through a pre-set fatigue state classification model to determine the target fatigue level with the largest value in the fatigue state confidence level, thereby obtaining the fatigue state classification result of the worker's current fatigue level.

7. The method for monitoring fatigue status of workers in high-risk industries based on EEG-eye fusion according to claim 6, characterized in that, Step 5: When the fatigue state is determined to be early fatigue, moderate fatigue or severe fatigue, the edge computing warning logic built into the smart safety helmet is triggered to obtain a fatigue warning instruction containing fatigue level and timestamp. Based on the fatigue warning command, the sound and light alarm component and vibration feedback component of the smart safety helmet are activated to obtain real-time fatigue warning signals, including: Based on the fatigue state classification results, fatigue level comparison processing is performed. When the fatigue state classification result is determined to be early fatigue, moderate fatigue or severe fatigue, the edge computing early warning logic built into the smart safety helmet is triggered. Based on the triggered edge computing early warning logic, the current timestamp and fatigue state classification result are obtained, and data encapsulation processing is performed to obtain a fatigue early warning instruction containing fatigue level and timestamp. The fatigue level parameters in the fatigue warning command are analyzed and extracted. The corresponding alarm intensity level is matched according to the fatigue level parameters to obtain the alarm control parameters. According to the alarm control parameters, the sound and light alarm component and vibration feedback component of the smart safety helmet are driven to execute the sound and light prompts and vibration feedback of the corresponding intensity to obtain the real-time fatigue warning signal.

8. A fatigue monitoring system for workers in high-risk industries based on EEG-eye tracking fusion, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to collect the raw physiological signals of workers in real time through the smart safety helmet. Based on the raw physiological signals, bandpass filtering and baseline drift removal processing are performed to obtain multimodal physiological monitoring data, which includes denoised EEG signals and primary eye movement feature data including blink frequency and gaze duration. The parsing module is used to extract blink interval temporal features and gaze point shift features from the primary eye movement feature data in the multimodal physiological monitoring data, and to extract the power spectral density of the EEG feature band within a preset time window; the power spectral density of the EEG feature band, blink interval temporal features and gaze point shift features are input into a preset lightweight feature extraction network for parsing to obtain EEG feature vectors and eye movement behavior feature vectors; The splicing module is used to perform cross-modal feature splicing and attention weight allocation calculation based on EEG feature vectors and eye-tracking behavior feature vectors to obtain multimodal fusion feature vectors; The calculation module is used to input the multimodal fusion feature vector into the preset fatigue state classification model for probability mapping calculation, obtain the fatigue state confidence score, and obtain the fatigue state classification result of the current fatigue level of the worker based on the maximum value of the fatigue state confidence score. The driving module is used to trigger the edge computing early warning logic built into the smart safety helmet when the fatigue state is determined to be early fatigue, moderate fatigue or severe fatigue, and obtain a fatigue warning command containing fatigue level and timestamp; according to the fatigue warning command, it drives the sound and light alarm component and vibration feedback component of the smart safety helmet to obtain a real-time fatigue warning signal.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.