A fatigue recognition method and system based on adaptive spatio-temporal graph convolution and multi-modal hierarchical fusion

CN122508481APending Publication Date: 2026-08-04SHANDONG BIG DATA MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG BIG DATA MEDICAL TECH CO LTD
Filing Date
2026-05-12
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0008]为了解决现有技术在复杂工业场景下进行人员疲劳识别时,存在的特征提取维度单一、长时序依赖特征易丢失,以及异构多模态数据融合不充分的问题,本发明提供一种基于自适应时空图卷积与多模态层级融合的疲劳识别方法及系统

Benefits of technology

[0063] By overcoming the limitations of physical skeleton topology, this invention significantly improves the sensitivity to latent fatigue. It innovatively introduces a data-driven, learnable adaptive adjacency matrix into the Adaptive Spatiotemporal Graph Convolutional Network (AGCN). This mechanism allows the model to move beyond explicit physical connections in human anatomy, autonomously mining and quantifying potential cross-modal relationships in high-dimensional space, such as "hand tremors and heart rate variability" or "sluggish movements and abnormalities in specific physiological indicators." This breakthrough effectively addresses the technical pain point of feature fragmentation in traditional methods, achieving highly sensitive capture of early, subtle latent fatigue features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122508481A_ABST
    Figure CN122508481A_ABST
Patent Text Reader

Abstract

This invention relates to the field of multimodal fatigue recognition technology, and in particular to a fatigue recognition method and system based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion. The method includes: deep spatiotemporal topology extraction of preprocessed data based on an improved adaptive graph convolutional network (AGCN); global temporal context modeling of the extracted features based on relative position encoding; multimodal hierarchical dynamic fusion of the modeling results based on nonlinear confidence penalties; adaptive dynamic threshold fatigue determination of the fusion results based on KL divergence and information theory; and closed-loop optimization of the determination results using a joint objective function based on multiple physical manifold constraints. This invention innovatively introduces a data-driven, learnable adaptive adjacency matrix into the adaptive spatiotemporal graph convolutional network (AGCN).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal fatigue recognition technology, and in particular to a fatigue recognition method and system based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion. Background Technology

[0002] With the deepening of Industry 4.0 and intelligent manufacturing, the automation level of production lines has significantly improved, but core and key positions still require a high degree of staff involvement. In high-intensity, long-cycle industrial work environments, the physiological and psychological fatigue of personnel has become a major factor inducing safety accidents, leading to decreased production efficiency and abnormal fluctuations in product quality. Therefore, achieving non-contact, high-precision, and real-time online monitoring of personnel fatigue in complex industrial scenarios has an urgent practical need and extremely high industrial application value.

[0003] Currently, existing fatigue detection technologies are mainly divided into two categories: contact and non-contact. Contact methods primarily rely on workers wearing specialized equipment to collect physiological signals such as electroencephalograms (EEGs) and electrocardiograms (ECGs). While these methods offer high accuracy, the equipment is expensive, highly invasive, and easily interferes with workers' normal operations. Non-contact methods are mostly based on computer vision technology, deploying cameras to capture facial features such as the frequency of yawning and eye closure. However, in complex industrial environments, existing machine learning-based fatigue detection methods still exhibit the following significant technical bottlenecks and shortcomings:

[0004] First, spatial feature mining relies on a single dimension, making it difficult to effectively identify latent fatigue. Existing conventional recognition algorithms (such as Support Vector Machines (SVM) or basic Convolutional Neural Networks (CNN)) typically treat image pixels or sequence data collected by sensors as isolated feature points. This approach severely severs the natural spatial topological relationship between human skeletal movement features and internal physiological signals (for example, slow limb movements are often spatially linked to abnormal rhythms of local electromyographic signals), resulting in the model lacking sufficient perception and recognition ability for "latent fatigue" that lacks obvious external manifestations.

[0005] Secondly, long-term time-dependent features are easily lost, and there is a lack of ability to track global state evolution. Human fatigue is essentially a gradual evolutionary process that accumulates over a long period. Existing time-series processing models based on recurrent neural networks are prone to gradient vanishing or memory forgetting when dealing with monitoring sequences over extremely long periods (such as several hours of continuous work). This makes it difficult for existing models to effectively capture the subtle evolutionary patterns of fatigue over long periods, and consequently, to accurately distinguish between brief pauses in normal human actions and true, deep-seated fatigue.

[0006] Third, the fusion layer of heterogeneous multimodal data is shallow, resulting in poor robustness against environmental interference. Current technologies for processing visual behavior and physiological data often employ simple feature-level splicing methods, failing to overcome the physical heterogeneity between different modalities at the deep semantic level. In real-world industrial scenarios, when a modality's data is severely affected by ambient noise (such as sudden changes in workshop lighting causing visual feature failure, or motion artifacts in physiological signals caused by personnel movements), the recognition accuracy and robustness of the entire model will significantly decrease due to the lack of adaptive modal complementarity and dynamic fault-tolerance mechanisms.

[0007] In summary, to overcome the aforementioned deficiencies in existing technologies, the industry urgently needs a fatigue recognition method and system that can deeply mine spatiotemporal topological linkage features, possess long-range temporal memory capabilities, and perform adaptive anti-interference multimodal hierarchical fusion. Summary of the Invention

[0008] To address the shortcomings of existing technologies in fatigue identification in complex industrial scenarios, such as single-dimensional feature extraction, easy loss of long-term dependent features, and insufficient fusion of heterogeneous multimodal data, this invention provides a fatigue identification method and system based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion. By constructing a heterogeneous spatiotemporal topology graph that integrates the human skeletal structure and multi-source physiological sensor nodes, and combining a deep adaptive graph convolutional neural network with a Transformer self-attention mechanism incorporating relative position encoding, the method deeply explores the potential high-dimensional spatial linkage features and long-term temporal cumulative evolution patterns between physiological micro-signals and external macro-behavioral actions. A hierarchical dynamic mechanism is employed to deeply fuse heterogeneous data, and an information-theoretic-based adaptive dynamic threshold strategy is used to calibrate the decision boundary in real time.

[0009] Firstly, the present invention provides a fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion, which adopts the following technical solution:

[0010] A fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion includes:

[0011] Acquire multi-source heterogeneous worker monitoring data in a pre-defined industrial scenario;

[0012] Data preprocessing is performed based on the acquired monitoring data;

[0013] Deep spatiotemporal topology extraction of preprocessed data is performed based on the improved adaptive graph convolutional network AGCN.

[0014] Global temporal context modeling is performed on the extracted features based on relative position encoding;

[0015] Multimodal hierarchical dynamic fusion of modeling results is performed based on nonlinear confidence penalty;

[0016] Adaptive dynamic threshold fatigue determination of fusion results is performed based on KL divergence and information theory;

[0017] The judgment results are subjected to closed-loop optimization of the joint objective function based on multiple physical manifold constraints.

[0018] Furthermore, the data preprocessing based on the acquired monitoring data includes high-dimensional space data cleaning and domain adaptation using the maximum mean difference algorithm, wherein the regenerating kernel Hilbert space mapping method is used, employing the Gaussian radial basis kernel function as the nonlinear mapping operator. The original low-dimensional and nonlinear physiological-behavioral mixed data is projected onto a high-dimensional feature space. The distribution distance between the source domain and the target domain is calculated in the high-dimensional space, using the following formula: Then, latent fatigue samples are mined. By constructing a variational autoencoder deep neural network, the first data filtered by MMD is input into the VAE network. The encoder compresses the high-dimensional input into latent variables that follow a normal distribution through a multi-layer fully connected network. The decoder then attempts to reconstruct the original input data from this latent variable, using the lower bound of evidence as the core loss criterion, expressed as:

[0019] .

[0020] Furthermore, the improved adaptive graph convolutional network (AGCN) is used to perform deep spatiotemporal topology extraction on the preprocessed data, including heterogeneous spatiotemporal topology graph construction and latent fatigue feature extraction. First, a graph structure is defined, equidistantly mapping key points of the human skeleton and multidimensional physiological monitoring channels in the image data to a set of nodes in the graph structure. ,in The total number of nodes is defined; physical edges are defined based on human physiological connections, and temporal edges are defined based on temporal continuity. Simultaneously, temporal edges are established by connecting different frame states of the same node along the time axis, ultimately instantiating a high-dimensional adjacency matrix. With graph signal tensor To refine the extraction of limb relaxation or stiffness features specific to fatigue states, a set of spatial partitioning strategies for graph convolution kernels was introduced. Based on the geodesic distances between nodes and the direction of gravity, the neighborhood of the central node is deconstructed into three orthogonal subspaces, each of which independently corresponds to a normalized adjacency matrix. The first data is input into a multi-layer graph convolutional neural network, its core... layer to the first The feature aggregation derivation of the layer is as follows:

[0021] ,in It is a degree matrix; The learnable weight matrix for feature channel mapping, matrix It is an unconstrained, zero-initialized, data-driven, learnable, adaptive adjacency matrix.

[0022] Furthermore, the global temporal context modeling of the extracted features based on relative position encoding includes addressing the memory forgetting defect of the Long Short-Term Memory (LSTM) network when faced with consecutive working data frames. First, the key feature sequences extracted by the preceding spatiotemporal graph convolutional network AGCN are introduced into a Transformer architecture with a pure attention mechanism for full temporal domain association. Here, sequence mapping and linear projection embedding are used. The tensor of the key feature sequence obtained after spatial feature extraction is... ,in The time step length, For the feature channel dimension, the sequence is mapped to the query matrix Query, key matrix Key, and value matrix Value required by the multi-head self-attention mechanism, for the first... The linear projection formula for an attention head is defined as:

[0023] ,

[0024] in, This is the corresponding learnable weight matrix. For a single attention head, the feature dimension is defined, and then the absolute position sine encoding in the Transformer is discarded. Instead, the time interval information is explicitly injected into the computation graph in the form of a relative position bias term, defining two time steps. and The relative distance index between them is ,in To maximize the truncation of relative distances, construct a dictionary of relative position embeddings. This generates a relative position offset matrix. Its elements That is, corresponding In the dictionary The scalar value in the middle, after introducing relative bias, the first Attention is focused on the time step time step Unnormalized attention score The calculation is as follows:

[0025] ,

[0026] The matrix form of this formula is the standard form of scaled dot product attention:

[0027] Among them, utilizing Tensors calculate the absolute semantic similarity between any two time steps within a sequence, while Then, semantically independent learnable bias compensation is performed based on the relative distance between the two time steps.

[0028] Furthermore, the global temporal context modeling of the extracted features based on relative position encoding also includes, in order to capture fatigue evolution features in different subspaces, [the following is missing from the original text]. The outputs of each attention head are concatenated and then linearly fused, as shown below:

[0029] ,

[0030] in To fuse the weights, residual connections and layer normalization are then used to prevent deep network degradation before feeding them into a position-fed forward neural network (FFN). To improve the non-linear smoothness of feature representation, the GELU activation function is used instead of the traditional ReLU. The mathematical derivation and network structure definition are as follows:

[0031] ,

[0032] ,

[0033] The GELU activation function introduces random regularization through Gaussian error linear units:

[0034] ,

[0035] Finally, the output matrix is ​​subjected to residual summation and layer normalization again:

[0036] ,

[0037] This output This is the second data that has been fully integrated with global context memory.

[0038] Furthermore, the multimodal hierarchical dynamic fusion of the modeling results based on nonlinear confidence penalty includes addressing high-frequency baseline drift and power frequency interference in physiological signals collected from industrial sites. First, wavelet packet decomposition is performed on the non-stationary physiological signals, assuming the discrete physiological signal sequence is... Using Daubechies wavelet basis functions with high-order vanishing moments, it is decomposed to the th order. Layers, generation The frequency bands are divided into orthogonal time-frequency sub-bands, and then the relative energy distribution of each sub-band is calculated: , ,in, For the first Layer Wavelet packet reconstruction coefficients for each frequency band The total energy of the frequency band, energy characteristic sequence Subsequently, it is mapped to a high-dimensional physiological feature vector that is perfectly aligned with the behavioral feature dimension through a multilayer perceptron (MLP). Then, to break down semantic barriers between modalities, a hierarchical bilinear cross-attention network is constructed, employing active interaction logic. Let the high-level behavioral features output after global temporal modeling by the Transformer be... Using behavioral features as the primary query vector and mapped physiological frequency band features as the key and value vectors, linear projections are performed respectively:

[0039] ,

[0040] in, The projected matrix is ​​then used for learning, followed by scaling dot product cross-attention computation: This enables deep semantic forced alignment across physical attributes, outputting highly purified cross-modal interaction features. .

[0041] Furthermore, the multimodal hierarchical dynamic fusion of the modeling results based on nonlinear confidence penalty also includes obtaining the original behavioral features. Cross-modal interaction features Then, adaptive information purification is performed using a feedforward gating unit, and the information is then spliced ​​together. Combine the two and utilize the learnable weight matrix and bias Generate dynamically adjusted weights :

[0042] ,

[0043] In the ideal fusion state, the weighted representation of features is as follows: , To achieve element-wise multiplication, and considering that the gated adaptive unit is highly dependent on the purity of the input features, a novel nonlinear penalty calibration formula based on prediction confidence is embedded in the multimodal fusion attention module (MFA) to perform a secondary forced intervention on the feature weights: ,in, The probability confidence level of the lightweight network output in real time for a single-modality side branch. For hyperparameters, for A very small constant of the order of magnitude, used to avoid logarithmic operations in When encountering a singularity explosion, when confidence level At that time, penalty items ,but The system maintains a normal adaptive fusion mechanism; however, when the confidence level... When the threshold for safety is breached, the logarithmic decay formula will produce a steep nonlinear penalty.

[0044] Furthermore, the adaptive dynamic threshold fatigue determination of the fusion result based on KL divergence and information theory includes introducing a dynamic boundary defense based on information theory. First, the network output is probabilistically calibrated. To prevent the deep neural network from making overconfident predictions when facing unknown interference, the independent behavioral feature vectors and physiological feature vectors extracted in the previous stage are respectively input into their respective lightweight classification heads, and a temperature-smoothing hyperparameter is introduced. Mapping using the Softmax function:

[0045] ,

[0046] Introduce a temperature parameter greater than 1 The extreme probabilities are smoothed, and then the probability distribution of biological signals is quantified in real time using the Kullback-Leibler divergence from information theory. With behavioral probability distribution Differences in asymmetric information between them:

[0047] After obtaining the real-time KL divergence, the model is used to calibrate the initial static baseline threshold by plotting the receiver operating characteristic curve and calculating the maximum point of the Youden exponent on the discrete validation set. Subsequently Construct a nonlinear dynamic threshold generation model as the starting point:

[0048] Among them, the index risk penalty item : This is the penalty coefficient. When intermodal information conflict intensifies, this term becomes negative and exponentially and rapidly lowers the discrimination threshold. Confidence-weighted reward items The summation term utilizes the current real-time normalized confidence level of each modality. Multidimensional fatigue measurement function Dynamic rewards will be relaxed.

[0049] Furthermore, the closed-loop optimization of the joint objective function based on multiple physical manifold constraints on the judgment result includes, to ensure convergence stability and generalization ability in the high-dimensional parameter space, constructing a joint objective loss function with three perspective physical constraints during the gradient backpropagation training phase of the deep neural network, specifically including non-equilibrium adversarial classification constraints. This is expressed as a focus loss function based on the modulation factor. , The modulation factor is the confidence level for the model to predict the probability that the current sample is fatigued. Healthy manifold fidelity regression constraint , indicating the introduction of based on The regression constraint formula for the norm:

[0050] , This represents the baseline feature space extracted from the target worker under optimal health conditions. The current monitoring extracts characterization features; high-dimensional space inter-class exclusion contrast constraint. This is represented by the introduction of a high-dimensional triplet margin loss logic:

[0051] , will the current sample As an anchor point, by forcibly bringing the anchor point sample closer to the positive sample... Euclidean distance At the same time, it violently pushed away its negative sample. distance Furthermore, the distance difference between the two must be greater than the set safety margin boundary. This ensures that micro-sleep and focused attention on feature clusters are clearly distinct.

[0052] Secondly, a fatigue recognition system based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion includes:

[0053] The data acquisition module is configured to acquire multi-source heterogeneous worker monitoring data in a preset industrial scenario;

[0054] The preprocessing module is configured to perform data preprocessing based on the acquired monitoring data;

[0055] The topology module is configured to perform deep spatiotemporal topology extraction on preprocessed data based on the improved adaptive graph convolutional network AGCN.

[0056] The context module is configured to perform global temporal context modeling on the extracted features based on relative position encoding;

[0057] The fusion module is configured to perform multimodal hierarchical dynamic fusion of modeling results based on nonlinear confidence penalty;

[0058] The determination module is configured to perform adaptive dynamic threshold fatigue determination on the fusion result based on KL divergence and information theory.

[0059] The optimization module is configured to perform closed-loop optimization of the joint objective function based on multiple physical manifold constraints on the judgment result.

[0060] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion.

[0061] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion.

[0062] In summary, the present invention has the following beneficial technical effects:

[0063] By overcoming the limitations of physical skeleton topology, this invention significantly improves the sensitivity to latent fatigue. It innovatively introduces a data-driven, learnable adaptive adjacency matrix into the Adaptive Spatiotemporal Graph Convolutional Network (AGCN). This mechanism allows the model to move beyond explicit physical connections in human anatomy, autonomously mining and quantifying potential cross-modal relationships in high-dimensional space, such as "hand tremors and heart rate variability" or "sluggish movements and abnormalities in specific physiological indicators." This breakthrough effectively addresses the technical pain point of feature fragmentation in traditional methods, achieving highly sensitive capture of early, subtle latent fatigue features.

[0064] This invention overcomes the forgetting defects of long-sequence memory and achieves accurate source tracing of cumulative fatigue. It utilizes a Transformer architecture incorporating relative positional encoding for global temporal modeling, endowing the system with the ability to correlate global contexts with continuous work data spanning several hours. This mechanism fundamentally overcomes the gradient vanishing and memory forgetting defects that traditional recurrent neural networks (such as LSTM / RNN) are prone to when processing ultra-long sequences, enabling precise tracking and quantification of the gradual cumulative evolution of fatigue states over long time spans.

[0065] A dynamic confidence-based penalty mechanism is constructed to significantly enhance the robustness against interference under complex operating conditions. Addressing frequent environmental noise issues in industrial environments, such as sudden changes in lighting and electromagnetic interference, this invention embeds a dynamic gating penalty unit based on predicted confidence into the multimodal hierarchical interaction. When a single sensor fails or a particular mode is severely noisy, the system can quantify its uncertainty in real time and automatically suppress the fusion weight of that noisy mode, forcing the decision-making focus to smoothly shift to a high-reliability mode, ensuring high availability and uninterrupted operation of the system in extremely harsh environments.

[0066] This invention proposes a dynamic boundary strategy based on information theory, endowing the system with high individual adaptability and generalization capabilities. Addressing the challenge of significant individual differences in fatigue performance among different personnel, this invention utilizes Kullback-Leibler (KL) divergence to quantify the consistency between the distribution of biological signals and behavioral performance in real time, thereby driving the adaptive dynamic adjustment of the discrimination threshold. This mechanism enables the system to possess extremely strong generalization capabilities, eliminating the need for costly data collection and model retraining for specific target personnel. It can automatically calibrate the safety defense line based on the real-time degree of "physiological-behavioral" disconnect, demonstrating significant engineering application value and promising industrial application prospects. Attached Figure Description

[0067] Figure 1 This is a schematic diagram of a machine learning-based method for identifying fatigue in target personnel, according to Embodiment 1 of the present invention.

[0068] Figure 2 This is a flowchart of the method in Embodiment 1 of the present invention;

[0069] Figure 3 This is a schematic diagram of a terminal device according to Embodiment 1 of the present invention;

[0070] Figure 4 This is a schematic diagram of mouth behavior features and facial expression recognition in Embodiment 1 of the present invention. Detailed Implementation

[0071] The present invention will be further described in detail below with reference to the accompanying drawings.

[0072] Example 1

[0073] Reference Figure 1 The present invention will be further described in detail below with reference to specific embodiments. It should be noted that the following embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion provided by the present invention is configured to run in a high-performance industrial edge computing server with heterogeneous computing units. The hardware configuration includes at least one central processing unit with a main frequency of 3.0 GHz or higher and a tensor processing unit (TPU) with a video memory of not less than 12 GB, so as to meet the huge computing power requirements of deep learning models to perform real-time floating-point operations on massive multi-source heterogeneous data. The application scenario of this embodiment is set as the SMT assembly line operation environment in a precision electronics manufacturing workshop.

[0074] Step 1: Acquisition and strict spatiotemporal synchronization of multi-source heterogeneous monitoring data

[0075] 1. Hardware Deployment and Data Acquisition Architecture

[0076] During system initialization and online operation, it is essential to acquire multi-source heterogeneous monitoring data of the target workers in a pre-defined industrial scenario in real time. An industrial-grade RGB-D depth camera is deployed at the optimal viewing angle directly in front of the workstation, capturing subtle facial expressions and 3D skeletal movements of the workers at a constant frame rate of 30fps. Simultaneously, to overcome the limitations of a single visual sensor, medical-grade non-invasive multimodal physiological sensors are attached to the wrists and chests of the target workers. This sensor array acquires photoplethysmography (PPG), electrical skin activity (EDA), and weak electrocardiogram (ECG) data in real time at a high sampling rate of 256Hz.

[0077] 2. Millisecond-level data synchronization and underlying infrastructure

[0078] Strict temporal alignment is a prerequisite for multimodal fusion. All sensor nodes in the industrial field are connected to the edge gateway via Industrial Ethernet and Bluetooth Low Energy (BLE) protocol. The system's underlying kernel runs Precision Time Protocol (PTP), which uses a hardware-level timestamp synchronization algorithm to perform millisecond-level interpolation and alignment resampling on the visual frame stream and high-frequency physiological one-dimensional time series, thereby constructing a heterogeneous data stream containing time-synchronized body monitoring data, image data, and work behavior data.

[0079] Step 2: Data Preprocessing and Implicit Fatigue Anomaly Detection Based on Manifold Learning

[0080] Industrial environments are complex and changeable, often characterized by sudden electromagnetic interference, large-scale mechanical vibrations, and drastic fluctuations in workshop lighting. To filter out false anomalies caused by sensor jitter and accurately locate early "hidden fatigue" that is difficult to detect with the naked eye, this system implements a highly innovative joint anomaly screening mechanism before entering the deep network:

[0081] 1. Distribution Alignment and Cleaning Based on Maximum Mean Difference (MMD)

[0082] Traditional data cleaning methods based on Euclidean distance often fail when processing high-dimensional, nonlinear physiological data. This invention innovatively employs a manifold learning approach, utilizing the Maximum Mean Difference (MMD) algorithm to perform a high-dimensional distribution comparison between the currently collected data and reference baseline data from a historical database showing the worker in an absolutely healthy state. The system selects the Gaussian radial basis function (RBF) kernel function as the nonlinear mapping operator. The low-dimensional physiological-behavioral hybrid features are projected into the infinite-dimensional regenerative nuclear Hilbert space (RKHS). The distributional distance between the source domain (health baseline) and the target domain (current monitoring) is calculated as follows:

[0083] ,

[0084] This metric calculates not only the deviation from the first-order mean but also the changes in higher-order statistical moments. It applies if and only if the calculated... Only when the distance exceeds the preset statistical significance confidence interval (e.g., the 0.05 threshold calibrated offline) does the system determine that the overall manifold topology of the current data stream has undergone a fundamental shift, and strictly mark it as "non-steady-state data" that deviates from the healthy baseline.

[0085] 2. Unsupervised Reconstruction and Latent Localization Based on Variational Inference

[0086] For the selected non-stationary data, the system introduces a variational autoencoder (VAE) constructed from a deep multilayer perceptron (MLP) for unsupervised parsing. The encoder network compresses and maps the high-dimensional input into low-dimensional latent variables that follow a prior normal distribution. Subsequently, the decoder network attempts to perfectly reconstruct the original input data from this latent space. The core optimization objective of this process is to maximize the evidence lower bound (ELBO) of the log marginal likelihood, and the joint loss function is defined as:

[0087] ,

[0088] In the formula, the reconstruction expectation term forces the network to retain core features, while The (KL divergence) term acts as a regularizer, forcing a smooth distribution rule in the latent space.

[0089] Industrial Real-World Application: In actual SMT placement workshop scenarios, the system monitors the reconstruction residuals of multimodal components in real time. If the reconstruction error at the decoding end of worker limb movement data (such as robotic arm trajectories) is found to be extremely small, it indicates that their macroscopic behavior is still at a proficient and normal level, supported by "muscle memory." However, simultaneously, the reconstruction errors of their heart rate variability (HRV) or skin conductance response (GSR) spike abnormally. This extremely subtle "reconstruction conflict" clearly reveals that the worker's endocrine system and external behavior are severely disconnected, indicating a state of "hidden fatigue" under high-load strain. The system utilizes this mechanism to achieve "zero missed detection" of early dangerous states and merges and extracts them as primary data for downstream network processing.

[0090] Step 3: Deep Spatiotemporal Topology Extraction Based on Adaptive Graph Convolutional Network (AGCN)

[0091] After acquiring high-quality initial data, the system enters the core stage of spatial feature extraction. Traditional convolutional neural networks (CNNs) treat pixels as isolated grids in Euclidean space, completely severing the spatial topological linkages between various joints in the human body and between "action and physiology".

[0092] 1. Mathematical definition of heterogeneous spatiotemporal topological graphs

[0093] The system first performs graph structure abstraction. Key human skeletal points extracted by the depth camera (such as acromion, elbow joint, and wrist) are equidistantly mapped to multi-dimensional physiological monitoring channels (chest ECG electrodes, wrist pulse sensor) as a set of nodes in the graph. ,in This represents the total number of heterogeneous nodes. In constructing edge relationships, the system not only establishes "physical edges" along the spatial dimension based on the natural anatomical structure of the human skeleton, but also connects different frame states of the same node along the time axis to establish "temporal edges," ultimately instantiating a massive high-dimensional adjacency matrix. With graph signal tensor .

[0094] 2. Spatial Partitioning Strategies and Adaptive Aggregation Operations

[0095] To refine the extraction of limb relaxation or stiffness features specific to fatigue states, the system introduces a set of spatial partitioning strategies for graph convolution kernels. Based on the geodesic distances between nodes and the direction of gravity, the neighborhood of the central node is deconstructed into three orthogonal subspaces: centripetal, centrifugal, and static. Each subspace independently corresponds to a normalized adjacency matrix. .

[0096] The first data is input into a multi-layer graph convolutional neural network, the core of which is the first... layer to the first The feature aggregation derivation of the layer is as follows:

[0097] ,

[0098] Deep deconstruction of the formula: the first half It is responsible for aggregating explicit motion information (such as the associated motion when an arm is raised) based on a fixed physical topology. This is the degree matrix, used to normalize the message passing process and prevent gradient explosion; This is the learnable weight matrix for the feature channel mapping. The latter part introduces... This is the core technological breakthrough of the present invention. Matrix It is an unconstrained, zero-initialized, data-driven, learnable, adaptive adjacency matrix. In the backpropagation iteration of massive industrial samples, It breaks through the strict limitations of the human body's anatomical and physical framework, and possesses the "global perception" capability to autonomously uncover the deep logical relationships between non-natural physical connection nodes.

[0099] Industrial Real-World Application: When SMT workshop workers experience minor hand tremors due to nervous fatigue, traditional networks tend to ignore these as operational noise. However, this system... The matrix automatically assigns and solidifies the extremely high non-zero weights between the "wrist joint node" and the "ECG sensor node" through backpropagation. The system mathematically quantifies the underlying central nervous system fatigue correlation between the two heterogeneous manifestations of "hand fine motor lag" and "autonomic nervous system-regulated heart rate imbalance," perfectly extracting this high-frequency cross-modal linkage feature imperceptible to the human eye. After multi-layer residual connections to prevent degradation and finally compression using a global average pooling layer, the system outputs a dense "key feature" vector representing the current high-order topological state.

[0100] Step 4: Global Temporal Context Modeling Based on Relative Position Encoding

[0101] Personnel fatigue is not an isolated, instantaneous event, but a long-term, gradual process with a significant "cumulative effect over time." Traditional recurrent neural network models such as Long Short-Term Memory (LSTM) networks inevitably encounter gradient vanishing and severe historical memory forgetting defects when faced with continuous 30fps working data frames spanning several hours. To address this, this system introduces the key feature sequences extracted by the pre-existing Spatiotemporal Graph Convolutional Network (AGCN) into a Transformer architecture with a pure attention mechanism for full temporal domain correlation.

[0102] 1. Sequence Mapping and Linear Projection Embedding: Let the key feature sequence tensor obtained after spatial feature extraction be... ,in The time step length, For the feature channel dimension. The system first maps the sequence into the query matrix, key matrix, and value matrix required by the multi-head self-attention mechanism. For the first... The linear projection formula for an attention head is defined as:

[0103] ,

[0104] in, This is the corresponding learnable weight matrix. For the feature dimension of a single attention head, satisfying ( (Total number of attention heads).

[0105] 2. Mathematical Construction of Relative Position Offset and Attention Calculation: Since fatigue judgment is highly dependent on relative work duration (such as "how long have you worked continuously") rather than absolute physical clock time, the system directly abandons the absolute position sine coding used in the original Transformer and innovatively injects the time interval information into the computation graph in the form of "relative position offset term".

[0106] Define two time steps and The relative distance index between them is ,in To maximize the truncation of relative distances and control the size of the position parameter table, the system constructs a relative position embedding dictionary. This generates a relative position offset matrix. Its elements That is, corresponding In the dictionary The scalar value in [the context]. After introducing relative bias, the [value]... Attention is focused on the time step time step Unnormalized attention score The calculation is as follows:

[0107] ,

[0108] The matrix form of this formula is the standard form of scaled dot product attention:

[0109] ,

[0110] (Note: scaling factor) This ensures the stability of the dot product variance and prevents the Softmax function from getting stuck in the gradient vanishing saturation region during backpropagation. In this structure, Tensors calculate the absolute semantic similarity between any two time steps within a sequence, while Then, semantically independent learnable bias compensation is performed based on the relative distance between these two time steps.

[0111] 3. Multi-head aggregation and nonlinear feedforward residual network: To capture the fatigue evolution characteristics in different subspaces, the system will... The outputs of each attention head are concatenated (Concat) and then subjected to a final linear fusion:

[0112] ,

[0113] in For weight fusion.

[0114] Subsequently, the model prevents deep network degradation through residual connections and layer normalization before feeding it into a position-based feedforward neural network (FFN). To improve the non-linear smoothness of feature representation, this invention uses the GELU activation function instead of the traditional ReLU, the mathematical derivation and network structure definition of which are as follows:

[0115] ,

[0116] ,

[0117] The GELU activation function introduces random regularization through Gaussian error linear units:

[0118] ,

[0119] Finally, the output matrix of this module is again processed by residual summation and layer normalization:

[0120] ,

[0121] This output This is the second data that has been fully integrated with global context memory.

[0122] Industrial Application Mapping: Based on the underlying mathematical logic of the pure attention architecture described above, the length of the information interaction path is always compressed to... Constant-level performance. In the real-world scenario of a continuously scheduled SMT workshop, the system, through GPU parallel computing, enables the feature slice at "3 PM" to instantly span millions of data frames, "backtracking" and highly weighting it to the physiological baseline and operational frequency at "9 AM when starting work." The model clearly identifies the sluggish grasping movements exhibited by workers in the afternoon as not an occasional pause caused by a single production line bottleneck, but a gradual cumulative fatigue process in which the feature amplitude increases logarithmically with working time, achieving precise identification and tracing of the hidden sources of fatigue.

[0123] Step 5: Multimodal hierarchical dynamic fusion based on nonlinear confidence penalty

[0124] Bridging the gap between discrete, structured external behavioral semantics (such as posture keypoint coordinate sequences) and continuous, non-stationary internal physiological levels (such as pulse waves and one-dimensional time series of skin conductance) is the ultimate challenge in the field of industrial multimodal fusion. Traditional simple feature-level concatenation or linear addition fusion fails to eliminate semantic barriers between modalities and is highly susceptible to introducing errors into the global system when a single modality is contaminated by environmental noise, leading to overall system collapse. To address this, this system designs a three-tiered fusion architecture consisting of "frequency domain feature reconstruction—cross-semantic alignment—dynamic confidence penalty".

[0125] 1. Frequency domain deconstruction and time-frequency energy characterization of physiological signals

[0126] Considering that physiological signals collected in industrial settings (such as PPG and EDA) are accompanied by significant high-frequency baseline drift and power frequency interference, directly concatenating them with highly abstract visual behavioral features would lead to severe semantic mismatch. Therefore, the system first performs wavelet packet decomposition (WPD) preprocessing on the non-stationary physiological signals.

[0127] Let the preprocessed discrete physiological signal sequence be... The system employs Daubechies wavelet basis functions (such as db4) with high-order vanishing moments to decompose them down to the th order. Layers, generation Each sub-band is orthogonal in time and frequency. Then, the relative energy distribution of each sub-band is calculated:

[0128] ,

[0129] ,

[0130] in, For the first Layer Wavelet packet reconstruction coefficients for each frequency band This represents the total energy of that frequency band. The system accurately extracts specific frequency band energies highly correlated with central nervous system fatigue, particularly alpha waves (8-13Hz) reflecting relaxation or decreased alertness and theta waves (4-8Hz) reflecting mild drowsiness. These normalized energy characteristic sequences... Subsequently, it is mapped using a multilayer perceptron (MLP) to a high-dimensional physiological feature vector that is perfectly aligned with the behavioral feature dimension. .

[0131] 2. Cross-modal bilinear cross-attention interaction

[0132] To break down semantic barriers between modalities, the system constructs a hierarchical bilinear cross-attention network. This mechanism breaks away from the conventional parallel and independent processing, innovatively adopting an active interaction logic of "retrieving micro-physiological information from macro-behavioral data." Let the high-level behavioral features output by global temporal modeling of the Transformer be... The system uses behavioral features as the primary query vector and mapped physiological frequency band features as the key and value vectors, respectively, and performs linear projection on each.

[0133] ,

[0134] in, This is a learnable projection matrix. Then, scaled dot product cross-attention calculation is performed:

[0135] ,

[0136] Physical derivation significance: This formula forces the model to search for the most matching response pattern in the "internal physiological frequency band" from the perspective of "external limb movements". For example, when querying the matrix... When the semantic feature of "the worker's head is continuously drooping" is extracted, The dot product operation automatically assigns higher attention weights to moments in the physiological sequence characterized by "alpha wave energy surges," thereby achieving deep semantic forced alignment across physical properties and outputting highly purified cross-modal interaction features. .

[0137] 3. Gated Adaptive Networks and Confidence-Based Nonlinear Penalty Mechanisms

[0138] After obtaining the original behavioral characteristics Cross-modal interaction features Subsequently, the system deployed a feedforward gating mechanism for adaptive information refinement. This was achieved through a splicing operation. Combine the two and utilize the learnable weight matrix and bias Generate dynamically adjusted weights :

[0139] ,

[0140] In the ideal fusion state, the weighted representation of features is as follows: ( (For element-wise multiplication). However, the gated adaptive unit is highly dependent on the purity of the input features. To cope with extremely harsh industrial environments (such as sudden physical-level failure of a single sensor), this invention innovatively incorporates a nonlinear penalty calibration formula based on prediction confidence into the multimodal fusion attention module (MFA) to perform a secondary forced intervention on the above weights:

[0141] ,

[0142] Formula parameter deconstruction:

[0143] The probability confidence level of a lightweight network that outputs a single modality side branch in real time (e.g., the softmax peak response of a visual network to the sharpness of a face in the current frame).

[0144] This is a hyperparameter that controls the intensity decay of the penalty curve (generally set to...). about).

[0145] for A very small constant of the order of magnitude, used to avoid logarithmic operations in It encounters a singularity explosion. From the perspective of the properties of the function's derivative, when the confidence level... At that time, penalty items ,but The system maintains a normal adaptive fusion mechanism; however, when the confidence level... When the threshold for safety is breached, the logarithmic decay formula will produce an extremely steep nonlinear penalty.

[0146] 4. Industrial Practical Mapping and Noise Resistance Verification

[0147] In a real-world SMT assembly line scenario, the high-power lighting equipment above the workshop experienced a sudden and severe voltage fluctuation and high-frequency flicker. This environmental catastrophic event caused the RGB-D depth camera deployed at the workstation to be instantly overexposed, resulting in the captured visual frames being covered by a large area of ​​pure white noise and artifacts.

[0148] When the classifier of the system's underlying visual sub-network faces such high-frequency disordered noise, its output Softmax probability distribution becomes extremely flat (information entropy surges), leading to the highest confidence of the visual branch. From normal within milliseconds A precipitous drop to .

[0149] At this point, the aforementioned nonlinear penalty calibration formula quickly intervenes: because Extremely low values ​​lead to a sharp expansion of the penalty factor term, resulting in an inefficient calculation of the calibration weights. Approaching The system instantly "melted" the computational weight channel of the visual behavior modality in the backpropagation computation graph, forcing the input information flow of the final decision engine to be smoothly and 100% transferred to the deep physiological level modality (PPG and EDA features) which is not affected by light.

[0150] This dynamic fault tolerance and penalty mechanism based on underlying confidence gives the system "extreme robustness" in the face of physical failure of some industrial sensors or in environments with extremely high signal-to-noise ratios, ensuring the high purity output of fused data and the continuity of fatigue monitoring services.

[0151] Step Six: Adaptive Dynamic Threshold Fatigue Determination Based on KL Divergence and Information Theory

[0152] Traditional fatigue assessment systems often employ a rigid prior value (e.g., triggering an alarm when the overall probability exceeds 0.8). However, in real industrial settings, different employees have varying physical limits, and the same employee exhibits significant individual differences in stress tolerance under varying environmental noise and workloads. Using a fixed threshold inevitably leads to severe false alarms (causing unnecessary downtime) and missed alarms (leading to safety accidents). Therefore, this invention innovatively introduces a dynamic boundary defense based on information theory.

[0153] 1. Probability calibration and the relative entropy quantification of mind-body disconnect.

[0154] This system abandons the rigid logic of a "one-size-fits-all" approach and first performs probability calibration on the network output. To prevent the deep neural network from making "overconfidence" predictions when faced with unknown interference, the system inputs the independent behavioral feature vectors and physiological feature vectors extracted in the previous stage into their respective lightweight classification heads, and introduces a temperature-smoothing hyperparameter. Mapping using the Softmax function:

[0155] ,

[0156] Introduce a temperature parameter greater than 1 It can effectively smooth out extreme probabilities, providing a more accurate probabilistic representation of the uncertainty in real physics for subsequent consistency measurements.

[0157] Subsequently, the system utilizes the Kullback-Leibler divergence (i.e., relative entropy) from information theory to quantify the probability distribution of biological signals in real time. With behavioral probability distribution Differences in asymmetric information between them:

[0158] ,

[0159] In this scenario, since physiological signals are often the most fundamental and genuine reflection of fatigue, while external behavior is the superficial output, the formula is based on physiological distribution. The KL divergence serves as the baseline metric. The physical meaning of this divergence directly points to the "risk level of mental-physical disconnect" among workers: the higher the divergence value, the more severe the internal physiological friction within the worker, which is extremely inconsistent with their outwardly exhibited smooth movements (i.e., a state of "hidden strain").

[0160] 2. Adaptive Exponential Tightening Network for Decision Boundaries

[0161] After acquiring the real-time KL divergence, the system uses the model to plot the receiver operating characteristic curve (ROC curve) on the discrete validation set and calculates the maximum point of the Youden index, thereby calibrating an initial static baseline threshold. Subsequently, with Starting from this point, a complex nonlinear dynamic threshold generation model is constructed as follows:

[0162] ,

[0163] Deconstructing the internal mechanism of the formula:

[0164] Index risk penalty items ( ): This is the penalty coefficient. When intermodal information conflict intensifies, this term becomes negative and exponentially and rapidly lowers the discrimination threshold. This means that in high-risk situations where the situation is extremely unclear, the system proactively adopts more stringent and conservative security standards.

[0165] Confidence-weighted reward item ( The system does not simply tighten the threshold. The summation term utilizes the current real-time normalized confidence level of each modality. Multidimensional fatigue measurement function (Including visually extracted indicators such as eye-closing duration, mouth-opening frequency, and limb stiffness) Dynamic reward relaxation is applied. When high-quality sensor data clearly indicates that the worker is in good condition, the system appropriately raises the judgment threshold, effectively reducing unnecessary downtime caused by fluctuations in normal operation.

[0166] Industrial Practical Mapping and Closed-Loop Verification:

[0167] Imagine an operator on an SMT (Surface Mount Technology) assembly line who is in their 9th hour of continuous work due to a shift change. At this point, because a skilled worker can maintain the mechanical grasping and placement actions solely through deep "muscle memory" in the cerebellum and spinal cord, the probability of fatigue based solely on visual behavioral modality is extremely low (only 0.4). However, at the same time, the PPG pulse wave and weak electrocardiogram characteristics calculated by the system's backend indicate that their sympathetic nervous system is extremely exhausted, and the probability of physiological fatigue has soared to 0.9.

[0168] The system calculates the current intermodal KL divergence in milliseconds. The value surged to 0.65, a clear alarm signal to the decision-making center: an extremely dangerous "behavioral deception" had occurred. The risk penalty factor in the dynamically generated model was strongly triggered, and the exponential decay term directly raised the alarm threshold for this workstation from the default value. Automatically and strictly tightened to the dynamic threshold. At this point, the system extracts the multimodal comprehensive fatigue metric value. Although it was only 0.68 (which would be considered safe according to traditional fixed standards), it had already ruthlessly crossed the tightened 0.55 defense line. The system instantly triggered the sound and light interlock to block and intervene in safety matters.

[0169] This mechanism endows industrial safety monitoring systems with an extremely rare individual adaptive capability, constructing a dynamic safety boundary that is "personalized to each individual" without the need for retraining. It enables precise intervention and proactive prevention tens of minutes before catastrophic workplace accidents (such as limbs being caught in heavy machine tools due to microsleep) occur.

[0170] Step 7: Closed-loop optimization of the joint objective function based on multiple physical manifold constraints

[0171] To ensure that the complex high-dimensional parameter space in this invention (including the topological weight matrix of the Adaptive Graph Convolutional Network AGCN, the relative position bias term of the Transformer, and the multimodal fusion gating factor, etc.) maintains excellent convergence stability and generalization ability during long-term online operation at industrial edge environments, this invention abandons the traditional single cross-entropy loss. During the gradient backpropagation training phase of the deep neural network, the system innovatively designs a joint objective loss function that incorporates physical constraints from three perspectives: classification adversarial, manifold regression, and contrastive repulsion.

[0172] 1. Macroscopic topology of the joint objective function

[0173] The system's overall objective function is constructed from three weighted terms, guiding the model to find the global optimum in massive high-dimensional heterogeneous data:

[0174] ,

[0175] in, is the hyperparameter balancing coefficient, used to dynamically adjust the gradient contribution of each loss term in the backpropagation computation graph.

[0176] 2. Non-equilibrium adversarial classification constraints ( )

[0177] Technical challenges: During most working hours in industrial settings, workers are in a state of "absolute wakefulness," making genuine "positive samples" of fatigue breakdown extremely scarce, resulting in a severe imbalance in the distribution of positive and negative samples. Traditional classifiers are overwhelmed by the gradients of massive amounts of easily classifiable samples.

[0178] Mathematical Reconstruction: The system introduces a focal loss function based on the modulation factor.

[0179] ,

[0180] In this formula, This sets the confidence level for the model's prediction that the current sample is fatigued. In this embodiment, a modulation factor is set. .

[0181] Physical meaning: When calculating the backward gradient through differentiation, for those... (Extremely alert and easily classified) normal samples, The decay term approaches 0, thus automatically suppressing and shielding gradient propagation for such samples at the lower level; conversely, the network is forcefully compelled to exhaust all computational power to characterize these samples. The critical "hidden fatigue" samples (near the decision boundary and with ambiguous features) significantly improve the ability to detect early signs of danger.

[0182] 3. Healthy manifold fidelity regression constraint ( )

[0183] Technical pain point: Multiple nonlinear mappings in deep neural networks are prone to falling into the "black box effect," leading to feature space drift and completely losing interpretability in medicine and physiology.

[0184] Mathematical Reconstruction: Introducing Based on The regression constraint formula for the norm:

[0185] ,

[0186] Physical meaning: The “benchmark feature space” is extracted for the target worker under excellent health conditions, namely the “physiological health manifold”. These are the characterization features extracted from the current monitoring. The L2 regression loss applies an "invisible spring" to the latent space, forcing the feature vectors of all current modal dimensions to fluctuate reasonably near a healthy manifold. If the extracted features drift unnaturally and drastically due to overfitting to strong workshop noise (such as artifacts caused by mechanical vibration), This will generate a huge penalty gradient, forcing the model back to the physically valid representation range.

[0187] 4. Inter-class repulsion contrast constraint in high-dimensional space ( )

[0188] Mathematical Reconstruction: Introducing High-Dimensional Triplet Margin Loss:

[0189] ,

[0190] Physical meaning: In the latent space of a high-dimensional manifold, the system will use the current sample As anchor points. The formula internally forces the anchor point samples to be closer to the positive samples. Euclidean distance (both under fatigue state) At the same time, it violently pushed away its negative sample. Distance (in a conscious state) Furthermore, the distance difference between the two must be greater than the set safety margin boundary. (Margin). This term does not participate in direct classification, but it greatly sharpens the decision-making discernibility of the model's final classification hyperplane, ensuring that "micro-sleep" and "focus" are clearly distinguished on the feature cluster.

[0191] 5. Optimizer Dynamics and Early Stopping Mechanism

[0192] Driven by the joint gradient, the system backend does not employ conventional SGD, but instead utilizes the AdamW optimizer, which features weight decay decoupling, for iterative updates of network parameters. To help the model escape local optima (saddle points) in the extremely complex non-convex loss surface, the system introduces a cosine annealing strategy to dynamically adjust the learning rate: the learning rate decays dynamically with a cosine wave over the training period, converging rapidly in the early stages and then finely optimizing near the optimum in the later stages.

[0193] Meanwhile, the system's backend is equipped with an extremely rigorous closed-loop monitoring system for the validation set. The system calculates the generalization loss on the validation set in real time after each epoch. Generalization is only considered complete when the validation set loss oscillates without decreasing for 10 consecutive epochs (indicating the model has reached its fitting bottleneck, and continued training will lead to severe overfitting), and the variance of the loss value converges to a very small value during this period (variance...). When the convergence is considered complete (ensuring a truly smooth and non-abrupt oscillation), the system immediately triggers the early stopping mechanism, automatically interrupting the computational graph resource usage of backpropagation and serializing and persistently saving the current optimal network snapshot parameter weights.

[0194] Industrial Practice Application and Implementation: This joint optimization mechanism plays a decisive role in the entire lifecycle operation of an SMT (Surface Mount Technology) workshop. Within a continuous 12-hour work cycle, the truly dangerous moments of fatigue may only last for 5 minutes. It successfully avoided the model becoming a waste classifier that only predicts "non-fatigue" items; and it also avoided dealing with random noise from workers moving around the workshop and bending down to pick up parts. This ensured that the network did not deviate from the diagnostic essence of "physiological health". The final early stop mechanism (monitoring for 10 consecutive epochs) not only significantly saved expensive GPU training computing power, but also established a solid engineering foundation for the extreme compression and lightweight deployment of the model to the edge computing gateway (such as a terminal board equipped with an NPU) at the front end of the pipeline, completely breaking down the last barrier from theoretical derivation to industrial-grade all-weather high-frequency online inference.

[0195] Visual behavioral feature extraction experiment and multidimensional fatigue measurement verification

[0196] To verify the accuracy of the high-level semantic extraction of behavioral state features in the multimodal preprocessing stage of this invention, particularly regarding the multidimensional fatigue measurement function defined in the claims... The system conducted a visual behavior feature classification experiment in a preset scenario.

[0197] Combination Figure 4 The experimental results of the fatigue testing method shown are presented below:

[0198] (1) Precise extraction and measurement of eye features: such as Figure 4 As shown in the first row, "Eyes Closed (Fatigue)" and the second row, "Eyes Open (Normal)," the visual processing module of this invention can accurately locate the eye area of ​​a target person via a network. Even under different skin tones, makeup, and lighting conditions, the system can accurately calculate the degree of eye closure, strictly distinguishing between the normal open eye state (green indicator box) and the fatigue-induced closed eye state (red indicator box). This high-precision classification result directly serves as a multi-dimensional fatigue measurement function. It provides key basic data for "closed eyes" (i.e., in the formula). (Item).

[0199] (2) Mouth behavior characteristics and facial expression recognition: such as Figure 4 As shown in the third row, "Yawning (fatigue)" and the fourth row, "No yawning (normal)," the system demonstrates strong robustness in handling facial expressions under complex conditions. Experiments have shown that even in challenging visual conditions where the target person's head is turned, they are wearing transparent glasses, or even dark sunglasses causing partial occlusion, the algorithm of this invention can still reliably identify typical fatigue expressions such as yawning (yellow markers) and normal facial states (green markers). This accurate output precisely supports the fatigue measurement function. The "mouth opening" indicator data in the formula (i.e., the data in the formula) (Item).

[0200] Experimental results show that the front-end vision module of this invention can output behavioral feature vectors with high quality. Even under common industrial noise interference such as illumination distortion and partial occlusion, the extracted fatigue appearance features (such as closed eyes and yawning) still maintain extremely high purity and confidence. These high-quality visual behavioral features are used as query vectors input into a hierarchical bilinear cross-attention mechanism, where they are deeply fused with physiological signal features. This ensures extremely high reliability and anti-interference capability of the final adaptive dynamic threshold determination from the underlying data source.

[0201] Example 2

[0202] This embodiment provides a fatigue recognition system based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion.

[0203] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device, the fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion.

[0204] A terminal device includes a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion.

[0205] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion, characterized in that, include: Acquire multi-source heterogeneous worker monitoring data in a pre-defined industrial scenario; Data preprocessing is performed based on the acquired monitoring data; Deep spatiotemporal topology extraction of preprocessed data is performed based on the improved adaptive graph convolutional network AGCN. Global temporal context modeling is performed on the extracted features based on relative position encoding; Multimodal hierarchical dynamic fusion of modeling results is performed based on nonlinear confidence penalty; Adaptive dynamic threshold fatigue determination of fusion results is performed based on KL divergence and information theory; The judgment results are subjected to closed-loop optimization of the joint objective function based on multiple physical manifold constraints.

2. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 1, characterized in that, The data preprocessing based on the acquired monitoring data includes high-dimensional data cleaning and domain adaptation using the maximum mean difference algorithm, wherein the regenerative kernel Hilbert space mapping method is used, employing the Gaussian radial basis kernel function as the nonlinear mapping operator. The original low-dimensional and nonlinear physiological-behavioral mixed data is projected onto a high-dimensional feature space. The distribution distance between the source domain and the target domain is calculated in the high-dimensional space, using the following formula: Then, latent fatigue samples are mined. By constructing a variational autoencoder deep neural network, the first data filtered by MMD is input into the VAE network. The encoder compresses the high-dimensional input into latent variables that follow a normal distribution through a multi-layer fully connected network. The decoder then attempts to reconstruct the original input data from this latent variable, using the lower bound of evidence as the core loss criterion, expressed as: 。 3. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 2, characterized in that, The improved adaptive graph convolutional network (AGCN) is used to perform deep spatiotemporal topology extraction on preprocessed data, including heterogeneous spatiotemporal topology graph construction and latent fatigue feature extraction. First, a graph structure is defined, and key points of the human skeleton and multidimensional physiological monitoring channels in the image data are equidistantly mapped to a set of nodes in the graph structure. ,in The total number of nodes is defined; physical edges are defined based on human physiological connections, and temporal edges are defined based on temporal continuity. Simultaneously, temporal edges are established by connecting different frame states of the same node along the time axis, ultimately instantiating a high-dimensional adjacency matrix. With graph signal tensor To refine the extraction of limb relaxation or stiffness features specific to fatigue states, a set of spatial partitioning strategies for graph convolution kernels was introduced. Based on the geodesic distances between nodes and the direction of gravity, the neighborhood of the central node is deconstructed into three orthogonal subspaces, each of which independently corresponds to a normalized adjacency matrix. The first data is input into a multi-layer graph convolutional neural network, its core... layer to the first The feature aggregation derivation of the layer is as follows: ,in It is a degree matrix; The learnable weight matrix for feature channel mapping, matrix It is an unconstrained, zero-initialized, data-driven, learnable, adaptive adjacency matrix.

4. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 3, characterized in that, The global temporal context modeling of the extracted features based on relative position encoding includes addressing the memory forgetting defect of the Long Short-Term Memory (LSTM) network when faced with consecutive working data frames. First, the key feature sequences extracted by the preceding spatiotemporal graph convolutional network AGCN are introduced into a Transformer architecture with a pure attention mechanism for full temporal association. Here, sequence mapping and linear projection embedding are used. The tensor of the key feature sequence obtained after spatial feature extraction is... ,in The time step length, For the feature channel dimension, the sequence is mapped to the query matrix Query, key matrix Key, and value matrix Value required by the multi-head self-attention mechanism, for the first... The linear projection formula for an attention head is defined as: , in, This is the corresponding learnable weight matrix. For a single attention head, the feature dimension is defined, and then the absolute position sine encoding in the Transformer is discarded. Instead, the time interval information is explicitly injected into the computation graph in the form of a relative position bias term, defining two time steps. and The relative distance index between them is ,in To maximize the truncation of relative distances, construct a dictionary of relative position embeddings. This generates a relative position offset matrix. Its elements That is, corresponding In the dictionary The scalar value in the middle, after introducing relative bias, the first Attention is focused on the time step time step Unnormalized attention score The calculation is as follows: The matrix form of this formula is the standard form of scaled dot product attention: Among them, utilizing Tensors calculate the absolute semantic similarity between any two time steps within a sequence, while Then, semantically independent learnable bias compensation is performed based on the relative distance between the two time steps.

5. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 4, characterized in that, The global temporal context modeling of the extracted features based on relative position encoding also includes capturing fatigue evolution features in different subspaces. The outputs of each attention head are concatenated and then linearly fused, as shown below: , in To fuse the weights, residual connections and layer normalization are then used to prevent deep network degradation before feeding them into a position-fed forward neural network (FFN). To improve the non-linear smoothness of feature representation, the GELU activation function is used instead of the traditional ReLU. The mathematical derivation and network structure definition are as follows: , , The GELU activation function introduces random regularization through Gaussian error linear units: , Finally, the output matrix is ​​subjected to residual summation and layer normalization again: , This output This is the second data that has been fully integrated with global context memory.

6. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 5, characterized in that, The multimodal hierarchical dynamic fusion of modeling results based on nonlinear confidence penalty includes addressing high-frequency baseline drift and power frequency interference in physiological signals collected from industrial sites. First, wavelet packet decomposition is performed on the non-stationary physiological signals. Let the discrete physiological signal sequence be... Using Daubechies wavelet basis functions with high-order vanishing moments, it is decomposed to the th order. Layers, generation The frequency bands are divided into orthogonal time-frequency sub-bands, and then the relative energy distribution of each sub-band is calculated: , ,in, For the first Layer Wavelet packet reconstruction coefficients for each frequency band The total energy of the frequency band, energy characteristic sequence Subsequently, it is mapped to a high-dimensional physiological feature vector that is perfectly aligned with the behavioral feature dimension through a multilayer perceptron (MLP). Then, to break down semantic barriers between modalities, a hierarchical bilinear cross-attention network is constructed, employing active interaction logic. Let the high-level behavioral features output after global temporal modeling by the Transformer be... Using behavioral features as the primary query vector and mapped physiological frequency band features as the key and value vectors, linear projections are performed respectively: , in, The projected matrix is ​​then used for learning, followed by scaling dot product cross-attention computation: This enables deep semantic forced alignment across physical attributes, outputting highly purified cross-modal interaction features. .

7. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 6, characterized in that, The method of performing multimodal hierarchical dynamic fusion of modeling results based on nonlinear confidence penalty also includes obtaining original behavioral features. Cross-modal interaction features Then, adaptive information purification is performed using a feedforward gating unit, and the information is then spliced ​​together. Combine the two and utilize the learnable weight matrix and bias Generate dynamically adjusted weights : , In the ideal fusion state, the weighted representation of features is as follows: , To achieve element-wise multiplication, and considering that the gated adaptive unit is highly dependent on the purity of the input features, a novel nonlinear penalty calibration formula based on prediction confidence is embedded in the multimodal fusion attention module (MFA) to perform a secondary forced intervention on the feature weights: ,in, The probability confidence level of the lightweight network output in real time for a single-modality side branch. For hyperparameters, for A very small constant of the order of magnitude, used to avoid logarithmic operations in When encountering a singularity explosion, when confidence level At that time, penalty items ,but The system maintains a normal adaptive fusion mechanism; however, when the confidence level... When the threshold for safety is breached, the logarithmic decay formula will produce a steep nonlinear penalty.

8. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 7, characterized in that, The adaptive dynamic threshold fatigue determination of the fusion results based on KL divergence and information theory includes introducing a dynamic boundary defense based on information theory. First, the network output is probabilistically calibrated. To prevent the deep neural network from making overconfident predictions when facing unknown interference, the independent behavioral feature vectors and physiological feature vectors extracted in the previous stage are respectively input into their respective lightweight classification heads, and a temperature-smoothing hyperparameter is introduced. Mapping using the Softmax function: Introduce a temperature parameter greater than 1 The extreme probabilities are smoothed, and then the probability distribution of biological signals is quantified in real time using the Kullback-Leibler divergence from information theory. With behavioral probability distribution Differences in asymmetric information between them: ; After obtaining the real-time KL divergence, the model is used to calibrate the initial static baseline threshold by plotting the receiver operating characteristic curve and calculating the maximum point of the Youden exponent on the discrete validation set. Subsequently Construct a nonlinear dynamic threshold generation model as the starting point: Among them, the index risk penalty item : This is a penalty coefficient; when intermodal information conflict intensifies, this term becomes negative and exponentially and rapidly lowers the discrimination threshold. Confidence-weighted reward items The summation term utilizes the current real-time normalized confidence level of each modality. Multidimensional fatigue measurement function Dynamic rewards will be relaxed.

9. The fatigue recognition method based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion according to claim 8, characterized in that, The closed-loop optimization of the joint objective function based on multiple physical manifold constraints on the judgment result includes, to ensure convergence stability and generalization ability in the high-dimensional parameter space, constructing a joint objective loss function with three perspective physical constraints during the gradient backpropagation training phase of the deep neural network, specifically including non-equilibrium adversarial classification constraints. This is expressed as a focus loss function based on the modulation factor. , The modulation factor is the confidence level for the model to predict the probability that the current sample is fatigued. ; Healthy manifold fidelity regression constraint , indicating the introduction of based on The regression constraint formula for the norm: , This represents the baseline feature space extracted from the target worker under optimal health conditions. The current monitoring extracts characterization features; high-dimensional space inter-class exclusion contrast constraint. This is represented by the introduction of a high-dimensional triplet margin loss logic: , will the current sample As an anchor point, by forcibly bringing the anchor point sample closer to the positive sample... Euclidean distance At the same time, it violently pushed away its negative sample. distance Furthermore, the distance difference between the two must be greater than the set safety margin boundary. This ensures that micro-sleep and focused attention on feature clusters are clearly distinct.

10. A fatigue recognition system based on adaptive spatiotemporal graph convolution and multimodal hierarchical fusion, characterized in that, include: The data acquisition module is configured to acquire multi-source heterogeneous worker monitoring data in a preset industrial scenario; The preprocessing module is configured to perform data preprocessing based on the acquired monitoring data; The topology module is configured to perform deep spatiotemporal topology extraction on preprocessed data based on the improved adaptive graph convolutional network AGCN. The context module is configured to perform global temporal context modeling on the extracted features based on relative position encoding; The fusion module is configured to perform multimodal hierarchical dynamic fusion of modeling results based on nonlinear confidence penalty; The determination module is configured to perform adaptive dynamic threshold fatigue determination on the fusion result based on KL divergence and information theory. The optimization module is configured to perform closed-loop optimization of the joint objective function based on multiple physical manifold constraints on the judgment result.