Multi-Feature Adaptive Fatigue Detection Method and System for Inland River Crew in Complex Environments
By combining a collaborative reflection sensing low-light enhancement network and an end-to-end spatiotemporal fatigue feature learning network with a two-branch adaptive individual feature calibration network, the problem of feature robustness and individual differences in fatigue detection of inland waterway vessels under complex environments is solved, and high-precision fatigue state monitoring with low false alarm rate is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO UNIV OF SCI & TECH
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-26
AI Technical Summary
In complex environments, existing non-contact fatigue detection methods exhibit poor robustness and high false alarm rates in inland waterway vessel scenarios, failing to adapt to individual differences among crew members and dynamic navigation conditions, resulting in insufficient detection sensitivity.
A collaborative reflection perception low-light enhancement network is used for video quality restoration. Combined with an end-to-end spatiotemporal fatigue feature learning network and a two-branch adaptive individual feature calibration network, reflection separation and illumination reconstruction are achieved, and fatigue state is adaptively determined.
It significantly improves video image quality and feature extraction accuracy, enhances the personalization and environmental adaptability of fatigue detection, reduces false alarm rate, and achieves real-time and accurate fatigue state monitoring.
Smart Images

Figure CN122090422A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision, video enhancement and fatigue detection technology, and in particular relates to a multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments. Background Technology
[0002] Inland waterway transportation boasts advantages such as large transport capacity, low cost, and green, low-carbon operation, making it a crucial link in my country's comprehensive transportation system and playing a vital role in serving national development strategies. As my country's inland waterway transportation gradually transforms towards intelligent operation, optimizing crew structure has become an inevitable trend in the industry. However, some vessels, due to crew optimization, have experienced extended single-shift working hours for crew members, leading to increased fatigue driving risks and posing a serious challenge to navigation safety. Crew fatigue driving is one of the main human factors causing maritime accidents, especially in complex navigation environments such as nighttime, dawn, and dusk. Low light intensity combined with the refraction of lights on the water and shore can cause blurred vision for crew members, further increasing safety risks. Poor data quality in these conditions also increases the difficulty of fatigue detection. Furthermore, unlike vehicle driving, ship driving typically employs a multi-person shift work model, and individual crew members exhibit significant differences in physiological rhythms and behavioral habits, making traditional methods of using uniform, fixed thresholds for fatigue assessment unsuitable. Therefore, establishing a real-time, personalized crew fatigue monitoring system to achieve accurate perception and early warning of crew fatigue is crucial for preventing maritime accidents and protecting lives and property.
[0003] Existing fatigue detection methods are divided into contact and non-contact methods. Contact detection methods rely on wearable sensor devices to collect physiological signals. While these methods offer high accuracy, their invasiveness, inconvenience of wearing them, and disruption to normal operations make them difficult to deploy in real-world navigation environments. With the development of computer vision technology, video-based non-contact fatigue detection methods have become the mainstream research direction. However, in the complex and realistic scenarios of inland waterway vessels, existing methods face two major problems: First, the lighting conditions inside and outside the cabin are complex and variable. Videos collected at night or in dim light suffer from severe quality degradation, loss of detail, and strong interference from water surface reflections, making subsequent feature extraction difficult. Second, most methods rely on a single fatigue index, such as PERCLOS or simple multi-feature weighted fusion, lacking modeling of behavioral benchmarks in an individual's conscious state. This makes them unable to adapt to the inherent differences among crew members, resulting in insufficient detection sensitivity and a high false alarm rate. Therefore, existing non-contact fatigue detection methods in the complex scenarios of inland waterway vessels generally suffer from poor feature robustness, high false alarm rates, and insufficient personalization, making it difficult to meet the reliability requirements of real-time, accurate, and adaptive assessment of crew fatigue status in inland waterway shipping.
[0004] This invention addresses the critical issues of drastically degraded facial video quality of crew members on inland waterway vessels during complex environments such as nighttime and dawn due to uneven lighting and severe water surface reflection, as well as the inability of existing fatigue detection methods to adapt to individual crew differences and dynamic navigation conditions. It proposes a fatigue detection system and method integrating physical enhancement, spatiotemporal modeling, and adaptive decision-making. By constructing a "cooperative reflection-sensing low-light enhancement network," a collaborative optimization of reflection separation and illumination reconstruction is achieved in a unified feature space, fundamentally improving image usability under low-light and reflective coupling interference. An "end-to-end spatiotemporal fatigue feature learning network" is designed, introducing Manhattan self-attention and retention mechanisms to model long-term fatigue characteristics with linear complexity and directly regress multi-dimensional quantitative indicators covering pupil movement, eye state, and facial expression. Furthermore, a "two-branch adaptive individual feature calibration and navigation scenario weighted network" is proposed, establishing an individual crew member alertness benchmark and combining AIS data to identify fatigue-prone navigation conditions, achieving personalized and contextualized adaptive determination of fatigue status. Finally, the above algorithms are integrated into a shipborne edge computing platform, forming a complete embedded detection system from video enhancement and feature extraction to intelligent decision-making. Summary of the Invention
[0005] To address the aforementioned problems, this invention proposes a multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments. The technical solution is as follows:
[0006] A multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments, including:
[0007] Step 1: In the video data preprocessing stage, cameras deployed in the ship's cockpit capture raw monitoring videos containing crew members. To address the degradation caused by low illumination and strong water surface reflection coupled in the video, a "cooperative reflection-sensing low-illuminance enhancement network" is designed for frame-by-frame processing. First, the input frames are analyzed using a physical prior encoder, simultaneously estimating the scene's transmission rate map and initial brightness map. Next, the raw frames are concatenated with these two physical prior maps, and a shared feature base is generated via a multi-scale reversible encoder. Then, this feature base is projected onto an improved HVI color space, and a reflection-illuminance cooperative attention module is used for interactive processing within the space to achieve reflection noise separation and dark area illumination enhancement. Finally, the enhanced features are inversely transformed back to the RGB space, and a residual reconstruction network outputs sharpened video frames. Ultimately, the original low-quality video is restored into a high-quality, clear video sequence free from reflection interference, providing reliable input for subsequent fatigue feature extraction.
[0008] Step 2: In the fatigue feature extraction stage, the high-quality crew facial video sequence preprocessed in Step 1 is input into the "end-to-end spatiotemporal fatigue feature learning network" deployed on the shipboard edge computing device for computation. First, the RMT-3DCNN backbone in the network spatiotemporally encodes continuous video frames, models long-range temporal dependencies through a forward linear recursive method by preserving mechanisms, and efficiently captures periodic fatigue patterns such as blink frequency and duration of fatigue expressions. At the same time, the fatigue region perception cross-head hybrid interaction mechanism integrated within the network enables different attention heads to collaboratively focus on and finely process features of key facial regions such as the eyes and mouth. Then, the high-dimensional spatiotemporal representation output by the network is fed into a multi-branch regression head, mapped into a structured nine-dimensional fatigue feature vector. Each element of this vector corresponds to a quantitative estimate of nine fatigue indicators: maximum pupil movement speed, average pupil movement speed, percentage of eyelid closure, longest duration of eye closure, blink frequency, frequency of moderate and severe fatigue expressions, and their respective longest duration.
[0009] Step 3: In the individual adaptive fatigue assessment stage, using the nine-dimensional fatigue feature vector generated in Step 2, the mean and variance of each dimension are calculated to construct a personalized "individual sobriety baseline vector," which serves as the baseline for fatigue state discrimination under normal conditions. Simultaneously, AIS data is accessed in real time to calculate the standard deviation of ship speed within the most recent time window. Maximum change in heading angle If within a set threshold, a weighted calculation based on the individual's alertness baseline and AIS is used. First, the offset between the current feature and the individual baseline is calculated. Then, based on the AIS contextual marker, a sensitivity coefficient greater than 1 is introduced to amplify the differential features in fatigue-prone situations. Finally, a comprehensive fatigue index is calculated using a classification function. Ultimately, the fatigue state is output.
[0010] Step 4: In the system integration and visualization output stage, the aforementioned enhanced preprocessing module CRL-Net, the spatiotemporal fatigue feature learning network RMT-3DCNN+LSTM, and the dual-branch adaptive decision network are integrated and deployed on the shipboard NVIDIA Jetson AGXOrin edge computing platform. This platform optimizes and accelerates the trained model using TensorRT and builds a software system based on the ROS2 framework to manage the real-time synchronous acquisition, modular processing, and message communication of camera video streams and ship AIS data. The system ultimately outputs a personalized fatigue level for each crew member: "Awake," "Mild Fatigue," "Moderate Fatigue," and "Severe Fatigue." The visualization warning interface, developed using PyQt5, displays the detected image on the monitoring screen in real time. Simultaneously, it dynamically visualizes the scene on the monitoring interface and triggers screen flashing warnings, sound alerts, and log recording when moderate or higher fatigue levels are detected. It supports historical data review and alarm information export, thereby achieving 24 / 7 automated monitoring and safety warning of crew member fatigue status.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0012] (1) Compared with traditional step-by-step image enhancement and fatigue detection methods, this invention achieves integrated restoration of low illumination and water surface reflection coupling degradation through the designed collaborative reflection sensing low illumination enhancement network, avoiding the problems of error accumulation and obvious artifacts in step-by-step processing, significantly improving the quality and stability of video images under complex lighting conditions, and providing reliable input for subsequent fatigue feature extraction;
[0013] (2) Compared with traditional methods that rely on multiple independent models or serial structures to extract fatigue features, the end-to-end spatiotemporal fatigue feature learning network adopted in this invention, by introducing Manhattan self-attention mechanism and retention mechanism, achieves unified modeling and direct regression of long-term multi-dimensional fatigue representation with linear computational complexity. This not only avoids the problems of complex coordination between models and high computational redundancy in traditional methods, but also significantly improves the accuracy of feature extraction and the overall efficiency of the system.
[0014] (3) Compared with the method of fatigue judgment using a fixed threshold or a single rule, the dual-branch adaptive individual feature calibration and navigation scenario weighted network proposed in this invention effectively overcomes the misjudgment problem caused by individual differences and changes in navigation conditions by establishing an individual alertness benchmark and combining it with real-time AIS data for situational awareness decision-making, which greatly improves the personalization of fatigue detection, environmental adaptability and overall early warning accuracy. Attached Figure Description
[0015] Figure 1 This is an overall architecture diagram of the multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments, as described in this invention example.
[0016] Figure 2 This is a hardware deployment diagram of the system in an example of the present invention, showing the connection relationship between the camera, edge computing device and display terminal;
[0017] Figure 3 This is an architectural diagram of the "cooperative reflectance sensing low-light enhancement network" in an example of the present invention;
[0018] Figure 4 This is an architecture diagram of the "end-to-end spatiotemporal fatigue feature learning network" in an example of the present invention;
[0019] Figure 5 This is a flowchart of the "two-branch adaptive individual feature calibration and navigation scenario weighting network" in the example of this invention;
[0020] Figure 6 This is a schematic diagram of the real-time output interface of the system in an example of the present invention, showing the fatigue classification status and the form of visual warnings. Detailed Implementation
[0021] To make the objectives and technical solutions of the present invention clearer, the present invention will now be described more completely in conjunction with the accompanying drawings.
[0022] This invention addresses the challenges of video image quality degradation caused by insufficient cabin lighting and strong water surface reflection during inland waterway vessel navigation, especially in complex environments such as nighttime and dusk. These challenges include difficulties in extracting fatigue features from crew members and a lack of individual differentiation in the detection results. A multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments are proposed. The specific implementation of this method and system is described in detail below.
[0023] Please see Figures 1-6 , Figure 1 This is an overall architecture diagram of an embodiment of the multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments proposed in this invention. The method acquires raw video through a camera deployed in the cockpit and integrates it with AIS data. The raw video stream is first restored to its quality by a cooperative reflectance perception low-light enhancement network, removing water surface reflections and enhancing details in dark areas. The enhanced, clear video sequence is then input into an end-to-end spatiotemporal fatigue feature learning network to extract a nine-dimensional structured feature vector covering pupil movement, eye state, and fatigue expression. Finally, through a dual-branch adaptive individual feature calibration and navigation scenario weighting network, combined with a pre-established individual crew member alertness benchmark and real-time AIS navigation scenario, personalized and contextualized fatigue state determination and graded output are performed, thereby achieving robust and adaptive detection of crew member fatigue states in complex inland waterway environments.
[0024] Figure 2This diagram illustrates the hardware deployment and data flow of the multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments, as described in this invention. The system uses the cockpit of an inland waterway vessel as the application scenario. The core computing unit is an edge computing device deployed in the cockpit. Cameras are responsible for capturing raw monitoring video streams containing crew members' faces, and the AIS data interface receives real-time information such as vessel speed and heading. The edge computing device integrates algorithm modules for video enhancement, feature extraction, and adaptive judgment. The processed fatigue level, alarm information, and enhanced video images are ultimately output to the monitoring display terminal in the cockpit.
[0025] Figure 3 The following is an architecture diagram of the collaborative reflectance sensing low-light enhancement network proposed in this invention. The specific steps are as follows:
[0026] Step 1: In the video data acquisition and enhancement preprocessing stage, the data acquisition module deployed in the ship's cockpit acquires the original monitoring video data containing the crew. To address the common problem of low light and strong water surface reflection coupling degradation in the video, the Cooperative Reflectance Perception Low Light Enhancement Network (CRL-Net) is used to repair and enhance the original video stream, obtaining a high-quality and clear video sequence, providing reliable input for subsequent feature extraction.
[0027] Let the acquired original image of frame t be... ,in and These represent the image height and width, respectively, and 3 represents the three RGB color channels. In low-light and reflection-coupled scenes, It can be modeled as the desired clear transport layer image. Image with interference reflection layer The complex mixture, accompanied by an overall decrease in brightness, i.e. ,in For an unknown degenerate function, These are scene-specific parameters. The goal of this step is to design the network. Learning from To high-quality, clear images Mapping: , These are network parameters.
[0028] Design a shared encoder with physical prior guidance. (Network) First, the input Physical degradation feature analysis is performed. This is specifically implemented through a joint transmittance-luminosity estimation module. This module is based on a modified lightweight ConvNeXt-Tiny architecture. Its input is... First, the input image undergoes initial downsampling and channel expansion using a 4×4 convolutional layer with a stride of 4, reducing the spatial size to 1 / 4 of its original size and expanding the number of channels to 96. Subsequently, the features are processed through four consecutive ConvNeXtBlock stages. Each ConvNeXtBlock mainly consists of depthwise separable convolutions, layer normalization, and the GELU activation function. After the first three stages, a 2×2 convolutional layer with a stride of 2 further halves the feature map spatial size. After these four stages, the input image is downsampled to 1 / 32 of its original size. Above these high-level semantic features, the module uses two parallel convolutional heads for regression prediction: one head uses a 3×3 convolutional kernel and outputs a transfer rate attenuation map. The graph is activated by the Sigmoid function, and each spatial location... and channels value In the [0,1] interval, the degree of light transmission attenuation due to the medium at the corresponding location is quantified; the lower the value, the more severe the attenuation. The other head also uses a 3×3 convolution kernel to output the initial brightness distribution map. Its pixel value Also in the [0,1] interval, it represents the inherent illumination intensity distribution of the scene before removing reflection interference.
[0029] Then, the original image frame Transmission rate diagram With brightness diagram The input representation is concatenated along the channel dimension to form an enhanced input representation. This tensor is fed into a multi-scale reversible encoder. This encoder aims to extract multi-scale features while ensuring that information is not lost during deep processing via a reversible neural network structure. The encoder contains three downsampling levels. Each level consists of multiple stacked reversible residual blocks. For a reversible block, its input features... Divided into two parts along the channel dimension: and The following transformations are performed within the block:
[0030]
[0031] in, and They are two structures with the same structure but different parameters ( , Independent residual subnetworks. Each subnetwork typically consists of two 3×3 convolutional layers, with a GELU activation function in between. The output is... This design eliminates the need to save intermediate activation values during backpropagation, allowing the input to be directly reconstructed from the output, thus ensuring the integrity of the feature information flow. The encoder ultimately outputs a shared feature base containing rich physical and semantic information. ,in This represents the number of feature channels (e.g., 256).
[0032] Adaptive HVI color space transformation and collaborative attention restoration. To effectively decouple reflection and illumination components, this module will share feature bases. The color appearance-related components are projected onto an improved HVI color space through an adaptive HVI transform layer. This transform is not fixed; the parameters of its transform matrix are dynamically generated by a lightweight multilayer perceptron, whose input is the transfer rate map. The global average pooling feature vectors are used. This allows the projection of the color space to adaptively adjust according to the overall reflection intensity of the current frame, enhancing the model's generalization ability to different degrees of degradation.
[0033] The obtained HVI space features Above, we constructed a reflection-lighting collaborative attention module. First, through two independent 1×1 convolutional layers (with weight matrices respectively) and )Will It is decomposed into two components that focus on different degradation factors:
[0034]
[0035] It is a reflection-sensitive component, which mainly encodes the mode of reflected noise; It is the light-sensitive component, which mainly encodes the brightness and darkness distribution information of the scene.
[0036] To enable the processing of the two components to guide and correct each other, we introduce a cross-attention mechanism. First, we... As a query These serve as keys and values, respectively. Specifically, they are achieved through three different 1×1 convolutional layers. Projection as ,Will Projection as and The calculation of attention weights and feature updates are as follows:
[0037]
[0038]
[0039] in This is the dimension of the key vector. This process utilizes global information about the illumination distribution to guide the separation of reflection components, for example, avoiding misclassifying shadows in dark areas as reflections. Similarly, with For query, Using keys and values, the update of the illumination component is calculated. This helps to avoid overexposing highly reflective areas when brightening dark areas. Finally, the two cross-attention corrected components are fused through a channel attention gating mechanism to obtain the enhanced HVI spatial features. .
[0040] Physically consistent image reconstruction and loss function. Enhanced HVI features. The image is mapped back to the standard RGB color space using an inverse HVI transform layer. Subsequently, a residual reconstruction network consisting of multiple residual blocks refines the RGB features to predict the residual image. Each residual block contains two 3×3 convolutional layers and identity shortcut connections across layers. Finally, the network outputs a sharp image by adding the original input, which has undergone slight gamma correction, to the predicted residual.
[0041]
[0042] in It is a learnable scalar parameter used for global brightness baseline adjustment of the raw input.
[0043] The entire CRL-Net is optimized end-to-end using a multi-task joint loss function. This loss function is defined as:
[0044]
[0045] This is the reflection separation loss. During the training phase, if a ground truth reflection layer exists... Then calculate loss If not, then prior constraints are used, such as encouraging the smoothness of the reflection component in the gradient domain.
[0046] It is illumination enhancement loss, which includes brightness uniformity loss and illumination smoothing loss.
[0047] It is a perceptual loss, using a VGG-19 network pre-trained on ImageNet to compute the augmented image. With clear image of target Between specific intermediate layer features Distance is used to preserve the realism of high-level semantics and texture structure. Hyperparameters Used to balance various losses.
[0048] The CRL-Net method primarily addresses image degradation caused by the complex environment during navigation. The network first estimates two physical priors: transmittance and brightness, providing the model with clear cues to distinguish between dark areas and reflections. Building upon this, a collaborative processing mechanism utilizes cross-attention within an improved color space to allow reflection removal and brightness enhancement processes to interact and mutually correct each other in real time, avoiding error accumulation and artifacts inherent in step-by-step processing. The entire architecture employs lightweight techniques such as reversible design to ensure real-time operation on shipboard edge devices, outputting a stable and clear video stream, providing crucial high-quality input for subsequent fatigue feature extraction.
[0049] Figure 4 The following is an architecture diagram of the "end-to-end spatiotemporal fatigue feature learning network" in an example of the present invention. The specific steps are as follows:
[0050] Step 2: This step aims to automatically and accurately extract key fatigue features of the crew members from the high-quality, clear video sequence output in Step 1. Using the RMT-3DCNN+LSTM network, with RMT as the backbone of the 3DCNN+LSTM, its innovative Manhattan self-attention mechanism introduces a distance-based spatial decay prior, enabling the network to focus more on key facial regions such as the eyes and mouth, while significantly reducing computational complexity and meeting the real-time processing requirements of edge devices. A deep learning model capable of integrating the spatiotemporal information of the video and directly regressing to calculate the values of nine key fatigue features was constructed.
[0051] The implementation details of step 2 are as follows: The network input is a sequence of T consecutive clear facial images processed in step 1. Each frame of the image has a resolution of The goal of the network is to learn a mapping function. The video sequence is mapped to a structured nine-dimensional fatigue feature vector: ,in For network parameters, The nine dimensions of this vector correspond to: (1) maximum pupillary speed (MSPM), (2) average pupillary speed (ASPM), (3) percentage of eyelid closure (PEC), (4) longest duration of eye closure (MDEC), (5) blink frequency (BF), (6) frequency of moderate fatigue expression (FMFE), (7) frequency of severe fatigue expression (FSFE), (8) longest duration of moderate fatigue expression (MDMFE), and (9) longest duration of severe fatigue expression (MDSFE).
[0052] This design employs an RMT backbone network based on an improved retention mechanism for feature extraction. Traditional Transformer self-attention lacks explicit spatial priors and has high computational complexity. This invention introduces Manhattan self-attention, injecting an explicit spatial decay prior based on Manhattan distance into the model.
[0053] For the input feature map First, the query (Q), key (K), and value (V) are obtained through linear projection. The core of MaSA is the introduction of a location-dependent spatial decay matrix when calculating attention weights. :
[0054]
[0055] in, This represents element-wise multiplication. Spatial decay matrix. Each element is defined as:
[0056]
[0057] here, and Representing the first The and the first The coordinates (integer indices) of each feature token on a two-dimensional feature map. It is a learnable decay factor ( The physical meaning of this formula is: the greater the Manhattan distance between two spatial locations (…), the more likely it is to be true. The larger the value, the greater its corresponding attenuation coefficient. The smaller the value, the more exponentially the contribution of distant feature points to the attention of the current point will decay. This gives the model a spatial locality prior, allowing it to focus more on key local regions such as the eyes and mouth, while perceiving the global context with linear attention complexity.
[0058] To handle high-resolution inputs, the network employs decomposed Manhattan self-attention in shallow layers, which decomposes the two-dimensional attention into two one-dimensional attention operations along the horizontal and vertical directions, further reducing the computational burden while maintaining the same receptive field shape.
[0059] The specific steps of feature extraction are as follows:
[0060] (1) Convolution Stem and Downsampling: The input frame sequence is first passed through a "convolution stem" consisting of multiple 3×3 convolutional layers and max pooling layers to downsample the single-frame image and embed it into a feature map. For a T-frame sequence, the initial spatiotemporal features are obtained. , where s is the downsampling factor.
[0061] (2) Multi-stage RMT encoding: features It is fed into a four-stage RMT encoder. Each stage consists of multiple stacked RMT blocks. Each RMT block contains:
[0062] Manhattan Self-Attention Layer (MaSA): Extracts global and local features with spatial priors.
[0063] Local Context Enhancement Module (LCE): A 5×5 depthwise separable convolutional layer used to further enhance local detail features.
[0064] Feedforward Network (FFN): A two-layer MLP that performs feature transformation.
[0065] Downsampling is performed between stages using 3×3 convolutions with a stride of 2 to progressively expand the receptive field and extract high-level semantic features. After four stages, high-level semantic features are obtained. .
[0066] (3) 3D convolutional spatiotemporal modeling: In order to explicitly capture short temporal motion patterns between frames, features are... The data is fed into a lightweight 3D convolutional module. This module contains two 3×3×3 3D convolutional layers, which fuse features from adjacent frames in the temporal dimension to initially extract motion-related spatiotemporal features. .
[0067] (4) LSTM long-term dependency modeling: In order to capture long-period patterns such as the duration of fatigue expression and periodic blinking, the features after flattening the spatial dimension are used. The sequence is considered as a length T and input into a bidirectional LSTM network. The hidden state update formula of the LSTM is:
[0068]
[0069] in Let be the hidden state at time t. This represents the cell state. The forward and backward hidden states of the final time step in the bidirectional LSTM are concatenated to form a feature vector containing global temporal information for the entire video sequence. .
[0070] (5) Multi-branch regression head: Finally, the feature vector The data is fed into a multi-branch regression head. This regression head consists of multiple parallel fully connected (FC) subnetworks, each responsible for regressing one or a set of relevant fatigue features. These subnetworks share... As input, but with independent parameters to ensure specific modeling of different types of features. Ultimately, the network outputs a nine-dimensional feature vector. .
[0071] The calculation methods for the nine features and their supervision during training are as follows:
[0072] (1) Pupil motion features: The network regresses the pupil center coordinates of each frame using an additional auxiliary head. .
[0073] Maximum Motion Speed (MSPM): Calculated by dividing the maximum Euclidean distance between consecutive frames by the frame interval. .
[0074]
[0075] During training, one output branch of the network directly regresses the MSPM estimate and calculates the mean squared error (MSE) loss with the ground truth.
[0076] Average Movement Speed (ASPM): Calculates the average movement distance across all consecutive frames.
[0077]
[0078] We also use MSE loss supervision.
[0079] (2) Eye condition characteristics: Network regression of an eyelid closure probability sequence (1 indicates fully open).
[0080] Percentage of eyelid closure (PEC): Statistics The proportion of frames below a threshold (e.g., 0.5) out of the total number of frames T.
[0081]
[0082] Longest duration of eye closure (MDEC): In the sequence, find the length of the longest consecutive subsequence below a threshold and convert it to time (frames ×). ).
[0083] Blink frequency (BF): Statistical frequency per unit of time, The number of complete "valleys" from above the threshold to below the threshold and then back above the threshold. The sequence needs to be smoothed and peak detected.
[0084] These features can be obtained from the network output via a dedicated time-series analysis module. The sequence shows that during training, the true PEC, MDEC, and BF values can be directly used as regression targets, or the loss function can be designed to optimize the network output. The sequence approximates the real eyelid opening and closing sequence.
[0085] (3) Fatigue expression features: Network regression of the probability distribution of fatigue expression state at each time step These correspond to the probabilities of "awake", "moderately fatigued", and "severely fatigued", respectively.
[0086] Moderate fatigue facial expression frequency (FMFE): Statistical frequency per unit time. The number of times the probability of the "moderate fatigue" category is greater than the threshold of 0.5.
[0087]
[0088] Frequency of severe fatigue expression (FSFE): Similarly, count the frequency of the "severe fatigue" category.
[0089]
[0090] The longest duration of moderate / severe fatigue expression (MDMFE / MDSFE): In the state sequences of "moderate fatigue" and "severe fatigue", find the length of the longest subsequence that is continuously judged as that state and convert it into time.
[0091] The calculation of these features depends on the network's accurate classification of facial expression states for each frame. During training, the network is supervised using frame-level fatigue expression labels (awake, moderate, severe), typically employing cross-entropy loss. Then, the four statistical features mentioned above are calculated from the network's output sequence through post-processing as the network's indirect output. Alternatively, multi-task learning can be designed, allowing one branch of the network to directly regress these statistical feature values.
[0092] During training, the entire RMT-3DCNN+LSTM network was pre-trained end-to-end on a crew fatigue video dataset. The loss function was a weighted sum of the losses for each feature regression task. .
[0093]
[0094] in Typically, smoothed L1 loss or mean squared error loss is used. After optimization, the network parameters are fixed, and it is deployed to an edge computing device. During real-time detection, the system feeds continuous video clips into the network in a sliding window manner. After forward propagation, the network directly outputs a nine-dimensional fatigue feature vector within the current time window. Step 3 involves personalized judgment.
[0095] The end-to-end spatiotemporal fatigue feature learning network designed in this step adopts the MaSA mechanism to introduce an explicit spatial decay prior to the model, enabling it to more accurately focus on subtle changes in key facial regions while maintaining linear computational complexity. Secondly, through a hybrid architecture of RMT, 3DCNN, and LSTM, comprehensive modeling from local spatial features and short-term temporal motion to long-term dependencies is achieved, effectively capturing fatigue patterns across various time scales, from instantaneous blinks to persistent fatigue expressions. Employing an end-to-end regression framework, it directly outputs standardized nine-dimensional feature vectors, avoiding the error accumulation and complex parameter tuning caused by multiple independent models in traditional methods, significantly improving the automation, accuracy, and overall system efficiency of feature extraction.
[0096] Figure 5 The flowchart of the "two-branch adaptive individual feature calibration and navigation scenario weighted network" in this invention example is as follows:
[0097] Step 3: In the fatigue determination stage based on individual adaptation and AIS perception, the nine-dimensional fatigue feature vector generated in real time in Step 2 is used for comprehensive analysis through a "two-branch adaptive individual feature calibration and navigation scenario weighted network," ultimately outputting the fatigue level as: "Sober," "Mild Fatigue," "Moderate Fatigue," and "Severe Fatigue." This step aims to solve the problem of misjudgment of fatigue status caused by individual physiological differences among crew members and different navigation conditions. By establishing an individual sobriety benchmark, integrating real-time Automatic Identification System (AIS) data, and performing adaptive weighted decision-making, individual adaptation of the detection results is achieved.
[0098] The specific steps are: first, establish a personal baseline; then, perceive the navigation environment; then, integrate and calculate the two; and finally, make a judgment.
[0099] The purpose of calculating the individual sobriety baseline during the system initialization phase is to obtain the individualized characteristic baseline of the crew member in a fatigue-free state.
[0100] (1) Data collection: During the initial period after the system confirms that the crew member is conscious. Within minutes, video is continuously acquired and a nine-dimensional fatigue feature vector is extracted using the network in step 2 to obtain the sequence. ,in The total number of samples, .
[0101] (2) Calculate the baseline vector: Calculate the mean of the sequence in nine feature dimensions to form the individual conscious baseline vector. .
[0102]
[0103] in, Representing the The individual's normal level of alertness is characterized by a certain feature.
[0104] (3) Calculate the confidence range: Simultaneously calculate the standard deviation of each feature dimension. This is used to define the normal range of fluctuations in the crew member's characteristics. and The characteristics of the crew member are jointly stored in the crew member's profile, serving as the basis for personalized calibration of all subsequent judgments.
[0105] During navigation, the current navigation environment is quantitatively assessed by parsing AIS data. The ship's speed, broadcast via AIS, is read in real time. (Section) and heading (Degree) information. In the sliding time window. Within minutes, calculate two key features:
[0106] Standard deviation of speed This reflects the stability of the ship's speed.
[0107]
[0108] in The number of data points within the window. This represents the average speed. The smaller the value, the more likely the ship is sailing at a constant speed.
[0109] Maximum change in heading This reflects the stability of the ship's course.
[0110]
[0111] The smaller this value, the more likely the ship is to maintain a straight course with minimal turning maneuvers. When a ship is in a constant straight-course state for an extended period, crew members are more prone to fatigue due to the reduced maneuvers. In such cases, a combination of AIS and individual alertness benchmarks should be used to diagnose fatigue.
[0112] The calculation results are compared with a preset threshold to generate a binary AIS context flag. Determine the scenario:
[0113]
[0114] in and Thresholds set based on inland waterway navigation experience Festival, Degree. When When the system determines that the current situation is one of easy fatigue; when When that happens, it is a normal situation.
[0115] According to AIS contextual markers The value is calculated using two different strategies:
[0116] In normal circumstances ( Fatigue assessment is performed using individual alertness benchmarks, and the specific steps are as follows:
[0117] (1) Calculate the differential eigenvectors :
[0118] For the nine-dimensional fatigue feature vector extracted at the current moment Calculate its relationship with the individual's conscious baseline vector. The absolute difference:
[0119]
[0120] in:
[0121] It is the fatigue feature vector extracted within the current 30-second time window;
[0122] It is the baseline vector of individual consciousness;
[0123] Indicates the first The absolute deviation of the current value of each feature from the benchmark value.
[0124] (2) Calculate the standardized difference score :
[0125] To eliminate the influence of different characteristic units and to consider the normal range of individual fluctuations, the difference is divided by the standard deviation of that characteristic under conscious conditions:
[0126]
[0127] in:
[0128] It is the first The standard deviation of each characteristic in a conscious state;
[0129] It is a very small constant (usually taken as...). ), used to ensure that the denominator is not zero;
[0130] This represents the standard deviation multiple by which the current characteristic value deviates from the individual's normal range; it is dimensionless.
[0131] (3) Calculate the overall difference score :
[0132] The weighted sum of the nine standardized difference scores yields the composite difference score:
[0133]
[0134] in It is the first The preset weights of each feature satisfy Weight It reflects the contribution of different features to fatigue, which can be learned through training data or set based on experience, with higher weights assigned to features with larger F values.
[0135] Fatigue level determination: using the benchmark grading threshold. ,in The judgment rule is as follows:
[0136]
[0137] In situations where fatigue is likely to occur ( Because prolonged, stable, straight voyages can easily lead to a decrease in crew alertness, it is necessary to be more sensitive to changes in fatigue characteristics. Therefore, based on calculations under normal conditions, a sensitivity coefficient greater than 1 is introduced to amplify the overall difference score, thereby lowering the threshold for determining fatigue.
[0138] (1) Calculate the differential eigenvectors Standardized difference score (Same as normal situation).
[0139] (2) Calculate the overall difference score (Same as normal situation).
[0140] (3) Introduce the situation sensitivity coefficient ( The value is determined through experimental optimization (usually set to 0.3-0.5), and the baseline threshold is compressed.
[0141]
[0142] Fatigue level determination:
[0143]
[0144] This mechanism makes the system more sensitive to differences in the same physiological characteristics under fatigue-prone conditions, enabling proactive early warning.
[0145] Sensitivity coefficient The settings can be optimized through experimental data, collecting fatigue data under different flight scenarios, and adjusting accordingly. This ensures that the false negative rate of the system under fatigued conditions is comparable to that under normal conditions.
[0146] Assume the baseline and standard deviation of a crew member's level of alertness are as follows:
[0147]
[0148]
[0149] Weight vector (The weights sum to 1).
[0150] The currently extracted feature vector is:
[0151]
[0152] If the current situation is normal ( ):
[0153] (1) Calculate the difference: Similarly, calculate other dimensions.
[0154] (2) Standardization: Similarly, calculate the others.
[0155] (3) Weighted summation:
[0156] If the current situation is one of easy fatigue ( ,Pick ):
[0157] Steps (1) to (3) are the same, yielding...
[0158] Normal situation (benchmark threshold) ):because The condition was diagnosed as "mild fatigue".
[0159] Fatigue-prone situations The threshold after compression ):because The condition was diagnosed as "moderate fatigue".
[0160] Grading threshold Optimization based on experimental data determined that during the training phase, leave-one-out cross-validation was used to calculate the data of each subject under different scenarios. Value distribution, select the threshold combination that maximizes the overall accuracy of the four-level classification.
[0161] The system ultimately outputs a personalized fatigue level for each crew member ("awake", "mildly fatigued", "moderately fatigued", "severely fatigued"), and simultaneously outputs the current judgment context (normal / fatigued) and a comprehensive difference index. For recording and analysis.
[0162] This step utilizes personal sobriety benchmarks. , Individual difference calibration was achieved, and contextual labels were generated by real-time parsing of AIS data. Furthermore, a context-adaptive threshold is employed to achieve individual difference adaptation. The combination of these two approaches enables the system to adapt to both individual crew member differences and changes in the navigation environment, significantly improving the accuracy and practicality of fatigue detection.
[0163] Step 4: Integrate the video enhancement module, spatiotemporal fatigue feature extraction module, and dual-branch adaptive decision module to construct a complete end-to-end crew fatigue monitoring system deployable on shipboard edge computing devices. The entire system is deployed on the shipboard NVIDIA Jetson AGXOrin edge computing platform. This platform, through its high-performance GPU and AI acceleration core, combined with the TensorRT inference optimization tool, quantizes and optimizes the trained model to generate an efficient inference engine, ensuring that all algorithm modules can meet real-time processing requirements. The software system is built on the ROS2 middleware framework, achieving loose coupling, highly reliable synchronization, and message communication between camera video streams, AIS data streams, and various algorithm processing nodes.
[0164] After the system starts up, the specific process is as follows:
[0165] 1. Data Synchronization Acquisition: The camera_driver_node drives the camera and publishes raw images at a rate of 30fps in the topic / image_raw; the ais_parser_node continuously parses the ship's AIS data and publishes the topic / ais / data containing speed, heading, and situation marker F_ais.
[0166] 2. Enhancement and Feature Extraction: The `image_enhancement_node` subscribes to ` / image_raw`, calls a TensorRT-accelerated CRL-Net model to enhance each frame of the image, and publishes the enhanced image to the topic ` / image_enhanced`. The `feature_extraction_node` subscribes to ` / image_enhanced`, maintains a 30-second sliding window buffer, and when the window is full, calls an accelerated RMT-3DCNN+LSTM network for forward computation, outputting a nine-dimensional fatigue feature vector `F_current` to the topic ` / fatigue_features`.
[0167] 3. Adaptive Status Determination: The `fatigue_detection_node` synchronously subscribes to ` / fatigue_features` and ` / ais / data`. This node loads the current crew member's personalized baseline profile `B_awake` and `σ_awake`, selects the corresponding determination strategy based on the value of `F_ais`, calculates the comprehensive difference index `S_raw`, and classifies it. Finally, this node publishes the four-category fatigue level determination result—"awaited," "mild fatigue," "moderate fatigue," or "severe fatigue"—along with metadata such as `S_raw`, AIS context markers, and confidence level, to the final status topic ` / fatigue_status`.
[0168] The visual alert interface is independently developed based on the PyQt5 framework, running as a standalone `monitor_ui_node` that subscribes to the ` / fatigue_status` topic. Its core display and alerting logic is as follows:
[0169] Main monitoring interface: The central area of the interface displays a real-time enhanced video stream of the crew's faces, taken from / image_enhanced, with key feature points such as the eyes and mouth overlaid and rendered. The right panel dynamically updates key information, including:
[0170] Status indicator lights and level display: Show the current fatigue level. Indicator light colors: green for "awake", yellow for "mild fatigue", orange for "moderate fatigue", and red for "severe fatigue".
[0171] Fatigue Index: Displays the current S_final value and a simplified historical trend graph.
[0172] Navigation scenario: Clearly indicate whether the current navigation is "normal navigation" or "navigation prone to fatigue".
[0173] Tiered alarm mechanism:
[0174] Tiered alarm mechanism: The system triggers differentiated alarm strategies based on the different fatigue levels determined in step 3.
[0175] (1) "Mild fatigue" prompt level: The status indicator light turns yellow, and a gentle text prompt bar appears at the bottom of the interface: "Mild fatigue signs detected, please take a rest", with no sound alarm.
[0176] (2) "Moderate Fatigue" Warning Level: The status indicator light turns orange and starts flashing, and the orange breathing light effect is triggered on the edge of the monitoring screen. At the same time, the system plays a pre-recorded voice reminder, such as: "Warning! Moderate fatigue state, it is recommended to take a proper rest."
[0177] (3) "Severe Fatigue" Alarm Level: The status indicator light turns red and flashes at a high frequency, and the screen border flashes bright red. The system immediately plays an emergency voice alarm, such as: "Alarm! Severe fatigue driving condition, please rest immediately!" If this condition lasts for more than 60 seconds, the system will generate the highest level warning log and send an emergency notification to the shore-based management platform through the network module.
[0178] (4) All alarm events, including timestamps, crew IDs, fatigue levels, S_raw, F_ais, etc., are automatically recorded to the local SQLite database.
[0179] Operators start the system through a visual interface and select or enter a crew member ID. If the system recognizes the crew member for the first time, it automatically enters baseline establishment mode, displaying the message "Personal baseline collection in progress, please remain alert," and runs silently in the background for 5 minutes to collect data and calculate B_awake and σ_awake. After data collection is complete, the system automatically switches to real-time monitoring mode, displaying "Monitoring," and begins executing the complete processing and alarm pipeline described above.
[0180] This integration process deeply merges core algorithms, edge computing hardware, robot middleware, and professional human-machine interfaces, forming a stable, real-time, and deployable intelligent monitoring and early warning system for fatigue among inland waterway vessel crew members.
[0181] In summary, the method described in this invention utilizes multi-level neural networks to collaboratively process CRL-Net and RMT-3DCNN+LSTM, along with a dual-branch adaptive decision algorithm. It also incorporates a physical prior-guided enhancement strategy and a dual individual-work condition calibration mechanism. This approach can fully extract multi-source dynamic facial behavioral information of crew members under complex environments, achieving highly robust and accurate real-time fatigue state identification. This is of great significance to the field of inland waterway vessel navigation safety monitoring.
[0182] The above discussion is merely a specific illustration of the present invention. Those skilled in the art, based on their understanding of the core technical ideas of the present invention, can make various appropriate changes, simplifications or combinations thereof, and these derivative solutions all fall within the protection scope of the claims of the present invention.
Claims
1. A multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments, characterized in that, include: S1: Video enhancement data preprocessing, design a "cooperative reflection perception low light enhancement" network to remove water surface reflection interference in the image, improve video quality, and obtain high-quality video input; S2: Design an "end-to-end spatiotemporal fatigue feature" learning network to learn and integrate three key fatigue representations: pupil movement, eye state, and fatigue expression; S3: Design a "two-branch adaptive individual feature calibration and navigation scenario weighted network" to calculate the final individual fatigue state; S4: Fatigue detection system integration and visualization output. The above modules are integrated to form an end-to-end detection system, which outputs the crew's fatigue status in real time and provides visual warnings on the monitoring interface.
2. The method and system for multi-feature adaptive fatigue detection of inland waterway crew members in complex environments according to claim 1, characterized in that, In step S1, video enhancement data preprocessing involves designing a "cooperative reflectivity-based low-light enhancement network" to remove water surface reflection interference, improve image quality, and achieve integrated restoration of coupled degraded videos. This includes: (1) Physically prior guided shared encoder Each frame of the input video sequence is first processed by a transmittance-luminance joint estimation module, which simultaneously outputs two physical prior maps: Transmission rate attenuation map Initial brightness distribution map These two prior images are concatenated with the original image of the frame and simultaneously input into a multi-scale reversible encoder to generate a set of shared feature bases containing physical degradation information. ; (2) Collaborative restoration in decoupled color spaces Color space transformation: sharing feature bases The RGB portion is projected onto the improved HVI color space through an adaptive HVI transform layer; Joint restoration reasoning: Within the HVI space, a reflection-illumination co-attention module is designed to decompose HVI features into reflection-sensitive and illumination-sensitive components, and these components interact through a cross-attention mechanism. Reflection sensitive component Transmission rate map Constraints are used to separate reflected noise and light-sensitive components. Subject to initial brightness map Guidance is used to enhance the brightness and contrast of dark areas, and the two correct each other under the attention mechanism; (3) Reconstruction of physical consistency Enhanced HVI spatial features The image is mapped back to RGB space through an inverse transform, and the final output image of the frame is generated by a residual reconstruction network. The entire network is optimized end-to-end through a joint loss function to obtain a clear preprocessed video sequence. To ensure the accuracy of reflection separation, Ensure the naturalness of the enhanced lighting. Ensure overall visual quality.
3. The method and system for multi-feature adaptive fatigue detection of inland waterway crew members in complex environments according to claim 1, characterized in that, In step S2, the design of the "end-to-end spatiotemporal fatigue feature" learning network replaces the traditional separate or concatenated design of 3DCNN and Transformer with RMT, realizing unified modeling of spatiotemporal features of video sequences and end-to-end regression calculation of nine types of fatigue feature values, specifically including: The backbone network is constructed based on the RMT preservation mechanism. It converts video frame sequences into feature vector sequences through spatiotemporal encoding, and then uses a linear recursive formula to perform forward scanning processing on the sequences, establishing long-range temporal dependencies and effectively capturing long-cycle fatigue features. The core state update process is as follows: The current hidden state. As a learnable retention factor, For the current input features, It is the characteristic transformation function; In the multi-head retention module of the backbone network, a cross-head hybrid interaction mechanism for fatigue area perception is introduced, which enables different attention heads to interact and collaboratively focus on key facial areas such as the eyes and mouth, thereby enhancing the ability to perceive subtle local changes. The network's output layer connects to a multi-branch regression head, mapping the global spatiotemporal representation generated by the backbone network into a structured nine-dimensional feature vector. .
4. The method and system for adaptive fatigue detection of inland waterway crew members under complex environments according to claim 1, characterized in that, In step S3, the design of the "dual-branch adaptive individual feature calibration and navigation scenario weighted network" calculates the final individual fatigue state. The key feature is that in step S3, the individual alertness benchmark and AIS navigation scenario weighted fusion module establishes an individual crew member alertness benchmark and combines it with real-time ship automatic identification system information to perform situation-aware adaptive fatigue state calculation, outputting a personalized fatigue state. Specifically, this includes: (1) Establishment of individual alertness benchmarks: In the initial monitoring phase, video data of the current crew member is continuously collected within a preset duration. The corresponding nine-dimensional fatigue feature vector sequence is extracted, and statistical analysis is performed on this sequence to calculate the mean and variance of each dimension of the feature, forming the individual alertness baseline vector for that crew member. Its confidence interval serves as a personalized reference baseline for subsequent fatigue state assessment; (2) AIS navigation context feature extraction and fatigue-prone state identification: Real-time access to the Automatic Identification System (AIS) data stream to detect ship speed. With heading angle Calculate the standard deviation of ship speed within the most recent time window. Maximum change in heading angle ; If all conditions are met and (in and If the threshold is set (for example, if the current navigation condition is a fatigue-prone situation), then the situation flag will be triggered. Otherwise, it is a normal situation. ; (3) Calculation of fatigue state based on context perception: For the fatigue feature vector extracted at the current moment Calculate its relationship with the individual's baseline level of alertness. The absolute differences in each dimension are divided by the standard deviation of the corresponding dimension to obtain the standardized difference score, and then the weighted sum is used to obtain the comprehensive difference index. According to AIS context markers Perform adaptive threshold adjustment; when (In normal circumstances) a baseline grading threshold is used. right The fatigue levels are determined by classification. when In situations prone to fatigue, a situational sensitivity coefficient is introduced. ( ), compressing the baseline grading threshold to Then use The fatigue level is obtained by comparing it with the compressed threshold. (4) Final personalized fatigue state output: The system ultimately outputs the fatigue status level and provides corresponding system prompts.
5. A multi-feature adaptive fatigue detection method and system for inland waterway crew members in complex environments, characterized in that, include: The system focuses on the monitoring environment of the cockpit of inland waterway vessels, integrating various modules into edge computing devices, and starting to work simultaneously. Cockpit Environment Monitoring Module: The physical entity of the system deployment, which includes cameras for capturing video and an AIS data interface for receiving ship navigation data, providing real-time data input; Image enhancement preprocessing module: It has a built-in cooperative reflectance perception low-light enhancement network, which receives the raw video stream from the camera, restores low-quality images, and outputs a clear video sequence to the subsequent feature extraction module; Spatiotemporal fatigue feature extraction and calculation module: It has a built-in improved RMT backbone network, receives the enhanced video sequence, and automatically locates and outputs a structured nine-dimensional fatigue feature vector through end-to-end forward computation. Individual baseline and situation-adaptive fusion module: responsible for establishing individual alertness baselines for crew members, while parsing AIS data in real time, comparing real-time fatigue characteristics with individual baselines, and adaptively weighting them in combination with the navigation situation determined by AIS, and outputting the fused comprehensive fatigue index; Fatigue classification and alarm module: Receives the comprehensive fatigue index, determines the current fatigue level of the crew through a classifier, and triggers audible and visual alarms and log recordings according to preset strategies; Human-computer interaction and system management module: Provides a visual interface for real-time display of crew status, system operation status and alarm information.