Driver fatigue state monitoring method and system, vehicle-mounted terminal equipment and storage medium

By combining parallel 3D and 2D convolutional neural network branches with residual connections and mask reconstruction processing, the contradictions in spatiotemporal feature extraction and environmental interference problems in existing technologies are solved, and efficient and accurate monitoring of driver fatigue is achieved.

CN121564693APending Publication Date: 2026-02-24HUACHEN XINYUAN CHONGQING AUTOMOBILE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511885584.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies struggle to balance the efficiency of spatiotemporal feature extraction with computational overhead in driver fatigue monitoring, and lack robustness in the face of in-vehicle environmental interference, resulting in high rates of missed and false alarms.

Method used

We employ parallel 3D and 2D convolutional neural network branches to extract spatiotemporal features, and enhance robustness through residual connections and mask reconstruction. We also fuse multimodal information from physiological signals and visual behavioral features.

Benefits of technology

It improves the accuracy and reliability of fatigue monitoring, reduces the rate of missed and false alarms, and can accurately monitor driver fatigue, especially in complex vehicle environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564693A_ABST
    Figure CN121564693A_ABST
Patent Text Reader

Abstract

According to the driver fatigue state monitoring method and system, the vehicle-mounted terminal equipment and the storage medium, the spatial-temporal features are extracted through parallel three-dimensional and two-dimensional convolutional neural network branches, residual connection and mask reconstruction processing are combined, the problems of spatial-temporal feature extraction contradiction and environmental interference in the prior art are solved, and the driver fatigue state monitoring accuracy is improved. Time sequence dynamic features and space static features of a video sequence are efficiently extracted through three-dimensional and two-dimensional convolutional neural network branches which are arranged in parallel, and robustness under shielding or illumination interference is remarkably enhanced through residual connection and mask reconstruction processing, so that accuracy and reliability of fatigue state monitoring are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a method, system, vehicle-mounted terminal device, and storage medium for monitoring driver fatigue. Background Technology

[0002] Driver fatigue is a significant human factor contributing to traffic accidents, and its real-time and accurate monitoring is crucial for ensuring driving safety. Current mainstream technologies employ vision-based, non-contact solutions, capturing driver facial image sequences using in-vehicle cameras and analyzing apparent features such as eye closure and mouth opening / closing frequency to infer fatigue levels. However, in real-world in-vehicle scenarios, these methods have significant limitations. First, there is a structural contradiction in spatiotemporal feature extraction: while 3D convolutional neural networks can effectively model dynamic changes in video sequences, their high computational complexity leads to large inference delays, making them unsuitable for the computational constraints of automotive-grade embedded platforms; while 2D convolutional neural networks, despite their lightweight advantage, cannot fully capture the continuous temporal patterns of fatigue state evolution and have weak representation capabilities for key dynamic features such as sudden drops in blink frequency. Second, the ability to handle environmental interference is severely inadequate: drastic fluctuations in in-vehicle lighting conditions, such as sudden changes in brightness when entering or exiting tunnels, glare from glasses obscuring the eye area, or partial occlusion due to the driver wearing a mask, can all cause distortion in the feature extraction signal. Current technologies lack mechanisms for repairing damaged features, relying solely on original image enhancement or post-processing filtering. This fails to reconstruct the context of low-confidence regions at the feature level, resulting in persistently high false alarm and false negative rates under complex operating conditions. Furthermore, the depth of multimodal information fusion is limited: physiological signals and visual behavioral features are typically fused using simple weighting or decision-level splicing, failing to establish fine-grained semantic relationships between physiological states and visual representations. This shallow fusion strategy disrupts the inherent synergy of multi-source data, weakening not only the early warning capability of fatigue states but also reducing the system's precision in classifying progressive fatigue processes. These issues collectively restrict the practicality of monitoring technology under limited onboard computing power, necessitating breakthroughs in precise temporal modeling, robust anti-interference, and deep fusion of multi-source features. Summary of the Invention

[0003] This invention provides a method for monitoring driver fatigue, addressing the problem of balancing spatiotemporal feature extraction efficiency with computational overhead in existing technologies. This results in high computational demands for 3D convolutional networks, while 2D convolutional networks lack the ability to capture core fatigue features. Furthermore, the system exhibits severely insufficient robustness in the face of common in-vehicle interferences such as strong sunlight, glare from glasses, and facial obstruction, lacking intelligent feature-level repair capabilities and prone to false negatives or missed detections.

[0004] To achieve the above objectives, the present invention provides the following technical solution: A method for monitoring driver fatigue includes: Obtain a continuous video sequence of the driver; The video sequence is input into a feature extraction network, which includes a parallel 3D convolutional neural network branch and a 2D convolutional neural network branch; the 3D convolutional neural network branch extracts the dynamic features of the video sequence in the temporal dimension; and the 2D convolutional neural network branch extracts the static features of the single frame image of the video sequence in the spatial dimension. The feature extraction network includes at least one intermediate layer, through which the dynamic features are pooled along the time dimension to obtain pooled features. The pooling features and the static features are joined across branches using residual connections to generate fused features; The fused features are subjected to mask reconstruction processing to enhance their robustness under occlusion or lighting interference, resulting in the processed fused features. Based on the processed fusion features, the driver's fatigue state is determined.

[0005] A driver fatigue monitoring system includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the driver fatigue monitoring method of the present invention.

[0006] An in-vehicle terminal device includes: an image acquisition device for acquiring a continuous video sequence of a driver; and the driver fatigue monitoring system of the present invention.

[0007] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the driver fatigue monitoring method of the present invention.

[0008] Beneficial effects This invention provides a driver fatigue monitoring method, system, vehicle-mounted terminal device, and storage medium. By extracting spatiotemporal features through parallel three-dimensional and two-dimensional convolutional neural network branches, combined with residual connections and mask reconstruction processing, it solves the problems of contradictions in spatiotemporal feature extraction and environmental interference in the prior art. The parallel three-dimensional and two-dimensional convolutional neural network branches efficiently extract the temporal dynamic features and spatial static features of video sequences, and the residual connections and mask reconstruction processing significantly enhance the robustness under occlusion or lighting interference, thereby improving the accuracy and reliability of fatigue monitoring. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0010] Figure 1 This is a flowchart of a driver fatigue monitoring method according to an embodiment of the present invention; Figure 2 This is a flowchart of a driver fatigue monitoring method according to another embodiment of the present invention; Figure 3 This is a schematic diagram of the driver fatigue monitoring system according to an embodiment of the present invention.

[0011] Structural diagram. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0014] For ease of understanding, some key terms in the embodiments of the present invention are explained below: A continuous video sequence refers to a set of image frames captured continuously over a period of time by an image acquisition device, containing the driver's face and upper body. This sequence is used to provide temporal information about the driver's behavior.

[0015] Feature extraction network: refers to a deep learning model designed to automatically learn and extract discriminative features from input video sequences. In this embodiment of the invention, the network includes a parallel structure to process spatiotemporal information simultaneously.

[0016] The 3D convolutional neural network branch refers to a module in the feature extraction network specifically designed to process temporal information from video sequences. It captures dynamic changes in driver behavior by performing convolutional operations in both spatial and temporal dimensions.

[0017] Two-dimensional convolutional neural network branch: This refers to a module in the feature extraction network specifically designed to process the spatial information of a single frame of a video sequence. It captures static details of the driver's facial features by performing convolution operations on the width and height dimensions of the image.

[0018] Dynamic features: These refer to the feature representations that reflect the changes in driver behavior over time, extracted from video sequences through branches of a three-dimensional convolutional neural network.

[0019] Static features: refers to the spatial structural features that reflect the driver's facial or body posture, extracted from a single frame of a video sequence through a two-dimensional convolutional neural network branch.

[0020] Intermediate layers: These refer to one or more network layers in the feature extraction network located between the branch outputs of the three-dimensional convolutional neural network and the pooling operation, used for preliminary processing or dimensional adjustment of dynamic features.

[0021] Pooled features: These are feature representations obtained by processing dynamic features through an intermediate layer and performing pooling operations along the time dimension. They effectively reduce feature dimensionality while retaining key temporal information.

[0022] Residual connection: refers to a common connection method in deep learning networks, which directly adds the input of a layer to its output to alleviate the gradient vanishing problem and promote the fusion of features from different layers or branches.

[0023] Fusion features: refers to the feature representation generated by connecting pooled features and static features through residuals across branches, which integrates the spatiotemporal information of driver behavior.

[0024] Mask reconstruction processing: refers to a post-processing operation performed on fused features, which aims to identify and repair unreliable areas in the features caused by occlusion or lighting interference, thereby improving the robustness of the features.

[0025] Processed fused features: refers to the feature representation obtained after the fused features have been processed by mask reconstruction, which has higher reliability and anti-interference ability.

[0026] Fatigue state: refers to the state of decline in physiological and psychological function of a driver due to prolonged driving, lack of sleep or other reasons. The embodiments of the present invention aim to quantify or classify it.

[0027] like Figure 1 As shown, this application proposes a method for monitoring driver fatigue, including: S101: Obtain a continuous video sequence of the driver; Specifically, the video sequence can be captured by a camera installed inside the vehicle. In practical applications, a wide-angle camera fixed above the dashboard can be used to continuously record the driver's face and upper body. Alternatively, an infrared camera can be used to capture data in low-light environments to ensure clear image data even at night or in tunnels. This continuous video sequence can be segmented at regular intervals to avoid additional computing power overhead.

[0028] S102: Input the video sequence into a feature extraction network, which includes a parallel three-dimensional convolutional neural network branch and a two-dimensional convolutional neural network branch; extract the dynamic features of the video sequence in the temporal dimension through the three-dimensional convolutional neural network branch; and extract the static features of the single frame image of the video sequence in the spatial dimension through the two-dimensional convolutional neural network branch. Specifically, the 3D convolutional neural network branch can employ existing 3D CNN or C3D network structures, which capture the temporal dependencies between video frames by stacking multiple 3D convolutional layers and 3D pooling layers. Simultaneously, the 2D convolutional neural network branch can employ existing 2D CNN or ResNet-18 network structures, which extract spatial features of single-frame images through 2D convolutional layers and residual blocks. Thus, the 3D convolutional neural network branch is used to extract dynamic features of the continuous video sequence in the temporal dimension, such as the driver's blinking, head turning, or yawning. Meanwhile, the 2D convolutional neural network branch is used to extract static features of single-frame images in the spatial dimension, such as the driver's eye opening and closing, mouth shape, or facial expressions. Through a parallelized spatiotemporal feature extraction network, a balance is effectively struck between computational overhead and feature capture capability.

[0029] S103: The feature extraction network includes at least one intermediate layer. The dynamic feature is passed through the intermediate layer and pooled along the time dimension to obtain pooled features. Specifically, this intermediate layer can be one or more convolutional layers, fully connected layers, or batch normalization layers from existing technologies, configured after the output of the branches of the three-dimensional convolutional neural network. The dynamic features, after passing through this intermediate layer, are pooled along the time dimension to obtain pooled features. In practical applications, average pooling or max pooling operations can be used to downsample the dynamic features within a preset time window to reduce feature dimensionality and retain key temporal information.

[0030] S104: Perform a cross-branch residual connection between the pooling feature and the static feature to generate a fused feature; Specifically, the dimensions of pooling features and static features are adjusted to be consistent before concatenation. In practice, this can be achieved using a linear projection layer or a convolutional layer to adjust the dimension of one feature to match that of the other. Subsequently, these two features are element-wise added to form a fused feature. This residual connection method helps preserve the original feature information and promotes the effective fusion of features from different modalities.

[0031] S105: Perform mask reconstruction on the fused feature to enhance its robustness under occlusion or lighting interference, and obtain the processed fused feature. Specifically, a mask with a fixed shape can be predefined. When the driver's face is detected to be obstructed or the local area is too bright, the mask is applied to the corresponding area of ​​the fused features, and simple interpolation or filling operations are performed on the feature values ​​of these areas to try to recover the disturbed feature information.

[0032] In other implementations, such as Figure 2 As shown, step S105, the mask reconstruction process for the fused feature includes: S501: Based on the brightness gradient and local texture changes of the current image frame of the video sequence, a soft mask matrix with the same spatial size as the fused feature is generated. The element values ​​in the soft mask matrix are used to characterize the credibility of the corresponding position in the fused feature. S502 suppresses low-confidence regions in the fused feature based on the soft mask matrix, and reconstructs the features of the low-confidence regions based on the context information of the high-confidence regions in the fused feature.

[0033] Specifically, in step S501, based on the brightness gradient and local texture changes of the current image frame in the video sequence, a soft mask matrix with the same spatial size as the fused feature is generated. The element values ​​in this soft mask matrix are used to characterize the reliability of the corresponding position in the fused feature, which means quantifying the reliability of the feature by analyzing the visual characteristics of the current image frame in the video sequence. The brightness gradient can reflect the drastic degree of illumination changes in the image, such as strong light illumination or shadow areas, where features are often disturbed. Local texture changes can indicate the presence of occlusions in the image, such as reflections from glasses or hand occlusions, where features may also be distorted. By calculating these visual characteristics, a soft mask matrix matching the spatial size of the fused feature can be generated, where each element value represents the reliability of the corresponding position of the fused feature. In practical applications, existing edge detection algorithms such as the Sobel operator and Canny operator can be used to calculate the brightness gradient, or methods such as Local Binary Pattern (LBP) and Gabor filters can be used to analyze local texture. These calculation results can be normalized to a range of 0 to 1, where higher values ​​represent higher confidence and lower values ​​represent lower confidence. Alternatively, a lightweight convolutional neural network can be trained to directly predict the soft mask matrix using the current image frame as input. This network can learn the complex relationship between lighting, occlusion, and other interference patterns and feature confidence.

[0034] In step S502, suppressing low-confidence regions in the fused feature based on the soft mask matrix means using the low-confidence information indicated in the soft mask matrix to reduce or eliminate the influence of interfering regions in the fused feature. For example, suppression can be achieved by multiplying the soft mask matrix element-wise with the fused feature. The lower the element value of the soft mask matrix, the lower the confidence level, and the smaller the contribution of the corresponding position in the fused feature. In this way, the weight of the feature information of the interfering region in subsequent processing is effectively reduced. Another implementation method is to set a confidence threshold and directly set the regions in the soft mask matrix below the threshold to zero, thereby completely eliminating the feature influence of these low-confidence regions.

[0035] Based on the contextual information of the high-confidence region in the fused features, the feature reconstruction of the low-confidence region refers to using the semantic information of the still reliable high-confidence region in the fused features to infer and restore the features of the suppressed region after suppressing the low-confidence region. In practical applications, diffusion-based or learning-based methods commonly used in existing image inpainting techniques can be adopted. Diffusion-based methods can use interpolation, mean filling, etc., to fill the low-confidence region with features of neighboring high-confidence regions. Learning-based methods can train a generative model, such as an autoencoder or generative adversarial network, to learn how to generate low-confidence region features consistent with the surrounding environment, using the features of the high-confidence region as input. Through the implementation of the above technical solution, this application can intelligently identify and quantify the credibility of driver facial features in complex in-vehicle environments, such as under strong light illumination, facial occlusion, and other interference, and can accurately locate the affected feature regions. By generating a soft mask matrix, this application achieves adaptive suppression of low-confidence regions in the fused features, effectively reducing the negative impact of noise and interference on fatigue state judgment. Meanwhile, by utilizing contextual information from high-confidence regions to reconstruct features from low-confidence regions, intelligent repair can be performed at the feature level. Even with missing or distorted information, a more complete and accurate representation of the driver's facial features can be recovered. This significantly enhances the robustness of fatigue monitoring in real-world, complex environments, reduces false alarms and false negatives, and thus improves the accuracy and reliability of driver fatigue monitoring.

[0036] In another embodiment, in step S502, the process of reconstructing the fused feature based on the soft mask matrix is ​​achieved by the following formula: F' = M⊙F + (1-M)⊙G(F) where F is the fused feature, M is the soft mask matrix, ⊙ represents element-wise multiplication, G(F) represents the feature reconstruction function based on context information, and F' is the reconstructed feature.

[0037] Specifically, F in the formula represents the fused feature generated in the feature extraction network by performing a cross-branch residual connection between the pooled feature and the static feature. This fused feature contains the dynamic features of the video sequence in the temporal dimension and the static features of a single frame image in the spatial dimension, which is the basis for subsequently determining the driver's fatigue state.

[0038] M represents the soft mask matrix, which is generated based on the brightness gradient and local texture changes of the current image frame in the continuous video sequence, and has the same spatial size as the fusion feature F. The element values ​​in this soft mask matrix M characterize the confidence level of the corresponding position in the fusion feature F; for example, element values ​​close to 1 represent high-confidence regions, and element values ​​close to 0 represent low-confidence regions. In the formula, the M⊙F part directly preserves the feature information of high-confidence regions in the fusion feature F through element-wise multiplication, avoiding modification or unnecessary processing of reliable information.

[0039] The symbol ⊙ represents element-wise multiplication, a matrix or tensor operation in existing technology used to multiply corresponding elements of two matrices or tensors of the same size. In this formula, it ensures that the soft mask matrix M can be precisely applied to each corresponding position of the fused feature F, achieving weighted or selective preservation of features from different regions.

[0040] G(F) represents a context-based feature reconstruction function. Its function is to intelligently reconstruct features of low-confidence regions by utilizing the contextual information of high-confidence regions in the fused feature F. This function can learn and infer reasonable feature representations of occluded or disturbed regions. In practical applications, G(F) can be a lightweight convolutional neural network that takes the fused feature F as input and learns the spatial dependencies of features through multiple convolutions and nonlinear activation functions, thereby predicting and generating features of low-confidence regions. Another implementation is that G(F) can employ an attention-based module. By calculating similarity weights between low-confidence and high-confidence regions, it weights and aggregates the feature information of high-confidence regions to reconstruct the features of low-confidence regions.

[0041] F' represents the reconstructed feature, which is the fused feature after mask reconstruction. This feature retains the original information of high-confidence areas while intelligently repairing low-confidence areas, thereby enhancing its robustness under occlusion or lighting interference.

[0042] S106: Based on the processed fusion features, determine the driver's fatigue state.

[0043] Specifically, the processed fused features are input into a classifier or regressor. In practical applications, this could be a multilayer perceptron (MLP), which outputs a discrete category representing the fatigue level or a continuous fatigue score. This classifier or regressor learns the mapping from the fused features to the fatigue state by being pre-trained on a large amount of fatigue data.

[0044] In other implementations, such as Figure 2 As shown, after step S106, the following steps are also included: S107: Extract physiological signal features reflecting chest cavity movement from the video sequence; S108: Input the physiological signal features and the visual behavior features extracted based on the fused features into different channels of the dual-channel Transformer encoder, respectively, and use the cross-attention mechanism to perform feature alignment and fusion to generate multimodal joint features; S109: Input the multimodal joint features into the regression network to output continuous fatigue level scores.

[0045] Specifically, in step S107, when extracting physiological signal features reflecting chest cavity movement from the continuous video sequence, the driver's chest cavity movement, such as respiratory rate, respiratory depth, and rhythm, is an important indicator reflecting their physiological state and is closely related to fatigue level. Extracting these physiological signal features through non-contact video analysis can provide physiological evidence independent of facial visual behavior for fatigue monitoring. In practical applications, an optical flow-based approach can be used to estimate the periodic movement of the chest cavity by analyzing the optical flow changes of pixels in the chest cavity region within the video sequence, thereby calculating the respiratory rate and amplitude. The driver's chest cavity region is selected within the video frame, and the optical flow vector of pixels within that region changes over time. Temporal analysis of the projection of the optical flow vector in the vertical direction yields the chest cavity movement curve. Alternatively, a deep learning-based approach can be used, training a dedicated convolutional neural network or recurrent neural network model to directly learn and extract chest cavity movement features from the video frames. This model can identify minute deformations in the chest cavity region and convert them into physiological signal waveforms, thereby calculating physiological parameters such as respiratory rate and respiratory depth.

[0046] In step S108, when the physiological signal features and the visual behavioral features extracted based on the fusion features are input into different channels of the dual-channel Transformer encoder, the physiological signal features, such as respiratory rate and depth, and the visual behavioral features, such as eye state, head posture, and facial expression, represent two different modalities of information, each reflecting different aspects of the driver's fatigue state. To fully utilize these two types of information and avoid mutual interference during early fusion, this step employs a dual-channel Transformer encoder for independent processing and preliminary encoding. The physiological signal features refer to the numerical values ​​or vectors extracted from the chest movement video sequence that characterize the driver's physiological state, such as respiratory rate, respiratory depth, and respiratory rhythm variability. The visual behavioral features refer to the facial and head behavioral features related to driver fatigue further extracted from the fusion features processed above, such as blinking frequency, eye closure duration, head drooping angle, and yawning frequency. The dual-channel Transformer encoder is a Transformer architecture with two independent input paths, each channel specifically designed to process a feature sequence of one modality. In practical applications, this dual-channel Transformer encoder uses one channel to receive physiological signal feature sequences and the other channel to receive visual behavioral feature sequences. Each channel contains a self-attention mechanism and a feedforward network, enabling it to capture the temporal dependencies and feature representations within its respective modality. Inputting the two types of features into different channels means that physiological signal features and visual behavioral features are fed into two independent, parallel input paths when entering the Transformer encoder. This design allows features from both modalities to undergo initial feature learning and encoding within their respective channels, maintaining their modality specificity and avoiding feature confusion or information loss due to significant modality differences before fusion.

[0047] When using cross-attention mechanisms for feature alignment and fusion to generate multimodal joint features, the cross-attention mechanism is an attention mechanism that allows information interaction between different sequences. In this embodiment, it is used to achieve deep interaction and alignment between physiological signal features and visual behavior features after they have been encoded through their respective channels, thereby generating a comprehensive joint feature that reflects multimodal correlation information. The cross-attention mechanism allows a query vector of one modality to focus on the key and value vectors of another modality. For example, visual behavior features can be used as a query to focus on the key and value of physiological signal features, thereby learning the parts of the physiological signal related to visual behavior; and vice versa. This mechanism enables features of two modalities to refer to and complement each other, capturing their potential correlations and synergies, thereby achieving feature alignment. Feature fusion refers to integrating the aligned and interacting physiological signal features and visual behavior features into a unified feature representation. This fusion is not a simple splicing, but a weighted summation through the cross-attention mechanism, so that the fused feature can simultaneously contain the key information of both modalities, and this information is mutually verified and reinforced. The resulting multimodal joint features are a comprehensive feature representation that integrates physiological signal features and visual behavioral features. It contains in-depth information about the driver's fatigue state in both physiological and behavioral dimensions, and can more comprehensively and accurately characterize the driver's fatigue level.

[0048] In step S109, when the multimodal joint features are input into the regression network to output a continuous fatigue level score, the regression network is a neural network model that can output continuous numerical values, unlike classification networks that output discrete categories. Inputting the multimodal joint features into the regression network can directly predict the driver's fatigue level, providing a continuous and refined fatigue level score, rather than a simple binary judgment of "fatigued" or "not fatigued." This regression network can be a multilayer perceptron (MLP), convolutional neural network (CNN), or recurrent neural network (RNN), with its last layer typically being a linear activation function to output continuous numerical values. The network predicts the driver's fatigue level by learning the mapping relationship between the multimodal joint features and the actual fatigue level score. A continuous fatigue level score refers to a floating-point value within a preset range, such as 0 to 100, or 0 to 1, used to quantify the driver's fatigue level. For example, 0 points represents complete alertness, and 100 points represents extreme fatigue. This continuous score provides a more detailed description of fatigue status, facilitating graded early warning and personalized intervention.

[0049] In other embodiments, after step S109, the fatigue level score is further stabilized, including: S901: aggregating the fatigue level scores of multiple consecutive time frames of the video sequence into a state queue; S902: calculating the smoothed fatigue level score at the current time point based on the changing trend of the scores in the state queue and / or the proportion of abnormal scores.

[0050] Specifically, stabilizing the fatigue level score involves further processing the output raw fatigue level score to eliminate transient noise, improve the continuity and stability of the score, and thus make the final fatigue state judgment more reliable. This processing aims to avoid frequent switching or misjudgment of fatigue state caused by occasional interference or minor fluctuations in the model.

[0051] In practical applications, digital filtering techniques, such as existing low-pass filters or Kalman filters, can be used to process continuous fatigue level scores to filter out high-frequency noise and preserve the long-term trend of fatigue status. Alternatively, a rule-based state machine can be designed to determine the final steady-state output based on the combination of historical and current scores.

[0052] Aggregating fatigue level scores from multiple consecutive time frames of the video sequence into a state queue means collecting and storing fatigue level scores generated over a period of time, such as the most recent N seconds or N frames, in an ordered data structure to form a time series. This state queue provides the necessary historical context information for subsequent steady-state processing, enabling the system to determine fatigue state based on overall performance over a period of time rather than a single instantaneous value.

[0053] In practical applications, a fixed-length First-In-First-Out (FIFO) queue can be used. Whenever a new fatigue level score is generated, it is added to the end of the queue, and the oldest score is removed, ensuring that the queue always contains scores from the latest N timeframes. Another approach is to use weighted historical storage, assigning different weights to scores at different points in time in the queue; for example, more recent scores have higher weights, to more precisely reflect the impact of recent conditions.

[0054] Calculating a smoothed fatigue level score for the current time point based on the changing trends of scores and / or the proportion of abnormal scores in the state queue involves using historical data accumulated in the state queue to analyze dynamic change patterns in scores and identify outlier data points to generate a more stable and accurate fatigue level score. This method comprehensively considers the overall trend of scores and the interference of outliers, thus outputting a smoothed value that better reflects the driver's true fatigue state.

[0055] In practical applications, trend analysis can be performed on the scores in the status queue first, such as predicting future trends through methods like linear regression or exponential smoothing. Simultaneously, the proportion of abnormal scores exceeding a preset threshold—for example, scores deviating from the queue average by more than a certain standard deviation—can be calculated. The final smoothed fatigue level score can be calculated by combining trend prediction and the proportion of abnormal scores. For instance, if the trend indicates worsening fatigue and the proportion of abnormal scores is low, the smoothed score will be closer to the current high score; if the proportion of abnormal scores is high, the smoothed score will be more inclined towards the queue average or the smoothed value from the previous moment. Alternatively, the exponential moving average (EMA) algorithm can be used to weight the scores in the queue, where the weights decay exponentially over time, making recent scores have a greater impact on the result while maintaining smoothness.

[0056] In another embodiment, in step S902, the smoothed fatigue level score at the current time point is calculated using an exponential moving average algorithm, and the smoothing coefficient is dynamically adjusted according to the real-time motion state of the vehicle or the light stability of the carriage.

[0057] Specifically, the exponential moving average algorithm is an existing time series data smoothing technique. Its core principle is to assign higher weights to recent data, while the weights of historical data decay exponentially. This algorithm can effectively filter out short-term noise while preserving the long-term trend of the data, thus enabling the smoothed fatigue level score to more stably reflect the driver's true fatigue state.

[0058] In practical applications, the existing EMA formula can be used for calculation: EMA_t = α * Current_Value_t + (1 - α) * EMA_{t-1}, where α is the smoothing coefficient, Current_Value_t is the fatigue level score at the current moment, and EMA_{t-1} is the smoothed score at the previous moment. Alternatively, variant algorithms such as double exponential moving average or adaptive exponential moving average can be used to further optimize the smoothing effect.

[0059] In this step, the dynamic adjustment of the smoothing coefficient aims to intelligently adapt the smoothing process of fatigue level scoring to complex driving environments. Real-time vehicle motion can be acquired through onboard sensors. For example, when the vehicle's acceleration sensor detects severe bumps or sharp turns, it indicates the vehicle is in an unstable state. In this case, the smoothing coefficient can be increased to enhance the smoothing effect and suppress visual feature fluctuations caused by vehicle swaying. When the vehicle is traveling smoothly, the smoothing coefficient can be decreased to respond more quickly to real changes in driver fatigue. The stability of cabin lighting can be evaluated using image brightness information obtained from in-vehicle lighting sensors or image acquisition devices. For example, when the rate of change in light intensity is large, such as when entering or exiting tunnels or under direct sunlight, it indicates unstable lighting. In this case, the smoothing coefficient can be increased to reduce the impact of lighting changes on visual feature extraction. When the lighting is stable, the smoothing coefficient can be decreased. Through this dynamic adjustment mechanism, the system can adaptively optimize the smoothing intensity according to environmental changes, ensuring the accuracy and environmental adaptability of fatigue level scoring.

[0060] In other embodiments, a dynamic sensitivity adjustment step is also included. This step can be initiated before or after any step of the driver fatigue state monitoring method of the present invention. In this embodiment, the dynamic sensitivity adjustment is initiated after step S109. The dynamic sensitivity adjustment step includes: S110: dynamically switching between a preset high sensitivity mode, a medium sensitivity mode and a low sensitivity mode based on the current driving conditions of the vehicle, the lighting conditions in the vehicle compartment or the stability information of the driver's facial image; wherein, different sensitivity modes correspond to different feature extraction intensities in the feature extraction network, different reconstruction intensities in the mask reconstruction process, or different judgment thresholds in the process of determining fatigue state.

[0061] Specifically, in step S110, the vehicle's current driving condition refers to its state while operating on the road, such as speed, acceleration, steering angle, road type, and traffic density. These conditions affect the driver's physiological and psychological state and may indirectly affect the image acquisition quality. Obtaining this information helps determine the complexity and potential interference of the current driving environment. In practical applications, vehicle speed, acceleration, steering angle, and other data can be obtained through the vehicle's CAN bus; or combined with GPS positioning and map data to determine the current road type and traffic conditions.

[0062] The lighting conditions inside the carriage refer to the intensity and uniformity of light, as well as the presence of strong light or backlighting. Lighting conditions directly affect the quality of visual images; excessive brightness can lead to overexposure, while excessive darkness can result in underexposure. Backlighting or strong light can create shadows or reflections on the driver's face, thus interfering with feature extraction. In practical applications, lighting conditions can be assessed through image sensor exposure parameters, image histogram analysis, or dedicated lighting sensors; or using image processing algorithms such as mean brightness, standard deviation, and highlight area ratio.

[0063] The stability information of the driver's facial image refers to the sharpness of the driver's face in the video sequence, the magnitude of pose changes, and the confidence level of key feature point detection. Unstable facial images may arise from driver head movement, vehicle bumps, image blurring, or occlusion, all of which reduce the accuracy of feature extraction. In practical applications, stability can be evaluated by calculating the average displacement or variance of facial key point positions between consecutive frames; or by using the confidence score output by the facial detection algorithm, or by using an image sharpness evaluation algorithm to quantify image quality.

[0064] The preset high-sensitivity, medium-sensitivity, and low-sensitivity modes are predefined operating states with different parameter configurations. High-sensitivity mode typically means a faster and more detailed response to fatigue signs, but may increase the risk of false alarms; low-sensitivity mode, on the other hand, prioritizes stability and reduces false alarms, but may delay warnings. Medium-sensitivity mode provides a balance. In practical applications, a fixed set of parameters can be configured for each mode during initialization, such as the weights of the feature extraction network, the threshold for mask reconstruction, and the fatigue determination threshold; or the parameters of these modes can be trained and optimized offline, adjusted and validated for datasets under different environmental conditions.

[0065] This dynamic switching refers to the automatic, real-time transition between the aforementioned preset sensitivity modes based on real-time environmental assessment information. This switching ensures adaptability to constantly changing driving environments and always operates with the optimal configuration. For example, a rule engine can be used to trigger mode switching when environmental indicators reach a specific threshold; or a state machine or reinforcement learning model can be employed to intelligently select the best mode based on the current environmental state and historical monitoring results.

[0066] Different sensitivity modes correspond to different feature extraction intensities in the feature extraction network, different reconstruction intensities in the mask reconstruction process, or different judgment thresholds in the process of determining fatigue state.

[0067] Specifically, under different sensitivity modes, the feature extraction network can adjust its ability to extract fatigue-related features from the input video sequence. In practical applications, in high-sensitivity mode, the depth or width of the feature extraction network can be increased, or its activation function and regularization parameters can be adjusted to enable it to capture more subtle fatigue features; or the weights or biases of certain layers in the feature extraction network can be adjusted to make it focus more on certain regions or features in a specific mode.

[0068] This mask reconstruction process aims to correct feature loss caused by occlusion or lighting interference. The aggressiveness of the reconstruction can be adjusted under different sensitivity modes. In practical applications, in high-sensitivity mode, the number of iterations of the reconstruction network can be increased, or the weights of the reconstruction loss function can be adjusted to more aggressively repair low-confidence regions; alternatively, the soft mask matrix generation strategy can be adjusted to make it more effective at suppressing low-confidence regions in specific modes, or to make fuller use of contextual information.

[0069] Determining fatigue status typically involves comparing extracted features with preset thresholds. These thresholds can be adjusted under different sensitivity modes to balance false positives and false negatives. In practical applications, a set of fatigue level rating thresholds is preset for each sensitivity mode; for example, in a high-sensitivity mode, a lower rating may trigger a fatigue warning. Alternatively, the thresholds can be dynamically adjusted based on historical data and expert experience, or optimized using a reinforcement learning model.

[0070] like Figure 3 As shown, this application also discloses a driver fatigue monitoring system 200, including a memory 202 and a processor 201. The memory 202 stores a computer program 203. When the computer program 203 is executed by the processor 201, the driver fatigue monitoring method of this application is implemented.

[0071] Specifically, memory 202 is a hardware unit used to store data and computer programs. In the vehicle terminal, memory 202 can take various forms. In practical applications, it can be volatile memory, such as dynamic random access memory (DRAM), used for temporary storage of runtime data and program instructions, or non-volatile memory, such as NAND Flash, embedded multimedia card eMMC, or automotive-grade solid-state drive (SSD), used for persistent storage of the operating system, application programs, and model parameters. Its main function is to provide stable storage space for computer programs, ensuring that the logic and data of the driver fatigue monitoring method can be reliably saved and loaded.

[0072] Processor 201 is the core computing unit that executes computer program instructions, processes data, and controls system operations. In an automotive environment, processor 201 is typically a high-performance, low-power automotive-grade system-on-a-chip (SoC), which may integrate multiple computing cores such as a central processing unit (CPU), graphics processing unit (GPU), neural network processor (NPU), or digital signal processor (DSP). The processor is responsible for parsing and executing the computer program in memory, driving the entire fatigue monitoring method, including complex feature extraction, data fusion, and state determination computational tasks.

[0073] Computer program 203 is a pre-written set of instructions used to guide the processor to complete a specific task. In this application, computer program 203 encapsulates all the logic of a driver fatigue state monitoring method, including all steps from video sequence acquisition, feature extraction, mask reconstruction processing, multimodal fusion to final fatigue state determination. This program can be part of firmware, an operating system, or run as a standalone application, and is designed to efficiently utilize the computing resources of the in-vehicle terminal.

[0074] When the computer program 203 is executed by the processor 201, the driver fatigue monitoring method of this application is implemented. This describes the working mechanism of the system, namely, by executing the computer program 203 in the memory 202 through the processor 201, the abstract fatigue monitoring method is transformed into a practical and operable function. The processor 201 reads data and instructions from the memory 202 according to the program instructions, performs calculations, and stores or outputs the results. This hardware and software collaborative approach enables complex algorithms to run efficiently and in real-time on the vehicle terminal, thereby achieving accurate monitoring of driver fatigue.

[0075] This application proposes an in-vehicle terminal device, including: an image acquisition device for acquiring a continuous video sequence of the driver; and a driver fatigue monitoring system 200 of this application.

[0076] Specifically, this vehicle-mounted terminal device is an integrated hardware platform designed and deployed specifically for the vehicle environment. Its main function is to provide a stable and reliable operating platform for driver fatigue monitoring, and to integrate the necessary data acquisition and processing modules. This device typically possesses characteristics adapted to the vehicle environment, such as vibration resistance, wide temperature range tolerance, and low power consumption design, and can seamlessly integrate with the vehicle's power system and communication network, ensuring stable operation even under complex driving conditions.

[0077] This image acquisition device is a key component for acquiring continuous video sequences of the driver. To address the complex and variable lighting conditions and potential occlusions in the in-vehicle environment, the device can be implemented using various existing technologies. For example, a CMOS image sensor equipped with wide dynamic range (WDR) can be used to effectively handle high-contrast scenes, ensuring clear facial images of the driver are captured in both bright and low-light conditions. Furthermore, the device can integrate an infrared (IR) supplemental light to provide stable illumination at night or in low-light conditions, guaranteeing image quality. To further enhance robustness, the image acquisition device can consist of multiple cameras, working collaboratively from multiple perspectives to reduce facial occlusion caused by the steering wheel, hands, or other objects, thereby providing a more complete and continuous video data stream.

[0078] The driver fatigue monitoring system 200 is the core processing unit of the vehicle-mounted terminal equipment. Its function is to receive continuous video sequences provided by the image acquisition device and perform fatigue state analysis and judgment based on these video data. The system typically includes one or more processors 201, a memory 202, and a computer program 203 running within them, which implements the aforementioned driver fatigue monitoring method. By integrating the image acquisition and monitoring system 200 into the same vehicle-mounted terminal equipment, data transmission latency can be effectively reduced, system response speed can be improved, and the real-time performance and accuracy of data processing can be ensured.

[0079] This application proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the driver fatigue monitoring method of this application.

[0080] Specifically, the computer-readable storage medium refers to a physical or logical carrier capable of storing digital data and readable by a computer. In practical applications, this medium can be, for example, a non-transitory storage medium such as a hard disk drive, solid-state drive, flash memory (e.g., USB flash drive, SD card), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), optical discs (CD-ROM, DVD, Blu-ray disc), etc. This storage medium provides persistent storage space for driver fatigue monitoring methods, ensuring that the program remains intact after the vehicle system is powered off and is always available for the processor to read and execute.

[0081] This computer program refers to a collection of instructions designed to instruct a computer to perform specific tasks or operations. The program can exist as a compiled program, such as a binary executable file or bytecode generated from compiling C / C++ or Java, or as an interpreted program, such as a script file from Python or JavaScript. This computer program encapsulates all the logic and algorithms of the driver fatigue monitoring method and is the core of implementing the method's functionality.

[0082] When a computer program is executed by a processor, the processor—such as a central processing unit (CPU), graphics processing unit (GPU), or dedicated AI acceleration chip (NPU)—operates according to the instruction sequence defined in the computer program to achieve the function described by the program. This includes a series of operations such as the processor reading instructions from storage media, decoding instructions, executing instructions, and writing the results back to memory. In an automotive environment, dedicated hardware units can also be used to accelerate specific computing tasks, such as inference of deep learning models. Through processor execution, the static computer program is transformed into a dynamic computational process, driving the entire fatigue monitoring method.

[0083] The method for monitoring driver fatigue refers to the complete reproduction and execution of all steps and functions related to driver fatigue monitoring through a computer program. This includes, but is not limited to, acquiring continuous video sequences of the driver, extracting temporal dynamic features and spatial static features using a feature extraction network, pooling the dynamic features, performing cross-branch residual connections between the pooled features and the static features to generate fused features, performing mask reconstruction processing on the fused features to enhance robustness, and determining the driver's fatigue state based on the processed fused features. During execution, the program can be executed sequentially according to the method steps, or it can be executed efficiently and accurately in an in-vehicle environment through modular calls or parallel processing.

[0084] The above technical solution will be explained in more detail below through a more specific embodiment: In an intelligent driving assistance system, an onboard terminal device continuously monitors the driver's fatigue level. The image acquisition unit within this device acquires continuous video sequences of the driver's face in real time, for example, at a frequency of 30 frames per second.

[0085] The acquired video sequence is input into a feature extraction network. This network comprises a 3D convolutional neural network (CNN) branch and a 2D CNN branch, operating in parallel. The 3D CNN branch processes consecutive video frames, for example, 16 frames per time segment, extracting temporal dynamic features such as the driver's blink frequency, head posture changes, and micro-expressions. This processing method can capture the unique temporal evolution patterns characteristic of fatigue states. Simultaneously, the 2D CNN branch processes each frame of the video sequence, extracting spatial static features of the driver's facial organs, such as the shape, position, and texture of the eyes and mouth. Through this parallel architecture, the system, with limited onboard computing power, can effectively capture key temporal information about fatigue while also considering spatial details, solving the problem of balancing spatiotemporal feature extraction performance and computational cost in existing technologies.

[0086] The dynamic features extracted by this 3D convolutional neural network branch are then pooled in the temporal dimension through an intermediate layer in the feature extraction network, such as an average pooling layer in the temporal dimension, to obtain a pooled feature with a lower dimensionality. This step helps to further reduce data redundancy and improve the efficiency of subsequent processing.

[0087] Subsequently, the pooled feature is joined with the static feature extracted by the branch of the two-dimensional convolutional neural network via a cross-branch residual connection. This connection method allows for deep fusion of feature information from two different dimensions, generating a fused feature containing rich spatiotemporal information, ensuring the integrity and complementarity of spatiotemporal information.

[0088] To enhance the system's robustness in complex in-vehicle environments, the generated fused features undergo mask reconstruction. Specifically, the system first generates a soft mask matrix with the same spatial size as the fused features, based on the brightness gradient and local texture changes of the current video sequence image frames. The element values ​​in this matrix represent the confidence level of corresponding locations in the fused features. For example, when the driver's face is obscured by a hand or there is glare from glasses, the element values ​​in the corresponding region of the soft mask matrix will be lower, indicating low confidence in that region; while the element values ​​in clear facial areas will be higher. Next, the system suppresses low-confidence regions in the fused features based on the soft mask matrix and reconstructs the features of low-confidence regions based on the contextual information of high-confidence regions in the fused features. This process is achieved using the formula F' = M⊙F + (1-M)⊙G(F), where F is the fused feature, M is the soft mask matrix, G(F) is the feature reconstruction function based on contextual information, and F' is the reconstructed feature. This mask reconstruction process can effectively repair feature loss or distortion caused by occlusion or lighting interference, significantly improving the system's anti-interference ability under common in-vehicle interferences such as strong light, glasses reflection, and facial occlusion, and solving the problems of insufficient robustness and lack of intelligent feature-level repair capabilities in existing technologies.

[0089] Based on the fused features after mask reconstruction, the system begins to determine the driver's fatigue state. In this process, the system not only utilizes visual behavioral features but also further extracts physiological signal features reflecting the driver's chest cavity movements from the video sequence; for example, it estimates respiratory rate by analyzing pixel changes in the chest area. Subsequently, the processed fused features—visual behavioral features and extracted physiological signal features—are input into different channels of a dual-channel Transformer encoder. This dual-channel Transformer encoder utilizes its internal cross-attention mechanism to achieve deep alignment and fusion of visual behavioral features and physiological signal features, generating multimodal joint features. This deep fusion approach overcomes the limitations of existing multimodal information fusion technologies that remain at a shallow decision-making level, achieving deep co-encoding of physiological and behavioral features, and improving the precision of fatigue judgment and the advance warning time.

[0090] The generated multimodal joint features are input into a regression network, which outputs a continuous fatigue level score, for example, from 0 to 100, with higher scores indicating greater driver fatigue. To ensure the stability of the fatigue level score, the system aggregates fatigue level scores from multiple consecutive time frames into a state queue. Based on the trend of score changes and / or the proportion of abnormal scores in this state queue, the system uses an exponential moving average algorithm to calculate the smoothed fatigue level score for the current time point. Notably, the smoothing coefficient is dynamically adjusted according to the real-time vehicle motion state, such as the degree of vehicle bumps or the stability of cabin lighting, such as the degree of drastic changes in lighting. For example, when the vehicle is moving smoothly and the lighting is stable, the smoothing coefficient may be smaller to respond more quickly to changes in fatigue state; while under complex conditions such as vehicle bumps or frequent changes in lighting, the smoothing coefficient will be appropriately increased to reduce the impact of instantaneous noise on the fatigue score and ensure the accuracy and stability of fatigue assessment.

[0091] Furthermore, the system includes a dynamic sensitivity adjustment step. Based on the vehicle's current driving conditions—such as highway driving or urban congestion, in-cabin lighting conditions (e.g., strong daylight or weak nightlight), or the stability of the driver's facial image (e.g., the degree of driver head movement)—the system dynamically switches between preset high-sensitivity, medium-sensitivity, and low-sensitivity modes. For example, when driving on a highway at night with stable lighting conditions, the system may switch to high-sensitivity mode. In this mode, the feature extraction network enhances feature extraction intensity, the mask reconstruction process achieves higher reconstruction intensity, and the threshold for determining fatigue status is lowered, making it easier to trigger fatigue warnings. Conversely, in urban congestion or conditions with frequently changing lighting, the system may switch to medium-sensitivity mode to avoid false alarms caused by oversensitivity. This dynamic adjustment mechanism further enhances the system's adaptability and practicality.

[0092] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

[0093] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A method for monitoring driver fatigue, characterized in that, include: Obtain a continuous video sequence of the driver; The video sequence is input into a feature extraction network, which includes a parallel 3D convolutional neural network branch and a 2D convolutional neural network branch; the 3D convolutional neural network branch extracts the dynamic features of the video sequence in the temporal dimension; and the 2D convolutional neural network branch extracts the static features of the single frame image of the video sequence in the spatial dimension. The feature extraction network includes at least one intermediate layer, through which the dynamic features are pooled along the time dimension to obtain pooled features. The pooling features and the static features are joined across branches using residual connections to generate fused features; The fused features are subjected to mask reconstruction processing to enhance their robustness under occlusion or lighting interference, resulting in the processed fused features. Based on the processed fusion features, the driver's fatigue state is determined.

2. The driver fatigue monitoring method according to claim 1, characterized in that, The mask reconstruction process for the fused features includes: Based on the brightness gradient and local texture changes of the current image frame of the video sequence, a soft mask matrix with the same spatial size as the fused feature is generated. The element values ​​in the soft mask matrix are used to characterize the credibility of the corresponding position in the fused feature. Based on the soft mask matrix, low-confidence regions in the fused features are suppressed, and features of the low-confidence regions are reconstructed based on the contextual information of the high-confidence regions in the fused features.

3. The driver fatigue monitoring method according to claim 2, characterized in that, The process of reconstructing the fused features based on the soft mask matrix is ​​achieved by the following formula: F' = M⊙F + (1-M)⊙G(F) where F is the fused feature, M is the soft mask matrix, ⊙ represents element-wise multiplication, G(F) represents the feature reconstruction function based on context information, and F' is the reconstructed feature.

4. The driver fatigue monitoring method according to claim 1, characterized in that, The process of determining the driver's fatigue state based on the processed fusion features includes: extracting physiological signal features reflecting chest cavity movement from the video sequence; inputting the physiological signal features and visual behavior features extracted based on the fusion features into different channels of a dual-channel Transformer encoder, aligning and fusing features using a cross-attention mechanism to generate multimodal joint features; and inputting the multimodal joint features into a regression network to output a continuous fatigue level score.

5. The driver fatigue monitoring method according to claim 4, characterized in that, After inputting the multimodal joint features into the regression network and outputting continuous fatigue level scores, the method further includes stabilizing the fatigue level scores, including: aggregating the fatigue level scores of multiple consecutive time frames of the video sequence into a state queue; and calculating the smoothed fatigue level score at the current time point based on the changing trend of the scores in the state queue and / or the proportion of abnormal scores.

6. The driver fatigue monitoring method according to claim 5, characterized in that, The fatigue level score at the current time point is calculated using an exponential moving average algorithm, and the smoothing coefficient is dynamically adjusted according to the real-time motion state of the vehicle or the stability of the lighting in the carriage.

7. The driver fatigue monitoring method according to claim 1, characterized in that, It also includes a dynamic sensitivity adjustment step, which includes: dynamically switching between a preset high sensitivity mode, a medium sensitivity mode and a low sensitivity mode based on the current driving conditions of the vehicle, the lighting conditions in the cabin or the stability information of the driver's facial image; wherein, different sensitivity modes correspond to different feature extraction intensities in the feature extraction network, different reconstruction intensities in the mask reconstruction process, or different judgment thresholds in the process of determining fatigue state.

8. A driver fatigue monitoring system, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the method as described in any one of claims 1 to 7.

9. A vehicle-mounted terminal device, characterized in that, include: Image acquisition device, used to acquire continuous video sequences of the driver; And the driver fatigue monitoring system as described in claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.