Driver fatigue detection method and system based on reconstructed augmented evidence network

By constructing a reconstruction-enhanced evidence network, the robustness of fatigue detection under the influence of noise and disturbances in intelligent transportation systems was solved, achieving high-precision and high-reliability driver fatigue state detection and improving the robustness and comprehensiveness of the detection.

CN120713525BActive Publication Date: 2025-11-11HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511164476.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-11
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

Existing driver fatigue detection methods suffer from robustness and reliability issues caused by noise and external disturbances in intelligent transportation systems, making it difficult to achieve high-precision fatigue state detection under multi-source heterogeneous data.

Method used

A reconstruction-enhanced evidence network is constructed, including a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module. Combining noise-robust training with a Dirichlet distribution uncertainty-guided fusion mechanism, the robustness and accuracy of the model are improved by simulating noise data augmentation and dynamic fusion of evidence perspectives.

Benefits of technology

Achieving robust and accurate driver fatigue state classification and detection in noisy environments enhances the reliability and interpretability of the detection. By extracting multi-angle information through modal independence and related branches, the comprehensiveness and accuracy of fatigue detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120713525B_ABST
    Figure CN120713525B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for driver fatigue detection based on a reconstruction-enhanced evidence network, belonging to the field of fatigue driving detection technology. The method includes: constructing a network model comprising a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module; firstly, reconstructing and restoring noisy physiological signals; then generating multimodal evidence from both modal-independent and cross-modal perspectives; and finally, introducing an uncertainty-guided fusion mechanism based on Dirichlet distribution to achieve robust and interpretable fatigue state identification. This invention achieves high-precision and highly robust detection of driver fatigue states in complex noise environments by constructing a deep network that integrates noise-resistant reconstruction, multimodal evidence generation, and uncertainty-guided dynamic consensus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fatigue driving detection technology, and specifically to a driver fatigue detection method and system based on a reconstructed enhanced evidence network. Background Technology

[0002] In intelligent transportation systems, driver fatigue detection is crucial for ensuring driving safety. Current fatigue detection methods are mainly divided into three categories: those based on driving behavior, those based on vehicle information, and those based on physiological signals. Among them, methods based on driving behavior and those based on vehicle driving information are highly dependent on road conditions, individual driving styles, or traffic flow, making it difficult to objectively and accurately reflect the driver's fatigue state and distinguish fatigue from other abnormal driving behaviors. In contrast, physiological signals directly reflect the activity state of the central nervous system, can capture the driver's cognitive and alert states, and show the potential to construct robust fatigue representations in cross-modal fusion, providing a new approach for accurate detection.

[0003] While physiological signals demonstrate great potential in driver fatigue detection, their practical application still faces significant challenges. First, intelligent transportation systems heavily rely on multi-source heterogeneous data for real-time perception and decision-making. This data, due to its diverse forms, sources, and acquisition conditions, inherently possesses ambiguity and uncertainty. Failure to effectively model and address these inherent uncertainties will directly limit the system's robustness and safety. More seriously, in real-world environments, these inherent uncertainties are easily amplified by external noise or malicious adversarial disturbances, leading to distortion of key physiological signals. This causes traditional detection methods to degrade drastically in noisy or attacked environments, severely threatening system reliability. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a driver fatigue detection method and system based on a reconstruction-enhanced evidence network. By constructing a reconstruction-enhanced evidence network comprising an anti-noise assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module, and combining noise-robust training with an uncertainty-guided fusion mechanism based on Dirichlet distribution, a highly robust and accurate classification and detection of driver fatigue state under noisy physiological signals is achieved.

[0005] The specific solution is as follows: On the one hand, a driver fatigue detection method based on reconstructed enhanced evidence networks includes:

[0006] S1. Construct a reconstruction-enhanced evidence network, which includes a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module; the multimodal evidence perspective generator includes modality-independent branches and modality-related branches;

[0007] S2, Data augmentation is performed on the original physiological data based on simulated noise to obtain noisy physiological data; The noisy physiological data and the original physiological data are used as inputs, and the mean square error is used as the loss function to pre-train the anti-noise-assisted reconstruction module to obtain the pre-trained anti-noise-assisted reconstruction module, and the weights of the pre-trained anti-noise-assisted reconstruction module are frozen.

[0008] S3, input the raw physiological data into the reconstructed enhanced evidence network, and extract features from the raw physiological data through the multimodal evidence perspective generator to obtain the first set of evidence perspectives; the first set of evidence perspectives includes the first EEG evidence perspective generated by the modality-independent branch, the first EEG evidence perspective, and the first cross-modal evidence perspective generated by the modality-related branch.

[0009] S4. Input the noisy physiological data into the reconstruction enhancement evidence network. First, it is reconstructed through the pre-trained noise-resistant assisted reconstruction module to obtain the reconstructed physiological data. Then, input the reconstructed physiological data into the multimodal evidence perspective generator to obtain the second set of evidence perspectives. The second set of evidence perspectives includes the second EEG evidence perspective generated by the modality-independent branch, the second EEG evidence perspective, and the second cross-modal evidence perspective generated by the modality-related branch.

[0010] S5, the dynamic evidence consensus module dynamically fuses the first and second sets of evidence perspectives based on the Dirichlet distribution using an uncertainty-guided fusion mechanism, to obtain the first set of fatigue detection consensus decisions, the first set of losses, the second set of fatigue detection consensus decisions, and the second set of losses for the original clean physiological data; the first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective, and the multimodal evidence perspective generator and the dynamic evidence consensus module are optimized through the overall learning objective to obtain the optimized reconstructed enhanced evidence network;

[0011] S6 receives the noisy physiological data to be tested and inputs it into the optimized reconstruction-enhanced evidence network. The network is then reconstructed using an anti-noise assisted reconstruction module to obtain the reconstructed physiological data. The reconstructed physiological data is input into the optimized multimodal evidence perspective generator to generate evidence perspectives. The uncertainty of the evidence perspectives is quantified based on the Dirichlet distribution using the optimized dynamic evidence consensus module to achieve a fatigue detection consensus decision. Based on the fatigue detection consensus decision, the final fatigue detection classification prediction result is obtained.

[0012] 2. The driver fatigue detection method based on reconstructed enhanced evidence network according to claim 1, characterized in that, in S2, the noisy physiological data includes noisy electroencephalogram data and noisy electrooculogram data, calculated using the following formula:

[0013] ;

[0014] ;

[0015] in, This represents noisy electroencephalogram (EEG) data; This represents the raw electroencephalogram (EEG) data; Represents a binary mask; Represents noise-induced electrooculography data. This represents the raw electrooculogram (EOG) data. Indicates Gaussian noise; This indicates element-wise multiplication.

[0016] Furthermore, in S2, the noise-resistant assisted reconstruction module is a Transformer-based symmetric encoder-decoder network. Both the encoder and decoder in the symmetric encoder-decoder network are four layers. The noisy physiological data is reconstructed using this network, and the calculation formula is as follows:

[0017] ;

[0018] in, Represents reconstructed physiological data. Represents physiological data with noise; Indicates based on Symmetric encoding and decoding network.

[0019] Furthermore, in S2, the mean square error is defined as follows:

[0020] ;

[0021] in, This represents the mean squared error loss; N represents the number of samples. Indicates the first Reconstructed physiological data; This represents the a-th raw physiological data; This represents the L2 norm.

[0022] Furthermore, feature extraction is performed on the raw physiological data using a multimodal evidence perspective generator to obtain a first set of evidence perspectives. This first set of evidence perspectives includes a first EEG evidence perspective generated by a modality-independent branch, a first EEG evidence perspective, and a first cross-modal evidence perspective generated by a modality-related branch, specifically including:

[0023] The multimodal evidence perspective generator takes EEG and EEG data from the raw physiological data and inputs them into the EEG and EEG branches respectively. After passing through a convolutional layer for feature extraction and dimensionality adjustment, an intermediate representation is obtained, as follows:

[0024] ;

[0025] ;

[0026] in, Intermediate representation of EEG data; This represents the intermediate representation of electrooculogram (EOG) data. Represents electroencephalogram (EEG) data; This represents electrooculogram (EOG) data; Indicates the max pooling layer; This represents the modified linear unit activation function; Indicates the batch normalization layer; surface Show a two-dimensional convolutional layer;

[0027] The residual routing branches, with kernel sizes of 3, 5, and 7 respectively, capture local patterns at different scales in the intermediate representations of EEG and EEG data. Each residual routing branch contains three levels of residual units, and the mathematical expression of the i-th level residual unit is:

[0028] ;

[0029] ;

[0030] ;

[0031] ;

[0032] in, An intermediate representation of the EEG data of the i-th level residual unit; The intermediate representation of the electrooculogram data of the i-th level residual unit; This represents the modified linear unit activation function; An intermediate representation of the EEG data of the (i-1)th level residual unit; An intermediate representation of the electrooculogram data of the (i-1)th level residual unit; Main path operations representing EEG data; This represents the main path operation of the electrooculogram data;

[0033] The intermediate representations are mapped to the same scale by a global average pooling layer. The intermediate representations of the original physiological data from the third layer, output by the three parallel residual routing branches, are concatenated to form a joint feature vector. and ;

[0034] Finally, through linear mapping, the evidence perspectives of the original physiological data and the first set of evidence perspectives of the reconstructed physiological data are output, as follows:

[0035] ; ;

[0036] in, This is the weight matrix; This is a transpose operation; For bias terms; SoftPlus() Indicates the activation function; This indicates the perspective of the first EEG evidence; This indicates the perspective from which the first electrocardiogram (ECG) evidence is presented; = It is a vector. The first perspective on evidence from EEG data k Evidence for the classification results; = For a vector, The evidence perspective of the electrooculogram data is used to support the classification result of category c.

[0037] Furthermore, the modality-related branch generates the first cross-modal evidence perspective, specifically including:

[0038] EEG and EEG data are input into the modality-related branch. The EEG and EEG data are mapped to a uniform sequence length and hidden dimension through a modality-specific linear layer, and local features are enhanced by a one-dimensional convolutional layer, as follows:

[0039] ;

[0040] ;

[0041] in, Indicates EEG characteristics, Indicates electrooculogram characteristics, This represents a one-dimensional convolutional layer. Indicates a linear layer;

[0042] Next, the EEG features and EEG features are concatenated to form a joint representation, as follows:

[0043] ;

[0044] in, Indicates joint expression, Indicates a splicing operation;

[0045] By inserting a learnable CLStoken at the beginning of the joint representation, we obtain the extended sequence as follows:

[0046] ;

[0047] in, Indicates an extended sequence. This indicates a splicing operation. Represents CLStoken;

[0048] The extended sequence is fed into the Transformer encoder as follows:

[0049] ;

[0050] in, This represents the output of the Transformer encoder;

[0051] Finally, extract The first line is mapped to the modality-related evidence perspective, as follows: ;

[0052] in, This represents the first cross-modal evidence perspective; express The first line. = Represents a vector; This represents the evidence from the first cross-modal evidence perspective regarding the classification result of the p-th class.

[0053] Furthermore, the first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective, calculated using the following formula:

[0054] First, define the following loss for each sample:

[0055] ;

[0056] in, t is the current iteration round; T is the annealing step size; The evidence cross-entropy loss term is expressed in the form of negative log-likelihood under the Dirichlet distribution:

[0057] ;

[0058] in, Indicates the true category label The value for category c; C represents the total number of alertness state categories. Indicates Dirichlet intensity; Represents the double gamma function;

[0059] Let represent the KL divergence regularization term. This loss term is used to penalize the model for overconfidence when it makes prediction errors, and its calculation formula is as follows:

[0060] ;

[0061] in, = Indicates from Dirichlet parameters The Dirichlet parameters after removing non-misleading evidence. express The c-th component, The simplex represents the probability of class assignment. express The determined Dirichlet distribution, Indicates a uniform Dirichlet distribution; It is a gamma function;

[0062] The consistency constraint regularization term is used to ensure consistency of evidence perspectives and suppress biases introduced by noise or adversarial disturbances in a single branch. The calculation formula is as follows:

[0063] Among them, the total number of C alertness state categories; It is the normalized probability value of the m-th evidence perspective on category i.

[0064] The first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective:

[0065] ;

[0066] In the formula, N is the number of samples. It's a hyperparameter. It is the i-th original physiological data sample. This is the i-th noisy physiological data sample. It is the label of the i-th sample.

[0067] Furthermore, the uncertainty of the evidence perspective is quantified based on the Dirichlet distribution through a dynamic evidence consensus module to achieve consensus decisions for fatigue detection. Specifically, this includes:

[0068] From the perspective of evidence Mapped to Dirichlet parameters The calculation formula is as follows:

[0069] ;

[0070] in, Let be the Dirichlet parameter for the c-th category; This indicates the total number of alertness state categories;

[0071] Through Dirichlet parameters Define a Dirichlet distribution to represent the distribution of beliefs about the true class assignment, and calculate the uncertainty as follows:

[0072] ;

[0073] ;

[0074] in, This indicates evidence that the sample was predicted to be of class c; Indicates Dirichlet intensity; This indicates the total number of alertness state categories; Indicates uncertainty;

[0075] Next, an uncertainty-guided fusion mechanism is used to integrate the Dirichlet parameters from the M evidence perspective generators, calculated as follows:

[0076] ;

[0077] in, This represents the aggregated Dirichlet parameters, i.e., the fatigue detection consensus decision; The Dirichlet parameter represents the m-th evidence perspective; M represents the total number of evidence perspectives. This represents the weight of the m-th evidence perspective; This represents the uncertainty of the m-th evidence perspective. On the other hand, a driver fatigue detection system based on a reconstructed enhanced evidence network includes:

[0078] The network construction module constructs a reconstructed enhanced evidence network, which includes a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module; the multimodal evidence perspective generator includes modality-independent branches and modality-related branches;

[0079] The pre-training module augments the original physiological data with simulated noise to obtain noisy physiological data. The noisy physiological data and the original physiological data are used as inputs, and the mean square error is used as the loss function to pre-train the anti-noise-assisted reconstruction module to obtain the pre-trained anti-noise-assisted reconstruction module. The weights of the pre-trained anti-noise-assisted reconstruction module are then frozen.

[0080] The first-person perspective generation module inputs raw physiological data into the reconstructed enhanced evidence network, and extracts features from the raw physiological data through a multimodal evidence perspective generator to obtain the first set of evidence perspectives. The first set of evidence perspectives includes the first EEG evidence perspective generated by the modality-independent branch, the first EEG evidence perspective, and the first cross-modal evidence perspective generated by the modality-related branch.

[0081] The second perspective generation module inputs noisy physiological data into the reconstruction enhancement evidence network. The network is first reconstructed by a pre-trained noise-resistant assisted reconstruction module to obtain reconstructed physiological data. The reconstructed physiological data is then input into the multimodal evidence perspective generator to obtain a second set of evidence perspectives. The second set of evidence perspectives includes a second EEG evidence perspective generated by a modality-independent branch, a second EEG evidence perspective, and a second cross-modal evidence perspective generated by a modality-related branch.

[0082] The optimization module and the dynamic evidence consensus module dynamically fuse the first and second sets of evidence perspectives based on the Dirichlet distribution using an uncertainty-guided fusion mechanism. This yields the first set of fatigue detection consensus decisions, the first set of losses, the second set of fatigue detection consensus decisions, and the second set of losses for the original clean physiological data. The first and second sets of losses are then weighted and fused to obtain the overall learning objective. This overall learning objective is then used to optimize the multimodal evidence perspective generator and the dynamic evidence consensus module for reconstructing the enhanced evidence network, resulting in an optimized reconstructed enhanced evidence network.

[0083] The prediction module receives the noisy physiological data to be tested and inputs it into an optimized reconstruction-enhanced evidence network. The network is then reconstructed using an anti-noise assisted reconstruction module to obtain the reconstructed physiological data. The reconstructed physiological data is then input into an optimized multimodal evidence perspective generator to generate evidence perspectives. An optimized dynamic evidence consensus module quantifies the uncertainty of the evidence perspectives based on the Dirichlet distribution to achieve a fatigue detection consensus decision. Based on the fatigue detection consensus decision, the final fatigue detection classification prediction result is obtained.

[0084] The present invention adopts the above technical solution and has the following beneficial effects:

[0085] (1) This invention designs a collaborative mechanism between an anti-noise auxiliary reconstruction module and a multimodal evidence perspective generator to first reconstruct physiological signals and then generate evidence perspectives, thereby recovering key features from the noise-interferenced physiological signals and improving the robustness of the model to noise in real driving environments.

[0086] (2) The present invention adopts a dynamic evidence consensus module based on Dirichlet distribution, and uses uncertainty quantification to guide evidence fusion, so that the model can automatically reduce the weight of high uncertainty evidence perspectives (such as modal branches affected by noise), thereby enhancing the credibility and interpretability of fatigue detection decisions.

[0087] (3) The present invention constructs a dual-path evidence generation architecture of modality independence and modality correlation. While retaining the specific information of each physiological modality (EEG / EEG), it introduces Transformer to fuse cross-modal interaction features to achieve a more comprehensive and complementary perception of fatigue state. Attached Figure Description

[0088] Figure 1 This is a flowchart of a driver fatigue detection method based on a reconstructed enhanced evidence network, according to an embodiment of the present invention.

[0089] Figure 2 This is a framework diagram of the driver fatigue detection method based on reconstructed enhanced evidence network according to an embodiment of the present invention;

[0090] Figure 3 This is a schematic diagram of the noise-resistant assisted reconstruction module according to an embodiment of the present invention;

[0091] Figure 4 This is a schematic diagram of modal independent branches according to an embodiment of the present invention;

[0092] Figure 5 This is a schematic diagram of large-modal correlation branches in an embodiment of the present invention;

[0093] Figure 6 This is a diagram of a driver fatigue detection system based on a reconstructed enhanced evidence network, according to an embodiment of the present invention. Detailed Implementation

[0094] The present invention will now be described in further detail with reference to embodiments and accompanying drawings, but the implementation of the present invention is not limited thereto. Figure 1 As shown, the driver fatigue detection method based on reconstructed enhanced evidence networks of the present invention includes:

[0095] S1. Construct a reconstruction-enhanced evidence network, which includes a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module; the multimodal evidence perspective generator includes modality-independent branches and modality-related branches.

[0096] Specifically, such as Figure 2 As shown, the deep learning framework for driver fatigue detection in this embodiment is illustrated. EEG and EEG data are first input into a noise-resistant reconstruction module for noise processing. Then, they are fed into the EEG branch and EEG branch (modality-independent branch) and the modality-related branch in the multimodal evidence perspective generator, respectively, to generate evidence from different perspectives. This evidence is then passed to the dynamic evidence consensus module, where a consensus decision is formed through a fusion mechanism of subjective logic modeling, Dirichlet distribution parameterization, and uncertainty guidance. Finally, the prediction result is output through a Softmax layer.

[0097] S2, perform data augmentation on the original physiological data based on simulated noise to obtain noisy physiological data; take the noisy physiological data and the original physiological data as input, use the mean square error as the loss function, pre-train the anti-noise-assisted reconstruction module to obtain the pre-trained anti-noise-assisted reconstruction module, and freeze the weights of the pre-trained anti-noise-assisted reconstruction module.

[0098] Specifically, the noisy physiological data includes noisy electroencephalogram (EEG) data and noisy electrooculogram (EOG) data, and the calculation formula is as follows:

[0099] ;

[0100] ;

[0101] in, This represents noisy electroencephalogram (EEG) data; This represents the raw electroencephalogram (EEG) data; Represents a binary mask; Represents noise-induced electrooculography data. This represents the raw electrooculogram (EOG) data. Indicates Gaussian noise; This indicates element-wise multiplication.

[0102] Specifically, the mean square error is defined as follows:

[0103] ;

[0104] in, This represents the mean squared error loss; N represents the number of samples.

[0105] Indicates the first Reconstructed physiological data; This represents the a-th raw physiological data; This represents the L2 norm.

[0106] Specifically, such as Figure 3 The diagram illustrates the structure of the noise-resistant assisted reconstruction module. Input data is first preprocessed through an embedding layer and positional encoding, then fed into an encoder unit containing a multi-head attention mechanism, addition and normalization, a feedforward layer, and further addition and normalization. This process is repeated four times to enhance feature extraction capabilities. Finally, the reconstructed data is output through a symmetrical decoder. The entire module aims to effectively remove noise and recover the original signal.

[0107] S3, input the raw physiological data into the reconstructed enhanced evidence network, and extract features from the raw physiological data through the multimodal evidence perspective generator to obtain the first set of evidence perspectives; the first set of evidence perspectives includes the first EEG evidence perspective generated by the modality-independent branch, the first EEG evidence perspective, and the first cross-modal evidence perspective generated by the modality-related branch.

[0108] Specifically, a multimodal evidence perspective generator extracts features from the raw physiological data to obtain a first set of evidence perspectives. This first set of evidence perspectives includes a first EEG evidence perspective generated by a modality-independent branch, a first EEG evidence perspective, and a first cross-modal evidence perspective generated by a modality-related branch, specifically including:

[0109] The multimodal evidence perspective generator takes EEG and EEG data from the raw physiological data and inputs them into the EEG and EEG branches respectively. After passing through a convolutional layer for feature extraction and dimensionality adjustment, an intermediate representation is obtained, as follows:

[0110] ;

[0111] ;

[0112] in, Intermediate representation of EEG data; This represents the intermediate representation of electrooculogram (EOG) data. Represents electroencephalogram (EEG) data; This represents electrooculogram (EOG) data; Indicates the max pooling layer; This represents the modified linear unit activation function; Indicates the batch normalization layer; Represents a two-dimensional convolutional layer;

[0113] The residual routing branches, with kernel sizes of 3, 5, and 7 respectively, capture local patterns at different scales in the intermediate representations of EEG and EEG data. Each residual routing branch contains three levels of residual units, and the mathematical expression of the i-th level residual unit is:

[0114] ; ;

[0115] ;

[0116] ;

[0117] in, An intermediate representation of the EEG data of the i-th level residual unit; The intermediate representation of the electrooculogram data of the i-th level residual unit; This indicates a modified linear unit activation function; An intermediate representation of the EEG data of the (i-1)th level residual unit; An intermediate representation of the electrooculogram data of the (i-1)th level residual unit; Main path operations representing EEG data; This represents the main path operation of the electrooculogram data;

[0118] The intermediate representations are mapped to the same scale by a global average pooling layer. The intermediate representations of the original physiological data from the third layer, output by the three parallel residual routing branches, are concatenated to form a joint feature vector. and ;

[0119] Finally, through linear mapping, the evidence perspectives of the original physiological data and the first set of evidence perspectives are output as follows:

[0120] ;

[0121] ;

[0122] in, This is the weight matrix; This is a transpose operation; For bias terms; SoftPlus() Indicates the activation function; This indicates the perspective of the first EEG evidence; This indicates the perspective from which the first electrocardiogram (ECG) evidence is presented; = It is a vector. This represents the evidence from the perspective of the k-th category classification result of the EEG data; = For a vector, The evidence perspective of the electrooculogram data is used to support the classification result of category c.

[0123] Specifically, the modality-related branch generates the first cross-modal evidence perspective, which includes:

[0124] EEG and EEG data are input into the modality-related branch. The EEG and EEG data are mapped to a uniform sequence length and hidden dimension through a modality-specific linear layer, and local features are enhanced by a one-dimensional convolutional layer, as follows: ;

[0125] ;

[0126] in, Indicates EEG characteristics, Indicates electrooculogram characteristics, This represents a one-dimensional convolutional layer. Indicates a linear layer;

[0127] Next, the EEG features and EEG features are concatenated to form a joint representation, as follows:

[0128] ;

[0129] in, Indicates joint expression, Indicates a splicing operation;

[0130] By inserting a learnable CLStoken at the beginning of the joint representation, we obtain the extended sequence as follows:

[0131] ;

[0132] in, Indicates an extended sequence. This indicates a splicing operation. Represents CLStoken;

[0133] The extended sequence is fed into the Transformer encoder as follows:

[0134]

[0135] in, This represents the output of the Transformer encoder;

[0136] Finally, extract The first line is mapped to the modality-related evidence perspective, as follows:

[0137] ;

[0138] in, This represents the first cross-modal evidence perspective; express The first line. = Represent a vector; This represents the evidence from the first cross-modal evidence perspective regarding the classification result of the p-th class.

[0139] Specifically, such as Figure 4As shown, the modality-independent branch structure is divided into an EEG branch and an EEG branch. Each branch first processes the input data through a convolutional block, then performs feature extraction through three residual blocks, and reduces dimensionality through an average pooling layer. The outputs of the two branches are concatenated and fed into a linear layer to ultimately generate the evidence perspective of their respective modalities. This design aims to independently extract key information from different physiological signals, providing multi-faceted support for subsequent fatigue detection.

[0140] Specifically, Figure 5 The structure of the modality-related branch is illustrated. Input data is first processed through linear and convolutional layers, then concatenated before entering the embedding layer, where it is added to the positional encoding. Next, the data passes through an encoder unit containing a multi-head attention mechanism, addition and normalization, a feedforward layer, and further addition and normalization—this process is repeated three times to capture cross-modal interaction features. Finally, the output is passed through linear layers to generate modality-related evidentiary perspectives.

[0141] S4. Input the noisy physiological data into the reconstruction enhancement evidence network. First, reconstruct the data through the pre-trained noise-resistant assisted reconstruction module to obtain the reconstructed physiological data. Then, input the reconstructed physiological data into the multimodal evidence perspective generator to obtain the second set of evidence perspectives. The second set of evidence perspectives includes the second EEG evidence perspective generated by the modality-independent branch, the second EEG evidence perspective, and the second cross-modal evidence perspective generated by the modality-related branch.

[0142] Specifically, noisy physiological data is input into the reconstruction enhancement evidence network, and then reconstructed using a noise-resistant assisted reconstruction module. This noise-resistant assisted reconstruction module is a Transformer-based symmetric encoder-decoder structure, where both the encoder and decoder are four layers.

[0143] ;

[0144] In the formula, Represents reconstructed physiological data. This represents physiological data with noise. This indicates the noise-resistant assisted reconstruction module;

[0145] A second set of evidence perspectives is obtained by extracting features from the reconstructed physiological data using a multimodal evidence perspective generator. This second set of evidence perspectives includes a second EEG evidence perspective generated by a modality-independent branch, a second EEG evidence perspective, and a second cross-modal evidence perspective generated by a modality-related branch. Specifically, it includes:

[0146] The multimodal evidence perspective generator inputs the reconstructed EEG and EEG data from the reconstructed physiological data into the EEG and EEG branches, respectively. After passing through a convolutional layer for feature extraction and dimensionality adjustment, an intermediate representation is obtained, as follows:

[0147] ;

[0148] ;

[0149] in, An intermediate representation of the reconstructed EEG data; An intermediate representation of the reconstructed electrooculogram data; Represents the reconstructed electroencephalogram (EEG) data; Represents the reconstructed electrooculogram data;

[0150] Indicates the max pooling layer; This represents the modified linear unit activation function; Indicates the batch normalization layer; Represents a two-dimensional convolutional layer;

[0151] The residual routing branches, with kernel sizes of 3, 5, and 7 respectively, capture local patterns at different scales in the intermediate representations of EEG and EEG data. Each residual routing branch contains three levels of residual units, and the mathematical expression of the i-th level residual unit is:

[0152] ;

[0153] ;

[0154] ;

[0155] ;

[0156] in, An intermediate representation of the reconstructed EEG data for the i-th level residual unit; The intermediate representation of the reconstructed electrooculogram data of the i-th residual unit; This represents the modified linear unit activation function; An intermediate representation of the reconstructed EEG data for the (i-1)th level residual unit; An intermediate representation of the reconstructed electrooculogram data for the (i-1)th level residual unit; This represents the main path operation of the reconstructed EEG data; surface The main path operation of the reconstructed electrooculogram data is shown;

[0157] The intermediate representations are mapped to the same scale by a global average pooling layer. The intermediate representations of the reconstructed physiological data from the third layer, output by the three parallel residual routing branches, are concatenated to form a joint feature vector. and ;

[0158] Finally, through linear mapping, the evidence perspectives of the original physiological data and the first set of evidence perspectives of the reconstructed physiological data are output, as follows:

[0159] ;

[0160] ;

[0161] in, This is the weight matrix; This is a transpose operation;

[0162] For bias terms; SoftPlus() Indicates the activation function;

[0163] This indicates the perspective of second EEG evidence;

[0164] This indicates the perspective of the second electrooculogram evidence;

[0165] It is a vector;

[0166] The evidence perspective for the reconstructed EEG data is the evidence for the k-th class classification result;

[0167] It is a vector;

[0168] The evidence perspective of the reconstructed electrooculogram data is used to support the classification results for category c.

[0169] Specifically, in this embodiment, the convolution kernel size of the i-th level residual unit in the residual routing branch is set to s, and its mathematical expression is updated as follows:

[0170] ;

[0171] ;

[0172] ;

[0173] ;

[0174] in, The intermediate representation of the EEG data of the i-th level residual unit of the residual routing branch with kernel size s; The intermediate representation of the electrooculogram data of the i-th level residual unit of the residual routing branch with kernel size s; This represents the modified linear unit activation function; The intermediate representation of the EEG of the (i-1)th level residual unit of the residual routing branch with kernel size s; The intermediate representation of the (i-1)th level residual unit of a residual routing branch with kernel size s; The main path operation of the EEG represents the (i-1)th level residual unit of the residual routing branch with a convolution kernel size of s; This represents the main path operation of the (i-1)th level residual unit in the residual routing branch with kernel size s; initially, Will as The inputs are fed into three parallel residual routing branches with kernel sizes of 3, 5, and 7, respectively. Will as The inputs are fed into three parallel residual routing branches with kernel sizes of 3, 5, and 7, respectively.

[0175] The intermediate representations are mapped to the same scale by a global average pooling layer. The original physiological data from the third layer, output by the three parallel residual routing branches, are concatenated to form a joint feature vector. and :

[0176] ;

[0177] ;

[0178] in, Intermediate representation of EEG data for the third-level residual unit of a residual routing branch with a kernel size of 3; Intermediate representation of EEG data for the third-level residual unit of a residual routing branch with a convolutional kernel size of 5; Intermediate representation of EEG data for the third-level residual unit of a residual routing branch with a kernel size of 7; The intermediate representation of the electrooculogram data of the third-level residual unit of the residual routing branch with a kernel size of 3; The intermediate representation of the electrooculogram data of the third-level residual unit of the residual routing branch with a convolution kernel size of 5; The intermediate representation of the electrooculogram data of the third-level residual unit of the residual routing branch with a kernel size of 7; Indicates a splicing operation;

[0179] Specifically, the modality-related branch generates a second cross-modal evidence perspective, which includes:

[0180] The reconstructed EEG and EEG data are input into the modality-related branch. The reconstructed EEG and EEG data are mapped to a uniform sequence length and hidden dimension through a modality-specific linear layer, and local features are enhanced by a one-dimensional convolutional layer, as follows:

[0181] ;

[0182] ;

[0183] in, The reconstructed EEG features, This represents the reconstructed electrooculogram features. This represents a one-dimensional convolutional layer. Indicates a linear layer;

[0184] Next, the EEG features and EEG features are concatenated to form a joint representation, as follows:

[0185] ;

[0186] in, A joint representation indicating reconstruction. Indicates a splicing operation;

[0187] By inserting a learnable CLStoken at the beginning of the reconstructed joint representation, we obtain the extended sequence of the reconstruction, as follows:

[0188] ;

[0189] in, Represents the extended sequence of reconstruction. This indicates a splicing operation. Represents CLStoken;

[0190] The reconstructed extended sequence is fed into the Transformer encoder as follows:

[0191]

[0192] in, This represents the output of the Transformer encoder;

[0193] Finally, extract The first line is mapped to the modality-related evidence perspective, as follows:

[0194] ;

[0195] in, This represents the second cross-modal evidence perspective;

[0196] express The first line; It is a vector. This represents the evidence for the classification result of the p-th class from the perspective of the second cross-modal evidence.

[0197] S5, the dynamic evidence consensus module dynamically fuses the first and second sets of evidence perspectives based on the Dirichlet distribution using an uncertainty-guided fusion mechanism, to obtain the first set of fatigue detection consensus decisions, the first set of losses, the second set of fatigue detection consensus decisions, and the second set of losses for the original clean physiological data. The first set of losses and the second set of losses are then weighted and fused to obtain the overall learning objective. The multimodal evidence perspective generator and the dynamic evidence consensus module are then optimized using the overall learning objective to obtain the optimized reconstructed enhanced evidence network.

[0198] Specifically, the first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective, calculated using the following formula:

[0199] First, define the following loss for each sample:

[0200] ;

[0201] in, t is the current iteration round; T is the annealing step size; The evidence cross-entropy loss term is expressed in the form of negative log-likelihood under the Dirichlet distribution:

[0202] ;

[0203] in, Indicates the true category label The value for category c; C represents the total number of alertness state categories. Indicates Dirichlet intensity; Represents the double gamma function;

[0204] Let represent the KL divergence regularization term. This loss term is designed to penalize the model for overconfidence when it makes prediction errors, and its calculation formula is as follows:

[0205] ;

[0206] in, = Indicates from Dirichlet parameters The Dirichlet parameters after removing non-misleading evidence. express The c-th component, The simplex represents the probability of class assignment. express The determined Dirichlet distribution, Indicates a uniform Dirichlet distribution; It is a gamma function;

[0207] The consistency constraint regularization term is used to ensure consistency of evidence perspectives and suppress biases introduced by noise or adversarial disturbances in a single branch. The calculation formula is as follows:

[0208] ;

[0209] Among them, C is the total number of alertness state categories; It is the normalized probability value of the m-th evidence perspective on category c.

[0210] The first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective:

[0211] ;

[0212] In the formula, N is the number of samples. It's a hyperparameter. It is the i-th original physiological data sample. This is the i-th noisy physiological data sample. It is the label of the i-th sample.

[0213] Specifically, the optimized reconstruction-enhanced evidence network comprises three modules: a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module. The noise-resistant assisted reconstruction module is pre-trained using S2, while the multimodal evidence perspective generator and dynamic evidence consensus module are optimized using S5. The first set of losses is obtained by inputting raw, clean physiological data into the network, while the second set of losses is obtained by inputting noisy physiological data obtained through data augmentation. The raw, clean physiological data does not require processing by the noise-resistant assisted reconstruction module, while the noisy physiological data does.

[0214] S6 receives the noisy physiological data to be tested and inputs it into the optimized reconstruction-enhanced evidence network. The network is then reconstructed using an anti-noise assisted reconstruction module to obtain the reconstructed physiological data. The reconstructed physiological data is input into the optimized multimodal evidence perspective generator to generate evidence perspectives. The uncertainty of the evidence perspectives is quantified based on the Dirichlet distribution using the optimized dynamic evidence consensus module to achieve a fatigue detection consensus decision. Based on the fatigue detection consensus decision, the final fatigue detection classification prediction result is obtained.

[0215] Specifically, the noise-resistant assisted reconstruction module is a Transformer-based symmetric encoder-decoder network. Both the encoder and decoder in the symmetric encoder-decoder network are four layers. The noisy physiological data is reconstructed using this network, and the calculation formula is as follows:

[0216] ;

[0217] In the formula, Represents reconstructed physiological data. Represents physiological data with noise; Indicates based on Symmetric encoding and decoding network.

[0218] Specifically, the dynamic evidence consensus module quantifies the uncertainty of the evidence perspective based on the Dirichlet distribution to achieve consensus decisions in fatigue detection. This includes:

[0219] From the perspective of evidence Mapped to Dirichlet parameters The calculation formula is as follows:

[0220] ;

[0221] in, Let be the Dirichlet parameter for the c-th category; This indicates the total number of alertness state categories;

[0222] Through Dirichlet parameters Define a Dirichlet distribution to represent the distribution of beliefs about the true class assignment, and calculate the uncertainty as follows:

[0223] ;

[0224] ;

[0225] in, This indicates evidence that the sample was predicted to be of class c; Indicates Dirichlet intensity; This indicates the total number of alertness state categories; Indicates uncertainty;

[0226] Next, an uncertainty-guided fusion mechanism is used to integrate the Dirichlet parameters from the M evidence perspective generators, calculated as follows:

[0227] ;

[0228] in, This represents the aggregated Dirichlet parameters, i.e., the fatigue detection consensus decision; The Dirichlet parameter represents the m-th evidence perspective; M represents the total number of evidence perspectives. This represents the weight of the m-th evidence perspective; This represents the uncertainty of the m-th evidence perspective.

[0229] Specifically, based on The calculated posterior expected probability will be used to derive the final classification prediction. The final predicted category. It is the category with the highest posterior expected probability, which is mathematically expressed as:

[0230] ;

[0231] in, Indicates the Dirichlet strength of the fusion. This represents the Dirichlet parameter for the c-th category.

[0232] Specifically, this experiment used the SEED-VIG dataset, which consisted of simultaneous EEG and EEG signals collected from 23 subjects in a simulated driving environment using the Neuroscan system. The dataset was labeled using PERCLOS, calculated based on the duration of eye closure recorded at 8-second intervals. Different PERCLOS values ​​were used for labeling: PERCLOS < 0.35 was labeled as awake; 0.35 <= PERCLOS < 0.7 was labeled as fatigued; and PERCLOS >= 0.7 was labeled as drowsy. Each subject's 2-hour data set was divided into 885 8-second samples, which, after artifact removal preprocessing, resulted in a total of 20355 labeled multimodal physiological signal fragments. For the EEG signals, a differential entropy linear dynamic system was used for preprocessing, extracting DE features from the full frequency band (1–50 Hz) with a frequency resolution of 2 Hz. The formula for calculating differential entropy is defined as follows:

[0233] ;in, It follows a Gaussian distribution. EEG samples, The mean of a Gaussian distribution is given. denoted by σ², where σ is the variance of the Gaussian distribution, e is the natural constant, log represents the natural logarithm, and π is pi. The experiment also used a forehead electrode shared with the electrooculogram (EOG) data, ultimately obtaining 21 channels of data.

[0234] For the electrooculogram signal, horizontal and vertical eye movements were extracted by subtraction rule and independent component analysis, or by using subtraction rule and independent component analysis simultaneously. After wavelet packet transformation, 36-dimensional features were extracted, and the features obtained by the three methods were stacked to construct a three-channel composite feature vector.

[0235] In the experiment, a random sampling strategy was used to divide all experimental data into training and test sets in a 16:7 ratio. To comprehensively evaluate the robustness of the model under different noise interferences, noise injections of varying intensities were simulated during the testing phase. The combinations of visible EEG channels and Gaussian noise standard deviation from EEG (channel number / standard deviation) used in the tests were set in increasing order of noise intensity as follows: 21 / 0, 17 / 0.3, 13 / 0.5, 9 / 0.7, 5 / 1.0. These combinations covered different scenarios from no interference to severe damage.

[0236] For the pre-training setup of the noise-resistant reconstruction module, the Adam optimizer was used with a learning rate of 5e-4 and 80 training epochs. For the reconstruction-enhanced evidence network, the Adam optimizer was used with a learning rate of 5e-4 and 30 training epochs. For data augmentation settings, the number of visible channels in the EEG data was randomly sampled between 5 and 21 in each training batch, and the standard deviation of the Gaussian noise was randomly sampled between 0 and 1. To ensure the reliability and statistical stability of the experimental results, all experiments were independently repeated 5 times with different random number seeds. The model performance was evaluated by classification accuracy and F1 score on the test set. The final performance metrics reported are the arithmetic mean and standard deviation of the results from 5 independent replicate experiments.

[0237] Table 1. Comparison of Classification Performance;

[0238]

[0239] Table 1 shows the results of driver fatigue detection on the SEED-VIG dataset, from which the following conclusions can be drawn:

[0240] REEN outperforms other methods under all noise conditions. In a noise-free environment (21 / 0), REEN achieves a classification accuracy of 96.69% and an F1 score of 96.89%; even under the most extreme condition (5 / 1.0), it maintains a classification accuracy of 91.76% and an F1 score of 92.24%, surpassing the second-best method. Its superior performance is attributed to the organic combination of reconstruction denoising and noise-based data augmentation, as well as dynamic fusion based on modal uncertainty. This allows the model to adaptively balance the contributions of each modality under different noise intensities, preserving the complementary advantages of multimodal information while effectively suppressing interference from unreliable sources.

[0241] The performance of all methods decreases as the number of channels decreases and the noise intensity increases, but REEN's decrease is the smallest. From no interference to severe damage, REEN's classification accuracy decreases by only 4.93%, demonstrating good robustness to missing channels and Gaussian noise.

[0242] In summary, by employing data augmentation and noise-resistant assisted reconstruction modules, the network's resilience to complex noise is enhanced from the source, significantly improving its robustness and accuracy in non-ideal environments. A multimodal evidence perspective generator generates modally independent and modally related evidence perspectives; an uncertainty-aware dynamic evidence consensus module quantifies the uncertainty of these perspectives and dynamically fuses them based on their respective uncertainties to achieve consensus, thereby enabling robust and quantifiable decisions with quantifiable uncertainty in noisy environments.

[0243] like Figure 6 As shown, this embodiment also discloses a driver fatigue detection system for reconstructing an enhanced evidence network, including:

[0244] Network construction module 61 constructs a reconstruction-enhanced evidence network, which includes a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module; the multimodal evidence perspective generator includes modality-independent branches and modality-related branches;

[0245] The pre-training module 62 performs data augmentation on the original physiological data based on simulated noise to obtain noisy physiological data; it takes the noisy physiological data and the original physiological data as input, uses the mean square error as the loss function, and pre-trains the anti-noise-assisted reconstruction module to obtain the pre-trained anti-noise-assisted reconstruction module, and freezes the weights of the pre-trained anti-noise-assisted reconstruction module.

[0246] The first-view generation module 63 inputs the raw physiological data into the reconstructed enhanced evidence network, and extracts features from the raw physiological data through the multimodal evidence perspective generator to obtain the first set of evidence perspectives; the first set of evidence perspectives includes the first EEG evidence perspective generated by the modality-independent branch, the first EEG evidence perspective, and the first cross-modal evidence perspective generated by the modality-related branch.

[0247] The second perspective generation module 64 inputs noisy physiological data into the reconstruction enhancement evidence network, first reconstructs it through a pre-trained noise-resistant assisted reconstruction module to obtain reconstructed physiological data; then inputs the reconstructed physiological data into the multimodal evidence perspective generator to obtain a second set of evidence perspectives, which includes a second EEG evidence perspective generated by a modality-independent branch, a second EEG evidence perspective, and a second cross-modal evidence perspective generated by a modality-related branch.

[0248] The optimization module 65 and the dynamic evidence consensus module dynamically fuse the first and second sets of evidence perspectives based on the Dirichlet distribution using an uncertainty-guided fusion mechanism. This yields a first set of fatigue detection consensus decisions, a first set of losses, a second set of fatigue detection consensus decisions, and a second set of losses for the original clean physiological data. The first and second sets of losses are then weighted and fused to obtain an overall learning objective. This overall learning objective is used to optimize the multimodal evidence perspective generator and the dynamic evidence consensus module for reconstructing the enhanced evidence network, resulting in an optimized reconstructed enhanced evidence network. The prediction module 66 receives the noisy physiological data to be tested and inputs it into the optimized reconstructed enhanced evidence network. The network is then reconstructed using an anti-noise assisted reconstruction module to obtain reconstructed physiological data. This reconstructed physiological data is input into the optimized multimodal evidence perspective generator to generate evidence perspectives. The optimized dynamic evidence consensus module quantifies the uncertainty of the evidence perspectives based on the Dirichlet distribution to achieve a fatigue detection consensus decision. Based on the fatigue detection consensus decision, the final fatigue detection classification prediction result is obtained.

[0249] The specific implementation of the driver fatigue detection system based on the reconstruction enhanced evidence network is the same as that of the driver fatigue detection method based on the reconstruction enhanced evidence network, and will not be described again in this embodiment.

[0250] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.

Claims

1. A method for driver fatigue detection based on reconstructed enhanced evidence networks, characterized in that, include: S1, Construct a reconstruction-enhanced evidence network, which includes a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module; The multimodal evidence perspective generator includes modality-independent branches and modality-related branches; S2, Data augmentation of raw physiological data based on simulated noise to obtain noisy physiological data; Noisy physiological data and raw physiological data are used as inputs, and mean squared error is used as the loss function to pre-train the anti-noise-assisted reconstruction module, obtain the pre-trained anti-noise-assisted reconstruction module, and freeze the weights of the pre-trained anti-noise-assisted reconstruction module. The noisy physiological data includes noisy electroencephalogram (EEG) data and noisy electrooculogram (EOG) data, and the calculation formula is as follows: in, Represents noisy electroencephalogram (EEG) data; X EEG Represents raw electroencephalogram (EEG) data; M channel Represents a binary mask; X represents noisy electrooculogram data. EOG This represents the raw electrooculogram data; ε represents Gaussian noise; ⊙ represents element-wise multiplication. S3, input the raw physiological data into the reconstructed enhanced evidence network, and extract features from the raw physiological data through the multimodal evidence perspective generator to obtain the first set of evidence perspectives; the first set of evidence perspectives includes the first EEG evidence perspective generated by the modality-independent branch, the first EEG evidence perspective, and the first cross-modal evidence perspective generated by the modality-related branch. S4. Input the noisy physiological data into the reconstruction enhancement evidence network. First, it is reconstructed through the pre-trained noise-resistant assisted reconstruction module to obtain the reconstructed physiological data. Then, input the reconstructed physiological data into the multimodal evidence perspective generator to obtain the second set of evidence perspectives. The second set of evidence perspectives includes the second EEG evidence perspective generated by the modality-independent branch, the second EEG evidence perspective, and the second cross-modal evidence perspective generated by the modality-related branch. S5, the dynamic evidence consensus module dynamically fuses the first and second sets of evidence perspectives based on the Dirichlet distribution using an uncertainty-guided fusion mechanism, to obtain the first set of fatigue detection consensus decisions, the first set of losses, the second set of fatigue detection consensus decisions, and the second set of losses for the original clean physiological data; the first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective, and the multimodal evidence perspective generator and the dynamic evidence consensus module are optimized through the overall learning objective to obtain the optimized reconstructed enhanced evidence network; S6 receives the noisy physiological data to be tested, inputs it into the optimized reconstruction-enhanced evidence network, and reconstructs it through the noise-resistant assisted reconstruction module to obtain the reconstructed physiological data; inputs the reconstructed physiological data into the optimized multimodal evidence perspective generator to generate evidence perspectives; and quantifies the uncertainty of the evidence perspectives based on the Dirichlet distribution through the optimized dynamic evidence consensus module to achieve consensus decision on fatigue detection. Based on the consensus decision of fatigue detection, the final fatigue detection classification prediction result is obtained.

2. The driver fatigue detection method based on reconstructed enhanced evidence network according to claim 1, characterized in that, In S2, the noise-resistant assisted reconstruction module is a Transformer-based symmetric encoder-decoder network. Both the encoder and decoder in the symmetric encoder-decoder network are four layers. The noisy physiological data is reconstructed using this network, and the calculation formula is as follows: in, Represents reconstructed physiological data. This represents noisy physiological data; Transformer represents a symmetric encoder-decoder network based on Transformer.

3. The driver fatigue detection method based on reconstructed enhanced evidence network according to claim 1, characterized in that, In S2, the mean square error is defined as follows: in, This represents the mean squared error loss; N represents the number of samples. X represents the physiological data of the a-th reconstruction; a This represents the a-th raw physiological data; ||·|| 2 This represents the L2 norm.

4. The driver fatigue detection method based on reconstructed enhanced evidence network according to claim 1, characterized in that, In S3, a multimodal evidence perspective generator extracts features from the raw physiological data to obtain the first set of evidence perspectives. This first set of evidence perspectives includes a first EEG evidence perspective generated by a modality-independent branch, a first EEG evidence perspective, and a first cross-modal evidence perspective generated by a modality-related branch. Specifically, it includes: The multimodal evidence perspective generator takes EEG and EEG data from the raw physiological data and inputs them into the EEG and EEG branches respectively. After passing through a convolutional layer for feature extraction and dimensionality adjustment, an intermediate representation is obtained, as follows: in, Intermediate representation of EEG data; X represents the intermediate representation of electrooculogram (EOG) data. EEG Represents electroencephalogram (EEG) data; X EOG Represents electrooculogram (EOG) data; MaxPool represents the max pooling layer; ReLU represents the modified linear unit activation function; BN represents the batch normalization layer; Conv2d represents the two-dimensional convolutional layer. The residual routing branches, with kernel sizes of 3, 5, and 7 respectively, capture local patterns at different scales in the intermediate representations of EEG and EEG data. Each residual routing branch contains three levels of residual units, and the mathematical expression of the i-th level residual unit is: in, An intermediate representation of the EEG data of the i-th level residual unit; The intermediate representation of the electrooculogram data of the i-th residual unit; ReLU represents the modified linear unit activation function; An intermediate representation of the EEG data of the (i-1)th level residual unit; An intermediate representation of the electrooculogram data of the (i-1)th level residual unit; Main path operations representing EEG data; This represents the main path operation of the electrooculogram data; The intermediate representations are mapped to the same scale by a global average pooling layer. The intermediate representations of the original physiological data from the third layer, output by the three parallel residual routing branches, are concatenated to form a joint feature vector I. EEG and I EOG ; Finally, through linear mapping, the evidence perspectives of the original physiological data and the first set of evidence perspectives of the reconstructed physiological data are output, as follows: e EEG =SoftPlus(W T AND EEG +b); e EOG =SoftPlus(W T AND EOG +b); Where W is the weight matrix; T is the transpose operation; b is the bias term; SoftPlus() represents the activation function; e EEG Indicates the perspective of first EEG evidence; e EOG Indicates the perspective of the first electrooculogram evidence; e EEG =(e1,...,e k ,...,e C ) EEG It is a vector, e k This represents the evidence from the perspective of the k-th class classification result of the EEG data; e EOG =(e1,...,e c ,...,e C ) EOG Let e ​​be a vector c The evidence perspective of the electrooculogram data is used to support the classification result of category c.

5. The driver fatigue detection method based on reconstructed enhanced evidence network according to claim 4, characterized in that, The modality-related branch generates the first cross-modal evidence perspective, specifically including: EEG and EEG data are input into the modality-related branch. The EEG and EEG data are mapped to a uniform sequence length and hidden dimension through a modality-specific linear layer, and local features are enhanced by a one-dimensional convolutional layer, as follows: Z EEG =Conv1D(Linear(X EEG )); Z EOG =Conv1D(Linear(X EOG )); Among them, Z EEG Z represents the characteristics of an electroencephalogram (EEG). EOG The features of electrooculogram are represented by Conv1D, which represents a one-dimensional convolutional layer, and Linear, which represents a linear layer. Next, the EEG features and EEG features are concatenated to form a joint representation, as follows: U=Concat([Z EEG ;Z EOG ]); Where U represents union representation and Concat represents concatenation operation; By inserting a learnable CLStoken at the beginning of the joint representation, we obtain the extended sequence as follows: U′ = Concat([S; U]); Where U′ represents the extended sequence, Concat represents the concatenation operation, and S represents CLStoken; The extended sequence is fed into the Transformer encoder as follows: O=TransformerEncoder(U′); Where O represents the output of the Transformer encoder; Finally, the first row of O is extracted and mapped to the modality-related evidence perspective, as follows: e cross-modal <SoftPlus(Linear(O0)); Among them, e cross-modal Indicates the first cross-modal evidence perspective; O0 indicates the first row of O; e cross-modal =(e1,...,e p ,...,e C ) cross-modal Represents a vector; e p This represents the evidence from the first cross-modal evidence perspective regarding the classification result of the p-th class.

6. The driver fatigue detection method based on reconstructed enhanced evidence network according to claim 1, characterized in that, In S5, the first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective. The calculation formula is as follows: First, define the following loss for each sample: in, t is the current iteration round; T is the annealing step size; The evidence cross-entropy loss term is expressed in the form of negative log-likelihood under the Dirichlet distribution: Among them, y c represents the value of the true category label y in category c; C represents the total number of alertness categories; S represents the Dirichlet intensity; ψ() represents the double gamma function. Let represent the KL divergence regularization term. This loss term is used to penalize the model for overconfidence when it makes prediction errors, and its calculation formula is as follows: in, This represents the Dirichlet parameter α after removing non-misleading evidence. express The c-th component, p, represents the simplex of the class assignment probability. express The determined Dirichlet distribution, D(p|1) represents a uniform Dirichlet distribution; Γ(·) is the gamma function; The consistency constraint regularization term is used to ensure consistency of evidence perspectives and suppress biases introduced by noise or adversarial disturbances in a single branch. The calculation formula is as follows: Among them, C represents the total number of alertness state categories; p m,c =Softmax(e m,c ) is the normalized probability value of the m-th evidence perspective on category i; The first set of losses and the second set of losses are weighted and fused to obtain the overall learning objective: In the formula, N is the number of samples, α is the hyperparameter, and x i It is the i-th original physiological data sample. It is the i-th noisy physiological data sample, y i It is the label of the i-th sample.

7. The driver fatigue detection method based on reconstructed enhanced evidence network according to claim 1, characterized in that, In S6, the uncertainty of the evidence perspective is quantified based on the Dirichlet distribution through the dynamic evidence consensus module to achieve consensus decision-making for fatigue detection. Specifically, this includes: From the perspective of evidence, e = (e1,…,e c ,…,e C The mapping is to Dirichlet parameters α = (α1, ... α2) c ,…,α C The calculation formula is as follows: α c =e c +1,c=1,2,...,C.←; Where, α c Here is the Dirichlet parameter for the c-th category; C.← represents the total number of alertness state categories; A Dirichlet distribution is determined using the Dirichlet parameter α. This distribution represents the distribution of beliefs about the true class assignment, and the uncertainty is calculated as follows: Among them, e c This represents the evidence that the sample is predicted to be of class c; S represents the Dirichlet intensity; C represents the total number of alertness state categories; u represents the uncertainty. Next, an uncertainty-guided fusion mechanism is used to integrate the Dirichlet parameters from the M evidence perspective generators, calculated as follows: Where, α fused This represents the aggregated Dirichlet parameters, i.e., the fatigue detection consensus decision; α m The Dirichlet parameter represents the m-th evidentiary perspective; M represents the total number of evidentiary perspectives; w m u represents the weight of the m-th evidence perspective; m This represents the uncertainty of the m-th evidence perspective.

8. A driver fatigue detection system based on a reconstructed enhanced evidence network, characterized in that, include: The network construction module constructs a reconstruction-enhanced evidence network, which includes a noise-resistant assisted reconstruction module, a multimodal evidence perspective generator, and a dynamic evidence consensus module. The multimodal evidence perspective generator includes modality-independent branches and modality-related branches; The pre-training module augments the original physiological data with simulated noise to obtain noisy physiological data. Noisy physiological data and raw physiological data are used as inputs, and mean squared error is used as the loss function to pre-train the anti-noise-assisted reconstruction module, obtain the pre-trained anti-noise-assisted reconstruction module, and freeze the weights of the pre-trained anti-noise-assisted reconstruction module. The noisy physiological data includes noisy electroencephalogram (EEG) data and noisy electrooculogram (EOG) data, and the calculation formula is as follows: in, Represents noisy electroencephalogram (EEG) data; X EEG Represents raw electroencephalogram (EEG) data; M channel Represents a binary mask; X represents noisy electrooculogram data. EOG This represents the raw electrooculogram data; ε represents Gaussian noise; ⊙ represents element-wise multiplication. The first-person perspective generation module inputs raw physiological data into the reconstructed enhanced evidence network, and extracts features from the raw physiological data through a multimodal evidence perspective generator to obtain the first set of evidence perspectives. The first set of evidence perspectives includes the first EEG evidence perspective generated by the modality-independent branch, the first EEG evidence perspective, and the first cross-modal evidence perspective generated by the modality-related branch. The second perspective generation module inputs noisy physiological data into the reconstruction enhancement evidence network. The network is first reconstructed by a pre-trained noise-resistant assisted reconstruction module to obtain reconstructed physiological data. The reconstructed physiological data is then input into the multimodal evidence perspective generator to obtain a second set of evidence perspectives. The second set of evidence perspectives includes a second EEG evidence perspective generated by a modality-independent branch, a second EEG evidence perspective, and a second cross-modal evidence perspective generated by a modality-related branch. The optimization module and the dynamic evidence consensus module dynamically fuse the first and second sets of evidence perspectives based on the Dirichlet distribution using an uncertainty-guided fusion mechanism. This yields the first set of fatigue detection consensus decisions, the first set of losses, the second set of fatigue detection consensus decisions, and the second set of losses for the original clean physiological data. The first and second sets of losses are then weighted and fused to obtain the overall learning objective. This overall learning objective is then used to optimize the multimodal evidence perspective generator and the dynamic evidence consensus module for reconstructing the enhanced evidence network, resulting in an optimized reconstructed enhanced evidence network. The prediction module receives the noisy physiological data to be tested and inputs it into an optimized reconstruction-enhanced evidence network. The network is then reconstructed using an anti-noise assisted reconstruction module to obtain the reconstructed physiological data. The reconstructed physiological data is then input into an optimized multimodal evidence perspective generator to generate evidence perspectives. An optimized dynamic evidence consensus module quantifies the uncertainty of the evidence perspectives based on the Dirichlet distribution to achieve a fatigue detection consensus decision. Based on the fatigue detection consensus decision, the final fatigue detection classification prediction result is obtained.

Citation Information

Patent Citations

  • Driving fatigue state recognition method based on multimode EEG signal and 1DCNN (one-dimensional convolutional neural network) migration

    CN110772268A

  • Pilot fatigue state monitoring method based on credible conflict multi-mode learning

    CN118228190A