A fatigue driving automatic monitoring system based on deep learning technology
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-08-11
AI Technical Summary
疲劳会严重削弱驾驶员的反应速度、判断力和专注度,导致操作失误、车道偏离甚至瞬间入睡,极易引发追尾、冲出路面等重大交通事故,威胁自身及他人生命安全
[0007] This invention discloses an automatic fatigue driving monitoring system based on deep learning technology. It accurately acquires and effectively identifies EEG signals, maps fatigue states, and outputs a continuous fatigue measurement value from 0-100%. When the user's fatigue value exceeds a preset safety threshold, the system triggers a voice prompt module to provide a voice alert. Compared to the traditional binary "fatigue/non-fatigue" judgment, this invention uses a continuous fatigue measurement from 0-100%, which accurately captures the fatigue accumulation process, adapts to the different tolerance thresholds of different drivers, and offers higher stability in EEG signal extraction and identification. The delay between threshold triggering and prompt output is low, and the voice prompt requires no manual operation from the driver, avoiding distraction from driving attention.
Smart Images

Figure CN121465592B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an automatic fatigue driving identification system. In particular, it relates to an automatic fatigue driving monitoring system based on deep learning technology. Background Technology
[0002] Fatigue driving is extremely dangerous. Fatigue severely impairs a driver's reaction time, judgment, and concentration, leading to operational errors, lane departures, and even momentary drowsiness. This greatly increases the risk of rear-end collisions, runaway accidents, and other serious traffic accidents, threatening the lives of the driver and others. Therefore, proactive fatigue driving detection can provide timely warnings of risks and effectively prevent tragedies, serving as a crucial line of defense for accident prevention and ensuring road safety.
[0003] In this regard, EEG signals can directly and objectively reflect changes in the brain's alertness state, demonstrating unique value in fatigue monitoring. Compared to subjective feelings or behavioral observations, EEG provides more objective and accurate physiological data, directly reflecting changes in the brain's alertness state and cognitive load, making it a core physiological indicator for fatigue monitoring. By analyzing EEG signals, we can identify fatigue-related characteristic patterns and detect the decline in driver alertness in different situations. This EEG-based monitoring technology allows us to continuously track the dynamic changes in fatigue levels. Mastering this key information enables timely warnings or interventions when a driver's fatigue level exceeds a safe threshold, effectively preventing accidents.
[0004] Deep learning, a crucial branch of machine learning, enables machines to automatically learn and extract features from massive amounts of data by constructing multi-layered neural networks. In the field of fatigue state monitoring based on EEG signals, deep learning plays a key role. By training deep models, a mapping relationship between EEG signals and fatigue states can be effectively established. This technology can automatically learn the complex spatiotemporal patterns contained in EEG signals, significantly improving the accuracy of emotion recognition. Furthermore, deep learning can flexibly integrate various architectures such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and graph neural networks (GNNs) to better capture the temporal characteristics of signals. Deep learning is an end-to-end learning method, and to date, many deep learning architectures have been proposed and applied in fields such as EEG analysis and fatigue state monitoring. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide an automatic fatigue driving monitoring system based on deep learning technology that can provide voice prompts when the user's fatigue value exceeds a preset safety threshold.
[0006] The technical solution adopted in this invention is: an automatic fatigue driving monitoring system based on deep learning technology, comprising an EEG signal acquisition device, a fatigue monitoring module, and a voice prompt module connected in series. The EEG signal acquisition module acquires EEG signals from the driver's scalp; the fatigue monitoring module integrates a multimodal attention coupled graph convolutional network MC-GCN to analyze the acquired EEG signals and determine the user's fatigue level in real time; the voice prompt module provides a voice prompt based on the recognition results of the fatigue monitoring module when the user's fatigue value exceeds a preset safety threshold.
[0007] This invention discloses an automatic fatigue driving monitoring system based on deep learning technology. It accurately acquires and effectively identifies EEG signals, maps fatigue states, and outputs a continuous fatigue measurement value from 0-100%. When the user's fatigue value exceeds a preset safety threshold, the system triggers a voice prompt module to provide a voice alert. Compared to the traditional binary "fatigue / non-fatigue" judgment, this invention uses a continuous fatigue measurement from 0-100%, which accurately captures the fatigue accumulation process, adapts to the different tolerance thresholds of different drivers, and offers higher stability in EEG signal extraction and identification. The delay between threshold triggering and prompt output is low, and the voice prompt requires no manual operation from the driver, avoiding distraction from driving attention. Attached Figure Description
[0008] Figure 1 This is a system block diagram of an automatic fatigue driving monitoring system based on deep learning technology according to the present invention.
[0009] Figure 2 This is a schematic diagram of the internal structure of the EEG signal acquisition device in this invention;
[0010] Figure 3 This is a schematic diagram of the structure of the Multimodal Attention Coupled Graph Convolutional Network (MC-GCN).
[0011] Figure 4 This is a schematic diagram of the five frequency bands;
[0012] Figure 5 This is a schematic diagram of the feature-enhanced attention coupling module;
[0013] Figure 6 This is a schematic diagram of the dual-modal functional connection topology building module;
[0014] Figure 7 This is a schematic diagram of the spatiotemporal adaptive feature hierarchical fusion module. Detailed Implementation
[0015] The fatigue driving automatic monitoring system based on deep learning technology of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.
[0016] like Figure 1 As shown, the present invention discloses an automatic fatigue driving monitoring system based on deep learning technology, comprising an EEG signal acquisition device 1, a fatigue monitoring module 2, and a voice prompt module 3 connected in series. The EEG signal acquisition module 1 acquires EEG signals from the driver's scalp; the fatigue monitoring module 2 integrates a multimodal attention-coupled graph convolutional network (MC-GCN) to analyze the acquired EEG signals and determine the user's fatigue level in real time; the voice prompt module 3 provides a voice prompt based on the recognition results of the fatigue monitoring module 2 when the user's fatigue value exceeds a preset safety threshold.
[0017] like Figure 2 As shown, the EEG signal acquisition device 1 includes: an acquisition module 11 for acquiring EEG signals, an analog-to-digital converter module 12 for amplifying and converting EEG signals, an MCU processor 13 for timing logic control, and a Wi-Fi module 14 for data transmission, all connected in series. A power supply module 15 is also provided, which is connected to the acquisition module 11, the analog-to-digital converter module 12, the MCU processor 13, and the Wi-Fi module 14 respectively for power supply.
[0018] The acquisition module 11 includes an EEG cap and connecting wires. The electrodes in the EEG cap are directly attached to the driver's scalp and collect EEG signals from the cerebral cortex in real time. The EEG cap is connected to the analog-to-digital conversion module 12 via the connecting wires to transmit the collected EEG signals to the analog-to-digital conversion module 12. The electrode distribution in the EEG cap conforms to the international 10-20 lead system and includes 34 electrodes: FP1, FP2, AF3, AF4, F7, F3, Fz, F4, F8, FC5, FC1, FC2, FC6, T7, C3, Cz, C4, T8, CP5, CP1, CP2, CP6, P7, P3, Pz, P4, P8, O1, Oz, O2, A1, and A2. Among them, A1 is the reference electrode, located at the left mastoid process, and A2 is ground (GND), located at the right mastoid process.
[0019] The input terminal of the analog-to-digital conversion module 12 is connected to the acquisition module 11, and the output terminal is connected to the MCU processor 13. It includes an analog input with high common-mode rejection ratio for receiving EEG signals, a low-noise programmable gain amplifier (PGA) for amplifying EEG signals, and a high-resolution synchronous sampling analog-to-digital converter (ADC) connected in sequence. The A / D conversion accuracy is 24 bits. It is used to receive the EEG signals acquired by the acquisition module 11 and amplify and convert the signals into analog signals.
[0020] The MCU processor 13 is used to adjust the PGA amplification factor and ADC sampling rate of the analog-to-digital converter module 12, control the analog-to-digital converter module 12 to sample at a fixed sampling frequency f = 200Hz, read the EEG signal collected by the analog-to-digital converter module 12 through the SPI bus, and communicate with the WiFi module 14 through the UART interface. The MCU processor 13 is equipped with a Cortex-M3 core, operates at a frequency of 72MHz, adopts an LQFP-48 package, has 48 pins, and has a size of 7mm×7mm. It is small and lightweight, has 64KB of Flash memory, and has SPI and UART communication interfaces.
[0021] The Wi-Fi module 14 includes an ESP8266 Wi-Fi chip. The Wi-Fi module operates in transparent transmission mode and transmits the EEG signal received by the MCU processor 13 to the fatigue monitoring module 2 wirelessly.
[0022] The power supply module 15 uses a 5V input voltage and obtains the different operating voltages required by different chips on each module through a voltage conversion chip to power each module.
[0023] The fatigue monitoring module 2 connects to the same wireless network as the EEG signal acquisition device 1 via the vehicle's built-in Wi-Fi module, establishing a communication link with the EEG signal acquisition device 1. It receives the driver's EEG signals from the EEG signal acquisition device 1 and monitors the driver's fatigue level in real time. The multimodal attention-coupled graph convolutional network (MC-GCN) deployed in the fatigue monitoring module 2 is a fatigue regression model that analyzes the driver's EEG signals in real time and outputs a continuous fatigue quantification value from 0-100%. When the driver's fatigue level is ≤40%, it indicates that the driver's physiological state is within a safe threshold and there is no risk of fatigue driving in the short term. When the driver's fatigue level is between 40% and 80%, it indicates that the driver is in a moderate fatigue state. When the driver's fatigue level is >80%, it indicates that the driver is in a severe fatigue state and there is a risk of a traffic accident at any time. The voice prompt module 3 then provides voice prompts.
[0024] like Figure 3As shown, the Multimodal Attention Coupled Graph Convolutional Network (MC-GCN) in the fatigue monitoring module 2 is cascaded together by a multimodal feature extraction module 21, a feature-enhancing attention coupling module 22, a bimodal functional connectivity topology construction module 23, and a spatiotemporal adaptive feature hierarchical fusion module 24. The multimodal feature extraction module 21 extracts the differential entropy (DE) and power spectral density (PSD) of the EEG signal. The feature-enhancing attention coupling module 22 performs complementary fusion of differential entropy (DE) and power spectral density (PSD) through a cross-attention mechanism to generate new enhanced features. The bimodal functional connectivity topology construction module 23 constructs an EEG channel topology by fusing two functional connectivity indices, phase lag index (PLI) and coherence (COH). The fused features are then input as node features into the spatiotemporal adaptive feature hierarchical fusion module 24 to achieve end-to-end fatigue regression.
[0025] The specific working steps of the Multimodal Attention Coupled Graph Convolutional Network (MC-GCN) are as follows:
[0026] 1) The raw EEG signal is input into the multimodal feature extraction module 21. The multimodal feature extraction module 21 preprocesses the EEG signal. First, it is filtered by a 50Hz power frequency notch filter to eliminate AC interference. Then, it is processed by independent component analysis (ICA) to remove artifacts from electrooculography (EOG) and electrocardiogram (ECG). Finally, a 1-49Hz bandpass filter is applied to filter the EEG signal after removing EOG and ECG artifacts, suppressing environmental noise and retaining the emotion-related physiological frequency bands to obtain the preprocessed EEG signal. The preprocessed EEG signal is then divided into segments of 2 seconds each. Data segmentation is performed by a sliding window of length 2, with no overlap between sliding windows. Each segment is considered as a sample.
[0027] 2) The multimodal feature extraction module 21 then extracts the differential entropy (DE) and power spectral density (PSD) of the preprocessed EEG signal; the preprocessed EEG signal approximates a Gaussian distribution. The differential entropy (DE) is expressed as: , For differential entropy, Let be the standard deviation of the EEG signal, and e be the natural constant with a value of 2.71828. For each sample, such as Figure 4 As shown, differential entropy (DE) is calculated independently for each of the five frequency bands. The dimensions of differential entropy (DE) are (number of samples, number of channels, number of frequency bands), where the number of channels is 32 and the number of frequency bands is 5. For power spectral density (PSD) extraction, the Welch method is used for calculation. First, each sample is divided into several sub-segments using a 50% overlapping Hanning window. Then, a Fourier transform is performed on each sub-segment to obtain the power at each frequency point. The formula is K is the number of segments, W is the normalization factor of the window function, and Fk(f) is the FFT value of the k-th segment at frequency f. Then, the power of all segments of each sample is averaged to obtain the power spectral density (PSD) estimate of the sample. The entire frequency band (1-49Hz) is divided into 5 frequency bands, and the average power (or integral power) of each frequency band is calculated as the power spectral density (PSD) feature of that frequency band. The dimension of the power spectral density (PSD) is the same as the dimension of the differential entropy (DE).
[0028] 3) The differential entropy (DE) and power spectral density (PSD) of each sample extracted by the multimodal feature extraction module 21 are input into the feature mutual enhancement attention coupling module 22. The feature mutual enhancement attention coupling module 22 performs feature fusion through dual-branch attention. One branch guides the enhancement of differential entropy (DE) through power spectral density (PSD), setting differential entropy (DE) as the key and power spectral density (PSD) as the query and value, thus realizing the scaling dot product attention mechanism. The expression is: ,specific: This represents the output of the attention mechanism, i.e., the enhanced features obtained after attention computation. Q represents the query matrix, which is the power spectral density (PSD) query matrix in branch one and the differential entropy (DE) query matrix in branch two. K represents the bond matrix, which is the differential entropy (DE) bond matrix in branch one and the power spectral density (PSD) bond matrix in branch two. V represents the value matrix, which is the power spectral density (PSD) value matrix in branch one and the differential entropy (DE) value matrix in branch two. T represents the matrix transpose operation. This represents the size of the last dimension of the query; Branch 2 is a power spectral density (PSD) enhancement guided by differential entropy (DE), setting the power spectral density (PSD) as the key and the differential entropy (DE) as the query and value, implementing the same scaling dot product attention mechanism as Branch 1, such as... Figure 5 As shown, the enhanced features of the two branches are then summed in the frequency dimension to obtain the fused features. The dimension of the fused features and the feature dimension of the differential entropy (DE) are the same as the dimension of the power spectral density (PSD).
[0029] 4) The EEG signal has 32 electrode channels, such as Figure 6As shown, the dual-modal functional connectivity topology construction module 23 treats each electrode channel as a node and determines the connections between nodes through dual-functional indices. The dual-functional indices are the Phase Lag Index (PLI) and the Coherence Index (COH). The PLI quantifies the non-zero phase difference synchronization between two EEG signals, with the core advantage of resisting volume conduction artifacts and avoiding false synchronization caused by scalp electric field diffusion. The COH measures the linear correlation of EEG signal amplitudes between two channels in the same frequency band, reflecting the intensity of coordinated energy oscillations in different brain regions in the same frequency band; essentially, it is a normalized cross-power spectral density. Based on the dual-functional indices, a PLI matrix and a COH matrix of size (32, 32) are calculated. The formula for calculating the PLI matrix is... , This represents the phase difference between channel i and channel j at time t obtained through the Hilbert transform. For the sign function, E[ The symbol represents the average value of the results of the sign function. The formula for calculating the coherence index matrix is: , This represents the cross-power spectral density between channel i and channel j. and Let represent the autopower spectral density of channel i and channel j. Each element of the phase hysteresis index matrix and the coherence index matrix represents the value of the phase hysteresis index (PLI) or coherence index (COH) between the two channels. The values of corresponding elements of the two matrices are nonlinearly fused, and the fusion formula is as follows: z is the sum of the elements of the two matrices, and F(z) represents the element value of the fused matrix. An adaptive threshold is set to the mean of the matrix elements plus twice the standard deviation. Only connections with F(z) values higher than the threshold are retained as valid edges, and elements with values lower than the threshold are set to 0, thereby generating a sparse adjacency matrix. The node features of each node are the fusion features of the corresponding EEG channels.
[0030] 5) such as Figure 7As shown, the spatiotemporal adaptive feature-level fusion module 24 uses the sparse adjacency matrix generated by the dual-modal functional connection topology construction module 23 as the adjacency matrix of the nodes, and the fused features obtained by the feature mutual enhancement attention coupling module 22 as the node features of the corresponding nodes. The spatiotemporal adaptive feature-level fusion module 24 contains two graph convolutional layers, two activation function layers, one global average pooling layer, and one fully connected layer. The first graph convolutional layer has an input dimension of 5 and an output dimension of 128, and the second graph convolutional layer has an input dimension of 128 and an output dimension of 128. The two activation function layers use the Elu activation function, enabling the network to learn and approximate complex nonlinear relationships. The core function of the global average pooling layer is to aggregate node-level features into graph-level features, providing a representation of the entire graph for the subsequent fully connected layer. The input of the fully connected layer is (x, batch), where x represents the output of the previous layer, i.e., the second graph convolutional layer, and batch is the number of samples in the batch. The input dimension of the fully connected layer is 128, and the output dimension is 1. The fully connected layer yields the final fatigue regression value of the driver.
[0031] During driving, the driver wears an EEG signal acquisition device 1 to collect EEG signals in real time and transmit the data to the fatigue monitoring module 2 via a Wi-Fi module 14. The fatigue monitoring module 2 analyzes the EEG signals in real time, monitors the driver's fatigue level, and provides safety assurance during driving.
[0032] The above description, in conjunction with the accompanying drawings, provides a detailed description of an automatic fatigue driving monitoring system based on deep learning technology, but is not intended to limit the scope of the invention. However, the invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various modifications can be made without departing from the spirit of the invention, and all such modifications should be covered within the scope of the claims of the invention.
Claims
1. A fatigue driving automatic monitoring system based on deep learning technology, characterized in that, The device comprises an EEG signal acquisition device (1), a fatigue monitoring module (2), and a voice prompt module (3) connected in series. The EEG signal acquisition device (1) acquires EEG signals from the driver's scalp; the fatigue monitoring module (2) integrates a multimodal attention-coupled graph convolutional network (MC-GCN) to analyze the acquired EEG signals and determine the user's fatigue level in real time; and the voice prompt module (3) provides a voice prompt based on the recognition results of the fatigue monitoring module (2) when the user's fatigue value exceeds a preset safety threshold. The multimodal attention-coupled graph convolutional network MC-GCN in the fatigue monitoring module (2) consists of a multimodal feature extraction module (21), a feature mutual enhancement attention coupling module (22), a bimodal functional connection topology construction module (23), and a spatiotemporal adaptive feature hierarchical fusion module (24) cascaded together. The multimodal feature extraction module (21) extracts the differential entropy and power spectral density of the EEG signal. The feature mutual enhancement attention coupling module (22) performs complementary fusion of differential entropy and power spectral density through a cross-attention mechanism to generate new enhancement features. The bimodal functional connection topology construction module (23) constructs an EEG channel topology by fusing two functional connection indices: phase lag index and coherence. The fused features are input as node features into the spatiotemporal adaptive feature hierarchical fusion module (24) to achieve end-to-end fatigue regression.
2. The fatigue driving automatic monitoring system based on deep learning technology according to claim 1, characterized in that, The EEG signal acquisition device (1) includes: an acquisition module (11) for acquiring EEG signals, an analog-to-digital converter module (12) for amplifying and converting EEG signals, an MCU processor (13) for timing logic control, and a Wi-Fi module (14) for data transmission, which are connected in series. A power supply module (15) is also provided to power the acquisition module (11), the analog-to-digital converter module (12), the MCU processor (13) and the Wi-Fi module (14) respectively. The acquisition module (11) includes an EEG cap and a connecting wire. The electrodes in the EEG cap are directly attached to the driver's scalp and collect EEG signals from the cerebral cortex in real time. The EEG cap is connected to the analog-to-digital conversion module (12) through the connecting wire to transmit the collected EEG signals to the analog-to-digital conversion module (12). The electrode distribution in the EEG cap conforms to the international 10-20 lead system and includes 34 electrodes: FP1, FP2, AF3, AF4, F7, F3, Fz, F4, F8, FC5, FC1, FC2, FC6, T7, C3, Cz, C4, T8, CP5, CP1, CP2, CP6, P7, P3, Pz, P4, P8, O1, Oz, O2, A1, and A2. Among them, A1 is the reference electrode, located at the mastoid process of the left ear, and A2 is ground (GND), located at the mastoid process of the right ear. The input end of the analog-to-digital converter module (12) is connected to the acquisition module (11), and the output end is connected to the MCU processor (13). It includes a high common-mode rejection ratio analog input for receiving EEG signals, a low-noise programmable gain amplifier for amplifying EEG signals, and a high-resolution synchronous sampling analog-to-digital converter connected in sequence. It is used to receive the EEG signals acquired by the acquisition module (11) and amplify and convert the signals into analog signals. The MCU processor (13) is used to adjust the PGA amplification factor and ADC sampling rate of the analog-to-digital converter module (12), control the analog-to-digital converter module (12) to sample at a fixed sampling frequency f = 200Hz, read the EEG signal collected by the analog-to-digital converter module (12) through the SPI bus, and communicate with the wifi module (14) through the UART interface. The MCU processor (13) is equipped with a Cortex-M3 core, operates at a frequency of 72MHz, uses an LQFP-48 package, has 48 pins, and has a size of 7mm×7mm. It is small and lightweight, has 64KB of Flash memory, and has SPI and UART communication interfaces. The Wi-Fi module (14) includes an ESP8266 Wi-Fi chip. The Wi-Fi module operates in transparent transmission mode and transmits the EEG signal received by the MCU processor (13) to the fatigue monitoring module (2) via wireless transmission. The power supply module (15) uses a 5V input voltage and obtains the different operating voltages required by different chips on each module through a voltage conversion chip to power each module.
3. The fatigue driving automatic monitoring system based on deep learning technology according to claim 1, characterized in that, The fatigue monitoring module (2) connects to the same wireless network as the EEG signal acquisition device (1) through the vehicle's built-in Wi-Fi module, establishes a communication link with the EEG signal acquisition device (1), receives the driver's EEG signal collected by the EEG signal acquisition device (1), and monitors the driver's fatigue level in real time. The multimodal attention coupled graph convolutional network MC-GCN deployed in the fatigue monitoring module (2) is a fatigue regression model that analyzes the driver's EEG signal in real time and outputs a continuous fatigue quantification value of 0-100%. When the driver's fatigue level is ≤40%, it means that the driver's physiological state is within the safe threshold and there is no risk of fatigue driving in the short term. When the driver's fatigue level is between 40% and 80%, it means that the driver is in a moderate fatigue state. When the driver's fatigue level is >80%, it means that the driver is in a severe fatigue state and there is a risk of car accident at any time. The voice prompt module (3) broadcasts the voice.
4. The fatigue driving automatic monitoring system based on deep learning technology according to claim 1, characterized in that, The specific working steps of the Multimodal Attention Coupled Graph Convolutional Network (MC-GCN) are as follows: 1) Input the original EEG signal into the multimodal feature extraction module (21). The multimodal feature extraction module (21) preprocesses the EEG signal. First, it eliminates AC interference by 50Hz power frequency notch filtering. Then, it removes EOG and ECG artifacts by independent component analysis (ICA). Finally, it applies a 1-49Hz bandpass filter to filter the EEG signal after removing EOG and ECG artifacts, suppresses environmental noise, and retains the emotion-related physiological frequency bands to obtain the preprocessed EEG signal. The preprocessed EEG signals were then divided into 2-second segments, and the data was segmented using a sliding window of length 2. The sliding windows did not overlap, and each segment was used as a sample. 2) Multimodal feature extraction module (21) then extracts the differential entropy and power spectral density of the preprocessed EEG signal; the preprocessed EEG signal approximates a Gaussian distribution. The differential entropy is expressed as: , For differential entropy, Let be the standard deviation of the EEG signal, and e be the natural constant with a value of 2.71828. For each sample, the differential entropy is calculated independently for each of the five frequency bands. The dimensions of the differential entropy are (number of samples, number of channels, number of frequency bands), where the number of channels is 32 and the number of frequency bands is 5. For power spectral density extraction, the Welch method is used. First, each sample is divided into several segments using a 50% overlapping Hanning window. Then, a Fourier transform is performed on each segment to obtain the power at each frequency point. The formula is K is the number of segments, W is the normalization factor of the window function, Fk(f) is the FFT value of the k-th segment at frequency f, and the power of all segments of each sample is averaged to obtain the power spectral density estimate of the sample. The entire frequency band is divided into 5 frequency bands, and the average power or integral power of each frequency band is calculated as the power spectral density feature of that frequency band. The dimension of the power spectral density is the same as the dimension of the differential entropy. 3) Input the differential entropy and power spectral density of each sample extracted by the multimodal feature extraction module (21) into the feature mutual enhancement attention coupling module (22). The feature mutual enhancement attention coupling module (22) performs feature fusion through dual-branch attention. The first branch guides the enhancement of differential entropy through power spectral density, sets the differential entropy as the key, and sets the power spectral density as the query and value, thereby realizing the scaling dot product attention mechanism. The expression is: ,specific: This represents the output of the attention mechanism, i.e., the enhanced features obtained after attention computation. Q represents the query matrix: power spectral density query matrix in branch one, and differential entropy query matrix in branch two. K represents the key matrix: Ki is the differential entropy key matrix in branch one, and power spectral density key matrix in branch two. V represents the index matrix: Vi is the power spectral density value matrix in branch one, and differential entropy value matrix in branch two. T represents the matrix transpose operation. The last dimension of the query is represented; the second branch is power spectral density enhancement guided by differential entropy. The power spectral density is set as the key and the differential entropy is set as the query and value, realizing the same scaling dot product attention mechanism as the first branch. Then, the enhanced features of the two branches are summed in the frequency dimension to obtain the fused features. The dimension of the fused features and the feature dimension of the differential entropy (DE) are the same as the dimension of the power spectral density. 4) The EEG signal has 32 electrode channels. The dual-modal functional connection topology construction module (23) treats each electrode channel as a node and determines the connection between nodes through dual-function indices. The dual-function indices are the phase lag index and the coherence index. The phase lag index is used to quantify the non-zero phase difference synchronization between two EEG signals. Its core advantage is that it is resistant to volume conduction artifacts and can avoid false synchronization caused by the diffusion of the scalp electric field. The coherence index is used to measure the linear correlation of the amplitude of the EEG signals between the two channels in the same frequency band, reflecting the intensity of energy coordinated oscillation in different brain regions in the same frequency band. The phase lag index matrix and coherence index matrix of size (32, 32) are calculated based on the bifunctional index. The formula for calculating the phase lag index matrix is as follows: , This represents the phase difference between channel i and channel j at time t obtained through the Hilbert transform. For the sign function, E[ The symbol represents the average value of the results of the sign function. The formula for calculating the coherence index matrix is: , This represents the cross-power spectral density between channel i and channel j. and Let represent the autopower spectral density of channel i and channel j. Each element of the phase lag exponent matrix and the coherence exponent matrix represents the value of the phase lag exponent (PLI) or coherence exponent between the two channels. The values of corresponding elements of the two matrices are nonlinearly fused using the following formula: z is the sum of the elements of the two matrices, and F(z) represents the element value of the fused matrix. An adaptive threshold is set to the mean of the matrix elements plus twice the standard deviation. Only connections with F(z) values higher than the threshold are retained as valid edges, and elements with values lower than the threshold are set to 0, thereby generating a sparse adjacency matrix. The node features of each node are the fusion features of the corresponding EEG channels. 5) The spatiotemporal adaptive feature hierarchical fusion module (24) uses the sparse adjacency matrix generated by the dual-modal functional connection topology construction module (23) as the adjacency matrix of the node, and the fused features obtained by the feature mutual enhancement attention coupling module (22) as the node features of the corresponding node. The spatiotemporal adaptive feature layer fusion module (24) contains two graph convolutional layers, two activation function layers, one global average pooling layer, and one fully connected layer. The first graph convolutional layer has an input dimension of 5 and an output dimension of 128. The second graph convolutional layer has an input dimension of 128 and an output dimension of 128. The two activation function layers use the Elu activation function, which enables the network to learn and approximate complex nonlinear relationships. The core function of the global average pooling layer is to aggregate node-level features into graph-level features, providing a representation of the entire graph for the subsequent fully connected layer. The input of the fully connected layer is (x, batch), where x represents the output of the previous layer, i.e., the second graph convolutional layer, and batch is the number of samples in the batch. The input dimension of the fully connected layer is 128, and the output dimension is 1. The fully connected layer yields the final fatigue regression value of the driver.
Citation Information
Patent Citations
Wearable fatigue driving monitoring system and method based on forehead electroencephalogram signals
CN111329497A
Channel screening method, emotion recognition method and system and storage medium
CN114818786A