Method for detecting fatigue state based on multi-modal data fusion technology
By integrating fNIRS, ECG, PPG, GSR, and Resp signals through multimodal data fusion technology and deep learning models, the accuracy problem of single-modal detection is solved, enabling efficient fatigue state identification and alerts, and ensuring driving safety.
Patent Information
- Application Number
- CN202511344081.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-01-16
AI Technical Summary
Existing fatigue detection methods mostly employ single-modal detection, and their accuracy is affected by individual differences, environmental conditions, and noise. The efficiency and accuracy of traditional methods need to be improved, and existing multimodal detection methods have failed to effectively integrate complementary information from multiple biological signals.
Employing multimodal data fusion technology, this system collects various physiological signals through an fNIRS headband, ECG sensor, PPG wristband, GSR ring, and respiratory wave sensor. It then combines these signals with a deep learning model for feature fusion and utilizes the cross-attention and self-attention mechanisms of the Transformer architecture to integrate the signals, thereby constructing a fatigue detection system that provides timely alerts when fatigue is detected.
It achieves effective integration of multimodal data, improves the accuracy and efficiency of fatigue detection, and can promptly identify and alert drivers to fatigue, ensuring driving safety.
Smart Images

Figure CN121337352A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fatigue detection technology, specifically a method for detecting fatigue state based on multimodal data fusion technology. Background Technology
[0002] Fatigue is one of the main factors causing serious accidents. In some fields requiring high alertness and precise operation, mental fatigue can lead to fatal accidents. Therefore, fatigue detection in the human body is urgent and necessary.
[0003] Currently, fatigue monitoring is mostly based on functional near-infrared brain imaging (fNIRS). After collecting relevant signal data, feature recognition is performed to ultimately predict fatigue. For example, vehicle fatigue monitoring systems provide early warnings by monitoring the driver's state in real time from multiple dimensions: 1. Facial feature analysis: using an infrared camera located on the A-pillar or above the steering wheel; 2. capturing the driver's blinking frequency, eye closure duration, and yawning; 3. Head posture detection: the camera simultaneously tracks the head's angle of movement; frequent head lowering or prolonged stillness is considered a sign of inattention. However, most existing solutions employ single-modal detection, while the human body contains multiple biosignals that can provide complementary information. Some solutions also utilize multimodal detection, extracting and integrating signals to detect the human state. Although fNIRS technology has shown potential in fatigue detection, the accuracy of fNIRS signals is affected by various factors, including individual differences, environmental conditions, and noise during the measurement process. Furthermore, traditional rule-based and threshold-based fatigue detection methods require improvement in both efficiency and accuracy. Therefore, a method based on multimodal data fusion technology for detecting fatigue state is proposed. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method for detecting fatigue state based on multimodal data fusion technology, thus solving the problems mentioned in the background section.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for detecting fatigue state based on multimodal data fusion technology, comprising the following steps: S1: Multimodal training data was collected through sleep deprivation experiments for model training. S11: Before the experiment began, the subjects closed their eyes and sat quietly. Two minutes later, the subjects were informed that the experiment was about to begin and that cognitive tests would be conducted and data would be collected at the same time. S12: During the fatigue induction phase, subjects who have a habit of taking a nap will be deprived of sleep and made to sit and read for 40 minutes, during which the subjects will gradually become fatigued. S13: After 40 minutes of quiet reading, participants were asked to repeat the cognitive test from the first phase and fill out the fatigue scale for the last time after the test. S14: Combine the assessment scales completed by the subjects with the collected behavioral data as data labels to determine the convergence direction of the behavioral data intervention model; S2: During driving, the driver uses multiple sensors to acquire data on hemoglobin concentration, electrocardiogram, pulse wave, skin conductance, and respiratory wave in the prefrontal cortex of the brain; S3: Preprocess the acquired physiological signals, including: removing abnormal channels from fNIRS, filtering the data to obtain processed data, segmenting the data, and using PCA to reduce the dimensionality of the segmented data to obtain preliminary feature data. The specific method is as follows: S31: Preprocess fNIRS; S32: Preprocessing of electrocardiogram (ECG); S33: Preprocess the PPG signal; S34: Preprocessing of the skin conductance response (GSR) signal; S35: Preprocess the respiratory wave (Resp) signal; S4: Construct a deep learning model to fuse multimodal (fNIRS, ECG, PPG, GSR, Resp) data to detect fatigue state. The specific method is as follows: S41: Divide the preprocessed multimodal data and the multimodal data collected in step one into a training set and a test set; S42: Building a deep learning model based on the Transformer architecture: First, the physiological signals are preprocessed by performing one-dimensional convolution on the input time series of each modality to handle the differences in sampling rate between different signals, ensuring that the input data have the same scale and characteristics in the time dimension during subsequent analysis or modeling. To mitigate inconsistencies between different auxiliary modalities, information needs to be transferred between them. A cross-attention mechanism is used to fuse features from each auxiliary modality. The formula for the cross-attention mechanism is as follows: (1) Where Q is the query vector, K is the key vector, V is the value vector, and dk is the dimension of the key vector; to enhance the fNIRS features by optimizing the features of each auxiliary modality, it is necessary to realize the information transfer from each auxiliary modality to the main modality. A cross-attention mechanism is adopted, taking the main modality as the main feature and embedding the other optimized auxiliary modality features into the main modality features to obtain the fused fNIRS features. In order to make full use of the common and complementary characteristics contained in the enhanced fNIRS features of different auxiliary modalities, it is subjected to self-attention transformation, that is, Q, K, V are all from the fused fNIRS modality. The self-attention mechanism allows each input feature of the sequence to interact with all other features in the sequence, thus capturing global contextual information, thereby obtaining the enhanced fNIRS features; Considering that different physiological signals are complementary in terms of the fatigue state of the subjects, we first use one-dimensional convolution to reduce the dimensionality, and then connect the inconsistent auxiliary modal features with the enhanced fNIRS features to obtain fused features. Finally, we use a classifier composed of fully connected layers and activation functions to predict and classify fatigue labels based on the fused features, which are divided into mild, moderate and severe. Step 4.3: Using the constructed deep learning model, train it with the preprocessed data and corresponding labels, evaluate and test the model's performance on the test set, confirm the model's effectiveness, and evaluate the model's performance using metrics such as accuracy, recall, and F1 score. Step 5: When the system detects that the user has entered a state of fatigue and this state persists for a period of time, it will automatically trigger a series of targeted reminder measures to ensure the driver's safety and intervene in the fatigue state in a timely manner. The detection process is as follows: Based on the fatigue level predicted by S4, the system will issue corresponding reminders to the driver. When the system detects that the driver is in a mild fatigue stage, it will issue a gentle voice reminder through the car audio system. If the system recognizes that the driver has entered a moderate fatigue stage, the reminder measures will be more serious. At this time, the system will still issue a voice prompt through the car audio system, but the tone and frequency will be more noticeable to ensure that the driver can perceive the change in his / her state in a timely manner. When the driver's fatigue level reaches a severe level, the system will apply a weak electrical stimulation through the headband worn by the driver to physically attract the driver's attention. At the same time, it will issue an emergency voice reminder, such as "You are in a state of severe fatigue, please stop and rest immediately."
[0006] Preferably, the cognitive test consists of two modules: a Schulte grid and an n-back module. The Schulte grid is composed of 6×6 and 7×7 number matrices, respectively. The test taker needs to quickly click on the numbers appearing on the screen in sequence. The n-back module consists of two parts: 1-back and 2-back. When performing the 1-back task, the test taker needs to determine whether the color position of the nth square on the screen is the same as that of the (n-1)th square. When performing the 2-back task, the test taker needs to determine whether the color position of the nth square is the same as that of the (n-2)th square. The test taker repeats the given task twice. The cognitive test lasts for approximately 20 minutes.
[0007] Preferably, in S2: During driving, the driver obtains data on hemoglobin concentration, electrocardiogram, pulse wave, skin conductance and respiratory wave in the prefrontal cortex of the brain through an fNIRS headband, an electrocardiogram (ECG) sensor (chest band), a PPG wristband, a GSR ring, and a respiratory wave (Resp) sensor (waistband).
[0008] Preferably, the preprocessing method for fNIRS in S31 is as follows: after frequency modulation, the collected fNIRS light intensity signal first needs to be orthogonally demodulated, and the original signal is decomposed using empirical mode decomposition (EMD) to remove motion artifacts and spur noise; after preprocessing the original light intensity signal, it is converted into the relative concentration of hemoglobin according to the modified Beer-Lambert formula.
[0009] Preferably, S32: the preprocessing method for the electrocardiogram (ECG) is as follows: the noise of the ECG signal includes power frequency noise, baseline drift noise and electromyographic interference noise. Wavelet transform is used to preprocess the ECG signal to remove noise and retain the main frequency band of the ECG signal; a low-order polynomial is used to fit and remove the baseline drift; frequency domain filtering, independent component analysis (ICA) and other methods are used to remove artifacts, and finally the processed ECG signal is obtained. The R wave position of the ECG signal is determined. Based on the detected R wave position, the start and end points of each cardiac cycle are determined. Based on the R wave position and the defined window, the corresponding cardiac cycle signal segments are extracted from the original signal, each extracted signal segment is saved and marked.
[0010] Preferably, the preprocessing method for the PPG signal in S33 is as follows: the noise in the PPG signal includes: baseline drift, power frequency interference, electromyographic noise, and motion artifacts; wavelet denoising is used to remove most of the high-frequency and low-frequency noise; the processed PPG signal is segmented, and the length of each PPG signal segment is set to N sampling points.
[0011] Preferably, the preprocessing method for the GSR signal in step S34 is as follows: The GSR signal mainly includes physiological noise, environmental noise, and motion artifacts. First, a bandpass filter is used to filter the GSR signal to remove high-frequency noise (such as electromagnetic interference) and low-frequency noise (such as baseline drift). Then, a polynomial curve is used to fit the baseline of the signal and remove it from the original signal. For motion shadows, an automated algorithm, such as a peak detection algorithm, is used to identify abnormal peaks and automatically remove or repair the data. Finally, the slowly changing skin conductance level and the rapidly changing skin conductance level data are obtained.
[0012] Preferably, the preprocessing method for the respiratory wave (Resp) signal in S35 is as follows: wavelet transform is performed on the acquired respiratory wave signal to remove noise from the signal and extract useful respiratory signals. In order to eliminate the difference in the amplitude of respiratory signals between different individuals, the signal is compressed to the range of 0 to 1 for normalization. The respiratory cycle is found by marking the inhalation and exhalation process of breathing by finding local maxima and minima in the respiratory signal. The respiratory wave signal is segmented based on the cycle to facilitate subsequent analysis of the respiratory frequency within the same time period.
[0013] This invention provides a method for detecting fatigue state based on multimodal data fusion technology, which has the following beneficial effects: This method for detecting fatigue state based on multimodal data fusion technology has the following advantages: 1. By collecting multimodal data, including fNIRS, ECG, PPG, GSP, Resp, and data on hemoglobin concentration in the prefrontal cortex obtained through an fNIRS headband, which is significantly correlated with fatigue, and by acquiring ECG, pulse wave, skin conductance, and respiratory wave data through ECG sensors (chest belt), PPG wristbands, GSR rings, and respiratory wave (Resp) sensors (waist belt), the data collected by different devices are objective, effectively supplementing the deficiencies of psychological experimental methods. Moreover, different modalities can provide complementary information and even capture the characteristics of the target from multiple angles. 2. Using fNIRS as the primary modal signal and others as auxiliary modal signals, a deep learning model based on the Transformer architecture is used to effectively integrate signals from different modalities. The auxiliary modalities interact using a cross-attention mechanism, and the results are weighted and integrated into the fNIRS modality. Thus, fatigue detection is performed through a multimodal data (fNIRS, ECG, PPG, GSR, Resp) fusion strategy. 3. Issue reminders and warnings upon detecting that a user is in a state of fatigue, and promptly notify and alert the user. Attached Figure Description
[0014] Figure 1This is a schematic diagram of the process of the present invention; Figure 2 This is a diagram of the deep learning model architecture of the present invention. Detailed Implementation
[0015] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0016] Please see Figures 1 to 2 This invention provides a technical solution: a method for detecting fatigue state based on multimodal data fusion technology, comprising the following steps: S1: Multimodal training data was collected through a sleep deprivation experiment for model training. The specific method is as follows: S11: Before the experiment began, the subjects sat quietly with their eyes closed. Two minutes later, the subjects were informed that the experiment had begun and data was collected at the same time. The cognitive test consisted of two modules: the Schulte Grid and the n-back. The Schulte Grid consisted of 6×6 and 7×7 number matrices, and the subjects had to quickly click on the numbers that appeared on the screen in sequence. The n-back consisted of two parts: 1-back and 2-back. In the 1-back task, the subjects had to determine whether the color and position of the nth square on the screen were the same as the (n-1)th square. In the 2-back task, the subjects had to determine whether the color and position of the nth square were the same as the (n-2)th square. The subjects repeated the given task twice. The cognitive test lasted approximately 20 minutes.
[0017] S12: During the fatigue induction phase, subjects who have a habit of taking a nap will be deprived of sleep and made to sit and read for 40 minutes, during which the subjects will gradually become fatigued.
[0018] S13: After 40 minutes of quiet reading, participants were asked to repeat the cognitive test from the first phase and fill out a fatigue scale for the last time after the test.
[0019] S14: Combine the assessment scales completed by the subjects with the collected behavioral data as data labels to achieve the convergence direction of the behavioral data intervention model.
[0020] S2: During driving, the driver obtains data on hemoglobin concentration, ECG, pulse wave, skin conductance, and respiratory wave from the prefrontal cortex of the brain through an fNIRS headband, an ECG sensor (chest band), a PPG wristband, a GSR ring, and a respiratory wave (Resp) sensor (waistband).
[0021] S3: Preprocessing of the acquired physiological signals includes: removing outlier channels from fNIRS, filtering the data to obtain processed data, segmenting the data, and using PCA to reduce the dimensionality of the segmented data to obtain preliminary feature data. The specific method is as follows: S31: The preprocessing method for fNIRS is as follows: The acquired fNIRS light intensity signal is frequency modulated, so orthogonal demodulation processing is required first. In the acquired raw fNIRS data, there are three main sources of noise: instrument noise, motion artifacts, and physiological noise. Instrument noise can be effectively suppressed by a low-pass filter. Physiological noise includes heartbeat, respiration, and Mayer waves; signals generated by these physiological processes need to be reduced using specific filtering techniques. For motion artifact correction, empirical mode decomposition (EMD) is used to decompose the raw signal to remove motion artifacts and spurious noise. After preprocessing the raw light intensity signal, it is converted into the relative concentration of hemoglobin according to the modified Beer-Lambert formula.
[0022] S32: The preprocessing method for electrocardiogram (ECG) is as follows: ECG signal noise includes power line noise, baseline drift noise, and electromyographic interference noise. Wavelet transform is used to preprocess the ECG signal to remove noise and retain the main frequency bands of the ECG signal. Low-order polynomials are used to fit and remove baseline drift. Frequency domain filtering and independent component analysis (ICA) are used to remove artifacts, finally obtaining the processed ECG signal. The R-wave position of the ECG signal is determined. Based on the detected R-wave position, the start and end points of each cardiac cycle are determined. According to the R-wave position and the defined window, the corresponding cardiac cycle signal segments are extracted from the original signal. Each extracted signal segment is saved and labeled.
[0023] S33: The preprocessing method for PPG signals is as follows: Noise in PPG signals includes baseline drift, power line interference, electromyographic noise, and motion artifacts; wavelet denoising is used to remove most of the high-frequency and low-frequency noise; the processed PPG signal is segmented, and the length of each PPG signal segment is set to N sampling points. The segmentation process can be defined as follows: The acquired PPG signal is ?(?), where ? represents the sequence of sampling points. The starting point representing each sub-signal segment in the PPG waveform is identified, usually by searching for the local minimum point of the waveform. This process is iteratively executed until the segmentation of the entire PPG signal is completed; S34: The preprocessing method for the Gestational Skin Response (GSR) signal is as follows: The GSR signal mainly includes physiological noise, environmental noise, and motion artifacts. First, a bandpass filter is used to filter the GSR signal to remove high-frequency noise (such as electromagnetic interference) and low-frequency noise (such as baseline drift). Then, a polynomial curve is used to fit the baseline of the signal and remove it from the original signal. For motion artifacts, an automated algorithm, such as a peak detection algorithm, is used to identify abnormal peaks and automatically remove or repair this part of the data. Finally, data on slowly changing and rapidly changing skin conductance levels are obtained. S35: The preprocessing method for the respiratory wave (Resp) signal is as follows: Wavelet transform is performed on the acquired respiratory wave signal to remove noise and extract useful respiratory signals. To eliminate differences in respiratory signal amplitude between individuals, the signal is compressed to the range of 0 to 1 for normalization. The respiratory cycle is determined by identifying local maxima and minima in the respiratory signal to mark the inhalation and exhalation processes. The respiratory wave signal is then segmented based on the cycle to facilitate subsequent analysis of respiratory frequencies within the same time period. S4: Construct a deep learning model to fuse multimodal (fNIRS, ECG, PPG, GSR, Resp) data to detect fatigue state. The specific method is as follows: S41: Divide the preprocessed multimodal data and the multimodal data collected in step one into a training set and a test set.
[0024] S42: Build a deep learning model based on the Transformer architecture, such as Figure 2 As shown: Given the significant differences in sampling rates among different physiological signals, preprocessing is necessary. One-dimensional convolution is performed on the input time series of each modality to address the differences in sampling rates between signals, ensuring that the input data maintains the same scale and characteristics in the time dimension during subsequent analysis or modeling. To mitigate inconsistencies between different auxiliary modalities, information transfer between them is necessary. This can be achieved through a cross-attention mechanism that fuses features from each auxiliary modality. The formula for the cross-attention mechanism is: (1) Where Q is the query vector, K is the key vector, V is the value vector, and dk is the dimension of the key vector; cross-attention is an attention mechanism used for feature interaction and fusion in multimodal or multi-source information. Its core idea is to facilitate information interaction between different auxiliary modalities. Specifically, it calculates Q from a specific auxiliary modality with K and V from another auxiliary modality to achieve feature fusion and obtain optimized auxiliary modality features. Secondly, to reduce information redundancy between specific auxiliary modality features and other auxiliary modality features, the model's learning rate needs to be adjusted to improve model performance and learning efficiency. Through feature fusion, the inconsistency and information redundancy between auxiliary modalities will be significantly reduced.
[0025] To enhance fNIRS features using features optimized from various auxiliary modalities, it's necessary to achieve information transfer from each auxiliary modality to the main modality. Similarly, a cross-attention mechanism is employed, using the main modality as the primary feature and embedding the optimized auxiliary modal features into it, resulting in the fused fNIRS features. Secondly, to fully utilize the common and complementary characteristics of the enhanced fNIRS features from different auxiliary modalities, a self-attention transformation is performed. This means that Q, K, and V are all derived from the fused fNIRS modality. The self-attention mechanism allows each input feature in the sequence to interact with all other features in the sequence, thus capturing global contextual information and obtaining the enhanced fNIRS features. Considering the complementarity of different physiological signals in terms of subject fatigue state, a one-dimensional convolutional model was first used to reduce dimensionality, and then inconsistent auxiliary modal features were concatenated with enhanced fNIRS features to obtain fused features. Finally, a classifier composed of fully connected layers and activation functions was used to predict and classify fatigue labels based on the fused features, classifying them into mild, moderate, and severe.
[0026] S43: Using the constructed deep learning model, train it with preprocessed data and corresponding labels. Evaluate and test the model's performance on the test set to confirm its effectiveness. Use metrics such as accuracy, recall, and F1 score to evaluate model performance.
[0027] S5: When the system detects that the user has entered a state of fatigue and this fatigue persists for a period of time, it will automatically trigger a series of targeted reminder measures to ensure the driver's safety and intervene in their fatigue state in a timely manner. The method is as follows: Based on the fatigue level predicted in step 4, the system will provide corresponding reminders to the driver. When the system detects that the driver is in a mild fatigue stage, it will issue a gentle voice reminder through the vehicle's audio system. If the system identifies that the driver has entered a moderate fatigue stage, the reminder measures will be more serious. At this time, the system will still issue a voice prompt through the vehicle's audio system, but the tone and frequency will be more noticeable to ensure that the driver can perceive the change in their state in a timely manner. When the driver's fatigue level reaches a severe level, the system will apply weak electrical stimulation through the headband worn by the driver to physically attract the driver's attention. At the same time, an emergency voice reminder will be issued, such as "You are in a state of severe fatigue, please stop and rest immediately."
[0028] Example 1: System hardware design based on this method: The system can use a microcontroller as the main control chip to directly control the system and realize the integration of components; for the headband device, a more light-shielding and comfortable flexible material can be selected; Example 2: Utilizing a newer and more advanced neural network framework to perform feature extraction and fatigue monitoring tasks, and employing new encoding methods, such as transforming one-dimensional time series into a two-dimensional graphical framework, and using CNNs, GNNs, etc. for high-dimensional feature extraction to achieve multi-class recognition of fatigue levels; Example 3: To improve the accuracy of fatigue monitoring, multimodal fusion can be explored, for example, by combining fNIRS technology with other technologies; Example 4: To improve portability: In the system design of Example 1, the system integration is increased and the overall size of the system hardware is reduced by improving the driving module, the acquisition module and the corresponding power supply circuit, and adopting a wireless transmission module, so as to make the system portable.
[0029] All standard parts used in this application can be purchased from the market, and can be customized according to the description and drawings. The specific connection methods of each part adopt conventional methods such as bolts, rivets, and welding that are mature in the prior art. The machinery, parts and equipment all adopt conventional models in the prior art. The installation methods between equipment are also the same as conventional installation methods in the prior art. For example, the two ends of shaft-shaped parts are connected by bearings, the connection position of valve components is provided with anti-leakage rubber strips, the outside of threaded rods or lead rods is provided with dust covers, and the equipment can be driven by either built-in batteries or external power supply. The control method is automatic control by a controller. The control circuit of the controller can be implemented by simple programming by those skilled in the art and is common knowledge in the field. Since this invention is mainly used to protect mechanical devices, this invention will not explain the control method and circuit connection in detail. The external controller mentioned in the specification can play a control role for the electrical components mentioned herein, and the external controller is a conventional known device.
[0030] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for detecting fatigue state based on multimodal data fusion technology, characterized in that: Includes the following steps: S1: Multimodal training data was collected through sleep deprivation experiments for model training. S11: Before the experiment began, the subjects closed their eyes and sat quietly. Two minutes later, the subjects were informed that the experiment was about to begin and that cognitive tests would be conducted and data would be collected at the same time. S12: During the fatigue induction phase, subjects who have a habit of taking a nap will be deprived of sleep and made to sit and read for 40 minutes, during which the subjects will gradually become fatigued. S13: After 40 minutes of quiet reading, participants were asked to repeat the cognitive test from the first phase and fill out the fatigue scale for the last time after the test. S14: Combine the assessment scales completed by the subjects with the collected behavioral data as data labels to determine the convergence direction of the behavioral data intervention model; S2: During driving, the driver uses multiple sensors to acquire data on hemoglobin concentration, electrocardiogram, pulse wave, skin conductance, and respiratory wave in the prefrontal cortex of the brain; S3: Preprocess the acquired physiological signals, including: removing abnormal channels from fNIRS, filtering the data to obtain processed data, segmenting the data, and using PCA to reduce the dimensionality of the segmented data to obtain preliminary feature data. The specific method is as follows: S31: Preprocess fNIRS; S32: Preprocessing of electrocardiogram (ECG); S33: Preprocess the PPG signal; S34: Preprocessing of the skin conductance response (GSR) signal; S35: Preprocess the respiratory wave (Resp) signal; S4: Construct a deep learning model to fuse multimodal (fNIRS, ECG, PPG, GSR, Resp) data to detect fatigue state. The specific method is as follows: S41: Divide the preprocessed multimodal data and the multimodal data collected in step one into a training set and a test set; S42: Building a deep learning model based on the Transformer architecture: First, the physiological signals are preprocessed by performing one-dimensional convolution on the input time series of each modality to handle the differences in sampling rate between different signals, ensuring that the input data have the same scale and characteristics in the time dimension during subsequent analysis or modeling. To mitigate inconsistencies between different auxiliary modalities, information needs to be transferred between them. A cross-attention mechanism is used to fuse features from each auxiliary modality. The formula for the cross-attention mechanism is as follows: ;(1) Where Q is the query vector, K is the key vector, V is the value vector, and dk is the dimension of the key vector; to enhance the fNIRS features by optimizing the features of each auxiliary modality, it is necessary to realize the information transfer from each auxiliary modality to the main modality. A cross-attention mechanism is adopted, taking the main modality as the main feature and embedding the other optimized auxiliary modality features into the main modality features to obtain the fused fNIRS features. In order to make full use of the common and complementary characteristics contained in the enhanced fNIRS features of different auxiliary modalities, it is subjected to self-attention transformation, that is, Q, K, V are all from the fused fNIRS modality. The self-attention mechanism allows each input feature of the sequence to interact with all other features in the sequence, thus capturing global contextual information, thereby obtaining the enhanced fNIRS features; Considering that different physiological signals are complementary in terms of the fatigue state of the subjects, we first use one-dimensional convolution to reduce the dimensionality, and then connect the inconsistent auxiliary modal features with the enhanced fNIRS features to obtain fused features. Finally, we use a classifier composed of fully connected layers and activation functions to predict and classify fatigue labels based on the fused features, which are divided into mild, moderate and severe. Step 4.3: Using the constructed deep learning model, train it with the preprocessed data and corresponding labels, evaluate and test the model's performance on the test set, confirm the model's effectiveness, and evaluate the model's performance using metrics such as accuracy, recall, and F1 score. Step 5: When the system detects that the user has entered a state of fatigue and this state persists for a period of time, it will automatically trigger a series of targeted reminder measures to ensure the driver's safety and intervene in the fatigue state in a timely manner. The detection process is as follows: Based on the fatigue level predicted by S4, the system will issue corresponding reminders to the driver. When the system detects that the driver is in a mild fatigue stage, it will issue a gentle voice reminder through the car audio system. If the system recognizes that the driver has entered a moderate fatigue stage, the reminder measures will be more serious. At this time, the system will still issue a voice prompt through the car audio system, but the tone and frequency will be more noticeable to ensure that the driver can perceive the change in his / her state in a timely manner. When the driver's fatigue level reaches a severe level, the system will apply a weak electrical stimulation through the headband worn by the driver to physically attract the driver's attention. At the same time, it will issue an emergency voice reminder, such as "You are in a state of severe fatigue, please stop and rest immediately." 2. The method for detecting fatigue state based on multimodal data fusion technology according to claim 1, characterized in that: The cognitive test consists of two modules: the Schulte Grid and the n-back. The Schulte Grid is composed of 6×6 and 7×7 number matrices, respectively. Participants need to quickly click on the numbers that appear on the screen in sequence. The n-back consists of two parts: 1-back and 2-back. In the 1-back task, participants need to determine whether the color and position of the nth square on the screen are the same as the (n-1)th square. In the 2-back task, participants need to determine whether the color and position of the nth square are the same as the (n-2)th square. Participants repeat the given task twice. The cognitive test lasts approximately 20 minutes.
3. The method for detecting fatigue state based on multimodal data fusion technology according to claim 1, characterized in that: S2: During driving, the driver obtains data on hemoglobin concentration, ECG, pulse wave, skin conductance, and respiratory wave in the prefrontal cortex of the brain through an fNIRS headband, an ECG sensor (chest band), a PPG wristband, a GSR ring, and a respiratory wave (Resp) sensor (waistband).
4. The method for detecting fatigue state based on multimodal data fusion technology according to claim 1, characterized in that: S31: The preprocessing method for fNIRS is as follows: After frequency modulation, the collected fNIRS light intensity signal first needs to be orthogonally demodulated. Empirical mode decomposition (EMD) is used to decompose the original signal to remove motion artifacts and spur noise. After preprocessing the original light intensity signal, it is converted into the relative concentration of hemoglobin according to the modified Beer-Lambert formula.
5. The method for detecting fatigue state based on multimodal data fusion technology according to claim 1, characterized in that: S32: The preprocessing method for electrocardiogram (ECG) is as follows: ECG signal noise includes power frequency noise, baseline drift noise, and electromyographic interference noise. Wavelet transform is used to preprocess the ECG signal to remove noise and retain the main frequency band of the ECG signal. Low-order polynomials are used to fit and remove baseline drift. Frequency domain filtering, independent component analysis (ICA), and other methods are used to remove artifacts. Finally, the processed ECG signal is obtained. The R-wave position of the ECG signal is determined. Based on the detected R-wave position, the start and end points of each cardiac cycle are determined. Based on the R-wave position and the defined window, the corresponding cardiac cycle signal segments are extracted from the original signal. Each extracted signal segment is saved and marked.
6. The method for detecting fatigue state based on multimodal data fusion technology according to claim 1, characterized in that: S33: The preprocessing method for the PPG signal is as follows: the noise in the PPG signal includes: baseline drift, power frequency interference, electromyographic noise, and motion artifacts; wavelet denoising is used to remove most of the high-frequency and low-frequency noise; the processed PPG signal is segmented, and the length of each PPG signal segment is set to N sampling points.
7. The method for detecting fatigue state based on multimodal data fusion technology according to claim 1, characterized in that: S34: The preprocessing method for the skin conductance response (GSR) signal is as follows: The skin conductance response (GSR) signal mainly includes physiological noise, environmental noise, and motion artifacts; First, a bandpass filter is used to filter the GSR signal to remove high-frequency noise (such as electromagnetic interference) and low-frequency noise (such as baseline drift). Then, a polynomial curve is used to fit the baseline of the signal and remove it from the original signal. For motion shadows, an automated algorithm, such as a peak detection algorithm, is used to identify abnormal peaks and automatically remove or repair this part of the data. Finally, the slowly changing skin conductance level and the rapidly changing skin conductance level data are obtained.
8. The method for detecting fatigue state based on multimodal data fusion technology according to claim 1, characterized in that: S35: The preprocessing method for the respiratory wave (Resp) signal is as follows: wavelet transform is performed on the acquired respiratory wave signal to remove noise from the signal and extract useful respiratory signals. In order to eliminate the difference in the amplitude of respiratory signals between different individuals, the signal is compressed to the range of 0 to 1 for normalization. The respiratory cycle is found by marking the inhalation and exhalation process of breathing by finding local maxima and minima in the respiratory signal. The respiratory wave signal is segmented based on the cycle to facilitate subsequent analysis of the respiratory frequency within the same time period.
Citation Information
Patent Citations
Physical fatigue detection method based on brain network topology law and system
CN112656373A
Driver state recognition system and method
CN120482064A
Cited By
User fatigue state intervention method and device, equipment and medium
CN121608751A