Vehicle-mounted radar driver face fatigue detection method based on deep learning
By applying deep learning technology in vehicle-mounted radar, combining mean decomposition and multi-sequence variational modal decomposition algorithms, the light dependence and low signal-to-noise ratio problems of existing driver fatigue detection methods are successfully solved, and high-precision detection of the tiny and fine-grained fatigue characteristics of the driver's face is achieved.
Patent Information
- Application Number
- CN202510081329.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
The existing driver fatigue detection methods have problems such as light dependence, privacy violation, discomfort and low signal-to-noise ratio, making it difficult to effectively identify the tiny and fine-grained fatigue characteristics of the driver's face.
The vehicle-mounted radar method based on deep learning is used to denoising and separation signal noise and separation through average decomposition and multi-sequence variational mode decomposition algorithms, and the micro Doppler feature map is obtained in combination with short-time Fourier transform, and fatigue detection is performed using a custom dual-stream fusion convolutional neural network.
High-precision detection of the driver's tiny and fine-grained fatigue characteristics is achieved, avoiding the problems of light dependence and privacy violations, and providing a comfortable and effective fatigue detection method.
Smart Images

Figure CN120071308A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fatigue detection, and particularly relates to a method for detecting driver facial fatigue based on deep learning by an in-vehicle radar. Background Art
[0002] Driver drowsiness or inattention is an important cause of motor vehicle accidents. Therefore, road safety has become a serious problem for the current generation of intelligent transportation systems. Existing driving safety research shows that driving fatigue is a transitional process. There are many early signs in the driver's behavior, called driver fatigue characteristics, and the driver exhibits typical fatigue characteristics before an accident occurs. Therefore, an effective method for detecting fatigue driving is needed, which can identify fatigue characteristics and remind the driver in advance to reduce the occurrence of traffic accidents. Currently, determining a person's fatigue level by detecting fatigue-related symptoms has become a hot research direction for many research institutions.
[0003] Traditional solutions mainly use cameras and wearable sensors for fatigue detection. Cameras are usually used to monitor the blinking and mouth states of the face and head movements. However, the quality of video images is easily affected by lighting conditions, such as darkness and strong light emitted by car lights, and it has the problem of privacy infringement. Fatigue driving detection methods based on sensors have been widely explored, such as electrocardiogram (ECG), electroencephalogram (EEG), electrooculogram (EOG), brain wave signal measurement, photoplethysmogram (PPG) detection, and these devices are all used to detect the degree of driving fatigue. However, wearable devices are not suitable for monitoring patients in specific populations, such as burn or skin disease patients; in addition, during the above measurement process, the driver is usually restricted by additional sensors and electrodes, which is uncomfortable and may have a negative impact on the driver's behavior.
[0004] Compared with the above methods, the latest progress in radio frequency (RF) sensing can achieve fatigue detection that is independent of light and non-invasive. The RF signal transmitter emits a signal to the driver, and then, the signal reflected by the driver's body is processed to detect the driver's fatigue characteristics. However, existing methods using long-wavelength RF signals, such as WiFi and RFID, can only capture coarse-grained fatigue characteristics (such as nodding and yawing), and cannot obtain minute and fine-grained characteristics (such as blinking or yawning) that are more strongly associated with fatigue. In addition, some studies use RF devices with a single transceiver design (such as ultra-wideband radar) to detect the characteristics of a single body area, that is, the fatigue characteristics around the face or chest, but it should be noted that due to its own high bandwidth, the UWB radar system will also have the problem of low signal-to-noise ratio. The fatigue detection method based on millimeter-wave radar has attracted great interest worldwide due to its high precision, high robustness, and potential for privacy protection.
[0005] The advantage of millimeter-wave radar in measuring human facial fatigue signals lies in its non-contact nature, without the need for direct contact with the human body, avoiding the discomfort or interference that traditional methods may bring. In addition, millimeter waves have high resolution, enabling high-precision imaging and detection of minute human movements. It has strong penetration power and can penetrate obstacles such as masks and glasses to achieve detection of human facial signals. Therefore, millimeter-wave radar, as an advanced wireless sensing technology, has great potential and advantages in the field of fatigue signal detection. Additionally, radar perception technology based on deep learning has been widely applied in fields such as assisted driving, intelligent security, and human-computer interaction. Therefore, the present invention proposes an effective method for detecting driver facial fatigue based on deep learning. Summary of the Invention
[0006] The object of the present invention is to provide an effective method for detecting driver facial fatigue. It realizes noise reduction and separation of facial signals through Mean Elimination (ME) and Multivariate Variational Mode Decomposition (MVMD). Then, useful modes are selected for target signal reconstruction based on energy ratio. Further, the Short Time Fourier Transform (STFT) is used to obtain the micro-Doppler feature map of the signal to form the original dataset. Then, the data augmentation algorithm (With MixupAu, WMA) is used for data expansion, and a custom Two-stream Feature fusion Convolutional Network (TFC-Net) is used for effective fatigue detection. The method for detecting driver facial fatigue based on deep learning according to the present invention includes the following steps:
[0007] Step 1: Oblique the frequency modulated continuous wave (FMCW) radar directly above the front of the human face, with the antenna facing the human face to collect human target information to obtain the radar intermediate frequency signal. Perform a fast Fourier transform on a single-frame intermediate frequency signal to obtain the distance vector matrix R M×1 , and then perform multi-frame accumulation over time. Concatenate the N-frame distance vectors by column to construct the distance-time matrix R M×N = [R 1 , R 2 ,... R N M×N , thereby obtaining the Range-Time-Map (RTM).
[0008] Step 2: Use the ME algorithm to remove static clutter, and then determine the distance range where the target to be detected is located by the maximum average power in different distance intervals. Extract the intermediate-frequency signal x(t) in the distance range where the target is located, and then use the MVMD decomposition algorithm to denoise x(t) to suppress dynamic clutter.
[0009] Step 3: Since the energies of both blink and yawn signals are within the range of 2 Hz, calculate the percentage of the energy of the components in this frequency band in each IMF signal in the total energy of all components, and then the correlation degree between each IMF signal and the target signal can be obtained. Select the component with the highest correlation degree as the potential target signal y(t) according to the energy ratio, and perform STFT on y(t) to obtain the micro-Doppler characteristics of the target signal.
[0010] Step 4: Use the WMA algorithm to perform data augmentation on the real sample dataset. Specifically, propose the MP algorithm and the AU algorithm for data enhancement respectively, and then obtain the mixed dataset using the weight factor.
[0011] Step 5: Input the mixed dataset into the constructed TFC-Net network for fatigue classification. Description of the Drawings
[0012] Figure 1 is the flow chart of the present invention;
[0013] Figure 2 is the time-frequency diagram of the blink signal decomposed by MVMD;
[0014] Figure 3 is the frequency-domain diagram of the blink signal after STFT;
[0015] Figure 4 is the TFC-Net network structure Specific Embodiment
[0016] The present invention will be further described below with reference to the drawings:
[0017] As Figure 1 shown, the technical solution adopted by the present invention is: a vehicle-mounted radar driver fatigue detection method based on deep learning, which mainly includes the following steps:
[0018] Step 1: Place the FMCW radar obliquely above the front of the human face at a certain angle. After collecting the radar intermediate-frequency signal of the human face data, perform a fast Fourier transform on the single-frame intermediate-frequency signal to obtain the distance vector matrix R M×1 , and then perform multi-frame accumulation through time. Concatenate the N-frame distance vectors by column to construct the distance-time matrix R M×N = [R 1 , R 2 ,... R N M×N , so as to obtain a Range-Time-Map (RTM).
[0019] Step 2: Use the average cancellation algorithm to remove stationary clutter, and its formula is:
[0020]
[0021] Furthermore, the distance range where the target to be detected is located is determined by the maximum average power in different distance intervals, and the intermediate frequency signal x(t) in the distance range where the target is located is extracted. Due to the complex environment inside the vehicle, the standard Variational Mode Decomposition (VMD) algorithm is not sufficient to accurately extract the target signal. Therefore, the traditional VMD algorithm is improved by merging multiple sequences. As Figure 2 shown, the main objective of the MVMD algorithm is to extract predefined K IMF components from the input data x(t) containing C data channels: where u k (t) = [u 1 (t), u 2 (t), … u C (t)]. Its objective is to extract a set of multivariate modulated oscillations from the input data. It is required to satisfy: i) the sum of the bandwidths of the extracted modes is minimized, and ii) the sum of the extracted modes exactly recovers the original signal u k (t). To achieve this goal, first use the Hilbert transform to calculate the analytical signal associated with each sub-signal in u k (t), and obtain u k (t) with a complex exponential center frequency ω k (t). Then, use the gradient of the L 2 norm to estimate the bandwidth of u k (t). This problem is decomposed into two sub-problems: i) mode update, and ii) center frequency correction.
[0022] For problem i), its equivalent optimization problem is:
[0023]
[0024] This sub-problem is formally similar to the mode update problem of the original VMD. Solve this problem in the Fourier domain, so as to obtain a concise update relationship in the frequency domain.
[0025]
[0026] For problem ii), its problem is simplified to:
[0027]
[0028] It can be more conveniently optimized in the frequency domain. Using the Plancherel theorem of the inner product of a function in the time domain and the frequency domain, the equivalent problem in the Fourier domain can be obtained as follows:
[0029]
[0030] By setting the first derivative of the above quadratic function to zero to minimize its sum, and then through algebraic simplification, the following relationship is obtained:
[0031]
[0032] Step 3: Since the energies of the blink and yawn signals are both within the range of 2 Hz, the percentage of the energy of the frequency components in this segment in the total energy of each IMF signal is calculated, and the correlation degree between each IMF signal and the target signal can be obtained:
[0033]
[0034] α j represents the correlation degree between the j-th IMF signal and the target signal. The component with the highest correlation degree is selected as the potential target signal y(t) through the energy ratio, and further STFT is performed on y(t) to obtain the Figure 3 micro-Doppler feature map as shown.
[0035] Step 4: The WMA algorithm is used to perform data augmentation on the real sample dataset. Specifically, the MP algorithm and the AU algorithm are respectively proposed for data enhancement, and then the weight factor is used to obtain the mixed dataset.
[0036] 4.1 The MP algorithm uses the linear interpolation method. Let x and y represent the data (parameter map) and the label (action type) respectively. (x i , y i ) and (x j , y j ) are two vectors randomly selected from the fatigue feature map. Then, virtual feature vectors are sampled from the mixed neighborhood distribution.
[0037]
[0038] 4.2 The AU algorithm can increase the diversity of the action feature map without changing the action label. It mainly includes random cropping, translation, scaling, contrast, and rotation transformations.
[0039] Step 5: Build as shown in Figure 4The shown Two-stream Featurefusion Convolutional Network (TFC-Net) inputs the micro-Doppler feature map obtained in Step 4 into the neural network for training, adjusts the network parameters, and performs recognition after the neural network converges.
[0040] 5.1 Set the number of convolutional kernels in the convolutional module to 32. Through a convolutional transformation with a convolutional kernel of 5×5 and a stride of 1 for 1 layer, while keeping the size of the input data unchanged, perform batch normalization (Batch Normalization, BN), max pooling with a stride of 2, and rectified linear unit (ReLU) activation operations on the result to obtain a shallow feature map with a size of 112×112.
[0041] 5.2 Input the shallow feature map into the S stream of TFC-Net to extract shallow features. The specific steps are as follows: (1) Set the number of convolutional kernels in the S stream to 48. Perform a convolutional operation with a convolutional kernel of 1×1 and a stride of 2 on the result of Step 5.1 to achieve a 2-fold downsampling effect and generate a feature map with a size of 56×56. (2) Input the feature map into the channel attention mechanism module. Use average pooling and max pooling to learn the background information and unique texture features of the target feature map, and aggregate the spatial information of the learned feature map through a shared multilayer perceptron (MLP). The specific calculation method is:
[0042]
[0043] where M c (F) represents the feature map generated after extracting features from the input feature map by the above method, δ(·) represents the ReLU function, represents the input feature vector, AvgPool and MaxPool respectively represent average pooling and max pooling, and respectively represent the feature maps obtained by using average pooling and max pooling on the input feature map to compress the spatial dimension, W 0 and W 1 respectively represent two-layer MLP hidden layers with 32 and 48 nodes; (3) Multiply M c (F) by the 56×56 feature map obtained in step (1) for adaptive learning of features, generate a feature map with a size of 56×56, and then perform global max pooling on it to extract more refined features at the key positions of the target feature map.
[0044] 5.3 Input the deep feature map into the D stream of TFC-Net to extract the deep features of the target feature map and obtain a feature map with a size of 28×28. The specific steps are as follows: (1) Set the number of convolution kernels of convolution module 2 to 32, the receptive field size of the max-pooling layer to 3×3, and the stride to 2. Perform a convolution operation with a convolution kernel of 3×3 and a stride of 1 on the result of step 5.2, and then perform BN, max-pooling, and ReLU activation operations to generate a feature map with a size of 56×56; (2) Set the number of convolution kernels of convolution module 3 to 48, the receptive field size of the max-pooling layer to 3×3, and the stride to 2. Perform a convolution operation with a convolution kernel of 3×3 and a stride of 1 on the result of step (1), and then perform BN, max-pooling, and ReLU activation operations to extract the deep features of the target feature map and obtain a feature map with a size of 28×28.
[0045] 5.4 Fuse the features extracted from the shallow layer and the deep layer, and introduce a Dropout layer with a probability of 0.2 to randomly disconnect the network connections to prevent overfitting during network training. Then, use two fully connected layers with 120 and 80 nodes respectively to further aggregate and mine the feature information of the target, and input the generated feature vector into the orthogonal Softmax function for calculation to obtain the probability result matrix X = [x 0 x 1 x 2 T . The subscript with the largest scalar value represents the result of the target feature map being classified as a certain class, thereby achieving accurate classification of fatigue actions.
Claims
1. A method for detecting driver facial fatigue based on vehicle-mounted radar, characterized in that The following steps are involved: Step 1: Place the frequency modulated continuous wave (FMCW) radar at a certain angle in front of the human face, collect the human face data to obtain the radar intermediate frequency signal, and then perform fast Fourier transform on the single frame intermediate frequency signal to obtain the distance vector matrix R M×1 , and then accumulate multiple frames through time, concatenate the N frame distance vectors by column, and construct the distance-time matrix R M×N =[R1,R2,...R N ] M×N , thus obtaining the Range-Time-Map (RTM). Step 2: Use the average cancellation algorithm to remove stationary clutter. The formula is: Then, the maximum average power in different distance intervals is used to determine the distance range of the target to be detected, and the intermediate frequency signal x(t) of the target distance range is extracted. Due to the complex environment inside the car, the standard variational mode decomposition algorithm (VMD) is not enough to accurately extract the target signal. Therefore, the traditional VMD algorithm is improved by merging multiple sequences. The main goal of multivariate variational mode decomposition (MVMD) is to extract predefined K IMF components from the input data x(t) containing C data channels: where u k (t)=[u1(t),u2(t),…u C (t)]. Its goal is to extract multivariate modulated oscillations in the input data The following needs to be satisfied: i) the sum of the bandwidths of the extracted modes is minimal, ii) the sum of the extracted modes accurately restores the original signal u k (t). To achieve this, we first use the Hilbert transform to calculate u k The analytical signal associated with each sub-signal in (t) is obtained with a complex exponential center frequency ω k t of u k (t). Then, the gradient of the L2 norm is used to estimate u k The problem is decomposed into two sub-problems: i) mode update and ii) center frequency correction. For problem i), the equivalent optimization problem is: This subproblem is similar in form to the mode update problem of the original VMD. Solving this problem in the Fourier domain leads to a concise update relation in the frequency domain. For problem ii), the problem is simplified to: It is more convenient to optimize in the frequency domain. Using Plancherel's theorem for the inner product of functions in the time domain and frequency domain, we can get the equivalent problem in the Fourier domain as follows: By setting the first-order derivative of the above quadratic function to zero to minimize its sum, and then simplifying it algebraically, we get the following relationship: Step 3: Since the energy of blinking and yawning signals is within the 2Hz range, the percentage of the frequency component in each IMF signal to the total component energy can be calculated to obtain the correlation between each IMF signal and the target signal: α j Indicates the correlation between the jth IMF signal and the target signal. The component with the highest correlation is selected as the potential target signal y(t) by energy proportion, and y(t) is subjected to short-time Fourier transform (STFT) to obtain the micro-Doppler characteristic map of the target signal. Step 4: Use the WMA algorithm to expand the real sample data set. Specifically, the MP algorithm and AU algorithm are proposed for data enhancement, and then the weight factor is used to obtain the mixed data set. The 4.1MP algorithm uses linear interpolation method, using x and y to represent data (parameter map) and labels (action type), respectively. i ,y i ) and (x j ,y j ) are two vectors randomly selected from the fatigue feature map, and then, a virtual feature vector is generated by sampling from the mixed neighborhood distribution. The 4.2AU algorithm can increase the diversity of action feature maps without changing the action labels. It mainly includes random cropping, translation, scaling, contrast and rotation transformations. Step 5: Build a two-stream feature fusion convolutional neural network (TFC-Net), input the micro-Doppler feature map obtained in step 4 into the neural network for training, adjust the network parameters, and perform recognition after the neural network converges. 5.1 Set the number of convolution kernels of the convolution module to 32, and use a convolution transformation with a convolution kernel of 5×5 and a stride of 1 to keep the size of the input data unchanged. Perform batch normalization (BN), maximum pooling with a stride of 2, and rectified linear unit (ReLU) activation operations on the result to obtain a shallow feature map of size 112×112. 5.2 Input the shallow feature map into the S stream of TFC-Net to extract the shallow features. The specific steps are as follows: (1) Set the number of convolution kernels in the S stream to 48, and perform a convolution operation with a convolution kernel of 1×1 and a step size of 2 on the result of step 5.1 to achieve a 2x downsampling effect and generate a feature map of size 56×56; (2) Input the feature map into the channel attention mechanism module, use average pooling and maximum pooling to learn the background information and unique texture features of the target feature map, and aggregate the learned feature map spatial information through a shared multilayer perceptron (MLP). The specific calculation method is: in, M c (F) represents the feature map generated after the input feature map is extracted by the above method, δ(·) represents the ReLU function, represents the input feature vector, AvgPool and MaxPool represent average pooling and maximum pooling respectively, and denote the feature maps obtained by compressing the spatial dimension of the input feature map using average pooling and maximum pooling, respectively. W0 and W1 denote two MLP hidden layers with 32 and 48 nodes, respectively. (3) c (F) Multiply the 56×56 feature map obtained in step (1) to perform adaptive feature learning and generate a 56×56 feature map. Then perform global maximum pooling on it to extract more detailed features at key positions of the target feature map. 5.3 Input the deep feature map into the D stream of TFC-Net, extract the deep features of the target feature map, and obtain a feature map of size 28×28. The specific steps are as follows: (1) Set the number of convolution kernels of convolution module 2 to 32, the receptive field size of the maximum pooling layer to 3×3, and the step size to 2. Perform a convolution operation with a convolution kernel of 3×3 and a step size of 1 on the result of step 5.2, and perform BN, maximum pooling and ReLU activation operations to generate a feature map of size 56×56; (2) Set the number of convolution kernels of convolution module 3 to 48, the receptive field size of the maximum pooling layer to 3×3, and the step size to 2. Perform a convolution operation with a convolution kernel of 3×3 and a step size of 1 on the result of step (1), and perform BN, maximum pooling and ReLU activation operations to extract the deep features of the target feature map, and obtain a feature map of size 28×28. 5.4 The features extracted from the shallow layer and the deep layer are fused, and a Dropout layer with a probability of 0.2 is introduced to randomly disconnect the network connection to prevent overfitting of the network training. Then, two fully connected layers with 120 and 80 nodes are used to further aggregate the feature information of the mining target, and the generated feature vector is input into the orthogonal Softmax function for calculation to obtain the probability result matrix X = [x0 x1 x2] that the current input sample belongs to each fatigue feature category T ,The scalar value of the largest subscript indicates that the target feature map is identified as a certain category, thus achieving accurate classification of fatigue actions.