Non-contact drunk driving detection method based on deep learning
Through multispectral cameras and deep learning technology, the facial micro-expression and physiological rhythm signals of drunk drivers are extracted, which solves the accuracy and environmental adaptability of existing drunk driving detection, and achieves efficient and safe contactless drunk driving detection.
Patent Information
- Application Number
- CN202510348451.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing drunk driving detection methods have problems such as low detection accuracy, poor environmental adaptability, easy to be disturbed and complex detection process. In particular, traditional contact detection has the risk of cross-infection, and non-contact detection is difficult to effectively extract physiological characteristics and insufficient environmental adaptability.
Multispectral cameras are used to collect video stream data, and facial micro-expression and physiological rhythm signals are extracted through spatiotemporal convolutional neural network and sparse coding algorithm. Combined with a cascaded deep learning classifier and dynamic threshold segmentation algorithm, multimodal feature fusion and decision optimization are achieved to generate alcohol concentration probability distribution.
It realizes high-precision, low-interference and safe drunk driving testing, reducing the complexity of the detection process and the risk of cross-infection, and improving the accuracy of the detection and environmental adaptability.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drunk driving detection, and specifically to a non-contact drunk driving detection method based on deep learning. Background Art
[0002] With the development of society, drunk driving behaviors pose a serious threat to public traffic safety. Traditional drunk driving detection methods, such as breath alcohol testers, blood alcohol tests, etc., although they can accurately detect the driver's alcohol intake to a certain extent, have many limitations.
[0003] A breath alcohol tester requires the driver to actively cooperate with the breath operation. Some drivers may have resistance or deliberately not cooperate, which affects the smooth progress of the detection. At the same time, this method requires the tester to be in close contact with the driver. In some special scenarios, such as during the epidemic of infectious diseases, there is a risk of cross-infection. Although the blood alcohol test has accurate results, it is an invasive detection method that requires professional medical staff to operate. The detection process is complex and time-consuming, and it cannot meet the needs of on-site rapid detection, making it difficult to be applied to the front line of traffic law enforcement.
[0004] At present, with the rapid development of computer vision and deep learning technologies, non-contact detection technologies have gradually become a research hotspot. In the field of drunk driving detection, the idea of using a camera for non-contact detection has emerged. However, existing non-contact drunk driving detection research still faces many challenges. From the perspective of data collection, the image information collected by ordinary cameras is limited and it is difficult to comprehensively capture the physiological characteristic changes related to alcohol intake. While multi-spectral cameras can collect richer information, there are problems such as difficulty in aligning images between different bands and noise interference, which affect subsequent feature extraction and analysis.
[0005] In terms of feature extraction, traditional machine learning methods are difficult to effectively extract the complex dynamic changes of facial micro-expressions and the weak alcohol metabolism-related information hidden in physiological signals. Facial micro-expression changes are affected by many factors, such as individual differences, environmental light changes, etc. How to accurately and stably extract these features for drunk driving detection is a major problem. For physiological rhythm signals, their components are complex and contain various interference signals unrelated to alcohol metabolism. Existing signal processing methods are difficult to accurately separate the signals related to alcohol metabolism and perform effective analysis.
[0006] In terms of model construction and decision-making, existing classification models are difficult to fully integrate the features of multi-modal data, resulting in low detection accuracy. Moreover, under different environmental conditions, such as different light intensities, different postures of the driver, etc., the generalization ability of the model is insufficient and misjudgments are likely to occur. When determining the decision boundary for drunk driving detection, there is a lack of scientific and effective optimization methods, which affects the reliability of the detection results. Summary of the Invention
[0007] The object of the present invention is to provide a non-contact drunk driving detection method based on deep learning to solve the problems raised in the above-mentioned background technology.
[0008] To achieve the above object, the present invention provides the following technical solution: A non-contact drunk driving detection method based on deep learning, the method comprising:
[0009] Collecting continuous video stream data of a target object through a multispectral camera, the multispectral camera including visible light band, near-infrared band and thermal infrared band sensors; performing inter-frame alignment and noise suppression processing on the continuous video stream data to extract a facial region image sequence of the target object;
[0010] Based on a pre-trained spatio-temporal convolutional neural network, extracting optical flow features from the facial region image sequence to generate a dynamic change vector of facial micro-expressions;
[0011] Fusing the optical flow features with the thermal infrared band data, and separating the physiological rhythm signal of the target object through a time series decomposition algorithm;
[0012] Based on an improved sparse coding algorithm, enhancing the features of the physiological rhythm signal to generate a heart rate variability index and an estimated value of blood oxygen saturation;
[0013] Constructing a multi-modal fusion feature vector, the multi-modal fusion feature vector including the optical flow features, the physiological rhythm signal, the heart rate variability index and the estimated value of blood oxygen saturation; inputting the multi-modal fusion feature vector into a cascaded deep learning classifier, the cascaded deep learning classifier consisting of a temporal convolutional network and a self-attention mechanism module, and outputting the probability distribution of the alcohol concentration of the target object;
[0014] Based on a dynamic threshold segmentation algorithm, optimizing the decision boundary of the alcohol concentration probability distribution to generate a drunk driving detection result.
[0015] Preferably, the inter-frame alignment and noise suppression processing of the continuous video stream data includes:
[0016] Calculating the translation and rotation parameters of adjacent video frames by using a phase correlation algorithm, and realizing pixel-level alignment between multi-spectral bands through bilinear interpolation; constructing a noise suppression model based on wavelet transform, performing adaptive threshold filtering on high-frequency sub-band coefficients, and retaining facial texture details; performing edge smoothing processing on the noise-suppressed image by using morphological closing operation.
[0017] Preferably, the time series decomposition algorithm includes:
[0018] Build a variational mode decomposition model, and decompose the physiological rhythm signal into multiple intrinsic mode functions by iteratively solving a constrained optimization problem; screen the mode components related to alcohol metabolism based on the joint criterion of kurtosis coefficient and energy entropy; perform Hilbert transform on the screened mode components to extract instantaneous frequency features and phase synchronization indexes.
[0019] Preferably, the improved sparse coding algorithm includes:
[0020] Build an asymmetric dictionary learning model, and optimize the joint objective function of dictionary atoms and sparse coefficients by the alternating direction multiplier method; introduce a dynamic weighting factor into the sparse constraint term, and the dynamic weighting factor is adaptively adjusted according to the importance of feature dimensions; adopt a layer-by-layer incremental coding strategy, and gradually optimize the sparse representation accuracy through a residual feedback mechanism.
[0021] Preferably, the construction of the cascaded deep learning classifier includes:
[0022] Design a temporal convolutional network module, use dilated convolutional kernels to capture multi-scale temporal dependencies, and suppress the weights of irrelevant time steps through a gating mechanism; build a self-attention mechanism module, calculate the correlation matrix between feature dimensions through multi-head attention layers, and adopt relative position encoding to enhance the temporal modeling ability; introduce a feature recalibration module at the cascaded connection, and dynamically adjust the channel weights of the feature map through a channel attention mechanism.
[0023] Preferably, the dynamic threshold segmentation algorithm includes:
[0024] Build a probability density estimation function based on a Gaussian mixture model to perform multi-modal modeling on the alcohol concentration probability distribution; use an improved mean shift algorithm to search for the local extreme points of the probability density function to generate a set of candidate decision boundaries; based on the KL divergence to measure the classification uncertainty of each candidate boundary, and select the segmentation threshold corresponding to the minimum uncertainty.
[0025] Preferably, it further includes building a data augmentation model based on a generative adversarial network, and synthesizing facial images under different lighting and pose conditions through a conditional generator; introducing a spectral consistency loss function into the discriminator to constrain the physical property consistency of the generated images among multi-spectral bands; adopting a curriculum learning strategy to gradually increase the complexity of the generated data and improve the generalization performance of the classifier.
[0026] Preferably, it further includes building a multi-task learning framework to jointly optimize the alcohol concentration prediction task and the physiological index regression task; introducing an orthogonal regularization term into the loss function to constrain the feature representation spaces of different tasks to be independent of each other; adopting a gradient reversal layer to implement adversarial training to eliminate the influence of environmental interference factors on feature extraction.
[0027] Preferably, it further includes designing an online incremental learning mechanism, collecting real-time detection data through a sliding window and calculating the feature drift amount; when the feature drift amount exceeds a preset threshold, triggering a local network parameter update process, and using the elastic weight consolidation algorithm to retain the prior knowledge of important parameters; introducing a momentum buffer mechanism during the model update process to smooth the parameter update direction.
[0028] Preferably, the present invention further includes a non-contact drunk driving detection system based on deep learning, and the system includes:
[0029] A data acquisition module for obtaining continuous video stream data of a target object through a multispectral camera;
[0030] A preprocessing module for performing inter-frame alignment, noise suppression, and facial region extraction;
[0031] A feature extraction module including a spatio-temporal convolutional neural network, a time series decomposition algorithm, and an improved sparse coding algorithm;
[0032] A classification and decision-making module composed of a cascaded deep learning classifier and a dynamic threshold segmentation algorithm;
[0033] A model optimization module for implementing functions such as generative adversarial network enhancement, multi-task learning, and online incremental learning.
[0034] Compared with the prior art, the beneficial effects of the present invention are:
[0035] Collecting continuous video stream data of a target object through a multispectral camera to achieve non-contact detection. Compared with breath alcohol testers and blood alcohol tests, there is no need for the driver to actively cooperate with specific actions, avoiding the situation where the driver's resistance affects the detection process. At the same time, it reduces the close contact between the tester and the driver, reducing the risk of cross-infection, which is particularly significant during special periods, improving the convenience and safety of detection.
[0036] The inter-frame alignment and noise suppression processing of continuous video stream data, using technologies such as phase correlation algorithm and wavelet transform, realizes pixel-level alignment between multispectral bands, effectively suppresses noise and retains facial texture details, laying a foundation for subsequent accurate feature extraction. Using a pre-trained spatio-temporal convolutional neural network to extract optical flow features, combined with a time series decomposition algorithm and an improved sparse coding algorithm, can accurately separate physiological rhythm signals, generate heart rate variability indicators and blood oxygen saturation estimation values, fully excavating weak information related to alcohol metabolism, and improving the accuracy and effectiveness of feature extraction.
[0037] The cascaded deep learning classifier consists of a temporal convolutional network and a self-attention mechanism module. Through designs such as dilated convolutional kernels, gating mechanisms, and multi-head attention layers, it can effectively capture multi-scale temporal dependencies, enhance the correlation analysis between feature dimensions and the temporal modeling ability, and accurately identify in complex environments. The dynamic threshold segmentation algorithm optimizes the decision boundary based on the Gaussian mixture model and the improved mean shift algorithm, reduces classification uncertainty, and improves the reliability of drunk driving detection results.
[0038] Based on the generative adversarial network, a data augmentation model is constructed to synthesize facial images under different lighting and pose conditions, improve the generalization performance of the classifier, and reduce misjudgments caused by environmental factors. The multi-task learning framework jointly optimizes the alcohol concentration prediction and physiological index regression tasks. The orthogonal regularization term and the gradient reversal layer eliminate environmental interference and improve the robustness of the model in complex environments. The online incremental learning mechanism can monitor data feature drift in real time, update the model in a timely manner, smooth the update direction while retaining the prior knowledge of important parameters, enabling the model to adapt to the continuously changing actual detection scenarios and maintain good performance in the long term. Brief Description of the Drawings
[0039] Figure 1 It is the working principle diagram of the non-contact drunk driving detection method described in the present invention;
[0040] Figure 2 It is the implementation step diagram of the time series decomposition algorithm;
[0041] Figure 3 It is the construction flow chart of the cascaded deep learning classifier. Detailed Embodiments
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0043] Please refer to Figures 1-3 , the present invention provides a non-contact drunk driving detection method based on deep learning, aiming to achieve efficient countermeasures against drones by constructing a dynamic defense model and optimizing countermeasure strategies, while reducing the consumption of countermeasure resources. Its overall implementation plan is as follows:
[0044] Use a multispectral camera to collect continuous video stream data of the target object. This multispectral camera integrates visible light band, near-infrared band, and thermal infrared band sensors, and can obtain facial information of the target object from multiple spectral dimensions, providing a rich data basis for subsequent analysis. Data in different bands have their own advantages in reflecting facial features and physiological information. The visible light band can clearly present the appearance features of the face, the near-infrared band helps to capture information such as blood vessels under the facial skin, and the thermal infrared band can reflect the temperature distribution of the face, which is closely related to the physiological state of the human body.
[0045] Perform inter-frame alignment and noise suppression processing on the collected continuous video stream data, and then extract the facial region image sequence of the target object. Inter-frame alignment is to eliminate the deviation caused by factors such as shooting device jitter and small movements of the target object between video frames, ensuring the accuracy of subsequent analysis; noise suppression processing is to remove various noises mixed in the video data to avoid noise interfering with subsequent feature extraction and analysis. Through this series of processes, a clear and accurate facial region image sequence is obtained, preparing for the subsequent feature extraction work.
[0046] Use a pre-trained spatio-temporal convolutional neural network to extract optical flow features from the extracted facial region image sequence, and generate a dynamic change vector of facial micro-expressions. The spatio-temporal convolutional neural network can effectively capture the dynamic features of the image sequence in the time and space dimensions. The changes in facial micro-expressions are often related to the physiological and psychological states of the human body. By extracting optical flow features, the dynamic changes of facial micro-expressions can be transformed into a quantifiable vector representation, providing an important feature basis for subsequent analysis.
[0047] Fuse the extracted optical flow features with the thermal infrared band data, and then use the time series decomposition algorithm to separate the physiological rhythm signal of the target object. The optical flow features reflect the dynamic changes of facial micro-expressions, and the thermal infrared band data contains the physiological heat information of the human body. The fusion of the two can more comprehensively reflect the physiological state of the human body. The time series decomposition algorithm can separate the signals related to the human physiological rhythm from the fused data. These signals contain rich physiological information and are of great significance for judging whether the target object is driving under the influence of alcohol.
[0048] Based on an improved sparse coding algorithm, perform feature enhancement on the separated physiological rhythm signal to generate a heart rate variability index and a blood oxygen saturation estimation value. The improved sparse coding algorithm can highlight the key features in the physiological rhythm signal and improve the expression ability of the features. The heart rate variability index reflects the rhythm changes of the heart beating, and the blood oxygen saturation estimation value reflects the oxygen content in the human blood. Both of these two indicators are closely related to the alcohol metabolism and physiological state of the human body and are important bases for judging whether a person is driving under the influence of alcohol.
[0049] Construct a multi-modal fusion feature vector, which includes optical flow features, circadian rhythm signals, heart rate variability metrics, and blood oxygen saturation estimates. Multi-modal fusion can make full use of the complementary information of different types of features to improve the accuracy of detection. Input the multi-modal fusion feature vector into a cascaded deep learning classifier composed of a temporal convolutional network and a self-attention mechanism module to output the probability distribution of the alcohol concentration of the target object. The temporal convolutional network and the self-attention mechanism module can effectively learn the complex patterns and temporal relationships in the multi-modal fusion feature vector, so as to accurately predict the probability distribution of the alcohol concentration.
[0050] Optimize the decision boundary for the alcohol concentration probability distribution based on the dynamic threshold segmentation algorithm to generate drunk driving detection results. The dynamic threshold segmentation algorithm can adaptively determine the optimal decision boundary according to the characteristics and distribution of the data, improving the accuracy and reliability of drunk driving detection. Through this algorithm, the alcohol concentration probability distribution is converted into clear drunk driving detection results to realize the judgment of whether the target object is driving under the influence of alcohol.
[0051] The following further illustrates the implementation of the present invention in conjunction with Examples 1 to 5.
[0052] Example 1:
[0053] This example mainly describes the specific implementation method in the data collection and preprocessing stage. This stage is the basis of the entire drunk driving detection system and directly affects the accuracy of subsequent feature extraction and detection results.
[0054] In the data collection link, a camera device with high-resolution and multi-spectral collection capabilities is selected. For example, select a multi-spectral camera with a resolution of [specific resolution value], whose visible light band wavelength range covers [visible light wavelength range], and can clearly capture details such as facial texture and skin color; the near-infrared band wavelength is in [near-infrared wavelength range], which can penetrate the skin surface layer to obtain information such as subcutaneous blood vessels; the thermal infrared band is set in [thermal infrared wavelength range] to monitor the change of facial temperature distribution. The camera is installed in a suitable position to ensure that it can stably and clearly capture the face of the target object, avoiding problems such as occlusion and reflection. In actual application scenarios, such as traffic law enforcement checkpoints, the camera can be installed directly in front of the vehicle inspection area, and the height and angle are precisely adjusted to adapt to the detected personnel with different heights and sitting postures.
[0055] After collecting continuous video stream data, it enters the data preprocessing stage. First, inter-frame alignment processing is performed, and the phase correlation algorithm is used to calculate the translation and rotation parameters of adjacent video frames. The phase correlation algorithm is based on the properties of the Fourier transform. By calculating the correlation between the phase spectra of two adjacent frames of images, it can quickly and accurately obtain the translation and rotation information between the images. Assume that two adjacent frames of images are I1(x,y) and I2(x,y) respectively, and their Fourier transforms are F1(u,v) and F2(u,v), then the phase correlation function is:
[0056]
[0057] where, is the conjugate complex number of F2(u,v). By performing the inverse Fourier transform on the phase correlation function, the translation parameters (Δx, Δy) and the rotation angle θ can be obtained. After obtaining the translation and rotation parameters, the bilinear interpolation algorithm is used to achieve pixel-level alignment between multi-spectral bands. Bilinear interpolation is a commonly used image interpolation method. For the target pixel position (x,y), its pixel value is obtained by the weighted average of the surrounding four adjacent pixels, and the weights are calculated according to the distance between the target pixel and the adjacent pixels.
[0058] After completing the inter-frame alignment, a noise suppression model based on wavelet transform is constructed. Wavelet transform can decompose an image into sub-bands of different frequencies, and the high-frequency sub-bands mainly contain the noise information in the image. Adaptive threshold filtering is performed on the high-frequency sub-band coefficients, and whether to retain a coefficient is determined according to the comparison result of the size of each coefficient and the adaptive threshold. The adaptive threshold can be calculated according to the statistical characteristics of the image, such as:
[0059]
[0060] where, σ is the standard deviation of the image noise, which can be estimated by statistical analysis of the image; N is the total number of pixels in the image. In this way, while removing the noise, the facial texture details are retained. Finally, morphological closing operation is used to perform edge smoothing on the noise-suppressed image. Morphological closing operation can fill small holes in the image and connect discontinuous edges through the operations of dilation first and then erosion, making the image edges smoother and avoiding interference to the subsequent facial region extraction caused by noise and edge discontinuity.
[0061] After the above data collection and preprocessing steps, a high-quality sequence of facial region images can be obtained, providing reliable data support for subsequent feature extraction and drunk driving detection.
[0062] Example 2:
[0063] This embodiment elaborates in detail the specific implementation process of the time series decomposition algorithm, which is used to separate the physiological rhythm signal of the target object from the fusion data and extract relevant features, and is of great significance for accurately judging drunk driving.
[0064] In the time series decomposition algorithm, a variational mode decomposition model (VMD) is first constructed. Variational mode decomposition is an adaptive signal decomposition method that decomposes a signal into multiple intrinsic mode functions (IMFs).
[0065] By introducing a quadratic penalty factor and a Lagrange multiplier, the constrained optimization problem is transformed into an unconstrained optimization problem, and the alternating direction multiplier method (ADMM) is used for iterative solution to obtain multiple intrinsic mode functions.
[0066] After obtaining the intrinsic mode functions, the modal components related to alcohol metabolism are screened based on the joint criterion of kurtosis coefficient and energy entropy. The kurtosis coefficient is used to measure the peak degree of a signal, and its calculation formula is:
[0067]
[0068] where E[(u―μ) 4 is the fourth-order central moment of the signal u, μ is the mean of the signal, and σ is the standard deviation of the signal. The energy entropy is used to measure the uniformity of the energy distribution of a signal, and its calculation formula is:
[0069]
[0070] where is the total energy of the signal u, and u i is the discrete sampling value of the signal. For each intrinsic mode function, its kurtosis coefficient and energy entropy are calculated, and according to the preset threshold and empirical judgment criterion, the modal components related to alcohol metabolism are screened. Generally speaking, the modal components related to alcohol metabolism will show specific distribution characteristics in terms of kurtosis coefficient and energy entropy. For example, a larger kurtosis coefficient and a smaller energy entropy indicate that this component has a higher peak and a relatively concentrated energy distribution.
[0071] The Hilbert transform is performed on the screened modal components to extract the instantaneous frequency feature and the phase synchronization index. The Hilbert transform can transform the real signal u(t) into an analytic signal where is the result of the Hilbert transform of u(t), and is calculated through The instantaneous frequency ω(t) can be obtained by taking the derivative of the phase of the analytic signal with respect to time, that is where The phase synchronization index is used to measure the degree of phase synchronization between different modal components, which is obtained by calculating the statistical characteristics of the phase differences between different analytic signals. For example, the standard deviation or cross-correlation coefficient of the phase differences is calculated. These instantaneous frequency characteristics and phase synchronization indexes can reflect the changes in the human physiological rhythm under the influence of alcohol, providing important characteristic bases for drunk driving detection.
[0072] Embodiment 3:
[0073] This embodiment mainly introduces the specific implementation details of the improved sparse coding algorithm. This algorithm enhances the features of the physiological rhythm signal to generate more representative heart rate variability indexes and blood oxygen saturation estimation values, improving the accuracy of drunk driving detection.
[0074] In the improved sparse coding algorithm, first, an asymmetric dictionary learning model is constructed. The purpose of dictionary learning is to learn an overcomplete dictionary D such that the input signal x can be represented by a sparse linear combination of this dictionary, that is, x = Da, where a is the sparse coefficient vector. The asymmetric dictionary learning model optimizes the joint objective function of the dictionary atoms and sparse coefficients on the basis of traditional dictionary learning, considering the importance differences of different feature dimensions.
[0075] A dynamic weighting factor is introduced into the sparse constraint term, and the dynamic weighting factor is adaptively adjusted according to the importance of the feature dimensions. For each feature dimension i, its importance measure w i is calculated. For example, it can be calculated according to the correlation between the feature dimension and the drunk driving detection result. The correlation can be obtained by calculating the Pearson correlation coefficient between the feature dimension and the known drunk driving sample labels. The Pearson correlation coefficient calculation formula is:
[0076]
[0077] where, x ij is the value of the i-th feature dimension in the j-th sample, is the mean of the i-th feature dimension, y j is the drunk driving label of the j-th sample (for example, drunk driving is 1 and non-drunk driving is 0), is the mean of the drunk driving labels. According to the calculated correlation r i , the dynamic weighting factor w i is adjusted. The higher the correlation, the larger w i is. The less the sparsity degree of this dimension is restricted in the sparse constraint, thus highlighting the important feature dimensions. The adjusted sparse constraint term becomes where m is the total number of feature dimensions.
[0078] Adopt a layer-by-layer incremental coding strategy and gradually optimize the sparse representation accuracy through a residual feedback mechanism. During the coding process, first, preliminarily encode the signal with the current dictionary to obtain sparse coefficients α1 and the reconstructed signal x1 = Dα1, and calculate the residual signal r1 = x - x1. Then, according to the characteristics of the residual signal, update the dictionary, for example, by adding new dictionary atoms or adjusting the parameters of existing dictionary atoms, to obtain the updated dictionary D1. Next, encode the residual signal with the updated dictionary to obtain new sparse coefficients α2, and calculate the reconstructed signal and the residual signal again. Repeat this process iteratively until the residual signal meets the preset accuracy requirements. Through this layer-by-layer incremental coding strategy and residual feedback mechanism, the sparse representation can be continuously optimized, more key features in the physiological rhythm signal can be accurately extracted, and more reliable heart rate variability indicators and blood oxygen saturation estimation values can be generated.
[0079] Example 4:
[0080] This example focuses on the construction and specific implementation process of a cascaded deep learning classifier, which is one of the core components of the drunk driving detection system, and its performance directly affects the accuracy of the detection results.
[0081] In the construction of the cascaded deep learning classifier, first design a temporal convolutional network module. The temporal convolutional network module uses dilated convolutional kernels to capture multi-scale temporal dependencies. The dilated convolutional kernel introduces a dilation rate parameter on the basis of the ordinary convolutional kernel, enabling the convolutional kernel to skip some pixels during convolution, thereby expanding the receptive field and being able to capture dependencies on longer time series. Assume the input feature sequence is X ∈ R T×C , where T is the time step and C is the feature dimension. The calculation formula for dilated convolution is:
[0082]
[0083] where y[i] is the value of the output feature at the i-th time step, w[j] is the convolutional kernel weight, k is the convolutional kernel size, d is the dilation rate, and x[i + dj] is the value of the input feature at the corresponding position. By setting different dilation rates, temporal features of different scales can be obtained. At the same time, to suppress the weights of irrelevant time steps, a gating mechanism is introduced. The gating mechanism calculates a gating signal g to weight the convolutional output. The calculation formula for the gating signal g is:
[0084] g = σ(W g x + b g )
[0085] where σ is the sigmoid function, W g and b gare learnable weights and biases, and x is the input feature. The value of the gating signal g ranges from 0 to 1, and by multiplying it with the convolutional output, it can selectively retain or suppress the feature information of certain time steps.
[0086] Construct a self-attention mechanism module to calculate the correlation matrix between feature dimensions through a multi-head attention layer. The multi-head attention mechanism maps the input features to multiple low-dimensional subspaces, calculates the attention weights separately in each subspace, and then concatenates the results. Assume the input features are Q, K, V ∈ R T×C , and the calculation process of the multi-head attention layer is as follows:
[0087] MultiHead(Q, K, V) = Concat(head1,..., head h )W o
[0088]
[0089] where h is the number of heads, and W O are learnable weight matrices, and d k is the dimension of K. Through the multi-head attention layer, the model can capture the correlation relationships between feature dimensions from different perspectives and enhance the learning ability for complex features. In addition, relative position encoding is adopted to enhance the temporal modeling ability. Relative position encoding calculates a relative position representation for each position pair and adds it to the original feature, enabling the model to better utilize the position information and capture the sequential relationships in the time series.
[0090] Introduce a feature recalibration module at the concatenation junction to dynamically adjust the channel weights of the feature map through the channel attention mechanism. The channel attention mechanism first performs global average pooling and global max pooling on the feature map in the spatial dimension to obtain the average pooling feature F avg and the max pooling feature F max . Then, these two features are respectively processed through a shared multi-layer perceptron (MIP) to obtain the channel attention weights M avg and M max . Finally, these two weights are added and activated through the sigmoid function to obtain the final channel attention weight M. In this way, the model can dynamically allocate weights according to the importance of different channel features, highlight the channel features that contribute significantly to drunk driving detection, and improve the performance of the classifier.
[0091] Example 5:
[0092] This embodiment details the dynamic threshold segmentation algorithm and the specific implementation methods of various functions in the model optimization module. These technologies can optimize the drunk driving detection results and improve the overall performance and generalization ability of the system.
[0093] In terms of the dynamic threshold segmentation algorithm, first, a probability density estimation function based on the Gaussian mixture model is constructed to perform multimodal modeling on the probability distribution of alcohol concentration. The Gaussian mixture model assumes that the data is composed of a mixture of multiple Gaussian distributions, and the parameters {π i , μ i , Σ i} of the Gaussian mixture model are estimated through the Expectation-Maximization algorithm (EM algorithm), so as to obtain an accurate modeling of the probability distribution of alcohol concentration. In practical applications, an appropriate K value is determined according to historical data and experience. For example, by trying different K values, observing the fitting effect of the model on the training data and the generalization ability on the test data, the optimal K is selected.
[0094] The improved mean shift algorithm is used to search for the local extreme points of the probability density function to generate a set of candidate decision boundaries. The mean shift algorithm is a density-based clustering algorithm, and its core idea is to find the direction of the density gradient ascent of data points in the data space, continuously move the window center until it converges to the local density maximum point. The improved mean shift algorithm introduces an adaptive bandwidth adjustment mechanism to address the problem that the traditional algorithm may fall into a local optimum when dealing with complex distributions. During the algorithm iteration, the bandwidth h is dynamically adjusted according to the distribution of data points within the current window. For example:
[0095] h t+1 = αh t + β
[0096] where h t is the bandwidth of the current iteration, h t+1 is the bandwidth of the next iteration, and α and β are parameters set according to data characteristics and experimental experience, with 0 < α < 1. In this way, the algorithm can more accurately search for the local extreme points of the probability density function, and the alcohol concentration values corresponding to these extreme points are the candidate decision boundaries.
[0097] Based on the KL divergence to measure the classification uncertainty of each candidate boundary, the segmentation threshold corresponding to the minimum uncertainty is selected. The KL divergence is used to measure the difference degree between two probability distributions. For two probability distributions P and Q, its KL divergence is defined as:
[0098]
[0099] In drunk driving detection, assuming that a certain candidate decision boundary divides the probability distribution of alcohol concentration into two sub-distributions P1 and P2, the KL divergences D between these two sub-distributions and the original distribution P are calculated respectively.KL (P1 ∥ P) and D KL (P2 ∥ P), and take the sum of the two as the classification uncertainty measure of this candidate boundary. Traverse all candidate decision boundaries and select the boundary that minimizes the classification uncertainty as the final segmentation threshold, so as to achieve the optimal decision boundary optimization for drunk driving detection results.
[0100] In the model optimization module, a data augmentation model is constructed based on the generative adversarial network. The generative adversarial network consists of a generator and a discriminator. The generator synthesizes facial images under different lighting and pose conditions through a conditional generator. The conditional generator takes the class label (whether drunk driving or not) of the original facial image and some random noise as inputs to generate diverse facial images. For example, when generating images under different lighting conditions, certain dimensions of the random noise are adjusted to simulate different lighting intensities and directions. A spectral consistency loss function is introduced in the discriminator to constrain the physical property consistency of the generated images among multi-spectral bands. The spectral consistency loss function can be constructed by calculating the correlation or difference of the generated images among different spectral bands.
[0101] Adopt the curriculum learning strategy to gradually increase the complexity of the generated data and improve the generalization performance of the classifier. The curriculum learning strategy first trains the generator and the discriminator starting from simple data. As the training progresses, the difficulty of the generated data is gradually increased. For example, first generate images with less lighting variation and simple pose changes, and then gradually increase the complexity of lighting and pose changes, so that the classifier can learn a wider range of features and improve the adaptability to drunk driving detection in different scenarios.
[0102] Construct a multi-task learning framework to jointly optimize the alcohol concentration prediction task and the physiological index regression task. Introduce an orthogonal regularization term in the loss function to constrain the feature representation spaces of different tasks to be independent. Assume that the loss function of the alcohol concentration prediction task is L1, the loss function of the physiological index regression task is L2, and the orthogonal regularization term is L ortho , then the total loss function is:
[0103] L = L1 + L2 + λ ortho L ortho
[0104] where, λ ortho is a hyperparameter that controls the intensity of orthogonal regularization. The orthogonal regularization term can be realized by calculating the inner product between the feature representation matrices of different tasks and performing constraints. For example:
[0105]
[0106] where, and They are the feature vectors of the alcohol concentration prediction task and the physiological index regression task respectively, and <·,·> represents the inner product operation. The gradient reversal layer is used to implement adversarial training to eliminate the influence of environmental interference factors on feature extraction. The gradient reversal layer keeps the input unchanged during forward propagation and multiplies the gradient by -1 during backward propagation, so that the two tasks are adversarial to each other during the training process, forcing the model to learn more robust features that are not affected by the environment.
[0107] An online incremental learning mechanism is designed to collect real-time detection data through a sliding window and calculate the feature drift amount. The size of the sliding window is set according to the time characteristics of the data and the computing resources, for example, set to include the data of the last T time steps. The feature drift amount can be measured by calculating the difference between the feature distribution of the data in the current window and the feature distribution of the historical data, for example, using the KL divergence or the Wasserstein distance. When the feature drift amount exceeds the preset threshold, the local network parameter update process is triggered, and the elastic weight consolidation algorithm is used to retain the prior knowledge of important parameters. The elastic weight consolidation algorithm adds a penalty term to the loss function to constrain the update of important parameters. The calculation formula of the penalty term is:
[0108]
[0109] where λ ewc is a hyperparameter that controls the penalty intensity, F w is the Fisher information matrix that measures the importance of the parameter w, and w0 is the initial value of the parameter. A momentum buffering mechanism is introduced during the model update process to smooth the parameter update direction. When calculating the parameter update, the momentum buffering mechanism not only considers the current gradient information but also combines the previous update direction. Its update formula is:
[0110]
[0111] w t = w t-1 + v t
[0112] where v t is the velocity of the t-th update, β is the momentum coefficient, 0 < β < 1, η is the learning rate, is the gradient of the t-th time. Through these optimization techniques, the system can continuously adapt to new data and environmental changes and maintain a high detection performance.
[0113] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0114] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A non-contact drunk driving detection method based on deep learning, characterized in that, The method includes: Collecting continuous video stream data of a target object through a multispectral camera, where the multispectral camera includes visible light band, near-infrared band, and thermal infrared band sensors; performing inter-frame alignment and noise suppression processing on the continuous video stream data to extract a sequence of facial region images of the target object; Extracting optical flow features from the sequence of facial region images based on a pre-trained spatio-temporal convolutional neural network to generate a dynamic change vector of facial micro-expressions; Fusing the optical flow features with the thermal infrared band data and separating the physiological rhythm signal of the target object through a time series decomposition algorithm; Enhancing the features of the physiological rhythm signal based on an improved sparse coding algorithm to generate a heart rate variability index and an estimated blood oxygen saturation value; Constructing a multi-modal fusion feature vector, where the multi-modal fusion feature vector includes the optical flow features, the physiological rhythm signal, the heart rate variability index, and the estimated blood oxygen saturation value; inputting the multi-modal fusion feature vector into a cascaded deep learning classifier, where the cascaded deep learning classifier consists of a temporal convolutional network and a self-attention mechanism module, and outputting the probability distribution of the alcohol concentration of the target object; Optimizing the decision boundary of the alcohol concentration probability distribution based on a dynamic threshold segmentation algorithm to generate a drunk driving detection result.
2. The non-contact drunk driving detection method according to claim 1, wherein The inter-frame alignment and noise suppression processing of the continuous video stream data includes: Calculating the translation and rotation parameters of adjacent video frames using a phase correlation algorithm and achieving pixel-level alignment between multi-spectral bands through bilinear interpolation; constructing a noise suppression model based on wavelet transform, performing adaptive threshold filtering on high-frequency sub-band coefficients, and retaining facial texture details; performing morphological closing operation on the noise-suppressed image for edge smoothing processing.
3. The non-contact drunk driving detection method according to claim 2, wherein The time series decomposition algorithm includes: Constructing a variational mode decomposition model and decomposing the physiological rhythm signal into multiple intrinsic mode functions by iteratively solving a constrained optimization problem; screening the mode components related to alcohol metabolism based on the joint criterion of kurtosis coefficient and energy entropy; performing Hilbert transform on the screened mode components to extract instantaneous frequency features and phase synchronization indexes.
4. The non-contact drunk driving detection method according to claim 1, characterized in that, The improved sparse coding algorithm includes: Constructing an asymmetric dictionary learning model and optimizing the joint objective function of dictionary atoms and sparse coefficients by the alternating direction method of multipliers; introducing a dynamic weighting factor into the sparse constraint term, where the dynamic weighting factor is adaptively adjusted according to the importance of feature dimensions; adopting a layer-by-layer incremental coding strategy and gradually optimizing the sparse representation accuracy through a residual feedback mechanism.
5. The non-contact drunk driving detection method according to claim 1, characterized in that, The construction of the cascaded deep learning classifier includes: Designing a temporal convolutional network module, using dilated convolutional kernels to capture multi-scale temporal dependencies and suppressing the weights of irrelevant time steps through a gating mechanism; constructing a self-attention mechanism module, calculating the correlation matrix between feature dimensions through a multi-head attention layer, and enhancing the temporal modeling ability by using relative position encoding; introducing a feature recalibration module at the cascade connection and dynamically adjusting the channel weights of the feature map through a channel attention mechanism.
6. The non-contact drunk driving detection method according to claim 1, wherein, The dynamic threshold segmentation algorithm includes: Construct a probability density estimation function based on the Gaussian mixture model to perform multimodal modeling on the probability distribution of the alcohol concentration; use an improved mean shift algorithm to search for the local extreme points of the probability density function and generate a set of candidate decision boundaries; based on the KL divergence, measure the classification uncertainty of each candidate boundary and select the segmentation threshold corresponding to the minimum uncertainty.
7. The non-contact drunk driving detection method according to claim 1, characterized in that, It further includes: Construct a data augmentation model based on the generative adversarial network, and synthesize facial images under different lighting and pose conditions through the conditional generator; introduce a spectral consistency loss function in the discriminator to constrain the physical property consistency of the generated images among multi-spectral bands; adopt a curriculum learning strategy to gradually increase the complexity of the generated data and improve the generalization performance of the classifier.
8. The non-contact drunk driving detection method according to claim 1, characterized in that, It further includes: Construct a multi-task learning framework to jointly optimize the alcohol concentration prediction task and the physiological index regression task; introduce an orthogonal regularization term in the loss function to constrain the feature representation spaces of different tasks to be independent of each other; use a gradient reversal layer to implement adversarial training to eliminate the influence of environmental interference factors on feature extraction.
9. The non-contact drunk driving detection method according to claim 1, characterized in that, It further includes: Design an online incremental learning mechanism, collect real-time detection data through a sliding window and calculate the feature drift amount; When the feature drift amount exceeds the preset threshold, trigger the local network parameter update process, and use the elastic weight consolidation algorithm to retain the prior knowledge of important parameters; introduce a momentum buffer mechanism during the model update process to smooth the parameter update direction.
10. A non-contact drunk driving detection system based on deep learning, characterized in that, It includes: A data acquisition module for obtaining continuous video stream data of the target object through a multi-spectral camera; A preprocessing module for performing inter-frame alignment, noise suppression, and facial region extraction; A feature extraction module containing a spatio-temporal convolutional neural network, a time series decomposition algorithm, and an improved sparse coding algorithm; A classification decision module composed of a cascaded deep learning classifier and a dynamic threshold segmentation algorithm; A model optimization module for implementing the functions of generative adversarial network enhancement, multi-task learning, and online incremental learning.
Citation Information
Cited By
Drunk driving automatic detection method and system based on deep learning model and robot system
CN120894671A
Soybean phenological period automatic identification method based on unmanned aerial vehicle multi-modal data
CN121564536A
Deep sleep recognition and intervention method and system based on heart rate variation analysis
CN121795847A