An abnormal driving recognition method based on pedal data and facial expression multi-modal
By integrating pedal data with facial expression data into a multimodal recognition method, and utilizing one-dimensional and three-dimensional convolutional networks to extract features, combined with physical rules and noise diffusion techniques, the problem of insufficient multimodal data fusion in existing technologies is solved, achieving more accurate and reliable abnormal driving recognition, and improving driving safety and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2025-08-22
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies for driver behavior recognition suffer from insufficient multimodal data fusion, inadequate in-depth analysis of driver behavior, and a lack of consideration for multiple factors, resulting in insufficient accuracy and reliability in identifying abnormal driving behavior and making it difficult to meet the safety requirements of real-world driving scenarios.
An abnormal driving identification method based on pedal data and facial expression multimodal data is adopted. By collecting vehicle driving data and facial video data, one-dimensional and three-dimensional convolutional networks are used to extract features. Combined with physical rules and noise diffusion technology, an abnormal driving model is constructed, and data filtering and model enhancement are performed to achieve deep fusion of multimodal data and consistent prediction judgment.
It improves the comprehensiveness of abnormal driving mode recognition, solves the problem of data imbalance, enhances the accuracy and reliability of recognition, and improves driving safety and user experience through a graded response mechanism.
Smart Images

Figure CN121188645B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of abnormal driving recognition methods, specifically an abnormal driving recognition method based on pedal data and facial expression multimodal data. Background Technology
[0002] With the continuous increase in car ownership, road traffic safety is becoming increasingly severe. Statistics show that improper driver operation is a major cause of traffic accidents, with pedal operation errors and abnormal driving behaviors accounting for a significant proportion of these accidents. To address this situation, numerous patented technologies related to driving behavior monitoring and anomaly recognition have emerged. However, current technologies still face many unresolved issues, primarily the following:
[0003] Limitations of Data Utilization and Analysis: Some existing technologies have shortcomings in data collection and analysis. For example, the "A Visual Perception-Based Method and System for Detecting Correct Pedal Use" primarily uses a facial recognition camera to monitor the driver's driving state and a leg camera to monitor leg state to determine pedal usage. This method relies solely on visual perception and does not fully integrate multimodal data from the vehicle's driving process. Key data such as pedal opening, vehicle speed, and motor torque are not included in the analysis system. When judging abnormal driving behavior, it relies solely on simple state monitoring and rule-based judgment, failing to delve into the potential connections between data. This makes it difficult to effectively identify complex and ever-changing abnormal driving patterns, significantly reducing the accuracy and reliability of the detection results.
[0004] The one-sidedness of behavioral judgment: The "Driving Behavior Detection System and Method" collects multi-source data from drivers, including facial expressions, eye movements, head posture, hand gestures, and pedal operations, through multiple modules. However, when processing this data, it only uses simple rules or threshold judgments to issue warnings, lacking in-depth data fusion and analysis. This results in problems such as low sensitivity, poor adaptability, and low intelligence. It cannot provide personalized and intelligent warnings based on individual differences among drivers and dynamic changes during the driving process, leading to unsatisfactory warning effects and failing to meet the stringent safety requirements of real-world driving scenarios.
[0005] The lack of multi-factor consideration: Existing driving behavior detection technologies often pay little attention to the impact of factors such as road environment, passenger behavior, and traffic conditions on the identification of abnormal driving behavior. The paper "A Method and System for Judging Abnormal Driving Behavior Based on Multimodality" points out that most studies only focus on human and vehicle factors for abnormal driving behavior detection, neglecting other important factors such as the road environment. In actual driving, environmental factors such as road conditions and traffic flow, as well as passenger behavior, all affect the driver's actions. Single-dimensional perception cannot provide a comprehensive basis for accurately judging abnormal driving behavior, hindering timely and accurate identification of abnormal driving behavior and making it difficult to effectively prevent traffic accidents.
[0006] In the field of driving behavior detection and abnormal driving behavior judgment, existing technologies have significant shortcomings in multimodal data fusion, in-depth analysis of driving behavior, and consideration of multiple factors. These deficiencies severely restrict the development and application of driving safety technologies, and there is an urgent need to develop innovative technical solutions to achieve accurate identification and early warning of abnormal driving behavior, improve driving safety, and reduce traffic accidents. Summary of the Invention
[0007] This invention provides an abnormal driving recognition method based on pedal data and facial expression multimodal analysis to solve the problems of existing technologies in multimodal data fusion, in-depth analysis of driving behavior, and multi-factor collaborative consideration.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0009] An abnormal driving recognition method based on pedal data and facial expression multimodal analysis is presented as follows:
[0010] Step 1: Collect vehicle driving data from multiple identical time observation windows. n Face video data V n Dr, the vehicle driving data in each observation time window n All include raw accelerator pedal opening data A n Original brake pedal opening data B n Original vehicle speed data S n Original motor torque data M n For each person's facial video data V n Label the abnormal facial expressions of real people with the tag Y. gt n Dr. n Face video data V n And the corresponding real-life abnormal facial expression label Y gt n The original multimodal samples D are respectively constructed. nAnd from the original multimodal samples D of each time observation window n The original multimodal dataset SetD0 is composed of these components.
[0011] Step 2: Based on the original multimodal dataset SetD0 obtained in Step 1, train an original abnormal driving model Model_Dr0;
[0012] Step 3: Maintain the original multimodal sample D n The real face abnormal expression tag Y gt n The original multimodal sample D remains unchanged. n Dr. n Face video data V n Noise diffusion was performed separately, and the corresponding vehicle driving noise diffusion data Dr' was obtained. n Face video noise diffusion data V' n ;
[0013] Dr's driving noise diffusion data for each vehicle n Face video noise diffusion data V' n And the corresponding real-life abnormal facial expression label Y gt n Each is used as a diffusion to generate multimodal samples D' n ;
[0014] Then, multimodal samples D' are generated based on each diffusion. n Vehicle driving noise diffusion data Dr' n Motor torque analysis is performed, and multimodal samples D' are generated from various diffusions based on the motor torque analysis results. n In the process, valid multimodal samples D that meet the requirements are selected. n To construct the multimodal dataset SetD”1 after motor torque analysis;
[0015] Step 4: Using the original abnormal driving model Model_Dr0 obtained in Step 2, analyze each valid multimodal sample D in the multimodal dataset SetD”1 obtained after motor torque analysis in Step 3. n Make predictions to obtain D for each effective multimodal sample. n Predicted labels Based on each valid multimodal sample D” n Predicted labels And the corresponding real-life abnormal facial expression label Y gt n Perform a prediction consistency assessment, and select the valid multimodal samples D that meet the requirements based on the assessment results. n As a multimodal sample D with predictive consistencyn And through each prediction consistent multimodal sample D”' n Construct a prediction consistency multimodal dataset SetD”'1;
[0016] Step 5: Using the prediction consistency multimodal dataset SetD”'1 obtained in Step 4, retrain the original abnormal driving model Model_Dr0 to enhance the original abnormal driving model Model_Dr0. The retrained and enhanced original abnormal driving model Model_Dr0 is used as the abnormal driving recognition model Model_Dr”.
[0017] Step 6: Input the current vehicle signal and the current face video signal into the abnormal driving recognition model "Model_Dr" obtained in Step 5. The abnormal driving prediction result is obtained by the abnormal driving recognition model "Model_Dr".
[0018] Furthermore, in step 1, multiple categories of abnormal facial expressions during driving are defined; a pre-trained ResNet model is used to predict the V-value of each face video. n The probability vector y of various abnormal facial expression labels for each frame of the expression. ti n Each face video V is output by a pre-trained ResNet model. n The probability vector y of similar expressions across all frames ti n The probability sum is calculated, and the category corresponding to the largest probability sum is taken as the corresponding face video V. n Realistic facial abnormal expression tag Y gt n .
[0019] Furthermore, in step 2, the original abnormal driving model Model_Dr0 includes a one-dimensional convolutional network, a three-dimensional convolutional network, a first fully connected layer, and a second fully connected layer;
[0020] During training, each original multimodal sample D from the original multimodal dataset SetD0 is processed through a one-dimensional convolutional network. n Dr. n Extracting vehicle driving features F_car n The original multimodal samples D from the original multimodal dataset SetD0 are extracted using a 3D convolutional network. n Face video data V n Extracting facial features F_face n ; Each original multimodal sample D n Vehicle driving characteristics F_car n Facial features F_face nAfter concatenation, the fused feature F_fuse is obtained by processing through the first fully connected layer. n Then, the fused features F_fuse are processed through a second fully connected layer. n Perform feature mapping to predict each original multimodal sample D. n The predicted probability vector p of abnormal facial expression labels n Finally, based on each original multimodal sample D n The predicted probability vector p n The real face abnormal expression label Y in the real face abnormal expression label set SetY0 gt n The loss function is calculated, and the parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer are updated by backpropagation based on the loss function calculation results, thereby training the original abnormal driving model Model_Dr0.
[0021] Furthermore, in step 3, the statistics of each original multimodal sample D are performed. n Dr. n In the original accelerator pedal opening, the maximum value maxA is... n The noise range coefficient δ_A is calculated, and the accelerator pedal opening noise ε_A is generated. Then, based on each original multimodal sample D, the noise range coefficient δ_A is calculated. n The original accelerator pedal opening data A n Generate accelerator pedal opening noise diffusion data A' n =A n +ε_A;
[0022] Statistical analysis of each original multimodal sample D n Dr. n In the original brake pedal opening, the maximum value maxB is... n The noise range coefficient δ_B is calculated, and the brake pedal opening noise ε_B is generated. Then, based on each original multimodal sample D, the noise range coefficient δ_B is calculated. n The original brake pedal opening data B n Generate brake pedal opening noise diffusion data B' n =B n +ε_B;
[0023] Based on the reasonable fluctuation range of vehicle speed in actual driving, a fixed tolerance δ_S for vehicle speed noise is set, and vehicle speed noise ε_S is generated. Then, based on each original multimodal sample D n Dr. n The original vehicle speed data S n Generate vehicle speed noise diffusion data S' n =S n +ε_S;
[0024] Based on the reasonable fluctuation range of the motor output torque, a fixed tolerance δ_M for the motor torque noise is set, and the motor torque noise ε_M is generated. Then, based on each original multimodal sample D n Dr. n The original motor torque data M n Generate motor torque noise diffusion data M' n =M n +ε_M;
[0025] Therefore, based on each original multimodal sample D n Dr. n The corresponding vehicle driving noise diffusion data Dr' were constructed and generated respectively. n ={A' n ,B' n ,S' n ,M' n};
[0026] Define the video noise variance σ_V and generate face video noise ε_V, then based on each original multimodal sample D n V facial video data n Generate face video noise diffusion data V' n =V n +ε_V;
[0027] Preserve realistic facial abnormal expression tags Y gt n Unchanged, serving as the label Y' for abnormal facial expressions after noise diffusion. n , i.e. Y' n =Y gt n ;
[0028] Finally, Dr' integrates vehicle driving noise diffusion data. n Face video noise diffusion data V' n And the corresponding facial abnormality label Y' after noise diffusion n Generate diffusion to generate multimodal samples D' n ={Dr' n ,V' n ,Y' n}
[0029] Furthermore, in step 3, the motor torque analysis process is as follows:
[0030] Based on the braking effectiveness rules, define the effective braking trigger condition and the abnormal vehicle speed without deceleration condition, and determine the multimodal sample D' generated by each diffusion. n Vehicle driving noise diffusion data Dr' nDoes the effective braking trigger condition and the abnormal vehicle speed without deceleration condition both meet simultaneously? If both conditions are met, then based on the vehicle driving noise diffusion data Dr' n Using the brake pedal opening data, vehicle speed data, and defined conditions, calculate each diffusion-generated multimodal sample D'. n Dr's data on vehicle driving noise diffusion n The brake effectiveness rule value Rule_brake n ;
[0031] Based on the physical basis of the dual-pedal operation conflict rule value, accelerator pedal activation threshold and brake pedal activation threshold are set, and acceleration operation intensity quantization function and braking operation intensity quantization function are defined; each diffusion generates a multimodal sample D'. n Vehicle driving noise diffusion data Dr' n The difference between the accelerator pedal opening data and the set accelerator pedal activation threshold is substituted into the acceleration operation intensity quantization function to calculate the accelerator pedal operation intensity; each diffusion generates a multimodal sample D'. n Vehicle driving noise diffusion data Dr' n The difference between the brake pedal opening data and the set brake pedal activation threshold is substituted into the brake operation intensity quantization function to calculate the brake pedal operation intensity. Finally, based on the accelerator pedal operation intensity and the brake pedal operation intensity, the multimodal sample D' generated by diffusion is calculated. n Dr's data on vehicle driving noise diffusion n The dual-pedal operation conflict rule value is Rule_conflict. n ;
[0032] Based on the physical basis of torque dynamics rules, an acceleration prediction model is established, and the least squares method is used to predict vehicle driving data from multiple observation time windows collected in step 1. n The acceleration prediction model is fitted to minimize the deviation between the predicted and actual accelerations, thereby obtaining the coefficients and corresponding acceleration errors in the model. Then, based on these coefficients and the acceleration errors, the multimodal sample D' generated by diffusion is calculated. n Dr's data on vehicle driving noise diffusion n Torque dynamics rule value Rule_torque n ;
[0033] Each diffusion generates a multimodal sample D'. n Dr's data on vehicle driving noise diffusion n The brake effectiveness rule value Rule_brake nDual-pedal operation conflict rule value: Rule_conflict n Torque dynamics rule value Rule_torque n Each rule is compared with its corresponding threshold value, and the rule value of brake effectiveness is determined by the value of brake_brake. n Dual-pedal operation conflict rule value: Rule_conflict n Torque dynamics rule value Rule_torque n When all values are less than or equal to their respective thresholds, then the corresponding diffusion generates multimodal samples D'. n Vehicle driving noise diffusion data Dr' n It conforms to physical laws, and determines that the corresponding diffusion generates multimodal samples D'. n For effective multimodal samples D” n .
[0034] Furthermore, in step 4, when a certain valid multimodal sample D” n Predicted labels Corresponding real-life abnormal facial expression label Y gt n If they are the same, then the valid multimodal sample D is determined. n If the prediction is consistent, then the valid multimodal sample is determined to be D. n This does not conform to the prediction.
[0035] Furthermore, in steps 2 and 5, the cross-loss function is used during training, and the parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer are updated by backpropagation based on the calculation results of the cross-loss function.
[0036] Furthermore, in step 6, based on the abnormal driving prediction results obtained from the prediction... Calculate the risks and implement corresponding graded controls based on the obtained risk levels.
[0037] Compared with the prior art, the advantages of the present invention are:
[0038] Deep fusion of multimodal data enhances the comprehensiveness of recognition: By fusing pedal operation data with facial expression data, we not only focus on the physical characteristics of vehicle operation behavior, but also combine the driver's physiological and emotional state. Compared with recognition methods based on a single data type, this approach can more comprehensively capture abnormal driving patterns and reduce misjudgments caused by incomplete information.
[0039] The combination of data generation and filtering addresses the problem of imbalanced data: Absorption generation technology is used to expand abnormal driving-related data, while motor torque analysis and prediction consistency verification are used to perform dual filtering on the generated data to ensure the physical rationality and label consistency of the generated data. This effectively alleviates the imbalance between normal and abnormal categories in the original data and improves the model's generalization ability.
[0040] Introducing physical rules to enhance data reliability: During the data screening process, torque dynamics rules are constructed based on physical principles such as Newton's second law to verify the physical feasibility of the generated data and eliminate abnormal samples that do not conform to the laws of vehicle motion. Compared with methods that rely solely on data-driven approaches, this reduces the interference of invalid data on model training and improves recognition accuracy.
[0041] A tiered response mechanism that balances safety and practicality: Different control commands are executed based on the level of abnormal driving risk, avoiding the limitations of a single intervention method. While ensuring driving safety, it reduces excessive interference with the normal driving process and improves the user experience. Attached Figure Description
[0042] Figure 1 This is a summary diagram of the overall technical flow of an embodiment of the present invention.
[0043] Figure 2 This is a schematic diagram of the acquisition of a multimodal dataset based on a pedal-face in an embodiment of the present invention.
[0044] Figure 3 This is a flowchart of the abnormal driving model training process based on initial data according to an embodiment of the present invention.
[0045] Figure 4 This is a schematic diagram of driving data filtering based on motor torque analysis according to an embodiment of the present invention.
[0046] Figure 5 This is a schematic diagram of driving data filtering based on predictive score analysis according to an embodiment of the present invention.
[0047] Figure 6 This is a flowchart of the abnormal driving model training process based on generated data according to an embodiment of the present invention.
[0048] Figure 7 This is a schematic diagram of abnormal driving model testing and response based on driving data according to an embodiment of the present invention. Detailed Implementation
[0049] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0050] like Figure 1 As shown in the figure, this embodiment discloses an abnormal driving recognition method based on pedal data and facial expression multimodality, the process of which is as follows:
[0051] Step 1, as follows Figure 2 As shown, multiple identical observation time windows t are set.
[0052] Dr collects vehicle driving data in each observation time window t n Dr. n Construct the original vehicle driving dataset SetDr0, where the vehicle driving data Dr in each observation time window n All include raw accelerator pedal opening data A n Original brake pedal opening data B n Original vehicle speed data S n Original motor torque data M n .
[0053] Collect face video data V for each observation time window t n Face video data V from each observation time window t n Construct the original human face video sample set SetV0.
[0054] For each face video data V in the original face video sample set SetV0 n Label the abnormal facial expressions of real people with the tag Y. gt n V, composed of all facial video data n Realistic facial abnormal expression tag Y gt n Construct a set of labels for abnormal facial expressions based on real human faces (SetY0).
[0055] Finally, the vehicle driving data Dr in each observation time window t n Face video data V n And the corresponding real-life abnormal facial expression label Y gt n The original multimodal samples D, which respectively constitute the observation time windows, n And composed of the original multimodal samples D from all observation time windows. n The original multimodal dataset SetD0 is constructed.
[0056] Specifically, the construction instructions for the original vehicle driving dataset SetDr0 are as follows:
[0057] In this embodiment, the observation time window is set to t = 1...T, where T = 20 frames, the sample sampling rate is 10Hz, the duration covered by the time window is 2 seconds, and it includes at least one expression change, with each expression lasting approximately 0.5 to 2 seconds.
[0058] In this embodiment, the vehicle driving data Dr collected in each observation time window t includes the original accelerator pedal opening data A, the original brake pedal opening data B, the original vehicle speed data S, and the original motor torque data M. Wherein:
[0059] The collected raw accelerator pedal opening data A = {a t}, t=1...T, a t The accelerator pedal opening value at time point t is expressed as a percentage, ranging from 0-100%, with a sampling rate of 10Hz. In this embodiment, the accelerator pedal opening is measured using a Hall effect position sensor in the vehicle. The rotation of the accelerator pedal changes the magnetic field, which the sensor converts into a voltage signal. The vehicle control unit then converts the voltage signal into the accelerator pedal opening value based on the correlation curve between the voltage signal and the opening.
[0060] The collected raw brake pedal opening data B = {b t}, t=1...T, b t The value of the brake pedal opening at time point t is expressed as a percentage, ranging from 0-100%, with a sampling rate of 10Hz. In this embodiment, the brake pedal opening signal is also measured using a Hall effect position sensor of the vehicle. The rotation of the brake pedal changes the magnetic field, which the sensor converts into a voltage signal. The vehicle control unit converts the voltage signal into the brake pedal opening based on the correlation curve between the voltage signal and the opening.
[0061] The collected raw vehicle speed data S = {s t}, t=1...T, s t The vehicle speed value is at time point t, in km / h, with a sampling rate of 10Hz. This embodiment measures the vehicle speed signal using a Hall effect vehicle speed sensor.
[0062] The collected raw motor torque data M = {m t}, t=1...T, m t The value represents the motor torque at time point t, in N·m, with a sampling rate of 10Hz. In this embodiment, the motor torque signal is measured by directly connecting a dynamic torque sensor via a coupling.
[0063] In this embodiment, the vehicle driving data Dr = {A, B, S, M} is collected in each observation time window t. Finally, by collecting vehicle driving data for a total of N0 observation time windows, the original vehicle driving dataset SetDr0 = {Dr... n}, n=1...N0. Where: N0 is the total number of observation time windows, and the subscript 0 indicates that the data is the original data; Dr n Let Dr be the vehicle driving data for the nth observation time window. n ={A n B n ,Sn M n}, A n For the raw accelerator pedal opening data of the nth observation time window, B n S represents the raw brake pedal opening data for the nth observation time window. n M represents the raw vehicle speed data for the nth observation time window. n This is the raw motor torque data for the nth observation time window.
[0064] The construction instructions for the original face video sample set SetV0 are as follows:
[0065] In this embodiment, the face video V = {v} collected for each observation time window t t}, t=1...T, where v t For video frames at time point t, the sampling rate is 10Hz, and each video frame includes three RGB channels, with each channel having a resolution of 224x224. By collecting face video data for a total of N0 observation time windows, the original face video sample set SetV0={V n}, n=1...N0.
[0066] The construction instructions for the SetY0 real-face abnormal expression label set are as follows:
[0067] In this embodiment, the abnormal facial expression label category C = {c1, c2, c3, c4} is defined when driving, where category c1 = fatigue, category c2 = anger, category c3 = tension, and category c4 = neutral.
[0068] For each face video V in the original face video sample set SetV0 n Predict the V value of each face video using a pre-trained ResNet model. n The probability vector y of various abnormal facial expression labels for each frame of the expression. ti n ,t=1...T,ci∈C, each face video V is output by a pre-trained ResNet model. n The probability vector y of similar expressions across all frames ti n The probability sum is calculated, and the category corresponding to the largest probability sum is taken as the corresponding face video V. n Realistic facial abnormal expression tag Y gt n That is, Y gt n =max i (∑ t y ti This completes the process of analyzing each person's facial video. n Labeling abnormal facial expressions using the Y tag.gt n , where max i This indicates the category with the highest sum of probabilities.
[0069] By analyzing all face videos V in the original face video sample set SetV0 n Annotation was performed to obtain the V of all face videos. n Realistic facial abnormal expression tag Y gt n and with all facial video V n Realistic facial abnormal expression tag Y gt n Construct a set of labels for abnormal facial expressions: SetY0 = {Y gt n}, n=1...N0, where Y gt n ∈C.
[0070] Ultimately, in this embodiment, the vehicle driving data Dr in each observation time window t is used. n Face video data V n Real facial abnormal expression tag Y gt n These constitute the original multimodal samples D. n ={Dr n V n ,Y gt n}. And composed of the original multimodal samples D from all observation time windows. n The original multimodal dataset SetD0 = {D} is composed of... n}={Dr n V n ,Y gt n}, n=1...N0.
[0071] Step 2: Based on the original multimodal dataset SetD0 obtained in Step 1, train an original abnormal driving model Model_Dr0.
[0072] In this embodiment, a one-dimensional convolutional network is used to extract each original multimodal sample D from the original multimodal dataset SetD0. n Dr. n Extracting vehicle driving features F_car n The original multimodal samples D from the original multimodal dataset SetD0 are extracted using a 3D convolutional network. n Face video data V n Extracting facial features F_face n Then, for each original multimodal sample D...n Vehicle driving characteristics F_car n Facial features F_face n After concatenation, the fused feature F_fuse is obtained by processing through the first fully connected layer. n Then, the fused features F_fuse are processed through a second fully connected layer. n Perform feature mapping to predict each original multimodal sample D. n The predicted probability vector p of abnormal facial expression labels n ={p n,i},∑ i p n,i =1, i=1,2,3,4. Finally, based on each original multimodal sample D n The predicted probability vector p n The real face abnormal expression label Y in the real face abnormal expression label set SetY0 gt n The loss function is calculated, and the parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer are updated by backpropagation based on the loss function calculation result. In this way, an original abnormal driving model Model_Dr0 is trained, which is composed of the trained one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer.
[0073] Specific examples Figure 3 As shown, the process of training the original abnormal driving model Model_Dr0 is as follows:
[0074] (A1) Input each original multimodal sample D into the one-dimensional convolutional network. n The original vehicle driving dataset SetDr0 = {Dr n}, n = 1...N0. A sample of vehicle driving data Dr n Includes four channels of signal, namely Dr n ={A n B n ,S n M n}where A n ={a t}, B n ={b t}, S n ={s t}, M n ={m t}, t=1...T, at this time Dr n The feature dimension is 4×T, where T = 20 frames.
[0075] A one-dimensional convolutional network consists of four layers. The first layer is a convolutional layer, the second layer is a pooling layer, the third layer is a convolutional layer, and the fourth layer is a fully connected layer. The first convolutional layer has 16 output channels, a kernel width of 3, and a stride of 1. The second pooling layer has a kernel size of 2. The third convolutional layer has 32 output channels, a kernel width of 3, and a stride of 1. The fourth layer stretches the features into one-dimensional features and uses a fully connected layer to output vehicle driving features.
[0076] Each vehicle driving data sample Dr in the original vehicle driving dataset SetDr0 n The input is fed into a one-dimensional convolutional network, which processes the data through the first layer of convolution, the second layer of pooling, the third layer of convolution, and the fourth layer, which pulls the data into one-dimensional features before fully connected processing. Finally, the fourth layer outputs the vehicle driving feature F_car. n Vehicle driving characteristics F_car n The feature dimension is 1×128.
[0077] (A2) Input each original multimodal sample D into the 3D convolutional network. n The original face video sample set SetV0 contains face video data V n ={V nt}, t=1...T, face video V n The feature dimensions are H×W×C×T, where H is the video frame height, W is the video frame width, C is the number of video channels, H = 244, W = 224, the number of RGB channels C = 3, and T = 20 frames.
[0078] The 3D convolutional network consists of eight layers. The first layer is a convolutional layer with 64 output channels, a kernel size of (3×3×3), and a stride of (1,2,2). The second layer is a pooling layer with a kernel size of (1×2×2). The third layer is a convolutional layer with 128 output channels, a kernel size of (3×3×3), and a stride of (1,1,1). The fourth layer is a convolutional layer with 256 output channels, a kernel size of (3×3×3), and a stride of (2,2,2). The fifth layer is a convolutional layer with 512 output channels, a kernel size of (3×3×3), and a stride of (2,2,2). The sixth layer is a convolutional layer with 1024 output channels, a kernel size of (3×3×3), and a stride of (2,2,2). The seventh layer uses global spatiotemporal average pooling. The eighth layer is a fully connected layer that outputs facial features.
[0079] Take each face video data V from the original face video sample set SetV0. nThe input is fed into a 3D convolutional network, and processed through the network's first convolutional layer, second pooling layer, third convolutional layer, fourth convolutional layer, fifth convolutional layer, sixth convolutional layer, seventh global spatiotemporal average pooling layer, and eighth connection layer. Finally, the eighth layer outputs the face feature F_face. n Facial features F_face n The feature dimension is 1×128.
[0080] (A3) The vehicle driving feature F_car with a feature dimension of 1×128 n Face features F_face with a feature dimension of 1×128 n Perform feature concatenation to obtain the concatenated feature F_concat n =Concat(F_car n ;F_face n ), where Concat() represents concatenation processing, and the concatenation feature is F_concat n The feature dimension is 1×256. Then, the first fully connected layer is used to process and concatenate the features F_concat. n This yields D for each original multimodal sample. n The fusion feature F_fuse n , F_fuse n The feature dimension is 1×128.
[0081] (A4) The fused feature F_fuse is applied through the second fully connected layer. n Perform feature mapping to predict each original multimodal sample D. n The predicted probability vector p of abnormal facial expression labels n ={p n,i},∑ i p n,i =1, i=1,2,3,4.
[0082] (A5) Input each original multimodal sample D n Realistic facial expression abnormality label set SetY0 Realistic facial expression abnormality label Y gt n Calculate the Y-label of each real human face abnormal expression. gt n one-hot vector Vet n ={vet n,i}, when the real human face shows an abnormal expression label Y gt n The corresponding index value is Y gt n =Ci, the abnormal facial expression label Y of the real human face gtn one-hot vector vet n,i =1, otherwise it is a one-hot vector vet n,i =0.
[0083] Based on various real-life facial abnormal expression labels Y gt n one-hot vectors and each original multimodal sample D n The predicted probability vector p n The cross-loss function Loss is calculated as follows:
[0084] Loss=-∑n∑ i=1 4 vet n,i ·log(p n,i )
[0085] Based on the calculation results of the cross loss function, the parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer are updated by backpropagation. The original abnormal driving model Model_Dr0 is composed of the trained one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer.
[0086] Step 3: For each original multimodal sample D in the original multimodal dataset SetD0 n Dr. n Face video data V n Noise diffusion was performed to obtain vehicle driving noise diffusion data Dr' n Face video noise diffusion data V' n ;Based on vehicle driving noise diffusion data Dr' n Face video noise diffusion data V' n The set of real-face abnormal expression tags SetY0 contains the real-face abnormal expression tag Y. gt n Composition diffusion generates multimodal samples D' n Multimodal samples D' are generated by diffusion. n Construct a diffusion-generated multimodal dataset SetD1. Based on the diffusion-generated multimodal dataset SetD1, generate vehicle driving noise diffusion data Dr' n Motor torque analysis is performed, and multimodal samples D' are generated from the diffusion based on the motor torque analysis results. n In the process, valid multimodal samples D that meet the requirements are selected. n To construct the multimodal dataset SetD”1 after motor torque analysis.
[0087] In this embodiment, the process of constructing the multimodal dataset SetD1 through diffusion generation is as follows:
[0088] (B1) Input each original multimodal sample D n ={Dr n V n ,Y gt n}, n=1...N0, where Dr n ={A n B n ,S n M n} represents vehicle driving data, V n Y is the raw face video data. gt n Tags for abnormal facial expressions on real people;
[0089] (B2) Generate accelerator pedal opening noise diffusion data A' n The details are as follows:
[0090] The maximum value among the original accelerator pedal openings is denoted as maxA. n Calculate the noise range factor δ_A = 0.1 × maxA n This generates accelerator pedal opening noise ε_A ~ Uniform(-δ_A,δ_A). Finally, accelerator pedal opening noise diffusion data A' is generated. n =A n +ε_A, where Uniform() indicates that the noise ε_A follows a uniform distribution in the interval (-δ_A, δ_A), that is, the value of ε_A is randomly distributed between -δ_A and δ_A, and the probability of taking any value in this interval is equal.
[0091] (B3) Generate brake pedal opening noise diffusion data B' n The details are as follows:
[0092] The maximum value among the original brake pedal openings is denoted as maxB. n Calculate the noise range factor δ_B = 0.1 × maxB n This generates brake pedal opening noise ε_B ~ Uniform(-δ_B,δ_B). Finally, brake pedal opening noise diffusion data B' is generated. n =B n +ε_B.
[0093] (B4) Generate vehicle speed noise diffusion data S' n The details are as follows:
[0094] Based on the reasonable fluctuation range of vehicle speed in actual driving, a fixed tolerance of vehicle speed noise δ_S = 5 km / h is set, and vehicle speed noise ε_S ~ Uniform(-δ_S,δ_S) is generated. Finally, vehicle speed noise diffusion data S' is generated. n =S n +ε_S.
[0095] (B5) Generate motor torque noise diffusion data M' n The details are as follows:
[0096] Based on the reasonable fluctuation range of the motor output torque, a fixed tolerance of δ_M = 10 N·m for motor torque noise is set, and motor torque noise ε_M ~ Uniform(-δ_M,δ_M) is generated. Finally, motor torque noise diffusion data M' is generated. n =M n +ε_M.
[0097] (B6) Generate face video noise diffusion data V' n The details are as follows:
[0098] Set the video noise variance σ_V = 0.01 2 This generates face video noise ε_V ~ N(0,σ_V). Finally, it generates face video noise diffusion data V'. n =V n +ε_V.
[0099] (B7) Integration generates diffusion to generate multimodal samples D' n The details are as follows:
[0100] Preserve realistic facial abnormal expression tags Y gt n Unchanged, serving as the label Y' for abnormal facial expressions after noise diffusion. n , i.e. Y' n =Y gt n .
[0101] Constructing and generating vehicle driving noise diffusion data Dr' n ={A' n ,B' n ,S' n ,M' n}
[0102] Finally, Dr' integrates vehicle driving noise diffusion data. n Face video noise diffusion data V' n Abnormal facial expressions after noise diffusion (label Y') n Generate diffusion to generate multimodal samples D' n ={Dr' n ,V'n ,Y' n}
[0103] (B8) Perform steps (B1)-(B7) on all original multimodal samples of the original multimodal dataset SetD0, performing the diffusion generation process 3 times to obtain N1 diffusion-generated multimodal samples D'. n , where N1 = 3 × N0.
[0104] Multimodal samples D' are generated using N1 diffusions. n Composition diffusion generates a multimodal dataset SetD1={D' n},D' n In the example, n = 1…N1, and the subscript 1 of the dataset represents the dataset generated by diffusion.
[0105] like Figure 4 As shown, in this embodiment, the motor torque analysis process is as follows:
[0106] (C1) Input each diffusion to generate multimodal samples D' n Vehicle driving noise diffusion data Dr' n ={A' n ,B' n ,S' n ,M' n}, where n = 1...N1.
[0107] (C2) Calculate the brake effectiveness rule value Rule_brake n In this embodiment, the physical basis for calculating the braking effectiveness rule value is as follows: the mechanical action of the brake pedal is to generate resistance through the braking system to decelerate the vehicle. When the brake pedal opening exceeds the activation threshold, i.e., the driver has a clear braking intention, the vehicle speed should show a decreasing trend or at least not increase significantly. The calculation process is as follows:
[0108] Extract each diffusion to generate a multimodal sample D' n Vehicle driving noise diffusion data Dr' n Brake pedal opening data {b' t n}∈SetB1 and vehicle speed data {s' t n}∈SetS1, t=1...T, n is the sample index. Where, b' t n Let s' be the brake pedal opening value of the nth sample generated by diffusion at time t, and SetB1 be the set of brake pedal opening signals of all diffused samples. t nSetS1 represents the vehicle speed value of the nth sample generated by diffusion at time point t, and SetS1 is the set of vehicle speed signals of all samples generated by diffusion.
[0109] Define effective braking trigger conditions: Set the braking activation threshold τ_B = 5%, when b' t n When ≥τ_B, it is denoted as condition A: effective braking intention activated.
[0110] Define the condition of abnormal vehicle speed without deceleration: Calculate the rate of change of vehicle speed Δs' t n =(s' t n -s' t-1 n ) / Δt, where Δt=0.1s, and the unit is km / h / s, when Δs' t n When +0.5>0, it is recorded as condition B: the vehicle speed did not decrease significantly.
[0111] Taking all the above conditions into account, the corresponding vehicle driving noise diffusion data Dr' at that moment will only be recorded if both conditions A and B are met simultaneously. n Brake validity rule value Rule_brake n Each diffusion generates a multimodal sample D' n Dr's data on vehicle driving noise diffusion n The brake effectiveness rule value Rule_brake n The calculation formula is:
[0112] Rule_brake n =∑ t=1 T (t:b' t n ≥τ_B)max(0,Δs' t n +0.5)
[0113] (C3) Calculate the rule value of dual-pedal operation conflict_conflict n In this embodiment, the physical basis for calculating the dual-pedal operation conflict rule value is as follows: During normal driving, the operating intentions of the accelerator pedal and the brake pedal are mutually exclusive. Simultaneously, pressing them sharply will cause energy conflict between the power system and the braking system, resulting in motor overload and decreased braking efficiency. The calculation process is as follows:
[0114] Extract each diffusion to generate a multimodal sample D' n Dr's data on vehicle driving noise diffusion n Accelerator pedal opening data {a' t n}∈SetA1 and brake pedal opening data{b' t n}∈SetB1, t=1...T. Where, a' t n SetA1 represents the accelerator pedal opening value of the nth sample generated by diffusion at time point t, and SetA1 is the set of accelerator pedal opening signals of all samples generated by diffusion.
[0115] Set the accelerator pedal activation threshold: Accelerator pedal τ_A = 10%, exceeding this threshold is considered a valid acceleration operation.
[0116] Define an acceleration operation intensity quantization function: use the Sigmoid function σ(x) = 1 / (1+exp(-5x)) to quantify the operation intensity when the accelerator pedal exceeds a threshold. Since the degree to which the accelerator pedal exceeds the threshold varies significantly among different samples, this function maps the difference to the 0-1 interval; the larger the input, the closer the output is to 1, thus quantifying the operation intensity.
[0117] Calculate accelerator pedal operation intensity: Calculate the difference a' between the accelerator pedal opening and the threshold. t n Substituting -τ_A into x in the Sigmoid function yields the acceleration intensity σ_accel. t n =σ(a' t n -τ_A).
[0118] Set the brake pedal activation threshold: brake pedal τ_B = 5%, exceeding this threshold is considered a valid braking operation.
[0119] Define a braking intensity quantification function: Using the same Sigmoid function as the acceleration function, σ(x) = 1 / (1 + exp(-5x)), this function quantifies the intensity of braking pedal operation exceeding a threshold. Since the degree of brake pedal opening exceeding the threshold varies, this function can quantify the intensity within the 0-1 range.
[0120] Calculate brake pedal operating intensity: Calculate the difference b' between the brake pedal opening and the threshold value. t n Substituting -τ_B into x in the Sigmoid function yields the braking intensity σ_brake. t n =σ(b') t n -τ_B).
[0121] The rule value is calculated by multiplying the acceleration intensity and braking intensity of each frame within the time window and summing them up to obtain the dual-pedal operation conflict rule value, Rule_conflict. n As shown in the following formula:
[0122] Rule_conflict n =∑ t=1 T (σ_accel t n ·σ_brake t n ).
[0123] (C4) Calculate the torque dynamics rule value Rule_torque n In this embodiment, the physical basis for calculating the torque dynamics rule value is as follows: based on Newton's second law and vehicle force analysis, vehicle acceleration is determined by both the motor driving force and resistance. Specifically, acceleration has a definite mathematical relationship with motor torque, vehicle speed, and brake pedal opening; deviations exceeding a reasonable range violate the laws of dynamics. The calculation process is as follows:
[0124] Based on Newton's second law and the force analysis of the vehicle, the acceleration prediction model is established as shown in the following equation:
[0125]
[0126] The physical meanings of each item are as follows:
[0127] To predict acceleration.
[0128] The first term K1·m t This represents the acceleration corresponding to the driving force generated by the motor's torque. (m) t is the motor torque, N·m; K1 is the torque-acceleration conversion coefficient.
[0129] The second term K2·v t 2 This indicates the deceleration caused by air resistance and rolling resistance. t K is the vehicle speed, m / s, and K2 is the drag coefficient;
[0130] The third item, K3, represents the deceleration caused by the vehicle's basic resistance. K3 is the basic resistance coefficient.
[0131] Item 4 K4·b t This indicates the deceleration corresponding to the braking force generated by the brake pedal opening. t K is the brake pedal opening, and K4 is the brake-deceleration conversion coefficient.
[0132] ∈ represents the modeling error, |∈|≤0.3m / s 2 It covers unmodeled factors such as slope and road surface friction;
[0133] v tVehicle speed, in m / s, converted from km / h: v t =s t ×1000 / 3600.
[0134] Actual acceleration a t ^GT is calculated from the vehicle speed difference: a t ^GT=(s t -s t-1 ) / (Δt×3.6), where Δt=0.1s, and the unit conversion factor 3.6 is used to convert km / h to m / s.
[0135] Next, the coefficients K1 to K4 in the acceleration prediction model are calibrated, as follows:
[0136] Input the vehicle driving data from multiple observation time windows collected in step 1 (Dr) n As samples, each sample contains four signals (A, B, S, M), and each signal is 20 frames long. Least squares fitting is used to predict the acceleration (K1·m). t -K2·v t 2 -K3-K4·b t ) and actual acceleration a t The deviation of ^GT is minimized, yielding the values of K1 to K4. The acceleration error under torque dynamics is recorded as |a t ^GT-(K1·m t -K2·v t 2 -K3-K4·b t )|, where GT is an abbreviation for "Ground Truth", representing the actual value calculated based on the original observation data.
[0137] Finally, calculate the multimodal sample D' generated by each diffusion. n Dr's data on vehicle driving noise diffusion n Torque dynamics rule value Rule_torque n As shown in the following formula:
[0138] Rule_torque n = (1 / T)×∑ t=1 T |a 1t n ^GT-(K1·m 1t n -K2·v 1t n2 -K3-K4·b 1t n )|,
[0139] Where T = 20, a 1t n ^GT represents the actual acceleration of the nth sample generated by diffusion at time point t; m 1t n The motor torque value at time t for the nth sample generated by diffusion; v 1t n2 b is the square of the vehicle speed at time point t for the nth sample generated by diffusion; 1t n The brake pedal opening value at time t for the nth sample generated by diffusion.
[0140] (C5) Based on steps (C2), (C3), and (C4), each diffusion generates a multimodal sample D'. n Dr's data on vehicle driving noise diffusion n The brake effectiveness rule value Rule_brake n Dual-pedal operation conflict rule value: Rule_conflict n Torque dynamics rule value Rule_torque n The diffusion generates multimodal samples D' n Conduct a comprehensive physical feasibility verification.
[0141] Specifically, a threshold is set for each individual rule, and each is compared with its corresponding calculated rule value to conduct a comprehensive verification of physical feasibility. The braking effectiveness rule threshold is set at 5 km / h; this threshold is used to accumulate unreasonable acceleration. The dual-pedal operation conflict rule threshold is set at 10; this threshold is used to accumulate high conflict duration and is a dimensionless relative indicator. The torque dynamics rule threshold is set at 3 m / s². 2 This threshold is used to accumulate acceleration error.
[0142] Through comprehensive verification of physical feasibility, diffusion generates multimodal samples D' in accordance with physical laws. n The judgment conditions are as follows:
[0143] The condition for brake effectiveness is: Rule_brake n ≤5km / h.
[0144] The condition for a dual-pedal conflict is: Rule_conflict n ≤10.
[0145] The torque dynamics condition is: Rule_torque n ≤3m / s 2 .
[0146] For each diffusion, a multimodal sample D' is generated. nFor n = 1...N1, if all three judgment conditions are met, then the vehicle driving noise diffusion data Dr' in the diffusion-generated multimodal sample is judged. n In accordance with physical laws, this diffusion generates multimodal samples D' n Effective samples that conform to physical laws, i.e., effective multimodal samples D” n And set a valid tag flag. n =1; otherwise, the multimodal samples generated by the diffusion are judged as invalid samples that do not conform to physical laws, marked as invalid, and a valid label flag is set. n =0.
[0147] In this embodiment, based on the motor torque analysis results, multimodal samples D' are finally generated from all diffusions. n In the middle, filter out all flags n The valid samples with a value of 1 are the valid multimodal samples D. n To construct the multimodal dataset SetD”1={D” after motor torque analysis. n}={Dr” n ,V” n ,Y gt n}. Among them, Dr” n ={Dr' n |flag n =1}, Dr” n In the n=1...N'1,N'1=∑ n flag n .
[0148] Step 4: Using the original abnormal driving model Model_Dr0 obtained in Step 2, analyze each valid multimodal sample D in the multimodal dataset SetD”1 obtained after motor torque analysis in Step 3. n Make predictions to obtain D for each effective multimodal sample. n The predicted probability vector p of abnormal facial expression labels for each frame of a human face video. n Calculate D for each valid multimodal sample. n The probability vector y of the same type of expression in all frames of a human face video ti n The probability sum is calculated, and the category corresponding to the largest probability sum is taken as the predicted label of a valid multimodal sample. Then, based on each valid multimodal sample D” n Predicted labels The tag Y for abnormal facial expressions on real faces gt n Perform a prediction consistency assessment, and select the valid multimodal samples D that meet the requirements based on the assessment results.n As a multimodal sample D with predictive consistency n And a prediction-consistent multimodal dataset SetD”'1 is constructed using each prediction-consistent multimodal sample. Figure 5 As shown, the specific process is as follows:
[0149] (4.1) Input the multimodal dataset SetD”1={Dr” obtained after the motor torque analysis in step 3. n ,V” n ,Y gt n}, where Dr” n ={A” n ,B” n ,S” n ,M” n The meanings of each parameter are as follows:
[0150] A” n B” n 、S” n M” n The effective multimodal samples D” are respectively n The data includes accelerator pedal opening, brake pedal opening, vehicle speed, and motor torque.
[0151] V” n For effective multimodal samples D” n Facial video data in the middle.
[0152] Y gt n For effective multimodal samples D” n The real facial abnormal expression labels in the data, and the facial abnormal expression labels Y' after noise diffusion in step 3. n Consistency, i.e., Y' n =Y gt n Y gt n ∈C, C={c1,c2,c3,c4}.
[0153] (4.2) Extract each valid multimodal sample D from the multimodal dataset SetD”1 after motor torque analysis. n Accelerator pedal opening data A” n Brake pedal opening data B” n Vehicle speed opening data S” n Motor torque data (M) n Face video data V” n Input the original abnormal driving model Model_Dr0 obtained in step 2.
[0154] Each valid multimodal sample D” in the multimodal dataset SetD”1 after analyzing the motor torque using the original abnormal driving model Model_Dr0. n Make predictions to obtain D for each effective multimodal sample. n The predicted probability vector p of abnormal facial expression labels for each frame of a human face video. n ={p n1 ,p n2 ,p n3 ,p n4}, c1 = fatigue, c2 = anger, c3 = tension, c4 = neutral, ∑p ni =1.
[0155] Calculate D” for each valid multimodal sample n The probability vector y of the same type of expression in all frames of a human face video ti n The probability sum is calculated, and the category corresponding to the largest probability sum is taken as the predicted label of a valid multimodal sample.
[0156] Therefore, all valid multimodal samples D” in the multimodal dataset SetD”1 after motor torque analysis are obtained. n Predicted labels
[0157] (4.3) Each valid multimodal sample D” n Predicted labels Corresponding real-life abnormal facial expression label Y gt n Compare them to determine the consistency of the predictions. Then determine the valid multimodal sample D” n Meets prediction consistency, and sets the prediction consistency flag to keep. n =1. When Then determine the valid multimodal sample D” n If the prediction is inconsistent, set the prediction consistency flag to keep. n =0.
[0158] (4.4) Based on the judgment result of step (4.3), retain the valid multimodal samples D” that meet the prediction consistency. n As a multimodal sample D with predictive consistency n .in:
[0159] Predictive consistency of multimodal samples D”' n The accelerator pedal signal set SetA”'1={A”' n}, A”' n To achieve consistent predictions, i.e., keepn =1 effective multimodal sample D” n Accelerator pedal opening data.
[0160] Predictive consistency of multimodal samples D”' n The brake pedal signal set SetB”'1={B”' n}, B”' 1n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n The data on brake pedal opening.
[0161] Predictive consistency of multimodal samples D”' n The vehicle speed signal set SetS”'1={S”' n}, S”' n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n Vehicle speed data.
[0162] Predictive consistency of multimodal samples D”' n The motor torque signal set SetM”'1={M”' n}, M”' n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n The motor torque data in the file.
[0163] Predictive consistency of multimodal samples D”' n The face video signal set SetV”'1={V”' in the middle n}, V”' n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n Facial video data in the middle.
[0164] Predictive consistency of multimodal samples D”' n The set of generated conditional tags in SetY gt ={Y gt n}, Y gt n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n The tags for abnormal facial expressions on real people.
[0165] Sample index n = 1...N”1, where N”1 = ∑ n keep n .
[0166] (4.4) Based on the prediction consistency multimodal sample D”' obtained in step (4.4) n Construct a prediction consistency multimodal dataset SetD”'1={D”' n}={Dr”' n ,V”' n ,Y gt n}, where Dr”' 1n ={A”' n ,B”' n ,S”' n ,M”' n}
[0167] Step 5: Using the prediction consistency multimodal dataset SetD”'1 obtained in Step 4, retrain the original abnormal driving model Model_Dr0 to enhance it. The retrained and enhanced original abnormal driving model Model_Dr0 will then be used as the abnormal driving recognition model Model_Dr.
[0168] In this embodiment, each prediction-consistent multimodal sample D”' from the prediction-consistent multimodal dataset SetD”'1 is extracted from the original abnormal driving model Model_Dr0 using a one-dimensional convolutional network. n Dr”' vehicle driving data n Extracting vehicle driving features F_car n The three-dimensional convolutional network is used to extract each prediction consistent multimodal sample D”' from the prediction consistent multimodal dataset SetD”'1. n Face video data V”' n Extracting facial features F_face n Then, for each prediction-consistent multimodal sample D”' n Vehicle driving characteristics F_car n Facial features F_face n After concatenation, the fused feature F_fuse is obtained by processing through the first fully connected layer. n Then, the fused features F_fuse are processed through a second fully connected layer. n Perform feature mapping to predict each predictive consistent multimodal sample D”' n The predicted probability vector p of abnormal facial expression labels n Finally, based on the consistent predictions of each multimodal sample D”' n The predicted probability vector p n and generate conditional label set SetY gt The real face abnormal expression tag Y gtn The loss function is calculated, and the parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer are updated by backpropagation based on the loss function calculation results. The original abnormal driving model Model_Dr0 is then trained and enhanced to obtain the abnormal driving recognition model Model_Dr.
[0169] Specific examples Figure 6 As shown, the process of retraining the original abnormal driving model Model_Dr0 is as follows:
[0170] (D1) Obtain the prediction consistency multimodal dataset SetD”'1={D”' n}={Dr”' n ,V”' n ,Y gt n}, Dr”' n ={A”' n ,B”' n ,S”' n ,M”' n}, n=1...N”1.
[0171] Among them, A”' n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n Accelerator pedal opening data in the data. (B”') 1n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n Brake pedal opening data in the image. (S”') n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n Vehicle speed data in the image. (M) n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n Motor torque data in the image. (V”') n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n Facial video data from Y. gt n To achieve consistent predictions, i.e., keep n =1 effective multimodal sample D” n The set of generated conditional tags SetY gt The tags for abnormal facial expressions on real people.
[0172] (D2) Take each prediction-consistent multimodal sample D”' from the prediction-consistent multimodal dataset SetD”'1. n Dr”' vehicle driving data n The input is fed into the original abnormal driving model Model_Dr0's one-dimensional convolutional network. The network undergoes a first convolutional layer, a second pooling layer, a third convolutional layer, and a fourth layer that pulls the features into one dimension before fully connected processing. Finally, the fourth layer outputs the vehicle driving features F_car. n Vehicle driving characteristics F_car n The feature dimension is 1×128.
[0173] (D3) Take each prediction-consistent multimodal sample D”' from the prediction-consistent multimodal dataset SetD”'1. n Face video data V”' n The input is fed into the original abnormal driving model Model_Dr0's 3D convolutional network. The network undergoes processing through the following convolutional layers: first, second, third, fourth, fifth, sixth, seventh (global spatiotemporal average pooling), and eighth (connection layer). Finally, the eighth layer outputs the facial features F_face. n Facial features F_face n The feature dimension is 1×128.
[0174] (D4) For each prediction-consistent multimodal sample D”' n Vehicle driving features F_car with a feature dimension of 1×128 n Face features F_face with a feature dimension of 1×128 n Perform feature concatenation to obtain the concatenated feature F_concat n =Concat(F_car n ;F_face n ), where Concat() represents concatenation processing, and the concatenation feature is F_concat n The feature dimension is 1×256. Then, the first fully connected layer is used to process and concatenate the features F_concat. n The fusion feature F_fuse is obtained. n , F_fuse n The feature dimension is 1×128.
[0175] (D5) Each prediction-consistent multimodal sample D”' is processed through the second fully connected layer. n The fusion feature F_fuse n Perform feature mapping to predict each predictive consistent multimodal sample D”' nThe predicted probability vector p of abnormal facial expression labels n ={p n1 ,p n2 ,p n3 ,p n4}, p n1 ,p n2 ,p n3 ,p n4 For c1 to c4, ∑p ni =1.
[0176] (D6) Input the true abnormal face expression labels Y from the predictive consistency face abnormal expression label set SetY”'1. gt n Calculate the Y-label of abnormal facial expressions in real people gt n one-hot vector Vet n ={vet n,i}, when the real human face shows an abnormal expression label Y gt n The corresponding index value is Y gt n =Ci, the abnormal facial expression label Y of the real human face gt n one-hot vector vet n,i =1, otherwise it is a one-hot vector vet n,i =0.
[0177] Based on the predictive consistency of the abnormal facial expression label set SetY”'1, the actual abnormal facial expression label Y is obtained. gt n The one-hot vector and each prediction consistent multimodal sample D”' n The predicted probability vector p of abnormal facial expression labels n The cross-loss function Loss is calculated as follows:
[0178] Loss=-∑n∑ i=1 4 vet n,i ·log(p n,i )
[0179] Based on the cross-loss function calculation results, backpropagation is used to update the parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer. This process enhances the original abnormal driving model Model_Dr0 through retraining, resulting in the abnormal driving recognition model Model_Dr.
[0180] Step 6: Use the abnormal driving recognition model "Model_Dr" obtained in Step 5 to predict abnormal driving, and perform hierarchical control based on the abnormal driving prediction results.
[0181] like Figure 7 As shown, in this embodiment, the current vehicle signal Dr_test={A_test,B_test,S_test,M_test} is collected in real time, where:
[0182] A_test represents the set of accelerator pedal opening signals acquired in real time, A_test = {a t}
[0183] B_test represents the set of brake pedal opening signals acquired in real time, B_test={b t}
[0184] S_test represents the set of real-time collected vehicle speed signals, S_test={s t}
[0185] M_test represents the set of motor torque signals acquired in real time, M_test={m t}
[0186] t = 1...T_test, the real-time sampling rate is 10Hz, and T_test is the number of frames in the test time window, which is 20 frames, corresponding to 2 seconds, consistent with the training window.
[0187] In this embodiment, the current face video signal V_test = {v} is acquired in real time. t}, t=1...T_test, where v t It is a real-time video frame with a sampling rate of 10Hz, including RGB three channels, and a resolution of 224×224.
[0188] Then, the current vehicle signal Dr_test and the current face video signal V_test are input into the abnormal driving recognition model Model_Dr, which then predicts and outputs the abnormal driving prediction result.
[0189] Finally, based on the abnormal driving prediction results, the risk level is calculated, and graded control is implemented based on the risk level, as follows:
[0190] (E1) Calculate the risk level.
[0191] If predicting labels It belongs to the category of abnormal expressions, that is Where c1 = fatigue, c2 = anger, and c3 = tension, then Risk = 1.0.
[0192] If predicting labels A neutral expression, that is Then Risk = 0.8 × max(p1,p2,p3). Where p1, p2, and p3 are the probability values of the three abnormal expressions of fatigue (c1), anger (c2), and tension (c3) in the prediction results, respectively, and max(p1,p2,p3) represents the maximum probability value among the three.
[0193] (E2) Execute hierarchical control instructions:
[0194] When Risk < 0.4, execute voice prompt control commands.
[0195] When 0.4 ≤ Risk < 0.7, execute the automatic speed reduction + play soothing music control command.
[0196] When Risk ≥ 0.7, execute the emergency braking + notify preset contact person control command.
[0197] In this embodiment, the final output is the abnormal driving prediction result. Risk value and the hierarchical control instructions executed.
[0198] The preferred embodiments of the present invention have been described in detail above with reference to the accompanying drawings. These embodiments are merely descriptions of preferred embodiments and are not intended to limit the scope or concept of the invention. The specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. Such combinations, as long as they do not violate the spirit of the present invention, should also be considered as part of this disclosure. To avoid unnecessary repetition, the present invention will not further describe the various possible combinations.
[0199] This invention is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this invention and without departing from the design idea of this invention, all modifications and improvements made by those skilled in the art to the technical solutions of this invention should fall within the protection scope of this invention. The technical content for which protection is sought in this invention has been fully described in the claims.
Claims
1. An abnormal driving recognition method based on pedal data and facial expression multimodal analysis, characterized in that, The process is as follows: Step 1: Collect vehicle driving data from multiple identical time observation windows. n Face video data V n Dr, the vehicle driving data in each observation time window n All include raw accelerator pedal opening data A n Original brake pedal opening data B n Original vehicle speed data S n Original motor torque data M n For each person's facial video data V n Label the abnormal facial expressions of real people with the tag Y. gt n Dr. n Face video data V n And the corresponding real-life abnormal facial expression label Y gt n The original multimodal samples D are respectively constructed. n And from the original multimodal samples D of each time observation window n The original multimodal dataset SetD0 is composed of these components. Step 2: Based on the original multimodal dataset SetD0 obtained in Step 1, train an original abnormal driving model Model_Dr0; Step 3: Maintain the original multimodal sample D n The real face abnormal expression tag Y gt n The original multimodal sample D remains unchanged. n Dr. n Face video data V n Noise diffusion was performed separately, and the corresponding vehicle driving noise diffusion data Dr' was obtained. n Face video noise diffusion data V' n ; Dr's driving noise diffusion data for each vehicle n Face video noise diffusion data V' n And the corresponding real-life abnormal facial expression label Y gt n Each is used as a diffusion to generate multimodal samples D' n ; Then, multimodal samples D' are generated based on each diffusion. n Vehicle driving noise diffusion data Dr' n Motor torque analysis is performed, and multimodal samples D' are generated from various diffusions based on the motor torque analysis results. n In the process, valid multimodal samples D'' that meet the requirements are selected. n To construct the multimodal dataset SetD''1 after motor torque analysis; Step 4: Using the original abnormal driving model Model_Dr0 obtained in Step 2, analyze each valid multimodal sample D'' in the multimodal dataset SetD''1 obtained after motor torque analysis in Step 3. n Make predictions to obtain each effective multimodal sample D'' n Predicted labels Based on each valid multimodal sample D'' n Predicted labels And the corresponding real-life abnormal facial expression label Y gt n Perform a prediction consistency assessment, and select the valid multimodal samples D'' that meet the requirements based on the assessment results. n As a predictive consistency multimodal sample D''' n And through each prediction consistent multimodal sample D''' n Construct a prediction consistency multimodal dataset SetD'''1; Step 5: Using the prediction consistency multimodal dataset SetD'''1 obtained in Step 4, retrain the original abnormal driving model Model_Dr0 to enhance the original abnormal driving model Model_Dr0. The retrained and enhanced original abnormal driving model Model_Dr0 is used as the abnormal driving recognition model Model_Dr''. Step 6: Input the current vehicle signal and the current face video signal into the abnormal driving recognition model Model_Dr'' obtained in Step 5, and use the abnormal driving recognition model Model_Dr'' to predict the abnormal driving result. ; In step 3, the motor torque analysis process is as follows: Based on the braking effectiveness rules, define the effective braking trigger condition and the abnormal vehicle speed without deceleration condition, and determine the multimodal sample D' generated by each diffusion. n Vehicle driving noise diffusion data Dr' n Does the effective braking trigger condition and the abnormal vehicle speed without deceleration condition both meet simultaneously? If both conditions are met, then based on the vehicle driving noise diffusion data Dr' n Using the brake pedal opening data, vehicle speed data, and defined conditions, calculate each diffusion-generated multimodal sample D'. n Dr's data on vehicle driving noise diffusion n The brake effectiveness rule value Rule_brake n ; Based on the physical basis of the dual-pedal operation conflict rule value, accelerator pedal activation threshold and brake pedal activation threshold are set, and acceleration operation intensity quantization function and braking operation intensity quantization function are defined; each diffusion generates a multimodal sample D'. n Vehicle driving noise diffusion data Dr' n The difference between the accelerator pedal opening data and the set accelerator pedal activation threshold is substituted into the acceleration operation intensity quantization function to calculate the accelerator pedal operation intensity; each diffusion generates a multimodal sample D'. n Vehicle driving noise diffusion data Dr' n The difference between the brake pedal opening data and the set brake pedal activation threshold is substituted into the brake operation intensity quantization function to calculate the brake pedal operation intensity. Finally, based on the accelerator pedal operation intensity and the brake pedal operation intensity, the multimodal sample D' generated by diffusion is calculated. n Dr's data on vehicle driving noise diffusion n The dual-pedal operation conflict rule value is Rule_conflict. n ; Based on the physical basis of torque dynamics rules, an acceleration prediction model is established, and the least squares method is used to predict vehicle driving data from multiple observation time windows collected in step 1. n The acceleration prediction model is fitted to minimize the deviation between the predicted and actual accelerations, thereby obtaining the coefficients and corresponding acceleration errors in the model. Then, based on these coefficients and the acceleration errors, the multimodal sample D' generated by diffusion is calculated. n Dr's data on vehicle driving noise diffusion n Torque dynamics rule value Rule_torque n ; Each diffusion generates a multimodal sample D'. n Dr's data on vehicle driving noise diffusion n The brake effectiveness rule value Rule_brake n Dual-pedal operation conflict rule value: Rule_conflict n Torque dynamics rule value Rule_torque n Each rule is compared with its corresponding threshold value, and the rule value of brake effectiveness is determined by the value of brake_brake. n Dual-pedal operation conflict rule value: Rule_conflict n Torque dynamics rule value Rule_torque n When all values are less than or equal to their respective thresholds, then the corresponding diffusion generates multimodal samples D'. n Vehicle driving noise diffusion data Dr' n It conforms to physical laws, and determines that the corresponding diffusion generates multimodal samples D'. n For effective multimodal samples D'' n .
2. The abnormal driving recognition method based on pedal data and facial expression multimodal analysis as described in claim 1, characterized in that, In step 1, multiple categories of abnormal facial expressions during driving are defined; a pre-trained ResNet model is used to predict the V value of each face video. n The probability vector y of various abnormal facial expression labels for each frame of the expression. ti n Each face video V is output by a pre-trained ResNet model. n The probability vector y of similar expressions across all frames ti n The probability sum is calculated, and the category corresponding to the largest probability sum is taken as the real facial abnormal expression label Y of the corresponding face video V. gt n .
3. The abnormal driving recognition method based on pedal data and facial expression multimodal analysis according to claim 1, characterized in that, In step 2, the original abnormal driving model Model_Dr0 includes a one-dimensional convolutional network, a three-dimensional convolutional network, a first fully connected layer, and a second fully connected layer; During training, each original multimodal sample D from the original multimodal dataset SetD0 is processed through a one-dimensional convolutional network. n Dr. n Extracting vehicle driving features F_car n The original multimodal samples D from the original multimodal dataset SetD0 are extracted using a 3D convolutional network. n Face video data V n Extracting facial features F_face n ; Each original multimodal sample D n Vehicle driving characteristics F_car n Facial features F_face n After concatenation, the fused feature F_fuse is obtained by processing through the first fully connected layer. n Then, the fused features F_fuse are processed through a second fully connected layer. n Perform feature mapping to predict each original multimodal sample D. n The predicted probability vector p of abnormal facial expression labels n Finally, based on each original multimodal sample D n The predicted probability vector p n The real face abnormal expression label Y in the real face abnormal expression label set SetY0 gt n The loss function is calculated, and the parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer are updated by backpropagation based on the loss function calculation results, thereby training the original abnormal driving model Model_Dr0.
4. The abnormal driving recognition method based on pedal data and facial expression multimodal analysis as described in claim 1, characterized in that, In step 3, the statistics of each original multimodal sample D are performed. n Dr. n In the original accelerator pedal opening, the maximum value maxA is... n The noise range coefficient δ_A is calculated, and the accelerator pedal opening noise ε_A is generated. Then, based on each original multimodal sample D, the noise range coefficient δ_A is calculated. n The original accelerator pedal opening data A n Generate accelerator pedal opening noise diffusion data A' n =A n +ε_A; Statistical analysis of each original multimodal sample D n Dr. n In the original brake pedal opening, the maximum value maxB is... n The noise range coefficient δ_B is calculated, and the brake pedal opening noise ε_B is generated. Then, based on each original multimodal sample D, the noise range coefficient δ_B is calculated. n The original brake pedal opening data B n Generate brake pedal opening noise diffusion data B' n = B n +ε_B; Based on the reasonable fluctuation range of vehicle speed in actual driving, a fixed tolerance δ_S for vehicle speed noise is set, and vehicle speed noise ε_S is generated. Then, based on each original multimodal sample D n Dr. n The original vehicle speed data S is used to generate vehicle speed noise diffusion data S'. n = S n +ε_S; Based on the reasonable fluctuation range of the motor output torque, a fixed tolerance δ_M for the motor torque noise is set, and the motor torque noise ε_M is generated. Then, based on each original multimodal sample D n Dr. n The original motor torque data M n Generate motor torque noise diffusion data M' n = M n +ε_M; Therefore, based on each original multimodal sample D n Dr. n The corresponding vehicle driving noise diffusion data Dr' were constructed and generated respectively. n = {A' n , B' n , S' n , M' n }; Define the video noise variance σ_V and generate face video noise ε_V, then based on each original multimodal sample D n V facial video data n Generate face video noise diffusion data V' n = V n +ε_V; Preserve realistic facial abnormal expression tags Y gt n Unchanged, serving as the label Y' for abnormal facial expressions after noise diffusion. n , i.e. Y' n = Y gt n ; Finally, Dr' integrates vehicle driving noise diffusion data. n Face video noise diffusion data V' n And the corresponding facial abnormality label Y' after noise diffusion n Generate diffusion to generate multimodal samples D' n = {Dr' n , V' n , Y' n } 5. The abnormal driving recognition method based on pedal data and facial expression multimodal analysis according to claim 1, characterized in that, In step 4, when a certain valid multimodal sample D'' n Predicted labels The corresponding real-life abnormal facial expression label Y gt n If they are the same, then the valid multimodal sample D'' is determined. n If the prediction is consistent, then the valid multimodal sample D'' is determined to be valid. n This does not conform to the prediction.
6. The abnormal driving recognition method based on pedal data and facial expression multimodal analysis according to claim 1, characterized in that, In steps 2 and 5, the cross-loss function is used during training. The parameters of the one-dimensional convolutional network, the three-dimensional convolutional network, the first fully connected layer, and the second fully connected layer are updated by backpropagation based on the calculation results of the cross-loss function.
7. The abnormal driving recognition method based on pedal data and facial expression multimodal analysis according to claim 1, characterized in that, In step 6, based on the abnormal driving prediction results obtained from the prediction... Calculate the risks and implement corresponding graded controls based on the obtained risk levels.
Citation Information
Patent Citations
Digital human emotion recognition and feedback system based on large model
CN119577557A
Fatigue driving detection method and fatigue driving detection system based on multi-feature fusion
CN120318802A