A Driving Behavior Analysis Method and System Based on Emotion Perception

By improving the emotion perception model, introducing micro-expression feature branches and spatiotemporal attention branches, and combining driver expression auxiliary labels and delay compensation mechanisms, the problem of emotion recognition for drivers with limited facial expressions was solved, achieving high-precision and real-time driving behavior analysis.

CN120877258BActive Publication Date: 2026-01-06JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511389711.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-01-06
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing technologies cannot effectively address the problem of drivers' limited facial expressions and unclear emotional expressions, resulting in insufficient accuracy and real-time performance in emotion recognition, and ignoring individual differences and the time delay of emotional responses.

Method used

By introducing micro-expression feature branches, spatiotemporal attention branches, and perception weight adjustment branches, the emotion perception model is improved. Combined with driver expression auxiliary labels and driving delay compensation mechanisms, the perception weights are adjusted in real time, driving behavior data is processed synchronously, and an emotion-behavior correlation model is established to achieve driving behavior analysis.

Benefits of technology

It improves the accuracy and reliability of emotion recognition, solves the problem of lag between emotional fluctuations and driving behavior, ensures the matching of emotions and behaviors, and enhances the timeliness and effectiveness of intervention measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877258B_ABST
    Figure CN120877258B_ABST
Patent Text Reader

Abstract

The application discloses a driving behavior analysis method and system based on emotion perception, and relates to the technical field of driver state perception.The driving behavior analysis system based on emotion perception comprises an emotion perception correction module and a driving behavior analysis module.The application introduces a micro-expression feature branch, a space-time attention branch and a perception weight adjustment branch, accurately captures small changes and time sequence features in the facial expression of a driver, and efficiently identifies emotional fluctuations of the driver even in the case that the facial expression is not rich or the emotional expression is weak, so that the emotion perception model can adapt to individualized emotional expression modes of different drivers, and the accuracy and reliability of emotion identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of driver state perception technology, and in particular to a driving behavior analysis method and system based on emotion perception. Background Technology

[0002] While significant progress has been made in current emotion perception and driving behavior analysis technologies, many challenges remain, especially when faced with differences in emotional expression among different drivers and the lag between emotion and behavior. Traditional methods often fail to accurately identify the driver's emotional state and intervene effectively in a timely manner. Most existing technologies rely on a single emotion perception model, typically using facial expression data or voice data for emotion recognition. However, these methods ignore individual differences in facial expression and the time delay in emotional responses, resulting in insufficient accuracy and real-time performance in emotion recognition. Furthermore, they ignore the varying intensity of emotional responses from different individuals, which can easily lead to misjudgments of emotions.

[0003] Therefore, it is necessary to consider personalized emotion perception for different drivers. By improving the emotion perception model, the problem that existing technologies cannot effectively deal with drivers' limited facial expressions and unclear emotional expressions has been solved. Summary of the Invention

[0004] This invention aims to provide a driving behavior analysis method and system based on emotion perception. By improving the emotion perception model, it solves the problem that existing technologies cannot effectively deal with drivers' limited facial expressions and unclear emotional expressions.

[0005] A driving behavior analysis method based on emotion perception includes the following steps:

[0006] Prioritize capturing high-frequency facial expressions of the current driver to obtain driver facial expression auxiliary labels; collect data on the current driver in real time to obtain driving environment data, driving behavior data, and driving facial data;

[0007] Driver facial data is input into a driver emotion perception model for analysis to obtain driver emotion perception results. Specifically, the driver emotion perception model is constructed based on the ResNet-50 architecture, with the addition of micro-expression feature branches, spatiotemporal attention branches, and perception weight adjustment branches. Simultaneously, the hyperparameters of the perception weight adjustment branch are adjusted in real time using driver facial auxiliary labels. A driving delay compensation mechanism is established to synchronously correct the real-time collected driving behavior data and corresponding driving behavior data streams, resulting in new driving behavior data.

[0008] Based on the driver's emotional perception results and corresponding new driving behavior data and driving environment data, the current driver's driving behavior is analyzed to obtain emotion-driving behavior analysis results, and driving behavior intervention operations based on driving behavior and emotional perception are realized.

[0009] As a preferred technical solution of the present invention, the specific steps for prioritizing high-frequency facial expression capture of the current driver include:

[0010] The facial expression detection algorithm at an ultra-high frequency is used to capture frames of the current driver's facial expression and obtain several facial key point data.

[0011] Based on facial key point data, a deep learning model combined with a time window is used to calculate the change amplitude of facial key points between the current frame and the previous frame to obtain the single-frame change amplitude; all single-frame change amplitudes are combined to obtain a frame change amplitude sequence.

[0012] The spatiotemporal feature extraction method is used to extract features from the frame change amplitude sequence to obtain micro-expression change time series features; auxiliary label mapping is performed based on the micro-expression change time series features to obtain driver expression auxiliary labels.

[0013] As a preferred technical solution of the present invention, the driving emotion perception model includes an input layer, a convolutional layer, a pooling layer, four residual blocks, a global average pooling layer, and a fully connected layer.

[0014] Each residual block consists of a 1×1 dimensionality reduction convolution, a 3×3 spatial convolution, and a 1×1 dimensionality increase convolution. A micro-expression feature branch is added in parallel after the third residual block, a spatiotemporal attention branch is inserted before the fourth residual block, and a perceptual weight adjustment branch is added before the global average pooling layer.

[0015] The micro-expression feature branch is used to capture the subtle movements of facial muscles in driver facial data to obtain driver micro-expression features;

[0016] The spatiotemporal attention branch is used to output driver-weighted facial features supplemented by driver micro-expression features;

[0017] The perception weight adjustment branch is used to adjust the perception weights of the driver's weighted facial features to obtain the driver's emotion perception feature map.

[0018] Finally, the driver's emotion perception result is output after feature classification in the fully connected layer.

[0019] As a preferred embodiment of the present invention, the specific steps for adjusting the hyperparameters of the perception weight adjustment branch in real time using driver expression auxiliary tags include:

[0020] Set the default hyperparameters in the perception weight adjustment branch;

[0021] The system uses a genetic algorithm to search for hyperparameter combinations adjusted based on driver facial expression auxiliary labels. At the same time, a reinforcement learning mechanism is introduced to assist the genetic algorithm in screening hyperparameter combinations during the emotion perception process, so as to obtain the optimal hyperparameters.

[0022] The perceptual weight adjustment branch is adjusted in real time based on the optimal hyperparameters.

[0023] As a preferred embodiment of the present invention, the specific steps for establishing a driving delay compensation mechanism and synchronously correcting the real-time collected driving behavior data and the corresponding driving behavior data stream include:

[0024] Real-time collected driving behavior data and driver emotion perception results are processed in time synchronization to obtain initial behavior-emotion alignment data;

[0025] Establish an emotion-behavior correlation model; use the emotion-behavior correlation model to predict the driver's emotion perception results and driving environment data to obtain the predicted driving behavior lag time; and use the emotion-behavior correlation model to perform emotion inversion on driving behavior data to obtain the predicted driving behavior emotion time; combine the predicted driving behavior lag time and the predicted driving behavior emotion time to obtain behavior-emotion prediction time delay data.

[0026] The behavior-emotion initial alignment data and behavior-emotion prediction latency data are matched in the driving behavior data stream to obtain the matching driving behavior data as new driving behavior data.

[0027] As a preferred embodiment of the present invention, the specific steps for analyzing the current driver's driving behavior based on the driver's emotional perception results and corresponding new driving behavior data and driving environment data include:

[0028] The aligned driver emotion perception results, new driving behavior data, and driving environment data are fused to obtain a comprehensive evaluation feature vector. Driving behavior analysis is performed based on the comprehensive evaluation feature vector to obtain emotion-driving behavior analysis results. Graded warnings are issued based on the emotion-driving behavior analysis results, and different feedback behaviors are executed on the current driver in the vehicle according to different graded warnings.

[0029] A driving behavior analysis system based on emotion perception includes:

[0030] The emotion perception correction module includes a data acquisition unit and a differential emotion perception unit. The data acquisition unit prioritizes capturing high-frequency facial expressions of the current driver to obtain driver expression auxiliary labels. It also collects real-time data on the current driver, acquiring driving environment data, driving behavior data, and driver facial data. The differential emotion perception unit inputs the driver facial data into the driving emotion perception model for analysis to obtain the driver's emotion perception results. The driving emotion perception model is constructed based on the ResNet-50 architecture, adding micro-expression feature branches, spatiotemporal attention branches, and perception weight adjustment branches. Simultaneously, the hyperparameters of the perception weight adjustment branch are adjusted in real-time using the driver expression auxiliary labels. A driving delay compensation mechanism is established to synchronously correct the real-time collected driving behavior data and the corresponding driving behavior data stream, obtaining new driving behavior data.

[0031] The driving behavior analysis module includes a behavior analysis unit. The behavior analysis unit is used to analyze the current driver's driving behavior based on the driver's emotional perception results and corresponding new driving behavior data and driving environment data, to obtain emotion-driving behavior analysis results, and to realize driving behavior intervention operations based on driving behavior and emotional perception.

[0032] The present invention has the following advantages:

[0033] 1. This invention introduces micro-expression feature branches, spatiotemporal attention branches, and perceptual weight adjustment branches to accurately capture subtle changes and temporal features in the driver's facial expressions. Especially when facial expressions are not rich or emotional expression is weak, it can still efficiently identify the driver's emotional fluctuations, ensuring that the emotion perception model can adapt to the personalized emotional expression of different drivers, thus improving the accuracy and reliability of emotion recognition.

[0034] 2. This invention establishes a driving delay compensation mechanism to synchronize driver emotional perception results and driving behavior data in real time, solving the problem of lag between emotional fluctuations and driving behavior. It accurately predicts the time delay impact of driver emotional changes on driving behavior and corrects the driving behavior data stream to ensure the matching of emotions and behaviors, thereby improving the timeliness and effectiveness of intervention measures. By fusing features from driver emotional perception results, driving behavior data, and driving environment data, it can comprehensively assess the driver's emotional and behavioral state, ensuring that the close correlation between emotions and driving behavior is fully explored, thereby improving the accuracy of driver emotional perception and behavioral intervention. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the structure of a driving behavior analysis system based on emotion perception used in an embodiment of the present invention. Detailed Implementation

[0036] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this invention.

[0037] Example 1: A driving behavior analysis method based on emotion perception, comprising the following steps:

[0038] Prioritize capturing high-frequency facial expressions of the current driver to obtain driver facial expression auxiliary labels; collect data on the current driver in real time to obtain driving environment data, driving behavior data, and driving facial data;

[0039] The specific steps for prioritizing the capture of high-frequency facial expressions from the current driver include:

[0040] The facial expression detection algorithm at an ultra-high frequency is used to capture frames of the current driver's facial expression and obtain several facial key point data.

[0041] Based on facial key point data, a deep learning model combined with a time window is used to calculate the change amplitude of facial key points between the current frame and the previous frame to obtain the single-frame change amplitude; all single-frame change amplitudes are combined to obtain a frame change amplitude sequence.

[0042] Spatiotemporal feature extraction methods are used to extract features from the frame change amplitude sequence to obtain micro-expression change time series features; auxiliary label mapping is performed based on the micro-expression change time series features to obtain driver expression auxiliary labels;

[0043] In actual emotion detection, different drivers have different facial expressions and express emotions in different ways. If the same emotion perception model is used for analysis, inaccurate recognition or error in emotion judgment may occur. This is because there are significant individual differences in facial expression: some drivers have rich facial expressions, and their emotional fluctuations are easily reflected through their facial expressions, while others may have stiff facial expressions or weak expressions, making it difficult to detect emotional changes through conventional expression analysis methods. In addition, factors such as the driver's cultural background, personality traits, and personal habits will also affect the way they express their facial expressions, thereby increasing the complexity and challenge of the emotion perception model.

[0044] By utilizing an ultra-high frequency facial expression detection algorithm to capture frames of the driver's facial expressions, this step involves inputting real-time facial image data captured from a camera. Real-time facial expression capture is required, recording the driver's expressions under different emotion labels. The captured expressions may be obtained through active or passive methods. Active capture involves guiding the driver to autonomously make specific facial expressions to obtain corresponding emotional expressions. By having the driver actively express different facial expressions in known situations, such as anger, happiness, and surprise, more standardized expression data can be collected. Passive capture involves monitoring the driver's facial expressions during natural driving without any intervention, automatically identifying the driver's emotional state. Passive capture can capture the driver's emotional fluctuations through long-term expression tracking, such as changes in facial expressions during rapid driving, congested traffic, or complex environments, capturing emotions like anxiety, irritability, and stress.

[0045] Facial image data consists of continuous image frames. A facial expression detection algorithm, such as Mediapipe or Dlib, is used to efficiently and accurately detect facial key points at high frame rates. Facial key point data represents the location of facial features, including the coordinates of areas such as the eyes, eyebrows, and corners of the mouth.

[0046] Based on facial key point data, a deep learning model combined with a time window method is used to calculate the change amplitude of facial key points between the current frame and the previous frame. It should be noted that in this step, the input data is the facial key point data of the previous frame and the current frame, and the output is the change amplitude of a single frame. Specifically, the deep learning model generates an amplitude value representing the change of facial expression by calculating the difference between the key points of the current frame and the previous frame. This step can capture subtle facial movements, such as a slight upturn of the corners of the mouth or a slight rise of the eyebrows, which is especially important for drivers whose facial expressions are not very rich.

[0047] All single-frame variation amplitude data are combined into a frame variation amplitude sequence, which records the changes in facial expressions over time. Features are extracted from this frame variation amplitude sequence using spatiotemporal feature extraction methods to obtain micro-expression change time series features. In this step, the spatiotemporal feature extraction method not only captures the amplitude of facial expression changes but also understands the emotional fluctuations behind these changes. Specifically, spatial features are extracted using convolutional layers, and then temporal modeling methods (such as LSTM) are used to analyze the evolution of facial features over time, thus forming a comprehensive micro-expression change time series feature.

[0048] Driver facial expression auxiliary labels represent the degree to which different drivers express emotions in their facial expressions, which is used to improve the accuracy of subsequent emotion perception models. Since each driver's facial expression habits, emotional expression methods, and the richness of their expressions are different, facial expression auxiliary labels are used to quantify the individual differences of drivers, so that the subsequent emotion perception can adapt to the emotional expression characteristics of each driver, thereby improving recognition accuracy.

[0049] By using a trained emotion intensity recognition model, the extracted micro-expression change features are mapped onto the driver's emotion state labels. The auxiliary labels are inferred from the subtle changes in the driver's facial expressions to indicate the degree of their current emotion. These emotion state intensities are converted into auxiliary labels, such as fatigue, anxiety, pleasure, or tension, to indicate the degree of expression of different drivers under different emotions. This process enables accurate assessment of the driver's emotional state in subsequent emotion perception and behavior analysis, especially when the driver's facial expressions are not obvious, it can still efficiently predict the emotional state through the time series features of micro-expression changes.

[0050] The core objective of training an emotion intensity recognition model is to enable it to accurately assess the intensity of emotions in a driver's facial expressions, especially when facial expressions are limited or emotions are subtle. Specifically, the training set should include facial expression data from multiple drivers in various emotional states, sourced from emotion-labeled databases such as FER-2013 and AffectNet, which contain facial images labeled with emotions. To improve the model's generalization ability, the dataset diversity is further enhanced by collecting driver-provided emotion expression data or expression data from simulated driving scenarios. The training of the emotion intensity recognition model typically employs supervised learning methods, where the input is facial... The model takes facial expression feature data (such as facial key points, the magnitude of changes in local areas, etc.) as input, and outputs auxiliary labels for emotion intensity. To extract effective features from the data, convolutional neural networks (CNNs) or other deep learning methods can be used to extract spatial and temporal features from facial images. Mean squared error is used as the loss function for model training. During training, the model evaluates the loss and accuracy on the validation set at the end of each epoch to ensure that the model does not overly rely on the features of the training set. If the loss on the validation set no longer decreases significantly and reaches a certain threshold (e.g., the validation loss does not decrease for 10 consecutive times), training can be stopped, and a well-trained emotion intensity recognition model is obtained.

[0051] Driver facial data is input into a driver emotion perception model for analysis to obtain driver emotion perception results. Specifically, the driver emotion perception model is constructed based on the ResNet-50 architecture, with the addition of micro-expression feature branches, spatiotemporal attention branches, and perception weight adjustment branches. Simultaneously, the hyperparameters of the perception weight adjustment branch are adjusted in real time using driver facial auxiliary labels. A driving delay compensation mechanism is established to synchronously correct the real-time collected driving behavior data and corresponding driving behavior data streams, resulting in new driving behavior data.

[0052] The driving emotion perception model includes an input layer, a convolutional layer, a pooling layer, four residual blocks, a global average pooling layer, and a fully connected layer.

[0053] Each residual block consists of a 1×1 dimensionality reduction convolution, a 3×3 spatial convolution, and a 1×1 dimensionality increase convolution. A micro-expression feature branch is added in parallel after the third residual block, a spatiotemporal attention branch is inserted before the fourth residual block, and a perceptual weight adjustment branch is added before the global average pooling layer.

[0054] The micro-expression feature branch is used to capture the subtle movements of facial muscles in driver facial data to obtain driver micro-expression features;

[0055] The spatiotemporal attention branch is used to output driver-weighted facial features supplemented by driver micro-expression features;

[0056] The perception weight adjustment branch is used to adjust the perception weights of the driver's weighted facial features to obtain the driver's emotion perception feature map.

[0057] Finally, the driver's emotion perception result is output in the fully connected layer after feature classification.

[0058] It should be noted that driving environment data refers to various information that reflects the current driving environment and external conditions. This typically includes information about the train's track, meteorological data (such as temperature, humidity, and wind speed), the train's current speed, track gradient, road curvature, and other external factors that may affect train operation. Acquisition of driving environment data usually relies on onboard sensors and external devices. Lighting conditions can be acquired in real-time through onboard cameras and light sensors. Cameras can detect the intensity and direction of external light to determine whether it is daytime, dusk, or nighttime, which is particularly important for emotion perception because lighting affects the driver's facial expression recognition. Weather condition data, such as sunny, rainy, or foggy weather, is usually obtained through real-time weather API data from the internet, providing current meteorological information. Road condition types include urban roads, highways, mountain roads, and rural roads. Driver behavior and emotional reactions differ under different road conditions, and this data can be obtained through GPS positioning and road type information. Time of day also affects driver emotions and is usually obtained through the onboard system's clock. The main function of driving environment data is to provide external contextual information for emotion perception models, helping to more accurately analyze the driver's emotional state.

[0059] Driving behavior data primarily refers to the driver's operational behavior data during train operation, such as acceleration, braking, and steering. To acquire this data, modern high-speed trains are equipped with vehicle control systems that record every operation of the control levers by the driver, down to the second of acceleration and braking force. Data recorders installed on the train can collect real-time data such as speed, acceleration, braking force, and steering angle. The vehicle control system also provides dynamic status data of the train, such as real-time speed, braking status, and power output. All this information accurately reflects the driver's operational behavior. Furthermore, high-speed trains are equipped with driver behavior monitoring systems that record every action the driver takes during driving, including manual acceleration, braking, and other important operations. This data helps the system understand the driver's reaction speed, operating habits, and ability to cope with the external environment. For example, when a driver is anxious, they may exhibit behaviors such as excessive acceleration or frequent braking. This data can provide intuitive behavioral characteristics for emotion analysis systems, helping the system determine the driver's emotional state and take appropriate driving behavior interventions.

[0060] Specifically, the driving emotion perception model is an improvement on the ResNet-50 architecture. The input layer receives the driver's facial data, which is usually a pre-processed image containing the driver's facial features. After the input layer, the facial data is processed by convolutional layers. The role of the convolutional layers is to extract features from local regions by sliding convolutional kernels across the image. This layer extracts different facial features, such as the shape, position, and dynamic changes of areas like the eyes, eyebrows, and mouth, through multiple convolutional kernels. Through convolutional operations, the network can extract low-level features useful for emotion perception from the input data, such as texture and contours. These features provide higher-dimensional input information for subsequent layers. Pooling layers downsample the feature maps output by the convolutional layers, reducing their spatial dimension. Pooling operations typically select the maximum or average value within a region to reduce the size of the feature map, reduce computation, and retain the most important feature information.

[0061] Subsequently, the data is fed into multiple residual blocks (ResNet structure), each consisting of a 1×1 dimensionality-reducing convolution, a 3×3 spatial convolution, and a 1×1 dimensionality-increasing convolution. The role of each residual block is to further extract high-level features of facial expressions, while reducing the gradient vanishing problem during training through residual connections, enabling the deep network to be trained more effectively. The 1×1 convolution in each residual block first reduces the dimension of the feature map, then the 3×3 convolution extracts the spatial features of facial expressions, and finally the 1×1 dimensionality-increasing convolution restores the number of channels of the features.

[0062] After the third residual block, a micro-expression feature branch is added in parallel to the model. This branch is specifically designed to capture subtle movements of the driver's facial muscles, especially micro-expression changes lasting 0.3-0.5 seconds. This branch extracts micro-expression features through small-sized convolutional kernels and further enhances the expressive power of subtle expression changes through convolutional layers. The output of the micro-expression feature branch will provide more detailed and accurate features of the driver's emotional state. Especially for drivers with limited facial expressions, the micro-expression branch can improve the accuracy of emotion recognition.

[0063] Meanwhile, a spatiotemporal attention branch is inserted before the fourth residual block to enhance the model's ability to perceive temporal data and spatial features. This branch uses channel attention and inter-frame attention mechanisms to weight the temporal and spatial information of each facial feature. Through channel attention, the spatiotemporal branch strengthens facial feature channels related to emotion intensity and suppresses noise information unrelated to emotion. Inter-frame attention, on the other hand, assigns higher weights to features at specific time steps based on the time sequence of facial expression changes. This mechanism ensures that key features are extracted more effectively during the dynamic changes of the driver's facial expressions, thereby optimizing the emotion perception effect.

[0064] The perceptual weight adjustment branch is inserted before the global average pooling layer. The main function of the perceptual weight adjustment branch is to adjust the weights of the driver's facial features based on the real-time emotion perception results. Specifically, the perceptual weight adjustment branch dynamically adjusts the weights of facial features based on the driver's emotional fluctuations, facial expression auxiliary labels, and micro-expression feature outputs. This branch can improve the sensitivity to subtle emotional changes, especially when facial expressions are relatively flat or do not change significantly, and can identify and effectively classify the driver's emotional changes.

[0065] Finally, after processing by the global average pooling layer, all feature maps are compressed into a single vector. The global average pooling layer compresses all spatial information of each feature channel into a mean, reducing data dimensionality while preserving global information. The pooled feature vector is then processed by a fully connected layer for emotion classification. The fully connected layer is responsible for linearly combining the compressed features and outputting the probability distribution of emotion categories through an activation function (such as Softmax). In the final output, the driver's emotional state, such as pleasure, fatigue, or anxiety, can be accurately identified based on facial expressions, providing a basis for subsequent driving behavior intervention decisions.

[0066] The specific steps for adjusting the hyperparameters of the perception weight adjustment branch in real time using driver facial expression auxiliary tags include:

[0067] Set default hyperparameters for the perceptual weight adjustment branch. Hyperparameters include learning rate, regularization coefficient, feature weights, etc. Default hyperparameters are usually preliminary settings obtained from previous experimental experience or pre-trained models, used to provide a reasonable starting point for the perceptual weight adjustment branch in the initial stage. The purpose of setting these default hyperparameters is to provide basic parameters for model training and emotion recognition process, ensuring that the model can run normally and perform the initial emotion recognition task without further optimization.

[0068] The system uses a genetic algorithm to search for hyperparameter combinations adjusted based on driver facial expression auxiliary labels. At the same time, a reinforcement learning mechanism is introduced to assist the genetic algorithm in screening hyperparameter combinations during the emotion perception process, so as to obtain the optimal hyperparameters.

[0069] Based on driver facial expression auxiliary labels, a genetic algorithm is used to search, automatically adjust, and optimize hyperparameter combinations. The goal of the genetic algorithm is to generate different hyperparameter combinations based on the emotional changes of different drivers and facial expression auxiliary labels, and to select the best combination based on its performance. Specifically, the genetic algorithm evaluates the effectiveness of each hyperparameter combination according to a preset fitness function, which typically measures the quality of the hyperparameter combination by the model's emotion recognition accuracy. In the genetic algorithm process, a hyperparameter population is first initialized, and then crossover and mutation operations are performed on each hyperparameter combination to generate new combinations. Through multiple iterations, the genetic algorithm gradually selects the optimal hyperparameter combination, enabling the emotion perception system to exhibit optimal performance under different driver emotional fluctuations. Simultaneously, a reinforcement learning mechanism is introduced to assist the genetic algorithm in further selecting hyperparameter combinations. During each emotion perception process, the reinforcement learning model adjusts the selection of hyperparameter combinations in real time based on the error between the actual emotion recognition result and the driver's facial expression auxiliary labels. Reinforcement learning mechanisms optimize the hyperparameter selection process through reward and punishment mechanisms. If a certain combination of hyperparameters makes the emotion perception result more accurate, the reinforcement learning model will give a higher reward to that combination; otherwise, it will give a punishment. This helps to adjust the hyperparameter combination in real time under dynamically changing emotional states, thereby improving the accuracy of driver emotion perception and the model's adaptability. By combining genetic algorithms and reinforcement learning, hyperparameters are continuously optimized to ensure efficient emotion recognition capabilities for different drivers.

[0070] The perception weight adjustment branch is adjusted in real time based on the optimal hyperparameters. The parameter settings in the perception weight adjustment branch are dynamically adjusted based on the optimal hyperparameters obtained from genetic algorithms and reinforcement learning to better adapt to the driver's emotional fluctuations. Specifically, the weights of facial expression features are adjusted according to the optimal hyperparameters, making the sensitivity to micro-expressions more accurate, especially when the driver's facial expressions are not rich or the emotional expression is weak, so as to better capture emotional changes. At the same time, these adjustments help the system flexibly respond to the driver's personalized emotional changes in different driving environments and emotional states, further improving the accuracy and real-time performance of emotion perception.

[0071] The specific steps for establishing a driving delay compensation mechanism to synchronously correct real-time collected driving behavior data and corresponding driving behavior data streams include:

[0072] Real-time driving behavior data and driver emotion perception results are synchronized in time to obtain initial behavior-emotion alignment data. In actual driving, there is usually a lag between driver emotions and behavior, with emotional fluctuations often preceding changes in driving behavior. Therefore, time synchronization ensures that the driving behavior data and emotion perception results match at the same point in time. The core of this step is to align the driver's emotion perception information with the real-time driving behavior data through the synchronization of timestamps and data streams, forming a set of data that can simultaneously reflect the driver's emotional and behavioral states.

[0073] Establish an emotion-behavior correlation model; use the emotion-behavior correlation model to predict the driver's emotion perception results and driving environment data to obtain the predicted driving behavior lag time; and use the emotion-behavior correlation model to perform emotion inversion on driving behavior data to obtain the predicted driving behavior emotion time; combine the predicted driving behavior lag time and the predicted driving behavior emotion time to obtain behavior-emotion prediction time delay data.

[0074] The main purpose of the emotion-behavior association model is to capture the relationship between a driver's emotional fluctuations and their driving behavior, and to predict how emotional changes affect driving behavior. In this process, the emotion-behavior association model is trained based on historical data to establish a dynamic relationship between emotional changes and driving behavior responses. Specifically, the emotion-behavior association model needs to use deep learning models (such as Long Short-Term Memory Networks (LSTM) or Convolutional Neural Networks (CNN)) to extract the temporal patterns between driver emotional fluctuations and driving behavior. The goal of this step is to build a model that can predict a driver's future driving behavior from their emotional state, providing an accurate basis for subsequent delay compensation and behavior prediction.

[0075] After establishing the emotion-behavior correlation model, this model is used to predict the driver's emotional perception results and driving environment data, obtaining the predicted driving behavior lag delay. In actual driving, the driver's emotional reaction usually precedes their actual driving behavior, and emotional changes introduce a certain delay to the driver's actions. Therefore, the emotion-behavior correlation model can predict the lag delay of driving behavior based on the current emotional perception results and driving environment data. This delay represents the time difference between the driver's emotional change and the actual behavior. By predicting the driving behavior lag delay, the system can more accurately adjust the perception and intervention strategies for driver behavior; simultaneously, the emotion-behavior correlation model can predict the lag delay of driving behavior. Behavioral association models can also perform emotion inversion on driving behavior data to obtain predicted driving behavior emotion latency. The purpose of this step is to infer the emotional state corresponding to a certain driving behavior through the model. Since there is a two-way relationship between driver behavior and emotion, the model can predict the driver's possible emotional fluctuations based on the driver's behavior. This process helps to capture the delay caused by the emotion not being immediately expressed. For example, when the driver suddenly accelerates or brakes hard, the system can infer the driver's emotional reaction (such as anxiety or tension) and provide feedback to the system. Through this inverse inference, the system can identify the time delay relationship between emotion and behavior, helping to understand the emotional changes behind the driver's behavior.

[0076] By combining the lag latency of predicted driving behavior and the latency of predicted driving behavior emotion, we obtain behavior-emotion prediction latency data. The purpose of this step is to combine the lag latency and the emotion latency to form a comprehensive prediction latency dataset, which is used to accurately calculate the temporal differences between driving behavior and emotion. By combining these two latencies, we can perform more accurate compensation in the behavior data stream.

[0077] The system matches initial behavior-emotion alignment data and behavior-emotion prediction latency data within the driving behavior data stream to obtain new driving behavior data. This process is repeated to obtain new driving behavior data. The key to this step is matching real-time collected driving behavior data with emotion perception results, incorporating the predicted latency data to correct the behavior data. Specifically, the system performs forward or backward corrections on the driving behavior data based on the predicted latency to ensure the behavior data matches the driver's actual emotional response. Through this latency compensation mechanism, the system dynamically adjusts behavior prediction and intervention strategies based on the driver's emotional state, ultimately generating more accurate driving behavior data and providing more precise emotion perception and driving assistance.

[0078] Based on the driver's emotional perception results and corresponding new driving behavior data and driving environment data, the driver's current driving behavior is analyzed to obtain the emotion-driving behavior analysis results, and driving behavior intervention operations based on driving behavior and emotional perception are realized.

[0079] The specific steps for analyzing the current driver's driving behavior based on the driver's emotional perception results and corresponding new driving behavior data and driving environment data include:

[0080] The aligned driver emotion perception results, new driving behavior data, and driving environment data are fused to obtain a comprehensive evaluation feature vector; driving behavior analysis is performed based on the comprehensive evaluation feature vector to obtain emotion-driving behavior analysis results; graded warnings are issued based on the emotion-driving behavior analysis results, and different feedback behaviors are executed on the current driver in the vehicle according to different graded warnings;

[0081] The aligned driver emotion perception results, new driving behavior data, and driving environment data are fused to obtain a comprehensive evaluation feature vector. Based on this comprehensive evaluation feature vector, driving behavior analysis is performed to obtain emotion-driving behavior analysis results. The core of this step is to analyze the relationship between the driver's emotions and behaviors. The emotion-driving behavior analysis results can provide an in-depth understanding of the driver's current emotional state, and combined with their driving behavior, predict the driver's possible next action. For example, if the driver exhibits anxiety or fatigue, and the current environment requires temporary control of the train, the system may predict that the driver will make a more hasty or dangerous driving behavior in a short period of time. Through the comprehensive analysis of these factors, a comprehensive view of the driver's behavior can be provided, and the impact of their behavior on driving safety can be assessed.

[0082] Based on the analysis of emotion and driving behavior, a tiered warning system is implemented. This process aims to provide appropriate safety warnings and feedback based on the driver's emotional and behavioral state. During this process, the driver's emotional and behavioral state is categorized into multiple levels according to the severity of the emotion and driving behavior. For example, if the driver is excessively anxious or fatigued and exhibits abnormal behavior, the system will classify it as a high-risk level, triggering a more urgent warning or intervention. Depending on the level, the system will automatically select different feedback methods, which may include audible warnings, visual cues, and vibration feedback. For instance, if abnormal driving behavior caused by driver fatigue is detected, the system may remind the driver to rest, or if the driver experiences severe emotional fluctuations, the system can automatically activate driver assistance systems (such as lane keeping assist) to reduce potential risks.

[0083] In this embodiment, it is assumed that the test is conducted on a high-speed train equipped with onboard sensors and cameras to monitor the driver's driving behavior and facial expressions in real time. Through an emotion perception system, the system can detect the driver's emotional state and combine it with driving behavior data (such as vehicle speed, acceleration, braking force, etc.) and external driving environment data (such as weather, road conditions, wind speed, etc.) to provide emotion-behavior analysis results. At the same time, the following scenario test is set: at a vehicle speed of 300 kilometers per hour, the driver's anxiety fluctuations caused by the external environment (such as complex tracks, bad weather) are simulated. During this process, onboard cameras and sensors continuously monitor the driver's facial expressions and analyze the driver's emotional state in real time through emotion recognition models, such as fatigue, anxiety, and stress. At the same time, data such as the train's driving status, acceleration, speed, steering angle, and braking force are also collected in real time through the vehicle control system. Driving environment data, such as weather conditions, road gradient, track conditions, and other external factors, are obtained through meteorological sensors, road condition monitoring systems, and GPS. After time synchronization, this information forms a unified emotion-driving behavior data stream, ensuring that the driver's emotional state, driving behavior, and environmental data can be accurately aligned at the same point in time.

[0084] Based on aligned data, a deep learning model is used to analyze the correlation between driver emotions and driving behavior. By fusing driver emotion perception results, driving behavior data, and driving environment information, the system forms a comprehensive evaluation feature vector and uses an emotion-behavior correlation model to predict the driver's possible future behavior. For example, when a driver exhibits high anxiety, the system predicts that he / she may perform high-risk behaviors such as sudden braking or sudden acceleration. Based on this analysis, the system can generate emotion-driving behavior analysis results and provide graded warnings to the driver based on changes in driver emotion and the danger of driving behavior.

[0085] Based on different emotional and behavioral analysis results, the system will issue different levels of warnings. For example, if it detects that the driver is extremely fatigued or emotionally unstable, accompanied by violent driving behaviors (such as frequent sudden braking or acceleration), the system will remind the driver to take a rest through audio prompts or the in-vehicle screen, and may intervene through the in-vehicle vibration system or autonomous driving assistance systems (such as automatic braking and lane keeping). If the driver only shows mild unease or anxiety, the system may guide the driver to relax through gentle reminders, such as playing soothing music or displaying emotion regulation suggestions. Through this real-time feedback, the system ensures that the driver's emotional fluctuations and unsafe driving behaviors can be identified and adjusted in a timely manner during vehicle operation, thereby improving driving safety and reducing the risk of accidents caused by changes in the driver's emotions.

[0086] Example 2, a driving behavior analysis system based on emotion perception, see [link to example]. Figure 1 As shown, it includes:

[0087] The emotion perception correction module includes a data acquisition unit and a differential emotion perception unit. The data acquisition unit prioritizes capturing high-frequency facial expressions of the current driver to obtain driver expression auxiliary labels. It also collects real-time data on the current driver, acquiring driving environment data, driving behavior data, and driver facial data. The differential emotion perception unit inputs the driver facial data into the driving emotion perception model for analysis to obtain the driver's emotion perception results. The driving emotion perception model is constructed based on the ResNet-50 architecture, adding micro-expression feature branches, spatiotemporal attention branches, and perception weight adjustment branches. Simultaneously, the hyperparameters of the perception weight adjustment branch are adjusted in real-time using the driver expression auxiliary labels. A driving delay compensation mechanism is established to synchronously correct the real-time collected driving behavior data and the corresponding driving behavior data stream, obtaining new driving behavior data.

[0088] The driving behavior analysis module includes a behavior analysis unit. The behavior analysis unit is used to analyze the current driver's driving behavior based on the driver's emotional perception results and corresponding new driving behavior data and driving environment data, to obtain emotion-driving behavior analysis results, and to realize driving behavior intervention operations based on driving behavior and emotional perception.

[0089] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims. Parts not described in detail in this specification are prior art known to those skilled in the art.

Claims

1. A driving behavior analysis method based on emotion perception, characterized in that, The method comprises the following steps: Preferentially capturing high-frequency expressions of the current driver to obtain driver expression auxiliary labels; Collecting data of the current driver in real time to obtain driving environment data, driving behavior data and driving facial data; Inputting the driving facial data into a driving emotion perception model for analysis to obtain driver emotion perception results; in the driving emotion perception model, a micro-expression feature branch, a space-time attention branch and a perception weight adjustment branch are added based on a Resnet-50 architecture to construct the model; meanwhile, the superparameters of the perception weight adjustment branch are adjusted in real time by using the driver expression auxiliary labels; a driving delay compensation mechanism is established to correct the real-time collected driving behavior data and the corresponding driving behavior data stream to obtain new driving behavior data; Based on the driver emotion perception results and the corresponding new driving behavior data and driving environment data, the driving behavior of the current driver is analyzed to obtain emotion-driving behavior analysis results, and driving behavior intervention operations based on driving behavior and emotion perception are realized; The specific steps of preferentially capturing high-frequency expressions of the current driver include: Using a super-high-frequency facial expression detection algorithm to capture the facial expressions of the current driver to obtain a plurality of facial key point data; Based on the facial key point data, a deep learning model is used to calculate the change amplitude of the facial key points of the current frame and the previous frame by combining a time window to obtain single-frame change amplitudes; all single-frame change amplitudes are combined to obtain a frame change amplitude sequence; The frame change amplitude sequence is subjected to feature extraction by using a space-time feature extraction method to obtain a micro-expression change time sequence feature; based on the micro-expression change time sequence feature, auxiliary label mapping is performed to obtain driver expression auxiliary labels; In the driving emotion perception model, there are an input layer, a convolution layer, a pooling layer, four residual blocks, a global average pooling layer and a full connection layer; The structure of each residual block is composed of a 1*1 dimension reduction convolution, a 3*3 spatial convolution and a 1*1 dimension increasing convolution; a micro-expression feature branch is added in parallel after the third residual block, a space-time attention branch is inserted before the fourth residual block, and a perception weight adjustment branch is added before the global average pooling layer; The micro-expression feature branch is used to capture the micro-movement of facial muscles in the driving facial data to obtain driver micro-expression features; The space-time attention branch is used to output driver weighted facial features supplemented based on the driver micro-expression features; The perception weight adjustment branch is used to adjust the perception weight of the driver weighted facial features to obtain driver emotion perception features; Finally, the driver emotion perception results of the current driver are output in the full connection layer. 2.The emotion-aware driving behavior analysis method of claim 1, wherein, The specific steps of adjusting the superparameters of the perception weight adjustment branch in real time by using the driver expression auxiliary labels include: Setting default superparameters in the perception weight adjustment branch; Searching for a superparameter combination adjusted based on the driver expression auxiliary labels based on a genetic algorithm; meanwhile, a reinforcement learning mechanism is introduced to assist the genetic algorithm in screening the superparameter combination during the emotion perception process to obtain optimal superparameters; Adjusting the perception weight adjustment branch in real time based on the optimal superparameters. 3.The emotion-aware driving behavior analysis method of claim 2, wherein, The specific steps of establishing a driving delay compensation mechanism and synchronously correcting the real-time collected driving behavior data and the corresponding driving behavior data stream include: synchronizing the real-time collected driving behavior data and the driver emotion perception result in time to obtain behavior-emotion initial alignment data; establishing an emotion-behavior correlation model; using the emotion-behavior correlation model to predict the driver emotion perception result and the driving environment data to obtain a predicted driving behavior lag time delay; and using the emotion-behavior correlation model to perform emotion inversion on the driving behavior data to obtain a predicted driving behavior emotion time delay; combining the predicted driving behavior lag time delay and the predicted driving behavior emotion time delay to obtain behavior-emotion prediction time delay data; matching the behavior-emotion initial alignment data and the behavior-emotion prediction time delay data in the driving behavior data stream to obtain matched corresponding driving behavior data as new driving behavior data. 4.The emotion-aware driving behavior analysis method of claim 3, wherein, The specific steps of performing driving behavior analysis on the current driver based on the driver emotion perception result and the corresponding new driving behavior data and driving environment data include: performing feature fusion on the aligned driver emotion perception result, new driving behavior data and driving environment data to obtain a comprehensive evaluation feature vector; performing driving behavior analysis based on the comprehensive evaluation feature vector to obtain an emotion-driving behavior analysis result; and performing hierarchical warning according to the emotion-driving behavior analysis result and executing different feedback behaviors in the vehicle for the current driver according to different hierarchical warnings.

5. A driving behavior analysis system based on emotion perception, characterized by, The system applies the emotion perception-based driving behavior analysis method of any one of claims 1-4, and includes: An emotion perception correction module includes a data acquisition unit and a differential emotion perception unit; the data acquisition unit is used to preferentially capture high-frequency expressions of the current driver to obtain driver expression auxiliary labels; the current driver is real-time data acquisition to obtain driving environment data, driving behavior data and driving facial data; the differential emotion perception unit is used to input the driving facial data into a driving emotion perception model for analysis to obtain a driver emotion perception result; wherein, in the driving emotion perception model, a micro-expression feature branch, a space-time attention branch and a perception weight adjustment branch are added based on the Resnet-50 architecture to construct the model; at the same time, the hyperparameters of the perception weight adjustment branch are adjusted in real time by using the driver expression auxiliary labels; a driving delay compensation mechanism is established to synchronously correct the real-time collected driving behavior data and the corresponding driving behavior data stream to obtain new driving behavior data; A driving behavior analysis module includes a behavior analysis unit; the behavior analysis unit is used to perform driving behavior analysis on the current driver based on the driver emotion perception result and the corresponding new driving behavior data and driving environment data to obtain an emotion-driving behavior analysis result, and realize driving behavior intervention operation based on driving behavior and emotion perception.

Citation Information

Patent Citations

  • Driver emotion recognition method, device and equipment and storage medium

    CN116994230A

  • Intelligent automobile dynamic adjustment method and system based on user experience

    CN118907118A