Driver Emotion Recognition Method, Device, Equipment and Storage Medium

By extracting feature of driver facial images and abstracting spatiotemporal feature sequences, the problem of low accuracy of driver emotion recognition is solved by using 3DCNN and ConvLSTM models, and efficient and accurate emotion recognition and safe driving assistance are achieved.

CN116994230BActive Publication Date: 2025-07-25WUHAN POLYTECHNIC UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310910029.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-21
Publication Date
2025-07-25
Estimated Expiration
2043-07-21

AI Technical Summary

Technical Problem

In the prior art, the driver's emotional recognition accuracy is low, and the driver's emotional state cannot be effectively identified, which affects driving safety.

Method used

By obtaining the driver's facial image, extracting facial features, and using 3DCNN and ConvLSTM models to extract spatiotemporal feature sequences and performing abstract processing, the driver's emotional state is obtained, and combined with the SPP layer to solve the image size problem and improve the recognition accuracy.

Benefits of technology

It improves the accuracy and efficiency of emotional recognition, and can provide early warning and comfort when the driver's emotions are extreme, ensuring driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994230B_ABST
    Figure CN116994230B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of emotion recognition, and discloses a method, device, equipment and storage medium for driver emotion recognition. The present invention obtains a facial image of a driver, extracts facial features of the facial image, extracts a spatio-temporal feature sequence of the facial features through a driver emotion recognition model, abstracts the spatio-temporal feature sequence to obtain an abstract feature, and obtains the emotional state of the driver according to the abstract feature. By performing feature extraction and abstraction processing on the facial image of the driver through the driver emotion recognition model, the emotion of the driver is obtained. Compared with traditional emotion recognition, the present invention can improve the accuracy of emotion recognition and reduce the response time, and can improve the emotion recognition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion recognition, and particularly to a method, device, equipment and storage medium for recognizing driver emotions. Background Art

[0002] The emotion of a driver is one of the factors affecting driving safety. For example, situations such as road congestion and other vehicles cutting in will cause the driver's emotion to be unstable and result in excessive behaviors. Therefore, recognizing the emotion of vehicle drivers has great safety significance. In the related art, during the process of recognizing the driver's emotion, by recognizing the driver's voice to determine the driver's emotional state, this solution cannot recognize all emotions, so it has low accuracy, or by recognizing the driver's face, using a single convolutional network for emotion recognition, the accuracy of this solution for emotion recognition is low, so it cannot meet the safety needs of drivers either. Summary of the Invention

[0003] The main purpose of the present invention is to provide a method, device, equipment and storage medium for recognizing driver emotions, aiming to solve the technical problem of low accuracy in emotion recognition in the prior art.

[0004] To achieve the above purpose, the present invention provides a method for recognizing driver emotions, the method comprising the following steps:

[0005] Obtain a facial image of a driver, and extract facial features of the facial image;

[0006] Pass the facial features through a driver emotion recognition model to extract a spatio-temporal feature sequence of the facial features;

[0007] Abstract the spatio-temporal feature sequence to obtain abstract features;

[0008] Obtain the emotional state of the driver according to the abstract features.

[0009] Optionally, before passing the facial features through a driver emotion recognition model to extract a spatio-temporal feature sequence of the facial features, it further includes:

[0010] Screen a face expression data set to obtain negative expressions;

[0011] Classify the negative emotions, and add expression labels to the negative expressions according to the classification results to obtain a negative emotion set;

[0012] Divide the face expression data set to obtain a training data set and a validation data set, and the face expression data set includes the negative emotion set;

[0013] Input the training data set into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training data set;

[0014] Input the local spatio-temporal sequence features into the ConvLSTM for abstraction to obtain abstract features;

[0015] Perform feature transformation on the abstract features to obtain the final features;

[0016] Make a judgment based on the final features to obtain the emotion recognition result and obtain an intermediate emotion recognition model;

[0017] Verify the intermediate emotion recognition model according to the validation data set to obtain the recognition error. When the recognition error is less than the error threshold, use the intermediate emotion recognition model as the emotion recognition model.

[0018] Optionally, classify the negative emotions, add expression labels to the negative expressions according to the classification results to obtain a negative emotion set, including:

[0019] Determine several negative emotion subsets according to the classification results and add emotion labels to the several negative emotion subsets correspondingly;

[0020] Screen the several emotion subsets to obtain strongly expressed samples;

[0021] Obtain an extreme expression set according to the strongly expressed samples and add emotion labels to the extreme expression set;

[0022] Add the extreme expression set to the negative emotion set.

[0023] Optionally, performing the feature transformation on the abstract features to obtain the final features includes:

[0024] Input the abstract features into a spatial pyramid pooling layer for feature transformation to obtain the final features, and the spatial pyramid pooling layer is connected to the last convolutional layer and the fully connected layer in the ConvLSTM.

[0025] Optionally, inputting the training data set into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training data set includes:

[0026] Separate the expression information and label information in the training data set to generate an expression data file and a label data file;

[0027] Establish a comparison table according to the expression data file and the label data file, and input the comparison table into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training data set.

[0028] Optionally, after obtaining the driver's emotional state according to the abstract feature, the method further includes:

[0029] When the driver's emotional state is an extreme expression, determining the expression category of the extreme expression;

[0030] Invoking a light warning and a voice warning for reminder according to the expression category;

[0031] Performing emotional comfort on the driver according to the expression category.

[0032] Optionally, after determining the expression category of the extreme expression when the driver's emotional state is an extreme expression, the method further includes:

[0033] Obtaining the duration of the extreme emotion;

[0034] When the duration of the extreme emotion exceeds a safety threshold, obtaining the current driving trajectory and judging the driving trajectory to determine the current driving behavior;

[0035] When the current driving behavior is a violent driving behavior, taking over the vehicle driving control right and driving the vehicle to a safe location to stop until the driver's emotion is stable.

[0036] In addition, to achieve the above object, the present invention further provides a driver emotion recognition device, where the driver emotion recognition device includes:

[0037] An image acquisition module, configured to acquire a facial image of a driver and extract facial features of the facial image;

[0038] A feature extraction module, configured to extract a spatio-temporal feature sequence of the facial features through a driver emotion recognition model;

[0039] A feature abstraction module, configured to abstract the spatio-temporal feature sequence to obtain an abstract feature;

[0040] An emotion recognition module, configured to obtain the driver's emotional state according to the abstract feature.

[0041] In addition, to achieve the above object, the present invention further provides a driver emotion recognition device, where the driver emotion recognition device includes: a memory, a processor, and a driver emotion recognition program stored on the memory and executable on the processor, and the driver emotion recognition program is configured to implement the steps of the driver emotion recognition method as described above.

[0042] In addition, to achieve the above object, the present invention also provides a storage medium, on which a driver emotion recognition program is stored. When the driver emotion recognition program is executed by a processor, the steps of the driver emotion recognition method described above are implemented.

[0043] The present invention obtains a facial image of a driver, extracts facial features of the facial image, extracts a spatio-temporal feature sequence of the facial features through a driver emotion recognition model, abstracts the spatio-temporal feature sequence to obtain an abstract feature, and obtains the emotional state of the driver according to the abstract feature. By performing feature extraction and abstraction processing on the facial image of the driver through the driver emotion recognition model, the emotion of the driver is obtained. Compared with traditional emotion recognition, the present invention can improve the accuracy of emotion recognition and reduce the response time, and can improve the emotion recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic structural diagram of a driver emotion recognition device in a hardware operating environment related to an embodiment of the present invention;

[0045] Figure 2 is a schematic flowchart of a first embodiment of the driver emotion recognition method of the present invention;

[0046] Figure 3 is a structural diagram of a driver emotion early warning assistance system in an embodiment of the driver emotion recognition method of the present invention;

[0047] Figure 4 is a schematic flowchart of a second embodiment of the driver emotion recognition method of the present invention;

[0048] Figure 5 is a block diagram of a first embodiment of the driver emotion recognition device of the present invention.

[0049] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] Refer to Figure 1 , Figure 1 is a schematic structural diagram of a driver emotion recognition device in a hardware operating environment related to an embodiment of the present invention.

[0052] As Figure 1As shown in the figure, the driver emotion recognition device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0053] Those skilled in the art can understand that Figure 1 the structure shown in the figure does not constitute a limitation on the driver emotion recognition device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0054] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and a driver emotion recognition program.

[0055] In Figure 1 the driver emotion recognition device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the driver emotion recognition device of the present invention may be provided in the driver emotion recognition device. The driver emotion recognition device calls the driver emotion recognition program stored in the memory 1005 through the processor 1001 and executes the driver emotion recognition method provided by the embodiments of the present invention.

[0056] The embodiments of the present invention provide a driver emotion recognition method. Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of a driver emotion recognition method of the present invention.

[0057] In this embodiment, the driver emotion recognition method includes the following steps:

[0058] Step S10: Obtain the driver's facial image and extract the facial features of the facial image.

[0059] It should be noted that the execution subject of this embodiment is a driver emotion recognition device. Among them, the driver emotion recognition device has functions such as data processing, data communication, and program operation. The driver emotion recognition device can be an integrated controller, a control computer, etc. Of course, it can also be other devices with similar functions. This embodiment does not limit this.

[0060] It can be understood that the driver's facial image is captured by an infrared camera, and the facial image is transmitted to the driver emotion recognition device in the vehicle through the vehicle computer. The infrared camera can obtain the driver's facial image both during the day and at night. The acquisition of the driver's facial image by the infrared camera can be triggered by the activation of the emotion detection function. When the vehicle is started and the driver is seated in the driver's seat, the infrared camera is activated to acquire the driver's facial image.

[0061] In a specific implementation, when the driver enters the driver emotion recognition area in the vehicle, the infrared camera can collect the driver's facial image at a preset frequency. Since emotions can almost all be reflected through facial expressions, it is possible to recognize emotions by processing the driver's facial image. During the recognition process, the facial features in the facial image are extracted, including relevant features such as eyes, mouth, eyebrows, nose, etc. This embodiment does not limit this, and the facial features related to facial expressions are extracted.

[0062] Step S20: Use the driver emotion recognition model to extract the spatio-temporal feature sequence of the facial features.

[0063] It should be noted that the driver emotion recognition model is used to recognize the driver's emotion state based on facial features. The driver emotion recognition model is composed of 3DCNN, ConvLSTM, and SVM. The facial features are subjected to a series of image conversions and feature extractions through the driver emotion recognition model to obtain the driver's emotion state.

[0064] It should be understood that the spatio-temporal feature sequence refers to the set of the driver's facial features within a certain time period. Since a single facial feature cannot be used as an accurate basis for emotion judgment, in order to ensure the accuracy of emotion recognition, emotion recognition can be performed through the features within a certain time period.

[0065] In a specific implementation, when the driver emotion recognition model obtains facial features, the 3DCNN in the driver emotion recognition model captures temporal features of the facial features. By using the facial features as input values within the 3DCNN, through a number of convolutional layers and fully connected layers, an output spatio-temporal feature sequence is obtained, and the spatio-temporal feature sequence reflects the spatio-temporal changes of the driver's facial features.

[0066] Step S30: Abstract the spatio-temporal feature sequence to obtain abstract features.

[0067] It should be noted that the abstract features refer to extracting and abstracting the emotion-related features in the spatio-temporal feature sequence, specifically referring to performing convolution on the spatio-temporal feature sequence.

[0068] In a specific implementation, abstracting the spatio-temporal feature sequence to obtain abstract features is completed in ConvLSTM. ConvLSTM is a deep learning model, which is a combination of the long short-term memory network LSTM and the convolutional neural network CNN. In ConvLSTM, the recurrent neural network RNN unit of LSTM is replaced by a convolutional operation. In LSTM, the recurrent neural network can store past information in the cell state and use this information in subsequent time steps. By performing convolutional abstraction on the spatio-temporal feature sequence data, abstract features are obtained, where the abstract features include all the necessary features for the expression recognition.

[0069] Step S40: Obtain the driver's emotional state based on the abstract features.

[0070] It should be noted that the driver's emotional state can include various types, such as different expressions like happy, surprised, sad, afraid, angry, etc. Based on these expressions, the current driver's emotional state can be determined.

[0071] In a specific implementation, the abstract features already contain the necessary features for expression recognition. Therefore, by mapping the necessary features in the abstract features to the features stored in the driver emotion recognition model, the matching degree between the abstract features and various stored emotions is determined. For a single facial feature, it may correspond to multiple different emotional states. Therefore, when obtaining the driver's emotional state based on the abstract features, each facial feature in the abstract features needs to be mapped, and then numerous emotional states with different matching degrees are obtained. The emotional state with the highest matching degree among the numerous emotional states is used as the final matching result, which is the emotional state recognized by the driver emotion recognition model.

[0072] Furthermore, after obtaining the driver's emotional state based on the abstract features, it further includes:

[0073] When the emotional state of the driver is an extreme expression, determine the expression category of the extreme expression;

[0074] Call light warning and voice warning for reminder according to the expression category;

[0075] Carry out emotional comfort for the driver according to the expression category.

[0076] In specific implementation, refer to Figure 3 , Figure 3 is the structural diagram of the driver emotion warning assistance system. When the driver is driving, the driver emotion recognition is started through the system startup module, and the emotion of the driver can be detected by the infrared camera. When the expression of the driver is in a non-extreme expression, the state of the detected driver is the normal state. The background system communicates with the background server through the WebSocket technology and detects the current emotional state of the driver in real time. When the driver is in the driving state, a connection request is sent to the WebSocket server. The WebSocket server then establishes a connection with the application, and the application sends the detected emotion, face frame, and the current position and driving trajectory of the vehicle to the background server in real time. At the same time, the background monitoring system also establishes a WebSocket connection with the server. Therefore, the server can synchronize the information received by the in-vehicle mobile terminal to the background monitoring system in real time. The background monitoring system can monitor the position of the driver and the driving trajectory of the vehicle in real time to determine whether there is violent driving behavior. By supervising and analyzing the driver database, it is judged whether the driver's emotion is in an abnormal state for a long time. After the driver of the vehicle has an out-of-control emotion for up to three minutes, the warning system will call the connected atmosphere light to flash red to remind the driver. After five minutes of continuous out-of-control emotion, a voice prompt will be sent through the in-vehicle voice to comfort the driver.

[0077] Further, after determining the expression category of the extreme expression when the emotional state of the driver is an extreme expression, it further includes:

[0078] Obtain the duration of the extreme emotion;

[0079] When the duration of the extreme emotion exceeds the safety threshold, obtain the current driving trajectory and judge the driving trajectory to determine the current driving behavior;

[0080] When the current driving behavior is a violent driving behavior, take over the vehicle driving control right and drive the vehicle to a safe place to stop until the driver's emotion is stable.

[0081] In a specific implementation, the server can synchronize the information received by the in-vehicle mobile terminal to the background monitoring system in real time. The background monitoring system can monitor the driver's location and the vehicle's driving trajectory in real time to determine whether there is violent driving behavior. When the voice warning is still ineffective, the system will call the assisted driving intelligent takeover module. By analyzing the driving trajectory in the out-of-control emotional state, such as whether there are violent driving behaviors such as malicious lane-changing and speeding, it is determined whether to take over the current vehicle's driving right. After determining to accept the vehicle driving, it maintains lane driving or makes a safe stop according to the current road conditions. After the driver's state is stable and requests to take over the driving right, the assisted driving is withdrawn.

[0082] In this embodiment, by obtaining the driver's facial image, extracting the facial features of the facial image, extracting the spatio-temporal feature sequence of the facial features through the driver emotion recognition model, abstracting the spatio-temporal feature sequence to obtain an abstract feature, and obtaining the driver's emotional state according to the abstract feature. Through the feature extraction and abstraction processing of the driver's facial image by the driver emotion recognition model, the driver's emotion is obtained. Compared with traditional emotion recognition, the present invention can improve the accuracy of emotion recognition and reduce the response time, and can improve the emotion recognition efficiency.

[0083] Reference Figure 4 , Figure 4 is a schematic flowchart of the second embodiment of a method for recognizing a driver's emotion according to the present invention.

[0084] Based on the above first embodiment, before the step S20 of this embodiment of the driver emotion recognition method, it further includes:

[0085] Step S201: Screen the facial expression data set to obtain negative expressions.

[0086] Step S202: Classify the negative emotions, and add an expression label to the negative expression according to the classification result to obtain a negative emotion set.

[0087] Step S203: Divide the facial expression data set to obtain a training data set and a validation data set, and the facial expression data set includes the negative emotion set.

[0088] Step S204: Input the training data set into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of the negative expressions in the training data set.

[0089] Step S205: Input the local spatio-temporal sequence features into ConvLSTM for abstraction to obtain abstract features.

[0090] Step S206: Perform feature transformation on the abstract features to obtain final features.

[0091] Step S207: Make a judgment based on the final feature to obtain an emotion recognition result and obtain an intermediate emotion recognition model.

[0092] Step S208: Verify the intermediate emotion recognition model according to the verification data set to obtain a recognition error. When the recognition error is less than the error threshold, use the intermediate emotion recognition model as the emotion recognition model.

[0093] It should be noted that negative expressions refer to bad emotions that will affect the driving safety of the driver, such as anger, disgust, fear, sadness, surprise, etc. This embodiment does not limit this. The negative emotion set is the set formed by negative emotions. That is to say, when the emotion corresponding to the facial feature falls into the negative emotion set, the emotion of the driver at this time is a negative emotion, while the human face expression data set contains all emotion types, and the negative emotion set belongs to a part of the human face expression data set.

[0094] In a specific implementation, a 3DCNN + ConvLSTM network is used to train the FER2013 dataset, and then a driver abnormal emotion recognition model is obtained. The FER2013 dataset is already a relatively standardized and unified dataset, and usually does not require too many image processing steps. Therefore, the face expression dataset is directly screened according to the FER2013 dataset to obtain negative expressions. The FER2013 dataset is sorted out, and samples containing five expressions of anger, disgust, fear, sadness, and surprise are selected. These five expressions correspond to digital labels and Chinese-English identifications respectively: 0 represents anger, 1 represents disgust, 2 represents fear, 3 represents sadness, and 4 represents surprise. Among the above five negative expression samples, samples with strong expressions are selected and formed into a separate category. The digital label and Chinese-English identification of this category are: '5uncontrolled out-of-control emotion'. Then, the obtained face expression dataset is divided into a training dataset and a validation dataset. The data in the training dataset is input into the initial three-dimensional convolutional neural network to extract the local spatio-temporal feature sequence features of the negative expressions in the training set, and the obtained local spatio-temporal sequence features are input into ConvLSTM for abstraction to obtain abstract features. After obtaining the abstract features, the abstract features are transformed to obtain the final features, and the emotion state corresponding to the current final features is determined according to the final features, and an intermediate emotion recognition model is obtained. The intermediate recognition model is verified according to the data in the validation dataset, and the recognition accuracy is statistically calculated. When the recognition accuracy is less than the preset threshold, the intermediate emotion recognition model can be used as the emotion recognition model, and the deep learning convolutional neural network is used for feature extraction and classification model training to optimize the parameters and adjust the model structure.

[0095] Further, for classifying the negative emotions, adding expression labels to the negative expressions according to the classification results to obtain a negative emotion set, including:

[0096] Determining several negative emotion subsets according to the classification results, and correspondingly adding emotion labels to the several negative emotion subsets;

[0097] Screening the several emotion subsets to obtain samples with strong expressions;

[0098] Obtaining an extreme expression set according to the samples with strong expressions, and adding emotion labels to the extreme expression set;

[0099] Adding the extreme expression set to the negative emotion set.

[0100] In specific implementation, the face expression dataset is screened according to the FER2013 dataset to obtain negative expressions. The FER2013 dataset is sorted out, and samples containing five expressions of anger, disgust, fear, sadness, and surprise are selected. These five expressions respectively correspond to digital labels and Chinese-English identifications: 0 represents anger, 1 represents disgust, 2 represents fear, 3 represents sadness, and 4 represents surprise. Among the above five negative expression samples, samples with strong expressions are selected and formed into a separate category. The digital label and Chinese-English identification of this category are: '5 uncontrolled out-of-control emotions'. The set formed by out-of-control emotions is the negative emotion set.

[0101] Further, the feature transformation of the abstract feature to obtain the final feature includes:

[0102] The abstract feature is input into a spatial pyramid pooling layer for feature transformation to obtain the final feature. The spatial pyramid pooling layer is connected to the last convolutional layer and the fully connected layer in the ConvLSTM.

[0103] In specific implementation, in a general CNN structure, a fully connected layer is usually connected after the convolutional layer. The number of features in the fully connected layer is fixed, so when inputting into the network, the input size (fixed-size) is fixed. However, in reality, the size of the input image always cannot meet the required size during input. The usual processing method is cropping and stretching. The disadvantages brought by this are that the aspect ratio and size of the input image will be changed, which will distort the original image, resulting in a reduction in image pixels or being blurred. Introducing a SPP (Spatial Pyramid Pooling) layer can well solve such problems. The SPP is usually connected to the last convolutional layer. Through the original weighted method of the spatial pyramid kernel and associating it with the SPP, where the high-resolution feature map has a higher weight than the low-resolution feature map, based on the output of the ConvLSTM layer, the SPP layer is used for multi-scale feature extraction. Through the feature scale transformation of the SPP layer, a weighted spatio-temporal feature vector is obtained. After the features abstracted in the previous steps, they are then transformed again through a fully connected layer. Compared with the features abstracted by the ConvLSTM, the features after the SPP transformation represent pattern information of different scales.

[0104] Further, inputting the training dataset into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training dataset includes:

[0105] Separate the expression information and label information in the training dataset to generate an expression data file and a label data file;

[0106] Establish a comparison table based on the expression data file and the label data file, and input the comparison table into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training dataset.

[0107] Use the above five negative expression samples of 0-4 as inputs to train an expression classifier to monitor out-of-control emotions, find out the out-of-control expression samples and label them. Further process the data, separate emotion and pixels, and generate two files, label.csv and data.csv. Convert the training, test, and validation data into image format and divide them into 0-6 categories, corresponding to images of different expressions. Among the 28,709 images in the training set, the first 24,000 will be used as the training dataset, and the remaining 4,709 will be used as the validation dataset, and they will be placed in the train and val folders respectively. To facilitate the input of the convolutional network program, a data-label comparison table will be established. Traverse all the files in the train and val folders, and write the names of the jpg format images and the corresponding labels into the comparison table as the input of the convolutional network program. To facilitate the input of the convolutional network program, a data-label comparison table will be established. Traverse all the files in the train and val folders, and write the names of the jpg format images and the corresponding labels into the comparison table as the input of the convolutional network program, and obtain the local spatio-temporal sequence features of negative expressions in the training dataset according to the input.

[0108] In this embodiment, a driver emotion recognition model is obtained through 3DCNN + ConvLSTM + SPP. By inputting the facial image obtained by the infrared camera into 3DCNN for convolution, the spatio-temporal sequence features of the driving facial features are obtained, and the obtained spatio-temporal sequence features are input into ConvLSTM. Since the RNN recurrent neural network in ConvLSTM is replaced by a convolutional neural network, and since the output of the fully connected layer is fixed, an SPP layer is added between the last convolutional layer and the fully connected layer to solve the problem that the image quality decreases due to the image size and affects emotion recognition. Therefore, the trained driver emotion recognition model provided in this embodiment has the advantages of not being limited to the input image size and having strong generalization ability through nonlinear processing, so it has higher recognition efficiency and emotion recognition accuracy.

[0109] In addition, an embodiment of the present invention also proposes a storage medium, on which a driver emotion recognition program is stored. When the driver emotion recognition program is executed by a processor, the steps of the driver emotion recognition method described above are implemented.

[0110] Referring to Figure 5 , Figure 5 which is the structural block diagram of the first embodiment of the driver emotion recognition device of the present invention.

[0111] As Figure 5 shown, the driver emotion recognition device proposed in the embodiment of the present invention includes:

[0112] An image acquisition module 10, configured to acquire a facial image of a driver and extract facial features of the facial image;

[0113] A feature extraction module 20, configured to extract a spatio-temporal feature sequence of the facial features through a driver emotion recognition model for the facial features;

[0114] A feature abstraction module 30, configured to abstract the spatio-temporal feature sequence to obtain an abstract feature;

[0115] An emotion recognition module 40, configured to obtain the emotion state of the driver according to the abstract feature.

[0116] In this embodiment, by acquiring a facial image of a driver, extracting facial features of the facial image, extracting a spatio-temporal feature sequence of the facial features through a driver emotion recognition model for the facial features, abstracting the spatio-temporal feature sequence to obtain an abstract feature, and obtaining the emotion state of the driver according to the abstract feature, through feature extraction and abstraction processing of the facial image of the driver by the driver emotion recognition model, the emotion of the driver is obtained. Compared with traditional emotion recognition, the present invention can improve the accuracy of emotion recognition and reduce the response time, and can improve the emotion recognition efficiency.

[0117] In one embodiment, the feature extraction module 20 is further configured to screen a face expression data set to obtain negative expressions; classify the negative emotions, add expression labels to the negative expressions according to the classification results to obtain a negative emotion set; divide the face expression data set to obtain a training data set and a validation data set, where the face expression data set includes the negative emotion set; input the training data set into an initial three-dimensional convolutional neural network to extract local spatio-temporal sequence features of negative expressions in the training data set; input the local spatio-temporal sequence features into a ConvLSTM for abstraction to obtain an abstract feature; perform feature transformation on the abstract feature to obtain a final feature; perform discrimination according to the final feature to obtain an emotion recognition result and obtain an intermediate emotion recognition model; verify the intermediate emotion recognition model according to the validation data set to obtain a recognition error, and when the recognition error is less than an error threshold, use the intermediate emotion recognition model as the emotion recognition model.

[0118] In one embodiment, the feature extraction module 20 is further configured to determine several subsets of negative emotions according to the classification result, and add emotion labels corresponding to the several subsets of negative emotions; screen the several subsets of emotions to obtain strongly-expressed emotion samples; obtain an extreme emotion expression set according to the strongly-expressed emotion samples, and add emotion labels to the extreme emotion expression set; and add the extreme emotion expression set to the negative emotion set.

[0119] In one embodiment, the feature extraction module 20 is further configured to input the abstract features into a spatial pyramid pooling layer for feature transformation to obtain final features, and the spatial pyramid pooling layer is connected to the last convolutional layer and the fully-connected layer in the ConvLSTM.

[0120] In one embodiment, the feature extraction module 20 is further configured to separate the expression information and the label information in the training dataset to generate an expression data file and a label data file; establish a comparison table according to the expression data file and the label data file, input the comparison table into an initial three-dimensional convolutional neural network, and extract the local spatio-temporal sequence features of negative expressions in the training dataset.

[0121] In one embodiment, the emotion recognition module 40 is further configured to determine the expression category of the extreme expression when the emotion state of the driver is an extreme expression; call a light warning and a voice warning for reminder according to the expression category; and perform emotion soothing on the driver according to the expression category.

[0122] In one embodiment, the emotion recognition module 40 is further configured to obtain the duration of the extreme emotion; when the duration of the extreme emotion exceeds a safety threshold, obtain the current driving trajectory, and judge the driving trajectory to determine the current driving behavior; when the current driving behavior is a violent driving behavior, take over the vehicle driving control right, and drive the vehicle to a safe location to stop until the driver's emotion is stable.

[0123] It should be understood that the above is only an example for illustration and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not make any restrictions thereon.

[0124] It should be understood that although the steps in the flowcharts in the embodiments of the present application are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0125] It should be noted that the workflow described above is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and no limitation is made here.

[0126] In addition, it should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0127] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a Read Only Memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0129] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present invention.

Claims

1. A driver emotion recognition method, characterized in that, The described driver emotion recognition method includes: Obtain the facial image of the driver and extract the facial features of the facial image; Extract the spatio-temporal feature sequence of the facial features through the driver emotion recognition model for the facial features; Abstract the spatio-temporal feature sequence to obtain an abstract feature; Obtain the emotion state of the driver according to the abstract feature; Before extracting the spatio-temporal feature sequence of the facial features through the driver emotion recognition model for the facial features, it further includes: Screen the human face expression dataset to obtain negative expressions; Classify the negative expressions, add expression labels to the negative expressions according to the classification results to obtain a negative emotion set; Divide the human face expression dataset to obtain a training dataset and a validation dataset, and the human face expression dataset includes the negative emotion set; Input the training dataset into an initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of the negative expressions in the training dataset; Input the local spatio-temporal sequence features into ConvLSTM for abstraction to obtain an abstract feature; Perform feature transformation on the abstract feature to obtain a final feature; Perform discrimination according to the final feature to obtain an emotion recognition result and obtain an intermediate emotion recognition model; Verify the intermediate emotion recognition model according to the validation dataset to obtain a recognition error. When the recognition error is less than the error threshold, use the intermediate emotion recognition model as the emotion recognition model; The classifying the negative expressions and adding expression labels to the negative expressions according to the classification results to obtain a negative emotion set includes: Determine several negative emotion subsets according to the classification results and add emotion labels to the several negative emotion subsets correspondingly; Screen the several emotion subsets to obtain strongly-expressed samples; Obtain an extreme expression set according to the strongly-expressed samples and add emotion labels to the extreme expression set; Add the extreme expression set to the negative emotion set; After obtaining the emotion state of the driver according to the abstract feature, it further includes: When the emotion state of the driver is an extreme expression, determine the expression category of the extreme expression; Call light warning and voice warning for reminder according to the expression category; Perform emotion soothing on the driver according to the expression category; After determining the expression category of the extreme expression when the emotion state of the driver is an extreme expression, it further includes: Obtain the duration of the extreme expression; When the duration of the extreme expression exceeds the safety threshold, obtain the current driving trajectory and judge the driving trajectory to determine the current driving behavior; When the current driving behavior is a violent driving behavior, take over the vehicle driving control right and drive the vehicle to a safe location to stop until the driver's emotion is stable.

2. The method according to claim 1, characterized in that, The performing feature transformation on the abstract feature to obtain a final feature includes: Input the abstract feature into a spatial pyramid pooling layer for feature transformation to obtain a final feature, and the spatial pyramid pooling layer is connected to the last convolutional layer and the fully-connected layer in the ConvLSTM.

3. The method according to claim 1, characterized in that, Inputting the training data set into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training data set includes: Separating the expression information and label information in the training data set to generate an expression data file and a label data file; Establishing a comparison table according to the expression data file and the label data file, and inputting the comparison table into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training data set.

4. A driver emotion recognition device, characterized in that, The driver emotion recognition device includes: An image acquisition module for acquiring the facial image of the driver and extracting the facial features of the facial image; A feature extraction module for extracting the spatio-temporal feature sequence of the facial features through the driver emotion recognition model; A feature abstraction module for abstracting the spatio-temporal feature sequence to obtain an abstract feature; An emotion recognition module for obtaining the emotion state of the driver according to the abstract feature; Before extracting the spatio-temporal feature sequence of the facial features through the driver emotion recognition model, it further includes: Screening the human face expression data set to obtain negative expressions; Classifying the negative expressions, adding expression labels to the negative expressions according to the classification results to obtain a negative emotion set; Dividing the human face expression data set to obtain a training data set and a validation data set, and the human face expression data set includes the negative emotion set; Inputting the training data set into the initial three-dimensional convolutional neural network to extract the local spatio-temporal sequence features of negative expressions in the training data set; Inputting the local spatio-temporal sequence features into ConvLSTM for abstraction to obtain an abstract feature; Performing feature transformation on the abstract feature to obtain a final feature; Making a judgment according to the final feature to obtain an emotion recognition result and obtaining an intermediate emotion recognition model; Validating the intermediate emotion recognition model according to the validation data set to obtain a recognition error. When the recognition error is less than the error threshold, using the intermediate emotion recognition model as the emotion recognition model; Classifying the negative expressions, adding expression labels to the negative expressions according to the classification results to obtain a negative emotion set, including: Determining several negative emotion subsets according to the classification results and adding emotion labels corresponding to the several negative emotion subsets; Screening the several emotion subsets to obtain strongly-expressed samples; Obtaining an extreme expression set according to the strongly-expressed samples and adding emotion labels to the extreme expression set; Adding the extreme expression set to the negative emotion set; After obtaining the emotion state of the driver according to the abstract feature, it further includes: When the emotion state of the driver is an extreme expression, determining the expression category of the extreme expression; Invoking light warning and voice warning for reminder according to the expression category; Performing emotion soothing on the driver according to the expression category; After determining the expression category of the extreme expression when the emotion state of the driver is an extreme expression, it further includes: Obtaining the duration of the extreme expression; When the duration of the extreme expression exceeds the safety threshold, obtain the current driving trajectory, and judge the driving trajectory to determine the current driving behavior; When the current driving behavior is a violent driving behavior, take over the vehicle driving control right, and drive the vehicle to a safe location to stop until the driver's emotion is stable.

5. A driver emotion recognition device, characterized in that The device includes: a memory, a processor, and a driver emotion recognition program stored on the memory and operable on the processor, and the driver emotion recognition program is configured to implement the steps of the driver emotion recognition method according to any one of claims 1 to 3.

6. A storage medium, characterized in that, A driver emotion recognition program is stored on the storage medium, and when the driver emotion recognition program is executed by the processor, the steps of the driver emotion recognition method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Subway driver vehicular driving behavior analysis method, vehicular terminal and system

    CN108216252A

  • Vehicle driver assistance system, vehicle, driver assistance method, computer program and computer-readable medium

    EP4015311A1